<?xml version="1.0" encoding="UTF-8"?>
<OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/"
         xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"
         xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd">
<ListRecords>
<oai_dc:dc xmlns="http://www.openarchives.org/OAI/2.0/oai_dc/"
           xmlns:oai_dc="http://www.openarchives.org/OAI/2.0/oai_dc/"
           xmlns:dc="http://purl.org/dc/elements/1.1/"
           xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"
           xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/oai_dc/ http://www.openarchives.org/OAI/2.0/oai_dc.xsd">
   	<dc:title>Hybrid decentralized optimization: Leveraging both first- and zeroth-order optimizers for faster convergence</dc:title>
   	<dc:creator>Talaei, Shayan</dc:creator>
   	<dc:creator>Ansaripour, Matin</dc:creator>
   	<dc:creator>Nadiradze, Giorgi ; https://orcid.org/0000-0001-5634-0731</dc:creator>
   	<dc:creator>Alistarh, Dan-Adrian ; https://orcid.org/0000-0003-3650-940X</dc:creator>
   	<dc:description>Distributed optimization is the standard way of speeding up machine learning training, and most of the research in the area focuses on distributed first-order, gradient-based methods. Yet, there are settings where some computationally-bounded nodes may not be able to implement first-order, gradient-based optimization, while they could still contribute to joint optimization tasks. In this paper, we initiate the study of hybrid decentralized optimization, studying settings where nodes with zeroth-order and first-order optimization capabilities co-exist in a distributed system, and attempt to jointly solve an optimization task over some data distribution. We essentially show that, under reasonable parameter settings, such a system can not only withstand noisier zeroth-order agents but can even benefit from integrating such agents into the optimization process, rather than ignoring their information. At the core of our approach is a new analysis of distributed optimization with noisy and possibly-biased gradient estimators, which may be of independent interest. Our results hold for both convex and non-convex objectives. Experimental results on standard optimization tasks confirm our analysis, showing that hybrid first-zeroth order optimization can be practical, even when training deep neural networks.</dc:description>
   	<dc:publisher>Association for the Advancement of Artificial Intelligence</dc:publisher>
   	<dc:date>2025</dc:date>
   	<dc:type>info:eu-repo/semantics/article</dc:type>
   	<dc:type>doc-type:Article</dc:type>
   	<dc:type>Article</dc:type>
   	<dc:type>http://purl.org/coar/resource_type/c_2df8fbb1</dc:type>
   	<dc:identifier>https://research-explorer.ista.ac.at/record/19713</dc:identifier>
   	<dc:source>Talaei S, Ansaripour M, Nadiradze G, Alistarh D-A. Hybrid decentralized optimization: Leveraging both first- and zeroth-order optimizers for faster convergence. &lt;i&gt;Proceedings of the 39th AAAI Conference on Artificial Intelligence&lt;/i&gt;. 2025;39(19):20778-20786. doi:&lt;a href=&quot;https://doi.org/10.1609/aaai.v39i19.34290&quot;&gt;10.1609/aaai.v39i19.34290&lt;/a&gt;</dc:source>
   	<dc:language>eng</dc:language>
   	<dc:relation>info:eu-repo/semantics/altIdentifier/doi/10.1609/aaai.v39i19.34290</dc:relation>
   	<dc:relation>info:eu-repo/semantics/altIdentifier/issn/2159-5399</dc:relation>
   	<dc:relation>info:eu-repo/semantics/altIdentifier/e-issn/2374-3468</dc:relation>
   	<dc:relation>info:eu-repo/semantics/altIdentifier/arxiv/2210.07703</dc:relation>
   	<dc:rights>info:eu-repo/semantics/openAccess</dc:rights>
</oai_dc:dc>
</ListRecords>
</OAI-PMH>
