<?xml version="1.0" encoding="UTF-8"?>

<modsCollection xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns="http://www.loc.gov/mods/v3" xsi:schemaLocation="http://www.loc.gov/mods/v3 http://www.loc.gov/standards/mods/v3/mods-3-3.xsd">
<mods version="3.3">

<genre>article</genre>

<titleInfo><title>Hybrid decentralized optimization: Leveraging both first- and zeroth-order optimizers for faster convergence</title></titleInfo>


<note type="publicationStatus">published</note>


<note type="qualityControlled">yes</note>

<name type="personal">
  <namePart type="given">Shayan</namePart>
  <namePart type="family">Talaei</namePart>
  <role><roleTerm type="text">author</roleTerm> </role></name>
<name type="personal">
  <namePart type="given">Matin</namePart>
  <namePart type="family">Ansaripour</namePart>
  <role><roleTerm type="text">author</roleTerm> </role></name>
<name type="personal">
  <namePart type="given">Giorgi</namePart>
  <namePart type="family">Nadiradze</namePart>
  <role><roleTerm type="text">author</roleTerm> </role><identifier type="local">3279A00C-F248-11E8-B48F-1D18A9856A87</identifier><description xsi:type="identifierDefinition" type="orcid">0000-0001-5634-0731</description></name>
<name type="personal">
  <namePart type="given">Dan-Adrian</namePart>
  <namePart type="family">Alistarh</namePart>
  <role><roleTerm type="text">author</roleTerm> </role><identifier type="local">4A899BFC-F248-11E8-B48F-1D18A9856A87</identifier><description xsi:type="identifierDefinition" type="orcid">0000-0003-3650-940X</description></name>







<name type="corporate">
  <namePart></namePart>
  <identifier type="local">DaAl</identifier>
  <role>
    <roleTerm type="text">department</roleTerm>
  </role>
</name>





<name type="corporate">
  <namePart>Elastic Coordination for Scalable Machine Learning</namePart>
  <role><roleTerm type="text">project</roleTerm></role>
</name>



<abstract lang="eng">Distributed optimization is the standard way of speeding up machine learning training, and most of the research in the area focuses on distributed first-order, gradient-based methods. Yet, there are settings where some computationally-bounded nodes may not be able to implement first-order, gradient-based optimization, while they could still contribute to joint optimization tasks. In this paper, we initiate the study of hybrid decentralized optimization, studying settings where nodes with zeroth-order and first-order optimization capabilities co-exist in a distributed system, and attempt to jointly solve an optimization task over some data distribution. We essentially show that, under reasonable parameter settings, such a system can not only withstand noisier zeroth-order agents but can even benefit from integrating such agents into the optimization process, rather than ignoring their information. At the core of our approach is a new analysis of distributed optimization with noisy and possibly-biased gradient estimators, which may be of independent interest. Our results hold for both convex and non-convex objectives. Experimental results on standard optimization tasks confirm our analysis, showing that hybrid first-zeroth order optimization can be practical, even when training deep neural networks.</abstract>

<originInfo><publisher>Association for the Advancement of Artificial Intelligence</publisher><dateIssued encoding="w3cdtf">2025</dateIssued>
</originInfo>
<language><languageTerm authority="iso639-2b" type="code">eng</languageTerm>
</language>



<relatedItem type="host"><titleInfo><title>Proceedings of the 39th AAAI Conference on Artificial Intelligence</title></titleInfo>
  <identifier type="issn">2159-5399</identifier>
  <identifier type="eIssn">2374-3468</identifier>
  <identifier type="arXiv">2210.07703</identifier><identifier type="doi">10.1609/aaai.v39i19.34290</identifier>
<part><detail type="volume"><number>39</number></detail><detail type="issue"><number>19</number></detail><extent unit="pages">20778-20786</extent>
</part>
</relatedItem>


<relatedItem type="Supplementary material">
  <location>
  
     <url>https://github.com/ShayanTalaei/HDO</url>
  
  </location>
</relatedItem>

<extension>
<bibliographicCitation>
<apa>Talaei, S., Ansaripour, M., Nadiradze, G., &amp;#38; Alistarh, D.-A. (2025). Hybrid decentralized optimization: Leveraging both first- and zeroth-order optimizers for faster convergence. &lt;i&gt;Proceedings of the 39th AAAI Conference on Artificial Intelligence&lt;/i&gt;. Association for the Advancement of Artificial Intelligence. &lt;a href=&quot;https://doi.org/10.1609/aaai.v39i19.34290&quot;&gt;https://doi.org/10.1609/aaai.v39i19.34290&lt;/a&gt;</apa>
<ista>Talaei S, Ansaripour M, Nadiradze G, Alistarh D-A. 2025. Hybrid decentralized optimization: Leveraging both first- and zeroth-order optimizers for faster convergence. Proceedings of the 39th AAAI Conference on Artificial Intelligence. 39(19), 20778–20786.</ista>
<ama>Talaei S, Ansaripour M, Nadiradze G, Alistarh D-A. Hybrid decentralized optimization: Leveraging both first- and zeroth-order optimizers for faster convergence. &lt;i&gt;Proceedings of the 39th AAAI Conference on Artificial Intelligence&lt;/i&gt;. 2025;39(19):20778-20786. doi:&lt;a href=&quot;https://doi.org/10.1609/aaai.v39i19.34290&quot;&gt;10.1609/aaai.v39i19.34290&lt;/a&gt;</ama>
<short>S. Talaei, M. Ansaripour, G. Nadiradze, D.-A. Alistarh, Proceedings of the 39th AAAI Conference on Artificial Intelligence 39 (2025) 20778–20786.</short>
<mla>Talaei, Shayan, et al. “Hybrid Decentralized Optimization: Leveraging Both First- and Zeroth-Order Optimizers for Faster Convergence.” &lt;i&gt;Proceedings of the 39th AAAI Conference on Artificial Intelligence&lt;/i&gt;, vol. 39, no. 19, Association for the Advancement of Artificial Intelligence, 2025, pp. 20778–86, doi:&lt;a href=&quot;https://doi.org/10.1609/aaai.v39i19.34290&quot;&gt;10.1609/aaai.v39i19.34290&lt;/a&gt;.</mla>
<chicago>Talaei, Shayan, Matin Ansaripour, Giorgi Nadiradze, and Dan-Adrian Alistarh. “Hybrid Decentralized Optimization: Leveraging Both First- and Zeroth-Order Optimizers for Faster Convergence.” &lt;i&gt;Proceedings of the 39th AAAI Conference on Artificial Intelligence&lt;/i&gt;. Association for the Advancement of Artificial Intelligence, 2025. &lt;a href=&quot;https://doi.org/10.1609/aaai.v39i19.34290&quot;&gt;https://doi.org/10.1609/aaai.v39i19.34290&lt;/a&gt;.</chicago>
<ieee>S. Talaei, M. Ansaripour, G. Nadiradze, and D.-A. Alistarh, “Hybrid decentralized optimization: Leveraging both first- and zeroth-order optimizers for faster convergence,” &lt;i&gt;Proceedings of the 39th AAAI Conference on Artificial Intelligence&lt;/i&gt;, vol. 39, no. 19. Association for the Advancement of Artificial Intelligence, pp. 20778–20786, 2025.</ieee>
</bibliographicCitation>
</extension>
<recordInfo><recordIdentifier>19713</recordIdentifier><recordCreationDate encoding="w3cdtf">2025-05-19T14:15:35Z</recordCreationDate><recordChangeDate encoding="w3cdtf">2026-02-16T12:34:44Z</recordChangeDate>
</recordInfo>
</mods>
</modsCollection>
