---
res:
  bibo_abstract:
  - As the size and complexity of models and datasets grow, so does the need for communication-efficient
    variants of stochastic gradient descent that can be deployed to perform parallel
    model training. One popular communication-compression method for data-parallel
    SGD is QSGD (Alistarh et al., 2017), which quantizes and encodes gradients to
    reduce communication costs. The baseline variant of QSGD provides strong theoretical
    guarantees, however, for practical purposes, the authors proposed a heuristic
    variant which we call QSGDinf, which demonstrated impressive empirical gains for
    distributed training of large neural networks. In this paper, we build on this
    work to propose a new gradient quantization scheme, and show that it has both
    stronger theoretical guarantees than QSGD, and matches and exceeds the empirical
    performance of the QSGDinf heuristic and of other compression methods.@eng
  bibo_authorlist:
  - foaf_Person:
      foaf_givenName: Ali
      foaf_name: Ramezani-Kebrya, Ali
      foaf_surname: Ramezani-Kebrya
  - foaf_Person:
      foaf_givenName: Fartash
      foaf_name: Faghri, Fartash
      foaf_surname: Faghri
  - foaf_Person:
      foaf_givenName: Ilya
      foaf_name: Markov, Ilya
      foaf_surname: Markov
  - foaf_Person:
      foaf_givenName: Vitalii
      foaf_name: Aksenov, Vitalii
      foaf_surname: Aksenov
      foaf_workInfoHomepage: http://www.librecat.org/personId=2980135A-F248-11E8-B48F-1D18A9856A87
  - foaf_Person:
      foaf_givenName: Dan-Adrian
      foaf_name: Alistarh, Dan-Adrian
      foaf_surname: Alistarh
      foaf_workInfoHomepage: http://www.librecat.org/personId=4A899BFC-F248-11E8-B48F-1D18A9856A87
    orcid: 0000-0003-3650-940X
  - foaf_Person:
      foaf_givenName: Daniel M.
      foaf_name: Roy, Daniel M.
      foaf_surname: Roy
  bibo_issue: '114'
  bibo_volume: 22
  dct_date: 2021^xs_gYear
  dct_isPartOf:
  - http://id.crossref.org/issn/1532-4435
  - http://id.crossref.org/issn/1533-7928
  dct_language: eng
  dct_publisher: Journal of Machine Learning Research@
  dct_title: 'NUQSGD: Provably communication-efficient data-parallel SGD via nonuniform
    quantization@'
...
