---
_id: '6558'
abstract:
- lang: eng
  text: This paper studies the problem of distributed stochastic optimization in an
    adversarial setting where, out of m machines which allegedly compute stochastic
    gradients every iteration, an α-fraction are Byzantine, and may behave adversarially.
    Our main result is a variant of stochastic gradient descent (SGD) which finds
    ε-approximate minimizers of convex functions in T=O~(1/ε²m+α²/ε²) iterations.
    In contrast, traditional mini-batch SGD needs T=O(1/ε²m) iterations, but cannot
    tolerate Byzantine failures. Further, we provide a lower bound showing that, up
    to logarithmic factors, our algorithm is information-theoretically optimal both
    in terms of sample complexity and time complexity.
article_processing_charge: No
arxiv: 1
author:
- first_name: Dan-Adrian
  full_name: Alistarh, Dan-Adrian
  id: 4A899BFC-F248-11E8-B48F-1D18A9856A87
  last_name: Alistarh
  orcid: 0000-0003-3650-940X
- first_name: Zeyuan
  full_name: Allen-Zhu, Zeyuan
  last_name: Allen-Zhu
- first_name: Jerry
  full_name: Li, Jerry
  last_name: Li
citation:
  ama: 'Alistarh D-A, Allen-Zhu Z, Li J. Byzantine stochastic gradient descent. In:
    <i>Advances in Neural Information Processing Systems</i>. Vol 2018. Neural Information
    Processing Systems Foundation; 2018:4613-4623.'
  apa: 'Alistarh, D.-A., Allen-Zhu, Z., &#38; Li, J. (2018). Byzantine stochastic
    gradient descent. In <i>Advances in Neural Information Processing Systems</i>
    (Vol. 2018, pp. 4613–4623). Montreal, Canada: Neural Information Processing Systems
    Foundation.'
  chicago: Alistarh, Dan-Adrian, Zeyuan Allen-Zhu, and Jerry Li. “Byzantine Stochastic
    Gradient Descent.” In <i>Advances in Neural Information Processing Systems</i>,
    2018:4613–23. Neural Information Processing Systems Foundation, 2018.
  ieee: D.-A. Alistarh, Z. Allen-Zhu, and J. Li, “Byzantine stochastic gradient descent,”
    in <i>Advances in Neural Information Processing Systems</i>, Montreal, Canada,
    2018, vol. 2018, pp. 4613–4623.
  ista: 'Alistarh D-A, Allen-Zhu Z, Li J. 2018. Byzantine stochastic gradient descent.
    Advances in Neural Information Processing Systems. NeurIPS: Conference on Neural
    Information Processing Systems vol. 2018, 4613–4623.'
  mla: Alistarh, Dan-Adrian, et al. “Byzantine Stochastic Gradient Descent.” <i>Advances
    in Neural Information Processing Systems</i>, vol. 2018, Neural Information Processing
    Systems Foundation, 2018, pp. 4613–23.
  short: D.-A. Alistarh, Z. Allen-Zhu, J. Li, in:, Advances in Neural Information
    Processing Systems, Neural Information Processing Systems Foundation, 2018, pp.
    4613–4623.
conference:
  end_date: 2018-12-08
  location: Montreal, Canada
  name: 'NeurIPS: Conference on Neural Information Processing Systems'
  start_date: 2018-12-02
date_created: 2019-06-13T08:22:37Z
date_published: 2018-12-01T00:00:00Z
date_updated: 2026-06-18T19:08:25Z
day: '01'
ddc:
- '000'
department:
- _id: DaAl
external_id:
  arxiv:
  - '1803.08917'
  isi:
  - '000461823304061'
intvolume: '      2018'
isi: 1
language:
- iso: eng
main_file_link:
- open_access: '1'
  url: https://arxiv.org/abs/1803.08917
month: '12'
oa: 1
oa_version: Published Version
page: 4613-4623
publication: Advances in Neural Information Processing Systems
publication_status: published
publisher: Neural Information Processing Systems Foundation
quality_controlled: '1'
scopus_import: '1'
status: public
title: Byzantine stochastic gradient descent
type: conference
user_id: 2DF688A6-F248-11E8-B48F-1D18A9856A87
volume: 2018
year: '2018'
...
---
_id: '7116'
abstract:
- lang: eng
  text: 'Training deep learning models has received tremendous research interest recently.
    In particular, there has been intensive research on reducing the communication
    cost of training when using multiple computational devices, through reducing the
    precision of the underlying data representation. Naturally, such methods induce
    system trade-offs—lowering communication precision could de-crease communication
    overheads and improve scalability; but, on the other hand, it can also reduce
    the accuracy of training. In this paper, we study this trade-off space, and ask:Can
    low-precision communication consistently improve the end-to-end performance of
    training modern neural networks, with no accuracy loss?From the performance point
    of view, the answer to this question may appear deceptively easy: compressing
    communication through low precision should help when the ratio between communication
    and computation is high. However, this answer is less straightforward when we
    try to generalize this principle across various neural network architectures (e.g.,
    AlexNet vs. ResNet),number of GPUs (e.g., 2 vs. 8 GPUs), machine configurations(e.g.,
    EC2 instances vs. NVIDIA DGX-1), communication primitives (e.g., MPI vs. NCCL),
    and even different GPU architectures(e.g., Kepler vs. Pascal). Currently, it is
    not clear how a realistic realization of all these factors maps to the speed up
    provided by low-precision communication. In this paper, we conduct an empirical
    study to answer this question and report the insights.'
article_processing_charge: No
author:
- first_name: Demjan
  full_name: Grubic, Demjan
  last_name: Grubic
- first_name: Leo
  full_name: Tam, Leo
  last_name: Tam
- first_name: Dan-Adrian
  full_name: Alistarh, Dan-Adrian
  id: 4A899BFC-F248-11E8-B48F-1D18A9856A87
  last_name: Alistarh
  orcid: 0000-0003-3650-940X
- first_name: Ce
  full_name: Zhang, Ce
  last_name: Zhang
citation:
  ama: 'Grubic D, Tam L, Alistarh D-A, Zhang C. Synchronous multi-GPU training for
    deep learning with low-precision communications: An empirical study. In: <i>Proceedings
    of the 21st International Conference on Extending Database Technology</i>. OpenProceedings;
    2018:145-156. doi:<a href="https://doi.org/10.5441/002/EDBT.2018.14">10.5441/002/EDBT.2018.14</a>'
  apa: 'Grubic, D., Tam, L., Alistarh, D.-A., &#38; Zhang, C. (2018). Synchronous
    multi-GPU training for deep learning with low-precision communications: An empirical
    study. In <i>Proceedings of the 21st International Conference on Extending Database
    Technology</i> (pp. 145–156). Vienna, Austria: OpenProceedings. <a href="https://doi.org/10.5441/002/EDBT.2018.14">https://doi.org/10.5441/002/EDBT.2018.14</a>'
  chicago: 'Grubic, Demjan, Leo Tam, Dan-Adrian Alistarh, and Ce Zhang. “Synchronous
    Multi-GPU Training for Deep Learning with Low-Precision Communications: An Empirical
    Study.” In <i>Proceedings of the 21st International Conference on Extending Database
    Technology</i>, 145–56. OpenProceedings, 2018. <a href="https://doi.org/10.5441/002/EDBT.2018.14">https://doi.org/10.5441/002/EDBT.2018.14</a>.'
  ieee: 'D. Grubic, L. Tam, D.-A. Alistarh, and C. Zhang, “Synchronous multi-GPU training
    for deep learning with low-precision communications: An empirical study,” in <i>Proceedings
    of the 21st International Conference on Extending Database Technology</i>, Vienna,
    Austria, 2018, pp. 145–156.'
  ista: 'Grubic D, Tam L, Alistarh D-A, Zhang C. 2018. Synchronous multi-GPU training
    for deep learning with low-precision communications: An empirical study. Proceedings
    of the 21st International Conference on Extending Database Technology. EDBT: Conference
    on Extending Database Technology, 145–156.'
  mla: 'Grubic, Demjan, et al. “Synchronous Multi-GPU Training for Deep Learning with
    Low-Precision Communications: An Empirical Study.” <i>Proceedings of the 21st
    International Conference on Extending Database Technology</i>, OpenProceedings,
    2018, pp. 145–56, doi:<a href="https://doi.org/10.5441/002/EDBT.2018.14">10.5441/002/EDBT.2018.14</a>.'
  short: D. Grubic, L. Tam, D.-A. Alistarh, C. Zhang, in:, Proceedings of the 21st
    International Conference on Extending Database Technology, OpenProceedings, 2018,
    pp. 145–156.
conference:
  end_date: 2018-03-29
  location: Vienna, Austria
  name: 'EDBT: Conference on Extending Database Technology'
  start_date: 2018-03-26
corr_author: '1'
date_created: 2019-11-26T14:19:11Z
date_published: 2018-03-26T00:00:00Z
date_updated: 2024-10-09T20:59:05Z
day: '26'
ddc:
- '000'
department:
- _id: DaAl
doi: 10.5441/002/EDBT.2018.14
file:
- access_level: open_access
  checksum: ec979b56abc71016d6e6adfdadbb4afe
  content_type: application/pdf
  creator: dernst
  date_created: 2019-11-26T14:23:04Z
  date_updated: 2020-07-14T12:47:49Z
  file_id: '7118'
  file_name: 2018_OpenProceedings_Grubic.pdf
  file_size: 1603204
  relation: main_file
file_date_updated: 2020-07-14T12:47:49Z
has_accepted_license: '1'
language:
- iso: eng
license: https://creativecommons.org/licenses/by-nc-nd/4.0/
month: '03'
oa: 1
oa_version: Published Version
page: 145-156
publication: Proceedings of the 21st International Conference on Extending Database
  Technology
publication_identifier:
  isbn:
  - '9783893180783'
  issn:
  - 2367-2005
publication_status: published
publisher: OpenProceedings
quality_controlled: '1'
scopus_import: 1
status: public
title: 'Synchronous multi-GPU training for deep learning with low-precision communications:
  An empirical study'
tmp:
  image: /images/cc_by_nc_nd.png
  legal_code_url: https://creativecommons.org/licenses/by-nc-nd/4.0/legalcode
  name: Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International
    (CC BY-NC-ND 4.0)
  short: CC BY-NC-ND (4.0)
type: conference
user_id: 2DF688A6-F248-11E8-B48F-1D18A9856A87
year: '2018'
...
---
_id: '7123'
abstract:
- lang: eng
  text: "Population protocols are a popular model of distributed computing, in which
    n agents with limited local state interact randomly, and cooperate to collectively
    compute global predicates. Inspired by recent developments in DNA programming,
    an extensive series of papers, across different communities, has examined the
    computability and complexity characteristics of this model. Majority, or consensus,
    is a central task in this model, in which agents need to collectively reach a
    decision as to which one of two states A or B had a higher initial count. Two
    metrics are important: the time that a protocol requires to stabilize to an output
    decision, and the state space size that each agent requires to do so. It is known
    that majority requires Ω(log log n) states per agent to allow for fast (poly-logarithmic
    time) stabilization, and that O(log2 n) states are sufficient. Thus, there is
    an exponential gap between the space upper and lower bounds for this problem.
    This paper addresses this question.\r\n\r\nOn the negative side, we provide a
    new lower bound of Ω(log n) states for any protocol which stabilizes in O(n1–c)
    expected time, for any constant c > 0. This result is conditional on monotonicity
    and output assumptions, satisfied by all known protocols. Technically, it represents
    a departure from previous lower bounds, in that it does not rely on the existence
    of dense configurations. Instead, we introduce a new generalized surgery technique
    to prove the existence of incorrect executions for any algorithm which would contradict
    the lower bound. Subsequently, our lower bound also applies to general initial
    configurations, including ones with a leader. On the positive side, we give a
    new algorithm for majority which uses O(log n) states, and stabilizes in O(log2
    n) expected time. Central to the algorithm is a new leaderless phase clock technique,
    which allows agents to synchronize in phases of Θ(n log n) consecutive interactions
    using O(log n) states per agent, exploiting a new connection between population
    protocols and power-of-two-choices load balancing mechanisms. We also employ our
    phase clock to build a leader election algorithm with a state space of size O(log
    n), which stabilizes in O(log2 n) expected time."
article_processing_charge: No
arxiv: 1
author:
- first_name: Dan-Adrian
  full_name: Alistarh, Dan-Adrian
  id: 4A899BFC-F248-11E8-B48F-1D18A9856A87
  last_name: Alistarh
  orcid: 0000-0003-3650-940X
- first_name: James
  full_name: Aspnes, James
  last_name: Aspnes
- first_name: Rati
  full_name: Gelashvili, Rati
  last_name: Gelashvili
citation:
  ama: 'Alistarh D-A, Aspnes J, Gelashvili R. Space-optimal majority in population
    protocols. In: <i>Proceedings of the 29th Annual ACM-SIAM Symposium on Discrete
    Algorithms</i>. ACM; 2018:2221-2239. doi:<a href="https://doi.org/10.1137/1.9781611975031.144">10.1137/1.9781611975031.144</a>'
  apa: 'Alistarh, D.-A., Aspnes, J., &#38; Gelashvili, R. (2018). Space-optimal majority
    in population protocols. In <i>Proceedings of the 29th Annual ACM-SIAM Symposium
    on Discrete Algorithms</i> (pp. 2221–2239). New Orleans, LA, United States: ACM.
    <a href="https://doi.org/10.1137/1.9781611975031.144">https://doi.org/10.1137/1.9781611975031.144</a>'
  chicago: Alistarh, Dan-Adrian, James Aspnes, and Rati Gelashvili. “Space-Optimal
    Majority in Population Protocols.” In <i>Proceedings of the 29th Annual ACM-SIAM
    Symposium on Discrete Algorithms</i>, 2221–39. ACM, 2018. <a href="https://doi.org/10.1137/1.9781611975031.144">https://doi.org/10.1137/1.9781611975031.144</a>.
  ieee: D.-A. Alistarh, J. Aspnes, and R. Gelashvili, “Space-optimal majority in population
    protocols,” in <i>Proceedings of the 29th Annual ACM-SIAM Symposium on Discrete
    Algorithms</i>, New Orleans, LA, United States, 2018, pp. 2221–2239.
  ista: 'Alistarh D-A, Aspnes J, Gelashvili R. 2018. Space-optimal majority in population
    protocols. Proceedings of the 29th Annual ACM-SIAM Symposium on Discrete Algorithms.
    SODA: Symposium on Discrete Algorithms, 2221–2239.'
  mla: Alistarh, Dan-Adrian, et al. “Space-Optimal Majority in Population Protocols.”
    <i>Proceedings of the 29th Annual ACM-SIAM Symposium on Discrete Algorithms</i>,
    ACM, 2018, pp. 2221–39, doi:<a href="https://doi.org/10.1137/1.9781611975031.144">10.1137/1.9781611975031.144</a>.
  short: D.-A. Alistarh, J. Aspnes, R. Gelashvili, in:, Proceedings of the 29th Annual
    ACM-SIAM Symposium on Discrete Algorithms, ACM, 2018, pp. 2221–2239.
conference:
  end_date: 2018-01-10
  location: New Orleans, LA, United States
  name: 'SODA: Symposium on Discrete Algorithms'
  start_date: 2018-01-07
date_created: 2019-11-26T15:10:55Z
date_published: 2018-01-30T00:00:00Z
date_updated: 2024-10-21T06:02:41Z
day: '30'
department:
- _id: DaAl
doi: 10.1137/1.9781611975031.144
external_id:
  arxiv:
  - '1704.04947'
  isi:
  - '000483921200145'
isi: 1
language:
- iso: eng
main_file_link:
- open_access: '1'
  url: https://arxiv.org/abs/1704.04947
month: '01'
oa: 1
oa_version: Preprint
page: 2221-2239
publication: Proceedings of the 29th Annual ACM-SIAM Symposium on Discrete Algorithms
publication_identifier:
  isbn:
  - '9781611975031'
publication_status: published
publisher: ACM
quality_controlled: '1'
scopus_import: '1'
status: public
title: Space-optimal majority in population protocols
type: conference
user_id: c635000d-4b10-11ee-a964-aac5a93f6ac1
year: '2018'
...
---
_id: '397'
abstract:
- lang: eng
  text: 'Concurrent sets with range query operations are highly desirable in applications
    such as in-memory databases. However, few set implementations offer range queries.
    Known techniques for augmenting data structures with range queries (or operations
    that can be used to build range queries) have numerous problems that limit their
    usefulness. For example, they impose high overhead or rely heavily on garbage
    collection. In this work, we show how to augment data structures with highly efficient
    range queries, without relying on garbage collection. We identify a property of
    epoch-based memory reclamation algorithms that makes them ideal for implementing
    range queries, and produce three algorithms, which use locks, transactional memory
    and lock-free techniques, respectively. Our algorithms are applicable to more
    data structures than previous work, and are shown to be highly efficient on a
    large scale Intel system. '
alternative_title:
- PPoPP
article_processing_charge: No
author:
- first_name: Maya
  full_name: Arbel Raviv, Maya
  last_name: Arbel Raviv
- first_name: Trevor A
  full_name: Brown, Trevor A
  id: 3569F0A0-F248-11E8-B48F-1D18A9856A87
  last_name: Brown
citation:
  ama: 'Arbel Raviv M, Brown TA. Harnessing epoch-based reclamation for efficient
    range queries. In: Vol 53. ACM; 2018:14-27. doi:<a href="https://doi.org/10.1145/3178487.3178489">10.1145/3178487.3178489</a>'
  apa: 'Arbel Raviv, M., &#38; Brown, T. A. (2018). Harnessing epoch-based reclamation
    for efficient range queries (Vol. 53, pp. 14–27). Presented at the PPoPP: Principles
    and Practice of Parallel Programming, Vienna, Austria: ACM. <a href="https://doi.org/10.1145/3178487.3178489">https://doi.org/10.1145/3178487.3178489</a>'
  chicago: Arbel Raviv, Maya, and Trevor A Brown. “Harnessing Epoch-Based Reclamation
    for Efficient Range Queries,” 53:14–27. ACM, 2018. <a href="https://doi.org/10.1145/3178487.3178489">https://doi.org/10.1145/3178487.3178489</a>.
  ieee: 'M. Arbel Raviv and T. A. Brown, “Harnessing epoch-based reclamation for efficient
    range queries,” presented at the PPoPP: Principles and Practice of Parallel Programming,
    Vienna, Austria, 2018, vol. 53, no. 1, pp. 14–27.'
  ista: 'Arbel Raviv M, Brown TA. 2018. Harnessing epoch-based reclamation for efficient
    range queries. PPoPP: Principles and Practice of Parallel Programming, PPoPP,
    vol. 53, 14–27.'
  mla: Arbel Raviv, Maya, and Trevor A. Brown. <i>Harnessing Epoch-Based Reclamation
    for Efficient Range Queries</i>. Vol. 53, no. 1, ACM, 2018, pp. 14–27, doi:<a
    href="https://doi.org/10.1145/3178487.3178489">10.1145/3178487.3178489</a>.
  short: M. Arbel Raviv, T.A. Brown, in:, ACM, 2018, pp. 14–27.
conference:
  end_date: 2018-02-28
  location: Vienna, Austria
  name: 'PPoPP: Principles and Practice of Parallel Programming'
  start_date: 2018-02-24
date_created: 2018-12-11T11:46:14Z
date_published: 2018-02-10T00:00:00Z
date_updated: 2023-09-11T14:10:25Z
day: '10'
department:
- _id: DaAl
doi: 10.1145/3178487.3178489
external_id:
  isi:
  - '000446161100002'
intvolume: '        53'
isi: 1
issue: '1'
language:
- iso: eng
month: '02'
oa_version: None
page: 14 - 27
publication_identifier:
  isbn:
  - 978-1-4503-4982-6
publication_status: published
publisher: ACM
publist_id: '7430'
quality_controlled: '1'
scopus_import: '1'
status: public
title: Harnessing epoch-based reclamation for efficient range queries
type: conference
user_id: c635000d-4b10-11ee-a964-aac5a93f6ac1
volume: 53
year: '2018'
...
---
_id: '43'
abstract:
- lang: eng
  text: 'The initial amount of pathogens required to start an infection within a susceptible
    host is called the infective dose and is known to vary to a large extent between
    different pathogen species. We investigate the hypothesis that the differences
    in infective doses are explained by the mode of action in the underlying mechanism
    of pathogenesis: Pathogens with locally acting mechanisms tend to have smaller
    infective doses than pathogens with distantly acting mechanisms. While empirical
    evidence tends to support the hypothesis, a formal theoretical explanation has
    been lacking. We give simple analytical models to gain insight into this phenomenon
    and also investigate a stochastic, spatially explicit, mechanistic within-host
    model for toxin-dependent bacterial infections. The model shows that pathogens
    secreting locally acting toxins have smaller infective doses than pathogens secreting
    diffusive toxins, as hypothesized. While local pathogenetic mechanisms require
    smaller infective doses, pathogens with distantly acting toxins tend to spread
    faster and may cause more damage to the host. The proposed model can serve as
    a basis for the spatially explicit analysis of various virulence factors also
    in the context of other problems in infection dynamics.'
acknowledgement: J.R. and J.V.A. were also supported by the Academy of Finland Grants
  1273253 and 267541.
article_processing_charge: No
author:
- first_name: Joel
  full_name: Rybicki, Joel
  id: 334EFD2E-F248-11E8-B48F-1D18A9856A87
  last_name: Rybicki
  orcid: 0000-0002-6432-6646
- first_name: Eva
  full_name: Kisdi, Eva
  last_name: Kisdi
- first_name: Jani
  full_name: Anttila, Jani
  last_name: Anttila
citation:
  ama: Rybicki J, Kisdi E, Anttila J. Model of bacterial toxin-dependent pathogenesis
    explains infective dose. <i>Proceedings of the National Academy of Sciences of
    the United States of America</i>. 2018;115(42):10690-10695. doi:<a href="https://doi.org/10.1073/pnas.1721061115">10.1073/pnas.1721061115</a>
  apa: Rybicki, J., Kisdi, E., &#38; Anttila, J. (2018). Model of bacterial toxin-dependent
    pathogenesis explains infective dose. <i>Proceedings of the National Academy of
    Sciences of the United States of America</i>. National Academy of Sciences. <a
    href="https://doi.org/10.1073/pnas.1721061115">https://doi.org/10.1073/pnas.1721061115</a>
  chicago: Rybicki, Joel, Eva Kisdi, and Jani Anttila. “Model of Bacterial Toxin-Dependent
    Pathogenesis Explains Infective Dose.” <i>Proceedings of the National Academy
    of Sciences of the United States of America</i>. National Academy of Sciences,
    2018. <a href="https://doi.org/10.1073/pnas.1721061115">https://doi.org/10.1073/pnas.1721061115</a>.
  ieee: J. Rybicki, E. Kisdi, and J. Anttila, “Model of bacterial toxin-dependent
    pathogenesis explains infective dose,” <i>Proceedings of the National Academy
    of Sciences of the United States of America</i>, vol. 115, no. 42. National Academy
    of Sciences, pp. 10690–10695, 2018.
  ista: Rybicki J, Kisdi E, Anttila J. 2018. Model of bacterial toxin-dependent pathogenesis
    explains infective dose. Proceedings of the National Academy of Sciences of the
    United States of America. 115(42), 10690–10695.
  mla: Rybicki, Joel, et al. “Model of Bacterial Toxin-Dependent Pathogenesis Explains
    Infective Dose.” <i>Proceedings of the National Academy of Sciences of the United
    States of America</i>, vol. 115, no. 42, National Academy of Sciences, 2018, pp.
    10690–95, doi:<a href="https://doi.org/10.1073/pnas.1721061115">10.1073/pnas.1721061115</a>.
  short: J. Rybicki, E. Kisdi, J. Anttila, Proceedings of the National Academy of
    Sciences of the United States of America 115 (2018) 10690–10695.
date_created: 2018-12-11T11:44:19Z
date_published: 2018-10-02T00:00:00Z
date_updated: 2025-06-03T11:16:28Z
day: '02'
ddc:
- '570'
- '577'
department:
- _id: DaAl
doi: 10.1073/pnas.1721061115
ec_funded: 1
external_id:
  isi:
  - '000447491300057'
file:
- access_level: open_access
  checksum: df7ac544a587c06b75692653b9fabd18
  content_type: application/pdf
  creator: dernst
  date_created: 2019-04-09T08:02:50Z
  date_updated: 2020-07-14T12:46:26Z
  file_id: '6258'
  file_name: 2018_PNAS_Rybicki.pdf
  file_size: 4070777
  relation: main_file
file_date_updated: 2020-07-14T12:46:26Z
has_accepted_license: '1'
intvolume: '       115'
isi: 1
issue: '42'
language:
- iso: eng
month: '10'
oa: 1
oa_version: Submitted Version
page: 10690 - 10695
project:
- _id: 260C2330-B435-11E9-9278-68D0E5697425
  call_identifier: H2020
  grant_number: '754411'
  name: ISTplus - Postdoctoral Fellowships
publication: Proceedings of the National Academy of Sciences of the United States
  of America
publication_identifier:
  eissn:
  - 1091-6490
  issn:
  - 0027-8424
publication_status: published
publisher: National Academy of Sciences
publist_id: '8011'
pubrep_id: '1063'
quality_controlled: '1'
scopus_import: '1'
status: public
title: Model of bacterial toxin-dependent pathogenesis explains infective dose
type: journal_article
user_id: 2DF688A6-F248-11E8-B48F-1D18A9856A87
volume: 115
year: '2018'
...
---
_id: '6589'
abstract:
- lang: eng
  text: Distributed training of massive machine learning models, in particular deep
    neural networks, via Stochastic Gradient Descent (SGD) is becoming commonplace.
    Several families of communication-reduction methods, such as quantization, large-batch
    methods, and gradient sparsification, have been proposed. To date, gradient sparsification
    methods--where each node sorts gradients by magnitude, and only communicates a
    subset of the components, accumulating the rest locally--are known to yield some
    of the largest practical gains. Such methods can reduce the amount of communication
    per step by up to \emph{three orders of magnitude}, while preserving model accuracy.
    Yet, this family of methods currently has no theoretical justification. This is
    the question we address in this paper. We prove that, under analytic assumptions,
    sparsifying gradients by magnitude with local error correction provides convergence
    guarantees, for both convex and non-convex smooth objectives, for data-parallel
    SGD. The main insight is that sparsification methods implicitly maintain bounds
    on the maximum impact of stale updates, thanks to selection by magnitude. Our
    analysis and empirical validation also reveal that these methods do require analytical
    conditions to converge well, justifying existing heuristics.
alternative_title:
- Advances in Neural Information Processing Systems
article_processing_charge: No
arxiv: 1
author:
- first_name: Dan-Adrian
  full_name: Alistarh, Dan-Adrian
  id: 4A899BFC-F248-11E8-B48F-1D18A9856A87
  last_name: Alistarh
  orcid: 0000-0003-3650-940X
- first_name: Torsten
  full_name: Hoefler, Torsten
  last_name: Hoefler
- first_name: Mikael
  full_name: Johansson, Mikael
  last_name: Johansson
- first_name: Nikola H
  full_name: Konstantinov, Nikola H
  id: 4B9D76E4-F248-11E8-B48F-1D18A9856A87
  last_name: Konstantinov
  orcid: 0009-0009-5204-7621
- first_name: Sarit
  full_name: Khirirat, Sarit
  last_name: Khirirat
- first_name: Cedric
  full_name: Renggli, Cedric
  last_name: Renggli
citation:
  ama: 'Alistarh D-A, Hoefler T, Johansson M, Konstantinov NH, Khirirat S, Renggli
    C. The convergence of sparsified gradient methods. In: <i>32nd Conference on Neural
    Information Processing Systems</i>. Neural Information Processing Systems Foundation;
    2018:5973-5983.'
  apa: 'Alistarh, D.-A., Hoefler, T., Johansson, M., Konstantinov, N. H., Khirirat,
    S., &#38; Renggli, C. (2018). The convergence of sparsified gradient methods.
    In <i>32nd Conference on Neural Information Processing Systems</i> (pp. 5973–5983).
    Montreal, Canada: Neural Information Processing Systems Foundation.'
  chicago: Alistarh, Dan-Adrian, Torsten Hoefler, Mikael Johansson, Nikola H Konstantinov,
    Sarit Khirirat, and Cedric Renggli. “The Convergence of Sparsified Gradient Methods.”
    In <i>32nd Conference on Neural Information Processing Systems</i>, 5973–83. Neural
    Information Processing Systems Foundation, 2018.
  ieee: D.-A. Alistarh, T. Hoefler, M. Johansson, N. H. Konstantinov, S. Khirirat,
    and C. Renggli, “The convergence of sparsified gradient methods,” in <i>32nd Conference
    on Neural Information Processing Systems</i>, Montreal, Canada, 2018, pp. 5973–5983.
  ista: 'Alistarh D-A, Hoefler T, Johansson M, Konstantinov NH, Khirirat S, Renggli
    C. 2018. The convergence of sparsified gradient methods. 32nd Conference on Neural
    Information Processing Systems. NeurIPS: Conference on Neural Information Processing
    Systems, Advances in Neural Information Processing Systems, , 5973–5983.'
  mla: Alistarh, Dan-Adrian, et al. “The Convergence of Sparsified Gradient Methods.”
    <i>32nd Conference on Neural Information Processing Systems</i>, Neural Information
    Processing Systems Foundation, 2018, pp. 5973–83.
  short: D.-A. Alistarh, T. Hoefler, M. Johansson, N.H. Konstantinov, S. Khirirat,
    C. Renggli, in:, 32nd Conference on Neural Information Processing Systems, Neural
    Information Processing Systems Foundation, 2018, pp. 5973–5983.
conference:
  end_date: 2018-12-08
  location: Montreal, Canada
  name: 'NeurIPS: Conference on Neural Information Processing Systems'
  start_date: 2018-12-02
corr_author: '1'
das_tickbox: '1'
date_created: 2019-06-27T09:32:55Z
date_published: 2018-12-01T00:00:00Z
date_updated: 2026-07-08T05:48:21Z
day: '01'
department:
- _id: DaAl
- _id: ChLa
ec_funded: 1
external_id:
  arxiv:
  - '1809.10505'
  isi:
  - '000461852000047'
isi: 1
language:
- iso: eng
main_file_link:
- open_access: '1'
  url: https://arxiv.org/abs/1809.10505
month: '12'
oa: 1
oa_version: Preprint
page: 5973-5983
project:
- _id: 2564DBCA-B435-11E9-9278-68D0E5697425
  call_identifier: H2020
  grant_number: '665385'
  name: International IST Doctoral Program
publication: 32nd Conference on Neural Information Processing Systems
publication_identifier:
  issn:
  - 1049-5258
publication_status: published
publisher: Neural Information Processing Systems Foundation
quality_controlled: '1'
scopus_import: '1'
status: public
title: The convergence of sparsified gradient methods
type: conference
user_id: 2DF688A6-F248-11E8-B48F-1D18A9856A87
year: '2018'
...
---
_id: '487'
abstract:
- lang: eng
  text: In this paper we study network architecture for unlicensed cellular networking
    for outdoor coverage in TV white spaces. The main technology proposed for TV white
    spaces is 802.11af, a Wi-Fi variant adapted for TV frequencies. However, 802.11af
    is originally designed for improved indoor propagation. We show that long links,
    typical for outdoor use, exacerbate known Wi-Fi issues, such as hidden and exposed
    terminal, and significantly reduce its efficiency. Instead, we propose CellFi,
    an alternative architecture based on LTE. LTE is designed for long-range coverage
    and throughput efficiency, but it is also designed to operate in tightly controlled
    and centrally managed networks. CellFi overcomes these problems by designing an
    LTE-compatible spectrum database component, mandatory for TV white space networking,
    and introducing an interference management component for distributed coordination.
    CellFi interference management is compatible with existing LTE mechanisms, requires
    no explicit communication between base stations, and is more efficient than CSMA
    for long links. We evaluate our design through extensive real world evaluation
    on of-the-shelf LTE equipment and simulations. We show that, compared to 802.11af,
    it increases coverage by 40% and reduces median flow completion times by 2.3x.
article_processing_charge: No
author:
- first_name: Ghufran
  full_name: Baig, Ghufran
  last_name: Baig
- first_name: Bozidar
  full_name: Radunovic, Bozidar
  last_name: Radunovic
- first_name: Dan-Adrian
  full_name: Alistarh, Dan-Adrian
  id: 4A899BFC-F248-11E8-B48F-1D18A9856A87
  last_name: Alistarh
  orcid: 0000-0003-3650-940X
- first_name: Matthew
  full_name: Balkwill, Matthew
  last_name: Balkwill
- first_name: Thomas
  full_name: Karagiannis, Thomas
  last_name: Karagiannis
- first_name: Lili
  full_name: Qiu, Lili
  last_name: Qiu
citation:
  ama: 'Baig G, Radunovic B, Alistarh D-A, Balkwill M, Karagiannis T, Qiu L. Towards
    unlicensed cellular networks in TV white spaces. In: <i>Proceedings of the 2017
    13th International Conference on Emerging Networking EXperiments and Technologies</i>.
    ACM; 2017:2-14. doi:<a href="https://doi.org/10.1145/3143361.3143367">10.1145/3143361.3143367</a>'
  apa: 'Baig, G., Radunovic, B., Alistarh, D.-A., Balkwill, M., Karagiannis, T., &#38;
    Qiu, L. (2017). Towards unlicensed cellular networks in TV white spaces. In <i>Proceedings
    of the 2017 13th International Conference on emerging Networking EXperiments and
    Technologies</i> (pp. 2–14). Incheon, South Korea: ACM. <a href="https://doi.org/10.1145/3143361.3143367">https://doi.org/10.1145/3143361.3143367</a>'
  chicago: Baig, Ghufran, Bozidar Radunovic, Dan-Adrian Alistarh, Matthew Balkwill,
    Thomas Karagiannis, and Lili Qiu. “Towards Unlicensed Cellular Networks in TV
    White Spaces.” In <i>Proceedings of the 2017 13th International Conference on
    Emerging Networking EXperiments and Technologies</i>, 2–14. ACM, 2017. <a href="https://doi.org/10.1145/3143361.3143367">https://doi.org/10.1145/3143361.3143367</a>.
  ieee: G. Baig, B. Radunovic, D.-A. Alistarh, M. Balkwill, T. Karagiannis, and L.
    Qiu, “Towards unlicensed cellular networks in TV white spaces,” in <i>Proceedings
    of the 2017 13th International Conference on emerging Networking EXperiments and
    Technologies</i>, Incheon, South Korea, 2017, pp. 2–14.
  ista: 'Baig G, Radunovic B, Alistarh D-A, Balkwill M, Karagiannis T, Qiu L. 2017.
    Towards unlicensed cellular networks in TV white spaces. Proceedings of the 2017
    13th International Conference on emerging Networking EXperiments and Technologies.
    CoNEXT: Conference on emerging Networking EXperiments and Technologies, 2–14.'
  mla: Baig, Ghufran, et al. “Towards Unlicensed Cellular Networks in TV White Spaces.”
    <i>Proceedings of the 2017 13th International Conference on Emerging Networking
    EXperiments and Technologies</i>, ACM, 2017, pp. 2–14, doi:<a href="https://doi.org/10.1145/3143361.3143367">10.1145/3143361.3143367</a>.
  short: G. Baig, B. Radunovic, D.-A. Alistarh, M. Balkwill, T. Karagiannis, L. Qiu,
    in:, Proceedings of the 2017 13th International Conference on Emerging Networking
    EXperiments and Technologies, ACM, 2017, pp. 2–14.
conference:
  end_date: 2017-12-15
  location: Incheon, South Korea
  name: 'CoNEXT: Conference on emerging Networking EXperiments and Technologies'
  start_date: 2017-12-12
corr_author: '1'
date_created: 2018-12-11T11:46:45Z
date_published: 2017-11-28T00:00:00Z
date_updated: 2025-09-18T09:50:43Z
day: '28'
department:
- _id: DaAl
doi: 10.1145/3143361.3143367
external_id:
  isi:
  - '000526087500002'
isi: 1
language:
- iso: eng
month: '11'
oa_version: None
page: 2 - 14
publication: Proceedings of the 2017 13th International Conference on emerging Networking
  EXperiments and Technologies
publication_identifier:
  isbn:
  - 978-145035422-6
publication_status: published
publisher: ACM
publist_id: '7333'
quality_controlled: '1'
scopus_import: '1'
status: public
title: Towards unlicensed cellular networks in TV white spaces
type: conference
user_id: 317138e5-6ab7-11ef-aa6d-ffef3953e345
year: '2017'
...
---
_id: '791'
abstract:
- lang: eng
  text: 'Consider the following random process: we are given n queues, into which
    elements of increasing labels are inserted uniformly at random. To remove an element,
    we pick two queues at random, and remove the element of lower label (higher priority)
    among the two. The cost of a removal is the rank of the label removed, among labels
    still present in any of the queues, that is, the distance from the optimal choice
    at each step. Variants of this strategy are prevalent in state-of-the-art concurrent
    priority queue implementations. Nonetheless, it is not known whether such implementations
    provide any rank guarantees, even in a sequential model. We answer this question,
    showing that this strategy provides surprisingly strong guarantees: Although the
    single-choice process, where we always insert and remove from a single randomly
    chosen queue, has degrading cost, going to infinity as we increase the number
    of steps, in the two choice process, the expected rank of a removed element is
    O(n) while the expected worst-case cost is O(n log n). These bounds are tight,
    and hold irrespective of the number of steps for which we run the process. The
    argument is based on a new technical connection between &quot;heavily loaded&quot;
    balls-into-bins processes and priority scheduling. Our analytic results inspire
    a new concurrent priority queue implementation, which improves upon the state
    of the art in terms of practical performance.'
article_processing_charge: No
arxiv: 1
author:
- first_name: Dan-Adrian
  full_name: Alistarh, Dan-Adrian
  id: 4A899BFC-F248-11E8-B48F-1D18A9856A87
  last_name: Alistarh
  orcid: 0000-0003-3650-940X
- first_name: Justin
  full_name: Kopinsky, Justin
  last_name: Kopinsky
- first_name: Jerry
  full_name: Li, Jerry
  last_name: Li
- first_name: Giorgi
  full_name: Nadiradze, Giorgi
  id: 3279A00C-F248-11E8-B48F-1D18A9856A87
  last_name: Nadiradze
  orcid: 0000-0001-5634-0731
citation:
  ama: 'Alistarh D-A, Kopinsky J, Li J, Nadiradze G. The power of choice in priority
    scheduling. In: <i>Proceedings of the ACM Symposium on Principles of Distributed
    Computing</i>. Vol Part F129314. ACM; 2017:283-292. doi:<a href="https://doi.org/10.1145/3087801.3087810">10.1145/3087801.3087810</a>'
  apa: 'Alistarh, D.-A., Kopinsky, J., Li, J., &#38; Nadiradze, G. (2017). The power
    of choice in priority scheduling. In <i>Proceedings of the ACM Symposium on Principles
    of Distributed Computing</i> (Vol. Part F129314, pp. 283–292). Washington, WA,
    USA: ACM. <a href="https://doi.org/10.1145/3087801.3087810">https://doi.org/10.1145/3087801.3087810</a>'
  chicago: Alistarh, Dan-Adrian, Justin Kopinsky, Jerry Li, and Giorgi Nadiradze.
    “The Power of Choice in Priority Scheduling.” In <i>Proceedings of the ACM Symposium
    on Principles of Distributed Computing</i>, Part F129314:283–92. ACM, 2017. <a
    href="https://doi.org/10.1145/3087801.3087810">https://doi.org/10.1145/3087801.3087810</a>.
  ieee: D.-A. Alistarh, J. Kopinsky, J. Li, and G. Nadiradze, “The power of choice
    in priority scheduling,” in <i>Proceedings of the ACM Symposium on Principles
    of Distributed Computing</i>, Washington, WA, USA, 2017, vol. Part F129314, pp.
    283–292.
  ista: 'Alistarh D-A, Kopinsky J, Li J, Nadiradze G. 2017. The power of choice in
    priority scheduling. Proceedings of the ACM Symposium on Principles of Distributed
    Computing. PODC: Principles of Distributed Computing vol. Part F129314, 283–292.'
  mla: Alistarh, Dan-Adrian, et al. “The Power of Choice in Priority Scheduling.”
    <i>Proceedings of the ACM Symposium on Principles of Distributed Computing</i>,
    vol. Part F129314, ACM, 2017, pp. 283–92, doi:<a href="https://doi.org/10.1145/3087801.3087810">10.1145/3087801.3087810</a>.
  short: D.-A. Alistarh, J. Kopinsky, J. Li, G. Nadiradze, in:, Proceedings of the
    ACM Symposium on Principles of Distributed Computing, ACM, 2017, pp. 283–292.
conference:
  end_date: 2017-07-27
  location: Washington, WA, USA
  name: 'PODC: Principles of Distributed Computing'
  start_date: 2017-07-25
date_created: 2018-12-11T11:48:31Z
date_published: 2017-07-26T00:00:00Z
date_updated: 2025-06-04T09:44:47Z
day: '26'
department:
- _id: DaAl
doi: 10.1145/3087801.3087810
external_id:
  arxiv:
  - '1706.04178'
  isi:
  - '000462995000035'
isi: 1
language:
- iso: eng
main_file_link:
- open_access: '1'
  url: https://arxiv.org/abs/1706.04178
month: '07'
oa: 1
oa_version: Submitted Version
page: 283 - 292
publication: Proceedings of the ACM Symposium on Principles of Distributed Computing
publication_identifier:
  isbn:
  - 978-145034992-5
publication_status: published
publisher: ACM
publist_id: '6864'
quality_controlled: '1'
scopus_import: '1'
status: public
title: The power of choice in priority scheduling
type: conference
user_id: 2DF688A6-F248-11E8-B48F-1D18A9856A87
volume: Part F129314
year: '2017'
...
---
_id: '431'
abstract:
- lang: eng
  text: 'Parallel implementations of stochastic gradient descent (SGD) have received
    significant research attention, thanks to its excellent scalability properties.
    A fundamental barrier when parallelizing SGD is the high bandwidth cost of communicating
    gradient updates between nodes; consequently, several lossy compresion heuristics
    have been proposed, by which nodes only communicate quantized gradients. Although
    effective in practice, these heuristics do not always converge. In this paper,
    we propose Quantized SGD (QSGD), a family of compression schemes with convergence
    guarantees and good practical performance. QSGD allows the user to smoothly trade
    off communication bandwidth and convergence time: nodes can adjust the number
    of bits sent per iteration, at the cost of possibly higher variance. We show that
    this trade-off is inherent, in the sense that improving it past some threshold
    would violate information-theoretic lower bounds. QSGD guarantees convergence
    for convex and non-convex objectives, under asynchrony, and can be extended to
    stochastic variance-reduced techniques. When applied to training deep neural networks
    for image classification and automated speech recognition, QSGD leads to significant
    reductions in end-to-end training time. For instance, on 16GPUs, we can train
    the ResNet-152 network to full accuracy on ImageNet 1.8 × faster than the full-precision
    variant. '
alternative_title:
- Advances in Neural Information Processing Systems
article_processing_charge: No
arxiv: 1
author:
- first_name: Dan-Adrian
  full_name: Alistarh, Dan-Adrian
  id: 4A899BFC-F248-11E8-B48F-1D18A9856A87
  last_name: Alistarh
  orcid: 0000-0003-3650-940X
- first_name: Demjan
  full_name: Grubic, Demjan
  last_name: Grubic
- first_name: Jerry
  full_name: Li, Jerry
  last_name: Li
- first_name: Ryota
  full_name: Tomioka, Ryota
  last_name: Tomioka
- first_name: Milan
  full_name: Vojnović, Milan
  last_name: Vojnović
citation:
  ama: 'Alistarh D-A, Grubic D, Li J, Tomioka R, Vojnović M. QSGD: Communication-efficient
    SGD via gradient quantization and encoding. In: Vol 2017. Neural Information Processing
    Systems Foundation; 2017:1710-1721.'
  apa: 'Alistarh, D.-A., Grubic, D., Li, J., Tomioka, R., &#38; Vojnović, M. (2017).
    QSGD: Communication-efficient SGD via gradient quantization and encoding (Vol.
    2017, pp. 1710–1721). Presented at the NIPS: Neural Information Processing System,
    Long Beach, CA, United States: Neural Information Processing Systems Foundation.'
  chicago: 'Alistarh, Dan-Adrian, Demjan Grubic, Jerry Li, Ryota Tomioka, and Milan
    Vojnović. “QSGD: Communication-Efficient SGD via Gradient Quantization and Encoding,”
    2017:1710–21. Neural Information Processing Systems Foundation, 2017.'
  ieee: 'D.-A. Alistarh, D. Grubic, J. Li, R. Tomioka, and M. Vojnović, “QSGD: Communication-efficient
    SGD via gradient quantization and encoding,” presented at the NIPS: Neural Information
    Processing System, Long Beach, CA, United States, 2017, vol. 2017, pp. 1710–1721.'
  ista: 'Alistarh D-A, Grubic D, Li J, Tomioka R, Vojnović M. 2017. QSGD: Communication-efficient
    SGD via gradient quantization and encoding. NIPS: Neural Information Processing
    System, Advances in Neural Information Processing Systems, vol. 2017, 1710–1721.'
  mla: 'Alistarh, Dan-Adrian, et al. <i>QSGD: Communication-Efficient SGD via Gradient
    Quantization and Encoding</i>. Vol. 2017, Neural Information Processing Systems
    Foundation, 2017, pp. 1710–21.'
  short: D.-A. Alistarh, D. Grubic, J. Li, R. Tomioka, M. Vojnović, in:, Neural Information
    Processing Systems Foundation, 2017, pp. 1710–1721.
conference:
  end_date: 2017-12-09
  location: Long Beach, CA, United States
  name: 'NIPS: Neural Information Processing System'
  start_date: 2017-12-04
corr_author: '1'
date_created: 2018-12-11T11:46:26Z
date_published: 2017-01-01T00:00:00Z
date_updated: 2025-09-18T10:07:20Z
day: '01'
department:
- _id: DaAl
external_id:
  arxiv:
  - '1610.02132'
  isi:
  - '000452649401072'
intvolume: '      2017'
isi: 1
language:
- iso: eng
main_file_link:
- open_access: '1'
  url: https://arxiv.org/abs/1610.02132
month: '01'
oa: 1
oa_version: Submitted Version
page: 1710-1721
publication_identifier:
  issn:
  - 1049-5258
publication_status: published
publisher: Neural Information Processing Systems Foundation
publist_id: '7392'
quality_controlled: '1'
status: public
title: 'QSGD: Communication-efficient SGD via gradient quantization and encoding'
type: conference
user_id: 317138e5-6ab7-11ef-aa6d-ffef3953e345
volume: 2017
year: '2017'
...
---
_id: '432'
abstract:
- lang: eng
  text: 'Recently there has been significant interest in training machine-learning
    models at low precision: by reducing precision, one can reduce computation and
    communication by one order of magnitude. We examine training at reduced precision,
    both from a theoretical and practical perspective, and ask: is it possible to
    train models at end-to-end low precision with provable guarantees? Can this lead
    to consistent order-of-magnitude speedups? We mainly focus on linear models, and
    the answer is yes for linear models. We develop a simple framework called ZipML
    based on one simple but novel strategy called double sampling. Our ZipML framework
    is able to execute training at low precision with no bias, guaranteeing convergence,
    whereas naive quanti- zation would introduce significant bias. We val- idate our
    framework across a range of applica- tions, and show that it enables an FPGA proto-
    type that is up to 6.5 × faster than an implemen- tation using full 32-bit precision.
    We further de- velop a variance-optimal stochastic quantization strategy and show
    that it can make a significant difference in a variety of settings. When applied
    to linear models together with double sampling, we save up to another 1.7 × in
    data movement compared with uniform quantization. When training deep networks
    with quantized models, we achieve higher accuracy than the state-of-the- art XNOR-Net. '
alternative_title:
- PMLR Press
article_processing_charge: No
author:
- first_name: Hantian
  full_name: Zhang, Hantian
  last_name: Zhang
- first_name: Jerry
  full_name: Li, Jerry
  last_name: Li
- first_name: Kaan
  full_name: Kara, Kaan
  last_name: Kara
- first_name: Dan-Adrian
  full_name: Alistarh, Dan-Adrian
  id: 4A899BFC-F248-11E8-B48F-1D18A9856A87
  last_name: Alistarh
  orcid: 0000-0003-3650-940X
- first_name: Ji
  full_name: Liu, Ji
  last_name: Liu
- first_name: Ce
  full_name: Zhang, Ce
  last_name: Zhang
citation:
  ama: 'Zhang H, Li J, Kara K, Alistarh D-A, Liu J, Zhang C. ZipML: Training linear
    models with end-to-end low precision, and a little bit of deep learning. In: <i>Proceedings
    of Machine Learning Research</i>. Vol 70. ML Research Press; 2017:4035-4043.'
  apa: 'Zhang, H., Li, J., Kara, K., Alistarh, D.-A., Liu, J., &#38; Zhang, C. (2017).
    ZipML: Training linear models with end-to-end low precision, and a little bit
    of deep learning. In <i>Proceedings of Machine Learning Research</i> (Vol. 70,
    pp. 4035–4043). Sydney, Australia: ML Research Press.'
  chicago: 'Zhang, Hantian, Jerry Li, Kaan Kara, Dan-Adrian Alistarh, Ji Liu, and
    Ce Zhang. “ZipML: Training Linear Models with End-to-End Low Precision, and a
    Little Bit of Deep Learning.” In <i>Proceedings of Machine Learning Research</i>,
    70:4035–43. ML Research Press, 2017.'
  ieee: 'H. Zhang, J. Li, K. Kara, D.-A. Alistarh, J. Liu, and C. Zhang, “ZipML: Training
    linear models with end-to-end low precision, and a little bit of deep learning,”
    in <i>Proceedings of Machine Learning Research</i>, Sydney, Australia, 2017, vol.
    70, pp. 4035–4043.'
  ista: 'Zhang H, Li J, Kara K, Alistarh D-A, Liu J, Zhang C. 2017. ZipML: Training
    linear models with end-to-end low precision, and a little bit of deep learning.
    Proceedings of Machine Learning Research. ICML: International Conference on Machine
    Learning, PMLR Press, vol. 70, 4035–4043.'
  mla: 'Zhang, Hantian, et al. “ZipML: Training Linear Models with End-to-End Low
    Precision, and a Little Bit of Deep Learning.” <i>Proceedings of Machine Learning
    Research</i>, vol. 70, ML Research Press, 2017, pp. 4035–43.'
  short: H. Zhang, J. Li, K. Kara, D.-A. Alistarh, J. Liu, C. Zhang, in:, Proceedings
    of Machine Learning Research, ML Research Press, 2017, pp. 4035–4043.
conference:
  end_date: 2017-08-11
  location: Sydney, Australia
  name: 'ICML: International Conference on Machine Learning'
  start_date: 2017-08-06
corr_author: '1'
date_created: 2018-12-11T11:46:26Z
date_published: 2017-01-01T00:00:00Z
date_updated: 2025-09-18T10:06:02Z
day: '01'
ddc:
- '000'
department:
- _id: DaAl
external_id:
  isi:
  - '000683309504015'
file:
- access_level: open_access
  checksum: 86156ba7f4318e47cef3eb9092593c10
  content_type: application/pdf
  creator: dernst
  date_created: 2019-01-22T08:23:58Z
  date_updated: 2020-07-14T12:46:26Z
  file_id: '5869'
  file_name: 2017_ICML_Zhang.pdf
  file_size: 849345
  relation: main_file
file_date_updated: 2020-07-14T12:46:26Z
has_accepted_license: '1'
isi: 1
language:
- iso: eng
month: '01'
oa: 1
oa_version: Submitted Version
page: 4035 - 4043
publication: Proceedings of Machine Learning Research
publication_identifier:
  isbn:
  - 978-151085514-4
publication_status: published
publisher: ML Research Press
publist_id: '7391'
quality_controlled: '1'
scopus_import: '1'
status: public
title: 'ZipML: Training linear models with end-to-end low precision, and a little
  bit of deep learning'
type: conference
user_id: 317138e5-6ab7-11ef-aa6d-ffef3953e345
volume: ' 70'
year: '2017'
...
