---
OA_place: repository
OA_type: green
_id: '18233'
abstract:
- lang: eng
  text: Neural network quantization enables the deployment of large models on resource-constrained
    devices. Current post-training quantization methods fall short in terms of accuracy
    for INT4 (or lower) but provide reasonable accuracy for INT8 (or above). In this
    work, we study the effect of quantization on the structure of the loss landscape.
    We show that the structure is flat and separable for mild quantization, enabling
    straightforward post-training quantization methods to achieve good results. We
    show that with more aggressive quantization, the loss landscape becomes highly
    non-separable with steep curvature, making the selection of quantization parameters
    more challenging. Armed with this understanding, we design a method that quantizes
    the layer parameters jointly, enabling significant accuracy improvement over current
    post-training quantization methods. Reference implementation is available at https://github.com/ynahshan/nn-quantization-pytorch/tree/master/lapq.
article_processing_charge: No
article_type: original
arxiv: 1
author:
- first_name: Yury
  full_name: Nahshan, Yury
  last_name: Nahshan
- first_name: Brian
  full_name: Chmiel, Brian
  last_name: Chmiel
- first_name: Chaim
  full_name: Baskin, Chaim
  last_name: Baskin
- first_name: Evgenii
  full_name: Zheltonozhskii, Evgenii
  last_name: Zheltonozhskii
- first_name: Ron
  full_name: Banner, Ron
  last_name: Banner
- first_name: Alexander
  full_name: Bronstein, Alexander
  id: 58f3726e-7cba-11ef-ad8b-e6e8cb3904e6
  last_name: Bronstein
  orcid: 0000-0001-9699-8730
- first_name: Avi
  full_name: Mendelson, Avi
  last_name: Mendelson
citation:
  ama: Nahshan Y, Chmiel B, Baskin C, et al. Loss aware post-training quantization.
    <i>Machine Learning</i>. 2021;110(11-12):3245-3262. doi:<a href="https://doi.org/10.1007/s10994-021-06053-z">10.1007/s10994-021-06053-z</a>
  apa: Nahshan, Y., Chmiel, B., Baskin, C., Zheltonozhskii, E., Banner, R., Bronstein,
    A. M., &#38; Mendelson, A. (2021). Loss aware post-training quantization. <i>Machine
    Learning</i>. Springer Nature. <a href="https://doi.org/10.1007/s10994-021-06053-z">https://doi.org/10.1007/s10994-021-06053-z</a>
  chicago: Nahshan, Yury, Brian Chmiel, Chaim Baskin, Evgenii Zheltonozhskii, Ron
    Banner, Alex M. Bronstein, and Avi Mendelson. “Loss Aware Post-Training Quantization.”
    <i>Machine Learning</i>. Springer Nature, 2021. <a href="https://doi.org/10.1007/s10994-021-06053-z">https://doi.org/10.1007/s10994-021-06053-z</a>.
  ieee: Y. Nahshan <i>et al.</i>, “Loss aware post-training quantization,” <i>Machine
    Learning</i>, vol. 110, no. 11–12. Springer Nature, pp. 3245–3262, 2021.
  ista: Nahshan Y, Chmiel B, Baskin C, Zheltonozhskii E, Banner R, Bronstein AM, Mendelson
    A. 2021. Loss aware post-training quantization. Machine Learning. 110(11–12),
    3245–3262.
  mla: Nahshan, Yury, et al. “Loss Aware Post-Training Quantization.” <i>Machine Learning</i>,
    vol. 110, no. 11–12, Springer Nature, 2021, pp. 3245–62, doi:<a href="https://doi.org/10.1007/s10994-021-06053-z">10.1007/s10994-021-06053-z</a>.
  short: Y. Nahshan, B. Chmiel, C. Baskin, E. Zheltonozhskii, R. Banner, A.M. Bronstein,
    A. Mendelson, Machine Learning 110 (2021) 3245–3262.
date_created: 2024-10-08T12:57:05Z
date_published: 2021-12-01T00:00:00Z
date_updated: 2024-10-15T07:33:28Z
day: '01'
doi: 10.1007/s10994-021-06053-z
extern: '1'
external_id:
  arxiv:
  - '1911.07190'
fulldoi: https://doi.org/10.1007/s10994-021-06053-z
intvolume: '       110'
issue: 11-12
language:
- iso: eng
main_file_link:
- open_access: '1'
  url: https://doi.org/10.48550/arXiv.1911.07190
month: '12'
oa: 1
oa_version: Preprint
page: 3245-3262
publication: Machine Learning
publication_identifier:
  eissn:
  - 1573-0565
  issn:
  - 0885-6125
publication_status: published
publisher: Springer Nature
quality_controlled: '1'
related_material:
  link:
  - relation: software
    url: https://github.com/ynahshan/nn-quantization-pytorch/tree/master/lapq
scopus_import: '1'
status: public
title: Loss aware post-training quantization
type: journal_article
user_id: 2DF688A6-F248-11E8-B48F-1D18A9856A87
volume: 110
year: '2021'
...
