---
OA_place: repository
OA_type: green
_id: '22827'
abstract:
- lang: eng
  text: Quantized training of Large Language Models (LLMs) remains an open challenge,
    as maintaining accuracy while performing all matrix multiplications in low precision
    has proven difficult. This is particularly the case when fine-tuning pre-trained
    models, which can have large weight, activation, and error (output gradient) outlier
    values that make lower-precision optimization difficult. To address this, we present
    HALO, a new quantization-aware training approach for Transformers that enables
    accurate and efficient low-precision training by combining 1) strategic placement
    of Hadamard rotations in both forward and backward passes, which mitigate outliers,
    2) high-performance kernel support, and 3) FSDP integration for low-precision
    communication. Our approach ensures that all large matrix multiplications during
    the forward and backward passes are executed in lower precision. Applied to LLaMa
    models, HALO achieves near-full-precision-equivalent results during fine-tuning
    on various tasks, while delivering up to 1.41x end-to-end speedup for full fine-tuning
    on RTX 4090 GPUs. HALO efficiently supports both standard and parameter-efficient
    fine-tuning (PEFT). Our results demonstrate the first practical approach to fully
    quantized LLM fine-tuning that maintains accuracy in INT8 and FP6 precision, while
    delivering performance benefits.
acknowledgement: "This project has received funding from the European Research Council
  (ERC) under the European\r\nUnion’s Horizon 2020 program (grant agreement PSAP,
  No. 101002047. This research also obtained\r\nfunding from the “UrbanTwin: An urban
  digital twin for climate action: Assessing policies and\r\nsolutions for energy,
  water and infrastructure” project, funded by the ETH-Domain Joint Initiative\r\nprogram
  in the Strategic Area Energy, Climate and Sustainable Environment."
alternative_title:
- Advances in Neural Information Processing Systems
article_processing_charge: No
arxiv: 1
author:
- first_name: Saleh
  full_name: Ashkboos, Saleh
  last_name: Ashkboos
- first_name: Mahdi
  full_name: Nikdan, Mahdi
  id: 66374281-f394-11eb-9cf6-869147deecc0
  last_name: Nikdan
- first_name: Soroush
  full_name: Tabesh, Soroush
  id: 06000900-6068-11ef-8d61-c2472ef2e752
  last_name: Tabesh
  orcid: 0009-0003-4119-6281
- first_name: Roberto
  full_name: Lopez Castro, Roberto
  id: 18495c31-fb57-11ef-ba0d-c290e10e394e
  last_name: Lopez Castro
- first_name: Torsten
  full_name: Hoefler, Torsten
  last_name: Hoefler
- first_name: Dan-Adrian
  full_name: Alistarh, Dan-Adrian
  id: 4A899BFC-F248-11E8-B48F-1D18A9856A87
  last_name: Alistarh
  orcid: 0000-0003-3650-940X
citation:
  ama: 'Ashkboos S, Nikdan M, Tabesh S, Lopez Castro R, Hoefler T, Alistarh D-A. HALO:
    Hadamard-assisted lower-precision optimization for LLMs. In: <i>39th Conference
    on Neural Information Processing Systems</i>. Vol 38. Neural Information Processing
    Systems Foundation; 2025:131755-131780. doi:<a href="https://doi.org/10.52202/085713-3966">10.52202/085713-3966</a>'
  apa: 'Ashkboos, S., Nikdan, M., Tabesh, S., Lopez Castro, R., Hoefler, T., &#38;
    Alistarh, D.-A. (2025). HALO: Hadamard-assisted lower-precision optimization for
    LLMs. In <i>39th Conference on Neural Information Processing Systems</i> (Vol.
    38, pp. 131755–131780). San Diego, CA, United States: Neural Information Processing
    Systems Foundation. <a href="https://doi.org/10.52202/085713-3966">https://doi.org/10.52202/085713-3966</a>'
  chicago: 'Ashkboos, Saleh, Mahdi Nikdan, Soroush Tabesh, Roberto Lopez Castro, Torsten
    Hoefler, and Dan-Adrian Alistarh. “HALO: Hadamard-Assisted Lower-Precision Optimization
    for LLMs.” In <i>39th Conference on Neural Information Processing Systems</i>,
    38:131755–80. Neural Information Processing Systems Foundation, 2025. <a href="https://doi.org/10.52202/085713-3966">https://doi.org/10.52202/085713-3966</a>.'
  ieee: 'S. Ashkboos, M. Nikdan, S. Tabesh, R. Lopez Castro, T. Hoefler, and D.-A.
    Alistarh, “HALO: Hadamard-assisted lower-precision optimization for LLMs,” in
    <i>39th Conference on Neural Information Processing Systems</i>, San Diego, CA,
    United States, 2025, vol. 38, pp. 131755–131780.'
  ista: 'Ashkboos S, Nikdan M, Tabesh S, Lopez Castro R, Hoefler T, Alistarh D-A.
    2025. HALO: Hadamard-assisted lower-precision optimization for LLMs. 39th Conference
    on Neural Information Processing Systems. NeurIPS: Neural Information Processing
    Systems, Advances in Neural Information Processing Systems, vol. 38, 131755–131780.'
  mla: 'Ashkboos, Saleh, et al. “HALO: Hadamard-Assisted Lower-Precision Optimization
    for LLMs.” <i>39th Conference on Neural Information Processing Systems</i>, vol.
    38, Neural Information Processing Systems Foundation, 2025, pp. 131755–80, doi:<a
    href="https://doi.org/10.52202/085713-3966">10.52202/085713-3966</a>.'
  short: S. Ashkboos, M. Nikdan, S. Tabesh, R. Lopez Castro, T. Hoefler, D.-A. Alistarh,
    in:, 39th Conference on Neural Information Processing Systems, Neural Information
    Processing Systems Foundation, 2025, pp. 131755–131780.
conference:
  end_date: 2025-12-07
  location: San Diego, CA, United States
  name: 'NeurIPS: Neural Information Processing Systems'
  start_date: 2025-12-02
das_tickbox: '0'
date_created: 2026-09-06T22:02:00Z
date_published: 2025-12-01T00:00:00Z
date_updated: 2026-09-17T06:39:18Z
day: '01'
department:
- _id: DaAl
- _id: GradSch
doi: 10.52202/085713-3966
external_id:
  arxiv:
  - '2501.02625'
fulldoi: https://doi.org/10.52202/085713-3966
intvolume: '        38'
language:
- iso: eng
main_file_link:
- open_access: '1'
  url: https://doi.org/10.48550/arXiv.2501.02625
month: '12'
oa: 1
oa_version: Preprint
page: 131755-131780
publication: 39th Conference on Neural Information Processing Systems
publication_identifier:
  eissn:
  - 1049-5258
  isbn:
  - '9798331338275'
publication_status: published
publisher: Neural Information Processing Systems Foundation
quality_controlled: '1'
researchdata_availability: no
scopus_import: '1'
status: public
supplementarymaterial: no
title: 'HALO: Hadamard-assisted lower-precision optimization for LLMs'
type: conference
user_id: 2DF688A6-F248-11E8-B48F-1D18A9856A87
volume: 38
year: '2025'
...
