---
OA_place: repository
OA_type: green
_id: '18976'
abstract:
- lang: eng
  text: We analyze asynchronous-type algorithms for distributed SGD in the heterogeneous
    setting, where each worker has its own computation and communication speeds, as
    well as data distribution. In these algorithms, workers compute possibly stale
    and stochastic gradients associated with their local data at some iteration back
    in history and then return those gradients to the server without synchronizing
    with other workers. We present a unified convergence theory for non-convex smooth
    functions in the heterogeneous regime. The proposed analysis provides convergence
    for pure asynchronous SGD and its various modifications. Moreover, our theory
    explains what affects the convergence rate and what can be done to improve the
    performance of asynchronous algorithms. In particular, we introduce a novel asynchronous
    method based on worker shuffling. As a by-product of our analysis, we also demonstrate
    convergence guarantees for gradient-type algorithms such as SGD with random reshuffling
    and shuffle-once mini-batch SGD. The derived rates match the best-known results
    for those algorithms, highlighting the tightness of our approach. Finally, our
    numerical evaluations support theoretical findings and show the good practical
    performance of our method.
acknowledgement: "The authors thank all anonymous reviewers for their valuable comments
  and suggestions on how to improve the manuscript. This work was done when Rustem
  Islamov was a Master’s student at Institut Polytechnique de Paris (IP Paris) and
  an intern at Institute of Science and Technology Austria (ISTA). The research of
  Rustem Islamov was supported by ISTA internship\r\nprogram. Mher Safaryan has received
  funding from the European Union’s Horizon 2020 research and innovation program under
  the Marie Skłodowska-Curie grant agreement No 101034413."
alternative_title:
- PMLR
article_processing_charge: No
arxiv: 1
author:
- first_name: Rustem
  full_name: Islamov, Rustem
  last_name: Islamov
- first_name: Mher
  full_name: Safaryan, Mher
  id: dd546b39-0804-11ed-9c55-ef075c39778d
  last_name: Safaryan
- first_name: Dan-Adrian
  full_name: Alistarh, Dan-Adrian
  id: 4A899BFC-F248-11E8-B48F-1D18A9856A87
  last_name: Alistarh
  orcid: 0000-0003-3650-940X
citation:
  ama: 'Islamov R, Safaryan M, Alistarh D-A. AsGrad: A sharp unified analysis of asynchronous-SGD
    algorithms. In: <i>Proceedings of The 27th International Conference on Artificial
    Intelligence and Statistics</i>. Vol 238. ML Research Press; 2024:649-657.'
  apa: 'Islamov, R., Safaryan, M., &#38; Alistarh, D.-A. (2024). AsGrad: A sharp unified
    analysis of asynchronous-SGD algorithms. In <i>Proceedings of The 27th International
    Conference on Artificial Intelligence and Statistics</i> (Vol. 238, pp. 649–657).
    Valencia, Spain: ML Research Press.'
  chicago: 'Islamov, Rustem, Mher Safaryan, and Dan-Adrian Alistarh. “AsGrad: A Sharp
    Unified Analysis of Asynchronous-SGD Algorithms.” In <i>Proceedings of The 27th
    International Conference on Artificial Intelligence and Statistics</i>, 238:649–57.
    ML Research Press, 2024.'
  ieee: 'R. Islamov, M. Safaryan, and D.-A. Alistarh, “AsGrad: A sharp unified analysis
    of asynchronous-SGD algorithms,” in <i>Proceedings of The 27th International Conference
    on Artificial Intelligence and Statistics</i>, Valencia, Spain, 2024, vol. 238,
    pp. 649–657.'
  ista: 'Islamov R, Safaryan M, Alistarh D-A. 2024. AsGrad: A sharp unified analysis
    of asynchronous-SGD algorithms. Proceedings of The 27th International Conference
    on Artificial Intelligence and Statistics. AISTATS: Conference on Artificial Intelligence
    and Statistics, PMLR, vol. 238, 649–657.'
  mla: 'Islamov, Rustem, et al. “AsGrad: A Sharp Unified Analysis of Asynchronous-SGD
    Algorithms.” <i>Proceedings of The 27th International Conference on Artificial
    Intelligence and Statistics</i>, vol. 238, ML Research Press, 2024, pp. 649–57.'
  short: R. Islamov, M. Safaryan, D.-A. Alistarh, in:, Proceedings of The 27th International
    Conference on Artificial Intelligence and Statistics, ML Research Press, 2024,
    pp. 649–657.
conference:
  end_date: 2024-05-04
  location: Valencia, Spain
  name: 'AISTATS: Conference on Artificial Intelligence and Statistics'
  start_date: 2024-05-02
corr_author: '1'
date_created: 2025-01-30T08:15:49Z
date_published: 2024-05-15T00:00:00Z
date_updated: 2025-04-14T07:54:52Z
day: '15'
department:
- _id: DaAl
ec_funded: 1
external_id:
  arxiv:
  - '2310.20452'
intvolume: '       238'
language:
- iso: eng
main_file_link:
- open_access: '1'
  url: https://doi.org/10.48550/arXiv.2310.20452
month: '05'
oa: 1
oa_version: Preprint
page: 649-657
project:
- _id: fc2ed2f7-9c52-11eb-aca3-c01059dda49c
  call_identifier: H2020
  grant_number: '101034413'
  name: 'IST-BRIDGE: International postdoctoral program'
publication: Proceedings of The 27th International Conference on Artificial Intelligence
  and Statistics
publication_identifier:
  eissn:
  - 2640-3498
publication_status: published
publisher: ML Research Press
quality_controlled: '1'
scopus_import: '1'
status: public
title: 'AsGrad: A sharp unified analysis of asynchronous-SGD algorithms'
type: conference
user_id: 2DF688A6-F248-11E8-B48F-1D18A9856A87
volume: 238
year: '2024'
...
---
OA_place: repository
OA_type: green
_id: '18977'
abstract:
- lang: eng
  text: "Recent advances in large language model (LLM) pretraining have led to high-quality
    LLMs with impressive abilities. By compressing such LLMs via quantization to 3-4
    bits per parameter, they can fit into memory-limited devices such as laptops and
    mobile phones, enabling personalized use. Quantizing models to 3-4 bits per parameter
    can lead to moderate to high accuracy losses, especially for smaller models (1-10B
    parameters), which are suitable for edge deployment. To address this accuracy
    issue, we introduce the Sparse-Quantized Representation (SpQR), a new compressed
    format and quantization technique that enables for the first time \\emph{near-lossless}
    compression of LLMs across model scales while reaching similar compression levels
    to previous methods. SpQR works by identifying and isolating \\emph{outlier weights},
    which cause particularly large quantization errors, and storing them in higher
    precision while compressing all other weights to 3-4 bits, and achieves relative
    accuracy losses of less than \r\n in perplexity for highly-accurate LLaMA and
    Falcon LLMs. This makes it possible to run a 33B parameter LLM on a single 24
    GB consumer GPU without performance degradation at 15% speedup, thus making powerful
    LLMs available to consumers without any downsides. SpQR comes with efficient algorithms
    for both encoding weights into its format, as well as decoding them efficiently
    at runtime. Specifically, we provide an efficient GPU inference algorithm for
    SpQR, which yields faster inference than 16-bit baselines at similar accuracy
    while enabling memory compression gains of more than 4x."
acknowledgement: "Denis Kuznedelev acknowledges the support from the Russian Ministry
  of Science and Higher\r\nEducation, grant No. 075-10-2021-068. Ruslan Svirschevski
  and Vage Egiazarian and Denis\r\nKuznedelev were supported by the grant for research
  centers in the field of AI provided by the\r\nAnalytical Center for the Government
  of the Russian Federation (ACRF) in accordance with the\r\nagreement on the provision
  of subsidies (identifier of the agreement 000000D730321P5Q0002) and the agreement
  with HSE University No. 70-2021-00139."
article_processing_charge: No
arxiv: 1
author:
- first_name: Tim
  full_name: Dettmers, Tim
  last_name: Dettmers
- first_name: Ruslan A.
  full_name: Svirschevski, Ruslan A.
  last_name: Svirschevski
- first_name: Vage
  full_name: Egiazarian, Vage
  last_name: Egiazarian
- first_name: Denis
  full_name: Kuznedelev, Denis
  last_name: Kuznedelev
- first_name: Elias
  full_name: Frantar, Elias
  id: 09a8f98d-ec99-11ea-ae11-c063a7b7fe5f
  last_name: Frantar
- first_name: Saleh
  full_name: Ashkboos, Saleh
  last_name: Ashkboos
- first_name: Alexander
  full_name: Borzunov, Alexander
  last_name: Borzunov
- first_name: Torsten
  full_name: Hoefler, Torsten
  last_name: Hoefler
- first_name: Dan-Adrian
  full_name: Alistarh, Dan-Adrian
  id: 4A899BFC-F248-11E8-B48F-1D18A9856A87
  last_name: Alistarh
  orcid: 0000-0003-3650-940X
citation:
  ama: 'Dettmers T, Svirschevski RA, Egiazarian V, et al. SpQR: A sparse-quantized
    representation for near-lossless LLM weight compression. In: <i>12th International
    Conference on Learning Representations</i>. OpenReview; 2024.'
  apa: 'Dettmers, T., Svirschevski, R. A., Egiazarian, V., Kuznedelev, D., Frantar,
    E., Ashkboos, S., … Alistarh, D.-A. (2024). SpQR: A sparse-quantized representation
    for near-lossless LLM weight compression. In <i>12th International Conference
    on Learning Representations</i>. Vienna, Austria: OpenReview.'
  chicago: 'Dettmers, Tim, Ruslan A. Svirschevski, Vage Egiazarian, Denis Kuznedelev,
    Elias Frantar, Saleh Ashkboos, Alexander Borzunov, Torsten Hoefler, and Dan-Adrian
    Alistarh. “SpQR: A Sparse-Quantized Representation for near-Lossless LLM Weight
    Compression.” In <i>12th International Conference on Learning Representations</i>.
    OpenReview, 2024.'
  ieee: 'T. Dettmers <i>et al.</i>, “SpQR: A sparse-quantized representation for near-lossless
    LLM weight compression,” in <i>12th International Conference on Learning Representations</i>,
    Vienna, Austria, 2024.'
  ista: 'Dettmers T, Svirschevski RA, Egiazarian V, Kuznedelev D, Frantar E, Ashkboos
    S, Borzunov A, Hoefler T, Alistarh D-A. 2024. SpQR: A sparse-quantized representation
    for near-lossless LLM weight compression. 12th International Conference on Learning
    Representations. ICLR: International Conference on Learning Representations.'
  mla: 'Dettmers, Tim, et al. “SpQR: A Sparse-Quantized Representation for near-Lossless
    LLM Weight Compression.” <i>12th International Conference on Learning Representations</i>,
    OpenReview, 2024.'
  short: T. Dettmers, R.A. Svirschevski, V. Egiazarian, D. Kuznedelev, E. Frantar,
    S. Ashkboos, A. Borzunov, T. Hoefler, D.-A. Alistarh, in:, 12th International
    Conference on Learning Representations, OpenReview, 2024.
conference:
  end_date: 2024-05-11
  location: Vienna, Austria
  name: 'ICLR: International Conference on Learning Representations'
  start_date: 2024-05-07
date_created: 2025-01-30T08:26:59Z
date_published: 2024-05-15T00:00:00Z
date_updated: 2025-01-30T08:27:47Z
day: '15'
department:
- _id: DaAl
external_id:
  arxiv:
  - '2306.03078'
language:
- iso: eng
main_file_link:
- open_access: '1'
  url: https://doi.org/10.48550/arXiv.2306.03078
month: '05'
oa: 1
oa_version: Preprint
publication: 12th International Conference on Learning Representations
publication_status: published
publisher: OpenReview
quality_controlled: '1'
scopus_import: '1'
status: public
title: 'SpQR: A sparse-quantized representation for near-lossless LLM weight compression'
type: conference
user_id: 2DF688A6-F248-11E8-B48F-1D18A9856A87
year: '2024'
...
---
OA_place: repository
_id: '18981'
abstract:
- lang: eng
  text: We establish several results combining discrete Morse theory and microlocal
    sheaf theory in the setting of finite posets and simplicial complexes. Our primary
    tool is a computationally tractable description of the bounded derived category
    of sheaves on a poset with the Alexandrov topology. We prove that each bounded
    complex of sheaves on a finite poset admits a unique (up to isomorphism of complexes)
    minimal injective resolution, and we provide algorithms for computing minimal
    injective resolution of an injective complex, as well as several useful functors
    between derived categories of sheaves. For the constant sheaf on a simplicial
    complex, we give asymptotically tight bounds on the complexity of computing the
    minimal injective resolution using those algorithms. Our main result is a novel
    definition of the discrete microsupport of a bounded complex of sheaves on a finite
    poset. We detail several foundational properties of the discrete microsupport,
    as well as a microlocal generalization of the discrete homological Morse theorem
    and Morse inequalities.
acknowledgement: "This project has received funding from the European Research Council
  (ERC) under the European\r\nUnion’s Horizon 2020 research and innovation programme,
  grant no. 788183, from the Wittgenstein Prize,\r\nAustrian Science Fund (FWF), grant
  no. Z 342-N31, and from the DFG Collaborative Research Center TRR\r\n109, ‘Discretization
  in Geometry and Dynamics’, Austrian Science Fund (FWF), grant no. I 02979-N35."
article_processing_charge: No
arxiv: 1
author:
- first_name: Adam
  full_name: Brown, Adam
  last_name: Brown
- first_name: Ondrej
  full_name: Draganov, Ondrej
  id: 2B23F01E-F248-11E8-B48F-1D18A9856A87
  last_name: Draganov
  orcid: 0000-0003-0464-3823
citation:
  ama: Brown A, Draganov O. Discrete microlocal Morse theory. <i>arXiv</i>. doi:<a
    href="https://doi.org/10.48550/arXiv.2209.14993">10.48550/arXiv.2209.14993</a>
  apa: Brown, A., &#38; Draganov, O. (n.d.). Discrete microlocal Morse theory. <i>arXiv</i>.
    <a href="https://doi.org/10.48550/arXiv.2209.14993">https://doi.org/10.48550/arXiv.2209.14993</a>
  chicago: Brown, Adam, and Ondrej Draganov. “Discrete Microlocal Morse Theory.” <i>ArXiv</i>,
    n.d. <a href="https://doi.org/10.48550/arXiv.2209.14993">https://doi.org/10.48550/arXiv.2209.14993</a>.
  ieee: A. Brown and O. Draganov, “Discrete microlocal Morse theory,” <i>arXiv</i>.
    .
  ista: Brown A, Draganov O. Discrete microlocal Morse theory. arXiv, <a href="https://doi.org/10.48550/arXiv.2209.14993">10.48550/arXiv.2209.14993</a>.
  mla: Brown, Adam, and Ondrej Draganov. “Discrete Microlocal Morse Theory.” <i>ArXiv</i>,
    doi:<a href="https://doi.org/10.48550/arXiv.2209.14993">10.48550/arXiv.2209.14993</a>.
  short: A. Brown, O. Draganov, ArXiv (n.d.).
corr_author: '1'
date_created: 2025-01-31T17:03:04Z
date_published: 2024-06-09T00:00:00Z
date_updated: 2026-04-07T11:47:29Z
day: '09'
department:
- _id: HeEd
doi: 10.48550/arXiv.2209.14993
ec_funded: 1
external_id:
  arxiv:
  - '2209.14993'
fulldoi: https://doi.org/10.48550/arXiv.2209.14993
language:
- iso: eng
main_file_link:
- open_access: '1'
  url: https://doi.org/10.48550/arXiv.2209.14993
month: '06'
oa: 1
oa_version: Preprint
project:
- _id: 266A2E9E-B435-11E9-9278-68D0E5697425
  call_identifier: H2020
  grant_number: '788183'
  name: Alpha Shape Theory Extended
- _id: 268116B8-B435-11E9-9278-68D0E5697425
  call_identifier: FWF
  grant_number: Z00342
  name: Mathematics, Computer Science
- _id: 2561EBF4-B435-11E9-9278-68D0E5697425
  call_identifier: FWF
  grant_number: I02979-N35
  name: Persistence and stability of geometric complexes
publication: arXiv
publication_status: draft
related_material:
  record:
  - id: '20323'
    relation: later_version
    status: public
  - id: '18979'
    relation: dissertation_contains
    status: public
status: public
title: Discrete microlocal Morse theory
tmp:
  image: /images/cc_by.png
  legal_code_url: https://creativecommons.org/licenses/by/4.0/legalcode
  name: Creative Commons Attribution 4.0 International Public License (CC-BY 4.0)
  short: CC BY (4.0)
type: preprint
user_id: 2DF688A6-F248-11E8-B48F-1D18A9856A87
year: '2024'
...
---
OA_place: repository
OA_type: green
_id: '18996'
abstract:
- lang: eng
  text: 'We consider the linear causal representation learning setting where we observe
    a linear mixing of d unknown latent factors, which follow a linear structural
    causal model. Recent work has shown that it is possible to recover the latent
    factors as well as the underlying structural causal model over them, up to permutation
    and scaling, provided that we have at least d environments, each of which corresponds
    to perfect interventions on a single latent node (factor). After this powerful
    result, a key open problem faced by the community has been to relax these conditions:
    allow for coarser than perfect single-node interventions, and allow for fewer
    than d of them, since the number of latent factors d could be very large. In this
    work, we consider precisely such a setting, where we allow a smaller than d number
    of environments, and also allow for very coarse interventions that can very coarsely
    \textit{change the entire causal graph over the latent factors}. On the flip side,
    we relax what we wish to extract to simply the \textit{list of nodes that have
    shifted between one or more environments}. We provide a surprising identifiability
    result that it is indeed possible, under some very mild standard assumptions,
    to identify the set of shifted nodes. Our identifiability proof moreover is a
    constructive one: we explicitly provide necessary and sufficient conditions for
    a node to be a shifted node, and show that we can check these conditions given
    observed data. Our algorithm lends itself very naturally to the sample setting
    where instead of just interventional distributions, we are provided datasets of
    samples from each of these distributions. We corroborate our results on both synthetic
    experiments as well as an interesting psychometric dataset. The code can be found
    at https://github.com/TianyuCodings/iLCS.'
alternative_title:
- Advances in Neural Information Processing Systems
article_processing_charge: No
arxiv: 1
author:
- first_name: Tianyu
  full_name: Chen, Tianyu
  last_name: Chen
- first_name: Kevin
  full_name: Bello, Kevin
  last_name: Bello
- first_name: Francesco
  full_name: Locatello, Francesco
  id: 26cfd52f-2483-11ee-8040-88983bcc06d4
  last_name: Locatello
  orcid: 0000-0002-4850-0683
- first_name: Bryon
  full_name: Aragam, Bryon
  last_name: Aragam
- first_name: Pradeep Kumar
  full_name: Ravikumar, Pradeep Kumar
  last_name: Ravikumar
citation:
  ama: 'Chen T, Bello K, Locatello F, Aragam B, Ravikumar PK. Identifying general
    mechanism shifts in linear causal representations. In: <i>38th Conference on Neural
    Information Processing Systems</i>. Vol 37. Neural Information Processing Systems
    Foundation; 2024.'
  apa: 'Chen, T., Bello, K., Locatello, F., Aragam, B., &#38; Ravikumar, P. K. (2024).
    Identifying general mechanism shifts in linear causal representations. In <i>38th
    Conference on Neural Information Processing Systems</i> (Vol. 37). Vancouver,
    Canada: Neural Information Processing Systems Foundation.'
  chicago: Chen, Tianyu, Kevin Bello, Francesco Locatello, Bryon Aragam, and Pradeep
    Kumar Ravikumar. “Identifying General Mechanism Shifts in Linear Causal Representations.”
    In <i>38th Conference on Neural Information Processing Systems</i>, Vol. 37. Neural
    Information Processing Systems Foundation, 2024.
  ieee: T. Chen, K. Bello, F. Locatello, B. Aragam, and P. K. Ravikumar, “Identifying
    general mechanism shifts in linear causal representations,” in <i>38th Conference
    on Neural Information Processing Systems</i>, Vancouver, Canada, 2024, vol. 37.
  ista: 'Chen T, Bello K, Locatello F, Aragam B, Ravikumar PK. 2024. Identifying general
    mechanism shifts in linear causal representations. 38th Conference on Neural Information
    Processing Systems. NeurIPS: Neural Information Processing Systems, Advances in
    Neural Information Processing Systems, vol. 37.'
  mla: Chen, Tianyu, et al. “Identifying General Mechanism Shifts in Linear Causal
    Representations.” <i>38th Conference on Neural Information Processing Systems</i>,
    vol. 37, Neural Information Processing Systems Foundation, 2024.
  short: T. Chen, K. Bello, F. Locatello, B. Aragam, P.K. Ravikumar, in:, 38th Conference
    on Neural Information Processing Systems, Neural Information Processing Systems
    Foundation, 2024.
conference:
  end_date: 2024-12-16
  location: Vancouver, Canada
  name: 'NeurIPS: Neural Information Processing Systems'
  start_date: 2024-12-16
date_created: 2025-02-04T13:09:34Z
date_published: 2024-09-25T00:00:00Z
date_updated: 2025-07-07T13:23:49Z
day: '25'
ddc:
- '000'
department:
- _id: FrLo
external_id:
  arxiv:
  - '2410.24059'
file:
- access_level: open_access
  checksum: 75c3091e70bd2916cd94afbf40a0c425
  content_type: application/pdf
  creator: dernst
  date_created: 2025-02-04T13:09:08Z
  date_updated: 2025-02-04T13:09:08Z
  file_id: '18997'
  file_name: 2024_NeurIPS_Chen.pdf
  file_size: 5659119
  relation: main_file
  success: 1
file_date_updated: 2025-02-04T13:09:08Z
has_accepted_license: '1'
intvolume: '        37'
language:
- iso: eng
month: '09'
oa: 1
oa_version: Published Version
publication: 38th Conference on Neural Information Processing Systems
publication_identifier:
  eissn:
  - 1049-5258
publication_status: published
publisher: Neural Information Processing Systems Foundation
quality_controlled: '1'
scopus_import: '1'
status: public
title: Identifying general mechanism shifts in linear causal representations
tmp:
  image: /images/cc_by.png
  legal_code_url: https://creativecommons.org/licenses/by/4.0/legalcode
  name: Creative Commons Attribution 4.0 International Public License (CC-BY 4.0)
  short: CC BY (4.0)
type: conference
user_id: 2DF688A6-F248-11E8-B48F-1D18A9856A87
volume: 37
year: '2024'
...
---
OA_place: publisher
OA_type: gold
_id: '18998'
abstract:
- lang: eng
  text: Word embeddings represent language vocabularies as clouds of d-dimensional
    points. We investigate how information is conveyed by the general shape of these
    clouds, instead of representing the semantic meaning of each token. Specifically,
    we use the notion of persistent homology from topological data analysis (TDA)
    to measure the distances between language pairs from the shape of their unlabeled
    embeddings. These distances quantify the degree of non-isometry of the embeddings.
    To distinguish whether these differences are random training errors or capture
    real information about the languages, we use the computed distance matrices to
    construct language phylogenetic trees over 81 Indo-European languages. Careful
    evaluation shows that our reconstructed trees exhibit strong and statistically-significant
    similarities to the reference.
article_processing_charge: No
arxiv: 1
author:
- first_name: Ondrej
  full_name: Draganov, Ondrej
  id: 2B23F01E-F248-11E8-B48F-1D18A9856A87
  last_name: Draganov
  orcid: 0000-0003-0464-3823
- first_name: Steven
  full_name: Skiena, Steven
  last_name: Skiena
citation:
  ama: 'Draganov O, Skiena S. The shape of word embeddings: Quantifying non-isometry
    with topological data analysis. In: <i>Findings of the Association for Computational
    Linguistics: EMNLP 2024</i>. Association for Computational Linguistics; 2024:12080-12099.
    doi:<a href="https://doi.org/10.18653/v1/2024.findings-emnlp.705">10.18653/v1/2024.findings-emnlp.705</a>'
  apa: 'Draganov, O., &#38; Skiena, S. (2024). The shape of word embeddings: Quantifying
    non-isometry with topological data analysis. In <i>Findings of the Association
    for Computational Linguistics: EMNLP 2024</i> (pp. 12080–12099). Miami, FL, United
    States: Association for Computational Linguistics. <a href="https://doi.org/10.18653/v1/2024.findings-emnlp.705">https://doi.org/10.18653/v1/2024.findings-emnlp.705</a>'
  chicago: 'Draganov, Ondrej, and Steven Skiena. “The Shape of Word Embeddings: Quantifying
    Non-Isometry with Topological Data Analysis.” In <i>Findings of the Association
    for Computational Linguistics: EMNLP 2024</i>, 12080–99. Association for Computational
    Linguistics, 2024. <a href="https://doi.org/10.18653/v1/2024.findings-emnlp.705">https://doi.org/10.18653/v1/2024.findings-emnlp.705</a>.'
  ieee: 'O. Draganov and S. Skiena, “The shape of word embeddings: Quantifying non-isometry
    with topological data analysis,” in <i>Findings of the Association for Computational
    Linguistics: EMNLP 2024</i>, Miami, FL, United States, 2024, pp. 12080–12099.'
  ista: 'Draganov O, Skiena S. 2024. The shape of word embeddings: Quantifying non-isometry
    with topological data analysis. Findings of the Association for Computational
    Linguistics: EMNLP 2024. EMNLP: Conference on Empirical Methods in Natural Language
    Processing, 12080–12099.'
  mla: 'Draganov, Ondrej, and Steven Skiena. “The Shape of Word Embeddings: Quantifying
    Non-Isometry with Topological Data Analysis.” <i>Findings of the Association for
    Computational Linguistics: EMNLP 2024</i>, Association for Computational Linguistics,
    2024, pp. 12080–99, doi:<a href="https://doi.org/10.18653/v1/2024.findings-emnlp.705">10.18653/v1/2024.findings-emnlp.705</a>.'
  short: 'O. Draganov, S. Skiena, in:, Findings of the Association for Computational
    Linguistics: EMNLP 2024, Association for Computational Linguistics, 2024, pp.
    12080–12099.'
conference:
  end_date: 2024-11-16
  location: Miami, FL, United States
  name: 'EMNLP: Conference on Empirical Methods in Natural Language Processing'
  start_date: 2024-11-12
corr_author: '1'
date_created: 2025-02-04T16:19:28Z
date_published: 2024-11-01T00:00:00Z
date_updated: 2025-02-10T08:21:37Z
day: '01'
ddc:
- '500'
department:
- _id: GradSch
- _id: HeEd
doi: 10.18653/v1/2024.findings-emnlp.705
external_id:
  arxiv:
  - '2404.00500'
file:
- access_level: open_access
  checksum: f4416a5962194f0181ab0dc7f9ef93c0
  content_type: application/pdf
  creator: dernst
  date_created: 2025-02-10T08:20:34Z
  date_updated: 2025-02-10T08:20:34Z
  file_id: '19016'
  file_name: 2024_EMNLP_Draganov.pdf
  file_size: 1312638
  relation: main_file
  success: 1
file_date_updated: 2025-02-10T08:20:34Z
fulldoi: https://doi.org/10.18653/v1/2024.findings-emnlp.705
has_accepted_license: '1'
language:
- iso: eng
month: '11'
oa: 1
oa_version: Published Version
page: 12080-12099
publication: 'Findings of the Association for Computational Linguistics: EMNLP 2024'
publication_status: published
publisher: Association for Computational Linguistics
quality_controlled: '1'
scopus_import: '1'
status: public
title: 'The shape of word embeddings: Quantifying non-isometry with topological data
  analysis'
tmp:
  image: /images/cc_by.png
  legal_code_url: https://creativecommons.org/licenses/by/4.0/legalcode
  name: Creative Commons Attribution 4.0 International Public License (CC-BY 4.0)
  short: CC BY (4.0)
type: conference
user_id: 2DF688A6-F248-11E8-B48F-1D18A9856A87
year: '2024'
...
---
OA_place: repository
OA_type: green
_id: '18999'
abstract:
- lang: eng
  text: Exploring the shape of point configurations has been a key driver in the evolution
    of TDA (short for topological data analysis) since its infancy. This survey illustrates
    the recent efforts to broaden these ideas to model spatial interactions among
    multiple configurations, each distinguished by a color. It describes advances
    in this area and prepares the ground for further exploration by mentioning unresolved
    questions and promising research avenues while focusing on the overlap with discrete
    geometry.
article_number: '2406.04102'
article_processing_charge: No
arxiv: 1
author:
- first_name: Sebastiano
  full_name: Cultrera di Montesano, Sebastiano
  id: 34D2A09C-F248-11E8-B48F-1D18A9856A87
  last_name: Cultrera di Montesano
  orcid: 0000-0001-6249-0832
- first_name: Ondrej
  full_name: Draganov, Ondrej
  id: 2B23F01E-F248-11E8-B48F-1D18A9856A87
  last_name: Draganov
  orcid: 0000-0003-0464-3823
- first_name: Herbert
  full_name: Edelsbrunner, Herbert
  id: 3FB178DA-F248-11E8-B48F-1D18A9856A87
  last_name: Edelsbrunner
  orcid: 0000-0002-9823-6833
- first_name: Morteza
  full_name: Saghafian, Morteza
  id: f86f7148-b140-11ec-9577-95435b8df824
  last_name: Saghafian
citation:
  ama: Cultrera di Montesano S, Draganov O, Edelsbrunner H, Saghafian M. Chromatic
    topological data analysis. <i>arXiv</i>. doi:<a href="https://doi.org/10.48550/ARXIV.2406.04102">10.48550/ARXIV.2406.04102</a>
  apa: Cultrera di Montesano, S., Draganov, O., Edelsbrunner, H., &#38; Saghafian,
    M. (n.d.). Chromatic topological data analysis. <i>arXiv</i>. <a href="https://doi.org/10.48550/ARXIV.2406.04102">https://doi.org/10.48550/ARXIV.2406.04102</a>
  chicago: Cultrera di Montesano, Sebastiano, Ondrej Draganov, Herbert Edelsbrunner,
    and Morteza Saghafian. “Chromatic Topological Data Analysis.” <i>ArXiv</i>, n.d.
    <a href="https://doi.org/10.48550/ARXIV.2406.04102">https://doi.org/10.48550/ARXIV.2406.04102</a>.
  ieee: S. Cultrera di Montesano, O. Draganov, H. Edelsbrunner, and M. Saghafian,
    “Chromatic topological data analysis,” <i>arXiv</i>. .
  ista: Cultrera di Montesano S, Draganov O, Edelsbrunner H, Saghafian M. Chromatic
    topological data analysis. arXiv, 2406.04102.
  mla: Cultrera di Montesano, Sebastiano, et al. “Chromatic Topological Data Analysis.”
    <i>ArXiv</i>, 2406.04102, doi:<a href="https://doi.org/10.48550/ARXIV.2406.04102">10.48550/ARXIV.2406.04102</a>.
  short: S. Cultrera di Montesano, O. Draganov, H. Edelsbrunner, M. Saghafian, ArXiv
    (n.d.).
corr_author: '1'
date_created: 2025-02-04T16:21:21Z
date_published: 2024-06-06T00:00:00Z
date_updated: 2025-02-10T08:14:27Z
day: '06'
ddc:
- '510'
department:
- _id: GradSch
- _id: HeEd
doi: 10.48550/ARXIV.2406.04102
external_id:
  arxiv:
  - '2406.04102'
fulldoi: https://doi.org/10.48550/ARXIV.2406.04102
has_accepted_license: '1'
language:
- iso: eng
main_file_link:
- open_access: '1'
  url: https://doi.org/10.48550/arXiv.2406.04102
month: '06'
oa: 1
oa_version: Preprint
publication: arXiv
publication_status: submitted
status: public
title: Chromatic topological data analysis
tmp:
  image: /images/cc_by.png
  legal_code_url: https://creativecommons.org/licenses/by/4.0/legalcode
  name: Creative Commons Attribution 4.0 International Public License (CC-BY 4.0)
  short: CC BY (4.0)
type: preprint
user_id: 2DF688A6-F248-11E8-B48F-1D18A9856A87
year: '2024'
...
---
OA_place: publisher
OA_type: gold
_id: '19005'
abstract:
- lang: eng
  text: "Causal representation learning promises to extend causal models to hidden
    causal\r\nvariables from raw entangled measurements. However, most progress has
    focused\r\non proving identifiability results in different settings, and we are
    not aware of any\r\nsuccessful real-world application. At the same time, the field
    of dynamical systems\r\nbenefited from deep learning and scaled to countless applications
    but does not allow\r\nparameter identification. In this paper, we draw a clear
    connection between the two\r\nand their key assumptions, allowing us to apply
    identifiable methods developed\r\nin causal representation learning to dynamical
    systems. At the same time, we can\r\nleverage scalable differentiable solvers
    developed for differential equations to build\r\nmodels that are both identifiable
    and practical. Overall, we learn explicitly controllable models that isolate the
    trajectory-specific parameters for further downstream\r\ntasks such as out-of-distribution
    classification or treatment effect estimation. We\r\nexperiment with a wind simulator
    with partially known factors of variation. We\r\nalso apply the resulting model
    to real-world climate data and successfully answer\r\ndownstream causal questions
    in line with existing literature on climate change.\r\nCode is available at https://github.com/CausalLearningAI/crl-dynamical-systems."
acknowledgement: "We thank Niklas Boers for recommending the SpeedyWeather simulator
  and Valentino Maiorca\r\nfor guidance on Fourier transformation for SST data. We
  are also grateful to Shimeng Huang and Riccardo Cadei for their feedback on the
  treatment effect estimation experiment and to Jiale Chen and Adeel Pervez for their
  assistance with the solver implementation. Finally, we appreciate the anonymous
  reviewers for their insightful suggestions, which helped improve the manuscript. "
alternative_title:
- Advances in Neural Information Processing Systems
article_processing_charge: No
arxiv: 1
author:
- first_name: Dingling
  full_name: Yao, Dingling
  id: d3e02e50-48a8-11ee-8f62-c108061797fa
  last_name: Yao
- first_name: Caroline J
  full_name: Muller, Caroline J
  id: f978ccb0-3f7f-11eb-b193-b0e2bd13182b
  last_name: Muller
  orcid: 0000-0001-5836-5350
- first_name: Francesco
  full_name: Locatello, Francesco
  id: 26cfd52f-2483-11ee-8040-88983bcc06d4
  last_name: Locatello
  orcid: 0000-0002-4850-0683
citation:
  ama: 'Yao D, Muller CJ, Locatello F. Marrying causal representation learning with
    dynamical systems for science. In: <i>38th Conference on Neural Information Processing
    Systems</i>. Vol 37. Neural Information Processing Systems Foundation; 2024.'
  apa: 'Yao, D., Muller, C. J., &#38; Locatello, F. (2024). Marrying causal representation
    learning with dynamical systems for science. In <i>38th Conference on Neural Information
    Processing Systems</i> (Vol. 37). Vancouver, Canada: Neural Information Processing
    Systems Foundation.'
  chicago: Yao, Dingling, Caroline J Muller, and Francesco Locatello. “Marrying Causal
    Representation Learning with Dynamical Systems for Science.” In <i>38th Conference
    on Neural Information Processing Systems</i>, Vol. 37. Neural Information Processing
    Systems Foundation, 2024.
  ieee: D. Yao, C. J. Muller, and F. Locatello, “Marrying causal representation learning
    with dynamical systems for science,” in <i>38th Conference on Neural Information
    Processing Systems</i>, Vancouver, Canada, 2024, vol. 37.
  ista: 'Yao D, Muller CJ, Locatello F. 2024. Marrying causal representation learning
    with dynamical systems for science. 38th Conference on Neural Information Processing
    Systems. NeurIPS: Neural Information Processing Systems, Advances in Neural Information
    Processing Systems, vol. 37.'
  mla: Yao, Dingling, et al. “Marrying Causal Representation Learning with Dynamical
    Systems for Science.” <i>38th Conference on Neural Information Processing Systems</i>,
    vol. 37, Neural Information Processing Systems Foundation, 2024.
  short: D. Yao, C.J. Muller, F. Locatello, in:, 38th Conference on Neural Information
    Processing Systems, Neural Information Processing Systems Foundation, 2024.
conference:
  end_date: 2024-12-16
  location: Vancouver, Canada
  name: 'NeurIPS: Neural Information Processing Systems'
  start_date: 2024-12-16
corr_author: '1'
date_created: 2025-02-05T07:49:00Z
date_published: 2024-12-01T00:00:00Z
date_updated: 2025-07-10T11:51:32Z
day: '01'
ddc:
- '000'
- '550'
department:
- _id: CaMu
- _id: FrLo
external_id:
  arxiv:
  - '2405.13888'
file:
- access_level: open_access
  checksum: fe8832367e7143876f178244385d859e
  content_type: application/pdf
  creator: dernst
  date_created: 2025-02-05T07:44:58Z
  date_updated: 2025-02-05T07:44:58Z
  file_id: '19006'
  file_name: 2024_NeurIPS_Yao.pdf
  file_size: 2595855
  relation: main_file
  success: 1
file_date_updated: 2025-02-05T07:44:58Z
has_accepted_license: '1'
intvolume: '        37'
language:
- iso: eng
month: '12'
oa: 1
oa_version: Published Version
publication: 38th Conference on Neural Information Processing Systems
publication_status: published
publisher: Neural Information Processing Systems Foundation
quality_controlled: '1'
related_material:
  link:
  - relation: software
    url: https://github.com/CausalLearningAI/crl-dynamical-systems
scopus_import: '1'
status: public
title: Marrying causal representation learning with dynamical systems for science
tmp:
  image: /images/cc_by.png
  legal_code_url: https://creativecommons.org/licenses/by/4.0/legalcode
  name: Creative Commons Attribution 4.0 International Public License (CC-BY 4.0)
  short: CC BY (4.0)
type: conference
user_id: 2DF688A6-F248-11E8-B48F-1D18A9856A87
volume: 37
year: '2024'
...
---
OA_place: publisher
OA_type: hybrid
_id: '19007'
abstract:
- lang: eng
  text: "Learning modular object-centric representations is crucial for systematic
    generalization. Existing methods show promising object-binding capabilities empirically,\r\nbut
    theoretical identifiability guarantees remain relatively underdeveloped. Understanding
    when object-centric representations can theoretically be identified is\r\ncrucial
    for scaling slot-based methods to high-dimensional images with correctness\r\nguarantees.
    To that end, we propose a probabilistic slot-attention algorithm that\r\nimposes
    an aggregate mixture prior over object-centric slot representations, thereby\r\nproviding
    slot identifiability guarantees without supervision, up to an equivalence\r\nrelation.
    We provide empirical verification of our theoretical identifiability result\r\nusing
    both simple 2-dimensional data and high-resolution imaging datasets.\r\n"
acknowledgement: A. Kori is supported by UKRI (grant number EP/S023356/1), as part
  of the UKRI Centre for Doctoral Training in Safe and Trusted AI. B. Glocker and
  F.D.S. Ribeiro acknowledge the support of the UKRI AI programme, and the Engineering
  and Physical Sciences Research Council, for CHAI - EPSRC Causality in Healthcare
  AI Hub (grant number EP/Y028856/1).
alternative_title:
- Advances in Neural Information Processing Systems
article_processing_charge: No
arxiv: 1
author:
- first_name: Avinash
  full_name: Kori, Avinash
  last_name: Kori
- first_name: Francesco
  full_name: Locatello, Francesco
  id: 26cfd52f-2483-11ee-8040-88983bcc06d4
  last_name: Locatello
  orcid: 0000-0002-4850-0683
- first_name: Ainkaran
  full_name: Santhirasekaram, Ainkaran
  last_name: Santhirasekaram
- first_name: Francesca
  full_name: Toni, Francesca
  last_name: Toni
- first_name: Ben
  full_name: Glocker, Ben
  last_name: Glocker
- first_name: Fabio
  full_name: De Sousa Ribeiro, Fabio
  last_name: De Sousa Ribeiro
citation:
  ama: 'Kori A, Locatello F, Santhirasekaram A, Toni F, Glocker B, De Sousa Ribeiro
    F. Identifiable object-centric representation learning via probabilistic slot
    attention. In: <i>38th Conference on Neural Information Processing Systems</i>.
    Vol 37. Neural Information Processing Systems Foundation; 2024.'
  apa: 'Kori, A., Locatello, F., Santhirasekaram, A., Toni, F., Glocker, B., &#38;
    De Sousa Ribeiro, F. (2024). Identifiable object-centric representation learning
    via probabilistic slot attention. In <i>38th Conference on Neural Information
    Processing Systems</i> (Vol. 37). Vancouver, Canada: Neural Information Processing
    Systems Foundation.'
  chicago: Kori, Avinash, Francesco Locatello, Ainkaran Santhirasekaram, Francesca
    Toni, Ben Glocker, and Fabio De Sousa Ribeiro. “Identifiable Object-Centric Representation
    Learning via Probabilistic Slot Attention.” In <i>38th Conference on Neural Information
    Processing Systems</i>, Vol. 37. Neural Information Processing Systems Foundation,
    2024.
  ieee: A. Kori, F. Locatello, A. Santhirasekaram, F. Toni, B. Glocker, and F. De
    Sousa Ribeiro, “Identifiable object-centric representation learning via probabilistic
    slot attention,” in <i>38th Conference on Neural Information Processing Systems</i>,
    Vancouver, Canada, 2024, vol. 37.
  ista: 'Kori A, Locatello F, Santhirasekaram A, Toni F, Glocker B, De Sousa Ribeiro
    F. 2024. Identifiable object-centric representation learning via probabilistic
    slot attention. 38th Conference on Neural Information Processing Systems. NeurIPS:
    Neural Information Processing Systems, Advances in Neural Information Processing
    Systems, vol. 37.'
  mla: Kori, Avinash, et al. “Identifiable Object-Centric Representation Learning
    via Probabilistic Slot Attention.” <i>38th Conference on Neural Information Processing
    Systems</i>, vol. 37, Neural Information Processing Systems Foundation, 2024.
  short: A. Kori, F. Locatello, A. Santhirasekaram, F. Toni, B. Glocker, F. De Sousa
    Ribeiro, in:, 38th Conference on Neural Information Processing Systems, Neural
    Information Processing Systems Foundation, 2024.
conference:
  end_date: 2024-12-16
  location: Vancouver, Canada
  name: 'NeurIPS: Neural Information Processing Systems'
  start_date: 2024-12-16
date_created: 2025-02-05T08:36:22Z
date_published: 2024-12-01T00:00:00Z
date_updated: 2025-05-14T11:29:10Z
day: '01'
ddc:
- '000'
department:
- _id: FrLo
external_id:
  arxiv:
  - '2406.07141'
file:
- access_level: open_access
  checksum: d27b3c7102adc28e798fe41001f0b919
  content_type: application/pdf
  creator: dernst
  date_created: 2025-02-05T08:34:25Z
  date_updated: 2025-02-05T08:34:25Z
  file_id: '19008'
  file_name: 2024_NeurIPS_Kori.pdf
  file_size: 6943800
  relation: main_file
  success: 1
file_date_updated: 2025-02-05T08:34:25Z
has_accepted_license: '1'
intvolume: '        37'
language:
- iso: eng
month: '12'
oa: 1
oa_version: Published Version
publication: 38th Conference on Neural Information Processing Systems
publication_status: published
publisher: Neural Information Processing Systems Foundation
quality_controlled: '1'
scopus_import: '1'
status: public
title: Identifiable object-centric representation learning via probabilistic slot
  attention
type: conference
user_id: 2DF688A6-F248-11E8-B48F-1D18A9856A87
volume: 37
year: '2024'
...
---
OA_place: publisher
OA_type: hybrid
_id: '19028'
abstract:
- lang: eng
  text: The stochastic nature of modern Monte Carlo (MC) rendering methods inevitably
    produces noise in rendered images for a practical number of samples per pixel.
    The problem of denoising these images has been widely studied, with most recent
    methods relying on data-driven, pretrained neural networks. In contrast, in this
    paper we propose a statistical approach to the denoising problem, treating each
    pixel as a random variable and reasoning about its distribution. Considering a
    pixel of the noisy rendered image, we formulate fast pair-wise statistical tests—based
    on online estimators—to decide which of the nearby pixels to exclude from the
    denoising filter. We show that for symmetric pixel weights and normally distributed
    samples, the classical Welch t-test is optimal in terms of mean squared error.
    We then show how to extend this result to handle non-normal distributions, using
    more recent confidence-interval formulations in combination with the Box-Cox transformation.
    Our results show that our statistical denoising approach matches the performance
    of state-of-the-art neural image denoising without having to resort to any computation-intensive
    pretraining. Furthermore, our approach easily generalizes to other quantities
    besides pixel intensity, which we demonstrate by showing additional applications
    to Russian roulette path termination and multiple importance sampling.
acknowledgement: 'We would like to thank Lukas Lipp for fruitful discussions, Károly
  Zsolnai-Fehér and Jaroslav Křivánek for valuable contributions to early versions
  of this work, and Bernhard Kerbl for help with our CUDA implementation. Moreover,
  we would like to thank the creators of the scenes we have used: Wig42 for “Wooden
  Staircase” (Fig. 1), “Grey and White Room” (Fig. S6), and “Modern Living Room” (Fig.
  S8); nacimus for “Bathroom” (Fig. 3, S5); NovaZeeke for “Japanese Classroom” (Fig.
  4, 6); Beeple for “Zero-Day” (Fig. 8); Jay-Artist for “White Room” (Fig. S7); Mareck
  for “Contemporary Bathroom” (Fig. 2); Christian Freude for “Glass Caustics” (Fig.
  S10); and Benedikt Bitterli for “Veach Ajar” (Fig. 7, S2), “Veach MIS” (Fig. S4),
  and “Fur Ball” (Fig. S11). This work has received funding from the Vienna Science
  and Technology Fund (WWTF) project ICT22-028 (“Toward Optimal Path Guiding for Photorealistic
  Rendering”) and the Austrian Science Fund (FWF) project F 77 (SFB “Advanced Computational
  Design”).'
article_number: '68'
article_processing_charge: Yes (in subscription journal)
author:
- first_name: Hiroyuki
  full_name: Sakai, Hiroyuki
  last_name: Sakai
- first_name: Christian
  full_name: Freude, Christian
  last_name: Freude
- first_name: Thomas
  full_name: Auzinger, Thomas
  id: 4718F954-F248-11E8-B48F-1D18A9856A87
  last_name: Auzinger
  orcid: 0000-0002-1546-3265
- first_name: David
  full_name: Hahn, David
  id: 357A6A66-F248-11E8-B48F-1D18A9856A87
  last_name: Hahn
- first_name: Michael
  full_name: Wimmer, Michael
  last_name: Wimmer
citation:
  ama: 'Sakai H, Freude C, Auzinger T, Hahn D, Wimmer M. A statistical approach to
    Monte Carlo denoising. In: <i>Proceedings - SIGGRAPH Asia 2024 Conference Papers</i>.
    Association for Computing Machinery; 2024. doi:<a href="https://doi.org/10.1145/3680528.3687591">10.1145/3680528.3687591</a>'
  apa: 'Sakai, H., Freude, C., Auzinger, T., Hahn, D., &#38; Wimmer, M. (2024). A
    statistical approach to Monte Carlo denoising. In <i>Proceedings - SIGGRAPH Asia
    2024 Conference Papers</i>. Tokyo, Japan: Association for Computing Machinery.
    <a href="https://doi.org/10.1145/3680528.3687591">https://doi.org/10.1145/3680528.3687591</a>'
  chicago: Sakai, Hiroyuki, Christian Freude, Thomas Auzinger, David Hahn, and Michael
    Wimmer. “A Statistical Approach to Monte Carlo Denoising.” In <i>Proceedings -
    SIGGRAPH Asia 2024 Conference Papers</i>. Association for Computing Machinery,
    2024. <a href="https://doi.org/10.1145/3680528.3687591">https://doi.org/10.1145/3680528.3687591</a>.
  ieee: H. Sakai, C. Freude, T. Auzinger, D. Hahn, and M. Wimmer, “A statistical approach
    to Monte Carlo denoising,” in <i>Proceedings - SIGGRAPH Asia 2024 Conference Papers</i>,
    Tokyo, Japan, 2024.
  ista: 'Sakai H, Freude C, Auzinger T, Hahn D, Wimmer M. 2024. A statistical approach
    to Monte Carlo denoising. Proceedings - SIGGRAPH Asia 2024 Conference Papers.
    SA: SIGGRAPH Asia, 68.'
  mla: Sakai, Hiroyuki, et al. “A Statistical Approach to Monte Carlo Denoising.”
    <i>Proceedings - SIGGRAPH Asia 2024 Conference Papers</i>, 68, Association for
    Computing Machinery, 2024, doi:<a href="https://doi.org/10.1145/3680528.3687591">10.1145/3680528.3687591</a>.
  short: H. Sakai, C. Freude, T. Auzinger, D. Hahn, M. Wimmer, in:, Proceedings -
    SIGGRAPH Asia 2024 Conference Papers, Association for Computing Machinery, 2024.
conference:
  end_date: 2024-12-06
  location: Tokyo, Japan
  name: 'SA: SIGGRAPH Asia'
  start_date: 2024-12-03
date_created: 2025-02-16T23:02:34Z
date_published: 2024-12-03T00:00:00Z
date_updated: 2025-12-02T13:58:56Z
day: '03'
ddc:
- '000'
doi: 10.1145/3680528.3687591
external_id:
  isi:
  - '001441591200068'
file:
- access_level: open_access
  checksum: 89f63b9237224362ec33430af9152700
  content_type: application/pdf
  creator: dernst
  date_created: 2025-04-15T12:53:24Z
  date_updated: 2025-04-15T12:53:24Z
  file_id: '19563'
  file_name: 2024_SIGGRAPH_Sakai.pdf
  file_size: 14791980
  relation: main_file
  success: 1
file_date_updated: 2025-04-15T12:53:24Z
fulldoi: https://doi.org/10.1145/3680528.3687591
has_accepted_license: '1'
isi: 1
language:
- iso: eng
month: '12'
oa: 1
oa_version: Published Version
publication: Proceedings - SIGGRAPH Asia 2024 Conference Papers
publication_identifier:
  isbn:
  - '9798400711312'
publication_status: published
publisher: Association for Computing Machinery
quality_controlled: '1'
scopus_import: '1'
status: public
title: A statistical approach to Monte Carlo denoising
tmp:
  image: /images/cc_by.png
  legal_code_url: https://creativecommons.org/licenses/by/4.0/legalcode
  name: Creative Commons Attribution 4.0 International Public License (CC-BY 4.0)
  short: CC BY (4.0)
type: conference
user_id: 2DF688A6-F248-11E8-B48F-1D18A9856A87
year: '2024'
...
---
OA_place: repository
OA_type: green
_id: '19307'
abstract:
- lang: eng
  text: "This repository contains the data, scripts, SAM codes and files required
    to reproduce the results of the manuscript \"The Unreasonable Efficiency of Total
    Rain Evaporation Removal in Triggering Convective Self-Aggregation\" submitted
    to the Geophysical Research Letters (GRL).\r\n\r\nBrief description of project:
    This project aims to examine the impact of rain evaporation removal or reduction
    in the planetary boundary layer (PBL) on convective self aggregation (CSA). Non-rotating
    radiative-convective equilibrium (RCE) simulations were conducted with the System
    for Atmospheric Modeling (SAM) cloud resolving model. Rain evaporation in the
    lowest 1 km was progressively reduced and the effect on CSA was investigated.
    The physical processes underlying this type of aggregation (referred to in the
    manuscript as no-evaporation CSA, or NE-CSA) were analyzed and described. \r\nThe
    default SAM code base (version 6.10.8) can be downloaded from here: http://rossby.msrc.sunysb.edu/~marat/SAM.html"
article_processing_charge: No
author:
- first_name: Yi-Ling
  full_name: Hwong, Yi-Ling
  id: 1217aa61-4dd1-11ec-9ac3-f2ba3f17ee22
  last_name: Hwong
  orcid: 0000-0001-9281-3479
- first_name: Caroline J
  full_name: Muller, Caroline J
  id: f978ccb0-3f7f-11eb-b193-b0e2bd13182b
  last_name: Muller
  orcid: 0000-0001-5836-5350
citation:
  ama: Hwong Y-L, Muller CJ. Data - The unreasonable efficiency of total rain evaporation
    removal in triggering convective self-aggregation. 2024. doi:<a href="https://doi.org/10.5281/ZENODO.10687169">10.5281/ZENODO.10687169</a>
  apa: Hwong, Y.-L., &#38; Muller, C. J. (2024). Data - The unreasonable efficiency
    of total rain evaporation removal in triggering convective self-aggregation. Zenodo.
    <a href="https://doi.org/10.5281/ZENODO.10687169">https://doi.org/10.5281/ZENODO.10687169</a>
  chicago: Hwong, Yi-Ling, and Caroline J Muller. “Data - The Unreasonable Efficiency
    of Total Rain Evaporation Removal in Triggering Convective Self-Aggregation.”
    Zenodo, 2024. <a href="https://doi.org/10.5281/ZENODO.10687169">https://doi.org/10.5281/ZENODO.10687169</a>.
  ieee: Y.-L. Hwong and C. J. Muller, “Data - The unreasonable efficiency of total
    rain evaporation removal in triggering convective self-aggregation.” Zenodo, 2024.
  ista: Hwong Y-L, Muller CJ. 2024. Data - The unreasonable efficiency of total rain
    evaporation removal in triggering convective self-aggregation, Zenodo, <a href="https://doi.org/10.5281/ZENODO.10687169">10.5281/ZENODO.10687169</a>.
  mla: Hwong, Yi-Ling, and Caroline J. Muller. <i>Data - The Unreasonable Efficiency
    of Total Rain Evaporation Removal in Triggering Convective Self-Aggregation</i>.
    Zenodo, 2024, doi:<a href="https://doi.org/10.5281/ZENODO.10687169">10.5281/ZENODO.10687169</a>.
  short: Y.-L. Hwong, C.J. Muller, (2024).
corr_author: '1'
date_created: 2025-03-07T08:39:40Z
date_published: 2024-02-21T00:00:00Z
date_updated: 2025-09-04T13:16:39Z
day: '21'
ddc:
- '550'
department:
- _id: CaMu
doi: 10.5281/ZENODO.10687169
fulldoi: https://doi.org/10.5281/ZENODO.10687169
has_accepted_license: '1'
main_file_link:
- open_access: '1'
  url: https://doi.org/10.5281/zenodo.8369509
month: '02'
oa: 1
oa_version: Published Version
publisher: Zenodo
related_material:
  record:
  - id: '15186'
    relation: used_in_publication
    status: public
status: public
title: Data - The unreasonable efficiency of total rain evaporation removal in triggering
  convective self-aggregation
tmp:
  image: /images/cc_by.png
  legal_code_url: https://creativecommons.org/licenses/by/4.0/legalcode
  name: Creative Commons Attribution 4.0 International Public License (CC-BY 4.0)
  short: CC BY (4.0)
type: research_data_reference
user_id: 2DF688A6-F248-11E8-B48F-1D18A9856A87
year: '2024'
...
---
OA_place: publisher
OA_type: diamond
_id: '19408'
abstract:
- lang: eng
  text: 'Continual learning is a subfield of machine learning, which aims to allow
    machine learning models to continuously learn on new data, by accumulating knowledge
    without forgetting what was learned in the past. In this work, we take a step
    back, and ask: "Why should one care about continual learning in the first place?".
    We set the stage by examining recent continual learning papers published at four
    major machine learning conferences, and show that memory-constrained settings
    dominate the field. Then, we discuss five open problems in machine learning, and
    even though they might seem unrelated to continual learning at first sight, we
    show that continual learning will inevitably be part of their solution. These
    problems are model editing, personalization and specialization, on-device learning,
    faster (re-)training and reinforcement learning. Finally, by comparing the desiderata
    from these unsolved problems and the current assumptions in continual learning,
    we highlight and discuss four future directions for continual learning research.
    We hope that this work offers an interesting perspective on the future of continual
    learning, while displaying its potential value and the paths we have to pursue
    in order to make it successful. This work is the result of the many discussions
    the authors had at the Dagstuhl seminar on Deep Continual Learning, in March 2023.'
alternative_title:
- TMLR
article_processing_charge: No
article_type: original
arxiv: 1
author:
- first_name: Eli
  full_name: Verwimp, Eli
  last_name: Verwimp
- first_name: Rahaf
  full_name: Aljundi, Rahaf
  last_name: Aljundi
- first_name: Shai
  full_name: Ben-David, Shai
  last_name: Ben-David
- first_name: Matthias
  full_name: Bethge, Matthias
  last_name: Bethge
- first_name: Andrea
  full_name: Cossu, Andrea
  last_name: Cossu
- first_name: Alexander
  full_name: Gepperth, Alexander
  last_name: Gepperth
- first_name: Tyler L.
  full_name: Hayes, Tyler L.
  last_name: Hayes
- first_name: Eyke
  full_name: Hüllermeier, Eyke
  last_name: Hüllermeier
- first_name: Christopher
  full_name: Kanan, Christopher
  last_name: Kanan
- first_name: Dhireesha
  full_name: Kudithipudi, Dhireesha
  last_name: Kudithipudi
- first_name: Christoph
  full_name: Lampert, Christoph
  id: 40C20FD2-F248-11E8-B48F-1D18A9856A87
  last_name: Lampert
  orcid: 0000-0001-8622-7887
- first_name: Martin
  full_name: Mundt, Martin
  last_name: Mundt
- first_name: Razvan
  full_name: Pascanu, Razvan
  last_name: Pascanu
- first_name: Adrian
  full_name: Popescu, Adrian
  last_name: Popescu
- first_name: Andreas S.
  full_name: Tolias, Andreas S.
  last_name: Tolias
- first_name: Joost
  full_name: Van De Weijer, Joost
  last_name: Van De Weijer
- first_name: Bing
  full_name: Liu, Bing
  last_name: Liu
- first_name: Vincenzo
  full_name: Lomonaco, Vincenzo
  last_name: Lomonaco
- first_name: Tinne
  full_name: Tuytelaars, Tinne
  last_name: Tuytelaars
- first_name: Gido M.
  full_name: Van De Ven, Gido M.
  last_name: Van De Ven
citation:
  ama: 'Verwimp E, Aljundi R, Ben-David S, et al. Continual learning: Applications
    and the road forward. <i>Transactions on Machine Learning Research</i>. 2024;2024.'
  apa: 'Verwimp, E., Aljundi, R., Ben-David, S., Bethge, M., Cossu, A., Gepperth,
    A., … Van De Ven, G. M. (2024). Continual learning: Applications and the road
    forward. <i>Transactions on Machine Learning Research</i>. Transactions on Machine
    Learning Research.'
  chicago: 'Verwimp, Eli, Rahaf Aljundi, Shai Ben-David, Matthias Bethge, Andrea Cossu,
    Alexander Gepperth, Tyler L. Hayes, et al. “Continual Learning: Applications and
    the Road Forward.” <i>Transactions on Machine Learning Research</i>. Transactions
    on Machine Learning Research, 2024.'
  ieee: 'E. Verwimp <i>et al.</i>, “Continual learning: Applications and the road
    forward,” <i>Transactions on Machine Learning Research</i>, vol. 2024. Transactions
    on Machine Learning Research, 2024.'
  ista: 'Verwimp E, Aljundi R, Ben-David S, Bethge M, Cossu A, Gepperth A, Hayes TL,
    Hüllermeier E, Kanan C, Kudithipudi D, Lampert C, Mundt M, Pascanu R, Popescu
    A, Tolias AS, Van De Weijer J, Liu B, Lomonaco V, Tuytelaars T, Van De Ven GM.
    2024. Continual learning: Applications and the road forward. Transactions on Machine
    Learning Research. 2024.'
  mla: 'Verwimp, Eli, et al. “Continual Learning: Applications and the Road Forward.”
    <i>Transactions on Machine Learning Research</i>, vol. 2024, Transactions on Machine
    Learning Research, 2024.'
  short: E. Verwimp, R. Aljundi, S. Ben-David, M. Bethge, A. Cossu, A. Gepperth, T.L.
    Hayes, E. Hüllermeier, C. Kanan, D. Kudithipudi, C. Lampert, M. Mundt, R. Pascanu,
    A. Popescu, A.S. Tolias, J. Van De Weijer, B. Liu, V. Lomonaco, T. Tuytelaars,
    G.M. Van De Ven, Transactions on Machine Learning Research 2024 (2024).
date_created: 2025-03-16T23:01:25Z
date_published: 2024-04-12T00:00:00Z
date_updated: 2025-03-20T09:21:02Z
day: '12'
ddc:
- '000'
department:
- _id: ChLa
external_id:
  arxiv:
  - '2311.11908'
file:
- access_level: open_access
  checksum: 0714e12f7423cd098976ed9974561155
  content_type: application/pdf
  creator: dernst
  date_created: 2025-03-20T09:02:18Z
  date_updated: 2025-03-20T09:02:18Z
  file_id: '19426'
  file_name: 2024_TMLR_Verwimp.pdf
  file_size: 1367966
  relation: main_file
  success: 1
file_date_updated: 2025-03-20T09:02:18Z
has_accepted_license: '1'
intvolume: '      2024'
language:
- iso: eng
month: '04'
oa: 1
oa_version: Published Version
publication: Transactions on Machine Learning Research
publication_identifier:
  eissn:
  - 2835-8856
publication_status: published
publisher: Transactions on Machine Learning Research
quality_controlled: '1'
scopus_import: '1'
status: public
title: 'Continual learning: Applications and the road forward'
tmp:
  image: /images/cc_by.png
  legal_code_url: https://creativecommons.org/licenses/by/4.0/legalcode
  name: Creative Commons Attribution 4.0 International Public License (CC-BY 4.0)
  short: CC BY (4.0)
type: journal_article
user_id: 2DF688A6-F248-11E8-B48F-1D18A9856A87
volume: 2024
year: '2024'
...
---
OA_place: repository
OA_type: green
_id: '19486'
abstract:
- lang: eng
  text: Consider the family of elliptic curves En:y2=x3+n2, where n varies over positive
    cubefree integers. There is a rational 3-isogeny ϕ from En to E^n:y2=x3−27n2 and
    a dual isogeny ϕ^:E^n→En. We show that for almost all n, the rank of Selϕ(En)
    is 0, and the rank of Selϕ^(E^n) is determined by the number of prime factors
    of n that are congruent to 2mod3 and the congruence class of nmod9.
acknowledgement: The author would like to thank Peter Koymans and Carlo Pagano for
  helpful discussions.
article_processing_charge: No
article_type: original
arxiv: 1
author:
- first_name: Yik Tung
  full_name: Chan, Yik Tung
  id: c4c0afc8-9262-11ed-9231-d8b0bc743af1
  last_name: Chan
  orcid: 0000-0001-8467-4106
citation:
  ama: Chan S. The 3-isogeny selmer groups of the elliptic curves y2=x3+n2. <i>International
    Mathematics Research Notices</i>. 2024;2024(9):7571-7593. doi:<a href="https://doi.org/10.1093/imrn/rnad266">10.1093/imrn/rnad266</a>
  apa: Chan, S. (2024). The 3-isogeny selmer groups of the elliptic curves y2=x3+n2.
    <i>International Mathematics Research Notices</i>. Oxford University Press. <a
    href="https://doi.org/10.1093/imrn/rnad266">https://doi.org/10.1093/imrn/rnad266</a>
  chicago: Chan, Stephanie. “The 3-Isogeny Selmer Groups of the Elliptic Curves Y2=x3+n2.”
    <i>International Mathematics Research Notices</i>. Oxford University Press, 2024.
    <a href="https://doi.org/10.1093/imrn/rnad266">https://doi.org/10.1093/imrn/rnad266</a>.
  ieee: S. Chan, “The 3-isogeny selmer groups of the elliptic curves y2=x3+n2,” <i>International
    Mathematics Research Notices</i>, vol. 2024, no. 9. Oxford University Press, pp.
    7571–7593, 2024.
  ista: Chan S. 2024. The 3-isogeny selmer groups of the elliptic curves y2=x3+n2.
    International Mathematics Research Notices. 2024(9), 7571–7593.
  mla: Chan, Stephanie. “The 3-Isogeny Selmer Groups of the Elliptic Curves Y2=x3+n2.”
    <i>International Mathematics Research Notices</i>, vol. 2024, no. 9, Oxford University
    Press, 2024, pp. 7571–93, doi:<a href="https://doi.org/10.1093/imrn/rnad266">10.1093/imrn/rnad266</a>.
  short: S. Chan, International Mathematics Research Notices 2024 (2024) 7571–7593.
date_created: 2025-04-05T10:50:33Z
date_published: 2024-05-01T00:00:00Z
date_updated: 2025-07-10T11:51:44Z
day: '01'
doi: 10.1093/imrn/rnad266
extern: '1'
external_id:
  arxiv:
  - '2211.06062'
fulldoi: https://doi.org/10.1093/imrn/rnad266
intvolume: '      2024'
issue: '9'
language:
- iso: eng
main_file_link:
- open_access: '1'
  url: https://doi.org/10.48550/arXiv.2211.06062
month: '05'
oa: 1
oa_version: Preprint
page: 7571-7593
publication: International Mathematics Research Notices
publication_identifier:
  eissn:
  - 1687-0247
  issn:
  - 1073-7928
publication_status: published
publisher: Oxford University Press
quality_controlled: '1'
scopus_import: '1'
status: public
title: The 3-isogeny selmer groups of the elliptic curves y2=x3+n2
type: journal_article
user_id: 2DF688A6-F248-11E8-B48F-1D18A9856A87
volume: 2024
year: '2024'
...
---
OA_place: repository
OA_type: green
_id: '19510'
abstract:
- lang: eng
  text: "We propose a new variant of the Adam optimizer [Kingma and Ba, 2014] called\r\nMICROADAM
    that specifically minimizes memory overheads, while maintaining\r\ntheoretical
    convergence guarantees. We achieve this by compressing the gradient\r\ninformation
    before it is fed into the optimizer state, thereby reducing its memory\r\nfootprint
    significantly. We control the resulting compression error via a novel\r\ninstance
    of the classical error feedback mechanism from distributed optimization [Seide
    et al., 2014, Alistarh et al., 2018, Karimireddy et al., 2019] in which\r\nthe
    error correction information is itself compressed to allow for practical memory\r\ngains.
    We prove that the resulting approach maintains theoretical convergence\r\nguarantees
    competitive to those of AMSGrad, while providing good practical performance. Specifically,
    we show that MICROADAM can be implemented efficiently\r\non GPUs: on both million-scale
    (BERT) and billion-scale (LLaMA) models, MICROADAM provides practical convergence
    competitive to that of the uncompressed\r\nAdam baseline, with lower memory usage
    and similar running time. Our code is\r\navailable at https://github.com/IST-DASLab/MicroAdam."
acknowledged_ssus:
- _id: CampIT
acknowledgement: The authors thank Razvan Pascanu, Mahdi Nikdan and Soroush Tabesh
  for their valuable feedback, the IT department from Institute of Science and Technology
  Austria for the hardware support and Weights and Biases for the infrastructure to
  track all our experiments. Mher Safaryan has received funding from the European
  Union’s Horizon 2020 research and innovation program under the Marie Sklodowska-Curie
  grant agreement No 101034413.
alternative_title:
- Advances in Neural Information Processing Systems
article_processing_charge: No
arxiv: 1
author:
- first_name: Ionut-Vlad
  full_name: Modoranu, Ionut-Vlad
  id: 449f7a18-f128-11eb-9611-9b430c0c6333
  last_name: Modoranu
- first_name: Mher
  full_name: Safaryan, Mher
  id: dd546b39-0804-11ed-9c55-ef075c39778d
  last_name: Safaryan
- first_name: Grigory
  full_name: Malinovsky, Grigory
  last_name: Malinovsky
- first_name: Eldar
  full_name: Kurtic, Eldar
  id: 47beb3a5-07b5-11eb-9b87-b108ec578218
  last_name: Kurtic
- first_name: Thomas
  full_name: Robert, Thomas
  id: de632733-1457-11f0-ae22-b5914b8c1c41
  last_name: Robert
- first_name: Peter
  full_name: Richtárik, Peter
  last_name: Richtárik
- first_name: Dan-Adrian
  full_name: Alistarh, Dan-Adrian
  id: 4A899BFC-F248-11E8-B48F-1D18A9856A87
  last_name: Alistarh
  orcid: 0000-0003-3650-940X
citation:
  ama: 'Modoranu I-V, Safaryan M, Malinovsky G, et al. MICROADAM: Accurate adaptive
    optimization with low space overhead and provable convergence. In: <i>38th Conference
    on Neural Information Processing Systems</i>. Vol 37. Neural Information Processing
    Systems Foundation; 2024.'
  apa: 'Modoranu, I.-V., Safaryan, M., Malinovsky, G., Kurtic, E., Robert, T., Richtárik,
    P., &#38; Alistarh, D.-A. (2024). MICROADAM: Accurate adaptive optimization with
    low space overhead and provable convergence. In <i>38th Conference on Neural Information
    Processing Systems</i> (Vol. 37). Neural Information Processing Systems Foundation.'
  chicago: 'Modoranu, Ionut-Vlad, Mher Safaryan, Grigory Malinovsky, Eldar Kurtic,
    Thomas Robert, Peter Richtárik, and Dan-Adrian Alistarh. “MICROADAM: Accurate
    Adaptive Optimization with Low Space Overhead and Provable Convergence.” In <i>38th
    Conference on Neural Information Processing Systems</i>, Vol. 37. Neural Information
    Processing Systems Foundation, 2024.'
  ieee: 'I.-V. Modoranu <i>et al.</i>, “MICROADAM: Accurate adaptive optimization
    with low space overhead and provable convergence,” in <i>38th Conference on Neural
    Information Processing Systems</i>, 2024, vol. 37.'
  ista: 'Modoranu I-V, Safaryan M, Malinovsky G, Kurtic E, Robert T, Richtárik P,
    Alistarh D-A. 2024. MICROADAM: Accurate adaptive optimization with low space overhead
    and provable convergence. 38th Conference on Neural Information Processing Systems.
    , Advances in Neural Information Processing Systems, vol. 37.'
  mla: 'Modoranu, Ionut-Vlad, et al. “MICROADAM: Accurate Adaptive Optimization with
    Low Space Overhead and Provable Convergence.” <i>38th Conference on Neural Information
    Processing Systems</i>, vol. 37, Neural Information Processing Systems Foundation,
    2024.'
  short: I.-V. Modoranu, M. Safaryan, G. Malinovsky, E. Kurtic, T. Robert, P. Richtárik,
    D.-A. Alistarh, in:, 38th Conference on Neural Information Processing Systems,
    Neural Information Processing Systems Foundation, 2024.
corr_author: '1'
date_created: 2025-04-06T22:01:32Z
date_published: 2024-12-20T00:00:00Z
date_updated: 2025-05-14T11:32:52Z
day: '20'
department:
- _id: DaAl
ec_funded: 1
external_id:
  arxiv:
  - '2405.15593'
intvolume: '        37'
language:
- iso: eng
main_file_link:
- open_access: '1'
  url: https://doi.org/10.48550/arXiv.2405.15593
month: '12'
oa: 1
oa_version: Preprint
project:
- _id: fc2ed2f7-9c52-11eb-aca3-c01059dda49c
  call_identifier: H2020
  grant_number: '101034413'
  name: 'IST-BRIDGE: International postdoctoral program'
publication: 38th Conference on Neural Information Processing Systems
publication_identifier:
  issn:
  - 1049-5258
publication_status: published
publisher: Neural Information Processing Systems Foundation
quality_controlled: '1'
related_material:
  link:
  - relation: software
    url: https://github.com/IST-DASLab/MicroAdam
scopus_import: '1'
status: public
title: 'MICROADAM: Accurate adaptive optimization with low space overhead and provable
  convergence'
type: conference
user_id: 2DF688A6-F248-11E8-B48F-1D18A9856A87
volume: 37
year: '2024'
...
---
OA_place: repository
OA_type: green
_id: '19511'
abstract:
- lang: eng
  text: We introduce QuaRot, a new Quantization scheme based on Rotations, which is
    able to quantize LLMs end-to-end, including all weights, activations, and KV cache
    in 4 bits. QuaRot rotates LLMs in a way that removes outliers from the hidden
    state without changing the output, making quantization easier. This computational
    invariance is applied to the hidden state (residual) of the LLM, as well as to
    the activations of the feed-forward components, aspects of the attention mechanism,
    and to the KV cache. The result is a quantized model where all matrix multiplications
    are performed in 4 bits, without any channels identified for retention in higher
    precision. Our 4-bit quantized LLAMA2-70B model has losses of at most 0.47 WikiText-2
    perplexity and retains 99% of the zero-shot performance. We also show that QuaRot
    can provide lossless 6 and 8 bit LLAMA-2 models without any calibration data using
    round-to-nearest quantization. Code is available at github.com/spcl/QuaRot.
alternative_title:
- Advances in Neural Information Processing Systems
article_processing_charge: No
arxiv: 1
author:
- first_name: Saleh
  full_name: Ashkboos, Saleh
  last_name: Ashkboos
- first_name: Amirkeivan
  full_name: Mohtashami, Amirkeivan
  last_name: Mohtashami
- first_name: Maximilian L.
  full_name: Croci, Maximilian L.
  last_name: Croci
- first_name: Bo
  full_name: Li, Bo
  last_name: Li
- first_name: Pashmina
  full_name: Cameron, Pashmina
  last_name: Cameron
- first_name: Martin
  full_name: Jaggi, Martin
  last_name: Jaggi
- first_name: Dan-Adrian
  full_name: Alistarh, Dan-Adrian
  id: 4A899BFC-F248-11E8-B48F-1D18A9856A87
  last_name: Alistarh
  orcid: 0000-0003-3650-940X
- first_name: Torsten
  full_name: Hoefler, Torsten
  last_name: Hoefler
- first_name: James
  full_name: Hensman, James
  last_name: Hensman
citation:
  ama: 'Ashkboos S, Mohtashami A, Croci ML, et al. QuaRot: Outlier-free 4-bit inference
    in rotated LLMs. In: <i>38th Conference on Neural Information Processing Systems</i>.
    Vol 37. Neural Information Processing Systems Foundation; 2024.'
  apa: 'Ashkboos, S., Mohtashami, A., Croci, M. L., Li, B., Cameron, P., Jaggi, M.,
    … Hensman, J. (2024). QuaRot: Outlier-free 4-bit inference in rotated LLMs. In
    <i>38th Conference on Neural Information Processing Systems</i> (Vol. 37). Vancouver,
    Canada: Neural Information Processing Systems Foundation.'
  chicago: 'Ashkboos, Saleh, Amirkeivan Mohtashami, Maximilian L. Croci, Bo Li, Pashmina
    Cameron, Martin Jaggi, Dan-Adrian Alistarh, Torsten Hoefler, and James Hensman.
    “QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs.” In <i>38th Conference
    on Neural Information Processing Systems</i>, Vol. 37. Neural Information Processing
    Systems Foundation, 2024.'
  ieee: 'S. Ashkboos <i>et al.</i>, “QuaRot: Outlier-free 4-bit inference in rotated
    LLMs,” in <i>38th Conference on Neural Information Processing Systems</i>, Vancouver,
    Canada, 2024, vol. 37.'
  ista: 'Ashkboos S, Mohtashami A, Croci ML, Li B, Cameron P, Jaggi M, Alistarh D-A,
    Hoefler T, Hensman J. 2024. QuaRot: Outlier-free 4-bit inference in rotated LLMs.
    38th Conference on Neural Information Processing Systems. NeurIPS: Neural Information
    Processing Systems, Advances in Neural Information Processing Systems, vol. 37.'
  mla: 'Ashkboos, Saleh, et al. “QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs.”
    <i>38th Conference on Neural Information Processing Systems</i>, vol. 37, Neural
    Information Processing Systems Foundation, 2024.'
  short: S. Ashkboos, A. Mohtashami, M.L. Croci, B. Li, P. Cameron, M. Jaggi, D.-A.
    Alistarh, T. Hoefler, J. Hensman, in:, 38th Conference on Neural Information Processing
    Systems, Neural Information Processing Systems Foundation, 2024.
conference:
  end_date: 2024-12-15
  location: Vancouver, Canada
  name: 'NeurIPS: Neural Information Processing Systems'
  start_date: 2024-12-09
date_created: 2025-04-06T22:01:32Z
date_published: 2024-12-20T00:00:00Z
date_updated: 2025-05-14T11:33:12Z
day: '20'
department:
- _id: DaAl
external_id:
  arxiv:
  - '2404.00456'
intvolume: '        37'
language:
- iso: eng
main_file_link:
- open_access: '1'
  url: https://doi.org/10.48550/arXiv.2404.00456
month: '12'
oa: 1
oa_version: Preprint
publication: 38th Conference on Neural Information Processing Systems
publication_identifier:
  issn:
  - 1049-5258
publication_status: published
publisher: Neural Information Processing Systems Foundation
quality_controlled: '1'
related_material:
  link:
  - relation: software
    url: https://github.com/spcl/QuaRot
scopus_import: '1'
status: public
title: 'QuaRot: Outlier-free 4-bit inference in rotated LLMs'
type: conference
user_id: 2DF688A6-F248-11E8-B48F-1D18A9856A87
volume: 37
year: '2024'
...
---
OA_place: repository
OA_type: green
_id: '19512'
abstract:
- lang: eng
  text: "Differential privacy with gradual expiration models the setting where data
    items\r\narrive in a stream and at a given time t the privacy loss guaranteed
    for a data item\r\nseen at time (t − d) is εg(d), where g is a monotonically non-decreasing
    function.\r\nWe study the fundamental continual (binary) counting problem where
    each data\r\nitem consists of a bit, and the algorithm needs to output at each
    time step the sum of\r\nall the bits streamed so far. For a stream of length T
    and privacy without expiration\r\ncontinual counting is possible with maximum
    (over all time steps) additive error\r\nO(log2\r\n(T)/ε) and the best known lower
    bound is Ω(log(T)/ε); closing this gap\r\nis a challenging open problem.\r\nWe
    show that the situation is very different for privacy with gradual expiration
    by\r\ngiving upper and lower bounds for a large set of expiration functions g.
    Specifically,\r\nour algorithm achieves an additive error of O(log(T)/ε) for a
    large set of privacy\r\nexpiration functions. We also give a lower bound that
    shows that if C is the additive\r\nerror of any ε-DP algorithm for this problem,
    then the product of C and the privacy\r\nexpiration function after 2C steps must
    be Ω(log(T)/ε). Our algorithm matches\r\nthis lower bound as its additive error
    is O(log(T)/ε), even when g(2C) = O(1).\r\nOur empirical evaluation shows that
    we achieve a slowly growing privacy loss\r\nwith significantly smaller empirical
    privacy loss for large values of d than a natural\r\nbaseline algorithm."
acknowledgement: 'Monika Henzinger: This project has received funding from the European
  Research Council (ERC) under the European Union’s Horizon 2020 research and innovation
  programme (Grant agreement No. 101019564) and the Austrian Science Fund (FWF) grant
  DOI 10.55776/Z422, grant DOI 10.55776/I5982, and grant DOI 10.55776/P33775 with
  additional funding from the netidee SCIENCE Stiftung, 2020–2024. Joel Daniel Andersson
  and Rasmus Pagh are affiliated with Basic Algorithms Research Copenhagen (BARC),
  supported by the VILLUM Foundation grant 16582, and are also supported by Providentia,
  a Data Science Distinguished Investigator grant from Novo Nordisk Fonden. Teresa
  Anna Steiner is supported by a research grant (VIL51463) from VILLUM FONDEN. This
  work was done while Teresa Anna Steiner was a Postdoc at the Technical University
  of Denmark. Jalaj Upadhyay’s research was funded by the Rutgers Decanal Grant no.
  302918 and an unrestricted gift from Google.'
alternative_title:
- Advances in Neural Information Processing Systems
article_processing_charge: No
arxiv: 1
author:
- first_name: Joel Daniel
  full_name: Andersson, Joel Daniel
  last_name: Andersson
- first_name: Monika H
  full_name: Henzinger, Monika H
  id: 540c9bbd-f2de-11ec-812d-d04a5be85630
  last_name: Henzinger
  orcid: 0000-0002-5008-6530
- first_name: Rasmus
  full_name: Pagh, Rasmus
  last_name: Pagh
- first_name: Teresa Anna
  full_name: Steiner, Teresa Anna
  last_name: Steiner
- first_name: Jalaj
  full_name: Upadhyay, Jalaj
  last_name: Upadhyay
citation:
  ama: 'Andersson JD, Henzinger M, Pagh R, Steiner TA, Upadhyay J. Continual counting
    with gradual privacy expiration. In: <i>38th Conference on Neural Information
    Processing Systems</i>. Vol 37. Neural Information Processing Systems Foundation;
    2024.'
  apa: 'Andersson, J. D., Henzinger, M., Pagh, R., Steiner, T. A., &#38; Upadhyay,
    J. (2024). Continual counting with gradual privacy expiration. In <i>38th Conference
    on Neural Information Processing Systems</i> (Vol. 37). Vancouver, Canada: Neural
    Information Processing Systems Foundation.'
  chicago: Andersson, Joel Daniel, Monika Henzinger, Rasmus Pagh, Teresa Anna Steiner,
    and Jalaj Upadhyay. “Continual Counting with Gradual Privacy Expiration.” In <i>38th
    Conference on Neural Information Processing Systems</i>, Vol. 37. Neural Information
    Processing Systems Foundation, 2024.
  ieee: J. D. Andersson, M. Henzinger, R. Pagh, T. A. Steiner, and J. Upadhyay, “Continual
    counting with gradual privacy expiration,” in <i>38th Conference on Neural Information
    Processing Systems</i>, Vancouver, Canada, 2024, vol. 37.
  ista: 'Andersson JD, Henzinger M, Pagh R, Steiner TA, Upadhyay J. 2024. Continual
    counting with gradual privacy expiration. 38th Conference on Neural Information
    Processing Systems. NeurIPS: Neural Information Processing Systems, Advances in
    Neural Information Processing Systems, vol. 37.'
  mla: Andersson, Joel Daniel, et al. “Continual Counting with Gradual Privacy Expiration.”
    <i>38th Conference on Neural Information Processing Systems</i>, vol. 37, Neural
    Information Processing Systems Foundation, 2024.
  short: J.D. Andersson, M. Henzinger, R. Pagh, T.A. Steiner, J. Upadhyay, in:, 38th
    Conference on Neural Information Processing Systems, Neural Information Processing
    Systems Foundation, 2024.
conference:
  end_date: 2024-12-15
  location: Vancouver, Canada
  name: 'NeurIPS: Neural Information Processing Systems'
  start_date: 2024-12-09
corr_author: '1'
date_created: 2025-04-06T22:01:32Z
date_published: 2024-12-20T00:00:00Z
date_updated: 2025-05-14T11:33:22Z
day: '20'
department:
- _id: MoHe
ec_funded: 1
external_id:
  arxiv:
  - '2406.03802'
intvolume: '        37'
language:
- iso: eng
main_file_link:
- open_access: '1'
  url: https://doi.org/10.48550/arXiv.2406.03802
month: '12'
oa: 1
oa_version: Preprint
project:
- _id: bd9ca328-d553-11ed-ba76-dc4f890cfe62
  call_identifier: H2020
  grant_number: '101019564'
  name: The design and evaluation of modern fully dynamic data structures
- _id: 34def286-11ca-11ed-8bc3-da5948e1613c
  grant_number: Z00422
  name: Efficient algorithms
- _id: bda196b2-d553-11ed-ba76-8e8ee6c21103
  grant_number: I05982
  name: Static and Dynamic Hierarchical Graph Decompositions
- _id: bd9e3a2e-d553-11ed-ba76-8aa684ce17fe
  grant_number: P33775
  name: Fast Algorithms for a Reactive Network Layer
publication: 38th Conference on Neural Information Processing Systems
publication_identifier:
  issn:
  - 1049-5258
publication_status: published
publisher: Neural Information Processing Systems Foundation
quality_controlled: '1'
scopus_import: '1'
status: public
title: Continual counting with gradual privacy expiration
type: conference
user_id: 2DF688A6-F248-11E8-B48F-1D18A9856A87
volume: 37
year: '2024'
...
---
OA_place: repository
OA_type: green
_id: '19515'
abstract:
- lang: eng
  text: "Neural models learn data representations that lie on low-dimensional manifolds,\r\nyet
    modeling the relation between these representational spaces is an ongoing challenge.
    By integrating spectral geometry principles into neural modeling, we show\r\nthat
    this problem can be better addressed in the functional domain, mitigating complexity,
    while enhancing interpretability and performances on downstream tasks.\r\nTo this
    end, we introduce a multi-purpose framework to the representation learning\r\ncommunity,
    which allows to: (i) compare different spaces in an interpretable way\r\nand measure
    their intrinsic similarity; (ii) find correspondences between them, both\r\nin
    unsupervised and weakly supervised settings, and (iii) to effectively transfer\r\nrepresentations
    between distinct spaces. We validate our framework on various\r\napplications,
    ranging from stitching to retrieval tasks, and on multiple modalities,\r\ndemonstrating
    that Latent Functional Maps can serve as a swiss-army knife for\r\nrepresentation
    alignment"
acknowledgement: MF is supported by the MSCA IST-Bridge fellowship which has received
  funding from the European Union’s Horizon 2020 research and innovation program under
  the Marie Skłodowska-Curie grant agreement No 101034413. ER and VM are supported
  by the PNRR MUR project PE0000013-FAIR. MP is supported by the Sapienza grant "Predicting
  and Explaining Clinical Trial Outcomes", prot. RG12218166FA3F13.
alternative_title:
- Advances in Neural Information Processing Systems
article_processing_charge: No
arxiv: 1
author:
- first_name: Marco
  full_name: Fumero, Marco
  id: 1c1593eb-393f-11ef-bb8e-ab4f1e979650
  last_name: Fumero
- first_name: Marco
  full_name: Pegoraro, Marco
  last_name: Pegoraro
- first_name: Valentino
  full_name: Maiorca, Valentino
  last_name: Maiorca
- first_name: Francesco
  full_name: Locatello, Francesco
  id: 26cfd52f-2483-11ee-8040-88983bcc06d4
  last_name: Locatello
  orcid: 0000-0002-4850-0683
- first_name: Emanuele
  full_name: Rodolà, Emanuele
  last_name: Rodolà
citation:
  ama: 'Fumero M, Pegoraro M, Maiorca V, Locatello F, Rodolà E. Latent functional
    maps: A spectral framework for representation alignment. In: <i>38th Conference
    on Neural Information Processing Systems</i>. Vol 37. Neural Information Processing
    Systems Foundation; 2024.'
  apa: 'Fumero, M., Pegoraro, M., Maiorca, V., Locatello, F., &#38; Rodolà, E. (2024).
    Latent functional maps: A spectral framework for representation alignment. In
    <i>38th Conference on Neural Information Processing Systems</i> (Vol. 37). Vancouver,
    Canada: Neural Information Processing Systems Foundation.'
  chicago: 'Fumero, Marco, Marco Pegoraro, Valentino Maiorca, Francesco Locatello,
    and Emanuele Rodolà. “Latent Functional Maps: A Spectral Framework for Representation
    Alignment.” In <i>38th Conference on Neural Information Processing Systems</i>,
    Vol. 37. Neural Information Processing Systems Foundation, 2024.'
  ieee: 'M. Fumero, M. Pegoraro, V. Maiorca, F. Locatello, and E. Rodolà, “Latent
    functional maps: A spectral framework for representation alignment,” in <i>38th
    Conference on Neural Information Processing Systems</i>, Vancouver, Canada, 2024,
    vol. 37.'
  ista: 'Fumero M, Pegoraro M, Maiorca V, Locatello F, Rodolà E. 2024. Latent functional
    maps: A spectral framework for representation alignment. 38th Conference on Neural
    Information Processing Systems. NeurIPS: Neural Information Processing Systems,
    Advances in Neural Information Processing Systems, vol. 37.'
  mla: 'Fumero, Marco, et al. “Latent Functional Maps: A Spectral Framework for Representation
    Alignment.” <i>38th Conference on Neural Information Processing Systems</i>, vol.
    37, Neural Information Processing Systems Foundation, 2024.'
  short: M. Fumero, M. Pegoraro, V. Maiorca, F. Locatello, E. Rodolà, in:, 38th Conference
    on Neural Information Processing Systems, Neural Information Processing Systems
    Foundation, 2024.
conference:
  end_date: 2024-12-15
  location: Vancouver, Canada
  name: 'NeurIPS: Neural Information Processing Systems'
  start_date: 2024-12-09
corr_author: '1'
date_created: 2025-04-06T22:01:32Z
date_published: 2024-12-20T00:00:00Z
date_updated: 2025-05-14T11:36:51Z
day: '20'
department:
- _id: FrLo
ec_funded: 1
external_id:
  arxiv:
  - '2406.14183'
intvolume: '        37'
language:
- iso: eng
main_file_link:
- open_access: '1'
  url: https://doi.org/10.48550/arXiv.2406.14183
month: '12'
oa: 1
oa_version: Preprint
project:
- _id: fc2ed2f7-9c52-11eb-aca3-c01059dda49c
  call_identifier: H2020
  grant_number: '101034413'
  name: 'IST-BRIDGE: International postdoctoral program'
publication: 38th Conference on Neural Information Processing Systems
publication_identifier:
  issn:
  - 1049-5258
publication_status: published
publisher: Neural Information Processing Systems Foundation
quality_controlled: '1'
scopus_import: '1'
status: public
title: 'Latent functional maps: A spectral framework for representation alignment'
type: conference
user_id: 2DF688A6-F248-11E8-B48F-1D18A9856A87
volume: 37
year: '2024'
...
---
OA_place: repository
OA_type: green
_id: '19517'
abstract:
- lang: eng
  text: "In this paper, we present a novel data-free method for merging neural networks
    in weight space. Differently from most existing works, our method optimizes for
    the permutations of network neurons globally across all layers. This allows us
    to enforce cycle consistency of the permutations when merging n ≥ 3 models, allowing
    circular compositions of permutations to be computed without accumulating error
    along the path. We qualitatively and quantitatively motivate the need for such
    a constraint, showing its benefits when merging sets of models in scenarios spanning
    varying architectures and datasets. We finally show that, when coupled\r\nwith
    activation renormalization, our approach yields the best results in the task."
acknowledgement: "This work is supported by the ERC grant no.802554 (SPECGEO), PRIN
  2020 project\r\nno.2020TA3K9N (LEGO.AI), and PNRR MUR project PE0000013-FAIR. Marco
  Fumero is supported by the MSCA IST-Bridge fellowship which has received funding
  from the European Union’s Horizon 2020 research and innovation program under the
  Marie Skłodowska-Curie grant agreement No 101034413. We thank Simone Scardapane
  for the helpful feedback on the paper."
alternative_title:
- Advances in Neural Information Processing Systems
article_processing_charge: No
arxiv: 1
author:
- first_name: Donato
  full_name: Crisostomi, Donato
  last_name: Crisostomi
- first_name: Marco
  full_name: Fumero, Marco
  id: 1c1593eb-393f-11ef-bb8e-ab4f1e979650
  last_name: Fumero
- first_name: Daniele
  full_name: Baieri, Daniele
  last_name: Baieri
- first_name: Florian
  full_name: Bernard, Florian
  last_name: Bernard
- first_name: Emanuele
  full_name: Rodolà, Emanuele
  last_name: Rodolà
citation:
  ama: 'Crisostomi D, Fumero M, Baieri D, Bernard F, Rodolà E. C2M3: Cycle-consistent
    multi-model merging. In: <i>38th Conference on Neural Information Processing Systems</i>.
    Vol 37. Neural Information Processing Systems Foundation; 2024.'
  apa: 'Crisostomi, D., Fumero, M., Baieri, D., Bernard, F., &#38; Rodolà, E. (2024).
    C2M3: Cycle-consistent multi-model merging. In <i>38th Conference on Neural Information
    Processing Systems</i> (Vol. 37). Vancouver, Canada: Neural Information Processing
    Systems Foundation.'
  chicago: 'Crisostomi, Donato, Marco Fumero, Daniele Baieri, Florian Bernard, and
    Emanuele Rodolà. “C2M3: Cycle-Consistent Multi-Model Merging.” In <i>38th Conference
    on Neural Information Processing Systems</i>, Vol. 37. Neural Information Processing
    Systems Foundation, 2024.'
  ieee: 'D. Crisostomi, M. Fumero, D. Baieri, F. Bernard, and E. Rodolà, “C2M3: Cycle-consistent
    multi-model merging,” in <i>38th Conference on Neural Information Processing Systems</i>,
    Vancouver, Canada, 2024, vol. 37.'
  ista: 'Crisostomi D, Fumero M, Baieri D, Bernard F, Rodolà E. 2024. C2M3: Cycle-consistent
    multi-model merging. 38th Conference on Neural Information Processing Systems.
    NeurIPS: Neural Information Processing Systems, Advances in Neural Information
    Processing Systems, vol. 37.'
  mla: 'Crisostomi, Donato, et al. “C2M3: Cycle-Consistent Multi-Model Merging.” <i>38th
    Conference on Neural Information Processing Systems</i>, vol. 37, Neural Information
    Processing Systems Foundation, 2024.'
  short: D. Crisostomi, M. Fumero, D. Baieri, F. Bernard, E. Rodolà, in:, 38th Conference
    on Neural Information Processing Systems, Neural Information Processing Systems
    Foundation, 2024.
conference:
  end_date: 2024-12-15
  location: Vancouver, Canada
  name: 'NeurIPS: Neural Information Processing Systems'
  start_date: 2024-12-09
corr_author: '1'
date_created: 2025-04-06T22:01:32Z
date_published: 2024-12-20T00:00:00Z
date_updated: 2025-05-14T11:36:59Z
day: '20'
department:
- _id: FrLo
ec_funded: 1
external_id:
  arxiv:
  - '2405.17897'
intvolume: '        37'
language:
- iso: eng
main_file_link:
- open_access: '1'
  url: https://doi.org/10.48550/arXiv.2405.17897
month: '12'
oa: 1
oa_version: Preprint
project:
- _id: fc2ed2f7-9c52-11eb-aca3-c01059dda49c
  call_identifier: H2020
  grant_number: '101034413'
  name: 'IST-BRIDGE: International postdoctoral program'
publication: 38th Conference on Neural Information Processing Systems
publication_identifier:
  issn:
  - 1049-5258
publication_status: published
publisher: Neural Information Processing Systems Foundation
quality_controlled: '1'
scopus_import: '1'
status: public
title: 'C2M3: Cycle-consistent multi-model merging'
type: conference
user_id: 2DF688A6-F248-11E8-B48F-1D18A9856A87
volume: 37
year: '2024'
...
---
OA_place: repository
OA_type: green
_id: '19518'
abstract:
- lang: eng
  text: "The rising footprint of machine learning has led to a focus on imposing model\r\nsparsity
    as a means of reducing computational and memory costs. For deep neural\r\nnetworks
    (DNNs), the state-of-the-art accuracy-vs-sparsity is achieved by heuristics\r\ninspired
    by the classical Optimal Brain Surgeon (OBS) framework [LeCun et al.,\r\n1989,
    Hassibi and Stork, 1992, Hassibi et al., 1993], which leverages loss curvature\r\ninformation
    to make better pruning decisions. Yet, these results still lack a solid\r\ntheoretical
    understanding, and it is unclear whether they can be improved by\r\nleveraging
    connections to the wealth of work on sparse recovery algorithms. In this\r\npaper,
    we draw new connections between these two areas and present new sparse\r\nrecovery
    algorithms inspired by the OBS framework that comes with theoretical\r\nguarantees
    under reasonable assumptions and have strong practical performance.\r\nSpecifically,
    our work starts from the observation that we can leverage curvature\r\ninformation
    in OBS-like fashion upon the projection step of classic iterative sparse\r\nrecovery
    algorithms such as IHT. We show for the first time that this leads both\r\nto
    improved convergence bounds under standard assumptions. Furthermore, we\r\npresent
    extensions of this approach to the practical task of obtaining accurate sparse\r\nDNNs,
    and validate it experimentally at scale for Transformer-based models on\r\nvision
    and language tasks."
acknowledged_ssus:
- _id: CampIT
acknowledgement: The authors thank the anonymous NeurIPS reviewers for their useful
  comments and feedback, the IT department from the Institute of Science and Technology
  Austria for the hardware support, and Weights and Biases for the infrastructure
  to track all our experiments. Mher Safaryan has received funding from the European
  Union’s Horizon 2020 research and innovation program under the Maria Skłodowska-Curie
  grant agreement No 101034413.
alternative_title:
- Advances in Neural Information Processing Systems
article_processing_charge: No
arxiv: 1
author:
- first_name: Diyuan
  full_name: Wu, Diyuan
  id: 1a5914c2-896a-11ed-bdf8-fb80621a0635
  last_name: Wu
- first_name: Ionut-Vlad
  full_name: Modoranu, Ionut-Vlad
  id: 449f7a18-f128-11eb-9611-9b430c0c6333
  last_name: Modoranu
- first_name: Mher
  full_name: Safaryan, Mher
  id: dd546b39-0804-11ed-9c55-ef075c39778d
  last_name: Safaryan
- first_name: Denis
  full_name: Kuznedelev, Denis
  last_name: Kuznedelev
- first_name: Dan-Adrian
  full_name: Alistarh, Dan-Adrian
  id: 4A899BFC-F248-11E8-B48F-1D18A9856A87
  last_name: Alistarh
  orcid: 0000-0003-3650-940X
citation:
  ama: 'Wu D, Modoranu I-V, Safaryan M, Kuznedelev D, Alistarh D-A. The iterative
    optimal brain surgeon: Faster sparse recovery by leveraging second-order information.
    In: <i>38th Conference on Neural Information Processing Systems</i>. Vol 37. Neural
    Information Processing Systems Foundation; 2024.'
  apa: 'Wu, D., Modoranu, I.-V., Safaryan, M., Kuznedelev, D., &#38; Alistarh, D.-A.
    (2024). The iterative optimal brain surgeon: Faster sparse recovery by leveraging
    second-order information. In <i>38th Conference on Neural Information Processing
    Systems</i> (Vol. 37). Vancouver, Canada: Neural Information Processing Systems
    Foundation.'
  chicago: 'Wu, Diyuan, Ionut-Vlad Modoranu, Mher Safaryan, Denis Kuznedelev, and
    Dan-Adrian Alistarh. “The Iterative Optimal Brain Surgeon: Faster Sparse Recovery
    by Leveraging Second-Order Information.” In <i>38th Conference on Neural Information
    Processing Systems</i>, Vol. 37. Neural Information Processing Systems Foundation,
    2024.'
  ieee: 'D. Wu, I.-V. Modoranu, M. Safaryan, D. Kuznedelev, and D.-A. Alistarh, “The
    iterative optimal brain surgeon: Faster sparse recovery by leveraging second-order
    information,” in <i>38th Conference on Neural Information Processing Systems</i>,
    Vancouver, Canada, 2024, vol. 37.'
  ista: 'Wu D, Modoranu I-V, Safaryan M, Kuznedelev D, Alistarh D-A. 2024. The iterative
    optimal brain surgeon: Faster sparse recovery by leveraging second-order information.
    38th Conference on Neural Information Processing Systems. NeurIPS: Neural Information
    Processing Systems, Advances in Neural Information Processing Systems, vol. 37.'
  mla: 'Wu, Diyuan, et al. “The Iterative Optimal Brain Surgeon: Faster Sparse Recovery
    by Leveraging Second-Order Information.” <i>38th Conference on Neural Information
    Processing Systems</i>, vol. 37, Neural Information Processing Systems Foundation,
    2024.'
  short: D. Wu, I.-V. Modoranu, M. Safaryan, D. Kuznedelev, D.-A. Alistarh, in:, 38th
    Conference on Neural Information Processing Systems, Neural Information Processing
    Systems Foundation, 2024.
conference:
  end_date: 2024-12-15
  location: Vancouver, Canada
  name: 'NeurIPS: Neural Information Processing Systems'
  start_date: 2024-12-09
corr_author: '1'
date_created: 2025-04-06T22:01:32Z
date_published: 2024-12-20T00:00:00Z
date_updated: 2025-05-14T11:37:10Z
day: '20'
department:
- _id: DaAl
- _id: MaMo
ec_funded: 1
external_id:
  arxiv:
  - '2408.17163'
intvolume: '        37'
language:
- iso: eng
main_file_link:
- open_access: '1'
  url: https://doi.org/10.48550/arXiv.2408.17163
month: '12'
oa: 1
oa_version: Preprint
project:
- _id: fc2ed2f7-9c52-11eb-aca3-c01059dda49c
  call_identifier: H2020
  grant_number: '101034413'
  name: 'IST-BRIDGE: International postdoctoral program'
publication: 38th Conference on Neural Information Processing Systems
publication_identifier:
  issn:
  - 1049-5258
publication_status: published
publisher: Neural Information Processing Systems Foundation
quality_controlled: '1'
scopus_import: '1'
status: public
title: 'The iterative optimal brain surgeon: Faster sparse recovery by leveraging
  second-order information'
type: conference
user_id: 2DF688A6-F248-11E8-B48F-1D18A9856A87
volume: 37
year: '2024'
...
---
OA_place: publisher
OA_type: gold
_id: '19519'
abstract:
- lang: eng
  text: There has been significant interest in "extreme" compression of large language
    models (LLMs), i.e. to 1-2 bits per parameter, which allows such models to be
    executed efficiently on resource-constrained devices. Existing work focused on
    improved one-shot quantization techniques and weight representations; yet, purely
    post-training approaches are reaching diminishing returns in terms of the accuracy-vs-bit-width
    trade-off. State-of-the-art quantization methods such as QuIP# and AQLM include
    fine-tuning (part of) the compressed parameters over a limited amount of calibration
    data; however, such fine-tuning techniques over compressed weights often make
    exclusive use of straight-through estimators (STE), whose performance is not well-understood
    in this setting. In this work, we question the use of STE for extreme LLM compression,
    showing that it can be sub-optimal, and perform a systematic study of quantization-aware
    fine-tuning strategies for LLMs.We propose PV-Tuning - a representation-agnostic
    framework that generalizes and improves upon existing fine-tuning strategies,
    and provides convergence guarantees in restricted cases.On the practical side,
    when used for 1-2 bit vector quantization, PV-Tuning outperforms prior techniques
    for highly-performant models such as Llama and Mistral. Using PV-Tuning, we achieve
    the first Pareto-optimal quantization for Llama-2 family models at 2 bits per
    parameter.
acknowledgement: "Authors would like to thank Vage Egiazarian, Andrei Panferov and
  Ruslan Svirschevski for their\r\nhelp and advice on AQLM codebase and running large-scale
  experiments. We also thank Philip\r\nZmushko and Artem Fedorov for helpful discussions
  during the early stages of our research. The research of Kai Yi, Konstantin Burlachenko,
  and Peter Richtárik reported in this publication was supported by funding from King
  Abdullah University of Science and Technology (KAUST) – Center of Excellence for
  Generative AI, under award number 5940. We would also like to thank our NeurIPS
  reviewers for their helpful suggestions, we specifically highlight p3Lv’s suggestions
  to consider smaller codebook sizes and evaluate PV-Tuning with QuIP#, both of which
  produced interesting findings. Finally, we thank the open-source contributors from
  llama.cpp9 and the LocalLlama10 community for discussions and inspirations on practical
  use cases of quantized language models, and in particular, Yalda Shabanzadeh and
  Arthur Aardvark for their help with improving the codebase."
alternative_title:
- Advances in Neural Information Processing Systems
article_processing_charge: No
arxiv: 1
author:
- first_name: Vladimir
  full_name: Malinovskii, Vladimir
  last_name: Malinovskii
- first_name: Denis
  full_name: Mazur, Denis
  last_name: Mazur
- first_name: Ivan
  full_name: Ilin, Ivan
  last_name: Ilin
- first_name: Denis
  full_name: Kuznedelev, Denis
  last_name: Kuznedelev
- first_name: Konstantin
  full_name: Burlachenko, Konstantin
  last_name: Burlachenko
- first_name: Kai
  full_name: Yi, Kai
  last_name: Yi
- first_name: Dan-Adrian
  full_name: Alistarh, Dan-Adrian
  id: 4A899BFC-F248-11E8-B48F-1D18A9856A87
  last_name: Alistarh
  orcid: 0000-0003-3650-940X
- first_name: Peter
  full_name: Richtarik, Peter
  last_name: Richtarik
citation:
  ama: 'Malinovskii V, Mazur D, Ilin I, et al. PV-tuning: Beyond straight-through
    estimation for extreme LLM compression. In: <i>38th Conference on Neural Information
    Processing Systems</i>. Vol 37. Neural Information Processing Systems Foundation;
    2024.'
  apa: 'Malinovskii, V., Mazur, D., Ilin, I., Kuznedelev, D., Burlachenko, K., Yi,
    K., … Richtarik, P. (2024). PV-tuning: Beyond straight-through estimation for
    extreme LLM compression. In <i>38th Conference on Neural Information Processing
    Systems</i> (Vol. 37). Vancouver, Canada: Neural Information Processing Systems
    Foundation.'
  chicago: 'Malinovskii, Vladimir, Denis Mazur, Ivan Ilin, Denis Kuznedelev, Konstantin
    Burlachenko, Kai Yi, Dan-Adrian Alistarh, and Peter Richtarik. “PV-Tuning: Beyond
    Straight-through Estimation for Extreme LLM Compression.” In <i>38th Conference
    on Neural Information Processing Systems</i>, Vol. 37. Neural Information Processing
    Systems Foundation, 2024.'
  ieee: 'V. Malinovskii <i>et al.</i>, “PV-tuning: Beyond straight-through estimation
    for extreme LLM compression,” in <i>38th Conference on Neural Information Processing
    Systems</i>, Vancouver, Canada, 2024, vol. 37.'
  ista: 'Malinovskii V, Mazur D, Ilin I, Kuznedelev D, Burlachenko K, Yi K, Alistarh
    D-A, Richtarik P. 2024. PV-tuning: Beyond straight-through estimation for extreme
    LLM compression. 38th Conference on Neural Information Processing Systems. NeurIPS:
    Neural Information Processing Systems, Advances in Neural Information Processing
    Systems, vol. 37.'
  mla: 'Malinovskii, Vladimir, et al. “PV-Tuning: Beyond Straight-through Estimation
    for Extreme LLM Compression.” <i>38th Conference on Neural Information Processing
    Systems</i>, vol. 37, Neural Information Processing Systems Foundation, 2024.'
  short: V. Malinovskii, D. Mazur, I. Ilin, D. Kuznedelev, K. Burlachenko, K. Yi,
    D.-A. Alistarh, P. Richtarik, in:, 38th Conference on Neural Information Processing
    Systems, Neural Information Processing Systems Foundation, 2024.
conference:
  end_date: 2024-12-15
  location: Vancouver, Canada
  name: 'NeurIPS: Neural Information Processing Systems'
  start_date: 2024-12-10
date_created: 2025-04-06T22:01:32Z
date_published: 2024-12-20T00:00:00Z
date_updated: 2025-05-14T10:49:20Z
day: '20'
ddc:
- '000'
department:
- _id: DaAl
external_id:
  arxiv:
  - '2405.14852'
file:
- access_level: open_access
  checksum: 54d36f947887e26d0e568b512167001a
  content_type: application/pdf
  creator: dernst
  date_created: 2025-04-07T09:17:10Z
  date_updated: 2025-04-07T09:17:10Z
  file_id: '19521'
  file_name: 2024_NeurIPS_Malinovskii.pdf
  file_size: 939712
  relation: main_file
  success: 1
file_date_updated: 2025-04-07T09:17:10Z
has_accepted_license: '1'
intvolume: '        37'
language:
- iso: eng
month: '12'
oa: 1
oa_version: Published Version
publication: 38th Conference on Neural Information Processing Systems
publication_identifier:
  isbn:
  - '9798331314385'
  issn:
  - 1049-5258
publication_status: published
publisher: Neural Information Processing Systems Foundation
quality_controlled: '1'
scopus_import: '1'
status: public
title: 'PV-tuning: Beyond straight-through estimation for extreme LLM compression'
type: conference
user_id: 2DF688A6-F248-11E8-B48F-1D18A9856A87
volume: 37
year: '2024'
...
---
OA_place: repository
OA_type: green
_id: '19520'
abstract:
- lang: eng
  text: Vertebrates exhibit a wide range of motor behaviors, ranging from swimming
    to complex limb-based movements. Here we take advantage of frog metamorphosis,
    which captures a swim-to-limb-based movement transformation during the development
    of a single organism, to explore changes in the underlying spinal circuits. We
    find that the tadpole spinal cord contains small and largely homogeneous populations
    of motor neurons (MNs) and V1 interneurons (V1s) at early escape swimming stages.
    These neuronal populations only modestly increase in number and subtype heterogeneity
    with the emergence of free swimming. In contrast, during frog metamorphosis and
    the emergence of limb movement, there is a dramatic expansion of MN and V1 interneuron
    number and transcriptional heterogeneity, culminating in cohorts of neurons that
    exhibit striking molecular similarity to mammalian motor circuits. CRISPR/Cas9-mediated
    gene disruption of the limb MN and V1 determinants FoxP1 and Engrailed-1, respectively,
    results in severe but selective deficits in tail and limb function. Our work thus
    demonstrates that neural diversity scales exponentially with increasing behavioral
    complexity and illustrates striking evolutionary conservation in the molecular
    organization and function of motor circuits across species.
acknowledged_ssus:
- _id: Bio
acknowledgement: "We would like to thank the members of the Sweeney Lab (especially
  Stavros Papadopoulos and\r\nSophie Gobeil) for their contributions to this project
  and, in addition to the lab, Graziana Gatto\r\nand Mario de Bono, for discussion,
  and support. We are also grateful to Tom Jessell and Chris\r\nKintner for their
  scientific insight and mentorship during the conception of this project. This\r\nproject
  would also not have been possible with the technical support of the Matthias Nowak,\r\nVerena
  Mayer and the Aquatics as well as the Imaging and Optics Facility support teams\r\n(ISTA).
  In addition, we thank our funding sources for providing the resources to do these\r\nexperiments:
  FTI Strategy Lower Austria Dissertation Grant Number FT121-D-046 (D.V.);\r\nHorizon
  Europe ERC Starting Grant Number 101041551 (L.B.S., F.A.T. and D.V); Special\r\nResearch
  Program (SFB) of the Austrian Science Fund (FWF) Project number F7814-B (L.B.S);\r\nNINDS
  5R35NS116858 (J.S.D); CZI grant DAF2020-225401 (DOI): 10.37921/120055ratwvi\r\n(R.H.);
  NIH grant number R01NS123116 (J.B.B); American Lebanese Syrian Associated\r\nCharities
  (ALSAC) (J.B.B.); German Academic Exchange Service (DAAD) IFI Grant Number\r\n57515251-91853472
  (Z.H.); and Project A.L.S. (S.B-M.). "
article_processing_charge: No
author:
- first_name: David
  full_name: Vijatovic, David
  id: cf391e77-ec3c-11ea-a124-d69323410b58
  last_name: Vijatovic
- first_name: 'Florina Alexandra '
  full_name: 'Toma, Florina Alexandra '
  id: 2f73f876-f128-11eb-9611-b96b5a30cb0e
  last_name: Toma
- first_name: Zoe P
  full_name: Harrington, Zoe P
  id: a8144562-32c9-11ee-b5ce-d9800628bda2
  last_name: Harrington
  orcid: 0009-0008-0158-4032
- first_name: Christoph M
  full_name: Sommer, Christoph M
  id: 4DF26D8C-F248-11E8-B48F-1D18A9856A87
  last_name: Sommer
  orcid: 0000-0003-1216-9105
- first_name: Robert
  full_name: Hauschild, Robert
  id: 4E01D6B4-F248-11E8-B48F-1D18A9856A87
  last_name: Hauschild
  orcid: 0000-0001-9843-3522
- first_name: Alexandra J.
  full_name: Trevisan, Alexandra J.
  last_name: Trevisan
- first_name: Phillip
  full_name: Chapman, Phillip
  last_name: Chapman
- first_name: Mara
  full_name: Julseth, Mara
  id: 1cf464b2-dc7d-11ea-9b2f-f9b1aa9417d1
  last_name: Julseth
- first_name: Susan
  full_name: Brenner-Morton, Susan
  last_name: Brenner-Morton
- first_name: Mariano I.
  full_name: Gabitto, Mariano I.
  last_name: Gabitto
- first_name: Jeremy S.
  full_name: Dasen, Jeremy S.
  last_name: Dasen
- first_name: Jay B.
  full_name: Bikoff, Jay B.
  last_name: Bikoff
- first_name: Lora Beatrice Jaeger
  full_name: Sweeney, Lora Beatrice Jaeger
  id: 56BE8254-C4F0-11E9-8E45-0B23E6697425
  last_name: Sweeney
  orcid: 0000-0001-9242-5601
citation:
  ama: Vijatovic D, Toma FA, Harrington ZP, et al. Spinal neuron diversity scales
    exponentially with swim-to-limb transformation during frog metamorphosis. <i>bioRxiv</i>.
    doi:<a href="https://doi.org/10.1101/2024.09.20.614050">10.1101/2024.09.20.614050</a>
  apa: Vijatovic, D., Toma, F. A., Harrington, Z. P., Sommer, C. M., Hauschild, R.,
    Trevisan, A. J., … Sweeney, L. B. (n.d.). Spinal neuron diversity scales exponentially
    with swim-to-limb transformation during frog metamorphosis. <i>bioRxiv</i>. <a
    href="https://doi.org/10.1101/2024.09.20.614050">https://doi.org/10.1101/2024.09.20.614050</a>
  chicago: Vijatovic, David, Florina Alexandra  Toma, Zoe P Harrington, Christoph
    M Sommer, Robert Hauschild, Alexandra J. Trevisan, Phillip Chapman, et al. “Spinal
    Neuron Diversity Scales Exponentially with Swim-to-Limb Transformation during
    Frog Metamorphosis.” <i>BioRxiv</i>, n.d. <a href="https://doi.org/10.1101/2024.09.20.614050">https://doi.org/10.1101/2024.09.20.614050</a>.
  ieee: D. Vijatovic <i>et al.</i>, “Spinal neuron diversity scales exponentially
    with swim-to-limb transformation during frog metamorphosis,” <i>bioRxiv</i>. .
  ista: Vijatovic D, Toma FA, Harrington ZP, Sommer CM, Hauschild R, Trevisan AJ,
    Chapman P, Julseth M, Brenner-Morton S, Gabitto MI, Dasen JS, Bikoff JB, Sweeney
    LB. Spinal neuron diversity scales exponentially with swim-to-limb transformation
    during frog metamorphosis. bioRxiv, <a href="https://doi.org/10.1101/2024.09.20.614050">10.1101/2024.09.20.614050</a>.
  mla: Vijatovic, David, et al. “Spinal Neuron Diversity Scales Exponentially with
    Swim-to-Limb Transformation during Frog Metamorphosis.” <i>BioRxiv</i>, doi:<a
    href="https://doi.org/10.1101/2024.09.20.614050">10.1101/2024.09.20.614050</a>.
  short: D. Vijatovic, F.A. Toma, Z.P. Harrington, C.M. Sommer, R. Hauschild, A.J.
    Trevisan, P. Chapman, M. Julseth, S. Brenner-Morton, M.I. Gabitto, J.S. Dasen,
    J.B. Bikoff, L.B. Sweeney, BioRxiv (n.d.).
corr_author: '1'
date_created: 2025-04-07T08:48:28Z
date_published: 2024-09-27T00:00:00Z
date_updated: 2025-05-14T11:40:13Z
day: '27'
department:
- _id: LoSw
- _id: TiVo
- _id: Bio
- _id: NiBa
doi: 10.1101/2024.09.20.614050
fulldoi: https://doi.org/10.1101/2024.09.20.614050
language:
- iso: eng
main_file_link:
- open_access: '1'
  url: https://doi.org/10.1101/2024.09.20.614050
month: '09'
oa: 1
oa_version: Preprint
project:
- _id: bd73af52-d553-11ed-ba76-912049f0ac7a
  grant_number: FTI21-D-046
  name: Development of V1 interneuron diversity during swim-to-walk transition of
    Xenopus metamorphosis
- _id: ebb66355-77a9-11ec-83b8-b8ac210a4dae
  grant_number: '101041551'
  name: Development and Evolution of Tetrapod Motor Circuits
- _id: c08e9ad1-5a5b-11eb-8a69-9d1cf3b07473
  grant_number: CZI01
  name: Tools for automation and feedback microscopy
publication: bioRxiv
publication_status: submitted
status: public
title: Spinal neuron diversity scales exponentially with swim-to-limb transformation
  during frog metamorphosis
type: preprint
user_id: 2DF688A6-F248-11E8-B48F-1D18A9856A87
year: '2024'
...
