---
OA_place: publisher
_id: '22857'
abstract:
- lang: eng
  text: "Artificial intelligence and machine learning have undergone an unprecedented
    evolution in the past decade, motivating a research effort toward a theory able
    to capture the qualitative behavior of large-scale neural systems. A central puzzle
    has been the clear benefit of scaling architecture size and overfitting the training
    set in supervised learning tasks. This evidence, in apparent contradiction with
    classical statistical learning theory, pushed researchers to develop a new theory
    capturing the interplay between the algorithmic and architectural bias of training
    and the specific target function, differently from previous methods rooted in
    uniform stability.\r\nThis approach has enabled a grounded understanding of novel
    learning regimes, typically through formal limits where the number of training
    samples $n$, data dimensions $d$, and model parameters $p$ grow to infinity at
    different rates. \\\\\r\nIn this thesis, we follow this approach, focusing on
    the trustworthiness of high-dimensional models: properties that are difficult
    to control during training or deployment and often emerge under unpredictable
    or adversarial conditions. In such settings, it is crucial to formally ensure
    a priori the reliability of machine learning systems.\r\nFirst, we study data
    memorization, both as label fitting and as the storage of private information
    about training samples in trained parameters. We prove that $p = \\Omega(n)$ parameters
    are sufficient for a deep neural network to memorize a generic set of labels,
    and for a model to memorize spurious features across training data. We then give
    evidence that $p = \\Omega(dn)$ parameters are instead necessary for an adversary
    to reconstruct the full training set from the trained parameters.\r\nSecond, we
    study robustness, both to adversarial perturbations and to distribution shift.
    We first prove that $p = \\Omega(dn)$ parameters can be sufficient for a class
    of neural networks to overfit the training data while guaranteeing robustness
    to adversarial perturbations. Then, we focus on spurious correlations learning
    in high-dimensional regression, studying the effect of the ridge regularization
    parameter in the proportional regime $n = \\Theta(d)$, and connecting it via an
    equivalence argument to the role of over-parameterization $p = \\Omega(n)$ in
    neural networks. We also investigate the architectural bias of attention-based
    networks, showing that they are sensitive to the replacement of individual words
    in an embedded sentence, allowing them to generalize on sentences where the contextual
    meaning depends on one or few words.\r\nFinally, we study differentially private
    optimization in high-dimensional regimes. We prove that standard private gradient
    methods do not suffer in the over-parameterized regime $p = \\Omega(n)$, challenging
    the current wisdom based on stability-derived generalization bounds. We then consider
    linear regression in the proportional regime $n = \\Theta(d)$, showing that standard
    private gradient descent can achieve optimal rates under appropriate hyper-parameter
    scaling, such as sufficiently small gradient clipping constants, whose role is
    still debated in practice."
acknowledged_ssus:
- _id: ScienComp
acknowledgement: "This project was partially supported by the 2019 Lopez-Loreta prize,\r\nthe
  European Union (ERC, INF2\r\n, project number 101161364), the Austrian Science Fund\r\n(FWF)
  10.55776/COE12, and a Google PhD fellowship in machine intelligence. Furthermore,\r\nthe
  candidate acknowledges the support from the Scientific Service Units of the Institute
  of\r\nScience and Technology Austria through resources provided by Scientific Computing."
alternative_title:
- ISTA Thesis
article_processing_charge: No
author:
- first_name: Simone
  full_name: Bombari, Simone
  id: ca726dda-de17-11ea-bc14-f9da834f63aa
  last_name: Bombari
citation:
  ama: Bombari S. Trustworthy machine learning in high dimensions. 2026. doi:<a href="https://doi.org/10.15479/AT-ISTA-22857">10.15479/AT-ISTA-22857</a>
  apa: Bombari, S. (2026). <i>Trustworthy machine learning in high dimensions</i>.
    Institute of Science and Technology Austria. <a href="https://doi.org/10.15479/AT-ISTA-22857">https://doi.org/10.15479/AT-ISTA-22857</a>
  chicago: Bombari, Simone. “Trustworthy Machine Learning in High Dimensions.” Institute
    of Science and Technology Austria, 2026. <a href="https://doi.org/10.15479/AT-ISTA-22857">https://doi.org/10.15479/AT-ISTA-22857</a>.
  ieee: S. Bombari, “Trustworthy machine learning in high dimensions,” Institute of
    Science and Technology Austria, 2026.
  ista: Bombari S. 2026. Trustworthy machine learning in high dimensions. Institute
    of Science and Technology Austria.
  mla: Bombari, Simone. <i>Trustworthy Machine Learning in High Dimensions</i>. Institute
    of Science and Technology Austria, 2026, doi:<a href="https://doi.org/10.15479/AT-ISTA-22857">10.15479/AT-ISTA-22857</a>.
  short: S. Bombari, Trustworthy Machine Learning in High Dimensions, Institute of
    Science and Technology Austria, 2026.
corr_author: '1'
das_tickbox: '0'
date_created: 2026-09-08T13:40:08Z
date_published: 2026-09-08T00:00:00Z
date_updated: 2026-09-21T13:07:01Z
day: '08'
ddc:
- '519'
degree_awarded: PhD
department:
- _id: GradSch
- _id: MaMo
doi: 10.15479/AT-ISTA-22857
doi_confirm: '1'
file:
- access_level: closed
  checksum: 8eab99d6dc6e4476826bdfc7c00fdebb
  content_type: application/zip
  creator: sbombari
  date_created: 2026-09-08T13:28:38Z
  date_updated: 2026-09-08T13:28:38Z
  file_id: '22859'
  file_name: Thesis copy.zip
  file_size: 99154325
  relation: source_file
- access_level: open_access
  checksum: 007033bafe4622ff4c2e758301354df1
  content_type: application/pdf
  creator: sbombari
  date_created: 2026-09-10T10:03:48Z
  date_updated: 2026-09-10T10:03:48Z
  file_id: '22898'
  file_name: 2026_Bombari_Simone_Thesis.pdf
  file_size: 12856211
  relation: main_file
  success: 1
file_date_updated: 2026-09-10T10:03:48Z
fulldoi: https://doi.org/10.15479/AT-ISTA-22857
has_accepted_license: '1'
keyword:
- machine learning
- high-dimensional statistics
- deep learning theory
- privacy
- memorization
- robustness
language:
- iso: eng
month: '09'
oa: 1
oa_version: Published Version
page: '446'
project:
- _id: 92099302-16d5-11f0-9cad-f9a785f54fbd
  name: 'Trustworthy Deep Learning Theory: Private Over-Parameterized Models and Robust
    LLMs'
- _id: 911e6d1f-16d5-11f0-9cad-c5c68c6a1cdf
  grant_number: '101161364'
  name: 'Inference in High Dimensions: Light-speed Algorithms and Information Limits'
- _id: 74caaef7-b034-11f1-8f2d-e0e993bb422e
  grant_number: COE12
  name: Bilateral Artificial Intelligence (Mondelli)
publication_identifier:
  isbn:
  - 978-3-99078-091-6
  issn:
  - 2663-337X
publication_status: published
publisher: Institute of Science and Technology Austria
related_material:
  record:
  - id: '12537'
    relation: part_of_dissertation
    status: public
  - id: '18972'
    relation: part_of_dissertation
    status: public
  - id: '18973'
    relation: part_of_dissertation
    status: public
  - id: '21324'
    relation: part_of_dissertation
    status: public
  - id: '22894'
    relation: part_of_dissertation
    status: public
  - id: '19627'
    relation: part_of_dissertation
    status: public
  - id: '12859'
    relation: part_of_dissertation
    status: public
researchdata_availability: no
status: public
supervisor:
- first_name: Marco
  full_name: Mondelli, Marco
  id: 27EB676C-8706-11E9-9510-7717E6697425
  last_name: Mondelli
  orcid: 0000-0002-3242-7020
supplementarymaterial: no
title: Trustworthy machine learning in high dimensions
type: dissertation
user_id: 8b945eb4-e2f2-11eb-945a-df72226e66a9
year: '2026'
...
---
OA_place: publisher
OA_type: gold
_id: '21326'
abstract:
- lang: eng
  text: 'Neural Collapse is a phenomenon where the last-layer representations of a
    well-trained neural network converge to a highly structured geometry. In this
    paper, we focus on its first (and most basic) property, known as NC1: the within-class
    variability vanishes. While prior theoretical studies establish the occurrence
    of NC1 via the data-agnostic unconstrained features model, our work adopts a data-specific
    perspective, analyzing NC1 in a three-layer neural network, with the first two
    layers operating in the mean-field regime and followed by a linear layer. In particular,
    we establish a fundamental connection between NC1 and the loss landscape: we prove
    that points with small empirical loss and gradient norm (thus, close to being
    stationary) approximately satisfy NC1, and the closeness to NC1 is controlled
    by the residual loss and gradient norm. We then show that (i) gradient flow on
    the mean squared error converges to NC1 solutions with small empirical loss, and
    (ii) for well-separated data distributions, both NC1 and vanishing test loss are
    achieved simultaneously. This aligns with the empirical observation that NC1 emerges
    during training while models attain near-zero test error. Overall, our results
    demonstrate that NC1 arises from gradient training due to the properties of the
    loss landscape, and they show the co-occurrence of NC1 and small test error for
    certain data distributions.'
acknowledgement: "This research was funded in whole or in part by the Austrian Science
  Fund (FWF) 10.55776/COE12. For the purpose of open access, the authors have applied
  a CC BY public\r\ncopyright license to any Author Accepted Manuscript version arising
  from this submission. The authors would like to thank Peter Sukenık for general
  helpful discussions and for pointing out that all the stationary points are approximately
  proportional in the case without entropic regularization. "
alternative_title:
- PMLR
article_processing_charge: No
arxiv: 1
author:
- first_name: Diyuan
  full_name: Wu, Diyuan
  id: 1a5914c2-896a-11ed-bdf8-fb80621a0635
  last_name: Wu
- first_name: Marco
  full_name: Mondelli, Marco
  id: 27EB676C-8706-11E9-9510-7717E6697425
  last_name: Mondelli
  orcid: 0000-0002-3242-7020
citation:
  ama: 'Wu D, Mondelli M. Neural collapse beyond the unconstrained features model:
    Landscape, dynamics, and generalization in the mean-field regime. In: <i>Proceedings
    of the 42nd International Conference on Machine Learning</i>. Vol 267. ML Research
    Press; 2025:67499-67536.'
  apa: 'Wu, D., &#38; Mondelli, M. (2025). Neural collapse beyond the unconstrained
    features model: Landscape, dynamics, and generalization in the mean-field regime.
    In <i>Proceedings of the 42nd International Conference on Machine Learning</i>
    (Vol. 267, pp. 67499–67536). Vancouver, Canada: ML Research Press.'
  chicago: 'Wu, Diyuan, and Marco Mondelli. “Neural Collapse beyond the Unconstrained
    Features Model: Landscape, Dynamics, and Generalization in the Mean-Field Regime.”
    In <i>Proceedings of the 42nd International Conference on Machine Learning</i>,
    267:67499–536. ML Research Press, 2025.'
  ieee: 'D. Wu and M. Mondelli, “Neural collapse beyond the unconstrained features
    model: Landscape, dynamics, and generalization in the mean-field regime,” in <i>Proceedings
    of the 42nd International Conference on Machine Learning</i>, Vancouver, Canada,
    2025, vol. 267, pp. 67499–67536.'
  ista: 'Wu D, Mondelli M. 2025. Neural collapse beyond the unconstrained features
    model: Landscape, dynamics, and generalization in the mean-field regime. Proceedings
    of the 42nd International Conference on Machine Learning. ICML: International
    Conference on Machine Learning, PMLR, vol. 267, 67499–67536.'
  mla: 'Wu, Diyuan, and Marco Mondelli. “Neural Collapse beyond the Unconstrained
    Features Model: Landscape, Dynamics, and Generalization in the Mean-Field Regime.”
    <i>Proceedings of the 42nd International Conference on Machine Learning</i>, vol.
    267, ML Research Press, 2025, pp. 67499–536.'
  short: D. Wu, M. Mondelli, in:, Proceedings of the 42nd International Conference
    on Machine Learning, ML Research Press, 2025, pp. 67499–67536.
conference:
  end_date: 2025-07-19
  location: Vancouver, Canada
  name: 'ICML: International Conference on Machine Learning'
  start_date: 2025-07-13
corr_author: '1'
date_created: 2026-02-18T12:02:45Z
date_published: 2025-07-30T00:00:00Z
date_updated: 2026-09-16T07:25:42Z
day: '30'
ddc:
- '000'
department:
- _id: MaMo
external_id:
  arxiv:
  - '2501.19104'
file:
- access_level: open_access
  checksum: c5ce8b1c83e33dc3a11122f4910deb67
  content_type: application/pdf
  creator: dernst
  date_created: 2026-02-19T08:28:22Z
  date_updated: 2026-02-19T08:28:22Z
  file_id: '21337'
  file_name: 2025_ICML_Wu.pdf
  file_size: 3994385
  relation: main_file
  success: 1
file_date_updated: 2026-02-19T08:28:22Z
has_accepted_license: '1'
intvolume: '       267'
language:
- iso: eng
month: '07'
oa: 1
oa_version: Published Version
page: 67499-67536
project:
- _id: 74caaef7-b034-11f1-8f2d-e0e993bb422e
  grant_number: COE12
  name: Bilateral Artificial Intelligence (Mondelli)
publication: Proceedings of the 42nd International Conference on Machine Learning
publication_identifier:
  eissn:
  - 2640-3498
publication_status: published
publisher: ML Research Press
quality_controlled: '1'
status: public
title: 'Neural collapse beyond the unconstrained features model: Landscape, dynamics,
  and generalization in the mean-field regime'
tmp:
  image: /images/cc_by.png
  legal_code_url: https://creativecommons.org/licenses/by/4.0/legalcode
  name: Creative Commons Attribution 4.0 International Public License (CC-BY 4.0)
  short: CC BY (4.0)
type: conference
user_id: 2DF688A6-F248-11E8-B48F-1D18A9856A87
volume: 267
year: '2025'
...
---
APC_amount: 2754,32 EUR
OA_place: publisher
OA_type: hybrid
_id: '19627'
abstract:
- lang: eng
  text: Differentially private gradient descent (DP-GD) is a popular algorithm to
    train deep learning models with provable guarantees on the privacy of the training
    data. In the last decade, the problem of understanding its performance cost with
    respect to standard GD has received remarkable attention from the research community,
    which formally derived upper bounds on the excess population risk  RP  in different
    learning settings. However, existing bounds typically degrade with over-parameterization,
    i.e., as the number of parameters  p  gets larger than the number of training
    samples  n  -- a regime which is ubiquitous in current deep-learning practice.
    As a result, the lack of theoretical insights leaves practitioners without clear
    guidance, leading some to reduce the effective number of trainable parameters
    to improve performance, while others use larger models to achieve better results
    through scale. In this work, we show that in the popular random features model
    with quadratic loss, for any sufficiently large  p , privacy can be obtained for
    free, i.e.,  |RP|=o(1) , not only when the privacy parameter  ε  has constant
    order, but also in the strongly private setting  ε=o(1) . This challenges the
    common wisdom that over-parameterization inherently hinders performance in private
    learning.
acknowledgement: This research was funded in whole, or in part, by the Austrian Science
  Fund (FWF) Grant number COE 12. For the purpose of open access, the author has applied
  a CC BY public copyright license to any Author Accepted Manuscript version arising
  from this submission. The authors were also supported by the 2019 Lopez-Loreta prize,
  and Simone Bombari was supported by a Google PhD fellowship. We thank Diyuan Wu,
  Edwige Cyffers, Francesco Pedrotti, Inbar Seroussi, Nikita P. Kalinin, Pietro Pelliconi,
  Roodabeh Safavi, Yizhe Zhu, and Zhichao Wang for helpful discussions.
article_number: e2423072122
article_processing_charge: Yes (in subscription journal)
article_type: original
arxiv: 1
author:
- first_name: Simone
  full_name: Bombari, Simone
  id: ca726dda-de17-11ea-bc14-f9da834f63aa
  last_name: Bombari
- first_name: Marco
  full_name: Mondelli, Marco
  id: 27EB676C-8706-11E9-9510-7717E6697425
  last_name: Mondelli
  orcid: 0000-0002-3242-7020
citation:
  ama: Bombari S, Mondelli M. Privacy for free in the overparameterized regime. <i>Proceedings
    of the National Academy of Sciences</i>. 2025;122(15). doi:<a href="https://doi.org/10.1073/pnas.2423072122">10.1073/pnas.2423072122</a>
  apa: Bombari, S., &#38; Mondelli, M. (2025). Privacy for free in the overparameterized
    regime. <i>Proceedings of the National Academy of Sciences</i>. National Academy
    of Sciences. <a href="https://doi.org/10.1073/pnas.2423072122">https://doi.org/10.1073/pnas.2423072122</a>
  chicago: Bombari, Simone, and Marco Mondelli. “Privacy for Free in the Overparameterized
    Regime.” <i>Proceedings of the National Academy of Sciences</i>. National Academy
    of Sciences, 2025. <a href="https://doi.org/10.1073/pnas.2423072122">https://doi.org/10.1073/pnas.2423072122</a>.
  ieee: S. Bombari and M. Mondelli, “Privacy for free in the overparameterized regime,”
    <i>Proceedings of the National Academy of Sciences</i>, vol. 122, no. 15. National
    Academy of Sciences, 2025.
  ista: Bombari S, Mondelli M. 2025. Privacy for free in the overparameterized regime.
    Proceedings of the National Academy of Sciences. 122(15), e2423072122.
  mla: Bombari, Simone, and Marco Mondelli. “Privacy for Free in the Overparameterized
    Regime.” <i>Proceedings of the National Academy of Sciences</i>, vol. 122, no.
    15, e2423072122, National Academy of Sciences, 2025, doi:<a href="https://doi.org/10.1073/pnas.2423072122">10.1073/pnas.2423072122</a>.
  short: S. Bombari, M. Mondelli, Proceedings of the National Academy of Sciences
    122 (2025).
corr_author: '1'
date_created: 2025-04-27T22:02:13Z
date_published: 2025-04-15T00:00:00Z
date_updated: 2026-09-21T13:07:01Z
day: '15'
ddc:
- '000'
department:
- _id: MaMo
doi: 10.1073/pnas.2423072122
external_id:
  arxiv:
  - '2410.14787'
  isi:
  - '001471214000001'
  pmid:
  - '40215275'
file:
- access_level: open_access
  checksum: 1ac6f78e368d35a0cafb4d2d9bd63443
  content_type: application/pdf
  creator: dernst
  date_created: 2025-05-05T07:27:54Z
  date_updated: 2025-05-05T07:27:54Z
  file_id: '19648'
  file_name: 2025_PNAS_Bombari.pdf
  file_size: 2328320
  relation: main_file
  success: 1
file_date_updated: 2025-05-05T07:27:54Z
fulldoi: https://doi.org/10.1073/pnas.2423072122
has_accepted_license: '1'
intvolume: '       122'
isi: 1
issue: '15'
language:
- iso: eng
month: '04'
oa: 1
oa_version: Published Version
pmid: 1
project:
- _id: 059876FA-7A3F-11EA-A408-12923DDC885E
  name: Prix Lopez-Loretta 2019 - Marco Mondelli
- _id: 92099302-16d5-11f0-9cad-f9a785f54fbd
  name: 'Trustworthy Deep Learning Theory: Private Over-Parameterized Models and Robust
    LLMs'
- _id: 74caaef7-b034-11f1-8f2d-e0e993bb422e
  grant_number: COE12
  name: Bilateral Artificial Intelligence (Mondelli)
publication: Proceedings of the National Academy of Sciences
publication_identifier:
  eissn:
  - 1091-6490
  issn:
  - 0027-8424
publication_status: published
publisher: National Academy of Sciences
quality_controlled: '1'
related_material:
  record:
  - id: '22857'
    relation: dissertation_contains
    status: public
scopus_import: '1'
status: public
title: Privacy for free in the overparameterized regime
tmp:
  image: /images/cc_by.png
  legal_code_url: https://creativecommons.org/licenses/by/4.0/legalcode
  name: Creative Commons Attribution 4.0 International Public License (CC-BY 4.0)
  short: CC BY (4.0)
type: journal_article
user_id: 2DF688A6-F248-11E8-B48F-1D18A9856A87
volume: 122
year: '2025'
...
