<?xml version="1.0" encoding="UTF-8"?>

<modsCollection xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns="http://www.loc.gov/mods/v3" xsi:schemaLocation="http://www.loc.gov/mods/v3 http://www.loc.gov/standards/mods/v3/mods-3-3.xsd">
<mods version="3.3">

<genre>conference paper</genre>

<titleInfo><title>Neural collapse is globally optimal in deep regularized ResNets and transformers</title></titleInfo>

  
  
<titleInfo type="alternative">
  
  <title>Advances in Neural Information Processing Systems</title>
</titleInfo>

<note type="publicationStatus">published</note>


<note type="qualityControlled">yes</note>

<name type="personal">
  <namePart type="given">Peter</namePart>
  <namePart type="family">Súkeník</namePart>
  <role><roleTerm type="text">author</roleTerm> </role><identifier type="local">d64d6a8d-eb8e-11eb-b029-96fd216dec3c</identifier></name>
<name type="personal">
  <namePart type="given">Christoph</namePart>
  <namePart type="family">Lampert</namePart>
  <role><roleTerm type="text">author</roleTerm> </role><identifier type="local">40C20FD2-F248-11E8-B48F-1D18A9856A87</identifier><description xsi:type="identifierDefinition" type="orcid">0000-0001-8622-7887</description></name>
<name type="personal">
  <namePart type="given">Marco</namePart>
  <namePart type="family">Mondelli</namePart>
  <role><roleTerm type="text">author</roleTerm> </role><identifier type="local">27EB676C-8706-11E9-9510-7717E6697425</identifier><description xsi:type="identifierDefinition" type="orcid">0000-0002-3242-7020</description></name>







<name type="corporate">
  <namePart></namePart>
  <identifier type="local">MaMo</identifier>
  <role>
    <roleTerm type="text">department</roleTerm>
  </role>
</name>

<name type="corporate">
  <namePart></namePart>
  <identifier type="local">GradSch</identifier>
  <role>
    <roleTerm type="text">department</roleTerm>
  </role>
</name>

<name type="corporate">
  <namePart></namePart>
  <identifier type="local">ChLa</identifier>
  <role>
    <roleTerm type="text">department</roleTerm>
  </role>
</name>



<name type="conference">
  <namePart>NeurIPS: Neural Information Processing Systems</namePart>
</name>



<name type="corporate">
  <namePart>Inference in High Dimensions: Light-speed Algorithms and Information Limits</namePart>
  <role><roleTerm type="text">project</roleTerm></role>
</name>



<abstract lang="eng">The empirical emergence of neural collapse—a surprising symmetry in the feature representations of the training data in the penultimate layer of deep neural
networks—has spurred a line of theoretical research aimed at its understanding.
However, existing work focuses on data-agnostic models or, when data structure is
taken into account, it remains limited to multi-layer perceptrons. Our paper fills
both these gaps by analyzing modern architectures in a data-aware regime: we
prove that global optima of deep regularized transformers and residual networks
(ResNets) with LayerNorm trained with cross entropy or mean squared error loss
are approximately collapsed, and the approximation gets tighter as the depth grows.
More generally, we formally reduce any end-to-end large-depth ResNet or transformer training into an equivalent unconstrained features model, thus justifying its
wide use in the literature even beyond data-agnostic settings. Our theoretical results
are supported by experiments on computer vision and language datasets showing
that, as the depth grows, neural collapse indeed becomes more prominent.</abstract>

<originInfo><publisher>Neural Information Processing Systems Foundation</publisher><dateIssued encoding="w3cdtf">2025</dateIssued><place><placeTerm type="text">San Diego, CA, United States</placeTerm></place>
</originInfo>
<language><languageTerm authority="iso639-2b" type="code">eng</languageTerm>
</language>



<relatedItem type="host"><titleInfo><title>39th Conference on Neural Information Processing Systems</title></titleInfo>
  <identifier type="issn">1049-5258</identifier>
  <identifier type="isbn">9798331338275</identifier>
  <identifier type="arXiv">2505.15239</identifier><identifier type="doi">10.52202/085713-1450</identifier>
<part><detail type="volume"><number>38</number></detail><extent unit="pages">48646-48677</extent>
</part>
</relatedItem>


<extension>
<bibliographicCitation>
<apa>Súkeník, P., Lampert, C., &amp;#38; Mondelli, M. (2025). Neural collapse is globally optimal in deep regularized ResNets and transformers. In &lt;i&gt;39th Conference on Neural Information Processing Systems&lt;/i&gt; (Vol. 38, pp. 48646–48677). San Diego, CA, United States: Neural Information Processing Systems Foundation. &lt;a href=&quot;https://doi.org/10.52202/085713-1450&quot;&gt;https://doi.org/10.52202/085713-1450&lt;/a&gt;</apa>
<short>P. Súkeník, C. Lampert, M. Mondelli, in:, 39th Conference on Neural Information Processing Systems, Neural Information Processing Systems Foundation, 2025, pp. 48646–48677.</short>
<chicago>Súkeník, Peter, Christoph Lampert, and Marco Mondelli. “Neural Collapse Is Globally Optimal in Deep Regularized ResNets and Transformers.” In &lt;i&gt;39th Conference on Neural Information Processing Systems&lt;/i&gt;, 38:48646–77. Neural Information Processing Systems Foundation, 2025. &lt;a href=&quot;https://doi.org/10.52202/085713-1450&quot;&gt;https://doi.org/10.52202/085713-1450&lt;/a&gt;.</chicago>
<ieee>P. Súkeník, C. Lampert, and M. Mondelli, “Neural collapse is globally optimal in deep regularized ResNets and transformers,” in &lt;i&gt;39th Conference on Neural Information Processing Systems&lt;/i&gt;, San Diego, CA, United States, 2025, vol. 38, pp. 48646–48677.</ieee>
<ama>Súkeník P, Lampert C, Mondelli M. Neural collapse is globally optimal in deep regularized ResNets and transformers. In: &lt;i&gt;39th Conference on Neural Information Processing Systems&lt;/i&gt;. Vol 38. Neural Information Processing Systems Foundation; 2025:48646-48677. doi:&lt;a href=&quot;https://doi.org/10.52202/085713-1450&quot;&gt;10.52202/085713-1450&lt;/a&gt;</ama>
<mla>Súkeník, Peter, et al. “Neural Collapse Is Globally Optimal in Deep Regularized ResNets and Transformers.” &lt;i&gt;39th Conference on Neural Information Processing Systems&lt;/i&gt;, vol. 38, Neural Information Processing Systems Foundation, 2025, pp. 48646–77, doi:&lt;a href=&quot;https://doi.org/10.52202/085713-1450&quot;&gt;10.52202/085713-1450&lt;/a&gt;.</mla>
<ista>Súkeník P, Lampert C, Mondelli M. 2025. Neural collapse is globally optimal in deep regularized ResNets and transformers. 39th Conference on Neural Information Processing Systems. NeurIPS: Neural Information Processing Systems, Advances in Neural Information Processing Systems, vol. 38, 48646–48677.</ista>
</bibliographicCitation>
</extension>
<recordInfo><recordIdentifier>22825</recordIdentifier><recordCreationDate encoding="w3cdtf">2026-09-06T22:01:59Z</recordCreationDate><recordChangeDate encoding="w3cdtf">2026-09-10T08:11:46Z</recordChangeDate>
</recordInfo>
</mods>
</modsCollection>
