<?xml version="1.0" encoding="UTF-8"?>

<modsCollection xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns="http://www.loc.gov/mods/v3" xsi:schemaLocation="http://www.loc.gov/mods/v3 http://www.loc.gov/standards/mods/v3/mods-3-3.xsd">
<mods version="3.3">

<genre>conference paper</genre>

<titleInfo><title>High-dimensional analysis of knowledge distillation: Weak-to-Strong generalization and scaling laws</title></titleInfo>


<note type="publicationStatus">published</note>


<note type="qualityControlled">yes</note>

<name type="personal">
  <namePart type="given">M.</namePart>
  <namePart type="family">Emrullah Ildiz</namePart>
  <role><roleTerm type="text">author</roleTerm> </role></name>
<name type="personal">
  <namePart type="given">Halil Alperen</namePart>
  <namePart type="family">Gozeten</namePart>
  <role><roleTerm type="text">author</roleTerm> </role></name>
<name type="personal">
  <namePart type="given">Ege Onur</namePart>
  <namePart type="family">Taga</namePart>
  <role><roleTerm type="text">author</roleTerm> </role></name>
<name type="personal">
  <namePart type="given">Marco</namePart>
  <namePart type="family">Mondelli</namePart>
  <role><roleTerm type="text">author</roleTerm> </role><identifier type="local">27EB676C-8706-11E9-9510-7717E6697425</identifier><description xsi:type="identifierDefinition" type="orcid">0000-0002-3242-7020</description></name>
<name type="personal">
  <namePart type="given">Samet</namePart>
  <namePart type="family">Oymak</namePart>
  <role><roleTerm type="text">author</roleTerm> </role></name>







<name type="corporate">
  <namePart></namePart>
  <identifier type="local">MaMo</identifier>
  <role>
    <roleTerm type="text">department</roleTerm>
  </role>
</name>



<name type="conference">
  <namePart>ICLR: International Conference on Learning Representations</namePart>
</name>



<name type="corporate">
  <namePart>Inference in High Dimensions: Light-speed Algorithms and Information Limits</namePart>
  <role><roleTerm type="text">project</roleTerm></role>
</name>



<abstract lang="eng">A growing number of machine learning scenarios rely on knowledge distillation where one uses the output of a surrogate model as labels to supervise the training of a target model. In this work, we provide a sharp characterization of this process for ridgeless, high-dimensional regression, under two settings: (i) model shift, where the surrogate model is arbitrary, and (ii) distribution shift, where the surrogate model is the solution of empirical risk minimization with out-of-distribution data. In both cases, we characterize the precise risk of the target model through non-asymptotic bounds in terms of sample size and data distribution under mild conditions. As a consequence, we identify the form of the optimal surrogate model, which reveals the benefits and limitations of discarding weak features in a data-dependent fashion. In the context of weak-to-strong (W2S) generalization, this has the interpretation that (i) W2S training, with the surrogate as the weak model, can provably outperform training with strong labels under the same data budget, but (ii) it is unable to improve the data scaling law. We validate our results on numerical experiments both on ridgeless regression and on neural network architectures.</abstract>

<relatedItem type="constituent">
  <location>
    <url displayLabel="2025_ICLR_Ildiz.pdf">https://research-explorer.ista.ac.at/download/20033/20112/2025_ICLR_Ildiz.pdf</url>
  </location>
  <physicalDescription><internetMediaType>application/pdf</internetMediaType></physicalDescription><accessCondition type="restrictionOnAccess">no</accessCondition>
</relatedItem>
<originInfo><publisher>ICLR</publisher><dateIssued encoding="w3cdtf">2025</dateIssued><place><placeTerm type="text">Singapore, Singapore</placeTerm></place>
</originInfo>
<language><languageTerm authority="iso639-2b" type="code">eng</languageTerm>
</language>



<relatedItem type="host"><titleInfo><title>13th International Conference on Learning Representations</title></titleInfo>
  <identifier type="isbn">9798331320850</identifier>
  <identifier type="arXiv">2410.18837</identifier>
<part><extent unit="pages">2967-3006</extent>
</part>
</relatedItem>


<extension>
<bibliographicCitation>
<ieee>M. Emrullah Ildiz, H. A. Gozeten, E. O. Taga, M. Mondelli, and S. Oymak, “High-dimensional analysis of knowledge distillation: Weak-to-Strong generalization and scaling laws,” in &lt;i&gt;13th International Conference on Learning Representations&lt;/i&gt;, Singapore, Singapore, 2025, pp. 2967–3006.</ieee>
<mla>Emrullah Ildiz, M., et al. “High-Dimensional Analysis of Knowledge Distillation: Weak-to-Strong Generalization and Scaling Laws.” &lt;i&gt;13th International Conference on Learning Representations&lt;/i&gt;, ICLR, 2025, pp. 2967–3006.</mla>
<chicago>Emrullah Ildiz, M., Halil Alperen Gozeten, Ege Onur Taga, Marco Mondelli, and Samet Oymak. “High-Dimensional Analysis of Knowledge Distillation: Weak-to-Strong Generalization and Scaling Laws.” In &lt;i&gt;13th International Conference on Learning Representations&lt;/i&gt;, 2967–3006. ICLR, 2025.</chicago>
<apa>Emrullah Ildiz, M., Gozeten, H. A., Taga, E. O., Mondelli, M., &amp;#38; Oymak, S. (2025). High-dimensional analysis of knowledge distillation: Weak-to-Strong generalization and scaling laws. In &lt;i&gt;13th International Conference on Learning Representations&lt;/i&gt; (pp. 2967–3006). Singapore, Singapore: ICLR.</apa>
<short>M. Emrullah Ildiz, H.A. Gozeten, E.O. Taga, M. Mondelli, S. Oymak, in:, 13th International Conference on Learning Representations, ICLR, 2025, pp. 2967–3006.</short>
<ama>Emrullah Ildiz M, Gozeten HA, Taga EO, Mondelli M, Oymak S. High-dimensional analysis of knowledge distillation: Weak-to-Strong generalization and scaling laws. In: &lt;i&gt;13th International Conference on Learning Representations&lt;/i&gt;. ICLR; 2025:2967-3006.</ama>
<ista>Emrullah Ildiz M, Gozeten HA, Taga EO, Mondelli M, Oymak S. 2025. High-dimensional analysis of knowledge distillation: Weak-to-Strong generalization and scaling laws. 13th International Conference on Learning Representations. ICLR: International Conference on Learning Representations, 2967–3006.</ista>
</bibliographicCitation>
</extension>
<recordInfo><recordIdentifier>20033</recordIdentifier><recordCreationDate encoding="w3cdtf">2025-07-20T22:02:02Z</recordCreationDate><recordChangeDate encoding="w3cdtf">2025-08-04T08:33:58Z</recordChangeDate>
</recordInfo>
</mods>
</modsCollection>
