<?xml version="1.0" encoding="UTF-8"?>

<modsCollection xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns="http://www.loc.gov/mods/v3" xsi:schemaLocation="http://www.loc.gov/mods/v3 http://www.loc.gov/standards/mods/v3/mods-3-3.xsd">
<mods version="3.3">

<genre>preprint</genre>

<titleInfo><title>Grounded object centric learning</title></titleInfo>


<note type="publicationStatus">submitted</note>



<name type="personal">
  <namePart type="given">Avinash</namePart>
  <namePart type="family">Kori</namePart>
  <role><roleTerm type="text">author</roleTerm> </role></name>
<name type="personal">
  <namePart type="given">Francesco</namePart>
  <namePart type="family">Locatello</namePart>
  <role><roleTerm type="text">author</roleTerm> </role><identifier type="local">26cfd52f-2483-11ee-8040-88983bcc06d4</identifier><description xsi:type="identifierDefinition" type="orcid">0000-0002-4850-0683</description></name>
<name type="personal">
  <namePart type="given">Fabio De Sousa</namePart>
  <namePart type="family">Ribeiro</namePart>
  <role><roleTerm type="text">author</roleTerm> </role></name>
<name type="personal">
  <namePart type="given">Francesca</namePart>
  <namePart type="family">Toni</namePart>
  <role><roleTerm type="text">author</roleTerm> </role></name>
<name type="personal">
  <namePart type="given">Ben</namePart>
  <namePart type="family">Glocker</namePart>
  <role><roleTerm type="text">author</roleTerm> </role></name>







<name type="corporate">
  <namePart></namePart>
  <identifier type="local">FrLo</identifier>
  <role>
    <roleTerm type="text">department</roleTerm>
  </role>
</name>








<abstract lang="eng">The extraction of modular object-centric representations for downstream tasks
is an emerging area of research. Learning grounded representations of objects
that are guaranteed to be stable and invariant promises robust performance
across different tasks and environments. Slot Attention (SA) learns
object-centric representations by assigning objects to \textit{slots}, but
presupposes a \textit{single} distribution from which all slots are randomly
initialised. This results in an inability to learn \textit{specialized} slots
which bind to specific object types and remain invariant to identity-preserving
changes in object appearance. To address this, we present
\emph{\textsc{Co}nditional \textsc{S}lot \textsc{A}ttention} (\textsc{CoSA})
using a novel concept of \emph{Grounded Slot Dictionary} (GSD) inspired by
vector quantization. Our proposed GSD comprises (i) canonical object-level
property vectors and (ii) parametric Gaussian distributions, which define a
prior over the slots. We demonstrate the benefits of our method in multiple
downstream tasks such as scene generation, composition, and task adaptation,
whilst remaining competitive with SA in popular object discovery benchmarks.</abstract>

<originInfo><dateIssued encoding="w3cdtf">2023</dateIssued>
</originInfo>
<language><languageTerm authority="iso639-2b" type="code">eng</languageTerm>
</language>



<relatedItem type="host"><titleInfo><title>arXiv</title></titleInfo>
  <identifier type="arXiv">2307.09437</identifier><identifier type="doi">10.48550/arXiv.2307.09437</identifier>
<part>
</part>
</relatedItem>


<extension>
<bibliographicCitation>
<apa>Kori, A., Locatello, F., Ribeiro, F. D. S., Toni, F., &amp;#38; Glocker, B. (n.d.). Grounded object centric learning. &lt;i&gt;arXiv&lt;/i&gt;. &lt;a href=&quot;https://doi.org/10.48550/arXiv.2307.09437&quot;&gt;https://doi.org/10.48550/arXiv.2307.09437&lt;/a&gt;</apa>
<ieee>A. Kori, F. Locatello, F. D. S. Ribeiro, F. Toni, and B. Glocker, “Grounded object centric learning,” &lt;i&gt;arXiv&lt;/i&gt;. .</ieee>
<short>A. Kori, F. Locatello, F.D.S. Ribeiro, F. Toni, B. Glocker, ArXiv (n.d.).</short>
<mla>Kori, Avinash, et al. “Grounded Object Centric Learning.” &lt;i&gt;ArXiv&lt;/i&gt;, 2307.09437, doi:&lt;a href=&quot;https://doi.org/10.48550/arXiv.2307.09437&quot;&gt;10.48550/arXiv.2307.09437&lt;/a&gt;.</mla>
<chicago>Kori, Avinash, Francesco Locatello, Fabio De Sousa Ribeiro, Francesca Toni, and Ben Glocker. “Grounded Object Centric Learning.” &lt;i&gt;ArXiv&lt;/i&gt;, n.d. &lt;a href=&quot;https://doi.org/10.48550/arXiv.2307.09437&quot;&gt;https://doi.org/10.48550/arXiv.2307.09437&lt;/a&gt;.</chicago>
<ama>Kori A, Locatello F, Ribeiro FDS, Toni F, Glocker B. Grounded object centric learning. &lt;i&gt;arXiv&lt;/i&gt;. doi:&lt;a href=&quot;https://doi.org/10.48550/arXiv.2307.09437&quot;&gt;10.48550/arXiv.2307.09437&lt;/a&gt;</ama>
<ista>Kori A, Locatello F, Ribeiro FDS, Toni F, Glocker B. Grounded object centric learning. arXiv, 2307.09437.</ista>
</bibliographicCitation>
</extension>
<recordInfo><recordIdentifier>14948</recordIdentifier><recordCreationDate encoding="w3cdtf">2024-02-07T14:47:04Z</recordCreationDate><recordChangeDate encoding="w3cdtf">2024-02-12T08:13:12Z</recordChangeDate>
</recordInfo>
</mods>
</modsCollection>
