---
OA_place: publisher
OA_type: diamond
_id: '20036'
abstract:
- lang: eng
  text: 'We introduce NeCo: Patch Neighbor Consistency, a novel self-supervised training
    loss that enforces patch-level nearest neighbor consistency across a student and
    teacher model. Compared to contrastive approaches that only yield binary learning
    signals, i.e. "attract" and "repel", this approach benefits from the more fine-grained
    learning signal of sorting spatially dense features relative to reference patches.
    Our method leverages differentiable sorting applied on top of pretrained representations,
    such as DINOv2-registers to bootstrap the learning signal and further improve
    upon them. This dense post-pretraining leads to superior performance across various
    models and datasets, despite requiring only 19 hours on a single GPU. This method
    generates high-quality dense feature encoders and establishes several new state-of-the-art
    results such as +2.3 % and +4.2% for non-parametric in-context semantic segmentation
    on ADE20k and Pascal VOC, +1.6% and +4.8% for linear segmentation evaluations
    on COCO-Things and -Stuff and improvements in the 3D understanding of multi-view
    consistency on SPair-71k, by more than 1.5%.'
article_processing_charge: No
arxiv: 1
author:
- first_name: Valentinos
  full_name: Pariza, Valentinos
  last_name: Pariza
- first_name: Mohammadreza
  full_name: Salehi, Mohammadreza
  last_name: Salehi
- first_name: Gertjan
  full_name: Burghouts, Gertjan
  last_name: Burghouts
- first_name: Francesco
  full_name: Locatello, Francesco
  id: 26cfd52f-2483-11ee-8040-88983bcc06d4
  last_name: Locatello
  orcid: 0000-0002-4850-0683
- first_name: Yuki M.
  full_name: Asano, Yuki M.
  last_name: Asano
citation:
  ama: 'Pariza V, Salehi M, Burghouts G, Locatello F, Asano YM. Near, far: Patch-ordering
    enhances vision foundation models’ scene understanding. In: <i>13th International
    Conference on Learning Representations</i>. ICLR; 2025:72303-72330.'
  apa: 'Pariza, V., Salehi, M., Burghouts, G., Locatello, F., &#38; Asano, Y. M. (2025).
    Near, far: Patch-ordering enhances vision foundation models’ scene understanding.
    In <i>13th International Conference on Learning Representations</i> (pp. 72303–72330).
    Singapore, Singapore: ICLR.'
  chicago: 'Pariza, Valentinos, Mohammadreza Salehi, Gertjan Burghouts, Francesco
    Locatello, and Yuki M. Asano. “Near, Far: Patch-Ordering Enhances Vision Foundation
    Models’ Scene Understanding.” In <i>13th International Conference on Learning
    Representations</i>, 72303–30. ICLR, 2025.'
  ieee: 'V. Pariza, M. Salehi, G. Burghouts, F. Locatello, and Y. M. Asano, “Near,
    far: Patch-ordering enhances vision foundation models’ scene understanding,” in
    <i>13th International Conference on Learning Representations</i>, Singapore, Singapore,
    2025, pp. 72303–72330.'
  ista: 'Pariza V, Salehi M, Burghouts G, Locatello F, Asano YM. 2025. Near, far:
    Patch-ordering enhances vision foundation models’ scene understanding. 13th International
    Conference on Learning Representations. ICLR: International Conference on Learning
    Representations, 72303–72330.'
  mla: 'Pariza, Valentinos, et al. “Near, Far: Patch-Ordering Enhances Vision Foundation
    Models’ Scene Understanding.” <i>13th International Conference on Learning Representations</i>,
    ICLR, 2025, pp. 72303–30.'
  short: V. Pariza, M. Salehi, G. Burghouts, F. Locatello, Y.M. Asano, in:, 13th International
    Conference on Learning Representations, ICLR, 2025, pp. 72303–72330.
conference:
  end_date: 2025-04-28
  location: Singapore, Singapore
  name: 'ICLR: International Conference on Learning Representations'
  start_date: 2025-04-24
date_created: 2025-07-20T22:02:03Z
date_published: 2025-04-01T00:00:00Z
date_updated: 2025-08-04T08:10:55Z
day: '01'
ddc:
- '000'
department:
- _id: FrLo
external_id:
  arxiv:
  - '2408.11054'
file:
- access_level: open_access
  checksum: ddbe981f3ad3f6cb6daf12c954822eb8
  content_type: application/pdf
  creator: dernst
  date_created: 2025-08-04T08:09:43Z
  date_updated: 2025-08-04T08:09:43Z
  file_id: '20109'
  file_name: 2025_ICLR_Pariza.pdf
  file_size: 37788223
  relation: main_file
  success: 1
file_date_updated: 2025-08-04T08:09:43Z
has_accepted_license: '1'
language:
- iso: eng
month: '04'
oa: 1
oa_version: Published Version
page: 72303-72330
publication: 13th International Conference on Learning Representations
publication_identifier:
  isbn:
  - '9798331320850'
publication_status: published
publisher: ICLR
quality_controlled: '1'
scopus_import: '1'
status: public
title: 'Near, far: Patch-ordering enhances vision foundation models'' scene understanding'
tmp:
  image: /images/cc_by.png
  legal_code_url: https://creativecommons.org/licenses/by/4.0/legalcode
  name: Creative Commons Attribution 4.0 International Public License (CC-BY 4.0)
  short: CC BY (4.0)
type: conference
user_id: 2DF688A6-F248-11E8-B48F-1D18A9856A87
year: '2025'
...
