---
_id: '18255'
abstract:
- lang: eng
  text: Learning an object detection or retrieval system requires a large data set
    with manual annotations. Such data sets are expensive and time consuming to create
    and therefore difficult to obtain on a large scale. In this work, we propose to
    exploit the natural correlation in narrations and the visual presence of objects
    in video, to learn an object detector and retrieval without any manual labeling
    involved. We pose the problem as weakly supervised learning with noisy labels,
    and propose a novel object detection paradigm under these constraints. We handle
    the background rejection by using contrastive samples and confront the high level
    of label noise with a new clustering score. Our evaluation is based on a set of
    11 manually annotated objects in over 5000 frames. We show comparison to a weakly-supervised
    approach as baseline and provide a strongly labeled upper bound.
article_number: '9022341'
article_processing_charge: No
arxiv: 1
author:
- first_name: Elad
  full_name: Amrani, Elad
  last_name: Amrani
- first_name: Rami
  full_name: Ben-Ari, Rami
  last_name: Ben-Ari
- first_name: Tal
  full_name: Hakim, Tal
  last_name: Hakim
- first_name: Alexander
  full_name: Bronstein, Alexander
  id: 58f3726e-7cba-11ef-ad8b-e6e8cb3904e6
  last_name: Bronstein
  orcid: 0000-0001-9699-8730
citation:
  ama: 'Amrani E, Ben-Ari R, Hakim T, Bronstein AM. Learning to detect and retrieve
    objects from unlabeled videos. In: <i>2019 IEEE/CVF International Conference on
    Computer Vision Workshop (ICCVW)</i>. IEEE; 2020. doi:<a href="https://doi.org/10.1109/iccvw.2019.00567">10.1109/iccvw.2019.00567</a>'
  apa: 'Amrani, E., Ben-Ari, R., Hakim, T., &#38; Bronstein, A. M. (2020). Learning
    to detect and retrieve objects from unlabeled videos. In <i>2019 IEEE/CVF International
    Conference on Computer Vision Workshop (ICCVW)</i>. Seoul, Korea (South): IEEE.
    <a href="https://doi.org/10.1109/iccvw.2019.00567">https://doi.org/10.1109/iccvw.2019.00567</a>'
  chicago: Amrani, Elad, Rami Ben-Ari, Tal Hakim, and Alex M. Bronstein. “Learning
    to Detect and Retrieve Objects from Unlabeled Videos.” In <i>2019 IEEE/CVF International
    Conference on Computer Vision Workshop (ICCVW)</i>. IEEE, 2020. <a href="https://doi.org/10.1109/iccvw.2019.00567">https://doi.org/10.1109/iccvw.2019.00567</a>.
  ieee: E. Amrani, R. Ben-Ari, T. Hakim, and A. M. Bronstein, “Learning to detect
    and retrieve objects from unlabeled videos,” in <i>2019 IEEE/CVF International
    Conference on Computer Vision Workshop (ICCVW)</i>, Seoul, Korea (South), 2020.
  ista: Amrani E, Ben-Ari R, Hakim T, Bronstein AM. 2020. Learning to detect and retrieve
    objects from unlabeled videos. 2019 IEEE/CVF International Conference on Computer
    Vision Workshop (ICCVW). 17th IEEE/CVF International Conference on Computer Vision
    Workshop, 9022341.
  mla: Amrani, Elad, et al. “Learning to Detect and Retrieve Objects from Unlabeled
    Videos.” <i>2019 IEEE/CVF International Conference on Computer Vision Workshop
    (ICCVW)</i>, 9022341, IEEE, 2020, doi:<a href="https://doi.org/10.1109/iccvw.2019.00567">10.1109/iccvw.2019.00567</a>.
  short: E. Amrani, R. Ben-Ari, T. Hakim, A.M. Bronstein, in:, 2019 IEEE/CVF International
    Conference on Computer Vision Workshop (ICCVW), IEEE, 2020.
conference:
  end_date: 2019-10-28
  location: Seoul, Korea (South)
  name: 17th IEEE/CVF International Conference on Computer Vision Workshop
  start_date: 2019-10-27
date_created: 2024-10-08T13:07:16Z
date_published: 2020-03-05T00:00:00Z
date_updated: 2024-12-05T16:04:03Z
day: '05'
doi: 10.1109/iccvw.2019.00567
extern: '1'
external_id:
  arxiv:
  - '1905.11137'
language:
- iso: eng
main_file_link:
- open_access: '1'
  url: https://doi.org/10.48550/arXiv.1905.11137
month: '03'
oa: 1
oa_version: Preprint
publication: 2019 IEEE/CVF International Conference on Computer Vision Workshop (ICCVW)
publication_identifier:
  eissn:
  - 2473-9944
  isbn:
  - '9781728150246'
publication_status: published
publisher: IEEE
quality_controlled: '1'
scopus_import: '1'
status: public
title: Learning to detect and retrieve objects from unlabeled videos
type: conference
user_id: 3E5EF7F0-F248-11E8-B48F-1D18A9856A87
year: '2020'
...
