---
res:
  bibo_abstract:
  - We study the problem of learning controllers for discrete-time non-linear stochastic
    dynamical systems with formal reach-avoid guarantees. This work presents the first
    method for providing formal reach-avoid guarantees, which combine and generalize
    stability and safety guarantees, with a tolerable probability threshold $p\in[0,1]$
    over the infinite time horizon. Our method leverages advances in machine learning
    literature and it represents formal certificates as neural networks. In particular,
    we learn a certificate in the form of a reach-avoid supermartingale (RASM), a
    novel notion that we introduce in this work. Our RASMs provide reachability and
    avoidance guarantees by imposing constraints on what can be viewed as a stochastic
    extension of level sets of Lyapunov functions for deterministic systems. Our approach
    solves several important problems -- it can be used to learn a control policy
    from scratch, to verify a reach-avoid specification for a fixed control policy,
    or to fine-tune a pre-trained policy if it does not satisfy the reach-avoid specification.
    We validate our approach on $3$ stochastic non-linear reinforcement learning tasks.@eng
  bibo_authorlist:
  - foaf_Person:
      foaf_givenName: Dorde
      foaf_name: Zikelic, Dorde
      foaf_surname: Zikelic
      foaf_workInfoHomepage: http://www.librecat.org/personId=294AA7A6-F248-11E8-B48F-1D18A9856A87
    orcid: 0000-0002-4681-1699
  - foaf_Person:
      foaf_givenName: Mathias
      foaf_name: Lechner, Mathias
      foaf_surname: Lechner
      foaf_workInfoHomepage: http://www.librecat.org/personId=3DC22916-F248-11E8-B48F-1D18A9856A87
  - foaf_Person:
      foaf_givenName: Thomas A
      foaf_name: Henzinger, Thomas A
      foaf_surname: Henzinger
      foaf_workInfoHomepage: http://www.librecat.org/personId=40876CD8-F248-11E8-B48F-1D18A9856A87
    orcid: 0000-0002-2985-7724
  - foaf_Person:
      foaf_givenName: Krishnendu
      foaf_name: Chatterjee, Krishnendu
      foaf_surname: Chatterjee
      foaf_workInfoHomepage: http://www.librecat.org/personId=2E5DCA20-F248-11E8-B48F-1D18A9856A87
    orcid: 0000-0002-4561-241X
  bibo_doi: 10.48550/ARXIV.2210.05308
  dct_date: 2022^xs_gYear
  dct_language: eng
  dct_title: Learning control policies for stochastic systems with reach-avoid guarantees@
...
