Activities per year
Abstract
Reinforcement learning (RL) is a powerful framework for learning complex behaviors, but lacks adoption in many settings due to sample size requirements. We introduce a framework for increasing sample efficiency of RL algorithms. Our approach focuses on optimizing environment rewards with high-level instructions. These are modeled as a high-level controller over temporally extended actions known as options. These options can be looped, interleaved and partially ordered with a rich language for high-level instructions. Crucially, the instructions may be underspecified in the sense that following them does not guarantee high reward in the environment. We present an algorithm for control with these so-called option machines (OMs), discuss option selection for the partially ordered case and describe an algorithm for learning with OMs. We compare our approach in zero-shot, single- and multi-task settings in an environment with fully specified and underspecified instructions. We find that OMs perform significantly better than or comparable to the state-of-art in all environments and learning settings.
| Original language | English |
|---|---|
| Title of host publication | Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence |
| Editors | Luc De Raedt |
| Publisher | International Joint Conferences on Artificial Intelligence Organization |
| Pages | 2909-2915 |
| Number of pages | 7 |
| ISBN (Electronic) | 9781956792003 |
| DOIs | |
| Publication status | Published - 2022 |
| Event | 31st International Joint Conference on Artificial Intelligence, IJCAI 2022 - Vienna, Austria Duration: 23 Jul 2022 → 29 Jul 2022 |
Publication series
| Name | IJCAI International Joint Conference on Artificial Intelligence |
|---|---|
| ISSN (Print) | 1045-0823 |
Conference
| Conference | 31st International Joint Conference on Artificial Intelligence, IJCAI 2022 |
|---|---|
| Country/Territory | Austria |
| City | Vienna |
| Period | 23/07/22 → 29/07/22 |
Bibliographical note
Publisher Copyright:© 2022 International Joint Conferences on Artificial Intelligence. All rights reserved.
Keywords
- reinforcement learning
- finite state transducers
- Multi-task environments
- artificial intelligence
- machine learning
Fingerprint
Dive into the research topics of 'Reinforcement Learning with Option Machines'. Together they form a unique fingerprint.-
Research visit DFKI Saarbrücken
den Hengst, F. (Speaker)
24 Aug 2024 → 27 Aug 2024Activity: Lecture / Presentation › Academic
-
Department of Computer Science, Saarland University
den Hengst, F. (Visiting researcher), Gros, T. (Visiting researcher) & Wolf, V. (Visiting researcher)
27 Aug 2024 → 28 Aug 2024Activity: Visiting an external institution › Visiting an external academic institution
Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver