Relative Entropy Inverse Reinforcement Learning

Boularias, A.; Kober, J.; Peters, J.

Datensatz

DATENSATZ AKTIONENEXPORT

Zur Ablage hinzufügen

Bitte beachten Sie, dass eine neuere Version dieses Datensatzes verfügbar ist:
https://pure.mpg.de/pubman/item/item_1577808_7

DetailsÜbersicht

Relative Entropy Inverse Reinforcement Learning

Boularias, A., Kober, J., & Peters, J. (2011). Relative Entropy Inverse Reinforcement Learning.

Item is Freigegeben

einblenden: alle ausblenden: alle

Basisdaten

einblenden: ausblenden:

Datensatz-Permalink: https://hdl.handle.net/11858/00-001M-0000-0010-4EFD-0 Versions-Permalink: https://hdl.handle.net/11858/00-001M-0000-0028-8047-0

Genre: Konferenzbeitrag

Alternativer Titel : JMLR: W&CP

ausblenden:

Urheber:
Boularias, A.¹, Autor
Kober, J.¹, Autor
Peters, J.¹, Autor

Affiliations:
1Dept. Empirical Inference, Max Planck Institute for Intelligent Systems, Max Planck Society, ou_1497647

Inhalt

einblenden:

ausblenden:

Schlagwörter: MPI für Intelligente Systeme; Abt. Schölkopf;

Zusammenfassung: We consider the problem of imitation learning where the examples, demonstrated by an expert, cover only a small part of a large state space. Inverse Reinforcement Learning (IRL) provides an efficient tool for generalizing the demonstration, based on the assumption that the expert is optimally acting in a Markov Decision Process (MDP). Most of the past work on IRL requires that a (near)-optimal policy can be computed for different reward functions. However, this requirement can hardly be satisfied in systems with a large, or continuous, state space. In this paper, we propose a model-free IRL algorithm, where the relative entropy between the empirical distribution of the state-action trajectories under a uniform policy and their distribution under the learned policy is minimized by stochastic gradient descent. We compare this new approach to well-known IRL algorithms using approximate MDP models. Empirical results on simulated car racing, gridworld and ball-in-a-cup problems show that our approach is able to learn good policies from a small number of demonstrations.

Details

einblenden:

ausblenden:

Sprache(n):

Datum: Erschienen: 2011

Publikationsstatus: Erschienen

Seiten: -

Ort, Verlag, Ausgabe: -

Inhaltsverzeichnis: -

Art der Begutachtung: -

Identifikatoren: eDoc: 596865
URI: http://www.kyb.tuebingen.mpg.de/fileadmin/user_upload/files/publications/2011/Aistats-2011-Boularias.pdf
Anderer: BoulariasKP2011

Art des Abschluß: -

Veranstaltung

einblenden:

ausblenden:

Titel: Fourteenth International Conference on Artificial Intelligence and Statistics (AISTATS)

Veranstaltungsort: Fort Lauderdale, FL, USA

Start-/Enddatum: 2011-04-11 - 2011-04-13

Entscheidung

einblenden:

Projektinformation

einblenden:

Quelle 1

einblenden:

ausblenden:

Titel: Proceedings of the 14th International Conference on Artificial Intelligence and Statistics (AISTATS)

Genre der Quelle: Heft

Urheber:
Gordon, G., Herausgeber
Dunson, D., Herausgeber
M., Dudík., Herausgeber

Affiliations:
-

Ort, Verlag, Ausgabe: -

Seiten: 7 Band / Heft: - Artikelnummer: - Start- / Endseite: 182 - 189 Identifikator: -

Quelle 2

einblenden:

ausblenden:

Titel: JMLR: Workshop and Conference Proceedings

Alternativer Titel : JMLR: W&CP

Genre der Quelle: Reihe

Urheber:

Affiliations:

Ort, Verlag, Ausgabe: -

Seiten: 7 Band / Heft: 15 Artikelnummer: - Start- / Endseite: - Identifikator: -

Datensatz

Basisdaten

Dateien

Externe Referenzen

Urheber

Inhalt

Details

Veranstaltung

Entscheidung

Projektinformation

Quelle 1

Quelle 2