English
 
Help Privacy Policy Disclaimer
  Advanced SearchBrowse

Item

ITEM ACTIONSEXPORT
  Automatic Task Decomposition using Compositional Reinforcement Learning

Tano, P., Dayan, P., & Pouget, A. (2022). Automatic Task Decomposition using Compositional Reinforcement Learning. Poster presented at Computational and Systems Neuroscience Meeting (COSYNE 2022), Lisboa, Portugal.

Item is

Files

show Files

Creators

show
hide
 Creators:
Tano, P, Author
Dayan, P1, Author           
Pouget, A, Author
Affiliations:
1Department of Computational Neuroscience, Max Planck Institute for Biological Cybernetics, Max Planck Society, ou_3017468              

Content

show
hide
Free keywords: -
 Abstract: Decomposing complex tasks into their simpler components is often the only way for animals to make any meaningful progress at all. We show that reusing the traditional reward prediction error machinery at multiple hierarchical levels allows complex tasks to be automatically decomposed in a compositional manner, leading to fast and flexible reinforcement learning. In this compositional reinforcement learning (CRL) framework, the agent computes a set of predictions for each state in the form of hierarchically organized general value functions (GVFs). Level 0 GVFs predict whether continuing straight along cardinal directions in the state space will lead to a rewarded location; while a level P GVF predicts whether the same simple straight ahead policy leads to any location with a high value in any of the level P-1 GVFs. Learning involves two steps: (1) learning the mapping from state to GVFs and (2) learning the policy from the GVFs. These steps are fast in environments with natural cardinal directions and strong compositional structure. Learning the mapping from states to the GVFs with TD learning is fast because it involves simple policies which have low entropy in their outcomes and are able to efficiently explore the state space; while learning the mapping from GVFs to policy is greatly simplified by the compositional structure of the GVFs and the simple mapping from the cardinal directions to available actions. In rapidly changing environments, as is typical for the real world, CRL leads to remarkably fast learning. For instance, CRL vastly outperforms traditional approaches in a maze task in which the maze changes frequently, or when learning to reach for an object, whose location varies over trials, with a robotic arm. This work provides a biologically plausible framework to study task decomposition in animals confronted with rapidly changing environments.

Details

show
hide
Language(s):
 Dates: 2022-03
 Publication Status: Published online
 Pages: -
 Publishing info: -
 Table of Contents: -
 Rev. Type: -
 Identifiers: -
 Degree: -

Event

show
hide
Title: Computational and Systems Neuroscience Meeting (COSYNE 2022)
Place of Event: Lisboa, Portugal
Start-/End Date: 2022-03-17 - 2022-03-20

Legal Case

show

Project information

show

Source 1

show
hide
Title: Computational and Systems Neuroscience Meeting (COSYNE 2022)
Source Genre: Proceedings
 Creator(s):
Affiliations:
Publ. Info: -
Pages: - Volume / Issue: - Sequence Number: 2-105 Start / End Page: 169 Identifier: -