Distributional reinforcement learning in prefrontal cortex.


Journal

Nature neuroscience
ISSN: 1546-1726
Titre abrégé: Nat Neurosci
Pays: United States
ID NLM: 9809671

Informations de publication

Date de publication:
10 Jan 2024
Historique:
received: 03 08 2022
accepted: 29 11 2023
medline: 11 1 2024
pubmed: 11 1 2024
entrez: 10 1 2024
Statut: aheadofprint

Résumé

The prefrontal cortex is crucial for learning and decision-making. Classic reinforcement learning (RL) theories center on learning the expectation of potential rewarding outcomes and explain a wealth of neural data in the prefrontal cortex. Distributional RL, on the other hand, learns the full distribution of rewarding outcomes and better explains dopamine responses. In the present study, we show that distributional RL also better explains macaque anterior cingulate cortex neuronal responses, suggesting that it is a common mechanism for reward-guided learning.

Identifiants

pubmed: 38200183
doi: 10.1038/s41593-023-01535-w
pii: 10.1038/s41593-023-01535-w
doi:

Types de publication

Journal Article

Langues

eng

Sous-ensembles de citation

IM

Subventions

Organisme : Leverhulme Trust
ID : DS-2017-026
Organisme : James S. McDonnell Foundation (McDonnell Foundation)
ID : JSMF220020372
Organisme : Wellcome Trust (Wellcome)
ID : 220296/Z/20/Z

Informations de copyright

© 2024. The Author(s).

Références

Walton, M. E., Behrens, T. E. J., Buckley, M. J., Rudebeck, P. H. & Rushworth, M. F. S. Separable learning systems in the macaque brain and the role of orbitofrontal cortex in contingent learning. Neuron 65, 927–939 (2010).
doi: 10.1016/j.neuron.2010.02.027 pubmed: 20346766 pmcid: 3566584
Kennerley, S. W., Walton, M. E., Behrens, T. E. J., Buckley, M. J. & Rushworth, M. F. S. Optimal decision making and the anterior cingulate cortex. Nat. Neurosci. 9, 940–947 (2006).
doi: 10.1038/nn1724 pubmed: 16783368
Bechara, A., Damasio, A. R., Damasio, H. & Anderson, S. W. Insensitivity to future consequences following damage to human prefrontal cortex. Cognition 50, 7–15 (1994).
doi: 10.1016/0010-0277(94)90018-3 pubmed: 8039375
Rudebeck, P. H. et al. Frontal cortex subregions play distinct roles in choices between actions and stimuli. J. Neurosci. 28, 13775–13785 (2008).
doi: 10.1523/JNEUROSCI.3541-08.2008 pubmed: 19091968 pmcid: 6671924
Fellows, L. K. & Farah, M. J. Different underlying impairments in decision-making following ventromedial and dorsolateral frontal lobe damage in humans. Cereb. Cortex 15, 58–63 (2005).
doi: 10.1093/cercor/bhh108 pubmed: 15217900
Fellows, L. K. & Farah, M. J. Ventromedial frontal cortex mediates affective shifting in humans: evidence from a reversal learning paradigm. Brain 126, 1830–1837 (2003).
doi: 10.1093/brain/awg180 pubmed: 12821528
Sutton, R. & Barto, A. Reinforcement Learning: An introduction (MIT, 1998).
Kennerley, S. W., Behrens, T. E. J. & Wallis, J. D. Double dissociation of value computations in orbitofrontal and anterior cingulate neurons. Nat. Neurosci. 14, 1581–1589 (2011).
doi: 10.1038/nn.2961 pubmed: 22037498 pmcid: 3225689
Rushworth, M. F. S., Noonan, M. A. P., Boorman, E. D., Walton, M. E. & Behrens, T. E. Frontal cortex and reward-guided learning and decision-making. Neuron 70, 1054–1069 (2011).
doi: 10.1016/j.neuron.2011.05.014 pubmed: 21689594
Rescorla, R. A. & Wagner, A. R. A theory of Pavlovian conditioning: Variations in the effectiveness of reinforcement and nonreinforcement. In Classical Conditioning II. Current Research and Theory 64–99 (Appleton-Century-Crofts, 1972).
Rigotti, M. et al. The importance of mixed selectivity in complex cognitive tasks. Nature 497, 585–590 (2013).
doi: 10.1038/nature12160 pubmed: 23685452 pmcid: 4412347
Wallis, J. D. & Kennerley, S. W. Heterogeneous reward signals in prefrontal cortex. Curr. Opin. Neurobiol. 20, 191–198 (2010).
doi: 10.1016/j.conb.2010.02.009 pubmed: 20303739 pmcid: 2862852
Dabney, W., Rowland, M., Bellemare, M. G. & Brain, G. Distributional reinforcement learning with quantile regression. In Proc. of the AAAI Conference on Artificial Intelligence Vol. 32, No. 1 (2018); https://doi.org/10.1609/aaai.v32i1.11791
Bellemare, M. G., Dabney, W. & Munos, R. A distributional perspective on reinforcement learning. In Proc. of the 34th International Conference on Machine Learning 70, 449–458 (PMLR, 2017).
Dabney, W. et al. A distributional code for value in dopamine-based reinforcement learning. Nature 577, 671–675 (2020).
doi: 10.1038/s41586-019-1924-6 pubmed: 31942076 pmcid: 7476215
Schultz, W., Dayan, P. & Montague, P. R. A neural substrate of prediction and reward. Science 275, 1593–1599 (1997).
doi: 10.1126/science.275.5306.1593 pubmed: 9054347
Kolling, N., Wittmann, M. & Rushworth, M. F. S. Multiple neural mechanisms of decision making and their competition under changing risk pressure. Neuron 81, 1190–1202 (2014).
doi: 10.1016/j.neuron.2014.01.033 pubmed: 24607236 pmcid: 3988955
Padoa-Schioppa, C. Neurobiology of economic choice: a good-based model. Annu. Rev. Neurosci. 34, 333–359 (2011).
doi: 10.1146/annurev-neuro-061010-113648 pubmed: 21456961 pmcid: 3273993
Hunt, L. T. et al. Triple dissociation of attention and decision computations across prefrontal cortex. Nat. Neurosci. 21, 1471–1481 (2018).
doi: 10.1038/s41593-018-0239-5 pubmed: 30258238 pmcid: 6331040
Matsumoto, M., Matsumoto, K., Abe, H. & Tanaka, K. Medial prefrontal cell activity signaling prediction errors of action values. Nat. Neurosci. 10, 647–656 (2007).
doi: 10.1038/nn1890 pubmed: 17450137
Bernacchia, A., Seo, H., Lee, D. & Wang, X. J. A reservoir of time constants for memory traces in cortical neurons. Nat. Neurosci. 14, 366–372 (2011).
doi: 10.1038/nn.2752 pubmed: 21317906 pmcid: 3079398
Cavanagh, S. E., Wallis, J. D., Kennerley, S. W. & Hunt, L. T. Autocorrelation structure at rest predicts value correlates of single neurons during reward-guided choice. eLife 5, 1–17 (2016).
doi: 10.7554/eLife.18937
Meder, D. et al. Simultaneous representation of a spectrum of dynamically changing value estimates during decision making. Nat. Commun. 8, 1942 (2017).
doi: 10.1038/s41467-017-02169-w pubmed: 29208968 pmcid: 5717172
Berger, B., Trottier, S., Verney, C., Gaspar, P. & Alvarez, C. Regional and laminar distribution of the dopamine and serotonin innervation in the macaque cerebral cortex: a radioautographic study. J. Comp. Neurol. 273, 99–119 (1988).
doi: 10.1002/cne.902730109 pubmed: 3209731
Williams, M. S. & Goldman-Rakic, P. S. Widespread origin of the primate mesofrontal dopamine system. Cereb. Cortex 8, 321–345 (1998).
doi: 10.1093/cercor/8.4.321 pubmed: 9651129
Haber, S. N. & Knutson, B. The reward circuit: linking primate anatomy and human imaging. Neuropsychopharmacology 35, 4–26 (2010).
doi: 10.1038/npp.2009.129 pubmed: 19812543
Louie, K. Asymmetric and adaptive reward coding via normalized reinforcement learning. PLoS Comput. Biol. 18, 1–15 (2022).
doi: 10.1371/journal.pcbi.1010350
Tano Retamales, P. E., Dayab, P. & Pouget, A. A local temporal difference code for distributional reinforcement learning. Adv. Neural Inf. Process. Syst. 33, 1–12 (2020).
Daw, N. D., Gershman, S. J., Seymour, B., Dayan, P. & Dolan, R. J. Model-based influences on humans’ choices and striatal prediction errors. Neuron 69, 1204–1215 (2011).
doi: 10.1016/j.neuron.2011.02.027 pubmed: 21435563 pmcid: 3077926
Miranda, B., Nishantha Malalasekera, W. M., Behrens, T. E., Dayan, P. & Kennerley, S. W. Combined model-free and model-sensitive reinforcement learning in non-human primates. PLoS Comput. Biol. 16, 1–25 (2020).
doi: 10.1371/journal.pcbi.1007944
Bayer, H. M. & Glimcher, P. W. Midbrain dopamine neurons encode a quantitative reward prediction error signal. Neuron 47, 129–141 (2005).
doi: 10.1016/j.neuron.2005.05.020 pubmed: 15996553 pmcid: 1564381
Caraco, T. Energy budgets, risk and foraging preferences in dark-eyed juncos (Junco hyemalis). Behav. Ecol. Sociobiol. 8, 213–217 (1981).
Mante, V., Sussillo, D., Shenoy, K. V. & Newsome, W. T. Context-dependent computation by recurrent dynamics in prefrontal cortex. Nature 503, 78–84 (2013).
doi: 10.1038/nature12742 pubmed: 24201281 pmcid: 4121670
Wang, J. X. et al. Prefrontal cortex as a meta-reinforcement learning system. Nat. Neurosci. 21, 860–868 (2018).
doi: 10.1038/s41593-018-0147-8 pubmed: 29760527
Kennerley, S. W., Dahmubed, A. F., Lara, A. H. & Wallis, J. D. Neurons in the frontal lobe encode the value of multiple decision variables. J. Cogn. Neurosci. 21, 1162–1178 (2009).
doi: 10.1162/jocn.2009.21100 pubmed: 18752411 pmcid: 2715848
Behrens, T. E. J., Woolrich, M. W., Walton, M. E. & Rushworth, M. F. S. Learning the value of information in an uncertain world. Nat. Neurosci. 10, 1214–1221 (2007).
doi: 10.1038/nn1954 pubmed: 17676057

Auteurs

Timothy H Muller (TH)

Department of Experimental Psychology, University of Oxford, Oxford, UK. timothymuller127@gmail.com.
Department of Clinical and Movement Neurosciences, University College London, London, UK. timothymuller127@gmail.com.

James L Butler (JL)

Department of Experimental Psychology, University of Oxford, Oxford, UK.
Department of Clinical and Movement Neurosciences, University College London, London, UK.

Sebastijan Veselic (S)

Department of Experimental Psychology, University of Oxford, Oxford, UK.
Department of Clinical and Movement Neurosciences, University College London, London, UK.
Wellcome Trust Centre for Human Neuroimaging, University College London, London, UK.

Bruno Miranda (B)

Department of Clinical and Movement Neurosciences, University College London, London, UK.
Institute of Physiology and Institute of Molecular Medicine, Lisbon School of Medicine, University of Lisbon, Lisbon, Portugal.

Joni D Wallis (JD)

Department of Psychology and Helen Wills Neuroscience Institute, University of California Berkeley, Berkeley, CA, USA.

Peter Dayan (P)

Max Planck Institute for Biological Cybernetics, Tübingen, Germany.
University of Tübingen, Tübingen, Germany.

Timothy E J Behrens (TEJ)

Wellcome Trust Centre for Human Neuroimaging, University College London, London, UK.
Wellcome Centre for Integrative Neuroimaging, University of Oxford, John Radcliffe Hospital, Oxford, UK.
Sainsbury Wellcome Centre for Neural Circuits and Behaviour, University College London, London, UK.

Zeb Kurth-Nelson (Z)

Google DeepMind, London, UK. zebkurthnelson@gmail.com.
Max Planck University College London Centre for Computational Psychiatry and Ageing Research, University College London, London, UK. zebkurthnelson@gmail.com.

Steven W Kennerley (SW)

Department of Experimental Psychology, University of Oxford, Oxford, UK. steven.kennerley@psy.ox.ac.uk.
Department of Clinical and Movement Neurosciences, University College London, London, UK. steven.kennerley@psy.ox.ac.uk.
Wellcome Centre for Integrative Neuroimaging, University of Oxford, John Radcliffe Hospital, Oxford, UK. steven.kennerley@psy.ox.ac.uk.

Classifications MeSH