Distributional reinforcement learning in prefrontal cortex.
Journal
Nature neuroscience
ISSN: 1546-1726
Titre abrégé: Nat Neurosci
Pays: United States
ID NLM: 9809671
Informations de publication
Date de publication:
10 Jan 2024
10 Jan 2024
Historique:
received:
03
08
2022
accepted:
29
11
2023
medline:
11
1
2024
pubmed:
11
1
2024
entrez:
10
1
2024
Statut:
aheadofprint
Résumé
The prefrontal cortex is crucial for learning and decision-making. Classic reinforcement learning (RL) theories center on learning the expectation of potential rewarding outcomes and explain a wealth of neural data in the prefrontal cortex. Distributional RL, on the other hand, learns the full distribution of rewarding outcomes and better explains dopamine responses. In the present study, we show that distributional RL also better explains macaque anterior cingulate cortex neuronal responses, suggesting that it is a common mechanism for reward-guided learning.
Identifiants
pubmed: 38200183
doi: 10.1038/s41593-023-01535-w
pii: 10.1038/s41593-023-01535-w
doi:
Types de publication
Journal Article
Langues
eng
Sous-ensembles de citation
IM
Subventions
Organisme : Leverhulme Trust
ID : DS-2017-026
Organisme : James S. McDonnell Foundation (McDonnell Foundation)
ID : JSMF220020372
Organisme : Wellcome Trust (Wellcome)
ID : 220296/Z/20/Z
Informations de copyright
© 2024. The Author(s).
Références
Walton, M. E., Behrens, T. E. J., Buckley, M. J., Rudebeck, P. H. & Rushworth, M. F. S. Separable learning systems in the macaque brain and the role of orbitofrontal cortex in contingent learning. Neuron 65, 927–939 (2010).
doi: 10.1016/j.neuron.2010.02.027
pubmed: 20346766
pmcid: 3566584
Kennerley, S. W., Walton, M. E., Behrens, T. E. J., Buckley, M. J. & Rushworth, M. F. S. Optimal decision making and the anterior cingulate cortex. Nat. Neurosci. 9, 940–947 (2006).
doi: 10.1038/nn1724
pubmed: 16783368
Bechara, A., Damasio, A. R., Damasio, H. & Anderson, S. W. Insensitivity to future consequences following damage to human prefrontal cortex. Cognition 50, 7–15 (1994).
doi: 10.1016/0010-0277(94)90018-3
pubmed: 8039375
Rudebeck, P. H. et al. Frontal cortex subregions play distinct roles in choices between actions and stimuli. J. Neurosci. 28, 13775–13785 (2008).
doi: 10.1523/JNEUROSCI.3541-08.2008
pubmed: 19091968
pmcid: 6671924
Fellows, L. K. & Farah, M. J. Different underlying impairments in decision-making following ventromedial and dorsolateral frontal lobe damage in humans. Cereb. Cortex 15, 58–63 (2005).
doi: 10.1093/cercor/bhh108
pubmed: 15217900
Fellows, L. K. & Farah, M. J. Ventromedial frontal cortex mediates affective shifting in humans: evidence from a reversal learning paradigm. Brain 126, 1830–1837 (2003).
doi: 10.1093/brain/awg180
pubmed: 12821528
Sutton, R. & Barto, A. Reinforcement Learning: An introduction (MIT, 1998).
Kennerley, S. W., Behrens, T. E. J. & Wallis, J. D. Double dissociation of value computations in orbitofrontal and anterior cingulate neurons. Nat. Neurosci. 14, 1581–1589 (2011).
doi: 10.1038/nn.2961
pubmed: 22037498
pmcid: 3225689
Rushworth, M. F. S., Noonan, M. A. P., Boorman, E. D., Walton, M. E. & Behrens, T. E. Frontal cortex and reward-guided learning and decision-making. Neuron 70, 1054–1069 (2011).
doi: 10.1016/j.neuron.2011.05.014
pubmed: 21689594
Rescorla, R. A. & Wagner, A. R. A theory of Pavlovian conditioning: Variations in the effectiveness of reinforcement and nonreinforcement. In Classical Conditioning II. Current Research and Theory 64–99 (Appleton-Century-Crofts, 1972).
Rigotti, M. et al. The importance of mixed selectivity in complex cognitive tasks. Nature 497, 585–590 (2013).
doi: 10.1038/nature12160
pubmed: 23685452
pmcid: 4412347
Wallis, J. D. & Kennerley, S. W. Heterogeneous reward signals in prefrontal cortex. Curr. Opin. Neurobiol. 20, 191–198 (2010).
doi: 10.1016/j.conb.2010.02.009
pubmed: 20303739
pmcid: 2862852
Dabney, W., Rowland, M., Bellemare, M. G. & Brain, G. Distributional reinforcement learning with quantile regression. In Proc. of the AAAI Conference on Artificial Intelligence Vol. 32, No. 1 (2018); https://doi.org/10.1609/aaai.v32i1.11791
Bellemare, M. G., Dabney, W. & Munos, R. A distributional perspective on reinforcement learning. In Proc. of the 34th International Conference on Machine Learning 70, 449–458 (PMLR, 2017).
Dabney, W. et al. A distributional code for value in dopamine-based reinforcement learning. Nature 577, 671–675 (2020).
doi: 10.1038/s41586-019-1924-6
pubmed: 31942076
pmcid: 7476215
Schultz, W., Dayan, P. & Montague, P. R. A neural substrate of prediction and reward. Science 275, 1593–1599 (1997).
doi: 10.1126/science.275.5306.1593
pubmed: 9054347
Kolling, N., Wittmann, M. & Rushworth, M. F. S. Multiple neural mechanisms of decision making and their competition under changing risk pressure. Neuron 81, 1190–1202 (2014).
doi: 10.1016/j.neuron.2014.01.033
pubmed: 24607236
pmcid: 3988955
Padoa-Schioppa, C. Neurobiology of economic choice: a good-based model. Annu. Rev. Neurosci. 34, 333–359 (2011).
doi: 10.1146/annurev-neuro-061010-113648
pubmed: 21456961
pmcid: 3273993
Hunt, L. T. et al. Triple dissociation of attention and decision computations across prefrontal cortex. Nat. Neurosci. 21, 1471–1481 (2018).
doi: 10.1038/s41593-018-0239-5
pubmed: 30258238
pmcid: 6331040
Matsumoto, M., Matsumoto, K., Abe, H. & Tanaka, K. Medial prefrontal cell activity signaling prediction errors of action values. Nat. Neurosci. 10, 647–656 (2007).
doi: 10.1038/nn1890
pubmed: 17450137
Bernacchia, A., Seo, H., Lee, D. & Wang, X. J. A reservoir of time constants for memory traces in cortical neurons. Nat. Neurosci. 14, 366–372 (2011).
doi: 10.1038/nn.2752
pubmed: 21317906
pmcid: 3079398
Cavanagh, S. E., Wallis, J. D., Kennerley, S. W. & Hunt, L. T. Autocorrelation structure at rest predicts value correlates of single neurons during reward-guided choice. eLife 5, 1–17 (2016).
doi: 10.7554/eLife.18937
Meder, D. et al. Simultaneous representation of a spectrum of dynamically changing value estimates during decision making. Nat. Commun. 8, 1942 (2017).
doi: 10.1038/s41467-017-02169-w
pubmed: 29208968
pmcid: 5717172
Berger, B., Trottier, S., Verney, C., Gaspar, P. & Alvarez, C. Regional and laminar distribution of the dopamine and serotonin innervation in the macaque cerebral cortex: a radioautographic study. J. Comp. Neurol. 273, 99–119 (1988).
doi: 10.1002/cne.902730109
pubmed: 3209731
Williams, M. S. & Goldman-Rakic, P. S. Widespread origin of the primate mesofrontal dopamine system. Cereb. Cortex 8, 321–345 (1998).
doi: 10.1093/cercor/8.4.321
pubmed: 9651129
Haber, S. N. & Knutson, B. The reward circuit: linking primate anatomy and human imaging. Neuropsychopharmacology 35, 4–26 (2010).
doi: 10.1038/npp.2009.129
pubmed: 19812543
Louie, K. Asymmetric and adaptive reward coding via normalized reinforcement learning. PLoS Comput. Biol. 18, 1–15 (2022).
doi: 10.1371/journal.pcbi.1010350
Tano Retamales, P. E., Dayab, P. & Pouget, A. A local temporal difference code for distributional reinforcement learning. Adv. Neural Inf. Process. Syst. 33, 1–12 (2020).
Daw, N. D., Gershman, S. J., Seymour, B., Dayan, P. & Dolan, R. J. Model-based influences on humans’ choices and striatal prediction errors. Neuron 69, 1204–1215 (2011).
doi: 10.1016/j.neuron.2011.02.027
pubmed: 21435563
pmcid: 3077926
Miranda, B., Nishantha Malalasekera, W. M., Behrens, T. E., Dayan, P. & Kennerley, S. W. Combined model-free and model-sensitive reinforcement learning in non-human primates. PLoS Comput. Biol. 16, 1–25 (2020).
doi: 10.1371/journal.pcbi.1007944
Bayer, H. M. & Glimcher, P. W. Midbrain dopamine neurons encode a quantitative reward prediction error signal. Neuron 47, 129–141 (2005).
doi: 10.1016/j.neuron.2005.05.020
pubmed: 15996553
pmcid: 1564381
Caraco, T. Energy budgets, risk and foraging preferences in dark-eyed juncos (Junco hyemalis). Behav. Ecol. Sociobiol. 8, 213–217 (1981).
Mante, V., Sussillo, D., Shenoy, K. V. & Newsome, W. T. Context-dependent computation by recurrent dynamics in prefrontal cortex. Nature 503, 78–84 (2013).
doi: 10.1038/nature12742
pubmed: 24201281
pmcid: 4121670
Wang, J. X. et al. Prefrontal cortex as a meta-reinforcement learning system. Nat. Neurosci. 21, 860–868 (2018).
doi: 10.1038/s41593-018-0147-8
pubmed: 29760527
Kennerley, S. W., Dahmubed, A. F., Lara, A. H. & Wallis, J. D. Neurons in the frontal lobe encode the value of multiple decision variables. J. Cogn. Neurosci. 21, 1162–1178 (2009).
doi: 10.1162/jocn.2009.21100
pubmed: 18752411
pmcid: 2715848
Behrens, T. E. J., Woolrich, M. W., Walton, M. E. & Rushworth, M. F. S. Learning the value of information in an uncertain world. Nat. Neurosci. 10, 1214–1221 (2007).
doi: 10.1038/nn1954
pubmed: 17676057