The value of confidence: Confidence prediction errors drive value-based learning in the absence of external feedback

Reinforcement learning algorithms have a long-standing success story in explaining the dynamics of instrumental conditioning in humans and other species. While normative reinforcement learning models are critically dependent on external feedback, recent findings in the field of perceptual learning point to a crucial role of internally-generated reinforcement signals based on subjective confidence, when external feedback is not available. Here, we investigated the existence of such confidence-based learning signals in a key domain of reinforcement-based learning: instrumental conditioning. We conducted a value-based decision making experiment which included phases with and without external feedback and in which participants reported their confidence in addition to choices. Behaviorally, we found signatures of self-reinforcement in phases without feedback, reflected in an increase of subjective confidence and choice consistency. To clarify the mechanistic role of confidence in value-based learning, we compared a family of confidence-based learning models with more standard models predicting either no change in value estimates or a devaluation over time when no external reward is provided. We found that confidence-based models indeed outperformed these reference models, whereby the learning signal of the winning model was based on the prediction error between current confidence and a stimulus-unspecific average of previous confidence levels. Interestingly, individuals with more volatile reward-based value updates in the presence of feedback also showed more volatile confidence-based value updates when feedback was not available. Together, our results provide evidence that confidence-based learning signals affect instrumentally learned subjective values in the absence of external feedback.

Download Full-text

The role of reinforcement learning models to assess decision-making in the Iowa Gambling Task under the influence of alcohol

Frontiers in Computational Neuroscience ◽

10.3389/conf.fncom.2012.55.00057 ◽

2012 ◽

Vol 6 ◽

Author(s):

Smolka Michael

Keyword(s):

Decision Making ◽

Reinforcement Learning ◽

Iowa Gambling Task ◽

Gambling Task ◽

Learning Models ◽

Reinforcement Learning Models

Download Full-text

Signed and unsigned reward prediction errors dynamically enhance learning and memory

eLife ◽

10.7554/elife.61077 ◽

2021 ◽

Vol 10 ◽

Author(s):

Nina Rouhani ◽

Yael Niv

Keyword(s):

Reinforcement Learning ◽

Locus Coeruleus ◽

Learning And Memory ◽

Learning Rate ◽

Prediction Errors ◽

Learning Models ◽

The Past ◽

Reward Prediction ◽

Midbrain Dopamine ◽

Reinforcement Learning Models

Memory helps guide behavior, but which experiences from the past are prioritized? Classic models of learning posit that events associated with unpredictable outcomes as well as, paradoxically, predictable outcomes, deploy more attention and learning for those events. Here, we test reinforcement learning and subsequent memory for those events, and treat signed and unsigned reward prediction errors (RPEs), experienced at the reward-predictive cue or reward outcome, as drivers of these two seemingly contradictory signals. By fitting reinforcement learning models to behavior, we find that both RPEs contribute to learning by modulating a dynamically changing learning rate. We further characterize the effects of these RPE signals on memory, and show that both signed and unsigned RPEs enhance memory, in line with midbrain dopamine and locus-coeruleus modulation of hippocampal plasticity, thereby reconciling separate findings in the literature.

Download Full-text