Open APA research · Launch collection CC BY 3.0

Domain-general enhancements of metacognitive ability through adaptive training.

Carpenter J, Sherman MT, Kievit RA, Seth AK, Lau H, Fleming SM.

Journal of experimental psychology. GeneralAmerican Psychological Association2019-01-01DOI 10.1037/xge0000505

Abstract

The metacognitive ability to introspect about self-performance varies substantially across individuals. Given that effective monitoring of performance is deemed important for effective behavioral control, intervening to improve metacognition may have widespread benefits, for example in educational and clinical settings. However, it is unknown whether and how metacognition can be systematically improved through training independently of task performance, or whether metacognitive improvements generalize across different task domains. Across 8 sessions, here we provided feedback to two groups of participants in a perceptual discrimination task: an experimental group (n = 29) received feedback on their metacognitive judgments, while an active control group (n = 32) received feedback on their decision performance only. Relative to the control group, adaptive training led to increases in metacognitive calibration (as assessed by Brier scores), which generalized both to untrained stimuli and an untrained task (recognition memory). Leveraging signal detection modeling we found that metacognitive improvements were driven both by changes in metacognitive efficiency (meta-d'/d') and confidence level, and that later increases in metacognitive efficiency were positively mediated by earlier shifts in confidence. Our results reveal a striking malleability of introspection and indicate the potential for a domain-general enhancement of metacognitive abilities. (PsycINFO Database Record (c) 2018 APA, all rights reserved).

Attribution and reuse record

Authors
Carpenter J, Sherman MT, Kievit RA, Seth AK, Lau H, Fleming SM.
Original journal
Journal of experimental psychology. General
Publisher
American Psychological Association
Publication date
2019-01-01
DOI
10.1037/xge0000505
License
CC BY 3.0
Open repository
Europe PMC · PMC6390881
Collection
School leadership launch collection

Presented by the Journal for School Superintendents under the license identified in the article’s open full-text record. The original authors and publisher do not endorse this journal or its agent.

Open full text

Read the scholarly record

Procedure

The experiment was divided into three phases: Phase 1, pretraining (one session) → Phase 2, training (eight sessions) → Phase 3, posttraining (one session), resulting in 10 sessions in total. Figure 1B provides an overview of the experiment timeline. Phase 1 consisted of stimulus titration and a pretraining session to evaluate baseline metacognitive accuracy in a series of two-alternative forced-choice (2AFC) discrimination tasks (see the Task section and Figure 1A ). One set of tasks assessed perceptual discrimination, the other set assessed recognition memory. The tasks followed a 2 × 2 factorial design crossing cognitive domain (perception or memory) with stimulus type (explained in detail subsequently). Each task consisted of 108 trials, giving 432 total trials in the pretraining session. The order of these tasks was counterbalanced such that each participant performed both tasks in one domain followed by both tasks in the other domain, and within each domain the order of stimulus types was also counterbalanced.

Task and session structure. Panel A: Participants were tested on both a perceptual discrimination and recognition memory task, each involving two stimulus types: abstract shapes and words. The perceptual task (left) comprised a two-alternative forced-choice discrimination judgment as to the brighter of two simultaneously presented stimuli on each trial. The memory task (right) comprised an encoding phase followed by a series of two-alternative forced-choice recognition memory judgments. Panel B: Experiment timeline. Each participant completed 10 sessions in total: a pretraining session, eight training sessions, and a posttraining session. All four conditions were assessed at pre- and posttraining, but only the perceptual task with a single stimulus type (shapes or words) was trained during Sessions 2 through 9. During training sessions, the control groups received feedback on their objective perceptual discrimination performance, whereas the experimental groups received feedback on their metacognitive calibration. In both groups, feedback was delivered every 27 trials.

At the start of Phase 2 participants were assigned to one of four groups. Each group formed a cell in a 2 × 2 factorial design crossing feedback type (control group vs. experimental group) and trained stimulus type (see the Training section to follow). All participants received training on the perceptual task only, with the recognition memory task introduced again at posttraining to assess transfer to a different task domain. During the training phase, each of the eight sessions consisted of 270 trials (2,160 trials total), and block-wise feedback was administered every 27 trials (see the Feedback section to follow).

Phase 3, the final posttraining session, was identical to the pretraining session Phase 1 except that stimulus titration was omitted. Task order was counterbalanced against that used in pretraining, such that each participant performed the task domains (memory, perception) in the opposite order to that seen in pretraining. The order of stimulus types within each domain remained the same.

Phase 1 lasted approximately 60 min, the eight training sessions in Phase 2 lasted approximately 25 min each, and Phase 3 lasted approximately 45 min. Participants were required to wait a minimum of 24 hr between each session and were asked via e-mail to complete each subsequent session within 48 hr to 72 hr of the previous session.

Tasks

Figure 1A displays example trial timelines for the perception and memory tasks. In the perception task, participants were presented with two images (i.e., words or shapes) and asked to respond to the following question: “Which [image] has brighter lines?” In the memory task, participants were first presented with a series of images to memorize (again, words or shapes). On each subsequent trial, one old image and one novel image were presented with the instruction to respond to the following question: “Which [image] have you seen before?” In all tasks, after each decision, participants were asked to rate their confidence on a four-point scale, whereon 1 = very low confidence , 2 = low confidence , 3 = high confidence , and 4 = very high confidence .

In the pretraining session, before beginning each task, participants completed three practice trials to become acquainted with making perception/memory judgments and using the confidence rating scale. Following the practice trials, we probed knowledge of how to perform the perception/memory judgments by asking the following comprehension question: “In the perception/memory task, how do you decide which image to choose?” The three response options were “which one you remember,” “which has more lines,” and “which is brighter.” If a participant answered either question incorrectly, they were excluded from further participation and offered a partial reimbursement determined by the proportion of the session completed. There were no practice trials or comprehension questions in the posttraining session.

Training

The second phase of the study involved eight training sessions of 270 trials each (2,160 trials in total), spread over 8 to 34 days. Participants were randomly allocated to one of four groups in a 2 × 2 factorial design crossing feedback type (control group vs. experimental group) and trained stimulus type (shapes or words). All groups received block-wise feedback in the form of reward (points) every 27 trials. The control groups (for both stimulus types) received feedback on their objective perceptual discrimination performance; the experimental groups (for both stimulus types) received feedback on their metacognitive calibration, as determined by the average quadratic scoring rule (QSR) score. The QSR provides a metric for how closely confidence ratings track accuracy ( Staël von Holstein, 1970 ) and is equal to one minus the Brier score ( Fleming & Lau, 2014 ). The rule underpinning each feedback type is described in more detail in the following Feedback section.

To ensure that each group fully understood how points could be earned, instructions were provided on the meaning of the feedback schedule. Participants completed eight demonstration trials which explained how earnings changed based on their objective performance (control group) or the correspondence between confidence and accuracy (experimental group). After the demonstration, participants performed 10 practice trials in which they received full feedback and a brief explanation. Note that in the demonstration and practice trials, feedback was calculated on a trial-by-trial basis and therefore differed from the block-wise feedback received in the training sessions (see the following Feedback section). After the demonstration and practice trials, participants were asked two comprehension questions probing their understanding of how to earn points. If they failed these questions they were asked to attempt them again until they were successful.

Task Performance Titration

Throughout the entire 10-session experiment, the performance of each participant was titrated online to achieve approximately 75% correct for all tasks except the memory-words task. This “threshold” level of percent correct produces sufficient trials for each signal detection theory outcome (hits, misses, false alarms and correct rejections) for analysis of d ′ and meta- d ′ ( Maniscalco & Lau, 2012 ), and ensured any changes in metacognitive sensitivity were not confounded by shifts in task performance.

Titration was accomplished in different ways for each task. In the perception tasks (for both word and shapes), we implemented two interleaved, weighted and transformed staircase procedures on the brightness of the images. We alternated two staircases with differently weighted step sizes. In the first staircase, after two consecutive correct responses the stimulus brightness was decreased by two steps; after one incorrect response the brightness was increased by four steps. In the second staircase, after three correct responses the brightness level was decreased by three steps, after 1 incorrect response the brightness was increased by four steps. Note that these are not traditional n -down/one-up procedures as the correct trial counter was not reset to zero after each pair or triplet of correct responses. However, we found in pilot work that this interleaved method stably converges to 75% correct. Brightness levels were adjusted independently for word and shape stimuli. In order to define initial brightness levels, participants performed a 60-trial titration block for each stimulus type after the practice trials and before beginning the pretraining session. The final brightness level at the end of the titration block acted as the initial brightness level for pretraining Session 1. Each subsequent Session 2 to 10 began with the final brightness level of the previous session.

In the memory-shapes task, the number of stimuli in the encoding period was adjusted based on the average percent correct recorded over the previous two blocks. If average performance exceeded 75% correct, one additional image was added to the encoding set. If performance dropped below 70% correct, one image was removed, down to a minimum of two images. We initialized the encoding set size at four images. Note that even though the minimum set size was two, the underlying staircase value had no minimum value.

For the memory-words task, we employed a fixed set size of 54 words. This larger set size was based on initial pilot data and the procedure of McCurdy et al. (2013) and reflects the fact that participants typically find encoding and remembering individual words significantly easier than encoding and remembering abstract shapes.

Feedback

Feedback in the form of points was given based on task performance in the control group and metacognitive calibration in the experimental group. We rewarded the control group on their achieved difficulty level, specified as the inverse distance between the current brightness level and the minimum brightness of 128:

where brightness level ϵ [128 − 256] → difficulty level ϵ [0 − 128]. We chose difficulty level instead of accuracy as the relevant performance measure because accuracy was titrated to ∼75% correct in each block.

We rewarded the experimental group using the QSR. The QSR is a proper scoring rule in the formal sense that maximum points are obtained by jointly maximizing the accuracy of choices and confidence ratings ( Staël von Holstein, 1970 ). We mapped each confidence rating onto a subjective probability correct using a linear transformation: p (correct)= −1/3 + confidence/3, where the confidence rating is ϵ [1 − 4] → p (correct) ϵ [0 − 1]. On each trial, i the QSR score is then obtained as follows:

where accuracy is ϵ [0, 1] and p (correct) ϵ [0 − 1] → QSR ϵ [0 − 1]. This rule ensures that people receive the highest number of points when they are highly confident and right, or unconfident and wrong (i.e., metacognitively accurate).

Despite feedback in each group being based on different variables, we endeavored to equate the distribution of points across groups. We used data from an initial pilot study (without feedback) to obtain distributions of expected difficulty level and QSR scores. We then calculated the average difficulty level/QSR score for each block, and fit Gaussian cumulative density functions (CDFs) to these distributions of scores. These CDFs were then used to transform a given difficulty or QSR score in the main experiment to a given number of points.

Compensation

Participants were compensated at approximately $4 per hr, plus a possible bonus on each session. Base pay for the 60-min pretraining session was $4, for the eight 25-min training sessions $2 each, and for the 45-min posttraining session $3. Participants were informed they had the opportunity to earn a session bonus if they outperformed a randomly chosen other participant on that session. In practice, bonuses were distributed pseudorandomly to ensure equivalent financial motivation irrespective of performance. All participants received in the range of four to seven bonuses throughout the course of the 10-session study. Bonuses comprised an additional 70% of the base payment received on any given session.

In addition to the pseudorandom bonuses, all participants received a $3 bonus for completing half (5) of the sessions and a $6 bonus for completing all (10) of the sessions. Total earnings ranged from $37.60 to $43.90 across participants, and income did not differ significantly between groups (control group: M = $41.47; experimental group: M = $40.98; t [59] = 0.94, p = .35). The base payment was paid immediately after completing each session and accumulated bonuses were paid only if the participant completed the full 10 session experiment.

Quantifying Metacognition

Our summary measure of metacognitive calibration was the QSR score achieved by participants before and after training. To separately assess effects of training on metacognitive bias (i.e., confidence level) and efficiency (i.e., the degree to which confidence discriminates between correct and incorrect trials), we also fitted meta- d ′ to the confidence rating data. The meta- d ′ model provides a bias-free method for evaluating metacognitive efficiency in a signal detection theory framework. Specifically, the ratio meta- d ′/ d ′ quantifies the degree to which confidence ratings discriminate between correct and incorrect trials while controlling for first-order performance ( d ′). Using this ratio as a measure of metacognition effectively eliminates performance and response bias confounds typically affecting other measures ( Barrett, Dienes, & Seth, 2013 ; Fleming & Lau, 2014 ). We conducted statistical analyses on log(meta- d ′/ d ′) as a logarithmic scale is appropriate for a ratio measure, giving equal weight to increases and decreases relative to the optimal value of meta- d ′/ d ′ = 1.

Meta- d ′ was fit to each participant’s confidence rating data on a per-session basis using maximum likelihood estimation as implemented in freely available MATLAB code ( http://www.columbia.edu/~bsm2105/Type2sdt/ ). Metacognitive bias was assessed as the average confidence level across a particular task and session, irrespective of correctness.

Analysis Plan

By employing a combination of frequentist and Bayesian statistics, we aimed to assess the differential impact of the training manipulation across groups and the transfer of training effects across domains. To model the dynamics of training, we additionally assessed the drivers of the training effect using latent change score modeling and mediation analysis.

We first applied mixed-effects analyses of variance (ANOVAs) to measures of metacognition including group as a between-subjects factor and task domain as a within-subjects factor. Complementary to classical ANOVAs, we also used a Bayesian “analysis of effects” that quantifies evidence in support of transfer of training effects across stimulus types and domains. Evidence in support of transfer is indicated by a simpler model, without stimulus or domain interaction terms, providing a better fit to the data. Finally, by modeling our data using latent changes scores, we gained insight into whether effects of training are dependent on baseline metacognitive abilities. In addition, we used mediation modeling to ask whether early shifts in confidence strategy facilitated later improvements in introspective ability.

In addition to the pretraining exclusion criteria detailed in the preceding text, the following set of predefined exclusion criteria was applied after data collection was complete. One participant was excluded for performing outside the range of 55% to 95% correct in at least one condition/session. One participant was excluded due to their average difficulty level calculated across all sessions dropping below 2.5 standard deviations below the group mean difficulty level. Five participants were excluded for reporting the same confidence level on 95% of trials for three or more sessions. Finally, trials in which either the participant did not respond in time (response times >2,000 ms) or response times were less than 200 ms were omitted from further analysis (0.98% of all trials).

To evaluate effects of training, we compared data from the pre- and posttraining sessions using mixed-model ANOVAs in JASP ( https://jasp-stats.org/ ) to assess the presence of training effects as a function of domain and stimulus type (factors: Training × Domain × Stimulus × Group). We coded the stimulus factor in terms of whether the stimulus encountered during the pre- and posttraining sessions was trained or untrained. We also used a Bayesian “analysis of effects” in JASP to quantify evidence for and against across-stimulus and across-domain transfer of training effects on confidence and metacognitive efficiency ( Rouder, Morey, Speckman, & Province, 2012 ).

Figures, tables, references, and supplementary files are best inspected in the licensed PDF or repository copy linked above.

Open Paper Agent