Open APA research · Launch collection CC BY 3.0

Child witness expressions of certainty are informative.

Winsor AA, Flowe HD, Seale-Carlisle TM, Killeen IM, Hett D, Jores T, Ingham M, Lee BP, Stevens LM, Colloff MF.

Journal of experimental psychology. GeneralAmerican Psychological Association2021-09-09DOI 10.1037/xge0001049

Abstract

Children are frequently witnesses of crime. In the witness literature and legal systems, children are often deemed to have unreliable memories. Yet, in the basic developmental literature, young children can monitor their memory. To address these contradictory conclusions, we reanalyzed the confidence-accuracy relationship in basic and applied research. Confidence provided considerable information about memory accuracy, from at least age 8, but possibly younger. We also conducted an experiment where children in young (4-6 years), middle (7-9 years), and late (10-17 years) childhood (N = 2,205) watched a person in a video and then identified that person from a police lineup. Children provided a confidence rating (an explicit judgment) and used an interactive lineup-in which the lineup faces can be rotated-and we analyzed children's viewing behavior (an implicit measure of metacognition). A strong confidence-accuracy relationship was observed from age 10 and an emerging relationship from age 7. A constant likelihood ratio signal-detection model can be used to understand these findings. Moreover, in all ages, interactive viewing behavior differed in children who made correct versus incorrect suspect identifications. Our research reconciles the apparent divide between applied and basic research findings and suggests that the fundamental architecture of metacognition that has previously been evidenced in basic list-learning paradigms also underlies performance on complex applied tasks. Contrary to what is believed by legal practitioners, but similar to what has been found in the basic literature, identifications made by children can be reliable when appropriate metacognitive measures are used to estimate accuracy. (PsycInfo Database Record (c) 2022 APA, all rights reserved).

Attribution and reuse record

Authors
Winsor AA, Flowe HD, Seale-Carlisle TM, Killeen IM, Hett D, Jores T, Ingham M, Lee BP, Stevens LM, Colloff MF.
Original journal
Journal of experimental psychology. General
Publisher
American Psychological Association
Publication date
2021-09-09
DOI
10.1037/xge0001049
License
CC BY 3.0
Open repository
Europe PMC · PMC8721974
Collection
School leadership launch collection

Presented by the Journal for School Superintendents under the license identified in the article’s open full-text record. The original authors and publisher do not endorse this journal or its agent.

Open full text

Read the scholarly record

Identification Parades and Memory Accuracy

When the identity of the culprit is unknown, a child witness may be asked to make an identification from a police identification parade (hereafter, a lineup). There are no official statistics on the number of children aged under 18 who view lineups each year, but given the proportion of children who experience crime there is reason to believe that the number is substantial. Recently, we surveyed 48 police officers from a U.K. metropolitan police force, and they estimated, on average, that 18% of child witnesses attempt to make an identification from a lineup. During a police lineup, the witness is shown images of the police suspect and other individuals who look similar to the suspect and are known to be innocent, called fillers . The police suspect may be innocent or may be guilty (i.e., may or may not be the culprit). It is the job of the witness to identify the culprit if they are present in the lineup or reject the lineup if the culprit is not present.

To determine the likely accuracy of children’s lineup identification decisions, applied research has largely focused on measuring average memory discrimination accuracy —that is, ability to discriminate between innocent and guilty suspects—in children of different age groups. This research suggests that memory discrimination accuracy improves with age ( Humphries & Flowe, 2015 ; Laurence & Mondloch, 2016 ; see Fitzgerald & Price, 2015 for a meta-analysis). There is some discussion about the mechanisms underlying the improvements in memory discrimination accuracy with age on lineup tasks. There has been a long tradition in the eyewitness literature of research concluding that children aged from about 5 years are just as likely as their older peers (and even adults) to make a correct identification of a guilty suspect in a target-present lineup, and that age differences in lineup identifications are limited to older children making fewer mistaken identifications of innocent suspects from target-absent lineups, perhaps due to younger children having difficulty withholding an inappropriate response (e.g., Dunlevy & Cherryman, 2013 ; Havard & Memon, 2013 ; Pozzulo & Lindsay, 1998 ; see also Roebers & Spiess, 2016 ; Schneider & Löffler, 2016 ; for discussion on maturation of monitoring and cognitive control in the basic science literature). Yet, some eyewitness research with children has found that correct identifications of guilty suspects in target-present lineups increase with age, possibly because memory mechanisms gradually mature throughout childhood (for example, Brewer & Day, 2005 ; Fitzgerald et al., 2014 ; Fitzgerald & Price, 2015 ; Keast et al., 2007 ; see also Crookes & McKone, 2009 ; McKone et al., 2012 ; for debate on development of face identification abilities in the basic science literature). Despite ongoing discussion about the underlying mechanisms, research has concluded that average memory discrimination accuracy improves throughout childhood.

Applied research with adult witnesses, however, indicates that average memory accuracy is not the most important metric for determining the reliability of eyewitness identifications. A better metric for legal decision-makers to decide how much trust to place in witness memory evidence is to use metacognitive measures, such as confidence judgments (e.g., Mickes, 2015 ). This is because, regardless of their average memory discrimination accuracy, a person with a reliable memory has good metacognitive ability 1 and is able to appropriately modulate their confidence in response to their memory performance, reporting higher confidence when likely to be correct and lower confidence when not likely to be correct ( Fleming & Lau, 2014 ). Even if memory discrimination accuracy is relatively poor, the reliability of memory evidence can be good, because people can be aware when their memories are inaccurate or accurate (e.g., Brewer & Wells, 2006 ; Sauer et al., 2010 ).

A key question is whether children can monitor their memory accuracy. Answering this question is practically important in determining how child witness memory evidence should be interpreted in legal systems and theoretically important in developing a unified theory of children’s metacognitive development. Currently, contradictory conclusions have been drawn in the applied witness and basic developmental literatures.

Memory Monitoring in Children

In the witness identification literature, the consensus is that children are unreliable witnesses, because their confidence judgments do not reflect their memory accuracy ( Keast et al., 2007 ; Powell et al., 2013 ). For example, one influential study asked children (aged 10–14) to watch a mock-crime video and later identify the culprit and another individual in the video ( Keast et al., 2007 ). The researchers found that the correspondence (called calibration ) between confidence and accuracy was poor and concluded that a child’s confidence provides no useful indicator of a suspect’s innocence or guilt. Similar conclusions have been reached in other research recruiting children who are between the ages of 8 and 11 ( Brewer & Day, 2005 ; Parker & Carranza, 1989 ; Parker & Ryan, 1993 ). Thus, the witness literature suggests that children who are younger than 12 have not yet fully developed the skills to monitor their memory, or to use confidence scales to indicate accuracy ( Powell et al., 2013 ; but see Bruer et al., 2017 for a notable exception). Critically, this conclusion has informed legal guidance worldwide. For example, Powell et al. (2013) state that confidence is not a useful guide to accuracy for children’s identification responses, and this book has been cited by superior courts in every jurisdiction in Australia and New Zealand.

Yet, a more positive picture emerges when the developmental literature is considered. Developmental research suggests that children from about age 4 or 5 can demonstrate memory-monitoring skills, which improve throughout childhood ( Sodian et al., 2012 ). For example, in one study that is representative of the basic literature, children aged 3–5 viewed objects and then subsequently identified which object of two they had seen before and provided a confidence judgment after each decision. Children from age 4 provided higher confidence judgments, on average, for correct answers than incorrect answers ( Hembacher & Ghetti, 2014 ), thereby demonstrating memory-monitoring skills. Moreover, instead of collecting confidence judgments (an explicit metacognitive judgment), other researchers have found that young children from age 3 can appropriately express uncertainty implicitly without full awareness, using gestures like shaking their head, shrugging their shoulders ( Kim et al., 2016 ), or asking for help when they are unsure ( Ghetti et al., 2013 ; Goupil et al., 2016 ). Although the memory task and test format can moderate the accuracy of children’s memory monitoring (e.g., Steiner et al., 2020 ), taken together, the developmental literature suggests that children from age 3 can monitor their performance when implicit measures of metacognition are collected, and that children from age 4 or 5 have developed at least some memory-monitoring skills and the ability to use explicit confidence scales to indicate accuracy.

Why has basic developmental research generally concluded that children’s expressions of certainty can be informative about memory accuracy, whereas applied witness research has concluded the opposite? There are at least three possible reasons. The first reason might be the task itself: Memories from complex witnessed events (e.g., the physical appearance of a culprit) may be more difficult for younger children to monitor, compared with the simple to-be-remembered stimuli (e.g., pictures) that children monitor in the developmental literature ( Harris, 1995 ). A second reason might be that different methods have been used to measure memory-monitoring across the literatures. For example, eyewitness researchers have seldom measured implicit metacognition, such as a child’s behavior during the lineup task, which might be more predictive of accuracy in younger children than explicit confidence judgments (e.g., Ghetti et al., 2013 ; Kim et al., 2016 ). A third reason is differences in statistical approach in analyzing explicit confidence judgments. A common approach in the developmental literature is to calculate average confidence for correct versus incorrect decisions, but this does not provide all of the information relevant to examine memory monitoring, because there could be a poor correspondence between confidence and accuracy, even if confidence is, on average, higher for correct than incorrect decisions. A good correspondence between confidence and accuracy occurs when high-confidence decisions are highly accurate, medium-confidence decisions are moderately accurate, and low-confidence decisions are of low accuracy. Conversely, because legal decision-makers (e.g., judges, jurors) are interested in determining the likelihood of accuracy of a single identification made with a particular level of certainty, eyewitness researchers have measured the typical correspondence between witnesses’ certainty judgments and their average accuracy. Examining the correspondence between certainty and accuracy provides comprehensive information about memory monitoring skills, but the applied literature has used approaches that can underestimate the relationship between confidence and accuracy. We explain this in more detail next.

Measuring the Relationship Between Confidence and Memory Accuracy

The witness identification literature has traditionally relied on statistical techniques which can underestimate the confidence–accuracy relationship. For example, the point biserial correlation coefficient has been used, but we now know that the correlation coefficient can vary dramatically, even when confidence and accuracy are perfectly calibrated, because it is affected by the distribution of correct and incorrect identification decisions across confidence levels ( Juslin et al., 1996 ). Compared with the point biserial correlation coefficient, a better way to assess the relationship between confidence and accuracy is to plot subjective confidence against objective performance (proportion correct) to construct calibration curves and calculate associated calibration statistics (e.g., Over/under confidence, C, Adjusted Normalized Resolution Index). More recent research has used the calibration approach to advance understanding about metacognition (e.g., Keast et al., 2007 ; Palmer et al., 2013 ; Sauer et al., 2010 ). From an applied perspective, however, calibration analyses may also underestimate the informativeness of confidence in criminal justice settings, because it includes filler IDs along with innocent suspect IDs to calculate errors ( Mickes, 2015 ; Wixted & Wells, 2017 ). When legal decision-makers are determining the likely accuracy of a witness’s identification, they are determining the likely accuracy of an identification of a police suspect. This is because only suspect identifications (and not filler identifications) are used as evidence of a suspect’s guilt or innocence in court ( Wixted & Wells, 2017 ). Consequently, instead of calibration analyses, researchers have recently begun to use confidence accuracy characteristic (CAC) analysis to examine the reliability of witness identification decisions.

In a CAC analysis, subjective confidence is plotted against objective performance, but only innocent suspect IDs (and not fillers) are included when calculating errors ( Mickes, 2015 ). Recent research in the adult witness literature using CAC analysis suggests that there is generally a strong relationship between confidence and suspect ID accuracy in adults (e.g., see Wixted & Wells, 2017 , for a review). Confidence typically tracks suspect ID accuracy, even in situations where overall memory discrimination accuracy is comparatively poor, such as in older adults compared with younger adults ( Colloff et al., 2017 ), or in those who experienced a longer delay between encoding and the identification test ( Wixted et al., 2016 ). To explain why confidence typically tracks suspect ID accuracy, even in situations where overall memory discrimination accuracy is comparatively poor, we need to consider theoretical models from basic science. In this regard, a constant likelihood ratio signal-detection model from the broader memory literature has recently been applied to account for adult witness memory performance ( Colloff et al., 2017 ; Semmler et al., 2018 ; Stretch & Wixted, 1998 ).

Constant Likelihood Ratio Signal-Detection Model

The constant likelihood ratio signal-detection model posits that adults “fan out” their confidence criteria across a memory strength continuum in conditions yielding poorer memory discriminability. The idea is that when discrimination accuracy is lower, adults place their most conservative decision criterion (e.g., 100% confidence) at a more conservative location on the memory strength continuum (requiring more memory evidence to make a recognition memory decision with high confidence), while placing their liberal decision criterion (e.g., 10% confidence) at a more liberal location (requiring less memory evidence to make a decision with low confidence). Behaving in this way means that adults place their decision criteria optimally to maintain a constant likelihood of accuracy at each level of confidence over hard (poorer discrimination) and easy (better discrimination) conditions. 2 It has been proposed that adults learn how to place their confidence criteria optimally through a lifetime of error feedback training about the circumstances in which their memories are and are not accurate ( Mickes et al., 2011 ; Stretch & Wixted, 1998 ). The constant likelihood ratio signal-detection model has been applied to account for performance of older adults, showing that they optimally place their criteria to compensate for age-related decline in memory performance ( Colloff et al., 2017 ) and also to show that adults optimally place their criteria to compensate for viewing distance impairments on memory performance ( Semmler et al., 2018 ). As such, theory predicts and data suggest that, at least as adults, eyewitnesses can be reliable; they have metacognitive skills to monitor memory and can usually assign appropriate confidence judgments that reflect their identification accuracy. We considered whether and at what age children optimally place their decision criterion and assign appropriate confidence judgments that correspond to their memory accuracy.

The Current Study

Currently, it is unclear why the basic and applied literatures have reached different conclusions regarding the informativeness of children’s expressions of certainty. It is important to reexamine memory-monitoring in children for both basic and applied researchers. First, for basic researchers, theories should account for monitoring performance across task domains. If it is the case that children can monitor their memory on a complex eyewitness identification task, and show a strong correspondence between certainty and accuracy, this suggests that the fundamental architecture of metacognition that has previously been evidenced in the developmental literature on relatively simple tasks also underlies performance on complex tasks. Conversely, if children do not have a good metacognitive awareness on a complex task, this suggests that the ability to monitor accuracy is dependent on the cognitive activity, or complexity of the memory, being monitored ( Ghetti et al., 2013 ). Second, for applied researchers, the correspondence between certainty and accuracy (i.e., the reliability of children’s identification decisions) may currently be underestimated in legal systems worldwide, because young children are able to monitor their memories according to studies in the basic developmental literature; and the most appropriate statistical techniques have not been used. Theoretically, a constant likelihood ratio signal-detection model predicts that people optimally adjust their criterion, and the correspondence between certainty and accuracy will improve with age, as the quantity of memory error-feedback training increases.

In this study, we first use CAC analysis to reanalyze children’s explicit confidence judgments in basic list-learning memory studies and an influential eyewitness identification study that sampled children in late childhood ( Keast et al., 2007 ). The data (both basic and applied) show a strong relationship between confidence and accuracy in children. We then present an original eyewitness study in which we asked more than 2,220 children in young (aged 4–6), middle (aged 7–9), and late (aged 10–17) childhood to watch a video of a complex event, then attempt to identify the person who was in the video from a police lineup, and provide a confidence judgment (explicit measure of metacognition). We used a novel interactive lineup—in which the lineup faces can be rotated and viewed from different angles—to record children’s viewing behavior moment by moment and explore whether viewing behavior (implicit measure of metacognition) differs in children who made correct versus incorrect identifications. Again, contrary to what is believed to be true in legal systems around the world, but consistent with the basic literature, we show that children’s expressions of certainty are informative even on a complex memory task.

Wilkinson et al. (2010)

Wilkinson et al. (2010) compared typically developing children between 9 and 17 years old with children with Autism Spectrum Disorder (ASD) of the same age using an old/new face recognition paradigm. They also compared adults (18 to 45 years) with and without ASD. During the learning phase, participants viewed 24 female faces sequentially. During the testing phase, participants were shown 48 faces, 24 old (i.e., shown in the learning phase) and 24 new (i.e., not shown in the learning phase). The faces were presented one at a time and participants decided if each face was old or new and rated their confidence ( guessing , somewhat certain , or certain for adults; and guessing , somewhat sure , or sure for children). We plotted proportion correct as a function of accuracy. Figure 1B illustrates that the typically developing children showed a strong confidence–accuracy relationship, while the children with ASD showed no relationship. Typically developing children were 25%, 60%, and 82% accurate at low, medium, and high confidence. Of note is that the CAC for the typically developing children mirrors that of the typically developed adults, although children were less accurate at low and medium confidence. In both typically developing children and adults, however, high-confidence decisions were likely to be accurate (85% and 82% correct in adults and children, respectively).

Hiller and Weber (2013)

Hiller and Weber (2013) tested 8- to 12-year-old children and adults (18–59 years) using an associative word-pair recognition paradigm. Participants were shown 28 word-pairs sequentially in the encoding phase and, after a two-minute delay, were given a memory test also consisting of 28 word-pairs. Participants had to recognize each word pair in the test as old or new, and after each decision rate their confidence using a confidence scale that ranged from 50 ( guessing ) to 100 ( certain ). Hiller and Weber plotted predicted log odds of a recognition decision being correct or incorrect as a function of confidence. We converted the predicted log odds to proportion correct and again focused on old decisions. Unsurprisingly, Figure 1C indicates that adults showed a strong confidence–accuracy relationship and were more accurate at each level of confidence than the children. Children’s accuracy was approximately 30% correct for low-confidence decisions, and accuracy increased monotonically with confidence, up to 86% correct for high-confidence decisions. Thus, the confidence–accuracy relationship was strong for both adults and children.

Hembacher and Ghetti (2014)

Hembacher and Ghetti (2014) tested uncertainty monitoring in 3- to 5-year-old children using a two-alternative forced-choice object recognition task. During the learning phase the children viewed 30 drawings of common objects. During the test phase, children decided which of two drawings they had seen in the learning phase and made a confidence judgment on a 3-point picture scale. Each point on the confidence scale was an illustration of a child displaying a facial and body expression, indicating either low, moderate, or high confidence. We obtained the data for this study through the Open Science Framework and plotted proportion correct as a function of confidence. Figure 1D illustrates that 3-year-olds show virtually no confidence–accuracy relationship, with both low-confidence and high-confidence responses resulting in similar levels of overall accuracy (both around 84%). Four-year-olds show a moderate confidence–accuracy relationship, with low-confidence responses being 68% correct and high-confidence responses being 86% correct. Five-year-olds showed a strong confidence–accuracy relationship with low-confidence responses being 69% correct and high-confidence responses being approximately 93% correct. These findings echo Hembacher and Ghetti’s conclusions about the developmental trajectory of uncertainty monitoring in their original analysis, whereby 3-year-olds were unable to monitor their own uncertainty, 5-year-olds were able to monitor their uncertainty, and 4-year-olds fell somewhere in between.

Figures, tables, references, and supplementary files are best inspected in the licensed PDF or repository copy linked above.

Open Paper Agent