Abstract
Auditory perceptual deficits are widely observed among children with developmental language disorder (DLD). Yet, the nature of these deficits and the extent to which they explain speech and language problems remain controversial. In this study, we hypothesize that disruption to the maturation of the basilar membrane may impede the optimization of the auditory pathway from brainstem to cortex, curtailing high-resolution frequency sensitivity and the efficient spectral decomposition and encoding of natural speech. A series of computational simulations involving deep convolutional neural networks that were trained to encode, recognize, and retrieve naturalistic speech are presented to demonstrate the strength of this account. These neural networks were built on top of biologically truthful inner ear models developed to model human cochlea function, which-in the key innovation of the present study-were scheduled to mature at different rates over time. Delaying cochlea maturation qualitatively replicated the linguistic behavior and neurophysiology of individuals with language learning difficulties in a number of ways, resulting in (a) delayed language acquisition profiles, (b) lower spoken word recognition accuracy, (c) word finding and retrieval difficulties, (d) "fuzzy" and intersecting speech encodings and signatures of immature neural optimization, and (e) emergent working memory and attentional deficits. These simulations illustrate many negative cascading effects that a primary maturational frequency discrimination deficit may have on early language development and generate precise and testable hypotheses for future research into the nature and cost of auditory processing deficits in children with language learning difficulties. (PsycInfo Database Record (c) 2024 APA, all rights reserved).
Attribution and reuse record
- Authors
- Jones SD, Stewart HJ, Westermann G.
- Original journal
- Psychological review
- Publisher
- American Psychological Association
- Publication date
- 2023-07-27
- DOI
- 10.1037/rev0000436
- License
- CC BY 3.0
- Open repository
- Europe PMC · PMC11115354
- Collection
- School leadership launch collection
Presented by the Journal for School Superintendents under the license identified in the article’s open full-text record. The original authors and publisher do not endorse this journal or its agent.
Open full text
Read the scholarly record
From Temporal to Spectral Processing Deficits in Language Disorder Research
A dominant view developed principally through the work of Tallal et al. is that children with language learning difficulties have a primary deficit affecting the perception of acoustic signals that change rapidly, something that these authors refer to as a temporal processing deficit 1 (e.g., Merzenich et al., 1996 ; Tallal et al., 1981 ). Much of the empirical research in this direction made use of the auditory repetition task, or ART, in which children press buttons to identify changes in frequency in a series of pure tones. In the ART, performance accuracy among children with DLD was regularly shown to decrease significantly when interstimulus interval (ISI; i.e., the gap between tones) was reduced to below approximately 250 ms, lending apparent support to the hypothesis that these children’s auditory processing systems were ill-equipped to accurately perceive and encode rapidly unfolding natural speech ( Merzenich et al., 1996 ; Tallal et al., 1981 ). This line of argument has been pursued in a significant body of research and has motivated the development of the Fast ForWord program of intervention, which claims to be able to train sensitivity to rapidly occurring auditory stimuli through the controlled manipulation of ISI and in doing so confer gains in speech and language abilities ( Tallal, 2013 ).
Despite the initial dominance of the temporal processing deficit hypothesis, however, a series of failed replications, both of the basic research and of the Fast ForWord intervention ( Strong et al., 2011 ; Bishop & McArthur, 2005 ; McArthur & Bishop, 2004 ; see Rosen, 2003 , for review), has motivated the search for alternative characterizations of the auditory perceptual deficits that appear to affect many children with speech and language problems. One promising, though comparatively underexamined, view is that such deficits are spectral rather than temporal in nature ( Bishop & McArthur, 2005 ; McArthur & Bishop, 2004 ; Mengler et al., 2005 ). That is, for many children the difficulty relates principally to distinguishing discrete sounds of similar frequency rather than discrete sounds that rapidly follow one another. For instance, across two studies, Bishop and McArthur presented children aged 10–19 with and without language disorder with a baseline tone of 600 Hz and a distinct tone, which was initialized at 700 Hz, but which was raised or lowered by increments of 25 Hz to determine the minimal frequency discrimination threshold, or limen, that participants could identify ( Bishop & McArthur, 2005 ; McArthur & Bishop, 2004 ; see also Mengler et al., 2005 ). These authors found that the minimal frequency discrimination threshold among children with severe language disorder was 750 Hz (i.e., a 150-Hz disparity) during an initial assessment and 674 Hz at follow-up (i.e., a 74-Hz disparity) compared to 629 and 624 Hz disparities, respectively, for control children. Readers may wish to visit one of the many freely available online pure tone generators to compare tones in this range themselves. For many, the average difference between the minimal threshold tones identified by children with DLD (i.e., 600 and 750 or 674 Hz) will appear striking, attesting to the difficulty such a deficit may cause during the analysis of the complex spectral profiles of natural speech ( Nuttall et al., 2018 ; Sumner et al., 2018 ).
Crucially, Bishop and McArthur found that this deficit in frequency discrimination was observed regardless of the rate of stimulus presentation, providing compelling evidence that the auditory processing difficulties of some children affected by language disorder are spectral rather than temporal in nature and perhaps explaining the failed replications of key studies in the temporal processing deficit literature ( Bishop & McArthur, 2005 ; McArthur & Bishop, 2004 ; Mengler et al., 2005 ; Rosen, 2003 ; Strong et al., 2011 ). What is more, even those children with DLD who performed well in the behavioral tone discrimination task nevertheless showed immature waveforms during electroencephalography (EEG) monitoring, providing tentative support for the maturational account that Bishop and McArthur (2005) then offer to explain their findings.
The Maturation of Frequency Discrimination Skills
Bishop and McArthur (2005) explained their results in terms of a disruption to the typical maturation of high-resolution frequency discrimination. In order to situate this account, upon which we intend to elaborate, it is useful to review key research on the early maturation of frequency discrimination skills and the neural basis of these skills. In younger children and infants, probing the maturation of frequency discrimination skills presents a significant challenge. Paradigms such as head turning and high-amplitude sucking have provided mixed results and are open to interpretation, not least that a failure to discriminate tones in such paradigms may be the result of immature motor skills or attention (see Burnham & Mattock, 2014 , for review). In response, some researchers have advocated the use of neuroimaging methods such as EEG and magnetoencephalography when studying frequency discrimination in neonates and infants (e.g., Novitski et al., 2007 ). Despite their own limitations, such neuroimaging methods are often considered to provide an index of neural activity that is relatively independent of motor and attentional factors ( Novitski et al., 2007 ).
Neuroimaging involving neonates and infants corroborates indications from behavioral research of an early maturation in frequency discrimination ability ( Jensen & Neff, 1993 ; Lopez-Poveda, 2014 ; Novitski et al., 2007 ; Shafer et al., 2000 ; Tharpe & Ashmead, 2001 ). This maturation is not uniform. High-frequency tone discrimination is approximately adult-like in apparently typically developing infants by 6 months of age. In contrast, low-frequency discrimination, in the range more regularly associated with speech signals (e.g., 400 Hz), develops more slowly, with continued maturation apparent in children up to ages 7–9 ( Burnham & Mattock, 2014 ; Jensen & Neff, 1993 ). While the empirical data vary somewhat, estimates from the “odd-one-out” paradigm (also known as the “mismatch negativity paradigm”) suggest that newborns can detect a 20% though not a 5% change in frequency in a 250–4,000 Hz window ( Novitski et al., 2007 ; see Burnham & Mattock, 2014 , for review). Such findings support the view that frequency resolution improves considerably from birth through childhood, making it increasingly easy to discriminate competing acoustic signals and thus to perform the complex spectral analysis that accurate and efficient natural speech perception and encoding requires ( Nuttall et al., 2018 ; Sumner et al., 2018 ).
The maturation of frequency discrimination skills reflects changes in neural architecture that, though many important questions remain, are now in a large part reasonably well understood. A key characteristic of the auditory perceptual system upon which speech representation and use is based is its tonotopic structure. That is, throughout the auditory pathway, from the inner ear to the auditory brainstem and on to the auditory cortex, we see selective responsivity to acoustic input of particular frequencies among sensory cells and neurons that constitute the neural basis of frequency resolution and the decomposition of auditory signals, including speech ( Echteler et al., 1989 ; Nuttall et al., 2018 ; Sumner et al., 2018 ). The characteristic “tonotopic” structure of the auditory pathway results predominantly from the physical properties of the basilar membrane, a 35-mm coiled membrane within the inner ear ( Figure 1A ).
The basilar membrane is narrow and firm at its base and as a result of these physical properties, fibers in this basal region vibrate maximally to the high frequencies in auditory input ( Figure 1A ; Sumner et al., 2018 ). The apex of the basilar membrane is, in contrast, wide and relatively slack and as a result, fibers in this apical region vibrate maximally to the low frequencies in auditory input ( Figure 1A ; Sumner et al., 2018 ). For instance, voiceless fricatives such as/ʃ/, which contain relatively high-frequency components, may stimulate basal regions of the membrane, while vowels such as /ɑ:/, which contain low-frequency components, may stimulate apical regions. Upon the basilar membrane sit a single row of approximately 3,500 inner hair cells, which become selectively responsive to specific frequencies—that is, they are “frequency-tuned”—as a result of their position on the basilar membrane ( Sumner et al., 2018 ; Tani et al., 2021 ). In turn, inner hair cells are innervated by spiral ganglion neurons, which project to the cochlear nucleus, with this and subsequent innervation conserving tonotopic sensitivity and resulting in the emergence of frequency sensitive “maps” throughout a complex array of subcortical structures of the auditory brainstem and on to the peripheral auditory cortex. The physical properties of the basilar membrane are, therefore, at the heart of frequency sensitivity and acoustic signal decomposition across the auditory pathway, and this itself underpins accurate and efficient speech processing and encoding ( Burnham & Mattock, 2014 ; Echteler et al., 1989 ; Nuttall et al., 2018 ; Sumner et al., 2018 ; Tani et al., 2021 ). From the third trimester to 6 months of age, structures from the auditory nerve throughout the auditory pathway to the auditory cortex undergo substantial changes in synaptic organization, myelination, and dendritic arborization, and this process of maturation continues through 2 years of age during a typically rich period of language development ( Chonchaiya et al., 2013 ). Work by Chonchaiya et al. (2013) indicates that, by 9 months of age, auditory brainstem responses continuous with relatively mature brainstem organization are predictive of better language outcomes.
Recent research has cast light on how the pre- and postnatal structural development of the basilar membrane underpins the emergence of high-resolution frequency tuning across the auditory-linguistic pathway. Studies using electron microscopy and polarized light microscopy have shown tha
Toward a Maturational Account of Frequency Resolution Deficits and Speech and Language Difficulties
Before stating our hypothesis, let us take stock of the key points reviewed so far: 1 Auditory processing deficits are widespread among children with DLD, and these deficits may be frequency-based rather than temporal in nature.
Evidence that deficits are related to frequency analysis points to specific cellular and neural structures of the auditory pathway. Specifically, the basilar membrane is at the heart of frequency tuning across the auditory pathway, with tonotopic maps emerging throughout the auditory brainstem and cortex predominantly as a result of dynamic adaptation to the structural properties—that is, the mechanical gradient—of the basilar membrane.
The basilar membrane undergoes crucial structural changes during prenatal development, with the fibers from which the membrane is composed increasing in diameter, density, and regularity. This process of maturation is integral to the emergence of tonotopic sensitivity across the auditory pathway.
Our hypothesis is, then, that: Early disruption to the maturation of the physical properties of the basilar membrane, which underpin that membrane’s mechanical gradient (i.e., increases in fiber density, diameter, and linear regularity), may disturb the optimization of the posterior auditory pathway from the brainstem to the cortex, curtailing high-resolution tonotopic sensitivity and contributing to speech and language difficulties in some children.
In what follows, we simulate and monitor the dynamic adaptation of an artificial auditory-linguistic pathway (broadly auditory brainstem to cortex) in response to biologically plausible representations of speech-elicited activation patterns in the developing cochlea, under (a) nondevelopmental, (b) regular, and (c) delayed maturational trajectories. We show how a disruption to the maturation of cochlea microarchitecture may result in the atypical optimization of subsequent neural pathways, qualitatively accounting for several commonly recorded characteristics of atypical human linguistic behavior and neurophysiology, namely, (a) delayed language acquisition profiles (e.g., Norbury et al., 2016 ), (b) spoken word recognition deficits ( Andreu et al., 2012 ; Evans et al., 2018 ; Rispens et al., 2015 ; Velez & Schwartz, 2010 ), (c) word finding or retrieval problems ( Kambanaros et al., 2015 ; Messer & Dockrell, 2006 ), (d) “fuzzy” long-term speech representations ( Claessen et al., 2009 ), (e) atypical neural signatures of auditory signal processing (e.g., Bishop & McArthur, 2005 ), and (f) apparent working memory deficits, attributable, we argue, to the imprecision of activated long-term speech representations ( Henry & Botting, 2017 ; Jones & Westermann, 2022 ).
Methods of Analysis
All postsimulation analyses were conducted in R ( RStudio Team, 2016 ). During training and testing, networks were presented with cochleagrams and in response output probability distributions over their 35-word lexicons. The word assigned the highest probability was taken as a network’s classification and where this corresponded to the true target cochleagram a “hit” was scored. The analysis of our training data involved measuring spoken word classification accuracy by training epoch. At test, we measured classification accuracy and the average maximum probability and probability distribution entropy output when a classification was made. These metrics provide a proxy for a network’s certainty in its classifications. A high probability, low entropy (i.e., low spread) distribution signals high certainty in a judgment, while a low probability, high entropy (i.e., high spread) distribution signals low certainty in a judgment.
We then teased apart item-specific effects, looking for subsets of words on which regular or delayed networks performed better or worse. As part of this analysis into item-specific effects, we ran a Bayesian regression model ( Burkner, 2017 ) in which the percentage of correct classifications per word was predicted by condition (i.e., regular, delayed) in interaction with two relevant independent variables that have generated considerable interest in developmental psycholinguistics: word frequency and word phonological neighborhood density (e.g., Ambridge et al., 2015 ; Jones & Brandt, 2019 ; Rispens et al., 2015 ). Word frequency quantifies how common the word is in the exposure language; here, the speech commands corpus from which training words were randomly sampled. Phonological neighborhood density meanwhile quantifies the average distance, calculated on the basis of phonological transcriptions, between each word and the other 34 words in the training data. Relatively high input frequency is regularly associated with better language learning in children ( Ambridge et al., 2015 ), while high phonological distance (i.e., phonemic dissimilarity) may improve speech classification accuracy among human listeners because potential between-item confusion is lower ( Karimi & Diaz, 2020 ). As our modeling approach did not involve semantic representations, it was not possible to include other variables of potential interest such as word concreteness, valence, or relevance to infants and babies ( Braginsky et al., 2018 ; Jones & Brandt, 2019 ).
Artificial neural networks are sometimes criticized for being inscrutable “black boxes.” Yet, there exist numerous methods that enable the researcher to go beyond performance metrics such as accuracy alone to peer inside the network and understand how it is representing information in the service of completing a certain task. Exploiting such methods is vital to the present study because our interest is in how a processing hierarchy modeling the auditory pathway from brainstem to cortex optimizes in the face of low-level constraints on frequency tuning in the cochlea. Convolutional neural network activation patterns have been shown to align broadly (i.e., not on a layer-to-structure level of granularity) with activation patterns in the biological brain ( Kell et al., 2018 ; cf. Thompson, 2020 ). Furthermore, Bishop and MacArthur’s work in this direction shows that even when there is apparently no group difference in performance metrics such as accuracy, frequency resolution deficits may be associated with different neural signatures across groups with and without language disorder ( Bishop & McArthur, 2005 ; McArthur & Bishop, 2004 ). Similarly, Chonchaiya et al. (2013) showed that auditory brainstem responses continuous with immature brainstem optimization predict relatively poor language outcomes. We wondered whether a similar neural signature of auditory processing impairments within the context of language learning deficits would emerge within our computational framework.
To better understand how our neural networks dynamically optimized to cochlea representations with varying spectral acuity ( Figure 3 ), we used a recently developed framework known as mean field theory-based manifold analysis (MFTMA; Figure 4 ; Chung & Abbott, 2021 ; Chung et al., 2018 ; Cohen et al., 2020 ). Under this approach, each neuron in any given structure of the auditory pathway, for instance the inferior colliculus, is configured as a single axis against which the spiking activity in that neuron can be plotted. Collectively, neurons in a given neural structure then define a neural state space ( Figure 4A ; graphically, a collection of axes) in which patterns of activation can be plotted either as trajectories through time or averaged spikes-per-second vectors. Given neural noise and variability in speaker and communicative context, no two instances of any given speech string stimulate the same response vector within that neural state space, that is, repeated spoken instances of a given linguistic structure never stimulate each neuron in the state space to the same degree. Repeated exposure to a range of exemplars from a single linguistic class, whether phoneme, word, or construction, therefore, stimulates a unified population response known as a “manifold,” which is a quasi-continuous subspace of the neural state space that can be considered the neural basis of the representation of that class ( Cohen et al., 2020 ). Implicitly estimating the bounds of this neural manifold is considered integral to recognizing and producing novel yet valid speech, as if recognizing that instances of this class may regularly stimulate activation patterns within but not substantially outside this region of the state space ( Cohen et al., 2020 ; DiCarlo & Cox, 2007 ; Stephenson et al., 2020 ; Yamins & DiCarlo, 2016 ).
The major contribution of the MFTMA method is enable us to treat distributed biological and artificial neural activation patterns as continuous geometric
Item-Specific Effects
We began our item-specific analyses by computing a by-item accuracy differential, calculated by subtracting the average percentage accurate at test for each word in the delay condition from the average percentage accurate for each word in the regular condition. The result is shown in Figure 6 . Here, a positive value indicates a performance advantage, as a percentage, for the regular network, and a negative value indicates a performance advantage for the delay network. Zero differential indicates no performance difference between conditions with respect to a particular word.
Networks in the regular condition outperformed networks in the delay condition with respect to 24 out of 35 words, sometimes reaching a differential of 24.6% (for the word cat). Networks in the delay condition, in contrast, performed better on eight words, with a maximum differential of −11.11% for the word wow. Clearly, then, error rates vary as a function of the target word. To better understand these effects, we looked at confusion matrices for predictions made during speech classification in each condition. The top 10 most confused words in the regular and delay conditions are presented in Tables 1 and 2 , respectively. These tables show the true word, the total number of misclassifications of that word, the most common misclassification of that word, the number of times that the most common misclassification occurred, and most common misclassification as a proportion of total misclassifications (%).
In many cases, the phonological overlap likely responsible for the misclassification is clear, for instance with respect to tree and three or no and go, and it is noteworthy that networks struggled by some margin with respect to these particular competitor words. Similar patterns are discussed by Karimi and Diaz (2020) who review classification disadvantages for near neighbors under certain experimental conditions. At first glance, then, networks appear to be broadly sensitive to similar spectral features input as human listeners (e.g., struggling with items like tree and three). Yet, Tables 1 and 2 also illustrate examples, which apparently deviate from this pattern, for instance the apparently high rates of misclassification of the word five as the word on or the misclassification of the word house as off. It is difficult to imagine this pattern performance in human participants, and this may attest to the fact that despite the many gross similarities between processing in artificial neural networks and the human brain, artificial neural networks may attend to different features of the input in the service of reducing error in a given task. We return to this point below.
To further understand the above disparities in item accuracy between conditions, we fitted a Bayesian regression model in which test phase accuracy (as a percentage) was predicted by standardized frequency and phonological distance, both in interaction with condition (i.e., regular, delay). We centered on frequency and phonological distance as predictor variables given their importance in the child language literature. However, alternative predictor variables of interest (e.g., orthographic word length) can be experimented with using the Jupyter Notebook and R script associated with this project. Frequency quantified the number of times that a word appeared in the randomly sampled training data. Meanwhile, phonological distance was computed as the mean optimal string alignment distance between a phonological transcription of each target word and of all other words in the speech commands corpus.
A range of diagnostics showed that this simple regression model with a skew normal likelihood and weakly informative priors fitted well (i.e., rhats at 1.0, a large number of effective samples, and credible posterior predictive checks; see the public repository associated with this project at https://osf.io/x2h8k/ and the brms documentation for further details; Burkner, 2017 ). Figure 7 shows the estimates from our Bayesian model.
In Figure 7 , Panels A and B show that across groups, classification accuracy was on average higher for high frequency (β = 2.11; 95% CI [−0.97, 5.45]) and phonologically distinctive (β = 2.82; 95% CI [−0.46, 6.46]) words. While the credible intervals (CIs) associated with these estimates cross zero, indicating that zero may be the true effect, a substantial proportion of probability mass is positively assigned, suggesting that a positive association is likely. Meanwhile, in Figure 7 , Panels C and D show that these effects interact slightly with condition but tend in the same positive direction (see R code for full estimates: https://osf.io/x2h8k/ ). In each case, networks with rapidly maturing high-resolution cochlea models benefitted slightly more from high frequency and greater phonologically distinctiveness.
In summary, item-specific analyses indicate that while networks struggled to different degrees with different words, they nevertheless struggled with broadly similar features of the data set, misclassifying close competitor words such as tree and three most frequently and performing best when words were highly frequent in the training data and phonologically distinctive. Higher resolution low-level auditory representations enabled networks in the regular condition to better exploit these input statistics. These results may be expected given that at any particular period, the regular and delay networks sit at different points on the same developmental trajectory. The resulting performance profiles are in agreement with the general observation that the language of children with DLD is delayed rather than deviant ( Kan & Windsor, 2010 ; see also Discussion section). That is, the language of children with DLD is often similar to that of younger children with typical language skills (though see Bishop, 2013 ). That said, our item-specific analysis also revealed potential discrepancies between artificial neural network per
Visualizing Internal Representations—Mean Field Theory-Based Manifold Analyses
The cochlea models that provide input to the deep convolutional neural networks used in these simulations were scheduled to mature according to one of two developmental time courses. In contrast, the neural networks into which cochleagrams were passed were provided with a randomized initial weight matrix, which was matched across networks and conditions, but which then optimized freely to solve the specific problems of speech encoding, recognition, and retrieval (note that the control network presents an optimal system, which is free to optimize in the absence of any significant low-level constraint). The performance profiles detailed above—specifically the disparities in accuracy, probability, entropy, and item-specific effects—point to systematic differences in dynamic optimization that, given matching across networks, can result only from these low-level maturational constraints on high-resolution frequency tuning. We are, therefore, modeling discrepancies in optimal adaptation in the face of different low-level constraints. But what does optimization in the face of a low-level frequency discrimination deficit look like? To better understand the optimization profiles of networks in our three conditions and therefore to unpick the representational basis of the performance discrepancies seen in networks across these conditions, we turned to mean field theory-based manifold analyses.
Variables of primary interest were (a) manifold dimensionality and (b) classification capacity. Manifold dimensionality quantifies how spread out through a neural state space long-term speech representations are—that is, how many artificial neurons (as a proportion of the layer size) are implicated in the representation of that speech string. Classification capacity quantifies the number of speech manifolds that can be linearly separated from all competitor representations, again standardized by network layer size. Analysis of biological and artificial neural networks suggests that dimensionality decreases across the auditory and visual perceptual systems, and accordingly, that system capacity increases ( Chung & Abbott, 2021 ; Chung et al., 2018 ; DiCarlo & Cox, 2007 ). This transformation reflects the gradual denoising of neural representations in a perceptual hierarchy. Speech representations, for instance, are shown to become decreasingly noise sensitive and increasingly speech selective during transformation from the basilar membrane to the peripheral auditory cortex and beyond ( Davis & Johnsrude, 2003 ; DeWitt & Rauschecker, 2012 ; Kaas et al., 1999 ; Okada et al., 2010 ).
System classification capacity has been interpreted as a measure of not only representation overlap, but also of attention or working memory load, given that calculating classification capacity involves linearly discriminating discrete representations from the system’s “long-term memory” in a manner continuous with cognitive recognition and retrieval ( Jones & Westermann, 2022 ). This view is in line with so-called state-based frameworks in which working memory is understood as activated long-term memory that must be optimized to “fit” within an attentional spotlight ( Adams et al., 2018 ; Oberauer, 2013 , 2019 ). Importantly, reducing manifold dimensionality in order to boost system classification capacity is a product of training in a given task, here speech encoding and classification. Training the same artificial neural network with the same data in a different task, for instance a speaker recognition task, would result in an internal network structure optimized for this task (i.e., activation patterns forming manifolds aligned with speaker voice characteristics; Stephenson et al., 2020 ). The result of this task-specific optimization process is presented in Figure 9 , which shows the average manifold dimensionality and classification capacity in networks’ penultimate layers antecedent to the classifier (see Figure 2 ) as a function of training epoch.
Figure 9 shows a clear disparity in the optimization of internal speech representations across conditions. Over 10 epochs, networks following the regular cochlea maturation schedule increasingly approached control standards of optimization supporting low-dimensional representation ( Figure 9A ). In contrast, despite an overall decrease across epochs, the average dimensionality of internal spoken word representations formed in networks in the delay condition remained significantly higher, that is, these representations were substantially more “spread out” in a relatively poorly optimized neural state space ( Figure 9A ). In Figure 9 , Panel B shows that this inability to optimize efficiently and reduce manifold dimensionality had a severe effect on the delay networks’ ability to retrieve any single representation from their internal “long-term memory” systems—what we interpret here as a form of simulated working memory or attentional capacity deficit. In essence, the delay networks optimized to noise, and this means that the artificial neural response patterns underpinning the long-term representations of different spoken words intersect substantially, making efficient recognition and retrieval difficult. Graphically, it is as though the delay networks remain in the suboptimal state shown in Figure 4C rather than approaching the relatively optimal state shown in Figure 4D alongside networks in the regular and control conditions.
The same representational disparity can be seen posttraining across the networks’ layers. In Figure 10 , we show the previously reported trajectory (e.g., Yamins & DiCarlo, 2016 ) across the auditory processing hierarchy from high-dimensional manifolds in a low-capacity system to low-dimensional manifolds in a high-capacity system. Again, this reflects the system optimizing to render initially noise-sensitive representations (i.e., waveform representations containing speaker effects, etc.) increasingly speech selective (i.e., word type representations in
Discussion
Frequency discrimination deficits are widely recognized among children with language learning difficulties ( Bishop & McArthur, 2005 ; McArthur & Bishop, 2004 ; Mengler et al., 2005 ). Yet, the nature of these deficits and their relation to speech processing problems remain unclear. The neural microarchitecture supporting high-resolution frequency discrimination matures from the prenatal period through to later childhood, and it is possible that the frequency discrimination deficits seen among some children with language learning difficulties stem from a disruption to this typical developmental trajectory ( Bishop & McArthur, 2005 ; McArthur & Bishop, 2004 ). Given that frequency tuning throughout the auditory pathway is predominantly attributable to the structural properties of the basilar membrane (i.e., the membrane’s mechanical gradient, including fiber diameter, density, and regularity; Tani et al., 2021 ), we hypothesized that disruption to the maturation of the structural properties of the basilar membrane may provide a good starting point for inquiry into the source of frequency discrimination deficits in children with neurodevelopmental disorder. Disruption to the structure of the basilar membrane has been demonstrated empirically in animal models manipulating emilin2 expression, which results in a deficient mechanical gradient and therefore suboptimal functioning of the auditory pathway not supporting high-resolution frequency processing ( Amma et al., 2003 ; Russell et al., 2020 ).
We developed this theoretical account through a series of computational simulations of speech encoding, recognition, and retrieval. The networks used in these simulations incorporated inner ear models developed to replicate human cochlea function ( McDermott & Simoncelli, 2011 ) that were fed into deep convolutional neural networks. Despite many important differences, for instance in scale, complexity, and the use of undifferentiated cell types, deep convolutional neural networks have demonstrated significant correspondences with human behavioral and neural responses across a range of tests of audition, including speech localization, pitch perception, and hearing in noise ( Francl & McDermott, 2022 ; Kell et al., 2018 ; Saddler et al., 2021 ). Our own innovation was to configure the cochlea models that formed a fundamental component of our networks to mature according to different developmental trajectories (i.e., baseline or optimal, regular, and delayed) and to analyze how the subsequent auditory-linguistic pathway optimized in the service of speech encoding, recognition, and retrieval.
Our analysis of networks in the delayed cochlea maturation condition qualitatively replicated the linguistic behavior and neurophysiology of individuals with language learning difficulties in a number of ways, showing (a) delayed acquisition profiles ( Norbury et al., 2016 ), (b) lower spoken word recognition accuracy ( Andreu et al., 2012 ; Evans et al., 2018 ; Rispens et al., 2015 ; Velez & Schwartz, 2010 ), (c) word finding and retrieval difficulties and uncertainty even when performing accurately, as evidenced, for instance, in eye tracking paradigms (i.e., Kambanaros et al., 2015 ; McMurray et al., 2019 ; Messer & Dockrell, 2006 ), (d) “fuzzy” long-term speech representations ( Claessen et al., 2009 , 2013 ; Claessen & Leitão, 2012a , 2012b ) and neurophysiological signatures of immature neural optimization that are associated with speech and language difficulties ( Bishop & McArthur, 2005 ; Chonchaiya et al., 2013 ; McArthur & Bishop, 2004 ), and (e) apparent working memory and attention deficits that are attributable, we believe, to the imprecision of long-term speech representations ( Gray et al., 2019 ; Henry & Botting, 2017 ; Jones & Westermann, 2022 ). Our results illustrate that optimizing to low-level, low-resolution spectral representations significantly curtails the capacity of the system to form speech representations supporting efficient recognition and retrieval.
We see, then, that some of the mechanisms widely thought to play a causal role in speech and language disorder may “come for free” if we assume a low-level frequency discrimination deficit. This includes not only the hypothesized working memory capacity bottleneck ( Archibald & Gathercole, 2006 ), which dominates DLD research but which we have argued to be a possible epiphenomenon (see also Jones & Westermann, 2022 ), but also the so-called lateral inhibition deficit suggested by McMurray et al. (2019) . McMurray et al. (2019) argued that a key feature of early language disorder may be an inability to inhibit activated competitor representations during speech recognition and retrieval. Our simulations suggest, however, that an apparent lateral inhibition deficit may be an emergent characteristic of a suboptimal auditory processing hierarchy. Networks in the delayed cochlea maturation condition of our simulations uniformly output predictive distributions with high spread (i.e., high entropy) and low maximum probability assignment, signaling heightened uncertainty and broader activation of the lexicon in response to speech stimuli. As in the case of the hypothesized working memory capacity limitation, then, we believe that evidence offered in support of a deficit in a functionally discrete lateral inhibition mechanism may instead reflect target isolation being overwhelmed due to the imprecision of activated long-term speech representations, a process illustrated in Figure 4C .
It may be argued that the results reported in the present study were inevitable. That is, disrupting the quality of the cochlea representations that networks could form would necessarily lead to worse performance. But this is not the case. Indeed, data disruption, for instance blurring, skewing, recoloring, or clipping the training data, is regularly used in machine learning, where the process is termed “data augmentation”, to boost network performance by preventi
Conclusion
Frequency discrimination is a core problem for many children with language learning difficulties and through computational simulation we have shown how this deficit would propagate problems with the encoding, recognition, and retrieval of natural speech. Our simulations provide proof of concept that the optimization of the auditory pathway to low-resolution cochlea representations—part of a typical maturational trajectory that may be disrupted in DLD—results in patterns of linguistic behavior that align qualitatively with a range of empirical findings observed among children with DLD. Our speculation that the locus of such deficits may be a disruption to the maturation of the basilar membrane during a sensitive period of auditory pathway optimization reflects the fact that the mechanical gradient of the basilar membrane provides the basis for the emergence of frequency sensitivity across the auditory pathway. Yet, this hypothesis of course requires empirical testing. The auditory pathway is a highly complex system, which could be disrupted at any level. Also in need of further scrutiny is our speculation, given the contemporary animal model literature, that atypicalities in emilin2 expression may be implicated in the disruption of the emergence of the mechanical gradient of the basilar membrane (i.e., the development of fibril microarchitecture supporting high-resolution processing, which promulgates the required tonotopic sensitivity through the auditory nerve, brainstem, and cortex). We fully recognize these elements of our argument to be speculation, albeit empirically driven speculation. Our view is simply that the empirical evidence with respect to structural changes in the basilar membrane suggests that this hypothesis constitutes a strong starting point for further inquiry into the nature of auditory processing deficits in children with language learning difficulties.
Figures, tables, references, and supplementary files are best inspected in the licensed PDF or repository copy linked above.