Open APA research · Launch collection CC BY 3.0

Occipitotemporal representations reflect individual differences in conceptual knowledge.

Braunlich K, Love BC.

Journal of experimental psychology. GeneralAmerican Psychological Association2018-11-01DOI 10.1037/xge0000501

Abstract

Through selective attention, decision-makers can learn to ignore behaviorally irrelevant stimulus dimensions. This can improve learning and increase the perceptual discriminability of relevant stimulus information. Across cognitive models of categorization, this is typically accomplished through the inclusion of attentional parameters, which provide information about the importance assigned to each stimulus dimension by each participant. The effect of these parameters on psychological representation is often described geometrically, such that perceptual differences over relevant psychological dimensions are accentuated (or stretched), and differences over irrelevant dimensions are down-weighted (or compressed). In sensory and association cortex, representations of stimulus features are known to covary with their behavioral relevance. Although this implies that neural representational space might closely resemble that hypothesized by formal categorization theory, to date, attentional effects in the brain have been demonstrated through powerful experimental manipulations (e.g., contrasts between relevant and irrelevant features). This approach sidesteps the role of idiosyncratic conceptual knowledge in guiding attention to useful information sources. To bridge this divide, we used formal categorization models, which were fit to behavioral data, to make inferences about the concepts and strategies used by individual participants during decision-making. We found that when greater attentional weight was devoted to a particular visual feature (e.g., "color"), its value (e.g., "red") was more accurately decoded from occipitotemporal cortex. We also found that this effect was sufficiently sensitive to reflect individual differences in conceptual knowledge, indicating that occipitotemporal stimulus representations are embedded within a space closely resembling that formalized by classic categorization theory. (PsycINFO Database Record (c) 2019 APA, all rights reserved).

Attribution and reuse record

Authors
Braunlich K, Love BC.
Original journal
Journal of experimental psychology. General
Publisher
American Psychological Association
Publication date
2018-11-01
DOI
10.1037/xge0000501
License
CC BY 3.0
Open repository
Europe PMC · PMC6586152
Collection
School leadership launch collection

Presented by the Journal for School Superintendents under the license identified in the article’s open full-text record. The original authors and publisher do not endorse this journal or its agent.

Open full text

Read the scholarly record

The SHJ Dataset

In the second dataset ( Mack et al., 2016 ), 23 right-handed participants (11 female, mean age = 22.3 years) categorized images of insects ( Figure 2B ) varying along three binary dimensions (legs: thick vs. thin, antennae: thick vs. thin, and mandible: pincer vs. shovel). We excluded data from two participants who each had corrupted data on one run. This resulted in 21 participants for the final analyses. During scanning, participants learned to categorize the stimuli according to the Types I, II, and VI problems described by Shepard et al. (1961 ; Table 2 ). In the Type I problem, the optimal strategy required attending to a single stimulus dimension (e.g., “legs”) that perfectly predicted the category label, while ignoring the other two dimensions. In the Type II problem, the optimal strategy was a logical XOR rule, in which two stimulus features had to be considered together. In the Type VI problem, all stimulus features were relevant to the decision, and participants had to learn the mapping between individual stimuli and the category label. To maximally differentiate endogenous and exogenous factors, the irrelevant feature in the Type II rule was used as a relevant feature of the Type I problem for each participant.

Each problem was performed across four scanner runs. Although all of the participants learned to perform the Type VI problem first, the order of the Types I and II problems was then counterbalanced across participants. Each trial consisted of a 3.5 s stimulus presentation period, a jittered 0.5–4.5 second fixation period, and feedback. Feedback was presented for 2 s and consisted of an image of the presented insect, as well as text indicating whether the response was correct or incorrect. Each trial was separated by jittered intertrial interval (4–8 s), which consisting of a fixation cross. Each run included four presentations of each of the eight stimuli.

For consistency across data sets, we used the group-derived region of interest (ROI) used in “5/4” dataset ( Figure 3B ) and performed a similar analysis. As participants in the SHJ experiment learned to perform the Types I, II, and VI problems during scanning, we mirrored the strategy used by the original authors, and divided the scanning sessions into early (first two runs of each problem) and late learning epochs (last two runs of each problem). We investigated the relationship between occipitotemporal representation and attention only during this late learning phase, in which behavior had largely stabilized.

(A) For the “5/4” dataset, a searchlight analysis indicated that binary perceptual dimensions could be decoded from widespread visual regions (including occipital, temporal and parietal cortex), right inferior frontal sulcus, and left postcentral motor cortex (the familywise error rate was controlled at the voxel-level p < .001). (B) To isolate voxels most strongly representing the stimulus features, we raised the statistical threshold, resulting in the region of interest (ROI) illustrated in yellow. (C) “5/4” dataset binary feature decoding. Red dots indicate scores from individual participants. (D) SHJ binary feature decoding. The same ROI (B) was used in both data sets.

Whole-brain images were acquired using a 3T Siemens Skyra Scanner. Anatomical images were collected using a T1-weighted MPRAGE sequence (TR = 1.9 s, TE = 2.43 ms, 256 × 256 matrix, 1 mm isotropic voxels, flip angle = 9°, field of vision [FOV] = 256 mm). Functional images were acquired using a T2*-weighted multiband (multiband factor = 3) accelerated Echo-planar imaging (EPI) sequence (TR = 2 s, TE = 31ms, flip angle = 73°, FOV = 220 mm, 128 × 128 matrix, 1.7 mm slice thickness, 1.7 mm isotropic voxels).

SUSTAIN was initialized with no clusters, and with equivalent weights assigned to each stimulus dimension. Its learning parameters were first fit to the learning performance of each participant using a maximum-likelihood genetic algorithm procedure. The model was fit in such a way that, after learning one problem, the resultant model state was used as the initial state for the subsequent problem. In this way, the model was fit under the assumption that learning of one task would influence later behavior. Once the learning parameters of the model were optimized, they were fixed, and the attentional parameters were extracted from the second two runs of each task (in which learning had largely stabilized). This yielded distinct sets of attentional parameters for each participant and each task. More information about the model can be found in Appendix B .

Image Processing

Preprocessing included motion correction, and coregistration of the anatomical images to the mean of the functional images (using statistical parametric mapping [SPM], Version 6470). All MVPA analyses were performed in native space without smoothing. For group-level analyses, the statistical maps from each participant were warped to Montreal Neurological Institute (MNI) atlas space using Advanced Normalization Tools (ANTs; Avants, Tustison, & Johnson, 2009 ), and then smoothed with a 6 mm full-width at half maximum Gaussian kernel. The ROI derived from group-level analyses were transformed back into each participants’ native space for ROI-level analyses. We performed MVPA on the unsmoothed, single-trial, t -statistic images ( Misaki, Kim, Bandettini, & Kriegeskorte, 2010 ) derived from the least-squares separate procedure (LSS; Mumford, Turner, Ashby, & Poldrack, 2012 ). We used SPM to estimate the LSS images for the “5/4” dataset but used the NiPy python package ( http://nipy.org/nipy/index.html ) for the SHJ dataset, as it tends to run more efficiently, and this study used a multiband sequence with smaller voxel dimensions.

SHJ dataset

First, we confirmed that each stimulus feature could be decoded significantly above chance from the ROI illustrated in Figure 3B . Using a fourfold, leave-one-run out cross-validation strategy, we used a linear support vector classifier (C = 1) to decode each visual feature across all runs (including both early and late learning epochs), retaining only estimates for the last two runs (which corresponded to the late-learning phase in which behavior had largely stabilized). This fourfold cross-validation strategy yielded better decoding accuracy than a twofold approach based on only the last two runs. This improvement reflects the increased amount of training data available in the fourfold approach and suggests that the multivariate patterns reflecting the individual visual features were stable across learning. Each feature could be decoded at rates significantly above chance ( Figure 3D ; antennae: M = 0.57, t (20) = 3.82, p = .001; mandibles: M = 0.56, t (20) = 3.22, p = .004; legs: M = 0.58, t (20) = 4.17, p < .001).

Next, we investigated whether the decoding accuracy associated with the features covaried with SUSTAIN’s attentional parameters. To do so, we used a mixed-effects linear regression analysis to predict decoding accuracy from attention weight, visual dimension, run and rule. As described in the Methods section, distinct attentional weights were derived for each subject and each rule. The decoding accuracy for each separate run was included in the analysis. The model included fixed-effects parameters for these four variables, and random-effects parameters for the intercept, attention weight, and run (which were free to vary by participant). This allowed us to control for differences in decoding accuracy across visual dimensions and participants (as with the model used for the “5/4” dataset), while additionally controlling for effects of rule and idiosyncratic differences in behavioral performance during the last two runs. Mirroring the findings from the “5/4” dataset, we found that the decoding accuracy of these patterns positively covaried with the attention parameters derived from SUSTAIN ( b = 0.09, 95% CI [0.004, 0.17], SE = 0.04), t (61) = 2.13, p = .038.

To investigate the sensitivity of occipitotemporal feature representations to individual differences in SUSTAIN’s attentional parameters, we conducted a permutation test similar to that described above (i.e., for the “5/4” experiment). This involved shuffling the attentional weight parameters between Participants 10,000 times (preserving the correspondence for both rule and abstract feature). This means that the attentional weight derived from the behavior of one participant, for one particular rule and one particular category feature, was assigned to the same rule and feature, but to a different participant. The slope parameter associated with the unpermuted data ( b = 0.09) was significantly greater than those composing the permuted null distribution (P = .979), suggesting that the visual feature representations were sensitive to idiosyncratic differences in attentional weights. A repeated measures ANOVA indicated that the perceptual dimensions did not influence the attentional parameters, F (2, 44) = 1.27, p = .291. A Bayesian repeated measures ANOVA additionally indicated that the null model was 1.98 times more likely than the alternative hypothesis, providing evidence that the attentional weights were not influenced by visual properties of the stimulus features.

Conclusions

Category training is known to induce changes in both perceptual ( Folstein, Gauthier, & Palmeri, 2012 ; Goldstone, 1994 ; Goldstone, Steyvers, & Larimer, 1996 ; Gureckis & Goldstone, 2008 ; Op de Beeck et al., 2003 ) and neural sensitivity (e.g., Dieciuc, Roque, & Folstein, 2017 ; Folstein et al., 2013 ; Folstein, Palmeri, Gauthier, & Van Gulick, 2015 ; Li et al., 2007 ; Sigala & Logothetis, 2002 ). In two data sets, we demonstrate that occipitotemporal stimulus representations covary with the attentional parameters derived from formal categorization theory. This effect was sufficiently sensitive to reflect individual differences in conceptual knowledge, which implies that these occipitotemporal representations are embedded within a space closely resembling that predicted by formal categorization theory (e.g., Kruschke, 1992 ; Love et al., 2004 ; Nosofsky, 1986 ).

By linking brain and behavior through the latent attentional parameters of cognitive models, we also link two (somewhat) disparate literatures. In the neuroscience literature, effects of selective attention are typically examined using highly structured decision problems, and selective attention is investigated by contrasting different aspects of the experimental design (i.e., relevant vs. irrelevant stimulus dimensions). In the cognitive categorization literature, researchers have focused on developing models that accurately account for behavioral patterns of generalization across different goals and tasks. Our results indicate that these cognitive models can be used to examine effects of selective attention in the brain. This is the case, even for ill-defined decision problems (such as the “5/4” task), as the models are able to successfully account for individual differences in conceptual knowledge.

Context

Brad Love has a longstanding interest in models of categorization. He developed the SUSTAIN model ( Love et al., 2004 ) used here and subsequently became interested in how to theoretically relate such models to the brain ( Love & Gureckis, 2007 ). Later, he used category learning models in model-based fMRI analyses, such as in the two papers from which this contribution draws its data ( Mack et al., 2013 , 2016 ). Through several papers, Kurt Braunlich has investigated neurobiological mechanisms associated with categorization and generalization. Recently ( Braunlich, Liu, & Seger, 2017 ), he found that occipitotemporal category representations are highly flexible, in that they are sensitive to transient generalization demands (i.e., strict vs. lax decision criteria). This dovetails with the present work, which examines attentional effects associated with task demands.

GCM

In the GCM ( Nosofsky, 1987 ), the psychological distance, d between stimuli i and j can be calculated as the attentionally-weighted sum of their unsigned differences across dimensions, k :

where w indicate the attentional parameters assigned to each dimension. The r parameter is set to 1 (city-block distance) for perceptually separable stimulus dimensions (as in the “5/4” dataset), and r is set to 2 (Euclidean distance) for integral dimensions ( Garner, 1976 ). Similarity is an exponentially-decaying function of psychological distance:

where the shape of the similarity gradient is influenced by the sensitivity parameter, c . The probability of choosing Category A, given stimulus, i , is given by the choice rule:

where γ governs the degree of deterministic responding.

SUSTAIN

SUSTAIN is a semisupervised clustering model, which incrementally learns to solve categorization problems by first applying simple solutions, and then increasing complexity as required. Through experience, the model can learn to group similar items into common clusters and can make inferences about novel stimuli based on its perceptual similarity to existing clusters (i.e., based on perceptual similarity, clusters compete to predict latent stimulus attributes). When unexpected feedback is received, the model can also learn in a supervised fashion by creating a new cluster to represent the novel stimulus.

In SUSTAIN, all clusters contain receptive fields (RF’s) for each stimulus dimension. As new stimuli are added to the cluster, the model learns by adjusting the position of each RF to best match the cluster’s expectation for novel stimuli. As the RF is an exponential function, a cluster’s activation, α, decreases exponentially with distance from its preferred value:

where μ represents the distance of the stimulus dimension value from the cluster’s preferred stimulus dimension value, and where λ represents the tuning (or width) of the RF. The λ parameters are specific to dimensions, but are shared across dimensions, and so, like the attentional parameters in the GCM, the λ parameters in SUSTAIN modulate the influence of each stimulus dimension on the overall decision outcome.

The overall activation of a cluster, H , involves consideration of each dimension, k :

where the γ parameter (which is always non-negative) modulates the influence of the λ parameters on the choice outcome. When γ is large, attended dimensions (which are associated with large λ values, and narrow RFs), dominate the activation function (Eq. 5); when γ is zero, the λ parameters are ignored, and all dimensions exert equal influence on the choice.

SUSTAIN was fit to the SHJ dataset in a supervised fashion, using the same trial order experienced by the participants; it was also fit across rule-switches, such that learning from one task was carried over to the next. Thus, SUSTAIN was capable of reflecting learning, as well as carry-over effects associated with previously learned rules.

Footnotes

In their paper, Mack et al. (2016) focused on effects associated with the Type I and Type II rules.

This involves moving an imaginary sphere throughout the brain; repeatedly investigating how well the voxels within the sphere can decode a variable of interest.

The C parameter modulates the penalty associated with training error. With large values, the classifier will choose a small-margin hyperplane, and training accuracy will be high. With smaller values, out-of-sample performance is often improved, but more training samples may be misclassified. C=1 is a common default setting for fMRI.

This provides a more conservative test than the likelihood ratio test or the Wald approximation ( Luke, 2016 ).

Figures, tables, references, and supplementary files are best inspected in the licensed PDF or repository copy linked above.

Open Paper Agent