Abstract
Contemporary models of categorization typically tend to sidestep the problem of how information is initially encoded during decision making. Instead, a focus of this work has been to investigate how, through selective attention, stimulus representations are "contorted" such that behaviorally relevant dimensions are accentuated (or "stretched"), and the representations of irrelevant dimensions are ignored (or "compressed"). In high-dimensional real-world environments, it is computationally infeasible to sample all available information, and human decision makers selectively sample information from sources expected to provide relevant information. To address these and other shortcomings, we develop an active sampling model, Sampling Emergent Attention (SEA), which sequentially and strategically samples information sources until the expected cost of information exceeds the expected benefit. The model specifies the interplay of two components, one involved in determining the expected utility of different information sources and the other in representing knowledge and beliefs about the environment. These two components interact such that knowledge of the world guides information sampling, and what is sampled updates knowledge. Like human decision makers, the model displays strategic sampling behavior, such as terminating information search when sufficient information has been sampled and adaptively adjusting the search path in response to previously sampled information. The model also shows human-like failure modes. For example, when information exploitation is prioritized over exploration, the bidirectional influences between information sampling and learning can lead to the development of beliefs that systematically differ from reality. (PsycInfo Database Record (c) 2022 APA, all rights reserved).
Attribution and reuse record
- Authors
- Braunlich K, Love BC.
- Original journal
- Psychological review
- Publisher
- American Psychological Association
- Publication date
- 2021-07-19
- DOI
- 10.1037/rev0000287
- License
- CC BY 3.0
- Open repository
- Europe PMC · PMC8766620
- Collection
- School leadership launch collection
Presented by the Journal for School Superintendents under the license identified in the article’s open full-text record. The original authors and publisher do not endorse this journal or its agent.
Open full text
Read the scholarly record
Optimal Experiment Design and Human Information Sampling
Several groups have used the principles of optimal experimental design (OED; Fedorov, 1972 , 2010 ; MacKay, 1992 ) to investigate whether humans strategically sample information to test specific hypotheses. Although the calculations underlying OED can be computationally prohibitive for cognitively limited human decision makers, these studies indicate that, despite being susceptible to perceptual ( Itti et al., 1998 ; Yamada & Cottrell, 1995 ; Zhang et al., 2008 ) and decisional ( Klayman, 1995 ) biases, we are often able to select information samples that resolve uncertainty about specific hypotheses. This effect is apparent both during the performance of traditional categorization tasks ( Markant et al., 2015 ; Markant & Gureckis, 2014 ), and during naturalistic behavior. Preschool children, for instance, spontaneously conduct “experiments” to test specific causal hypotheses about the objects they are playing with ( Cook et al., 2011 ). Hypothesis-dependent sampling strategies have also been identified through study of human eye movements. During categorization, for instance, we tend to selectively fixate on stimulus locations that resolve uncertainty about the potential category label ( Nelson & Cottrell, 2007 ; Yang et al., 2016 ). During visual search, we similarly tend to fixate on locations expected to maximize information about the target location ( Najemnik & Geisler, 2005 ).
To select useful information sources to sample, a decision maker must be able to simulate future events. This capacity for preposterior analysis 6 involves predicting the probability and utility of future states. When diagnosing a patient, for instance, doctors must have sufficient knowledge of human pathology to identify plausible diagnoses. They must also be able to use this knowledge to select medical tests that efficiently differentiate between the most probable diagnoses. To reflect the fact that some some results can be more informative than others, 7 full preposterior analysis aggregates information about both the probability and usefulness of each expected result. In practice, this forward-search process can be computationally prohibitive for large problems, necessitating an approximation to the full preposterior search performed by SEA.
What Is a “Useful” Question?
A number of different sampling norms have been used to define the usefulness of sampling a particular dimension (see Nelson, 2005 ). Disinterested sampling norms seek to maximize decision accuracy. One way to define the usefulness of a particular medical test, for instance, would be the degree to which it is expected to improve the probability of making a correct diagnosis. 8 In contrast, situation-specific sampling norms maximize reward rather than accuracy, and may be preferable when payoffs are asymmetric (i.e., when the maximization of accuracy differs from the maximization of reward; Meder & Nelson, 2012 ). For example, incorrectly diagnosing a malignant tumor as benign can be more costly than incorrectly diagnosing a benign tumor as malignant.
Utility-sensitive decision makers should also consider the costs associated with sampling each information source. Invasive medical tests (e.g., a biopsy), for example, can be more informative than noninvasive tests (e.g., an X-ray). As a result, doctors must determine whether the benefits of a particular test are outweighed by its cost. A purely exploitative decision maker should stop deliberating and commit to a choice when the expected gain in value from a particular test is outweighed by its cost. An exploratory decision maker, however, might be willing to tolerate a small cost to learn about the environment. Agents must, therefore, carefully balance demands for exploration and exploitation when learning about a domain, or risk developing inaccurate beliefs (as depicted in Figure 2 ). Although medical decisions are often extended in time, we face the same challenges when making rapid decisions (i.e., deciding what information should be sampled), even about which eye movements to make, as evaluated in category learning experiments.
Self-Termination and Branching
As its beliefs are updated after observing each sample, SEA can display “branching” and “self-termination.” Branching involves changes in sampling strategy based on the values of the incoming information. Self-termination occurs when decision makers decide to commit to a choice, rather than selecting additional samples.
Such decisions about when to commit to a choice are a fundamental component of many temporally extended decisions ( Figure 1B ). Decision makers may fail to capitalize on transient opportunities for reward (or accrue excessive costs associated with deliberation) if they wait too long before committing to a choice. Conversely, if they respond too quickly, they may fail to collect enough evidence to support a desirable level of accuracy. We propose that the depth of forward search, which varies from myopic search to full preposterior analysis ( Figure 1B ), can be adjusted based on contextual demands on response timing. As clusters are “activated” based on the observed features, and the cluster representations predict the appropriate final choice (e.g., the category label), as more information is accumulated/sampled, inferences about the correct response become more accurate (assuming that clusters reflect relevant aspects of the environment).
Several models have been proposed to address the question of self-termination. “Integrate-to-bound” models, such as the Sequential Probability Ratio Test (SPRT; Wald & Wolfowitz, 1948 ) and the Drift Diffusion Model (DDM; Ratcliff, 1978 ), for example, operate by collecting evidence for competing hypotheses over time (in the form of a log-likelihood ratio), and committing to a choice when the strength of the cumulative evidence exceeds a predefined threshold. In typical implementations of these models, the threshold remains stationary during deliberation, and is chosen to balance the trade-off between accuracy and deliberation cost. Unlike SEA, however, these models act as passive observers, as they do not select the samples from which they learn.
In contrast, SEA selects samples sources of information through consideration of its beliefs about the environment, and updates these beliefs following the observation of each sample. Incidentally, the calculations involved in this procedure provide a principled way to define the termination criterion. Although a purely exploitative decision maker should commit to a choice when the expected gain in utility for each sample is outweighed by its cost, an exploratory decision maker may be willing to bear some sampling cost to learn about the environment. Whereas the DDM and SPRT define the termination criterion to balance demands for accuracy with missed opporunity costs, in SEA the termination criterion is calculated with regards to expected information gain, and a heuristic that strives to balance the trade-off between exploration and exploitation. 9
Because SEA strives to sample the most informative information source at each step, successive samples tend to become less informative. Concurrently, costs associated with deliberation tend to accrue. The likelihood of committing to a final choice, therefore, tends to increase with the number of samples observed. The dynamic nature of this decision boundary resembles that of some integrate-to-bound models (e.g., Cisek et al., 2009 ; Niyogi & Wong-Lin, 2013 ; Standage et al., 2013 ; Thura et al., 2012 ), which have been developed to account for modulation of the speed-accuracy trade-off during decision making. In both frameworks, the collection of additional information (which can be perceptual and/or mnemonic) generally improves decision accuracy, but also tends to increase costs associated with deliberation. However, whereas integrate-to-bound models tend to describe the decision process as the diffusion of a variable through time, SEA tracks expected information gain in conjunction with the accruing costs associated with collecting information samples. SEA additionally proposes that the depth of decision planning (i.e, ranging from myopic to full-preposterior forward search) influences the trade-off between decision accuracy and cost.
As discussed above, although leading contemporary models provide a compelling account for how decision makers organize information during decision making, they tend to sidestep questions relating to how decision makers choose information sources to sample, how they sequentially update their representations during deliberation, and how they terminate this deliberative process (for experimental evidence of sequential processing during human categorization, see: Milton & Wills, 2004 ; Wills et al., 2015 ). There are, however, a few notable exceptions. The Exemplar-Based Random Walk model (EBRW; Nosofsky & Palmeri, 1997 ), for example, sequentially retrieves items from memory until the available evidence exceeds a decision threshold. The EBRW does not, however, selectively encode stimulus information, but rather initially encodes all stimulus representations considered during the decision. Similarly, the extended generalized context model (EGCM-RT; Lamberts, 2000 ) stores representations of individual exemplars, but sequentially encodes stimulus feature values. As the feature values are encoded, the similarity between the stimulus and exemplars stored in memory is updated. This process resembles the sequential sampling problem faced by human learners, but information sampling is not strategic (i.e., it does not reflect previously retrieved information). In addition, the EGCM-RT will consider all stimulus features instead of self-terminating.
The Proposed Model
Here, we introduce a novel model of categorization, SEA, which is designed to treat decision making as an active sampling problem (in which decisions are made, not only about the final choice ( Figure 1A ), but also about what information to sample; Figure 1B ). It combines two normatively motivated components. The first is a concept-learning component that reflects the decision maker’s knowledge of the world. The second is a utility-sensitive sampling component that calculates the expected utility of particular states. The two components interact to perform preposterior analysis. These interactions allow the model to selectively sample from information sources that are expected to be useful for differentiating a set of “active” hypotheses.
Strategically sampling learners, such as SEA, can easily learn representations that deviate from reality ( Figure 2 ; Rich & Gureckis, 2018 ). This can happen when the learner fails to balance demands for exploration and exploitation. For instance, when a number of costly experiences with a stochastic variable are encountered early in training, a cost-sensitive decision maker may choose to avoid it, and never learn that it actually yields net long-term gain. 10 To encourage exploration of undersampled information sources, SEA can include exploration bonuses for undersampled information sources. As the partially observable Markov decision process (POMDP) can only be solved for relatively simple problems ( Knox et al., 2012 ), this mechanism can be seen as a heuristic linking the concept-learning and utility-sensitive sampling components.
Although category learning with feedback is typically treated as a supervised learning task, the present work recasts it as a problem in which the agent learns to traverse a series of probabilistic states (i.e., information samples) while minimizing sampling costs and maximizing reward (similar to reinforcement learning; Kaelbling et al., 1996 ; Sutton, 1990 ). Although SEA will initially sample uniformly across dimensions, it will gradually learn to sample selectively from dimensions expected to provide useful information. The resulting representational structure is efficient, in that it minimizes both the amount of information encoded across experiences, and the amount of information considered during individual decisions.
In SEA, the effects of selective sampling emerge with learning, and so sampling strategies change as the model learns about the environment. These bidirectional interactions between information sampling and concept-learning result in high-density representations along dimensions SEA believes are useful, and low-density representations along dimensions SEA deem irrelevant (reflecting the relative sampling frequency of these dimensions). This is analogous to the effects shown in Figure 1A , which are captured by “single-step” categorization models, which sidestep the information sampling stage of decision making, and selectively weight dimensions through attentional processes (e.g., Kruschke, 1992 ; Love et al., 2004 ; Nosofsky, 1986 ). In both frameworks, behaviorally relevant stimulus dimensions have greater influence on the final choice than do irrelevant dimensions.
Active sampling can lead to a self-enforcing pattern of belief updating, where beliefs about the world influence the information that is sampled from it, and this information is used to update beliefs. This can have important consequences on learning efficiency. When decision makers are free to select the stimuli from which to learn, they often learn more efficiently than when stimuli are presented in a predetermined order ( Castro et al., 2009 ; Gureckis & Markant, 2009 ; Markant & Gureckis, 2010 , 2014 ; Markant et al., 2015 ). This effect, however, depends on the structure of the problem being learned ( Enkvist et al., 2006 ; Markant & Gureckis, 2010 , 2014 ). Bidirectional interactions between information sampling and learning can also determine what concepts are ultimately learned. One example is the blocking effect ( Kamin, 1969 ), wherein after learning that a particular dimension is informative, a decision maker will tend to exploit this knowledge rather than continue to explore other information sources. To avoid these kinds of “knowledge traps” ( Rich & Gureckis, 2018 ), decision makers must successfully balance demands for exploration and exploitation ( Kaelbling et al., 1996 ; Sutton & Barto, 1998 ).
Model Overview
In this section, we present SEA, and its potential variations. SEA’s information-value component determines which (if any) features should be sampled. Its learning component provides the information-value component with the probabilities required to determine the sampling policy, and is updated based on the information sampled. Below, we specify these components, outline their interactions, and consider model variants that incorporate mechanisms that reflect the constraints of human decision makers.
The concept learning component we use is closely related to the Rational Model of Categorization (RMC; 1991b ; Anderson & Matessa, 1990 ), although any generative probabilistic model would also likely be appropriate. The RMC incrementally learns to sort stimuli into appropriate clusters, and can make near-optimal use of past information during learning and prediction. Here, we provide an overview of the RMC. Additional details can be found in the original articles.
The RMC is a flexible clustering model, which learns to parcelate representational space into clusters based on its experience with the normative characteristics of the task environment. Formally, the probability that any unobserved stimulus dimension, F i , will take a particular value, j , can be inferred by weighting the prediction of each cluster, P ( F i = j | k ), by the probability of the cluster given the observed features, P ( k | F O ): P ( F i = j | F O ) = ∑ k P ( F i = j | k ) P ( k | F O ) . 1 where P ( F i = j | k ) is calculated using Equation 2 , and P ( k | F O ) is estimated using Equation 3 . By this notation (which we will use throughout the article) the subscript, O , denotes the index of the observed features of a given stimulus, and i denotes the index of the considered feature. For instance, given a stimulus (including both observed and unobserved dimensions) defined as vector F = [2, 1, 1, 2], if the second feature was under consideration, and the third and fourth features were known, then i would be 2, O would be [False, False, True, True], and F O would be [?, ?, 1, 2].
For each dimension, discrete feature values are assumed to be distributed according to a Dirichlet density characterized by dimension-value parameters α j , and dimension-wide parameters, α 0 (where α 0 = Σ j α j ). The Dirichlet distribution allows the data to determine the number of clusters (as in SUSTAIN; Love et al., 2004 ), and allows for a potentially infinite number of clusters. However, between one and three clusters per category is typical. These desirable characteristics of the Dirichlet distribution have led to it being used in many categorization models (e.g., Anderson, 1991a ; Griffiths et al., 2007 ).
Across learning, SEA tracks the number of items in cluster k with the same value, j , on feature i in C ij . The posterior is also Dirichlet-distributed, and the probability that a feature will take a particular value within a cluster is as follows: P ( F i = j | k ) = α j + C i j α o + ∑ j C i j . 2
As C ij becomes populated through experience, it exerts stronger influence on P ( F i = j | k ) relative to the prior. The prior parameters (the α’s), therefore play an important role during early learning, as they allow SEA to appropriately estimate its uncertainty when few samples have been observed. After a single trial, for example, it would be erroneous to infer that all future objects will display the observed values.
Bayes’ theorem can be used to calculate the last term in Equation 1 , P ( k | F O ). This term represents the probability (or “activation”) of each cluster given the observed features: P ( k | F O ) = P ( F O | k ) P ( k ) ∑ k P ( F O | k ) P ( k ) , 3 where P ( F O | k ) is calculated using Equation 2 , and P ( k ) represents the prior probability that any stimulus will be assigned to cluster k . This probability is calculated as follows: P ( k ) = c n k ( 1 − c ) + c n , 4 where c denotes the coupling probability (a parameter that determines the probability that two objects come from the same category), n k is the number of items already assigned to cluster k , and n is the total number of stimuli observed. The prior probability that a stimulus will be assigned to a novel cluster is as follows: P ( 0 ) = ( 1 − c ) ( 1 − c ) + c n . 5
As no clusters have yet been created on the first trial, the model will start with a single cluster with each feature initialized with a uniform probability of occurring (as in Equation 5 ). With greater experience, the model will incrementally learn a single partition of stimuli into clusters. 11 Although the fully normative solution would be to consider all possible partitions of stimuli into clusters ( Anderson, 1991a ), this approach is intractable for all but the simplest problems. 12 The incremental approach may also be more psychologically valid ( Love et al., 2004 ). With the parameters set as in the simulations described below, the model tends to sample all features before selectively sampling from those expected to provide useful information.
Combining Concept-Learning With a Utility-Sensitive Sampling Norm
When facing a choice with an uncertain outcome, the expected utility of a particular action, a , can be calculated by weighting the utility of each resulting state by its probability. In a category learning experiment, for example, one category label may be more probable than the other, but yield lesser reward. The action-utility function shown in Table 1 corresponds to a contingency table in which two states (or categories), s p and s q , are mutually exclusive and exhaustive (i.e., P ( s p ∪ s q ) = P ( s p ) + P ( s q ) = 1), and the decision maker must choose the appropriate action ( a p or a q ; in a categorization experiment, this corresponds to the category label). The table depicts a hypothetical action-utility function reflecting the utility for two actions: a p and a q . For this particular example, maximizing utility is equivalent to maximizing accuracy, as correct responses are rewarded with 100 utility units, and incorrect responses are awarded zero units. The table could be expanded to include more than two actions and states.
For the action-utility function shown in Table 1 , the expected utility, E ( U ) of action a p can be calculated as follows: E ( U ( a p ) ) = U ( a p | s p ) P ( s p ) + U ( a p | s q ) P ( s q ) . 6
For example, if P ( s p ) = 0.7, and P ( s q ) = 0.3, the expected utility of choosing a p would be 70 and that of a q would be 30. In this case, the utility-maximizing action would be to choose a p . As mentioned, payoffs can also be asymmetric. For instance, if the lower-left entry in Table 1 was −1,000, there would be a high penalty associated with a p when state s q holds, and the optimal choice would switch to a q .
The above examples describe problems involving a single feature with two possible values ( s p and s q ). Real-world decisions typically require decision makers to integrate evidence across multiple features, which often have more than two possible values. When diagnosing a tumor, for instance, it might be necessary to consider results from blood tests as well as from CT-scans or MRI. Categorization tasks are often designed to reflect this aspect of real-world decisions; participants must integrate information across relevant stimulus features.
In SEA, as in Anderson’s RMC (1991b ; Anderson & Matessa, 1990 ), the category label is treated like any other cluster feature, and Equation 1 can be used to calculate the probability of each label, given the observed feature values. The value of action, a , given the observed features, F O , 13 can be estimated by summing over states, s , and subtracting the costs associated with sampling each observed feature, ℒ o : 14 E ( U ( a | F O ) ) = ∑ s U ( a | s ) P ( s | F O ) − ∑ o ∈ O ℒ o , 7 where P ( s | F O ) is provided by Equation 1 , and U ( a | s ) was introduced in Equation 6 . Before learning about the environment, a uniform prior (resulting from Equations 4 and 5 ) drives probabilistic sampling of each stimulus feature.
The estimated utility of the current state, F O , can be estimated by maximizing over possible actions: E ( U ( F O ) ) = argmax a ∈ Actions ( E ( U ( a | F O ) ) ) . 8
As discussed, real-world decisions often require decision makers to decide what information should be sampled. This is important, as the information that is sampled can influence the final choice. The results of a blood test, for instance, can influence a doctor’s decision about whether to suggest chemotherapy for a patient. To estimate the utility of a test that reveals the value of an unknown feature (e.g., “cancer antigen present” vs. “cancer antigen absent”), we consider how much the results of the test would improve the utility of the current state (where the current state is defined by the vector of observed features, F O ). The expected utility of the state after sampling unobserved feature i , can be estimated by summing across its possible values, j : E ( U ( F O , F i ) ) = ∑ j ∈ F i E ( U ( F O , j ) ) P ( F i = j ) , 9 where E ( U ( F O , j ) ) denotes the expected utility of the state if value j (of unobserved feature F i ) was included in the vector of observed features. Equation 9 demonstrates how the expected utility of the state can be calculated for a single feature. As each feature can have multiple values (two in the simulations described below), the model explores each branch for each feature. During “myopic” decisions, the model considers only a single step into the future. Preposterior analysis (which is implemented in SEA as a “depth-first” search process), involves imagining each branch several steps into the future. The computational demands of preposterior analysis, therefore, are high, even for the relatively low-dimensional decision problems commonly considered in the categorization literature.
Equation 9 contributes to the calculation of the gain in utility (cf., Nelson, 2005 ; Nelson et al., 2010 ) from sampling unobserved feature i : G ( F i ) = E ( U ( F O , F i ) ) − E ( U ( F O ) ) . 10
SEA proposes that this expected increase in utility from sampling F i is the key variable to consider when deciding what feature to sample, or whether to stop sampling and commit to a final choice. When G ( F i ) for all features is less than, or equal to zero, a cost-sensitive decision maker should stop sampling and commit to a final choice. When G( F i ) for at least one feature is greater than zero, an exploitative strategy would be to sample the feature with the greatest expected gain.
Importantly, costs are often dependent across features. The cost of a blood test, for instance, can be substantially less if other blood tests have already been ordered. A normative strategy therefore requires the consideration of all possible sequences of tests to account for these potential dependencies. As the computational demands of this approach increase exponentially with the number of features considered, it can only be justified when decisions involve a low number of stimulus features (as
Balancing Demands for Exploration and Exploitation
Precisely determining the optimal balance of exploration and exploitation is intractable for most tasks and is only possible for special cases ( Kaelbling, 1993 ; Kaelbling et al., 1996 ; Simsek & Barto, 2006 ). To derive the optimal solution, one would need to make several assumptions. It would be necessary, for instance, to estimate the number of trials left in the study (as the negative consequences of choosing a suboptimal strategy increases with the number of trials on which it is applied). It would also be necessary to estimate how rewarding the environment is (as optimal inference requires normalizing estimates based on environmental characteristics). It would also be necessary to estimate the probabilities of different category structures, which represents uncertainty about the appropriate categorization strategy (alternatively, one could restrict the possible forms of the environment, as in Stankiewicz et al., 2006 ). Finally, it would also be necessary to consider the probability of any of these factors changing over time (cf., Brown & Steyvers, 2009 ; Gittins & Jones, 1979 ; Steyvers et al., 2009 ).
Fortunately, a number of heuristic methods exist ( Kaelbling, 1993 ; Kearns & Singh, 2002 ; Moore & Atkeson, 1993 ; Schmidhuber, 1991 ; Sutton, 1990 ). We combine two of these heuristic methods: stochastic choice via a softmax choice rule, and exploration bonuses for underexplored options ( Kaelbling, 1993 ). The exploration bonus, E , could take many forms. In Kalman filter models, this term often takes the form of an uncertainty bonus that reflects the standard deviation of the choice’s utility ( Daw et al., 2006 ). In the current model, E is calculated for each feature separately: E i = max ( U ) − E ( U ( F O ) ) ( 1 + n i ) ϕ , 11 where max( U ) denotes the maximum utility possible irrespective of sampling costs, E ( U ( F O ) ) will always be less than or equal to max( U ). n i denotes the number of previous observations of feature i , and ϕ denotes a fixed parameter modulating the influence of n i on E i . In the case that F O supports perfect prediction, the comparison of E ( U ( F O ) ) (i.e., the expected utility, including the costs of sampling each feature given the observed features) to max( U ) encourages the model to explore when sampling F O is costly.
Combining the exploration bonus with a softmax choice rule, the probability of sampling feature m is: P ( F m ) = e β ( G ( F m ) + E m ) ∑ n e β ( G ( F n ) + E n ) , 12 where β denotes a nonnegative temperature parameter that modulates the stochasticity of the decision process (i.e., “how often is the feature with highest expected profit chosen?” ). When E m = 0, and G ( F m ) ≤ 0, the model stops deliberating and commits to a final choice.
Summary
SEA interleaves concept learning and information sampling, such that they mutually influence one another. Information sampling is akin to a dynamic planning process in which SEA’s concept learning component (i.e., the RMC) serves as an internal model of the environment. For instance, the RMC may learn that red objects tend to be heavy 90% of the time. After observing that an object is red, the RMC would update its expectation that the object is heavy to 90% ( Equation 1 ). Before learning this relationship between color and weight, the RMC would rely on its uninformative prior (50% of objects are heavy, 50% of objects are light) to guide its predictions.
Calculating the probabilities of these unobserved features (e.g., weight) is critical for planning which feature to sample next. The expected utility of a possible state is calculated by combining the probabilities of these states with their utilities (e.g., Equation 7 ). Importantly, the utility sensitive sampling component does not learn utilities of various states. Instead, SEA is initialized with a utility table (as in Table 1 ) and with the costs associated with sampling each information source. The conjunction of the concept-learning and utility-sensitive sampling component allows the model to perform active sampling.
Equation 8 is used to calculate the expected utility of states (i.e., specific stimulus feature configurations), abstracting beyond specific actions (or choices). Equation 9 is used to calculated the expected utility associated with sampling an unseen feature, abstracting beyond its possible values. This helps the model to determine if, after sampling a single feature, it should sample another feature. To make this determination, SEA considers information gain ( Equation 10 ) and the exploration bonus for each unsampled stimulus feature ( Equation 11 ) and combines them using a softmax choice rule ( Equation 12 ).
In deciding which feature to sample, SEA plans ahead for the maximal number of steps, like an adult might when playing a simple game such as tic-tac-toe. In the simulations, we compared SEA to variants that are “myopic” in that they only consider the next step or move ( Figure 3 ). To clarify how the equations interact to support myopic decision making and preposterior analysis, we describe SEA’s behavior in a two-class categorization problem involving stimuli with three binary stimulus features. We assume that the model has already been trained.
Myopic decision making involves simulating the sampling of single unobserved features. Before sampling any stimulus features F O is [?, ?, ?]. To determine what feature to sample first, the model simulates the effects of sampling each. For instance, the model might calculate the expected utility of the possible states after sampling the first feature (i.e., F O = [0, ?, ?] or [1, ?, ?]) using Equation 8 . The expected utility of sampling this particular feature can be calculated by combining these expected utilities across feature values ( Equation 9 ). The gain in utility from sampling the feature can then be calculated using Equation 10 . The exploration bonus for this feature could then be calculated using Equation 11 . After performing these calculations for each feature, the decision of what feature to sample would be made using the softmax choice rule ( Equation 12 ).
Conceptually, preposterior analysis is an extension of the myopic algorithm that involves simulation of multiple unobserved features. As in the previous example, the model might begin a trial by calculating the expected utility of the first feature (feature “1”) being “0” (i.e., F O = [0, ?, ?]). Holding this imaginary feature-value constant, SEA would then simulate the expected utility if other features were subsequently sampled. For instance, to simulate sampling feature “2”, SEA would consider the expected utility of states F O = [0, 0, ?] and F O = [0, 1, ?], using Equation 8 . It would then abstract over these possible values using Equation 9 . It would then calculate the gain in utility and the exploration bonus associated with sampling this feature using Equations 10 and 11 .
The process would then be repeated with the value of feature “1” set to 1 (i.e., F O = [1, ?, ?]). When the depth of the forward search is limited to two steps, the algorithm would commit to sampling a feature after simulating the sampling of two features. 15 For a three-feature categorization problem, SEA would simulate the sampling of all three features before sampling the first. After sampling one feature, SEA would simulate sampling both remaining features. Importantly, SEA does not learn anything during simulation. The concept learning component is updated only after the final choice is made, and this learning changes behavior on future trials only.
Although the myopic decision algorithm requires minimal computational demands, it lacks the sophisticated behavior that forward search enables (i.e., strategic self-termination and branching). When SEA employs a myopic strategy, it tends to sample more dimensions, and to be less accurate (in terms of its categorization decisions), than when preposterior analysis is employed. These limitations of the myopic algorithm are illustrated in the simulation of experiment performed by Blair et al. (2009 ; see Strategic Attention Within Individual Trials section).
Figures, tables, references, and supplementary files are best inspected in the licensed PDF or repository copy linked above.