Theories of Intelligence
Linas JuozenasShare
Intelligence Unleashed · Foundations
Theories of intelligence from g to CHC—and beyond
Why do performances on different mental tasks tend to move together? Why can the same person reason brilliantly in one setting yet struggle in another? Spearman, Sternberg and the Cattell–Horn–Carroll tradition begin with different questions. Read together, they reveal both the structure that cognitive tests reliably detect and the parts of intelligent action that no single score can contain.
A guide to the argument—not a contest with one winner
The models overlap, but they do not make the same kind of claim
Spearman’s tradition asks what explains the positive correlations among cognitive tests. CHC classifies the general, broad and narrow abilities visible in those correlations. Sternberg asks what it takes to pursue worthwhile goals when novelty, experience and social context matter. Confusion begins when an answer at one level is treated as the whole definition of intelligence.
Psychometric models are powerful because they compress repeatable patterns in data. They can help predict learning and performance, design balanced test batteries and describe relative cognitive strengths. They do not, by themselves, explain the biological or developmental cause of every correlation; measure wisdom, motivation, character or human value; or tell a teacher, employer or clinician what decision is fair.
01 · Orient the debate
Three questions, three maps
The frameworks are easiest to understand when each is allowed to answer the question it was built to answer.
An intelligence theory can describe the structure of individual differences, propose mental processes, explain development, or define effective action. These are connected projects, but they are not interchangeable. Spearman and CHC are principally psychometric; Sternberg’s theory is broader, functional and contextual.
Why do test scores correlate?
Spearman proposed that every cognitive test reflects a general factor, g, plus influences specific to that task. His lasting discovery is the positive manifold: people who perform well on one cognitive task tend, on average, to perform well on others.1
What makes thought effective?
Triarchic theory examines the interaction of mental components, experience and context. Its familiar analytical, creative and practical abilities describe ways of deploying intelligence—not simply three independent scores that replace IQ.
| Level | The question | What evidence can establish | What it cannot establish alone |
|---|---|---|---|
| Observation | Which tasks covary, and by how much? | A reproducible correlation pattern in a defined sample. | The cause of the pattern or its moral meaning. |
| Statistical or structural model | Can a specified factor, network or other model represent those associations parsimoniously? | How well the chosen structure reproduces the observed data and where residual misfit remains. | That statistical fit identifies one unique causal mechanism. |
| Explanatory theory | Which processes and developmental interactions generate performance? | Predictions tested across tasks, time, interventions and populations. | A fair decision rule merely because the explanation fits. |
| Application | How should a score inform teaching, diagnosis or selection? | Validity for that use, reliability, fairness, consequences and alternatives. | A universal policy from a correlation in another setting. |
The frameworks are lenses, not human hierarchies
A person may show strong abstract reasoning but little relevant knowledge, deep expertise but slower processing, or original practical judgement that a timed battery barely samples. These patterns matter. None determines whose interests count, whose voice deserves respect or who possesses equal human rights.
02 · Read the statistics correctly
What factor analysis can—and cannot—tell us
Before arguing about g, it helps to distinguish raw performance, standardised scores, correlations and latent factors.
Imagine a battery containing vocabulary, matrices, spatial rotation, memory and speed tasks. Each produces an observed score. When scores from many people are correlated, a familiar pattern usually appears: most correlations are positive. Factor analysis represents that covariance using a smaller number of unobserved variables. The result is a disciplined summary of differences among people in that sample—not a scan of a substance called intelligence.
-
Sample performances. Tasks differ in content, instructions, language, time limits, sensory and motor demands. The battery defines what has an opportunity to count.
-
Estimate relationships. Researchers examine whether people’s relative performances move together. A correlation describes association across people; it does not say that one task causes another.
-
Specify a model. A one-factor, correlated-factors, higher-order, bifactor or network model can be fitted. Choices about tasks, scoring, estimators and constraints affect the result.
-
Test generalisation. Fit in one dataset is not enough. A useful model should be interpretable, replicable, stable enough for its purpose and tested across relevant ages, languages, groups and conditions.
-
Validate the intended use. The final question is not “Is this test valid?” in the abstract, but “Are these score interpretations and decisions supported for this population and purpose?”37
The positive manifold
Across varied batteries, cognitive task scores generally correlate positively. The exact correlations and factor loadings change with the task set and population, but the broad phenomenon is one of psychology’s most replicated patterns.5
Why the manifold exists
A common cause is one explanation. Overlapping processes, developmental mutualism and networks of interacting abilities can also generate positive correlations. Statistical fit alone rarely chooses conclusively among these causal stories.
“Latent” does not mean hidden biological organ
A latent variable is not observed directly. It earns meaning from the model, indicators and external evidence. Treating g as a useful statistical dimension is compatible with disagreement about whether it reflects one causal capacity, many overlapping processes, reciprocal development or some combination.
03 · Spearman’s starting point
A general factor and task-specific influences
Spearman’s 1904 paper supplied a method and a hypothesis that reshaped psychological measurement.
Working with school marks and sensory-discrimination measures, Charles Spearman argued that correlations among intellectual performances could be explained by a common general factor plus a specific factor for each task. He called the general factor g and the task-specific component s.1
General variance
The portion of individual differences that diverse tasks have in common. In a factor model, each task has a loading that expresses its relationship with the general factor.
Specific and group variance
Performance also depends on narrower abilities, learned methods, task content and response demands. Later evidence required intermediate “group” factors that Spearman’s simplest two-factor account did not fully accommodate.
Error and occasion
Fatigue, distraction, misunderstanding, practice, scoring imprecision and chance influence observed scores. Reliable assessment separates signal from noise as far as possible; it never eliminates uncertainty.
The durable insight is not that every mind contains a fixed quantity of “mental horsepower.” It is that diverse cognitive performances share variance—and any adequate theory must explain why.
What changed after Spearman
Researchers repeatedly found clusters that were broader than one task but narrower than g: verbal, spatial, memory, speed and other domains. Louis Thurstone emphasised several primary mental abilities; later work showed that even these abilities correlated. Hierarchical models therefore became a natural compromise: specific tasks at the bottom, broad abilities in the middle and general covariance at the top. Carroll’s later three-stratum synthesis made this architecture especially influential.24
Two historical shortcuts to avoid
Spearman helped pioneer factor-analytic reasoning, but modern factor analysis is the result of many contributors and methods that did not exist in 1904. And later IQ batteries did not simply turn his original theory into one pure number: they sample multiple domains, use different scoring models and yield both composites and narrower indexes.
04 · The general factor now
A stable pattern with several possible meanings
Modern researchers agree more readily on the covariance than on the ontology—what, exactly, the statistical factor represents.
In a sufficiently broad, balanced battery, a general dimension can usually be extracted, and general-factor scores from well-constructed batteries often rank people similarly.58 Yet a g score is not entirely test-free: its content and precision depend on which indicators are included, which population is sampled and which model is fitted.
A reflective model treats general ability as a common influence on task performance. Process accounts can instead treat the general index as an aggregate emerging from many overlapping demands rather than as one cause with interchangeable effects. Either representation can summarise covariance; distinguishing them requires evidence beyond the correlation matrix.8
| Model | Basic structure | What the general factor means statistically | Important caution |
|---|---|---|---|
| Single factor | All tasks load on one common dimension. | One factor reproduces part of their shared covariance. | Usually too simple for a broad modern battery; residual clusters remain. |
| Correlated abilities | Verbal, spatial, memory and other factors correlate. | No explicit higher-order factor is required. | It describes relations among broad domains but leaves their correlations to be explained. |
| Higher order | Tasks load on broad abilities, which in turn load on g. | g accounts for covariance among the broad factors. | General influence is mediated through the broad abilities in the model. |
| Bifactor | Each task loads directly on a general factor and usually one group factor. | General and group dimensions partition item or subtest variance. | Extra flexibility can improve fit even when the true process is not bifactor; factors may become unstable or hard to interpret.944 |
| Network or process model | Abilities or processes interact, overlap or reinforce one another. | A general factor may summarise an emergent pattern rather than act as one common cause. | Promising explanations still require discriminating longitudinal and experimental tests. |
Why “the percentage explained by g” has no universal answer
Reports that g explains a fixed share of “intelligence” mix together different denominators and methods. A first principal component, a common factor, a higher-order factor and a bifactor general factor do not partition variance in the same way. Estimates also depend on the breadth and reliability of the tests. The safe conclusion is qualitative: general covariance is substantial, but broad and specific variance—and measurement error—remain.
Use an index as an index
Like an economic index, g compresses many correlated observations into a useful summary. An index can predict and compare without being a single physical ingredient. Biological research may identify distributed systems that support broad performance, but no one brain region, neurotransmitter or neural-speed measure is simply identical to psychometric g.
05 · Useful, not destiny
What general cognitive ability predicts
Prediction is one of the strongest reasons general scores remain useful—and one of the easiest reasons to overreach.
General cognitive ability is associated with how quickly people learn many complex tasks and with average performance in education and work. Childhood cognitive scores also predict later health and mortality at the population level. These findings are important. None turns a score into a complete causal explanation or a certain forecast for one person.
Learning and achievement
Cognitive ability and school achievement are substantially related, partly because reasoning supports learning and accumulated learning changes the knowledge and strategies that tests sample. Motivation, instruction, opportunity, prior knowledge and self-regulation also matter.11
Work performance
Ability tests predict job performance. A 2026 synthesis found little evidence that cognitive or data complexity moderated that relationship and produced estimates lower than classic headline figures. Training performance is a separate criterion, so validity should be estimated for the actual outcome. Structured interviews, work samples and relevant experience can add information.10
Health and longevity
Early cognitive scores correlate with later morbidity and mortality. Education, working conditions, resources, health literacy, behaviour, reverse causation and shared biological or social influences can contribute. A test score is not a diagnosis or treatment target.12
| Evidence | Defensible reading | Overclaim | Missing information |
|---|---|---|---|
| A group correlation | Higher scores tend to accompany higher outcomes on average. | Every higher-scoring person will outperform every lower-scoring person. | Overlap, uncertainty, other predictors and the person’s circumstances. |
| Incremental validity | A score improves prediction beyond specified existing information. | The test is the single best or only necessary decision method. | Comparison set, criterion quality, fairness, costs and consequences. |
| Longitudinal association | An earlier score forecasts a later outcome. | The measured ability directly caused the outcome. | Confounding, mediation, selection and reciprocal effects. |
| Average stability | Rank ordering is moderately or strongly persistent in a population. | An individual score cannot change. | Measurement error, development, education, health and altered opportunity. |
06 · Beneath the correlations
Competing explanations for the positive manifold
The existence of general covariance is less controversial than its cause. Several accounts can be partly true at once.
One broad capacity influences many tasks
Reflective factor models are often read as if a general ability causes performance across domains. This is intuitive and useful for prediction, but covariance alone cannot prove that causal direction or identify the mechanism.
Tests overlap in elementary processes
In Thomson-style sampling accounts, tasks draw partly overlapping samples of many basic processes. The shared sampling produces positive correlations without requiring a single mental energy.
Abilities strengthen one another during development
Early cognitive processes may begin less correlated, then become coupled: better attention aids learning; knowledge improves reasoning in familiar domains; success creates more opportunities. Reciprocal growth can produce a manifold and rising stability.67
Domain-general control is repeatedly required
Process Overlap Theory proposes that tests sample domain-specific processes alongside executive processes that are broadly required and capacity-limited. A general score emerges from overlapping demands rather than one unitary process.8
Psychometric network models represent observed scores as nodes linked by conditional associations; those edges alone do not establish causal influence. Other accounts emphasise working memory, attentional control, processing speed, neural efficiency or their combinations. The strongest future evidence will come from models that make different predictions about development, intervention, brain–behaviour relations and performance under carefully changed task demands—not from fit contests alone.528
Structure and mechanism can coexist
CHC or a hierarchical g model may organise scores well even if mutualism or process overlap better explains how the pattern arose. A map of covariance and a theory of development need not use the same entities. The key is to state which job each model is doing.
07 · Intelligence in action
Sternberg’s triarchic theory
Triarchic theory broadens the frame from correlations among test scores to the processes, experiences and environments through which people act intelligently.
Robert Sternberg’s 1985 theory is often reduced to analytical, creative and practical intelligence. That is a useful introduction, but the original architecture is more precise: a componential subtheory describes information processing, an experiential subtheory describes novelty and automatization, and a contextual subtheory describes adaptation, environmental shaping and selection. The three familiar abilities cut across these concerns.55
Examine, compare and evaluate
Analytical intelligence is prominent when a problem is clearly posed and a person must identify relations, judge alternatives, monitor progress or justify a conclusion. Conventional academic tests sample much of this territory, though no item is process-pure.
Respond to novelty and generate
Creative intelligence concerns coping with relative novelty, finding a fruitful representation, imagining alternatives and eventually automating familiar operations so attention can move to harder levels. Originality alone is not enough; an idea must also be relevant to the problem.
Make knowledge work in context
Practical intelligence concerns pursuing goals amid incomplete instructions, local norms, competing interests and real consequences. Research often operationalises it through tacit knowledge or situational judgements—not a generic inventory of “street smarts” or repair skills.
The three subtheories underneath the popular summary
| Subtheory | Central idea | Key mechanisms | Example |
|---|---|---|---|
| Componential | Intelligent behaviour depends on coordinated information-processing components. | Metacomponents plan and monitor; performance components execute; knowledge-acquisition components encode and learn. | Recognising what a problem asks, choosing a strategy, carrying it out and checking the result. |
| Experiential | The intelligence of an act depends partly on a person’s experience with the task. | Handling relative novelty and automatising recurring operations. | An expert quickly recognises a familiar pattern, freeing capacity for the genuinely new part. |
| Contextual | Intelligence serves purposeful fit between person and environment. | Adapting to conditions, shaping them or selecting a different environment. | Learning an organisation’s unwritten constraints, changing a broken process or leaving for a better fit. |
Sternberg later called the framework the theory of successful intelligence: the ability to achieve personally meaningful goals within a sociocultural context by capitalising on strengths, correcting or compensating for weaknesses and balancing analytical, creative and practical abilities. “Success” here is contextual, not merely wealth or status, and adaptation need not mean passive conformity.17
Adapt, shape or select
A student may adapt by learning an existing format, shape the environment by requesting an accessible alternative, or select a different route that better supports the goal. Intelligent action can involve changing an unjust or ineffective context—not simply becoming better at enduring it.
This is not a learning-styles prescription
Teaching learners through varied explanations, examples, creation and application can enrich instruction. That does not establish that every student has one fixed analytical, creative or practical “style,” or that matching lessons to a labelled style improves learning. All learners may benefit from practising all three modes.
08 · Promise and contest
What the evidence supports—and where it resists
Sternberg’s programme generated imaginative assessments and classroom studies. It also illustrates how difficult it is to measure broad, contextual abilities as distinct constructs.
The strongest charitable reading is neither “triarchic theory disproved IQ” nor “it has no evidence.” Creative and practical tasks can add useful information and diversify what people are invited to demonstrate. Whether their scores form three independent intelligences, predict broadly beyond existing measures and scale fairly remains contested.
| Programme | What was reported | What it supports | What remains uncertain |
|---|---|---|---|
| Rainbow Project | Among 1,013 participants at 15 sites, adding creative and practical tasks to SAT scores increased author-reported, in-sample explained variance in college GPA for the analysed subgroup; several added measures showed smaller mean group gaps.18 | Broader performance tasks can contribute incremental information in a research admissions setting. | Large independent operational replication, out-of-sample performance, coaching, cost, reliability and consequences. Smaller score gaps alone do not prove fairness. |
| MBA tacit knowledge | Case and situational-judgement measures showed small-to-moderate relations with academic and consulting criteria, with some small increments beyond GMAT and prior grades.19 | Job-relevant judgement can add contextual evidence. | Whether the score reflects a distinct intelligence rather than job knowledge, general ability, personality, reading and method variance. |
| STAT factor studies | Some analyses found that response mode or content—verbal, quantitative and figural—organised scores as much as, or better than, analytical, creative and practical labels.21 | The tasks can sample varied performances. | A clean three-factor structure and independence from general cognitive ability. |
| Aurora gifted assessment | Developer studies reported promising reliability and identification patterns; an independent study of 499 Dutch pupils found poor fit for the proposed triarchic structure.22 | Nontraditional tasks may reveal candidates missed by one conventional cutoff. | Whether additional classifications are valid, stable and predictive of meaningful long-term outcomes. |
| Triarchic teaching | Early studies were favourable, while a later study of 7,702 fourth-graders produced only a small number of significant advantages across 23 comparisons.23 A meta-analysis found a positive average that became smaller after publication-bias adjustment.24 | Teaching for analysis, creation and application may be a productive design principle. | The size, durability, active ingredient and generalisability of benefits; evidence does not justify rigid learner typing. |
Why practical intelligence is hard to isolate
Tacit-knowledge inventories typically present a workplace or life scenario and ask respondents to rate possible actions. That can have face validity and criterion validity while still mixing several constructs. Situational-judgement research finds that scores vary with instructions and relate to cognitive ability and personality.20 A useful assessment need not measure one pure faculty, but its interpretation must match what the evidence supports.
A design lens for broader opportunity
Invite people to analyse evidence, produce alternatives and apply ideas. Use portfolios, work samples, structured scenarios and conventional measures when each is relevant. Validate the combination for the intended decision.
A replacement taxonomy of three separate minds
The evidence does not warrant assigning everyone a fixed triarchic type, treating the three abilities as independent of g, or inferring practical intelligence from charisma, rule-breaking or social advantage.
From successful to adaptive intelligence
Sternberg’s newer theory of adaptive intelligence asks whether intelligence helps people and communities adapt to real, shared problems—not merely optimise personal outcomes within existing systems. It brings wisdom, ethical consequences and the common good closer to the centre.25 This is a provocative normative framework, especially for education, but it is not yet a mature psychometric taxonomy with the measurement base of g or CHC.
Sternberg’s enduring challenge is not that test scores are useless. It is that solving a decontextualised problem and choosing a worthwhile problem to solve are different achievements.
09 · From one factor to differentiated abilities
Cattell, Horn and fluid–crystallised intelligence
The route to CHC began with a distinction between solving relatively novel problems and deploying knowledge built through learning and culture.
Raymond Cattell proposed fluid ability (Gf) and crystallised ability (Gc) in the 1940s; John Horn later tested, refined and expanded the model. The familiar contrast remains valuable, but neither ability is pure, context-free or isolated from the other.2627
Reason with relationships that are relatively new
Induction, deduction, quantitative relations and identifying rules in unfamiliar material are central examples. A matrix problem can reduce explicit schooling demands, but it still requires instructions, visual access, familiarity with testing conventions, attention and strategies developed through experience.
Use language and culturally acquired knowledge
Vocabulary, lexical knowledge, listening comprehension and general verbal knowledge are common examples. Gc is not a store of trivia: it reflects the breadth, depth and organisation of knowledge available through a person’s learning history and language community.
Investment is a relationship, not a one-way destiny
Cattell’s investment idea proposed that fluid capacity is “invested” in learning, helping produce crystallised knowledge. Contemporary development is more reciprocal. Prior knowledge makes new reasoning more efficient; interests influence what is practised; schooling changes both knowledge and test strategies; motivation and opportunity govern whether potential becomes expertise. A person can therefore have similar broad reasoning to a peer but very different domain knowledge—or deep expertise that compensates for slower novel processing.
| Question | Gf | Gc | Nuance |
|---|---|---|---|
| Typical demand | Discover a relation or solve a relatively unfamiliar problem. | Understand and apply learned language or knowledge. | Most real tasks require both: one must represent a problem using something already learned. |
| Common indicators | Matrix reasoning, series, induction and quantitative reasoning. | Vocabulary, verbal comprehension and general information. | Indicators also contain method, sensory, speed and schooling demands. |
| Experience | Designed to reduce dependence on specific prior content. | Explicitly reflects acculturation and education. | No complex human task is literally experience-free or culture-free. |
| Development | Often rises earlier and shows average adult decline sooner. | Often grows longer as knowledge accumulates. | Abilities and tasks follow different trajectories; cohort, health and measurement matter. |
Horn did not simply endorse Carroll’s top-level g
Horn expanded Gf–Gc theory into multiple broad abilities and generally resisted placing one general factor above them. Carroll’s hierarchy did include g. Modern CHC practice often draws the hierarchy with g at the top, but the name preserves a productive disagreement as well as a synthesis.
Historical note: Donald Hebb’s distinction between biologically grounded “Intelligence A” and developed “Intelligence B” influenced Cattell’s formulation. Archival work therefore cautions against a simple lone-inventor story.42
10 · A hierarchy assembled from decades of data
Carroll’s three strata and the CHC synthesis
John Carroll’s 1993 work reorganised a vast historical factor-analytic literature into one influential hierarchy.
Carroll reanalysed 461 datasets spanning more than seven decades of cognitive-ability research. His three-stratum theory placed narrow abilities at Stratum I, broader domains at Stratum II and a general factor at Stratum III. The achievement was a systematic synthesis of covariance evidence—not the discovery of three literal floors inside the brain.24
Stratum III · General. The broadest shared dimension across a sufficiently varied set of cognitive measures: psychometric g.
Stratum II · Broad. Relatively general domains such as fluid intelligence, crystallised intelligence, memory and learning, broad visual perception, broad auditory perception, broad retrieval ability, cognitive speediness and processing speed in Carroll’s original formulation.
Stratum I · Narrow. More specific abilities—such as induction, memory span, spatial relations or word fluency—identified through performances that share a tighter set of demands.
Carroll’s hierarchy substantially overlapped with the expanded Cattell–Horn model. Kevin McGrew and other assessment scholars progressively aligned the taxonomies around the development of later Woodcock–Johnson batteries, and “CHC” became the umbrella term. The synthesis informed test blueprints, score classifications and cross-battery interpretation.3
Broad abilities recur
Fluid reasoning, acquired knowledge, visual processing, auditory processing, memory and speed repeatedly emerge when appropriate tasks are sampled.
The place of g
Carroll regarded a general stratum as important; Horn preferred explanations in terms of broad abilities and their developmental relations. CHC can be presented with or without a causal reading of the top factor.
The taxonomy keeps changing
Contemporary CHC splits, renames and tentatively adds abilities as evidence develops. Exact counts vary with the edition and with whether sensory and motor domains are included.
A hierarchy is not a value ladder
“Higher” means more general in a statistical model, not more important in life. A narrow skill can decide whether a person reads music, detects a subtle sound, performs surgery or retrieves the right word. General scores often predict broadly; specialised abilities and knowledge determine what performance looks like in a particular domain.
11 · The contemporary ability map
What CHC includes
CHC is best treated as a living taxonomy. The most useful level depends on the question, and no ordinary battery measures the full map.
Contemporary summaries describe roughly 17–18 broad abilities and many narrower ones, depending partly on how proposed domains such as Gkn and Gei are counted. School and clinical batteries concentrate on a subset with established measurement and relevance. The codes are compact labels, not genes, brain modules or guarantees that every published test isolates the ability cleanly.29
| Family | Broad ability | Plain-language description | Interpretive boundary |
|---|---|---|---|
| Acquired knowledge | Gc · Comprehension–knowledge | Breadth and depth of culturally acquired knowledge, including language, general knowledge and verbal comprehension. | Strongly shaped by language, culture, education and opportunity. |
| Gkn · Domain-specific knowledge | Knowledge developed deeply within specialised fields. | A general battery cannot substitute for direct assessment of expertise. | |
| Grw · Reading and writing | Knowledge and skills used in literacy. | Often treated as achievement as well as cognitive ability; conventions vary. | |
| Gq · Quantitative knowledge | Learned mathematical knowledge and procedures. | Distinguish acquired mathematics from novel quantitative reasoning. | |
| Reasoning and control | Gf · Fluid reasoning | Induction, deduction and reasoning with novel relations. | Performance still draws on representation, strategies and task familiarity. |
| Gwm · Working-memory capacity | Maintain and manipulate information under attentional control. | Short tasks mix storage, control, strategy and modality; terminology has changed from older Gsm. | |
| Learning and retrieval | Gl · Learning efficiency | Acquire and consolidate new information efficiently. | Contemporary revisions separate this from fluent retrieval; older sources combine them as Glr. |
| Gr · Retrieval fluency | Rapidly and strategically produce stored ideas or information. | Fluency is not the same as the accuracy or depth of what is retrieved. | |
| Perceptual processing | Gv · Visual processing | Perceive, transform, remember and reason with visual patterns. | Not identical to eyesight; access and motor response demands must be separated. |
| Ga · Auditory processing | Analyse and discriminate sound patterns, including speech-related signals. | Not identical to hearing acuity or musical accomplishment. | |
| Cognitive speed | Gs · Processing speed | Perform simple or overlearned cognitive operations fluently under time constraints. | Scores can include visual scanning, graphomotor, decision and speed–accuracy trade-offs. |
| Gt · Reaction and decision speed | Speed of very simple reactions or choices, often measured in fractions of a second. | Being faster is not always better; accuracy and task purpose matter. | |
| Sensory and motor domains | Go · Olfactory; Gh · Tactile; Gk · Kinesthetic | Broad individual differences in smell, touch and body-sensation processing. | These remain less central in ordinary IQ batteries and parts of the taxonomy are tentative. |
| Gp · Psychomotor ability | Precision and coordination of deliberate bodily movement. | Motor disability must not be mistaken for lower reasoning. | |
| Gps · Psychomotor speed | Speed and fluency of bodily movements. | Separate motor speed from cognitive processing when interpreting timed output. |
Some contemporary versions additionally include emotional intelligence (Gei), provisionally defined as ability in perceiving, understanding, using and managing emotion-related information.
Why the number of abilities varies
Carroll’s original broad stratum, early integrated CHC diagrams and the 2018 revision do not contain identical labels. For example, older short-term memory (Gsm) was reconceptualised as working-memory capacity (Gwm), while long-term storage and retrieval (Glr) was split into learning efficiency (Gl) and retrieval fluency (Gr). Some summaries omit sensory and motor abilities; others distinguish quantitative reasoning differently. A source that says “about ten,” “sixteen,” “seventeen” or “eighteen” may be describing a different edition rather than making a simple error.
A common language for test coverage
CHC helps researchers and practitioners ask whether a battery samples the domains required by the referral question and whether two differently named scores may reflect similar abilities.
Reading every label as a separate brain system
Factor labels are theoretical summaries. Broad abilities correlate, tasks cross-load, factor definitions evolve and neural implementation is distributed. A neat chart should not be mistaken for settled causal anatomy.
A 2023 psychometric-network analysis supported several familiar CHC clusters while suggesting revisions around working memory, attentional control, learning and retrieval. That is exactly how a living taxonomy should behave: evidence can preserve useful regularities while changing their organisation.28
12 · Put the maps side by side
Where the frameworks converge—and where they do not
No framework is “best” without specifying the task. Their claims differ in level, evidence base and intended use.
| Question | Spearman / g | Sternberg | CHC |
|---|---|---|---|
| Primary aim | Explain or summarise positive covariance across mental tests. | Explain intelligent self-management across processing, experience and context. | Classify the hierarchical structure of cognitive abilities. |
| Unit of analysis | Shared individual differences among test performances. | Goal-directed behaviour by a person in a sociocultural environment. | General, broad and narrow dimensions of performance. |
| Core strength | Parsimony and broad prediction. | Attention to novelty, application, goals and context. | Detailed, shared vocabulary for test design and interpretation. |
| Main evidence | Correlations, factor models and criterion prediction. | Task batteries, tacit-knowledge studies, teaching and admissions programmes. | Large factor-analytic literature across cognitive and achievement batteries. |
| Recurring criticism | Reifying a statistical factor and neglecting noncognitive or contextual competence. | Unstable construct separation, method variance, overlap with g and limited independent replication. | Taxonomic complexity, changing labels and overinterpretation of specific profiles. |
| Best practical role | A broad summary when the outcome truly requires general learning or reasoning. | A design lens for asking people to analyse, create and apply in authentic contexts. | A map for balanced assessment and cautious hypotheses about broad- and narrow-ability performance. |
| Does it measure the whole person? | No. | No, even though it deliberately widens the construct. | No. |
What can be integrated
A rigorous assessment can use a general score for broad prediction, CHC-informed domains to understand the task coverage, and Sternberg-inspired work samples to observe generation and application. These sources need not be forced into one scale. A student’s matrix reasoning, vocabulary, practical design project, persistence and interests are different forms of evidence with different reliability and relevance.
What should remain distinct
Analytical, creative and practical performances are not simple synonyms for Gf, divergent thinking and Gc. Likewise, CHC’s broad abilities are not fixed learning styles, and g is not the average of every admirable human capacity. Integration is most honest when it preserves the constructs’ boundaries.
One decision may need three lenses
Structure: what abilities does the evidence sample? Function: can the person use them in a meaningful task? Fairness: did everyone have an accessible, relevant opportunity to demonstrate what the decision requires?
13 · From a theory to a score report
What modern intelligence tests actually measure
A battery is an engineered sample of behaviour. It may be informed by a theory without becoming a pure embodiment of that theory.
A standardised cognitive assessment asks a person to solve selected tasks under controlled rules. Raw performances are converted using an appropriate norm sample, combined into subtest and index scores and sometimes summarised by a full-scale composite. Each step makes the result more interpretable—and each rests on assumptions.
-
Task performance. The person hears, sees or reads instructions; reasons with the content; and responds by speaking, pointing, writing, manipulating objects or using a device.
-
Raw score. Correct responses, quality, time or another rule yields a total. Raw scores have little meaning without the administration rules and reference data.
-
Normed subtest score. Performance is compared with a defined age or other reference group. The norm date, representativeness and version matter.
-
Index or broad composite. Two or more subtests are combined to improve reliability and sample a proposed domain such as verbal comprehension, working memory or processing speed.
-
General composite. A full-scale score combines performance across domains. It is usually the most reliable broad summary, but it can be incomplete when the referral question concerns a barrier or an unusual pattern.
Every reported number is an estimate
A score should be accompanied by a confidence interval, the relevant norms and an explanation of what could have affected access to the construct. A percentile is also an estimate of relative standing, not a statement that the person answered that percentage of items correctly or possesses that percentage of “intelligence.”
How major batteries relate to CHC
| Battery | Connection to contemporary ability theory | What a careful description says | Caution |
|---|---|---|---|
| WISC–V | Primary indexes correspond approximately to comprehension–knowledge, visual processing, fluid reasoning, working memory and processing speed. | CHC was one important influence alongside developmental, neuropsychological and clinical considerations.30 | Independent factor studies often find dominant general variance; an index label does not guarantee a pure or fully separable broad ability. |
| WAIS–5 | The 2024 U.S. edition reports five primary indexes and a seven-subtest Full Scale IQ, with ancillary options for specific questions.31 | Its structure reflects contemporary models and practical assessment goals. | An independent analysis of the U.S. standardisation correlation matrix did not support five latent factors or a distinct Visual Spatial–Fluid Reasoning split; it favoured four group factors under a dominant g. Replication is needed.32 |
| Woodcock–Johnson V | The current edition explicitly uses an updated CHC framework, including separate Long-Term Storage (Gl) and Retrieval Fluency (Gr) clusters.33 | It offers a broad CHC-oriented set of cognitive, oral-language and achievement measures. | No battery covers every CHC ability equally, and publisher alignment with a taxonomy does not by itself establish the distinctness of every cluster. |
| Other batteries | KABC, Stanford–Binet, DAS, nonverbal measures and neuropsychological instruments draw on different combinations of hierarchical, processing and developmental theory. | Selection should follow the referral question, examinee and supporting evidence—not brand prestige. | Scores with similar names can use different tasks and norms; scores with different names may overlap substantially. |
Most contemporary IQ composites are scaled so that the age-normed mean is 100 and the standard deviation is 15, but the test manual—not a generic chart—defines a particular score. Age-norming means a 10-year-old and a 60-year-old can have the same standard score while completing different amounts or kinds of raw performance relative to same-age peers.
Broad, efficient and usually precise
A Full Scale IQ or general-ability composite often provides the most precise single summary of performance on that battery. It is useful only when the construct is relevant, the composite is interpretable for the individual and the included tasks were accessible.
Closer to the question, but not automatically purer
An index can matter when a referral specifically concerns language, working memory, visual processing or speed. Its interpretation needs reliability, unique variance beyond g, adequate indicators and relevant criterion evidence.
Start with the referral question
“Measure intelligence” is too vague. Is the goal to understand access to instruction, intellectual disability, acquired cognitive change, gifted-program fit, workplace learning, language effects or a specific functional problem? The purpose determines the constructs, methods, accommodations and evidence that belong in the assessment.
14 · Strengths, weaknesses and uncertainty
How to interpret a cognitive profile without inventing a diagnosis
Profiles can organise useful hypotheses. Their differences are generally less precise than their component scores, and a pattern rarely identifies one cause.
A low working-memory index does not, by itself, diagnose ADHD. A low processing-speed index does not prove slow thinking in everyday life. A high visual score does not prescribe “visual learning.” Interpretation becomes stronger only when a pattern is reliable, unusual, stable, relevant to the question and corroborated by behaviour, history and other measures.
| Check | Question | Why it matters | Common error |
|---|---|---|---|
| Precision | Are the scores and their difference estimated reliably? | Difference scores combine error from both components. | Treating a one-point rank order as an exact distinction. |
| Base rate | How often does a discrepancy this large occur in the norm group or among people with similar overall scores? | Statistically significant differences can still be common. | Calling ordinary scatter clinically exceptional. |
| Construct | Does the index contain dependable variance specific to the named ability beyond g? | A reliable score can owe much of its reliability to general variance. | Assuming every index is a pure module. |
| Replication | Does the pattern recur across tasks, occasions or qualitatively different evidence? | One unusual subtest can reflect occasion, method or content. | Building a life story around one outlier. |
| Criterion | Does the specific score improve understanding or prediction of the relevant real outcome? | Domain-matched abilities sometimes add modest information beyond g.36 | Inferring educational treatment from structural fit alone. |
| Convergence | Do history, achievement, observation, adaptive behaviour and the person’s report support the hypothesis? | Functional meaning lives outside the score report. | Letting test labels overrule direct evidence. |
Research on WISC–V profiles illustrates the difficulty. Index scores can show high total systematic variance while retaining much less variance uniquely attributable to their intended broad factor once general variance is removed. In simulation, an observed index discrepancy did not consistently reproduce the corresponding latent-factor discrepancy.35 Long-term WISC–V scores in one clinical sample nevertheless showed substantial rank-order stability.34 In a separate one-year study of German students in grades 7–9, specific scores were moderately to highly stable, while incremental prediction beyond general ability was mostly small but sometimes domain-matched.36 The evidence supports disciplined, question-specific interpretation—not the blanket claim that profiles are either diagnostic fingerprints or completely useless.
Say what happened
“Performance was markedly slower on timed visual–motor tasks today” stays close to the evidence and makes relevant conditions visible.
Name alternatives
Possible contributors could include speed, motor output, visual scanning, anxiety, perfectionism, fatigue, medication or unfamiliarity. Test them rather than choosing one by label.
Avoid diagnosis by scatter
No distinctive index pattern, by itself, establishes dyslexia, ADHD, autism, brain injury, giftedness or a particular teaching method.
Cross-battery assessment: shared language is not score exchangeability
CHC codes can reveal that a primary battery under-sampled an ability and help select targeted follow-up measures. But two composites labelled Ga or Gf may use different tasks, norms and weighting. Research has found that supposedly equivalent CHC composites across batteries are not always interchangeable.43 Do not average them casually or invent a new composite without evidence for its reliability and validity.
- Anchor the battery. Use a coherent, preplanned set of measures; add tests only to answer a defined gap.
- Prefer multiple indicators. One task rarely supports a broad-ability inference.
- Check the norms. Age, date, language, country and administration mode matter.
- Preserve task differences. Similar labels do not make distinct methods equivalent.
- Explain uncertainty. Report confidence, base rates and plausible alternatives.
- Return to function. Ask what the person needs to learn, communicate, participate or decide.
Scatter does not automatically invalidate the overall score
An uneven profile can make one global number incomplete for a particular question without making it meaningless. The broad composite may remain the most precise summary of overall test performance. Report both the stable summary and the relevant, well-supported qualifications.
15 · Abilities in time
Development, ageing and changing norms
Cognitive abilities do not rise and fall on one clock. Raw performance, age-normed standing and stability relative to peers answer different questions.
The simple story—fluid intelligence peaks in youth, crystallised intelligence rises through middle age—is directionally useful and individually unreliable. Processing speed, memory, reasoning, vocabulary, social understanding and specialised knowledge have different average trajectories. Education, cohort, health, practice and the exact task change the curve.
Rapid differentiation and learning
Language, knowledge, attention and strategies develop together. Test–retest rank ordering generally becomes more stable with age, but early scores remain especially sensitive to development and access.
Knowledge and control expand
Formal learning, interests and domain practice create increasingly differentiated profiles. A single test occasion can still be affected by sleep, stress, language, engagement and health.
No universal cognitive peak
Some speeded abilities peak relatively early; other tasks plateau later, while vocabulary and expertise may grow for decades. Different measures reached their maxima from late adolescence into the forties or later in a large study.15
Average change is heterogeneous
Novel reasoning and speed tend to decline earlier on average than accumulated knowledge. Individuals vary greatly, and illness, sensory access, cohort and terminal decline complicate any age-only explanation.16
Three kinds of change that are easy to confuse
| Change | What it means | Example | What it does not mean |
|---|---|---|---|
| Raw or absolute | The person can solve more, harder or faster tasks than before. | A child’s vocabulary and reasoning expand substantially over five years. | That the age-normed standard score must rise; peers are developing too. |
| Age-normed | Relative standing compared with same-age peers changes. | A standard score rises after unusually rapid learning or falls after a period of disrupted access. | That an exact latent capacity changed by the same number of “points.” |
| Rank-order stability | People tend to preserve their ordering within a group over time. | A cohort’s scores are strongly correlated across two occasions. | That the group mean stayed constant or that no individual changed. |
| Cohort change | People born or tested in different periods perform differently. | Later cohorts outperform earlier ones on some abstract reasoning measures. | That ageing caused the difference or that every cognitive domain moved together. |
Education can change measured cognitive performance
A meta-analysis of 42 datasets and more than 600,000 participants used designs including policy changes and school-entry cutoffs. It estimated that an additional year of education improved intelligence-test performance by roughly one to five IQ points across the studied designs.14 The range is not a guaranteed dose for an individual, and schooling can affect test strategies, knowledge and reasoning in different proportions. It nevertheless refutes the idea that a highly predictive score must be environmentally inert.
The Flynn effect is not a metronome
Across much of the twentieth century, average scores rose in many populations. A meta-analysis covering nearly four million participants found gains that varied by country, period and ability, with larger average gains in some fluid and spatial measures than in crystallised measures and evidence of slowing.13 Gains have slowed in some recent data. Causes may include education, health, family size, nutrition, technology, abstract problem environments and measurement changes; no single account explains every pattern.
Renorming changes the comparison, not yesterday’s raw performance
When a test receives newer norms, the same raw score can convert to a different standard score because the reference population changed. Near a programme or diagnostic cutoff, that can alter classification.50 Report the edition and norm date, avoid false precision and reconsider decisions when a tiny difference carries a large consequence.
Stable does not mean fixed
A trait can show substantial rank-order stability while every person learns. A measure can be heritable within a population while environmental change shifts its mean. Predictive validity, stability, heritability and malleability are different statistical questions.
16 · Measurement inside society
Culture, language, disability and fairness
No complex cognitive assessment is culture-free. Fairness requires relevant constructs, comparable meaning, accessible administration and defensible consequences.40
Reducing vocabulary or learned-content demands can make a task less dependent on one kind of experience. It cannot remove culture from instructions, symbols, speed expectations, testing relationships, problem-solving conventions, technology, motivation or norms. The responsible aim is not a mythical context-free test; it is evidence that the intended interpretation is sufficiently comparable and useful.
The measurement-invariance ladder46
| Level | What is held comparable | What it generally supports | What it still does not prove |
|---|---|---|---|
| Configural | Approximately the same factor pattern. | The construct may be organised similarly. | That scores have the same units or group means are comparable. |
| Metric | Factor pattern and loadings. | Comparisons of associations or relations with other variables. | Unqualified latent-mean comparisons. |
| Scalar | Loadings plus item or test intercepts/thresholds. | Latent-mean comparison under the model. | Equal opportunity, prediction, treatment or consequences. |
| Strict | Loadings and intercepts plus residual variances. | Stronger score comparability. | That every item is culturally fair or every use is equitable. |
A general factor has been recovered in many non-Western datasets, but recurrence of a structure is not the same as scalar invariance, equivalent norms or fair decisions.45 When invariance fails, the cause might involve translation, schooling, familiarity, response strategy, item content, opportunity or the construct itself. Translation and adaptation therefore require a documented, iterative validation process.53 When invariance holds, unequal access and harmful selection rules may still remain.46
| Concept | What it asks | Correct interpretation | Mistake to avoid |
|---|---|---|---|
| Group mean difference | Do average observed scores differ? | A descriptive result that requires contextual investigation. | Assuming it proves either biological causation or psychometric bias. |
| Differential item functioning | Do examinees from different groups who are matched on the target construct have different response probabilities? | A diagnostic flag requiring substantive study. | Calling every DIF item biased—or assuming no DIF makes the whole system fair. |
| Measurement invariance | Does a model operate comparably across groups? | Evidence for specified score comparisons under assumptions. | Treating model fit as proof of equal opportunity or consequence. |
| Differential prediction | Do prediction slopes, intercepts or errors differ? | Evidence about how scores relate to a criterion by group. | Comparing subgroup correlations alone and declaring fairness. |
| Adverse impact | Does a selection rule produce materially different selection rates? | A legal and policy concern distinct from measurement bias. | Assuming one threshold rule proves either legality or equity. |
| Accessibility | Can the person engage the intended construct without irrelevant barriers? | Equivalent access may require different accommodations. | Calling identical conditions automatically fair. |
Psychometric bias and outcome inequality are different questions. Either can exist without the other, and a responsible evaluation examines both.
Accommodations follow the construct
Make the target ability reachable
Braille, large print, screen readers, sign or communication supports, accessible input, breaks, a reduced-distraction setting or an alternative response can improve access when sensory, motor or communication speed is not the intended construct. Proper accommodations reduce construct-irrelevant barriers; they do not create the target ability.
Document what changed
Extra time may preserve access when speed is incidental, but can change interpretation when speed is the target. A calculator may be appropriate for algebraic reasoning and inappropriate for arithmetic computation. When a modification changes the construct, do not assume the original norms or score meaning still apply; document the change and interpret the result using evidence for that modified use.
U.S. special-education rules require varied tools, prohibit using one measure as the sole criterion, call for assessment in the language and form most likely to yield accurate information, and require that sensory, manual or speaking limitations not contaminate results unless those skills are the target.47 DOJ guidance for ADA-covered examinations says accommodated scores should not be flagged and cautions that prior academic success does not disprove disability.48
Neurodiversity and diagnosis
Cognitive tests can contribute to understanding support needs; they cannot diagnose a complex condition from a score pattern alone. For intellectual disability, current professional definitions require significant limitations in both intellectual functioning and adaptive behaviour, originating before age 22—not an IQ cutoff by itself.39 ADHD, autism, dyslexia, language disorder, sensory disability and brain injury likewise require evidence beyond an intelligence profile.
Do not infer an individual from a group—or a cause from a gap
Within-group variation is usually large. A population average cannot tell you why one person obtained a score, what support would help, or what they can accomplish. Heritability within a population does not identify the cause of a difference between populations, and a score gap cannot partition genetic, educational, economic, health, discrimination and measurement pathways.
17 · Decisions, not labels
How to use intelligence evidence responsibly
A theory earns practical value only when the construct, population, purpose, precision, alternatives and consequences are made explicit.
The professional standard is use-specific validity: evidence must support the proposed interpretation and the decision made from it. Reliability is necessary but not sufficient; a test can consistently measure the wrong thing. A publisher’s general validation does not transfer automatically to every school, clinic, language, job or cutoff.3738
| Setting | Define first | Use evidence well | Avoid |
|---|---|---|---|
| Education | Is the target achievement, reasoning, readiness, disability-related need, access or programme fit? | Use multiple sources; when relevant, combine appropriate cognitive and achievement measures with evidence about opportunity to learn, classroom work, observation and family and learner input; provide review and appeal. | Sole-score placement, permanent tracking, diagnosing from a profile or testing knowledge that was never taught. |
| Gifted identification | Which talents and services is the programme designed to develop? | Use broad or universal access to screening where referrals create barriers; include domain achievement, work and local opportunity where relevant. Universal screening has increased identification of underserved pupils in a natural experiment.49 | Treating one cutoff as a natural boundary, equating giftedness with worth or assuming a broadened test is fair without outcome evidence. |
| Clinical assessment | What functional change, developmental question or support need prompted the evaluation? | Use appropriate norms and language; integrate history, observation, adaptive behaviour, health, sensory access, medication, testing conditions, engagement and daily functioning.52 | One-score diagnosis, obsolete norms, automatic reports presented as judgement or a percentile presented as exact. |
| Employment | Which demonstrable job demands require learning, reasoning or specific abilities? | Conduct job analysis; validate the actual cutoff or ranking; monitor adverse impact and differential prediction; provide a clear, confidential process for requesting reasonable accommodations; compare substantially equally effective alternatives with less adverse impact.41 | Generic “smartness” screening, universal validity coefficients, hidden proxy variables or using a vendor claim as local validation. |
| Research and policy | What construct, population-level question and causal estimand are intended? | Pre-register models, test measurement equivalence, report uncertainty and heterogeneity, protect privacy and separate description from causal and normative claims. | Ranking social groups as essences, interpreting latent factors as direct biology or converting a mean association into an individual rule. |
A twelve-question audit before a score affects a life
- Construct: What exact ability or performance is intended?
- Purpose: What decision will the result inform?
- Population: Does validation cover this age, language, culture and setting?
- Norms: Are the reference data current and appropriate?
- Precision: What is the confidence interval, especially at a cutoff?
- Access: Could language, disability, interface or speed add irrelevant difficulty?
- Accommodation: Which change preserves the target construct?
- Convergence: What other evidence supports or challenges the result?
- Prediction: Is the criterion relevant, reliable and independently validated?
- Fairness: Were DIF, invariance, differential prediction, adverse impact and access considered separately?
- Alternative: Is a less exclusionary method similarly effective?
- Rights and recourse: Where required, are consent or authorisation, privacy and data access handled appropriately; will results and limits be explained accessibly; and can the person correct, appeal or retest?
Match the method to what you want to know
Use a balanced cognitive composite
When the criterion requires learning across varied content, a reliable general score can supply relevant evidence. Pair it with direct achievement or work evidence and do not treat a modest group correlation as certain individual prediction.
Sample products across occasions
Divergent-thinking tasks indicate creative potential rather than completed creative achievement. Multiple prompts, trained raters, domain products, portfolios, expertise, persistence and implementation provide a stronger picture.51
Build authentic, analysed scenarios
Create situations from the actual domain, use diverse experts, document scoring rationales and test reliability, subgroup outcomes and incremental value. Do not generalise “street smarts” beyond the sampled situations.
Use current occupational evidence
A 2026 synthesis of 212 samples and 109,396 people estimated contemporary operational validity for general mental ability and job performance at roughly .18 to .22 and recommended .20 as a broad estimate.10 That remains useful at scale, but it is far below the classic .51 headline repeated as a universal constant. Job-demand mix, criterion choice, range-restriction correction and local design still matter.
The decision-maker retains responsibility
A school, psychologist, employer or agency cannot outsource the validity and fairness of its actual use to a theory, test publisher or algorithm.374152 Monitor outcomes as populations, curricula, jobs, norms and technology change, and retire a rule when its justification no longer holds.
18 · The debate is still alive
Open questions and future directions
The next advance will need more than a new label. It must connect structure, mechanism, development, ecology and fair use.
A compact history of widening the lens
| Period | Development | Lasting contribution | Myth to avoid |
|---|---|---|---|
| 1904 | Spearman formalises a general-plus-specific account. | The positive manifold and a testable covariance model. | One paper proved a single biological mental energy. |
| 1930s | Thurstone foregrounds primary mental abilities.54 | Verbal, number, spatial, memory and other domains deserve measurement. | The abilities are wholly independent or permanently eliminate g.24 |
| 1940s–1960s | Cattell and Horn develop and expand Gf–Gc theory. | Novel reasoning and acquired knowledge are distinguishable and develop differently. | Fluid tests are culture-free, or Horn accepted a mandatory top-level g. |
| 1980s | Gardner and Sternberg challenge narrow definitions in different ways. | Human capability includes contextual, creative and culturally valued performance. | Multiple intelligences, triarchic abilities and learning styles are the same claim—or equally validated psychometric factors. |
| 1993 | Carroll synthesises factor-analytic evidence into three strata. | A comprehensive hierarchy of narrow, broad and general covariance. | The strata are uncontested literal brain layers. |
| Late 1990s onward | CHC becomes a common taxonomy for test development. | Shared language makes battery coverage and research easier to compare. | The taxonomy is frozen or every same-coded score is interchangeable. |
| 2000s onward | Mutualism, process-overlap and network accounts formalise alternatives. | The manifold may emerge from development and interacting processes. | Questioning a causal g denies the general statistical pattern. |
Six questions that could move the science forward
Which processes generate general covariance?
Models should make competing predictions under changes in attention, memory load, knowledge and strategy—not merely reproduce the same correlation matrix.
How do abilities become coupled?
Dense longitudinal studies can test whether early processes cause, constrain or mutually reinforce one another and how schools, health and opportunity alter the network.
What survives outside the test room?
Research needs authentic tasks, delayed transfer, daily-life function and long-term outcomes—not only faster completion of closely related items.
When do broad and narrow abilities add value?
Pre-registered prediction and intervention studies can identify when a domain score changes a conclusion beyond a general composite and direct achievement evidence.
Which constructs travel across contexts?
Translation, cognitive interviewing, local norms, invariance, prediction and consequences must be studied together rather than using one statistical test as a fairness certificate.
Who benefits from a classification?
A technically accurate ranking can still narrow opportunity. Research should track false decisions, access, privacy, appeal and whether assessment actually improves support.
Dynamic and computational assessment
Traditional tests sample what a person can do now under standard support. Dynamic assessment asks how performance changes with prompts, teaching or feedback, potentially revealing learning processes as well as attained level. Computer-adaptive testing can target difficulty efficiently, and process data can record timing or strategy traces. These methods also create new dependencies: device familiarity, proprietary models, surveillance, security and algorithmic drift. More data do not remove the need to define the construct or validate the decision.
Biology should constrain—not replace—psychometrics
Brain imaging, genetics and computational neuroscience may clarify distributed mechanisms that support reasoning, memory and knowledge. But a neural correlate of a score does not prove the factor is one biological cause; biological measures have their own reliability, selection and interpretation problems. Strong theory will link levels while preserving the difference between a statistical summary, cognitive process, developmental pathway and social outcome.
The most useful future is plural and disciplined
Retain general scores when broad prediction is the goal. Use CHC domains when specific coverage matters. Observe creative and practical performance when the context demands it. Add motivation, knowledge, personality, adaptive functioning and opportunity as separate evidence rather than stretching “intelligence” until it means every valued human quality.
19 · Questions people ask
Clear answers to common intelligence-theory questions
The right answer usually begins by separating the observed score, the statistical model, the proposed mechanism and the decision someone wants to make.
Is g the same thing as IQ?
No. g is a latent statistical dimension used to summarise covariance among cognitive tasks. An IQ is a norm-referenced score produced by a particular battery and scoring system. A Full Scale IQ is usually strongly related to psychometric g, but it also reflects the battery’s task content, weighting, broad and specific abilities, measurement error and administration conditions.
Does the positive manifold prove there is only one intelligence?
No. It shows that diverse cognitive performances tend to correlate positively. A common-cause g model can represent that pattern, but sampling, process-overlap, mutualism and network models can also generate it. The manifold is strong evidence that abilities are not wholly independent; it is not proof that one mental substance causes every correlation.
How many intelligences are there?
There is no theory-neutral count. At a broad psychometric level, one can summarise performance with a general factor. At a finer level, CHC distinguishes numerous broad and narrow abilities. Sternberg describes modes of intelligent action, while Gardner proposed culturally valued intelligences. The useful number depends on the question, evidence and level of description—not on discovering a final set of boxes.
Do Spearman and CHC contradict one another?
Not necessarily. Hierarchical CHC models retain a general factor while also describing broad and narrow abilities. The historical tension lies more clearly between Horn, who resisted a substantive higher-order g, and Carroll, whose three-stratum model included it. Modern CHC is an umbrella taxonomy that can tolerate more than one account of why abilities correlate.
Did Sternberg show that analytical, creative and practical intelligence are independent?
No settled evidence supports three cleanly independent faculties. Some performance tasks add information and separate better than multiple-choice tasks, but many studies find overlap with general cognitive ability, content and method. Sternberg’s framework remains valuable for broadening what education and assessment invite people to do; its factor structure and incremental validity should be judged study by study.
Is CHC the scientifically proven “best” theory?
CHC is the most influential comprehensive taxonomy in much contemporary psychometric assessment, supported by a large structural literature. “Best” becomes misleading when it implies a complete causal theory or a frozen final map. Exact abilities and labels continue to change; no battery measures them all; and process, developmental, contextual and network theories address questions the taxonomy does not settle.
What is the difference between fluid and crystallised intelligence?
Fluid reasoning concerns discovering and applying relations in relatively novel problems. Crystallised or comprehension–knowledge ability concerns language and knowledge developed through culture and education. Most meaningful tasks require both. A fluid task can reduce explicit prior-content demands but cannot be literally culture- or experience-free; crystallised intelligence is organised understanding, not mere trivia.
At what age does intelligence peak?
There is no single age. Processing speed and some memory or reasoning tasks tend to peak earlier; vocabulary, knowledge and some socially informed judgements mature later. Even within these categories, tasks differ. Raw ability, same-age normed standing and position relative to peers are also different. Health, education, cohort and practice change the trajectory.
Can intelligence scores change?
Yes. Scores can change through development, education, illness or recovery, sensory access, language, practice, motivation, test version and measurement error. Education has a causal effect on cognitive-test performance in quasi-experimental evidence.14 Moderate or high stability means people often retain a similar ranking relative to peers; it does not mean a person cannot learn or that a score is a biological ceiling.
Are nonverbal or “culture-fair” tests culture-free?
No. They may reduce vocabulary and specific curricular knowledge, which can be useful. They still depend on instructions, symbols, visual access, test-taking conventions, speed expectations, strategies and norms. Cross-cultural interpretation needs appropriate adaptation, local evidence and examination of measurement equivalence and consequences.
Can a cognitive profile diagnose ADHD, autism, dyslexia or brain injury?
Not by itself. Conditions can influence performance, and a profile may suggest hypotheses worth following. Similar patterns can arise from many causes, and people with the same diagnosis can have different profiles. Diagnosis requires relevant developmental, medical, educational, behavioural and functional evidence using condition-specific standards.
How large must a score difference be to matter?
No universal number works. Ask whether the difference exceeds expected measurement error, how often it occurs in the relevant norm group, whether it replicates, whether the compared scores measure sufficiently distinct constructs and whether it changes understanding of the real referral question. Statistical significance alone does not establish clinical or educational importance.
Do testing accommodations make scores unfairly high?
Appropriate accommodations remove barriers unrelated to the target construct. A screen reader may allow access to reasoning when ordinary print is incidental; extra time may be inappropriate if speed itself is being measured. Fairness means equivalent access to the intended construct, not necessarily identical administration. Changes and interpretive limits should be documented.
Does a high IQ guarantee academic, career or life success?
No. General cognitive ability predicts learning and some performance outcomes on average, but distributions overlap widely. Knowledge, motivation, self-regulation, interests, creativity, health, social support, discrimination, resources, opportunity and chance also shape outcomes. A correlation improves group-level prediction; it does not write one person’s future.
Does a group score difference prove that a test is biased?
No, and the reverse is also false: lack of psychometric bias does not prove a selection system is equitable. A gap is a descriptive finding. Investigate measurement equivalence, item functioning, prediction, opportunity to learn, language, access, norms, selection rules and consequences separately. Group averages cannot identify an individual’s cause or potential.
Are Gardner’s multiple intelligences the same as Sternberg’s theory?
No. Gardner proposed several relatively autonomous human capacities, initially drawing strongly on neuropsychology and culturally valued performance.56 Sternberg organised intelligence around components, experience and context, later emphasising successful and adaptive action. Both challenge overly narrow definitions, but they propose different constructs; psychometric tests of the claimed independence among Gardner’s intelligences have also produced substantial general-factor overlap.57 Neither theory is equivalent to the claim that lessons should be matched to fixed learning styles.
Can an online IQ test provide a diagnosis or official score?
An online score is not diagnostic merely because it is labelled “IQ.” Professionally supervised remote assessment may support an official use only when evidence covers that instrument, delivery mode, population and decision, with secure standardisation, identity and access checks, and contextual interpretation. A recreational website score should not guide diagnosis, placement or employment.
Will adaptive testing or AI eliminate bias?
No. Adaptive delivery can target item difficulty and reduce testing time; automated scoring can improve consistency for defined responses. Both inherit choices about constructs, training data, items, interfaces, stopping rules and criteria. They can introduce new inequities through device access, language models, surveillance or drift. Automation changes where judgement occurs; it does not remove the need for validation, transparency and appeal.
The conclusion
Intelligence has structure, context and consequences
The theories become more useful—not less—when their claims are kept at the right level.
Spearman identified a general pattern that any serious account of cognitive differences must confront. Cattell and Horn differentiated reasoning from acquired knowledge and widened the map of broad abilities. Carroll organised decades of evidence into a hierarchy, while CHC gave assessment a shared, evolving vocabulary. Sternberg insisted that intelligent behaviour also involves novelty, application, goals and environment.
Use the broad pattern honestly
General cognitive scores can be reliable and predictive. Report their uncertainty, norms, task demands and validated purpose. Do not reify an index into a fixed essence.
Look closer when the question requires it
Broad and narrow abilities, knowledge and performance context can matter. Interpret profiles as corroborated hypotheses rather than diagnostic fingerprints.
Keep values outside the score
Assessment can guide support or opportunity; it cannot decide a person’s dignity. Fair decisions require access, relevant evidence, proportional consequences and recourse.
There is no need to choose between recognising general cognitive ability and appreciating diverse forms of competence. The disciplined position is additive: use g for what it summarises; use CHC for the structure it maps; use contextual tasks for what they uniquely reveal; and measure achievement, creativity, motivation, adaptive behaviour, knowledge and opportunity directly when those are what matter.
The wisest theory of intelligence is not the one that turns every human difference into one number. It is the one that tells us which evidence a number contains, which evidence it leaves out and how to act without confusing prediction with destiny.
Primary research, technical documentation and official guidance
Sources
Link behaviour: External source links open in a new tab.
- Spearman, C. (1904), “General Intelligence,” Objectively Determined and Measured, American Journal of Psychology.
- Carroll, J. B. (1993), A Theory of Cognitive Abilities: The Three-Stratum Theory, in Human Cognitive Abilities.
- McGrew, K. S. (2009), CHC theory and the human cognitive abilities project: Standing on the shoulders of the giants of psychometric intelligence research, Intelligence.
- McGrew, K. S. (2023), Carroll’s Three-Stratum (3S) Cognitive Ability Theory at 30 Years: Impact, 3S-CHC Theory Clarification, Structural Replication, and Cognitive–Achievement Psychometric Network Analysis Extension, Journal of Intelligence.
- Savi, A. O. et al. (2019), The Wiring of Intelligence, Perspectives on Psychological Science.
- van der Maas, H. L. J. et al. (2006), A dynamical model of general intelligence: the positive manifold of intelligence by mutualism, Psychological Review.
- van der Maas, H. L. J. et al. (2017), Network Models for Cognitive Development and Intelligence, Journal of Intelligence.
- Kovacs, K. & Conway, A. R. A. (2016), Process Overlap Theory: A Unified Account of the General Factor of Intelligence, Psychological Inquiry.
- Kan, K.-J. et al. (2024), Why Do Bi-Factor Models Outperform Higher-Order g Factor Models? A Network Perspective, Journal of Intelligence.
- Demeke, S. et al. (2026), Is General Mental Ability Still the Best Predictor of Job Performance? Integrating Contemporary Meta-Analytic Evidence, Journal of Business and Psychology.
- Roth, B. et al. (2015), Intelligence and school grades: a meta-analysis, Intelligence.
- Deary, I. J., Hill, W. D. & Gale, C. R. (2021), Intelligence, health and death, Nature Human Behaviour.
- Pietschnig, J. & Voracek, M. (2015), One Century of Global IQ Gains: A Formal Meta-Analysis of the Flynn Effect (1909–2013), Perspectives on Psychological Science.
- Ritchie, S. J. & Tucker-Drob, E. M. (2018), How Much Does Education Improve Intelligence? A Meta-Analysis, Psychological Science.
- Hartshorne, J. K. & Germine, L. T. (2015), When Does Cognitive Functioning Peak? The Asynchronous Rise and Fall of Different Cognitive Abilities Across the Life Span, Psychological Science.
- Tucker-Drob, E. M. et al. (2022), A strong dependency between changes in fluid and crystallized abilities in human cognitive aging, Science Advances.
- Sternberg, R. J. (1999), The Theory of Successful Intelligence, Review of General Psychology.
- Sternberg, R. J. & the Rainbow Project Collaborators (2006), The Rainbow Project: Enhancing the SAT through assessments of analytical, practical, and creative skills, Intelligence.
- Hedlund, J. et al. (2006), Assessing practical intelligence in business school admissions: A supplement to the graduate management admissions test, Learning and Individual Differences.
- McDaniel, M. A. & Whetzel, D. L. (2005), Situational judgment test research: Informing the debate on practical intelligence theory, Intelligence.
- Chooi, W.-T., Long, H. E. & Thompson, L. A. (2014), The Sternberg Triarchic Abilities Test (Level-H) Is a Measure of g, Journal of Intelligence.
- Gubbels, J. et al. (2016), The Aurora-a Battery as an Assessment of Triarchic Intellectual Abilities in Upper Primary Grades, Gifted Child Quarterly.
- Sternberg, R. J. et al. (2014), Testing the Theory of Successful Intelligence in Teaching Grade 4 Language Arts, Mathematics, and Science, Journal of Educational Psychology.
- Saw, K. N. N. & Han, B. (2021), Effectiveness of successful intelligence training program: A meta-analysis, PsyCh Journal.
- Sternberg, R. J. (2021), Adaptive Intelligence: Intelligence Is Not a Personal Trait but Rather a Person × Task × Situation Interaction, Journal of Intelligence.
- Cattell, R. B. (1943), The measurement of adult intelligence, Psychological Bulletin.
- Horn, J. L. & Cattell, R. B. (1966), Refinement and test of the theory of fluid and crystallized general intelligences, Journal of Educational Psychology.
- McGrew, K. S. et al. (2023), A Psychometric Network Analysis of CHC Intelligence Measures: Implications for Research, Theory, and Interpretation of Broad CHC Scores “Beyond g”, Journal of Intelligence.
- Schneider, W. J. & McGrew, K. S. (2018), The Cattell–Horn–Carroll Theory of Cognitive Abilities, in Contemporary Intellectual Assessment, 4th ed.
- Pearson, official WISC–V information.
- Pearson, official WAIS–5 information.
- Canivez, G. L. et al. (2026), Construct Validity of the WAIS-5: Complementary Exploratory and Confirmatory Factor Analyses of the 20 Primary and Secondary Subtests, Assessment.
- LaForte, E. M., Dailey, D. & McGrew, K. S. (2025), WJ V Technical Abstract, Riverside Insights.
- Watkins, M. W. et al. (2022; online 2021), Long-term stability of Wechsler Intelligence Scale for Children–Fifth Edition scores in a clinical sample, Applied Neuropsychology: Child.
- de Jong, P. F. (2023), The Validity of WISC–V Profiles of Strengths and Weaknesses, Journal of Psychoeducational Assessment.
- Breit, M. et al. (2024), How useful are specific cognitive ability scores? An investigation of their stability and incremental validity beyond general intelligence, Intelligence.
- AERA, APA & NCME (2014), Standards for Educational and Psychological Testing.
- International Test Commission, Guidelines for Test Use and Guidelines for Translating and Adapting Tests.
- American Association on Intellectual and Developmental Disabilities, Definition of Intellectual Disability.
- Pedraza, O. & Mungas, D. (2008), Measurement in Cross-Cultural Neuropsychology, Neuropsychology Review.
- U.S. Equal Employment Opportunity Commission, Employment Tests and Selection Procedures.
- Brown, R. E. (2016), Hebb and Cattell: The Genesis of the Theory of Fluid and Crystallized Intelligence, Frontiers in Human Neuroscience.
- Floyd, R. G. et al. (2005), Are Cattell–Horn–Carroll broad ability composite scores exchangeable across batteries?, School Psychology Review.
- Rodriguez, A., Reise, S. P. & Haviland, M. G. (2016), Applying bifactor statistical indices in the evaluation of psychological measures, Journal of Personality Assessment.
- Warne, R. T. & Burningham, C. (2019), Spearman’s g Found in 31 Non-Western Nations: Strong Evidence That g Is a Universal Phenomenon, Psychological Bulletin.
- Wicherts, J. M. (2016), The importance of measurement invariance in neurocognitive ability testing, The Clinical Neuropsychologist.
- U.S. Department of Education, 34 CFR §300.304—Evaluation procedures under IDEA.
- U.S. Department of Justice, ADA Requirements: Testing Accommodations.
- Card, D. & Giuliano, L. (2016), Can universal screening increase the representation of low-income and minority students in gifted education?, PNAS.
- Kanaya, T., Scullin, M. H. & Ceci, S. J. (2003), The Flynn effect and U.S. policies: the impact of rising IQ scores on special education classifications, American Psychologist.
- Runco, M. A. & Acar, S. (2012), Divergent Thinking as an Indicator of Creative Potential, Creativity Research Journal.
- American Psychological Association, Ethical Principles of Psychologists and Code of Conduct, especially assessment standards 9.01–9.10.
- International Test Commission (2018), ITC Guidelines for Translating and Adapting Tests, second edition, International Journal of Testing.
- Thurstone, L. L. (1938), Primary Mental Abilities.
- Sternberg, R. J. (1985), Beyond IQ: A Triarchic Theory of Human Intelligence, Cambridge University Press.
- Gardner, H. (1983), Frames of Mind: The Theory of Multiple Intelligences, Basic Books.
- Visser, B. A., Ashton, M. C. & Vernon, P. A. (2006), Beyond g: Putting multiple intelligences theory to the test, Intelligence.
- Breit, M., Conway, A. R. A. & Kovacs, K. (2025), 99 Ways to Obtain a General Factor of Intelligence—And They Are Not All Created Equal, Journal of Psychoeducational Assessment.
Evidence reviewed through 4 September 2026. Source dates do not indicate equal evidential weight: original theories establish historical claims; psychometric studies test structure; meta-analyses synthesise associations; and professional or legal guidance governs particular uses. Readers applying an assessment should consult the current manual and requirements in their own jurisdiction.
Continue the Intelligence Unleashed series
Explore intelligence from foundations to application
Written as an evidence-aware educational guide. Last evidence review: 4 September 2026.