Measuring Intelligence
Linas JuozenasShare
Measure intelligence without shrinking the mind
Intelligence matters. Stronger reasoning, memory, knowledge and learning capacity can help people understand more, learn faster, solve harder problems and navigate life with greater agency. Good IQ tests can measure important parts of that capacity. They become most useful when we read them as carefully constructed estimates—not identities, verdicts or ceilings.
A useful signal, interpreted in context
An intelligence-test result is a sample of performance on selected cognitive tasks under defined conditions. Well-designed batteries can provide evidence about broad abilities, possible relative strengths and support needs, and they can help guide education or clinical follow-up. Their scores are meaningful precisely because their claims are limited and testable.1
The right question is not “Does IQ matter, or does the whole person matter?” Both do. Cognitive ability influences learning and many life outcomes, while dignity, moral worth, purpose and the total richness of a mind cannot be reduced to a score.
Exceptional intelligence can carry exceptional importance
A higher valid score is meaningful evidence of stronger performance in the abilities the test measures. That difference matters and should not be trivialized. Joined with curiosity, discipline, knowledge and wisdom, exceptional cognitive ability can make a person unusually important to other people and to society—helping them see hidden connections, solve difficult problems, preserve knowledge and guide work that might otherwise never be done.
A score alone does not prove wisdom, character or accomplishment, and basic human dignity belongs equally to everyone. But equal dignity does not require pretending that all abilities or contributions are identical. Rare intelligence and hard-earned intellectual achievement deserve genuine recognition, protection and respect.
Reading a familiar IQ scale
Many contemporary composite scores are normed to a mean of 100 and a standard deviation of 15. The illustration below shows approximate positions under a normal distribution; labels and interpretations vary by test, edition, age and norm group.
Norm-referenced
The score compares performance with an appropriate standardization sample, usually people of a similar age. It does not directly count a fixed quantity of “intelligence.”
An estimate, not a verdict
Reliability is never perfect. Reports should show a confidence interval and explain whether small differences are larger than expected measurement error.
A profile may add important context
Reliable differences across broad domains may suggest hypotheses about relative strengths and support needs. Isolated subtest scatter should not be treated as diagnostic.
What an IQ test actually measures
An IQ battery does not peer directly into a mind. It uses standardized tasks to elicit behavior, scores that behavior consistently and evaluates whether the resulting pattern supports a particular interpretation.
Most comprehensive intelligence tests sample several cognitive domains: reasoning with new information, acquired verbal knowledge, visual-spatial analysis, working memory and processing efficiency. Performance across those tasks is positively correlated, allowing a broad composite to summarize a substantial common dimension of cognitive functioning.2
That composite is powerful because it condenses information. It is incomplete for the same reason. Motivation, curiosity, creativity, emotional skills, practical knowledge, perseverance, health, values and opportunity can all influence what a person does with their abilities, yet are not identical to general cognitive ability.
Construct before number
Every score needs a sentence that begins: “This result is evidence about…” The ending should name the construct, population, conditions and intended use. If the sentence becomes “this number tells us everything about this person,” the interpretation has exceeded the evidence.
A short language key: a construct is the ability being inferred; a battery is a collection of tasks; a subtest samples a narrower skill; an index combines related subtests; a composite summarizes a broader pattern; and the norm group is the reference sample used to interpret standing.
| Term | What it means | What it does not mean |
|---|---|---|
| Raw score | Items answered correctly or points earned before norm conversion. | A directly comparable measure across ages, editions or different tests. |
| Standard score | A transformed score located within a reference distribution. | A percentage of all intelligence possessed. |
| Percentile rank | The approximate percentage of the norm group scoring at or below that result. | Percent correct, or the percentage by which one person is “smarter.” |
| Confidence interval | A range acknowledging uncertainty around an observed score. | A guarantee that a hidden “true score” is permanently fixed inside the range. |
How a trustworthy test earns its authority
A polished interface, difficult puzzles or an “AI-powered” label do not make an intelligence test valid. Trust must be built with documented evidence.
Clear construct
The developer defines which abilities the test is intended to measure, what it deliberately leaves out and which interpretations are supported.
Standard administration
Instructions, timing, prompts, scoring rules and testing conditions are controlled so results remain comparable.
Appropriate norms
A sufficiently large, relevant sample anchors interpretation. Old or poorly matched norms can distort standing.
Reliability
Scores are sufficiently consistent across items, forms, raters or occasions for the intended decision. Required precision rises with the stakes.
Validity evidence
Content, response processes, internal structure, relations with other variables and consequences support the proposed interpretation and use.
Fairness & accessibility
Developers investigate barriers, subgroup functioning and accommodations, then state where evidence is weak or use is inappropriate.
Reliability is necessary—but not sufficient
A bathroom scale that is always five kilograms wrong is reliable but inaccurate. Likewise, a cognitive test may rank performance consistently without justifying every interpretation made from it. Professional standards treat validity as evidence supporting a specific interpretation for a specific use, not a permanent sticker attached to the instrument.1
Standard error and confidence intervals
Because testing samples behavior, an observed score includes uncertainty. A confidence interval expresses that imprecision using the test’s reliability and score scale. The interval does not capture every possible source of error—such as poor sleep, misunderstood instructions, language mismatch or an unrepresentative testing session—but it is more honest than presenting a single number as exact.
A simplified example
In classical test theory, the standard error of measurement can be illustrated as SD × √(1 − reliability). On an SD-15 scale with reliability .95, that is about 3.35 points. In this simplified example, an observed 120 would have an approximate 95% interval of 113–127.
Why “simplified” matters
Real interpretation should use the manual’s age-, score- and edition-specific error estimates. Precision may vary across the scale, especially near extremes. Change and difference scores also have their own error and cannot be judged from either score’s reliability alone.1
When two index scores differ, the examiner should ask more than whether the difference reaches a statistical threshold. How often does a difference that large occur in the norm sample? Does it persist across related tasks and history? Is it relevant to the referral question? A dramatic-looking profile can otherwise become a story built from ordinary variation.
A tool for support, a history of progress—and a warning about power
Intelligence testing did not emerge in one finished form. Its history contains humane educational aims, major technical advances, unjustified extrapolations and coercive misuse. Understanding all four helps us use measurement better today.
A 30-task precursor develops into a fuller age-graded scale for identifying children who may need different instruction.
William Stern proposes “intelligence quotient”: estimated mental age divided by chronological age.
Lewis Terman adapts and expands the scale for the United States and popularizes whole-number ratio IQ.
Adult testing combines verbal and performance tasks and advances age-normed deviation scoring.
Modern batteries report composites and domain scores against defined norms, with reliability and validity evidence.
Binet and Simon: present performance, not permanent destiny
In 1905 Alfred Binet and Théodore Simon published an influential 30-task precursor; their fuller 1908 revision organized tasks by age level and more closely resembled the practical intelligence scale later textbooks describe. The work grew from French debates about how children struggling in ordinary classrooms should be identified and educated. Its tasks sampled abilities such as comprehension, memory and judgment so that current needs could be recognized more systematically.34
This was not yet a modern IQ test, nor did it produce Stern’s later quotient. Binet described the scale as a practical hierarchy rather than a literal measurement like physical length, and he opposed treating low performance as proof of an unchangeable quantity. That caution should not be romanticized into a claim that every ability can grow without limit. Its enduring value is simpler: a result should guide support, not become a sentence.
Stern, Terman and the number called IQ
In 1912 German psychologist William Stern proposed the term Intelligenzquotient: estimated mental age divided by chronological age. Lewis Terman’s 1916 Stanford revision adopted the quotient, multiplied it by 100 and helped make “IQ” prominent in American education.56 These early scores were ratio IQs. They were not yet today’s deviation scores, which locate performance within an age-matched normative distribution.
Mass testing and the danger of extrapolation
During the First World War, Army Alpha tested English-literate recruits in groups, while the largely pictorial Army Beta was designed for recruits with limited literacy or English. The program demonstrated that psychological tests could be administered at enormous scale and served practical classification purposes.7
The historical failure came when scores affected by language, schooling, acculturation, selection and testing conditions were stretched into claims about innate racial or national hierarchy. Carl Brigham’s 1923 interpretation of Army data became a prominent example; by 1930, he rejected the comparative foundation of those conclusions. Racialized test interpretations gave scientific-looking authority to a wider eugenic and nativist culture, although claims that Army testing directly caused the United States Immigration Act of 1924 overstate the archival evidence.89
The lesson is precision, rights and agency
Eugenics ranked lives and removed agency through segregation, exclusion and reproductive coercion. Helping people strengthen cognition through education, nutrition, health care, safety and freely chosen practice does the opposite: it expands agency. Celebrating intellectual growth is compatible with human equality when ability is cultivated and used well—not confused with anyone’s right to dignity.
Wechsler and the modern age-normed profile
David Wechsler’s 1939 Wechsler–Bellevue scale was built for adult clinical use. It combined verbal and performance tasks and compared adults with others of the same age, helping shift the field away from literal mental-age ratios. The WISC followed for children in 1949 and the first WAIS in 1955.10 Later batteries increasingly paired a broad composite with interpretable domain scores—an architecture that remains central today.
How general and specific abilities fit together
A widely used psychometric view is hierarchical: cognitive tasks share something broad, yet different tasks also require more specific capacities.
The general factor, g
People who perform well on one sufficiently demanding cognitive task tend, on average, to perform well on others. Factor analysis summarizes this positive pattern with a general factor commonly called g. The pattern is one of psychology’s most replicated findings.2
g is a statistical regularity, not a tiny substance located in one brain region, and theories differ about the mechanisms that produce it. Its usefulness does not require pretending those mechanisms are completely settled.
Broad and narrow abilities
Hierarchical frameworks such as Cattell–Horn–Carroll organize performance into a general level, broad abilities and narrower skills. Fluid reasoning and acquired knowledge are prominent examples; other domains may include visual processing, short-term working memory, retrieval fluency and processing speed.11
CHC is best understood as a taxonomy informed by psychometric evidence—not a claim that every battery implements the same map perfectly.
A composite asks, “What broad pattern is most defensible?” A profile asks, “Are any broad-domain differences dependable, uncommon and corroborated enough to guide a hypothesis?”
Specific scores warrant attention when sufficiently reliable, represented by more than one task and shown to add relevant information beyond the composite for the intended use. They deserve less attention when an apparent peak or dip depends on one short subtest. Broad scores are often more reliable; narrower scores may guide targeted follow-up, but instructional or clinical utility should not be assumed.
What contemporary intelligence tests may assess
Batteries differ by age range, purpose and theory. The following domains illustrate common targets rather than a universal structure shared by every test.
| Domain | Typical demand | Interpretive caution |
|---|---|---|
| Fluid reasoning | Discover rules, infer relations and solve unfamiliar problems. | Prior familiarity with puzzle formats, strategy and instruction still matters. |
| Acquired knowledge / verbal comprehension | Use vocabulary, concepts and culturally transmitted knowledge. | Language, schooling and opportunity to learn are integral to performance. |
| Visual-spatial processing | Analyze, transform or construct spatial patterns. | Vision, motor response and experience with diagrams may affect results. |
| Working memory | Hold and manipulate information over short intervals. | Attention, hearing, anxiety and strategy can influence a brief sample. |
| Processing speed | Perform simple visual decisions accurately and efficiently under time limits. | Motor speed, visual access and cautious response style may enter the score. |
| Long-term retrieval / fluency | Retrieve learned associations or generate information efficiently. | Knowledge base, language and task-specific practice influence output. |
Examples of professionally developed instruments include the Stanford–Binet, Wechsler scales, Woodcock–Johnson cognitive battery, Kaufman batteries and Raven’s matrices. A brand name alone is not enough. The appropriate edition, age range, language, norms, examiner qualifications and intended use all matter. The current adult Wechsler battery, for example, is WAIS–5 in the United States, but availability and approved versions vary by country.12
Screening is not comprehensive assessment
A short screener can flag whether fuller evaluation may be useful. A nonverbal matrix task can sample abstract reasoning with reduced language demand. An online quiz can be entertaining or educational. None automatically carries the breadth, supervised conditions, secure items, representative norms and interpretive evidence of a professionally administered battery.
How to read a cognitive assessment report
A responsible report turns numbers into bounded conclusions, checks rival explanations and connects findings to useful action.
Start with the referral question
- Why was the assessment requested?
- Which test and edition were used, and is it appropriate here?
- Who was in the norm group, and when were norms collected?
- Were administration and scoring standardized?
- Was performance representative of ordinary functioning?
Then read from broad to specific
- Composite score and confidence interval
- Reliable, meaningful index differences
- Converging subtests—not isolated highs or lows
- Behavioral observations and testing conditions
- Agreement with learning, work and daily-life history
Full-scale score versus index profile
A broad composite is often the most reliable summary of general performance. Large, dependable domain differences may add clinically relevant information alongside it, but scatter does not by itself invalidate the composite. An uneven profile does not automatically establish a named disorder, nor is the highest score necessarily the person’s “true IQ.” The examiner should determine whether differences are statistically dependable, uncommon enough to merit attention and supported by independent evidence.
Performance is situated
Sleep loss, pain, acute illness, anxiety, depression, medication effects, hearing or vision barriers, motor demands, unfamiliar language, interrupted education, low engagement and misunderstanding instructions can all influence observed performance. Some of these conditions are part of the real-world functioning being investigated; others are irrelevant obstacles. Interpretation must distinguish them rather than automatically adding or subtracting points.
Retesting and practice effects
A higher retest score can reflect remembered formats, new strategies, reduced anxiety or overlapping items. These practice effects are real improvements in test performance but do not necessarily represent an equivalent increase in broad underlying ability. Size depends on interval, age, test and prior exposures; alternative forms and appropriate retest guidance help.17
Ask for useful uncertainty
A good explanation sounds like: “Performance in this domain was estimated in this range; the result is likely meaningful because these tasks agree with one another and with the person’s history. These conditions may have affected it. Here is what the finding suggests we should try next.”
What IQ predicts—and what it does not
General cognitive ability is consequential. It predicts probabilities across groups, not predetermined biographies for individuals.
Learning
Reasoning, knowledge and memory make it easier to understand instruction, connect ideas and acquire complex skills. Cognitive ability is substantially associated with school performance.2
Complex work
General mental ability predicts job performance, though revised meta-analytic estimates are more moderate than some older claims.14
Life navigation
Across populations, stronger cognitive performance is associated with educational attainment, occupational complexity and some health and longevity outcomes. These links arise through many individual and social pathways; they do not determine a person’s future.2
Correlation is not a personal prophecy
A correlation describes how variables tend to move together within a population. It does not say that a score causes an outcome by itself, or that everyone with the same score will live the same life. Education quality, health, family resources, personality, interests, discrimination, motivation, relationships, values and chance all affect what opportunities appear and how abilities are used.
Even a strong relationship leaves enormous room for individual variation. Conversely, saying that IQ is “not everything” should never be twisted into saying it is nothing. Abilities that support comprehension, reasoning and learning are worth protecting, developing and celebrating.
| Evidence can support | Evidence cannot support by itself |
|---|---|
| An estimate of selected general and domain-specific cognitive abilities. | A complete description of consciousness, personality, wisdom or creativity. |
| Probabilistic forecasts for relevant outcomes in validated populations. | Certain predictions about one person’s future. |
| Identification of some learning strengths, needs or discrepancies. | An automatic diagnosis or educational placement without other evidence. |
| Evidence of change when methods, norms, uncertainty and retest effects are addressed. | A claim that every score change is broad intellectual growth—or that growth is impossible. |
| A reason to provide challenge, opportunity, recognition or targeted support—and to protect rare ability from being wasted or needlessly harmed. | A complete ranking of human dignity or basic rights; those are not awarded by a psychometric score. |
Culture, language, disability and access
Fairness is not a decorative final check. It is part of whether the proposed score interpretation is valid for this person, population and purpose.
No assessment is culture-free. Language, schooling, familiarity with testing, visual conventions, technology access, speed expectations, motivation and relationships with authority can affect performance. That does not make every score difference “bias,” nor does a group difference prove either bias or innate ability differences. The empirical question is whether irrelevant barriers materially distort the construct the test is meant to measure.
Social context can influence performance—but effect sizes are not universal
Concern about confirming a negative group stereotype may sometimes consume attention or alter engagement during testing. Research called this stereotype threat. A meta-analysis of girls younger than 18 in mathematics, science and spatial testing found a small average effect but serious signs of publication bias; results should not be generalized into one large explanation for every group or score difference.34 The practical lesson is broader: respectful administration, psychological safety and a fair opportunity to understand the task improve the quality of evidence for everyone.
| Level | Question to ask |
|---|---|
| Access & administration | Can the person perceive, understand and respond without an irrelevant language, sensory, motor or technology barrier? |
| Item | Do equally able members of different groups have different probabilities of answering a particular item correctly? |
| Scale / construct | Does the underlying structure function comparably enough across groups for the intended comparison? |
| Norms | Is the reference sample relevant, recent, sufficiently representative and described transparently? |
| Decision / outcome | Does the score predict relevant criteria and classify people with acceptable accuracy, including near thresholds? |
Differential item functioning is a signal to investigate
Differential item functioning (DIF) occurs when people from different groups who are matched on the ability being measured nevertheless have different probabilities of responding to an item in a particular way. DIF does not automatically prove bias: reviewers examine its size, content, possible cause and effect on total scores. Automatically deleting every flagged item can also damage content coverage.
Measurement invariance asks whether comparisons mean the same thing
Invariance analyses examine whether a test’s factor structure, relationships and score levels operate comparably across groups. Different levels of invariance support different claims; evidence sufficient for comparing associations may not justify comparing group means. Invariance is valuable model-based evidence, not proof that every social or measurement problem has disappeared.21
Translation is not adaptation
A translated assessment must preserve more than words. Construct meaning, instructions, examples, graphics, response formats, scoring, norms and administration all need cultural and empirical review. Back-translation can help find discrepancies, but it cannot establish equivalence by itself. International Test Commission guidance recommends multidisciplinary expertise, pilot work and evidence at the item, method and construct levels.20
“Nonverbal” means reduced language demand—not neutral context
Tests such as matrix reasoning can reduce dependence on vocabulary and acquired verbal knowledge. Their instructions, abstract visual conventions, speed, motor response and puzzle familiarity still reflect experience. They require the same population-specific reliability, validity, norming and fairness evidence as verbal measures.
Accommodation
A change intended to remove a barrier irrelevant to the construct while preserving the score’s meaning. It should be individualized, documented and supported by evidence.
Modification
A change that alters the target construct or standardized comparison. It may be humane and useful, but the resulting score may not be comparable with original norms.
Extra time may preserve comparability when speed is irrelevant; when processing speed is the intended construct, it may change what the score means or make standard timed norms inapplicable. Similar care applies to interpreters, readers, scribes, alternative interfaces and simplified instructions. Identical administration is not always fair; altered administration is not automatically valid.22
Fairness needs both community knowledge and psychometric evidence
Consultation with teachers, families and community members may help identify inaccessible wording, missing local knowledge or harmful assumptions that distant developers overlook. Such input should generate hypotheses for documented accessibility, standardization, reliability, validity and outcome studies; it does not establish fairness by itself.120
The higher the stakes, the stronger the safeguards
The same score may be adequate for research on group patterns yet far too imprecise to decide one person’s education, diagnosis, employment or access to services.
Psychological testing is one component of assessment. A consequential conclusion should integrate the referral question, personal and developmental history, observed behavior, relevant records and other measures rather than allowing one number to silently become the decision.13
Use multiple sources
Combine cognitive results with relevant achievement, history, observations, interviews, work samples or adaptive-function evidence.
Respect uncertainty
Examine confidence intervals, reliability of differences and classification error near any threshold.
Check fit
Confirm language, norms, accessibility, examiner competence and validity evidence for this population and use.
Create a path forward
A result should lead to suitable challenge, support, instruction or further inquiry—not a dead-end label.
Intellectual disability is not diagnosed from IQ alone
Contemporary definitions require significant limitations in both intellectual functioning and adaptive behavior, with onset during the developmental period. Adaptive behavior covers conceptual, social and practical skills used in everyday life. Culture, language, community expectations, strengths and support needs also matter.23 A score near 70 may contribute important evidence; it is not a self-sufficient diagnosis.
Giftedness and advanced opportunity also need converging evidence
High cognitive scores can identify exceptional reasoning and learning capacity that deserves challenge. But rigid use of one cutoff may miss learners whose language, disability, educational opportunity or uneven profile suppresses a composite. Multiple measures can improve both discovery and planning: domain scores, achievement, portfolios, teacher evidence, rate of learning and demonstrated creative work may all be relevant to a program’s stated purpose.
Selection and employment
When cognitive tests are used in employment, the measured abilities should be demonstrably related to the work, procedures should follow applicable law, and adverse outcomes should be examined. A predictive correlation at the group level does not absolve an organization from accessibility, fairness, privacy, transparency or human review.
Never let a cutoff pretend to be nature
Administrative thresholds are human decisions applied to uncertain measurements. Near a cutoff, small shifts in health, norms, test form or measurement error can change a category. The closer the score is to a consequential boundary, the more important corroborating evidence and review become.
Complementary assessments answer different questions
The answer to an incomplete measure is not an even vaguer “whole-person score.” It is a purposeful set of measures, each used for the construct it can validly illuminate.
Keep the measures distinct
Keep distinct evidence distinct. Cognitive ability, achievement, executive function, adaptive behavior, creativity, emotional skills and demonstrated work are not interchangeable ingredients to average into one grand number. A dashboard preserves what each result means—and exposes the value judgments behind any decision.
| Method | Best question | Important limitation |
|---|---|---|
| Achievement assessment | What knowledge or skill has this person learned in reading, mathematics, science or another domain? | Reflects curriculum, instruction, language and opportunity to learn—not pure ability. |
| Executive-function tasks & ratings | How does the person inhibit, shift, maintain goals and manage working information in tasks or everyday life? | Brief laboratory tasks and real-life ratings often capture different aspects and should not be treated as equivalent.25 |
| Adaptive-behavior scales | How independently and effectively does the person use conceptual, social and practical skills in daily life? | Ratings depend on opportunity, context, expectations and informants; multiple sources are valuable. |
| Dynamic assessment | How does performance change with prompts, teaching or scaffolding? | Administration is longer and less uniform; improvement may be task-specific and examiner-dependent.26 |
| Creativity assessment | Can the person generate, evaluate and improve ideas that are original and effective in a defined domain? | Divergent-thinking tasks sample creative potential, not a complete creative life or future achievement. |
| Emotional-skill measures | How well does the person perceive, understand or manage emotional information—or how do they describe typical emotional tendencies? | Ability tests, trait self-reports and mixed “EQ” models measure different constructs. |
| Portfolios & work samples | What has the person actually written, designed, built, performed or solved over time? | Selection, coaching, resources and rater judgment can affect results; common rubrics and multiple samples help. |
Executive function: task performance and daily regulation
Executive functions include separable but related aspects of goal-directed control, often described through inhibition, shifting and working-memory updating. Performance tasks can reveal processing under controlled conditions; self-, parent- or teacher ratings can show success in everyday goal pursuit. Their modest convergence means disagreement is informative, not necessarily an error: the methods may be sampling different levels of behavior.2425
Dynamic assessment: observe learning, not only arrival
Test–teach–retest procedures and graduated prompts examine how a learner responds to guidance. This can reveal strategies, kinds of support and responsiveness that a static score cannot. It is especially useful when prior opportunity, language or instruction complicates interpretation. Evidence varies by method and population, so dynamic assessment supplements standardized evidence rather than eliminating bias.
Creativity: originality joined with effectiveness
Creativity is not merely producing many unusual answers. In serious assessment, ideas must also be useful, effective or fitting within a domain. Open-ended tasks, actual products, portfolios, expert ratings and a record of creative activity can be combined. Cognitive ability and creative achievement are positively related but far from identical, which is exactly why both can add information.2728
Emotional intelligence: specify which model
Ability EI
Performance tasks intended to assess perceiving, using, understanding or managing emotional information.
Trait EI
Self-reported typical tendencies and self-perceptions related to emotion.
Mixed models
Broader combinations of skills, personality, motivation and social behavior, often marketed under one EQ label.
These families should not be compared as though they were one scale. Emotional skills can contribute to relationships, learning and performance, while cognitive ability measures something different. Evidence suggests a modest association between EI measures and academic performance, with limited added prediction after cognitive ability and personality depending on the instrument.2930
Multiple intelligences: a humane prompt, not a validated diagnostic map
Howard Gardner’s framework encouraged educators to notice musical, bodily, interpersonal, spatial and other strengths too often ignored by narrow classroom routines. That educational message can be valuable. Evidence does not, however, establish the proposed intelligences as independent psychometric systems, and popular online “MI tests” should not fix children into learning-style labels. Use the framework to broaden opportunity, not to replace validated cognitive or skill assessment.31
IQ estimates a broad, useful pattern of cognitive performance. Achievement shows learned mastery; executive measures sample control; adaptive scales show daily functioning; dynamic assessment shows response to teaching; creativity and emotional measures sample other capacities; portfolios show what a person has actually made or done.
Can intelligence—and IQ—grow?
Yes—but these are different claims. Cognitive abilities can develop, and age-normed IQ can change. Showing durable broad growth requires evidence beyond maturation, measurement error, changed norms and familiarity with one test.
Growth deserves celebration
When people expand knowledge, learn to reason across unfamiliar problems, improve memory strategies or gain the cognitive tools to understand the world faster and more accurately, the benefits can reach education, work, decisions, relationships and independence.
Measurement should help us recognize that achievement—not deny it because intelligence is partly stable, and not exaggerate it because one practiced score rose.
Education has the strongest broad evidence
A large meta-analysis emphasizing quasi-experimental designs found that an additional year of education produced an estimated average benefit of roughly 1–5 IQ points, depending on design and assumptions, across more than 600,000 participants. Effects appeared across broad cognitive categories and across the lifespan.15
This is a population estimate, not a yearly promise to every individual and not an unlimited linear law. It is strong evidence against treating measured intelligence as frozen.
Stability and change are compatible
Cognitive abilities show increasing rank-order stability from childhood into adulthood: people often retain a broadly similar position relative to peers. That does not mean absolute performance never changes, environments are irrelevant or the group average cannot move. The most comprehensive recent meta-analysis found high—but not perfect—stability from late adolescence through much of adulthood.19
Population scores have changed across generations
Across much of the twentieth century, performance rose on many intelligence tests—the Flynn effect. A meta-analysis spanning 31 countries found gains that differed by cognitive domain, place and period, with weaker trends in more recent decades.16 This is why tests require periodic renorming. It does not mean every newer generation is wiser in every sense, and there is no universal “three points per decade” correction for every score.
| Kind of gain | What changed | Best evidence |
|---|---|---|
| Practice gain | Performance on repeated or highly similar items. | Useful for mastering that task; do not infer broad transfer without more. |
| Skill gain | A trained domain such as arithmetic, vocabulary, spatial strategy or attention routine. | Unpracticed measures of the same skill and sustained real-world performance. |
| Broad cognitive gain | Abilities generalize across substantially different tasks or domains. | Alternate measures, active controls where possible, and follow-up over time. |
| Normative IQ change | Standing changes relative to same-age peers. | Comparable, current norms; confidence intervals; reliable-change and retest data. |
Brain training: value the real gain, name it honestly
Commercial cognitive training often improves the practiced task and sometimes closely related tasks. Far transfer to general intelligence is usually weak or inconsistent in well-controlled reviews.18 A domain-specific gain may still be useful. Precision protects progress: calling every gain “higher IQ” makes genuine improvement harder to recognize.
Conditions that support learning and cognitive functioning
Learn deeply
Build knowledge, retrieve it repeatedly, explain it, connect it and use it in new problems.
Protect the brain
Prioritize sleep, physical health, movement, nutrition, safety and treatment of barriers to attention or learning.
Seek challenge
Work slightly beyond current mastery with feedback, scaffolding and enough time to consolidate.
Track transfer
Look for durable improvement on new tasks, real work and everyday understanding—not only a familiar exercise.
Genes influence cognitive differences, but heritability is a population statistic within a particular range of environments; it is not the percentage of one person’s intelligence “caused by genes,” nor evidence that education cannot work.2 Stability describes a pattern. It does not mean growth is impossible.
Solitude can form an idea; community can help it live
Original thought and social connection are not enemies. They often belong to different phases of the same creative cycle.
Protected wandering
Time alone can free attention from immediate instructions, group expectations and the pressure to respond. A person can follow a strange question, tolerate an unfinished thought and connect ideas before anyone else defines the task.
In groups, simultaneous idea generation can be limited by production blocking—waiting to speak while a thought fades—and by evaluation apprehension or reduced individual contribution.32 Early examples can also produce collaborative fixation, narrowing the range explored before a distinctive line of thought has had time to form.35
Chosen return
Once an idea has enough shape to share, other people often become valuable—and may become indispensable for testing, development and diffusion. They can question assumptions, contribute knowledge, test the work, supply resources, refine communication, open doors and protect the creator from blind spots. Across three controlled brainwriting experiments, alternating individual and group sessions produced more ideas than staying in either mode—evidence for a flexible cycle in laboratory ideation, not one universally best setting for all creative work.33
Returning by choice can be a celebration: the private spark meets recognition, support and shared effort. A mind need not remain alone to remain original.
Give originality enough solitude to form—and enough community to grow, travel, be supported and be celebrated.
Neither phase should become a command. Solitude can become isolation; collaboration can become conformity. Some people think best through dialogue, some in quiet, and many alternate between both. The healthy pattern is flexible: withdraw without shame when undirected thought needs protection, then return without shame when relationship, critique and shared creation are desired.
When minds are protected from humiliation, coercion and constant interruption, people can support one another’s intellectual growth without erasing difference. Intelligence becomes more than private advantage: it becomes a capacity shared through teaching, invention, care, honest disagreement and collective problem-solving.
Choosing—and preparing for—an assessment
Begin with the decision that must be improved. The “best test” in the abstract may be the wrong assessment for the person, question or jurisdiction.
Define the purpose
Clinical diagnosis, educational planning, gifted identification, disability support, research and personal curiosity require different evidence and levels of precision.
Match the person
Confirm age, language, cultural context, sensory and motor access, communication needs, digital familiarity and relevant norm representation.
Match the decision
Ask whether the test manual and independent standards support this exact interpretation and whether other evidence is required.
Before the session
- Use a qualified professional for clinical, educational or other consequential interpretation.
- Share relevant language history, education, diagnoses, medication, accommodations and previous testing.
- Ask what constructs, scores and comparison groups the selected battery uses.
- Protect ordinary sleep, food, prescribed treatment and sensory aids; do not pursue last-minute “IQ preparation.”
- Explain unusual illness, pain, distress or interruption so postponement can be considered when appropriate.
During the session
Try sincerely, ask when instructions are unclear and report access barriers. Do not seek leaked items or rehearse secure test content: that compromises the meaning of the result. Ordinary familiarity with puzzles is part of life; coaching on protected items changes the inference the score can support.
After the session: questions worth asking
- What was the composite estimate and its confidence interval?
- Which profile differences are reliable, uncommon and practically meaningful?
- What testing conditions may have influenced performance?
- How well did results agree with history and other evidence?
- What does the test not measure?
- Which decisions can these scores legitimately inform?
- What challenge, support or further assessment should follow?
- When would retesting be useful, and how will practice effects be handled?
Personal curiosity deserves context too
A reputable assessment may offer valuable self-knowledge. Treat the number as one coordinate, then ask what it helps you do: choose appropriate challenge, build weaker skills, use strengths responsibly and continue learning. Chasing tiny score differences can replace development with measurement anxiety.
Faster testing should not mean thinner evidence
Adaptive platforms, process data and artificial intelligence may improve measurement—but efficiency is not validity, and prediction is not understanding.
Computerized adaptive testing
An algorithm selects the next item from a calibrated bank based on previous responses. This can reach similar precision with fewer items, but it is a delivery method—not a new form of intelligence.
Process data
Timing, revisions, strategy changes and response sequences may reveal how a solution emerged. These signals need independent validation and can introduce technology or motor demands.
AI-assisted interpretation
Systems may detect patterns, draft reports or suggest hypotheses. They must not hide uncertainty, invent diagnoses, reproduce training-data bias or replace accountable professional judgment.
What responsible innovation must make visible
- Construct: what is inferred from which observable behavior.
- Training and norm data: who is represented, who is missing and how recently data were collected.
- Uncertainty: precision by score range, subgroup and decision threshold.
- Fairness: item functioning, accessibility, differential prediction and downstream errors.
- Privacy: what sensitive cognitive and behavioral data are stored, linked, retained or shared.
- Accountability: who can explain, challenge and correct a consequential result.
More behavioral data can produce a more detailed model and a larger privacy risk at the same time. Voice, keystroke, gaze, latency and interaction patterns may reveal disability, health or identity information beyond the stated purpose. Data minimization, informed consent, security and meaningful appeal therefore belong inside test design—not in an unread final paragraph.
The standard for the future
A better assessment is not the one that produces the most numbers. It is the one that answers an important question with sufficient precision, creates less irrelevant burden, reveals its limits and leads to better opportunities or support.
Common myths—corrected
Most arguments about IQ become clearer when an absolute claim is replaced with a bounded one.
Conclusion: measure what matters, then help it grow
IQ testing is neither a verdict nor an illusion. A well-designed, appropriately normed assessment can provide a useful estimate of general and specific cognitive abilities. Those abilities matter: they support faster learning, deeper understanding, adaptive reasoning and the power to solve increasingly difficult problems.
The estimate remains conditional. It has uncertainty, reflects a particular sample of behavior and cannot contain a whole person. Fair use requires suitable norms, accessible administration, evidence for the intended interpretation and restraint when consequences are high. Complementary assessments should broaden the picture without pretending that all valued capacities are one hidden score.
Most importantly, measurement should serve development. Education can strengthen cognitive performance; health and safety protect the brain; challenge, feedback and practice build knowledge and skill; solitude can shelter an original thought; community can help it mature and reach the world. Celebrate durable cognitive growth—and verified IQ gains when transfer, persistence, practice effects and measurement error have been addressed—then use greater ability to deepen understanding, expand opportunity and improve lives.
Measure intelligence carefully, help it grow and protect the people in whom it lives. Celebrate exceptional ability, lifelong learning, original thought and accumulated wisdom. Give developed minds the health, freedom, solitude, community, challenge and respect they need—because when such a mind is neglected or lost, humanity may lose guidance, understanding and possibilities that no possession or database can simply replace.
Sources and further reading
Primary standards, official guidance, historical records and peer-reviewed research used for this guide.
- AERA, APA & NCME. Standards for Educational and Psychological Testing (2014).
- Deary, Cox & Hill. Genetic variation, brain, and intelligence differences. Molecular Psychiatry (2022).
- Brysbaert & Nicolas. Two persistent myths about Binet and the beginnings of intelligence tests in psychology textbooks. Collabra: Psychology (2024).
- Binet & Simon. The development of intelligence in children (1908, French original).
- Kovacs & Pléh. William Stern: the relevance of his program of “differential psychology” for contemporary intelligence measurement and research. Journal of Intelligence (2023).
- Terman. The Measurement of Intelligence (1916).
- U.S. Department of Defense. History of military testing.
- Warne et al. Stephen Jay Gould’s analysis of the Army Beta test in The Mismeasure of Man: distortions and misconceptions regarding a pioneering mental test. Journal of Intelligence (2019); Brigham. Intelligence tests of immigrant groups. Psychological Review (1930).
- National Human Genome Research Institute. Eugenics and scientific racism; Snyderman & Herrnstein. Intelligence tests and the Immigration Act of 1924 (1983).
- Boake. From the Binet–Simon to the Wechsler–Bellevue. Journal of Clinical and Experimental Neuropsychology (2002).
- McGrew. Carroll’s three-stratum (3S) cognitive ability theory at 30 years: impact, 3S-CHC theory clarification, structural replication, and cognitive–achievement psychometric network analysis extension. Journal of Intelligence (2023).
- Pearson Assessments. Wechsler Adult Intelligence Scale, Fifth Edition.
- National Academies. Psychological testing in the service of disability determination (2015).
- Sackett et al. Revisiting meta-analytic estimates of validity in personnel selection: addressing systematic overcorrection for restriction of range. Journal of Applied Psychology (2022).
- Ritchie & Tucker-Drob. How much does education improve intelligence? A meta-analysis. Psychological Science (2018).
- Pietschnig & Voracek. One century of global IQ gains: a formal meta-analysis of the Flynn effect. Perspectives on Psychological Science (2015).
- Scharfen, Peters & Holling. Retest effects in cognitive ability tests: a meta-analysis. Intelligence (2018).
- Simons et al. Do “brain-training” programs work? Psychological Science in the Public Interest (2016).
- Breit et al. The stability of cognitive abilities: a meta-analytic review of longitudinal studies. Psychological Bulletin (2024).
- International Test Commission. ITC guidelines for translating and adapting tests (second edition) (2017).
- Putnick & Bornstein. Measurement invariance conventions and reporting. Developmental Review (2016).
- European Federation of Psychologists’ Associations. EFPA Test Review Model, Version 2025.
- American Association on Intellectual and Developmental Disabilities. Definition of intellectual disability.
- Miyake et al. The unity and diversity of executive functions. Cognitive Psychology (2000).
- Toplak, West & Stanovich. Performance-based and rating measures of executive function. Journal of Child Psychology and Psychiatry (2013).
- Swanson & Lussier. A selective synthesis of the experimental literature on dynamic assessment. Review of Educational Research (2001).
- OECD. PISA 2022 creative-thinking framework.
- Karwowski et al. How is intelligence test performance associated with creative achievement? A meta-analysis. Journal of Intelligence (2021).
- O’Connor et al. The measurement of emotional intelligence: a critical review of the literature and recommendations for researchers and practitioners. Frontiers in Psychology (2019).
- MacCann et al. Emotional intelligence predicts academic performance: a meta-analysis. Psychological Bulletin (2020).
- Waterhouse. Multiple intelligences, the Mozart effect, and emotional intelligence: a critical review. Educational Psychologist (2006).
- Mullen, Johnson & Salas. Productivity loss in brainstorming groups: a meta-analytic integration. Basic and Applied Social Psychology (1991).
- Korde & Paulus. Alternating individual and group idea generation: finding the elusive synergy. Journal of Experimental Social Psychology (2017).
- Flore & Wicherts. Does stereotype threat influence performance of girls in stereotyped domains? A meta-analysis. Journal of School Psychology (2015).
- Kohn & Smith. Collaborative fixation: effects of others’ ideas on brainstorming. Applied Cognitive Psychology (2011).
Educational and assessment note: This article provides general information, not an individual diagnosis, score interpretation, legal opinion or substitute for evaluation by a qualified psychologist or other appropriately credentialed professional. Test selection, accommodations, eligibility rules and legal requirements vary by location and purpose.
Intelligence Unleashed series
- Definitions and Perspectives on Intelligence
- Brain Anatomy and Function
- Types of Intelligence
- Theories of Intelligence
- Neuroplasticity and Lifelong Learning
- Cognitive Development Across the Lifespan
- Genetics and Environment in Intelligence
- Measuring Intelligence
- Brain Waves and States of Consciousness
- Cognitive Functions