Critical Thinking and Problem-Solving
Linas JuozenasShare
Reason clearly. Imagine widely. Decide responsibly.
Powerful problem-solving is not one flash of cleverness. It is a disciplined cycle: understand the situation, examine the evidence, generate alternatives, test what matters, choose under uncertainty and learn from the result.
- FrameDefine the real problem and its constraints
- InvestigateMap claims, evidence and uncertainty
- GenerateCreate genuinely different possibilities
- Test and decideLearn cheaply, choose and update
Critical thinking is not automatic disbelief
Critical thinking judges a claim in proportion to its support. Sometimes that warrants scepticism; sometimes strong evidence warrants confidence. Reflexively rejecting experts or unfamiliar ideas is no more rigorous than accepting them without examination.
A widely used expert-consensus account emphasizes interpretation, analysis, evaluation, inference, explanation and self-regulation—the willingness to inspect one’s own assumptions as well as another person’s.1
Creativity is more than producing something unusual
An idea can be novel yet useless, harmful or impossible. Creative problem-solving joins possibility generation with knowledge, judgement and testing. Divergent-thinking tasks measure part of creative potential, not the completed achievement.15
The goal is a productive exchange between expansion and selection: enlarge the search, then improve the standards by which options survive.
A complete thinker changes the question at the right time
The mind must sometimes narrow, sometimes widen, sometimes calculate and sometimes observe. Failure often comes from using a valuable mode at the wrong stage.
- ObserveWhat is happening?
- FrameWhat problem are we solving?
- ExplainWhat might cause it?
- GenerateWhat else could work?
- EvaluateWhich option survives?
- LearnWhat changed our model?
| Question | Mode | Typical failure | Evidence of success |
|---|---|---|---|
| What is true? | Analysis | Familiarity or confidence is treated as proof | Claims are traceable and rivals were considered |
| What could be? | Generation | The first workable option ends the search | Options differ in mechanism and perspective |
| What should we do? | Judgement | One metric replaces the whole objective | Criteria, trade-offs and affected people are explicit |
| What did we learn? | Updating | The outcome is rewritten as obvious | Predictions, surprises and model changes are recorded |
Meta-analyses indicate that critical-thinking instruction can improve measured skills, particularly when reasoning is taught explicitly, discussed substantively and practised within real subject matter.2 A newer synthesis also reports positive effects, while high heterogeneity shows that programmes and outcomes differ substantially.3
Knowledge and general reasoning are partners
A person cannot evaluate evidence they do not understand. Domain knowledge determines which variables, mechanisms and exceptions are visible; general reasoning helps compare them, detect inconsistency and transfer lessons. The practical goal is not “skills instead of facts” or “facts instead of skills,” but richly organized knowledge used through deliberate inquiry.
An argument is a support structure, not a quarrel
Arguments connect reasons to conclusions. Mapping that connection exposes where disagreement really lives: the evidence, the bridge from evidence to claim, the confidence level or an unaddressed exception.
| Part | Question | Checkout example | Weakness |
|---|---|---|---|
| Claim | What is asserted? | “This redesign will reduce abandonment.” | Vague or broader than the evidence |
| Grounds | What supports it? | Recordings show failure at one step. | Poor measurement or comparison |
| Warrant | Why does support connect? | Removing the obstacle should let willing buyers continue. | The causal bridge is assumed |
| Backing | What supports the warrant? | Prior tests and relevant usability research. | Authority replaces relevance |
| Qualifier | How confident and where? | “Probably, for mobile users reaching this step.” | Certainty hides limited scope |
| Rebuttal | What could reverse it? | The change may reduce trust or shift abandonment. | Counterevidence is ignored |
Toulmin's framework was designed for practical arguments whose force depends on field, context and qualification, not only on formal validity.28 It is especially useful when someone presents evidence but leaves the warrant invisible.
Deduction
If true premises and a valid form guarantee the conclusion, ask both whether the logic is valid and whether the premises are sound.
Induction
Observations support a conclusion with a degree of strength. Sampling, selection and measurement determine how far it travels.
Abduction
Choose the best current explanation by fit, mechanism and predictions that rivals do not explain as well.
- Claim and premises: write the conclusion precisely; separate observations, assumptions and interpretations.
- Warrant and scope: explain why the premises support the claim and under which population, time and conditions.
- Strongest rival: construct the best competing explanation rather than the easiest version to defeat.
- Discriminating evidence: name what would favour either account, what would change your mind and what remains unknown.
Do not use argument analysis as covert control
The purpose is transparent understanding, not extracting disclosure, manufacturing dependence or wearing someone down until they comply. Ethical reasoning preserves the other person's right to question the frame, withhold private information, disagree, pause or leave.
Ask what the evidence can establish—and what it cannot
A graph may describe a sample, an experiment may estimate a causal effect under particular conditions, and a review may summarize a literature. None automatically proves every claim built around it.
- Measurement: does the measure capture the claimed construct or only a convenient proxy?
- Selection: who entered, left or never appeared in the sample?
- Comparison: compared with nothing, usual practice, placebo or a credible active alternative?
- Magnitude and uncertainty: is the effect useful, and which values remain compatible with the data?
- Multiplicity and transparency: how many outcomes or analyses were tried, and were methods specified in advance?
- Transport: why should this result apply to another population, setting or decision?
Correlation is a clue, not a causal verdict
If people who use a tool perform better, the tool may help—but capable people may also choose it, training may accompany it, or both may be influenced by a third factor. Reverse causation, confounding, selection and measurement can all produce or distort an association. Causal diagrams force these assumptions into view and warn that controlling for every available variable can itself introduce bias when the variable is a mediator or collider.10
- Cause?A may change B
- Reverse?B may change A
- Common cause?C may change both
- Selection or measurement?The observed sample or proxy may distort the link
A meta-analysis can increase precision and reveal variation, but it cannot rescue biased underlying studies by arithmetic alone. Check inclusion rules, study quality, heterogeneity, outcome definitions and publication bias. Large literatures can still be fragile: the Open Science Collaboration's coordinated replication project helped make reproducibility a central part of evidence evaluation, while also showing that replication outcomes require careful interpretation rather than a pass/fail mythology.11
Calibrate language to evidence
Observed describes what was measured. Associated describes co-variation. Consistent with means a result fits an explanation but may fit rivals. Caused requires a defensible causal design and assumptions. Established should be reserved for conclusions supported across strong, converging evidence—not one dramatic study.
Cognitive bias is not a collection of insults
Heuristics make judgement possible under limited time and information. Bias research identifies conditions in which shortcuts predictably mislead; it does not justify diagnosing every disagreeing person from a distance.9
Ability matters—but does not decide every bias task
General intelligence supports complex reasoning and valid cognitive tests provide meaningful evidence about what they measure. Yet some myside and framing tendencies relate only weakly to cognitive ability, while other rational-thinking tasks relate more strongly.45 Knowledge, reflection, motivation and thinking dispositions add further information.
Reflection is used selectively
Fast answers can be efficient or mistaken. Cognitive-reflection tasks test whether a person interrupts an attractive first response when the structure signals a trap.6 Real skill is not universal slowness, but recognizing when novelty, stakes, conflict or uncertainty deserve another pass.
| Risk | When to suspect it | Countermeasure | Caution |
|---|---|---|---|
| Myside processing | Identity or prior commitment is at stake | Write the strongest opposing case and evidence that would change your mind | Do not manufacture balance when evidence is genuinely unequal |
| Availability | A vivid event replaces frequency information | Retrieve a relevant base rate and comparison class | The base rate may be outdated or mismatched |
| Anchoring | The first number dominates later estimates | Estimate independently before seeing others; compare outside ranges | A relevant anchor may contain information |
| Sunk cost | Past investment is the reason to continue | Ask what you would choose if inheriting the project today | Switching costs can be genuine future costs |
| Outcome bias | Luck is confused with decision quality | Score the process using information available at the time | Repeated outcomes still update process quality |
“Consider the opposite” can recover neglected evidence.7 Natural frequencies—such as “8 out of 100”—can clarify Bayesian relationships.8 Forecasting practice shows the value of decomposition, probabilistic estimates, updating and feedback.12
Debiasing is targeted practice, not permanent immunity
A 2025 review of 54 randomized trials and 10,941 participants found a small average reduction in targeted biases, while transfer to real decisions remained uncertain and all included studies had unclear or high risk of bias.29 A field study offers encouraging evidence that carefully designed training can transfer beyond its original task, but one intervention cannot certify a person as unbiased.13
Fallacy names should begin analysis, not end conversation
A label is useful only when it identifies what went wrong and how the conclusion should be reconsidered. Context can also make information relevant that would be irrelevant elsewhere.
| Pattern | What goes wrong | Diagnostic question |
|---|---|---|
| Straw man | A weaker substitute is attacked | Would the other person accept this summary? |
| Ad hominem | A personal attack replaces evaluation of the claim | Would the evidence change if another person presented it? |
| False dilemma | More than two options are compressed into two | What mixed, staged or third alternatives exist? |
| Post hoc | Sequence is treated as sufficient proof of cause | What comparison, mechanism and rival causes were tested? |
| Cherry-picking | Supporting cases are selected while relevant failures disappear | What rule determined which evidence entered? |
| Moving the goalposts | The standard changes after evidence arrives | Was the success criterion stated beforehand? |
| Appeal to authority | Status substitutes for relevant evidence | Is the expertise relevant, accurately represented and supported? |
Relevance depends on the claim
Source incentives, conflicts and reliability matter when they change the probability that evidence is accurate. They do not logically refute a claim by themselves. Examine the evidence and the source conditions.
Principles need boundary conditions
Slopes, slippery or otherwise, require mechanisms and probabilities. Exceptions do not erase a general rule unless the rule claimed universality; anecdotes can reveal possibilities without estimating frequency.
A brilliant answer to a mistaken frame is still a failure
Problem statements quietly choose the goal, unit of analysis, time horizon and people whose interests count. Before generating solutions, expose those choices.
- Observed state: what happens, where, when, how often and for whom—without assuming a cause?
- Desired state: what outcome would be better, and how would it be recognized?
- Mechanisms: which competing causal pathways might produce the pattern?
- Constraints: which limits are physical, legal, ethical or resource-based, and which are habits?
- Stakeholders: who benefits, pays, carries risk, supplies knowledge or has a right to consent?
- Unknowns and reversibility: what information would change the choice, and can it be piloted or undone?
Experts see deeper structure
Classic comparisons found that physics experts grouped problems by underlying principles, while novices relied more on surface features.30 Analogies transfer when their relational structure is recognized,20 yet expertise can also create fixation when a familiar good solution blocks a better one.21 The remedy is not less expertise, but explicit searches for alternative representations.
Change the level
Is this an individual error, interface problem, incentive problem, supply problem or system interaction? Each level implies different evidence.
Change the verb
Replace “stop” or “fix” with “enable,” “detect,” “contain,” “learn,” “adapt” or “make safer.” New verbs expose new mechanisms.
Change the horizon
Examine immediate response, adaptation, maintenance, second-order effects and exit conditions—not only this week’s metric.
Decomposition is useful only if you preserve the interfaces
Breaking a complex problem into parts reduces load, but the parts may interact. Optimize each department separately and the customer journey can worsen; optimize speed alone and quality can collapse. Record dependencies, feedback loops and shared constraints before treating subproblems as independent.
Generate alternatives that differ in mechanism, not merely wording
Divergence is a search strategy. It aims to escape the first available cluster of ideas and explore other assumptions, scales, users, technologies and causal routes.
Creativity training can improve creative performance, especially when it teaches explicit cognitive strategies, uses realistic exercises and provides practice in applying them.14 Yet fluency on an alternate-uses test is not identical to producing an original scientific theory, business model or work of art. Domain knowledge, taste, persistence, resources and implementation determine what becomes real.
SCAMPER
Substitute, combine, adapt, modify, repurpose, eliminate or rearrange. Use the letters as prompts, not proof of originality.
Mechanism quota
Require more options than feel necessary and label each by how it works so synonyms do not inflate the count.
Invert a constraint
Remove, reverse, tighten or transfer a presumed fixed boundary, then ask which consequences are useful.
Build a random bridge
Import a process from another field and search for a genuine structural analogy rather than decorative similarity.
Change the search, not only the wording
Invert a deliberately bad solution to expose hidden assumptions, or search biology, logistics, accessibility, games, history and adjacent industries for mechanisms worth adapting. Keep real harms outside the joke and affected people inside the evaluation.
Movement and incubation: useful, not mystical
Walking improved divergent ideation in four experiments, although not every task benefited.16 A meta-analysis found incubation benefits under some conditions, especially after an initial attempt and with a modest intervening task.17 Neither replaces knowledge or evaluation. Neuroscience likewise rejects a “creative right brain”: creative cognition reflects changing cooperation among spontaneous-association, salience and control systems.1819
Protect the unfinished thought from constant direction. Give it enough knowledge to grow, enough solitude to become distinct and enough criticism to become real.
A decision is a commitment made with incomplete information
The best available choice can still produce a bad outcome, and a reckless choice can occasionally succeed. Judge process using what was knowable at the time, then use outcomes to update the process.
- State the decision and deadlineDefine what must be chosen, by whom and when.
- Set thresholds before scoresRemove options that violate safety, legality, consent or indispensable requirements.
- Compare value, evidence and riskUse explicit criteria, confidence ranges and affected stakeholders.
- Stress-test the leaderRun sensitivity checks, a premortem and a search for disconfirming evidence.
- Choose, record and revisitDocument assumptions, ownership, stop rules and the evidence that will trigger revision.
| Criterion | Weight | Small pilot | Full launch | Do nothing |
|---|---|---|---|---|
| Learning value | High | High | Medium | Low |
| Reversibility | High | High | Low | High |
| Potential benefit | High | Medium | High | Low |
| Cost and disruption | Medium | Low | High | Low now |
| Risk to affected people | Threshold | Must pass | Must pass | May preserve existing harm |
Weighted matrices are valuable for exposing criteria and disagreement, but weights remain judgements and correlated criteria can double-count the same value. Expected-value calculations can discipline choices when probabilities and consequences are meaningful, yet precision should not exceed the evidence.
Premortem
Imagine the plan has failed and write the plausible history. This makes neglected threats easier to name before commitment.
Reversibility
Stage uncertain choices, preserve options and spend irreversible resources only when evidence justifies the commitment.
Ethics belongs inside the decision model
Do not treat consent, privacy, coercion, environmental cost or unequal exposure to risk as decorative concerns added after optimization. Ask who receives the benefit, who absorbs error, whose data are used, who can refuse and whether the decision is reversible for the people with the least power.
A prototype is a question made inspectable
Testing is not a final ceremony after the “real” thinking. It is part of thinking: a way to make assumptions collide with the world before the cost of correction becomes enormous.
Test the riskiest assumption
Identify the belief whose failure would collapse the plan. Do not spend weeks polishing features that depend on an untested demand, mechanism or permission.
Use the smallest informative test
A mock-up, sample, simulation, interview protocol or limited pilot should discriminate between plausible stories—not merely create activity.
Predefine interpretation
Record the expected signal, comparison, minimum meaningful difference and decision rule before seeing results. Otherwise every outcome can be narrated as success.
- Which precise belief should the test update?
- What result is expected if the preferred explanation is correct—and if its strongest rival is?
- Which measure, comparison and sample can distinguish them?
- What will count as success, stopping, revision or abandonment?
- What harms, privacy costs or consent requirements must be controlled?
- Will the result transfer to the real setting, or only to the prototype?
Failure can be productive only under designed conditions
Research on productive failure shows that attempting complex problems before direct instruction can prepare learners to notice important features and learn from subsequent explanation.22 This does not mean that suffering, random failure or withholding necessary support is educational. Productive failure protects learners from catastrophic stakes, supplies feedback and connects the attempt to later instruction.
Improve the model, not only the product
After a test, separate four outcomes: what happened, what the original model predicted, which assumption was wrong and what new prediction follows. A project that fails cheaply while correcting a major misconception may create more value than a lucky pilot interpreted carelessly.
Original thought needs solitude—and a community worthy of its return
Other minds can widen evidence, challenge assumptions and develop an insight. They can also interrupt, conform, dominate or erase it. Good collaboration is designed, not presumed.
Protect independent generation
Chosen solitude can preserve attention and allow distant associations to coexist before social expectations narrow them. Quietness, introversion, autism, unconventional communication or limited social polish do not invalidate intelligence or contribution. A person may need uninterrupted time precisely because the idea is not yet easy to explain.
Return for testing and construction
Trustworthy collaborators contribute missing knowledge, detect errors, supply tools, protect creators, implement discoveries and give fair credit. Community should strengthen an original mind without demanding the surrender of individuality.
Traditional face-to-face brainstorming often underperforms the combined output of people generating separately, partly because only one person can speak at a time, ideas are forgotten while waiting, evaluation is feared and responsibility diffuses.23 Yet groups can gain from idea exchange when the process reduces blocking and lets members build on one another's knowledge.
- Generate independently: record ideas and estimates before status or charisma anchors the room.
- Pool visibly: use a shared document or anonymous board so ideas are not lost while someone else speaks.
- Clarify, then challenge: understand mechanisms before ranking and invite genuine dissent rather than theatrical devil’s advocacy.24
- Combine and decide: join compatible fragments, move disagreement toward evidence and name decision rights, objections and test ownership.
Do not confuse harmony with intellectual health
Psychological safety is not freedom from criticism. It is freedom to report uncertainty, error and dissent without humiliation or retaliation. Respectful challenge and humane treatment are compatible; coercive consensus and personal attack are not.
Use AI to enlarge inquiry—not to outsource ownership of belief
Generative systems can produce explanations, alternatives, code, summaries and counterarguments rapidly. Fluency is useful, but fluent output is not verified evidence, original understanding or accountable judgement.
Good uses
Ask for competing frames, missing assumptions, test cases, search terms, analogies, critiques or a clearer structure. Use the model to multiply questions and expose what requires verification.
Characteristic risks
Outputs may invent citations, reproduce bias, hide uncertainty, leak sensitive information, homogenize ideas or encourage acceptance because the answer is polished. Automation bias predates generative AI and affects both novices and experts.26
Human responsibility
The person or institution making the decision remains responsible for source inspection, domain review, privacy, consequences and the final claim. Delegating text generation does not delegate accountability.
A 2025 CHI study surveyed 319 knowledge workers about 936 real-world uses. Participants reported less critical-thinking effort when confidence in the AI was higher, while higher confidence in their own ability was associated with more critical engagement. Because this was self-report research, it identifies relationships and design questions—not proof that AI use caused cognitive decline.25
- Voice the problem yourself: state the decision, constraints and current model before prompting.
- Elicit rivals: ask for counterexamples, alternative frames, failure modes and test cases.
- Reach original sources: open papers, official documents, data and quotations; citation-shaped text is only a lead.
- Inspect scope and calculations: recompute important numbers and verify the exact population and claim.
- Protect people: filter sensitive data, examine harms and record who owns the final judgement.
The NIST Generative AI Profile organizes risk work around governance, mapping, measurement and management.27 For an individual thinker: define the use, anticipate failure, test against reality and preserve a route for correction.
Fluency is not verification
Confirm that every important source exists and supports the sentence, inspect methods when the claim depends on them, and keep responsibility with the human or institution making the decision.
Train the complete cycle, then test transfer
Knowing the vocabulary of logic can improve explanation without changing behaviour. Practice must require decisions, feedback, revision and application beyond the exercise that taught the rule.
Arguments
Map claims, premises, warrants and rivals. Rewrite an opponent’s case until they would accept the summary.
Evidence
Audit measurement, selection, comparison and causation. Translate one probability problem into natural frequencies.
Creation
Reframe one problem three ways and generate options through at least five different mechanisms.
Decisions
Use thresholds, a matrix, premortem and prediction log; review whether new evidence changed confidence appropriately.
Six exercises with visible outputs
- Opposite file: collect the strongest evidence against a belief you care about.
- Three-frame rule: express one problem as prevention, enablement and system redesign.
- Mechanism quota: produce options through at least five distinct causal routes.
- Red-team contract: define the requested criticism and protect the critic from retaliation.
- Sensitivity pass: change weights and assumptions until the choice changes; inspect why.
- Transfer test: apply the method to a new domain without the teaching prompts visible.
Measure transfer, not familiarity
Track whether you identify the real claim, notice missing evidence, generate non-redundant options, calibrate confidence and revise after feedback in unfamiliar problems. Build domain knowledge alongside the framework; fast judgement becomes valuable when expertise—not haste—supports it.
Feedback must reach the reasoning, not only the answer
When reviewing a mistake, ask which representation, assumption, search strategy or decision rule produced it. Correct answers reached through defective reasoning remain dangerous because the same process may fail when luck changes.
Strong thinking is less theatrical than its myths
It rarely looks like instant certainty. It looks like precise questions, organized knowledge, visible assumptions, provisional confidence and correction.
Think in cycles: understand, imagine, test, choose and revise
Critical thinking protects belief from weak inference. Creative thinking prevents the first available option from becoming destiny. Decision discipline makes trade-offs visible. Experimentation lets reality correct imagination. Reflection preserves what the result taught.
None of these eliminates uncertainty, and no checklist replaces intelligence, expertise or moral responsibility. Their value is that they make powerful minds more inspectable and collaborative without flattening genuine differences in ability. Protect the quiet in which original thought forms, the knowledge from which it grows, the people capable of developing it and the honest community that can challenge, build and celebrate what survives.
Sources and further reading
Foundational works, major reviews, meta-analyses and authoritative guidance supporting this article.
- Facione. Critical Thinking: A Statement of Expert Consensus for Purposes of Educational Assessment and Instruction (1990).
- Abrami et al. Strategies for Teaching Students to Think Critically: A Meta-Analysis (2015).
- Batdı et al. Evaluation of the effectiveness of critical thinking training on critical thinking skills and academic achievement (2024).
- Stanovich, West & Toplak. Myside Bias, Rational Thinking, and Intelligence (2013).
- Stanovich & West. On the relative independence of thinking biases and cognitive ability (2008).
- Toplak, West & Stanovich. The Cognitive Reflection Test as a predictor of performance on heuristics-and-biases tasks (2011).
- Lord, Lepper & Preston. Considering the opposite: A corrective strategy for social judgment (1984).
- Gigerenzer & Hoffrage. How to improve Bayesian reasoning without instruction: Frequency formats (1995).
- Tversky & Kahneman. Judgment under uncertainty: Heuristics and biases (1974).
- Rohrer. Thinking Clearly About Correlations and Causation (2018).
- Open Science Collaboration. Estimating the reproducibility of psychological science (2015).
- Mellers et al. Psychological strategies for winning a geopolitical forecasting tournament (2014).
- Sellier, Scopelliti & Morewedge. Debiasing Training Improves Decision Making in the Field (2019).
- Scott, Leritz & Mumford. The effectiveness of creativity training: A quantitative review (2004).
- Runco & Acar. Divergent Thinking as an Indicator of Creative Potential (2012).
- Oppezzo & Schwartz. Give your ideas some legs: The positive effect of walking on creative thinking (2014).
- Sio & Ormerod. Does incubation enhance problem solving? A meta-analytic review (2009).
- Beaty, Benedek, Silvia & Schacter. Creative Cognition and Brain Network Dynamics (2016).
- Beaty et al. Robust prediction of individual creative ability from brain functional connectivity (2018).
- Gick & Holyoak. Schema induction and analogical transfer (1983).
- Bilalić, McLeod & Gobet. Why good thoughts block better ones: The mechanism of the pernicious Einstellung effect (2008).
- Kapur. Productive Failure (2008).
- Mullen, Johnson & Salas. Productivity loss in brainstorming groups: A meta-analytic integration (1991).
- Nemeth. Differential Contributions of Majority and Minority Influence (1986).
- Lee et al. The Impact of Generative AI on Critical Thinking (2025).
- Parasuraman & Manzey. Complacency and Bias in Human Use of Automation (2010).
- Autio et al., National Institute of Standards and Technology. Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (2024).
- Toulmin. The Uses of Argument (updated edition, 2003).
- Swaryandini et al. Educational approaches to reduce cognitive biases among students: systematic review and meta-analysis (2025).
- Chi, Feltovich & Glaser. Categorization and representation of physics problems by experts and novices (1981).
Educational note: High-stakes scientific, medical, legal, financial, engineering or public-safety decisions require relevant expertise, suitable evidence, independent review and responsibility proportionate to possible harm.
Intelligence Unleashed series
- Cognitive Training and Mental Exercises
- Learning New Skills
- Mindfulness and Meditation
- Memory Improvement Techniques
- Critical Thinking and Problem-Solving
- Healthy Lifestyle Habits
- Social Engagement
- Technology and Tools
- Nootropics and Supplements