Artificial Intelligence Assistants
Linas JuozenasShare
AI assistants:
amplify the mind
An AI assistant can compress hours of searching, drafting or routine transformation into minutes. That is real power. The highest standard, however, is not whether the machine produced something quickly—it is whether the person became more capable, the decision became better, the truth remained intact and human agency stayed in command.
agency
goal
context
generate
evidence
decision
learning
What an AI assistant is—and is not
The word “assistant” describes an interface and a role, not a guarantee of knowledge, loyalty or judgment.
A modern assistant may combine a language model with speech recognition, image understanding, search or retrieval, personal memory, software tools and permission to take actions. The model predicts useful continuations from patterns in data; retrieval can supply documents; tools can calculate, browse, write files or operate services. The result may feel like one fluent mind, but it is better understood as a stack of components with different failure modes.
Interface
Voice, chat, camera, screen or embedded controls translate a person’s request into machine-readable input.
Model
A model interprets context and generates language, code, images or a proposed plan. Fluency is not proof.
Grounding
Search, databases and supplied files can anchor an answer in evidence—but retrieval can still be incomplete or wrong.
Tools & permissions
Calendars, email, code runners and other tools turn words into consequences. More autonomy demands tighter controls.
A conversation can create a false sense of certainty
Natural language is socially powerful. A warm tone, confident explanation or apparent memory can make a system feel as if it understands, witnessed or verified more than it actually did. Treat personality as interface design. Judge claims by evidence, and actions by permission, reversibility and review.
An autonomy ladder
| Mode | What the system does | Human responsibility | Minimum safeguard |
|---|---|---|---|
| Answer | Explains, translates, summarizes or generates options. | Interprets and decides. | Check material facts and source quality. |
| Draft | Creates text, code, analysis or plans for review. | Edits, tests and accepts authorship responsibility. | Keep an inspectable draft and validation step. |
| Recommend | Ranks choices or proposes a decision. | Examines assumptions, affected people and alternatives. | Independent evidence for consequential choices. |
| Act with approval | Prepares an email, purchase, change or workflow and waits. | Reviews the exact action and target. | Clear preview, scoped permissions and confirmation. |
| Act autonomously | Executes multiple steps without reviewing each one. | Defines boundaries, monitors and remains accountable. | Least privilege, logs, spending limits, stop controls and recovery. |
What productivity studies actually show
AI can create substantial gains on some tasks. The size and direction of the effect depend on the task, system, person, workflow and definition of quality.
Faster and better within a defined task set
In a preregistered experiment with 453 college-educated professionals completing mid-level writing tasks, access to ChatGPT reduced completion time by 40% and increased independently rated output quality by 18% on average.[1] This supports a real task-performance benefit; it does not establish the same effect for every occupation or long-term learning.
Large benefits for newer workers
A study of 5,172 customer-support agents found a 15% average increase in issues resolved per hour. Less-experienced and lower-performing agents gained most; the final paper reports about a 30% increase for the less-skilled group, while highly skilled agents saw little productivity change and a small decline in conversation quality.[2]
Powerful inside the capability frontier
Among 758 consultants, those using GPT-4 completed 12.2% more tasks within the tested capability frontier, finished them 25.1% faster and produced higher-quality work. On a task deliberately outside that frontier, AI users were 19% less likely to reach the correct answer.[3]
Expert work can resist simple speed claims
In a 2025 randomized study of 16 experienced open-source developers performing 246 tasks in repositories they knew, early-2025 AI tools made completion 19% slower on average—even though developers expected a speed-up.[4] A 2026 follow-up found signs that newer tools may help more, but strong selection and measurement problems prevented a reliable current effect estimate.[5]
The honest conclusion
“AI increases productivity” is too broad. A defensible statement is: on some well-matched tasks, with competent review, particular systems have improved speed, throughput or rated quality. That is useful and important. It is also conditional.
Human plus AI is not automatically the strongest team
A preregistered meta-analysis of 106 experiments and 370 effects found that human–AI combinations outperformed unaided humans on average (g = 0.64), yet underperformed whichever was stronger—the human or AI alone—on average (g = −0.23). Results were extremely heterogeneous, and the included studies ended in June 2023, before many current assistants.[6] Collaboration must be designed and tested; adding a person to a model, or a model to a person, does not guarantee synergy.
Measure value, not visible busyness
The jagged frontier: nearby tasks, opposite outcomes
A system can be astonishing on one task and confidently unreliable on the next, even when both look equally difficult to a person.
Transformation
Reformatting supplied material, generating alternatives, translating tone, extracting fields and drafting from a verified brief.
Synthesis
Summarizing long documents, comparing evidence and coding within an unfamiliar system can work well—if context is complete and review is skilled.
Truth and judgment
Novel facts, causal claims, ambiguous requirements, value conflicts and rare edge cases may exceed the model’s reliable frontier.
The boundary moves as models and tools change, and it differs by language, domain and context. Familiarity can also deceive: a polished response may resemble the form of expert work without containing the expert’s causal understanding. The safest user is neither reflexively trusting nor reflexively dismissive. They learn where a particular system helps, design checks around its weaknesses and remain willing to work without it.
A quick frontier test before delegation
- Define successSpecify the answer, evidence and acceptable error.
- Estimate stakesWhat happens if the output is subtly wrong?
- Probe the edgeTry known cases, adversarial cases and missing context.
- Assign reviewUse a person capable of detecting the relevant failure.
- Record learningUpdate the workflow when the system or task changes.
The review paradox
The more difficult an error is to see, the more expertise review requires. AI can help novices approach established patterns, but experts remain essential precisely where the pattern breaks. A society that benefits from distilled expertise should protect the people who developed it, preserve their authority to challenge automation and credit their contribution.
Learning, memory and intelligence
The most valuable assistant is not always the one that gives the answer fastest. It may be the one that helps a mind form stronger models, retrieve knowledge and solve the next problem independently.
Immediate performance and durable capability are different outcomes. A completed essay, correct equation or working program shows that a human–tool system produced a result. It does not by itself show what the person learned. To measure growth, remove the tool later and test recall, explanation, transfer to a new problem and error detection.
This distinction matters because intelligence is not decorative. Reasoning, memory, learning speed, knowledge and the ability to transfer understanding can influence how well a person navigates education, work and complex life decisions. Those capacities deserve cultivation and celebration. They arise from interacting influences—including biology, development, health, opportunity, instruction and years of deliberate learning—and preserving them can require sustained effort. AI should help more people build intellectual strength, while never erasing the rare expertise and original thought from which others benefit.
A higher score on a valid cognitive assessment is real evidence of stronger performance in the abilities that assessment measures. That evidence can matter profoundly: exceptional reasoning and accumulated knowledge may let a person see a danger, solution or possibility that no available tool or committee can recover without them. A score is not a shop-bought ornament, and a model’s borrowed fluency is not a substitute for a brain shaped by aptitude, development and lifelong learning. Respect intelligence by protecting cognitive health, giving gifted and dedicated minds room to work, listening when their expertise is relevant and recognizing the value they create. Careful measurement limits what a score proves; it should never be used to pretend that measured ability is unimportant.
The goal is not to prove that the tool can think for you. The goal is to leave you able to think farther than before.
A human-development standard for AI assistanceHelp during practice can hide weaker learning
In a class-randomized study of 839 students at one Turkish high school, an unrestricted GPT-4 interface improved supported practice performance, but those students later performed 17% worse than the control group on an unassisted examination. A guarded tutor using teacher-authored material, hints and refusal to supply full answers raised practice performance while eliminating that exam penalty; it did not significantly outperform the control on the exam.[7] The study covered four lessons and an immediate test, not long-term development.
Carefully engineered tutoring can improve immediate learning
Among 194 eligible Harvard physics students across two lessons, a bespoke GPT-4 tutor produced an estimated 0.63-standard-deviation immediate post-test advantage over an expert active-learning class, with a median study time of 49 rather than 60 minutes.[8] This is evidence for a deliberately designed tutor in a specific setting—not proof that a generic chatbot replaces teachers. Delayed retention and semester outcomes were not tested.
Attempting first can change what an explanation teaches
Across two experiments with 1,818 U.S. adults, correct LLM explanations improved immediate performance on similar unassisted problems compared with receiving answers alone, particularly when participants attempted the problem before seeing the explanation.[9] The transfer was near, brief and measured with only a small number of items, so it supports a learning design—not a claim of broad or lasting intelligence growth.
Confidence changes where thinking happens
A 2025 study surveyed 319 knowledge workers about 936 real uses. Greater confidence in generative AI was associated with less self-reported critical-thinking effort; greater confidence in one’s own ability was associated with more. Participants also described a shift toward verification, integration and stewardship.[10] The study reports perceptions and associations, not proof that AI caused cognitive decline.
Turn the assistant into a tutor, not an answer dispenser
Weak learning loop
- Ask for a finished answer immediately.
- Read passively and accept fluent wording.
- Copy the result without reconstructing it.
- Repeat dependence on the next similar task.
Growth-oriented learning loop
- Attempt the problem and expose your reasoning first.
- Ask for one hint, a counterexample or a diagnostic question.
- Explain the answer back in your own words.
- Close the assistant and solve a varied example from memory.
Can AI raise IQ?
There is not yet good evidence that ordinary use of a general AI assistant reliably produces lasting gains in general intelligence or IQ. It can help people acquire knowledge, practise reasoning, receive feedback, translate difficult material and access instruction—ingredients that may improve particular cognitive skills and real-world competence. Any claim of IQ growth should use valid, repeated assessment, account for practice effects and show transfer beyond tasks rehearsed with the system. The absence of a proven universal effect is a reason to measure carefully, not a reason to dismiss intellectual growth.
Protect desirable difficulty
- Retrieve before revealing. Write what you remember before asking for a summary.
- Generate before comparing. Form your own hypothesis, outline or solution before viewing the model’s.
- Ask for causes. Request mechanisms, assumptions, boundary cases and a test that could prove the answer wrong.
- Vary the problem. Change the context and solve again without assistance to test transfer.
- Keep a competence floor. Regularly perform core skills unaided so emergency, novel and adversarial situations remain manageable.
Originality, chosen solitude and community
A mind sometimes needs freedom from prompts—including human and machine prompts—to discover what it actually sees.
Original thought often begins before consensus. When other people—or an AI trained on prevailing patterns—immediately suggest what to notice, how to frame the problem or what a “good answer” looks like, attention can be anchored prematurely. The influence may be well meant and still narrow the search space. Chosen solitude creates a protected interval in which curiosity can wander, unusual associations can form and an idea can become distinct enough to survive comparison.
That does not make isolation the final ideal. Once an original idea exists, other people become indispensable: they challenge weak assumptions, contribute missing knowledge, build what one person cannot, introduce the work to wider communities and celebrate genuine achievement. Healthy intellectual life therefore alternates between private formation and generous collaboration—not permanent withdrawal, and not permanent exposure to other people’s directions.
Step away long enough to hear your own thought. Return when you choose—and let the return become a celebration: testing, supporting, extending, crediting and sharing what one mind began.
Solitude and community are complementary intellectual conditionsHigher individual ratings, lower collective diversity
In a 2024 experiment, 293 writers produced short stories and 600 blinded readers evaluated them. Access to one AI idea increased average novelty ratings by 5.4% and usefulness by 3.7%; access to as many as five increased them by 8.1% and 9.0%. Gains were concentrated among writers who scored lower on the study’s narrow baseline creativity task. Yet AI-assisted stories became more similar to one another and about 5% more similar to the supplied ideas.[11] The task was eight-sentence fiction, not creativity as a whole, but it shows how individual uplift can coexist with collective convergence.
A private first pass can preserve agency and distance from the model
In a small 2025 experiment, 60 participants either used GPT-4 from the beginning of a 20-minute health-topic ideation task or only after creating at least three ideas independently. Delayed-AI participants produced ideas less similar to the model’s output and reported greater creative self-efficacy, autonomy, ownership and self-credit.[12] The study did not include a no-AI group, and its idea-count difference was not conventionally significant. It makes independent first, assistance second a promising protective design—not a universal law.
A four-phase originality protocol
- WanderSpend protected time observing and thinking without incoming suggestions.
- CaptureRecord the idea, reasoning, date and early evidence in your own words.
- ChallengeUse people and AI to find precedent, objections, tests and missing expertise.
- BuildCollaborate with clear roles and preserve a trace of who contributed what.
- CelebrateCredit originators and collaborators; help strong ideas reach the people they can serve.
Authorship is more than pressing “generate”
Responsibility and credit should follow meaningful human contribution: conceiving the purpose, choosing or rejecting material, making expressive decisions, revising, verifying and accepting accountability. Keep version history when provenance matters. Current U.S. Copyright Office guidance centres copyright protection on human authorship and asks applicants to disclose more-than-minimal AI-generated material in a registered work.[13]
Respect intelligence without ranking human dignity
Cognitive abilities and contributions are not interchangeable. Some people develop exceptional reasoning, knowledge or creative power through rare aptitude and a lifetime of work; society benefits when it protects, listens to and fairly credits them. Equal human dignity does not require pretending all performance is equal. Recognition should honour real achievement without turning a score into permission to mistreat anyone.
Accessibility and inclusion: more ways to think, speak and act
For many people, an assistant is not a novelty. It can be a bridge across a mismatch between a person’s abilities and an inflexible environment.
Language and cognition
Plain-language rewriting, stepwise instructions, vocabulary support and adjustable pacing can reduce unnecessary barriers without reducing the seriousness of the idea.
Sensory access
Captions, text-to-speech, image descriptions and multimodal input can translate information between visual, auditory and textual forms.
Motor and speech access
Voice control, switch-compatible interfaces, text generation and augmentative communication can reduce the effort required to express intent or operate software.
Good accessibility expands control; it does not force everyone through the same “intelligent” channel. A person should be able to type instead of speak, read instead of listen, slow or stop output, correct recognition, preserve their own voice and choose whether personalization is worth the data it requires. The Web Content Accessibility Guidelines provide a durable foundation for perceivable, operable, understandable and robust digital experiences; an AI feature does not replace accessible structure, keyboard support, labels, captions or testing with disabled users.[14]
Recognition is not equally accurate for everyone
Speech systems can work differently across accents, dialects, speaking styles, ages, noise conditions and disabilities. A 2020 study of five commercial speech-recognition systems found substantially higher word-error rates for Black speakers than for White speakers in the tested U.S. English datasets.[15] The products have since changed, so those figures are not a current scorecard; the study remains a clear demonstration that average accuracy can conceal unequal failure.
The person relying on an output may be least able to check it
A 2025 U.S. survey shaped by 19 disability organizations included 1,735 adults, 1,070 of them disabled. Among 290 blind and low-vision respondents using AI visual description, 21% reported an error that had harmed them; the rate was 9% in each sighted comparison group. Among 54 Deaf or hard-of-hearing caption users, 24% reported harm from an error.[16] Subgroup sizes and self-selection limit population estimates, but the asymmetry is unmistakable: plausible guessing can carry more risk when the output substitutes for unavailable perception.
Faster expression must remain the person’s expression
In a CHI study with 12 people who use augmentative and alternative communication, live language-model suggestions could reduce typing time and effort, while participants emphasized preserving their own style and preferences.[17] Suggestions should therefore be optional, editable and easy to reject or undo—never spoken or sent automatically.
For product teams
- Test with disabled people and compensate their expertise.
- Publish results across relevant user groups and conditions.
- Provide non-voice and non-AI alternatives for essential actions.
- Make errors easy to detect, correct and undo.
- Do not infer incompetence from a recognition failure.
For schools and workplaces
- Distinguish legitimate accommodation from hidden substitution.
- Evaluate the person’s knowledge through an accessible route.
- Do not require disclosure of sensitive conditions to a chatbot.
- Keep human support available when automation fails.
- Let the user control tone, pace, format and memory.
Accessibility needs vary. A feature that helps one person may burden another; user choice and direct testing are more reliable than assumptions.
Truth, calibration and the discipline of checking
An assistant can state a true fact, a plausible inference and an invented detail in the same calm voice. Verification must therefore follow the claim, not the tone.
NIST uses confabulation for confidently presented false or erroneous content from generative systems. Its Generative AI Profile also identifies privacy, harmful bias, information integrity, cybersecurity and human–AI configuration as risks that require context-specific measurement and management.[18] The profile is a voluntary risk-management resource, not a law or a certificate that any system is safe.
Models can invent citations, misread a source, perform an invalid calculation, omit contrary evidence or agree with the user’s premise. Older benchmarks make the phenomenon visible but should not be reused as current product ratings. For example, TruthfulQA tested 817 misconception-sensitive questions across 38 categories and found large truthfulness gaps in the models available to its researchers at the time.[19] It does not tell us the accuracy of every assistant in 2026.
Confabulation
The model produces unsupported facts, quotations, links, legal cases, data or reasoning that fit the requested shape.
Sycophancy
The system follows a user’s expressed belief or desired conclusion instead of maintaining an evidence-based position; controlled studies have demonstrated this tendency in tested assistants.[20]
Unequal error
Performance can differ across groups, languages and contexts. Benchmark results such as BBQ reveal stereotype-related patterns under controlled conditions, but are not universal discrimination rates.[21]
A verification ladder matched to consequence
| Consequence if wrong | Examples | Appropriate checking | Do not rely on |
|---|---|---|---|
| Low | Brainstorming names, changing tone, reformatting your own notes. | Read for fit; keep or discard. | The assistant’s taste as an objective fact. |
| Moderate | Public article, analysis, routine code, study explanation. | Open cited sources, test examples, review by a knowledgeable person. | A reference merely because it has a realistic title or DOI. |
| High | Medical, legal, financial, employment, safety or security decisions. | Use authoritative current sources and an appropriately qualified human; document assumptions and review. | An AI disclaimer, majority vote among chatbots or confident wording. |
| Irreversible | Publication, deletion, payment, account change, disclosure or physical action. | Preview the exact action and target; require fresh confirmation and a recovery plan. | Broad standing permission or an ambiguous command. |
Calibrated friction protects thought
In a controlled experiment with 199 participants, interfaces that required people to engage with a problem before receiving or accepting AI advice reduced overreliance on incorrect recommendations. The most effective friction was also liked less.[22] Convenience and good judgment can pull in different directions; excellent design makes important thinking unavoidable without making every harmless task difficult.
The claim-by-claim check
- Separate facts from synthesis. Mark dates, quantities, quotations and causal claims that can be checked.
- Go to the source. Confirm that a cited paper or official document exists and supports the exact sentence.
- Recalculate. Run numbers independently; inspect units, denominators, missing values and base rates.
- Search for disconfirmation. Ask what evidence would reverse the conclusion, then look for it outside the conversation.
- State uncertainty honestly. “Unknown,” “mixed” and “not tested” are stronger than invented precision.
Privacy, security and permission
Before asking what an assistant can do, ask what it can see, remember, send and change.
Input
Prompts, voice, images, files, screens, location and surrounding conversations.
Derived data
Transcripts, embeddings, inferred interests, summaries, profiles and safety labels.
Retention
Conversation history, service logs, backups, human review queues and connected-app records.
Action
Messages, purchases, code, calendars, records and devices reachable through tools or integrations.
Voice assistants: listening is not one single state
Many voice assistants run a local wake-word detector over a short rolling audio buffer; that does not necessarily mean every ambient conversation is uploaded. Detection is probabilistic, however. A false activation can cause unintended speech to be transmitted, transcribed or retained depending on the service and settings. European data-protection guidance highlights accidental activation, bystander data, purpose limitation, retention and meaningful deletion as central design concerns.[23]
A controlled 2020 study played 134 hours of television material near tested smart speakers. False activations varied by device; almost all activations observed locally in that laboratory setup were sent to the cloud.[24] The study demonstrates a real mechanism, not the current activation frequency of every product in an ordinary home.
Private by default is stronger than “be careful”
Do not paste confidential work, health records, identity documents, private messages, unpublished inventions, children’s data or another person’s information into a service unless you know the governing agreement, purpose, retention, training use, deletion process and access controls. De-identification is harder than deleting a name: combinations of details can identify people.
When an assistant can use tools, untrusted content can become an instruction
An assistant may read a webpage, email or document that contains malicious text telling it to ignore the user, disclose data or call a tool. Research demonstrations of indirect prompt injection have shown manipulation, data exfiltration and misuse of connected functions in vulnerable LLM-integrated applications.[25] The attack does not magically create access: consequences are constrained—or amplified—by the data, credentials, tools and permissions the assistant already has.
Personal privacy controls
- Review history, retention, training and connected-app settings.
- Delete recordings and revoke integrations you no longer need.
- Prefer on-device processing where it fits—but still inspect telemetry and sync.
- Use temporary chats or separate accounts for bounded contexts when available.
- Ask before recording or summarizing other people.
Tool-security controls
- Give the minimum permission for the shortest practical time.
- Separate read access from write access.
- Keep secrets outside model-visible context.
- Preview recipients, destinations, data and cost before action.
- Retain logs carefully, because logs can become sensitive data too.
A prompt is a behavioural instruction, not an access-control boundary. Enforce security with architecture, permissions, isolation, confirmation and recovery.
Attention, dependence and safety
Offloading is not automatically harmful. It becomes a problem when the saved effort is never reinvested in judgment, learning, rest or work that matters more.
Humans have always used notebooks, maps, calculators and other people to extend memory and action. Cognitive offloading can free limited working capacity and reduce avoidable effort. The cost appears when a person no longer knows what was delegated, cannot notice a bad output, loses a core skill that remains necessary, or begins consulting the system before forming any intention of their own.
Healthy augmentation
You can explain the result, reject the assistant, work without it when needed and use saved time for higher-value thought.
Growing dependence
You ask before trying, accept outputs you cannot inspect, feel unable to begin alone or lose track of what you actually believe.
Workflow failure
No competent person owns the answer; automation errors pass silently; essential work stops when the service, account or connection fails.
Hands-free does not mean attention-free
Voice interaction can reduce visual and manual demands, but complex dialogue still occupies attention. In a 2015 AAA Foundation study of 257 participants using ten model-year-2015 vehicle systems, voice tasks produced moderate-to-high cognitive workload, with measurable residual effects lasting as long as 27 seconds after some tasks.[26] The maximum was not the duration after every interaction, and systems have changed. The practical rule endures: set navigation and media before moving, keep essential voice use brief, avoid it in demanding traffic and pull over for complex tasks.
Rebalance before convenience becomes compulsion
- Create assistant-free starts. Begin important thinking, writing and planning from your own attention.
- Choose the offload. Delegate repetitive transformation; keep goal-setting, value judgments and learning-critical steps active.
- Practise recovery. Maintain offline copies, manual procedures and core knowledge for outages and novel failures.
- Notice emotional substitution. A conversational system can rehearse a difficult conversation or help organize thoughts, but it cannot consent, care, witness or accept responsibility as a person can.
- Return to people. Mentors, peers and loved ones provide reciprocal challenge, lived context and shared responsibility that generated conversation cannot replace.
The personal playbook: attempt, assist, verify, retain
Good AI use is not a collection of clever prompts. It is a workflow in which the human purpose remains visible from beginning to end.
- AttemptThink, recall, sketch or solve enough to expose your own model.
- AssistDelegate a bounded role: tutor, critic, editor, simulator or transformer.
- VerifyCheck claims, calculations, sources, constraints and affected people.
- DecideExercise human judgment and own the final action or publication.
- RetainExplain, practise, document and test what remains after the tool closes.
Choose a role before you open the conversation
Tutor
“Do not solve it yet. Ask one diagnostic question, then give the smallest hint that lets me continue.”
Best when learning and independent transfer matter.Critic
“Identify the strongest objection, hidden assumption and boundary case. Do not rewrite my position.”
Best when protecting authorship and testing reasoning.Research aide
“Separate verified facts, plausible interpretation and unknowns. Give primary sources for each material claim.”
Best for mapping evidence; every source still needs inspection.Editor
“Preserve my argument and voice. Flag ambiguity and propose local edits with reasons.”
Best after the human has formed the substance.Simulator
“Respond as a sceptical reader with these stated assumptions. Show where this framing fails.”
Useful for rehearsal; not a substitute for real stakeholders.Transformer
“Convert only the supplied material into this structure. Mark anything that requires new information.”
Best for bounded, inspectable changes.A strong prompt contains a stop condition
State what the assistant may use, what it must not invent, which uncertainties it should surface and when it should ask instead of act. But remember: instructions improve behaviour; they do not enforce security. Technical permissions must still constrain consequential actions.
Match the workflow to the purpose
| Purpose | Human-first step | Useful AI role | Final test |
|---|---|---|---|
| Learn | Attempt and state where understanding breaks. | Socratic questions, examples, feedback and adaptive practice. | Solve and explain a new case without AI. |
| Create | Capture your own observations, taste and early alternatives. | Research precedent, stress-test and expand execution options. | Can you identify the human idea and defend every retained choice? |
| Decide | Define values, stakeholders, evidence and acceptable risk. | Compare scenarios and expose assumptions. | Independent evidence plus accountable human approval. |
| Communicate | Know the meaning, audience and desired effect. | Structure, translate, clarify and test interpretations. | Read as the audience; verify every factual claim and quotation. |
| Automate | Map the process, exceptions and failure cost. | Execute narrow repeatable steps within explicit permissions. | Logs, sampling, stop thresholds, rollback and a named owner. |
Before
- What am I trying to achieve?
- Which part must I personally understand?
- What data am I allowed to share?
- What would a dangerous error look like?
After
- Which claims did I verify independently?
- What did I change, reject or contribute?
- Can I explain the result without the chat?
- What should be remembered—and what should be deleted?
Teams, education and governance
“A human is in the loop” is not a safeguard if that human lacks time, authority, evidence or the skill to detect the error.
Organizations should govern the whole sociotechnical system: model, data, interface, permissions, workers, incentives, affected people and appeals. NIST’s AI Risk Management Framework organizes this work around governing, mapping, measuring and managing risk across the lifecycle.[27] It is voluntary guidance; laws, contracts and professional duties may impose additional requirements.
Inventory
Know which systems, models, integrations, datasets and unofficial “shadow” uses exist.
Classify
Separate low-stakes drafting from decisions affecting rights, access, livelihood, health or safety.
Validate
Test on real tasks, languages and user groups; measure quality, error distribution and overrides.
Own
Name a responsible person, provide an appeal route, monitor change and learn from incidents.
A governance baseline
- Purpose before procurement. Define the problem and comparison baseline before selecting a product.
- Approved data boundaries. Specify which data classes may enter which systems; enforce the rule technically.
- Consequential decisions receive competent review. Reviewers need domain knowledge, time, authority and access to the underlying evidence.
- Learning is a measured outcome. Schools test unaided transfer; employers preserve pathways from novice to expert instead of removing every practice opportunity.
- Expertise is protected and credited. Reward the people whose tacit knowledge improves workflows; do not use automation to erase attribution or silence dissent.
- Accessibility is verified. Maintain alternative routes and test the complete workflow, not a feature checklist.
- Incidents have a response path. Stop, contain, notify where required, correct affected records and update controls.
- Sunset is possible. Retain export, deletion and migration options so a vendor or model change does not own the institution’s memory.
Do not automate away the apprenticeship ladder
If AI performs all entry-level research, drafting and diagnosis, novices may lose the repetitions through which expert judgment is formed. Redesign roles so learners still predict, practise, receive feedback and examine failures. Senior experts should spend less time on mechanical repetition—but more time teaching judgment, defining standards, handling frontier cases and originating new methods.
Rights, data protection and transparency
In the European Union, identifiable prompts, recordings and transcripts may be personal data under the GDPR. A voice recording becomes special-category biometric data when processed through specific technical means for unique identification; not every recording is automatically biometric. Lawful basis, transparency, purpose limitation, minimization, security, retention and data-subject rights still require context-specific analysis.[28]
The EU AI Act applies obligations according to the actor, intended purpose and risk category. A general chatbot is not automatically a high-risk system. As of September 2026, most provisions apply, including relevant transparency duties; particular high-risk requirements follow the amended schedule. AI Act compliance does not replace data protection, consumer, employment, safety or sector-specific law.[29] Organizations should obtain current, jurisdiction-specific advice for real deployments.
Evidence beats adoption theatre
Licences purchased, prompts sent and documents generated are activity measures. A serious deployment reports task quality, time including correction, learning, unequal errors, incidents, employee and user experience, privacy cost and outcomes after the tool changes. It also permits the conclusion that some tasks should not use AI.
What comes next
The important shift is from systems that answer a prompt to systems that perceive context, remember preferences, coordinate tools and pursue multi-step goals.
Multimodal assistance
Text, voice, images, video, screens and sensor input will increasingly share one conversational interface.
Tool-using agents
Assistants will plan and execute longer workflows, making permissions, observability, confirmation and recovery central design problems.
Local and hybrid AI
More work can happen on-device for latency, resilience or privacy, while demanding tasks may still use remote models and services.
Deep personalization
Persistent tutors and collaborators may adapt to a person over years—but memory can also magnify profiling, stale assumptions and dependency.
Agent standards and evaluation are still developing. In February 2026, NIST launched an AI Agent Standards Initiative focused on interoperability, security, identity and adoption; it is an ongoing initiative, not a completed safety standard.[30] Benchmarks will also need to move beyond isolated answers toward long-horizon reliability, recovery from interruption, permission handling, privacy and the quality of human oversight.
The questions worth carrying forward
Does this system extend what people can understand and do—or only produce the appearance of competence?
Does repeated use strengthen memory, reasoning and learning transfer, or remove the practice that builds them?
Does assistance expand the search space while protecting independent first thought, or converge everyone on the same patterns?
Who controls the model, memory, data and permissions—and who can inspect, refuse, correct or leave?
Are originators, experts, workers and communities fairly credited for the knowledge the system makes useful?
Does the saved time return to learning, creation, relationship and rest—or merely raise the volume expected from each person?
The best future is not one in which human thought becomes unnecessary. It is one in which more people can learn deeply, original minds can travel farther and powerful tools remain answerable to the lives they affect.
Augmentation worthy of human intelligenceFrequently asked questions
Short answers first; the surrounding article provides the evidence and conditions.
Are AI assistants making people less intelligent?
There is no sound basis for a universal claim. Effects depend on use. Passive substitution can remove practice and conceal weak understanding; tutoring, feedback, accessible explanation and deliberate retrieval can support learning. Judge durable capability with unassisted recall, explanation and transfer—not by how polished an AI-assisted product looks.
Can an AI assistant help someone grow their intelligence or IQ?
It can help build knowledge and practise particular cognitive skills, especially when it asks questions, adapts examples and gives corrective feedback. But there is not yet strong evidence that ordinary assistant use reliably raises general intelligence or produces durable IQ gains. That important possibility should be tested with valid assessments, long follow-up and transfer to genuinely new tasks.
Should experts avoid AI because it can deskill people?
No. Experts may gain greatly from automation, search, simulation and rapid drafting. They should decide which skills can safely be offloaded, preserve the ability to handle edge cases and review quality, and protect time for original work. Organizations should respect expert refusal when a system crosses its reliable frontier.
Is it better to brainstorm alone, with people or with AI?
Use sequence rather than a single rule. Begin alone when independent observation and originality matter. Capture your reasoning. Then involve AI and people to find precedent, challenge assumptions, combine expertise and build the idea. Return to solitude when you need to integrate what you heard.
Does a voice assistant upload everything said nearby?
Not necessarily. Many devices detect a wake word locally using a rolling buffer. False activations can nevertheless transmit unintended speech, and retention differs by service and settings. Review the product’s current controls, delete history you do not need and obtain consent before recording or summarizing other people.
Can I trust links and citations supplied by an AI?
Use them as leads, not evidence. Open each source; confirm author, title, date and publication; then check that it supports the exact claim. Search independently for omitted or contrary evidence. Invented references can look entirely plausible.
Is AI-generated work plagiarism?
Rules differ by school, workplace, publication and jurisdiction. Undisclosed substitution can violate authorship or assessment rules even when wording is new. Disclose use when required, cite underlying sources, preserve your contribution and never present another person’s ideas or labour as your own. Copyright, plagiarism and academic integrity are related but not identical questions.
When should I not use a general AI assistant?
Avoid it when you cannot lawfully or ethically share the necessary data, when nobody can verify the result, when failure would be unacceptable, or when the task’s purpose is to assess your unaided ability. For high-stakes medical, legal, financial, safety or rights-affecting decisions, use current authoritative sources and appropriately qualified human judgment.
Evidence and sources
Primary studies, peer-reviewed papers and official frameworks are prioritized. Results describe the tested people, tasks, systems and dates—not every possible AI assistant.
- Noy, S. & Zhang, W. (2023). Experimental evidence on the productivity effects of generative artificial intelligence. Science. Randomized experiment
- Brynjolfsson, E., Li, D. & Raymond, L. (2025). Generative AI at Work. The Quarterly Journal of Economics. Field study
- Dell’Acqua, F. et al. (2026). Navigating the Jagged Technological Frontier. Organization Science. Randomized experiment
- Becker, J. et al. (2025). Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity. METR. Randomized experiment
- METR (2026). We Are Changing Our Developer Productivity Experiment Design. Study update
- Vaccaro, M., Almaatouq, A. & Malone, T. (2024). When combinations of humans and AI are useful: A systematic review and meta-analysis. Nature Human Behaviour. Preregistered meta-analysis
- Bastani, H. et al. (2025). Generative AI without guardrails can harm learning: Evidence from high school mathematics. PNAS. Field experiment
- Kestin, G. et al. (2025). AI tutoring outperforms in-class active learning: an RCT introducing a novel research-based design in an authentic educational setting. Scientific Reports. Randomized crossover
- Kumar, H. et al. (2025). Math education with large language models: Peril or promise?. AIED 2025. Preregistered experiments
- Lee, H.-P. et al. (2025). The Impact of Generative AI on Critical Thinking. CHI 2025. Survey study
- Doshi, A. R. & Hauser, O. P. (2024). Generative AI enhances individual creativity but reduces the collective diversity of novel content. Science Advances. Controlled experiment
- Qin, P. et al. (2025). Timing Matters: How Using LLMs at Different Timings Influences Writers’ Perceptions and Ideation Outcomes in AI-Assisted Ideation. CHI 2025. Randomized experiment
- U.S. Copyright Office (2023). Copyright Registration Guidance: Works Containing Material Generated by Artificial Intelligence. Official guidance
- W3C (2023). Web Content Accessibility Guidelines (WCAG) 2.2. Standard
- Koenecke, A. et al. (2020). Racial disparities in automated speech recognition. PNAS. Comparative study
- American Foundation for the Blind (2026). AI & Disability: Research Results—Innovation & Access. Disability-led survey
- Valencia, S. et al. (2023). The Less I Type, the Better: How AI Language Models Can Enhance or Impede Communication for AAC Users. CHI 2023. User study
- NIST (2024). Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile. NIST AI 600-1. Official framework
- Lin, S., Hilton, J. & Evans, O. (2022). TruthfulQA: Measuring How Models Mimic Human Falsehoods. ACL 2022. Benchmark study
- Sharma, M. et al. (2024). Towards Understanding Sycophancy in Language Models. ICLR 2024. Evaluation study
- Parrish, A. et al. (2022). BBQ: A Hand-Built Bias Benchmark for Question Answering. Findings of ACL 2022. Benchmark study
- Buçinca, Z., Malaya, M. B. & Gajos, K. Z. (2021). To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-Assisted Decision-Making. Controlled experiment
- European Data Protection Board (2021). Guidelines 02/2021 on Virtual Voice Assistants. Official guidance
- Dubois, D. J. et al. (2020). When Speakers Are All Ears: Characterizing Misactivations of IoT Smart Speakers. Proceedings on Privacy Enhancing Technologies. Controlled study
- Greshake, K. et al. (2023). Not What You’ve Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection. Security research
- AAA Foundation for Traffic Safety (2015). Measuring Cognitive Distraction in the Automobile III. Experimental study
- NIST (2023). Artificial Intelligence Risk Management Framework 1.0. Official framework
- European Union (2016). General Data Protection Regulation (EU) 2016/679. Law
- European Commission (updated 2026). AI Act regulatory framework and application timeline. Official overview
- NIST (2026). AI Agent Standards Initiative. Standards initiative