Once a recommendation appears, the person is no longer solving the original problem.
Anatomy of decision support
A decision passes through both technical and human transformations.
01
World
Events, people, goals
→
02
Measurement
Data, sensors, records
→
03
Model
Score, forecast, option
→
04
Human judgment
Interpret, question, choose
→
05
Action
Consequences and feedback
Outcomes become future data, incentives, and expectations.
A recommendation has layers
Good support separates the answer from the evidence behind it.
Prediction
What does the model expect?
“Flood risk: 0.78”
Uncertainty
How stable is that estimate here?
Sensor coverage is incomplete
Rationale
What evidence influenced it?
Rainfall, elevation, road capacity
Action
What should happen next?
Evacuate, inspect, wait, or escalate
Authority is a design choice
Decision support can inform, recommend, constrain, or execute.
01
Inform
Show evidence or predictions.
Human initiates
02
Recommend
Propose a preferred option.
Human accepts or rejects
03
Constrain
Block, rank, or limit choices.
Human chooses within bounds
04
Execute
Act unless stopped or reversed.
System initiates
A person in the loop is not a safety mechanism
Oversight only works when the person can detect, understand, and correct failure.
SEE
Relevant evidence and uncertainty are visible.
→
JUDGE
The person has skill, time, and an independent basis.
→
ACT
They have authority and a usable correction path.
If any gate fails, “human oversight” may be ceremonial.
Humans are not the gold standard
Human judgment is adaptive, contextual, and often inconsistent.
HUMAN
Limited attentionWe cannot inspect everything.
Fatigue and workloadPerformance changes with conditions.
Heuristics and biasShortcuts can help or mislead.
Variable expertisePeople differ across cases and time.
Memory and consistencyWe forget and apply criteria unevenly.
AI is not the gold standard either
AI can be consistent and scalable while remaining systematically wrong.
AI
Partial dataThe model only receives encoded evidence.
Objective mismatchThe optimized score may be a poor proxy.
Distribution shiftDeployment differs from training.
Hidden uncertaintyConfident outputs can still be wrong.
Missing values and contextPrediction cannot settle what should matter.
The opportunity
Neither partner is perfect, but they may be wrong on different cases.
Human errorsfatigue · rare patterns · inconsistency
AI errorsshift · missing context · proxy failure
shared blind spots
Collaboration has potential only when one partner can recognize or repair the other’s error.
The evidence is a warning
Human plus AI does not automatically beat the better solo performer.
106
experiments
370 effect sizes · preregistered meta-analysis
g = −0.23Team versus the better solo performer 95% CI [−0.39, −0.07]
g = +0.64Team versus the human alone 95% CI [+0.53, +0.74]
I² = 97.7%Very high heterogeneity across settings
On average, AI helped people, but the combined workflow still lost value compared with using the stronger partner alone.
The moderators matter
The average changed sharply with the task and with who was stronger alone.
Decision tasks
g = −0.27
Teams underperformed the better solo partner.
95% CI [−0.44, −0.10]
Creation tasks
g = +0.19
Direction was positive, but not statistically different from zero.
95% CI [−0.09, +0.48]
Human stronger alone
g = +0.46
The team beat both solo partners on average.
95% CI [+0.28, +0.66]
AI stronger alone
g = −0.54
Adding the human reduced available AI performance.
95% CI [−0.71, −0.37]
Boundary: About 85% of effect sizes involved decision tasks, and more than 95% used a human-final-decision workflow. These are moderators, not universal laws.
Define the target clearly
Complementary performance means the team beats both partners alone.
Human alone
AI alone
Human + AI
TEAM > MAX(HUMAN, AI)
Clinical evidence: benefit and danger
The same clinicians gained from good AI and lost accuracy when its top answer was corrupted.
In a separate 600-image aggregation, human collectives plus AI reached 81.0% accuracy, versus 73.7% for human collectives and 76.9% for the CNN alone.
Boundary: This was an image-based experimental benchmark, not evidence about patient outcomes or deployment in a clinic.
Two sources of complementarity
Partners add value when they possess different information or different capabilities.
INFORMATION
They see different evidence.
The model sees thousands of historical patterns. The person sees a new constraint, local context, or unrecorded goal.
Can each partner reveal what the other lacks?
+
CAPABILITY
They process evidence differently.
The model searches consistently at scale. The person reasons about values, exceptions, causal stories, and consequences.
Can the workflow allocate work to the stronger contributor?
Build an error map
The useful question is not “Who is better?” but “Who is better here?”
Case conditionsHuman advantageAI advantage
Novel local change●○
Large repetitive search○●
Value conflict●○
Subtle historical pattern○●
Shared missing data××
Averages hide the cases that should be routed, escalated, or deferred.
Appropriate reliance
A good decision-maker accepts correct advice and rejects incorrect advice.
AI correct
AI wrong
Human accepts
Appropriate relianceThe advice repairs or confirms the decision.
OverrelianceThe person follows the system into error.
Human rejects
UnderrelianceUseful advice is ignored.
Appropriate self-relianceThe person catches the model’s error.
Trust and reliance are related, not identical
Trust is an attitude. Reliance is an action.
TRUST
“I expect this system to help me under uncertainty.”
survey · belief · expectation
≠
RELIANCE
“I used, accepted, deferred to, or acted on its output.”
choice · behavior · delegation
People may report skepticism and still follow a default under time pressure.
Reliance should be local
“I trust the AI” is too broad to guide a decision.
Trust it...
for what task?classification, prediction, ranking, generation, action
on which cases?routine, rare, ambiguous, shifted, adversarial
using what evidence?validation, calibration, provenance, recent monitoring
with what stakes?reversible suggestion or consequential action
Automation bias
People often treat an automated cue as a substitute for searching and checking.
AI recommendation
→
attention narrows
contradictory evidence receives less search
the recommendation becomes the default frame
→
error travels through the human
Two signatures of automation bias
Automation can make us miss what it omits and do what it wrongly suggests.
OMISSION
The system stays silent.
The person fails to notice or act because no alert appeared.
No warning → hazard goes unexamined
COMMISSION
The system recommends the wrong action.
The person follows the recommendation despite conflicting evidence.
Wrong warning → unnecessary action
Evidence across domains
Automation bias is common in the literature, but the evidence base has real limitations.
40
studies met review criteria
Human factors and health-care research, 1983–2015
81%25 of 31 studies testing omission errors reported evidence of bias
91%21 of 23 studies testing commission errors reported evidence of bias
9studies tested statistical significance against a manual control
Interpret carefully: Samples were often small and homogeneous, measures varied widely, and most studies did not report an effect size against an unassisted control.
Why overreliance feels reasonable
The problem is rarely “people are lazy.” The system changes the economics of thinking.
FluencyClear language feels easier to accept.
Authority cuesScores and precision look objective.
Time pressureVerification competes with throughput.
Low self-confidencePeople defer when they doubt themselves.
High base accuracyRare failures train complacency.
Default designAccepting is easier than objecting.
Overreliance is produced by people, interfaces, organizations, and incentives together.
A crucial distinction
Cognitive offloading can help. Cognitive surrender gives up the judgment itself.
OFFLOADING
Move part of the cognitive work into an external tool.
Use a calculator, map, checklist, search tool, or AI draft.
The person still frames, evaluates, and owns the decision.
VS
SURRENDER
Adopt the external answer with little independent scrutiny.
The system supplies the frame, conclusion, and stopping point.
The person cannot explain, test, or recover the reasoning.
Cognitive surrender
The danger is not that the AI thinks. It is that the person stops deciding how to think.
“
An emerging term for adopting AI output with minimal scrutiny while intuition, deliberation, or verification is bypassed.
Useful as a diagnostic idea. Not a medical diagnosis and not every instance of AI assistance.
Separate the evidence from the label
The mechanisms are better established than the term “cognitive surrender.”
ESTABLISHED
Cognitive offloading
People routinely move memory and computation into external tools. Benefits and costs depend on the task and what remains internal.
Review literature · Risko & Gilbert, 2016
GROWING
Passive AI reliance
Experiments increasingly connect passive use with overreliance, weaker error detection, and lower self-efficacy or ownership.
Multiple behavioral studies · 2021–2026
EMERGING
Cognitive surrender
A proposed umbrella theory for external AI reasoning displacing intuitive and deliberative judgment.
New theory preprint · Shaw & Nave, 2026
Teach surrender as a diagnostic hypothesis, not as a settled syndrome or a validated measurement scale.
How surrender develops
A fluent answer can create closure before understanding has formed.
01
Need
A difficult or uncertain task
→
02
Fluent answer
Fast, complete, confident output
→
03
Closure
The problem feels resolved
→
04
Skipped work
No independent frame or search
→
05
Adoption
Output becomes judgment and action
Five warning signs
You may have surrendered when the answer survives but your understanding disappears.
1
You cannot explain why the recommendation fits this case.
2
You did not form a view before seeing the output.
3
You searched for confirmation, not contradiction.
4
You cannot continue if the system becomes unavailable.
5
You still own the consequences but no longer feel ownership of the reasoning.
The cost unfolds over time
Overreliance can improve today’s throughput while weakening tomorrow’s judgment.
NOW
Immediate performance
Faster output
Lower effort
More consistency
Possible accuracy gain
→
LATER
Capability and agency
Reduced practice
Weaker situation awareness
Lower self-efficacy
Harder recovery when AI fails
Evidence on passive versus active use
How people used AI mattered more than whether AI was present.
N = 269
Preregistered experiment
Occupation-specific writing tasks
No AIWrite independently
Passive AICopy AI-generated content
Active collaborationDraft first, then use AI to refine
Passive useLower self-efficacy, psychological ownership, and meaningfulness. Self-efficacy and meaningfulness remained lower on a later manual task.
Active usePsychological outcomes were statistically similar to the no-AI condition.
Boundary: This is one writing experiment plus a correlational follow-up survey of 270 workers. It does not establish long-term skill loss across occupations.
The irony of automation
The more routine work automation absorbs, the less prepared people may be for the rare failure left behind.
99%
Normal operation
The system handles familiar cases. Human skill receives little practice.
failure
1%
Abnormal condition
The human must diagnose quickly with incomplete awareness and decayed skill.
The oversight trap
A human can become a rubber stamp with full responsibility and little real control.
AI recommendation
arrives first · looks precise · usually correct
APPROVED
Reviews hundreds of cases under time pressure
When the decision fails: “A human approved it.”
Study 1: explanations and team accuracy
Explanations increased acceptance, but did not improve performance over simply showing confidence.
1,626 participants3 tasksAI matched to human accuracy
Compared with confidence only
No gain
No explanation condition produced significantly higher team accuracy.
When explanations were present
More reliance
Accuracy rose when AI was correct and fell when AI was wrong.
The explanations changed agreement more reliably than they changed the team’s ability to discriminate correct from incorrect advice.
Boundary: Crowdworkers completed beer-review, book-review, and LSAT-style tasks. The study tested local explanations under controlled accuracy matching, not every explanation type or domain.
Study 2: transparency and error detection
A model can be easier to understand and still harder to correct.
N = 3,800Four preregistered apartment-price experiments
Functionally identical predictions, varied transparency and number of features
SimulationClear two-feature models were easier for people to predict.
Useful relianceThat clarity did not make people follow the model more when doing so was beneficial.
Error correctionClear models sometimes made people less able to correct large mistakes on unusual apartments.
Targeted cueAn explicit outlier warning improved correction on both unusual cases, p < .001.
Boundary: Lay participants estimated apartment prices. Simulatability, reliance, and mistake detection are different outcomes and should be measured separately.
Confidence is not enough either
A number is useful only if it is calibrated, understandable, and relevant to this case.
87%
Calibrated? Across similar cases, does 87% mean roughly 87 out of 100?
In distribution? Is this case represented by the validation evidence?
Decision relevant? Does the remaining uncertainty change the action threshold?
Comparable? Does the human have a credible estimate of their own uncertainty?
Study 3: forcing people to think
Cognitive forcing reduced overreliance, but it did not solve team performance.