HCC 3030 · Week 3

Human-AI Interaction
& Decision Support

Overreliance, cognitive surrender, imperfect partners, and the design of better joint decisions.

Opening conflict

The AI says evacuate the east district. The emergency manager says west.

Emergency manager
WEST

A nursing home cannot move quickly, and one river sensor has been unreliable all morning.

WHO
LEADS?
Flood model
EAST

Rainfall, elevation, drainage, and traffic patterns indicate the largest expected loss there.

What would you need to know before choosing?

Decision support changes the decision

Advice is never “just information.”

RECOMMENDATION
changes attentionanchors expectationsredistributes effortshifts responsibility

Once a recommendation appears, the person is no longer solving the original problem.

Anatomy of decision support

A decision passes through both technical and human transformations.

01

World

Events, people, goals

02

Measurement

Data, sensors, records

03

Model

Score, forecast, option

04

Human judgment

Interpret, question, choose

05

Action

Consequences and feedback

Outcomes become future data, incentives, and expectations.
A recommendation has layers

Good support separates the answer from the evidence behind it.

Prediction
What does the model expect?
“Flood risk: 0.78”
Uncertainty
How stable is that estimate here?
Sensor coverage is incomplete
Rationale
What evidence influenced it?
Rainfall, elevation, road capacity
Action
What should happen next?
Evacuate, inspect, wait, or escalate
Authority is a design choice

Decision support can inform, recommend, constrain, or execute.

01

Inform

Show evidence or predictions.

Human initiates
02

Recommend

Propose a preferred option.

Human accepts or rejects
03

Constrain

Block, rank, or limit choices.

Human chooses within bounds
04

Execute

Act unless stopped or reversed.

System initiates
A person in the loop is not a safety mechanism

Oversight only works when the person can detect, understand, and correct failure.

SEE

Relevant evidence and uncertainty are visible.

JUDGE

The person has skill, time, and an independent basis.

ACT

They have authority and a usable correction path.

If any gate fails, “human oversight” may be ceremonial.
Humans are not the gold standard

Human judgment is adaptive, contextual, and often inconsistent.

HUMAN
Limited attentionWe cannot inspect everything.
Fatigue and workloadPerformance changes with conditions.
Heuristics and biasShortcuts can help or mislead.
Variable expertisePeople differ across cases and time.
Memory and consistencyWe forget and apply criteria unevenly.
AI is not the gold standard either

AI can be consistent and scalable while remaining systematically wrong.

AI
Partial dataThe model only receives encoded evidence.
Objective mismatchThe optimized score may be a poor proxy.
Distribution shiftDeployment differs from training.
Hidden uncertaintyConfident outputs can still be wrong.
Missing values and contextPrediction cannot settle what should matter.
The opportunity

Neither partner is perfect, but they may be wrong on different cases.

Human errorsfatigue · rare patterns · inconsistency
AI errorsshift · missing context · proxy failure
shared
blind spots

Collaboration has potential only when one partner can recognize or repair the other’s error.

The evidence is a warning

Human plus AI does not automatically beat the better solo performer.

106
experiments

370 effect sizes · preregistered meta-analysis

g = −0.23Team versus the better solo performer
95% CI [−0.39, −0.07]
g = +0.64Team versus the human alone
95% CI [+0.53, +0.74]
I² = 97.7%Very high heterogeneity across settings
On average, AI helped people, but the combined workflow still lost value compared with using the stronger partner alone.
The moderators matter

The average changed sharply with the task and with who was stronger alone.

Decision tasks
g = −0.27

Teams underperformed the better solo partner.

95% CI [−0.44, −0.10]
Creation tasks
g = +0.19

Direction was positive, but not statistically different from zero.

95% CI [−0.09, +0.48]
Human stronger alone
g = +0.46

The team beat both solo partners on average.

95% CI [+0.28, +0.66]
AI stronger alone
g = −0.54

Adding the human reduced available AI performance.

95% CI [−0.71, −0.37]
Boundary: About 85% of effect sizes involved decision tasks, and more than 95% used a human-final-decision workflow. These are moderators, not universal laws.
Define the target clearly

Complementary performance means the team beats both partners alone.

Human alone
AI alone
Human + AI
TEAM > MAX(HUMAN, AI)
Clinical evidence: benefit and danger

The same clinicians gained from good AI and lost accuracy when its top answer was corrupted.

Skin-lesion classification155 raterswithin-person comparison
Good AI probabilities
+9.5%

Median accuracy gain after support

Faulty top prediction
−6.3%

Median accuracy change after support

In a separate 600-image aggregation, human collectives plus AI reached 81.0% accuracy, versus 73.7% for human collectives and 76.9% for the CNN alone.
Boundary: This was an image-based experimental benchmark, not evidence about patient outcomes or deployment in a clinic.
Two sources of complementarity

Partners add value when they possess different information or different capabilities.

INFORMATION

They see different evidence.

The model sees thousands of historical patterns. The person sees a new constraint, local context, or unrecorded goal.

Can each partner reveal what the other lacks?
+
CAPABILITY

They process evidence differently.

The model searches consistently at scale. The person reasons about values, exceptions, causal stories, and consequences.

Can the workflow allocate work to the stronger contributor?
Build an error map

The useful question is not “Who is better?” but “Who is better here?”

Case conditionsHuman advantageAI advantage
Novel local change
Large repetitive search
Value conflict
Subtle historical pattern
Shared missing data××
Averages hide the cases that should be routed, escalated, or deferred.
Appropriate reliance

A good decision-maker accepts correct advice and rejects incorrect advice.

AI correct
AI wrong
Human accepts
Appropriate relianceThe advice repairs or confirms the decision.
OverrelianceThe person follows the system into error.
Human rejects
UnderrelianceUseful advice is ignored.
Appropriate self-relianceThe person catches the model’s error.
Trust and reliance are related, not identical

Trust is an attitude. Reliance is an action.

TRUST

“I expect this system to help me under uncertainty.”

survey · belief · expectation
RELIANCE

“I used, accepted, deferred to, or acted on its output.”

choice · behavior · delegation
People may report skepticism and still follow a default under time pressure.
Reliance should be local

“I trust the AI” is too broad to guide a decision.

Trust it...
for what task?classification, prediction, ranking, generation, action
on which cases?routine, rare, ambiguous, shifted, adversarial
using what evidence?validation, calibration, provenance, recent monitoring
with what stakes?reversible suggestion or consequential action
Automation bias

People often treat an automated cue as a substitute for searching and checking.

AI recommendation
attention narrows
contradictory evidence receives less search
the recommendation becomes the default frame
error travels
through the human
Two signatures of automation bias

Automation can make us miss what it omits and do what it wrongly suggests.

OMISSION

The system stays silent.

The person fails to notice or act because no alert appeared.

No warning → hazard goes unexamined
COMMISSION

The system recommends the wrong action.

The person follows the recommendation despite conflicting evidence.

Wrong warning → unnecessary action
Evidence across domains

Automation bias is common in the literature, but the evidence base has real limitations.

40
studies met review criteria

Human factors and health-care research, 1983–2015

81%25 of 31 studies testing omission errors reported evidence of bias
91%21 of 23 studies testing commission errors reported evidence of bias
9studies tested statistical significance against a manual control
Interpret carefully: Samples were often small and homogeneous, measures varied widely, and most studies did not report an effect size against an unassisted control.
Why overreliance feels reasonable

The problem is rarely “people are lazy.” The system changes the economics of thinking.

FluencyClear language feels easier to accept.
Authority cuesScores and precision look objective.
Time pressureVerification competes with throughput.
Low self-confidencePeople defer when they doubt themselves.
High base accuracyRare failures train complacency.
Default designAccepting is easier than objecting.
Overreliance is produced by people, interfaces, organizations, and incentives together.
A crucial distinction

Cognitive offloading can help. Cognitive surrender gives up the judgment itself.

OFFLOADING

Move part of the cognitive work into an external tool.

Use a calculator, map, checklist, search tool, or AI draft.
The person still frames, evaluates, and owns the decision.
VS
SURRENDER

Adopt the external answer with little independent scrutiny.

The system supplies the frame, conclusion, and stopping point.
The person cannot explain, test, or recover the reasoning.
Cognitive surrender

The danger is not that the AI thinks. It is that the person stops deciding how to think.

An emerging term for adopting AI output with minimal scrutiny while intuition, deliberation, or verification is bypassed.

Useful as a diagnostic idea. Not a medical diagnosis and not every instance of AI assistance.
Separate the evidence from the label

The mechanisms are better established than the term “cognitive surrender.”

ESTABLISHED

Cognitive offloading

People routinely move memory and computation into external tools. Benefits and costs depend on the task and what remains internal.

Review literature · Risko & Gilbert, 2016
GROWING

Passive AI reliance

Experiments increasingly connect passive use with overreliance, weaker error detection, and lower self-efficacy or ownership.

Multiple behavioral studies · 2021–2026
EMERGING

Cognitive surrender

A proposed umbrella theory for external AI reasoning displacing intuitive and deliberative judgment.

New theory preprint · Shaw & Nave, 2026
Teach surrender as a diagnostic hypothesis, not as a settled syndrome or a validated measurement scale.
How surrender develops

A fluent answer can create closure before understanding has formed.

01

Need

A difficult or uncertain task

02

Fluent answer

Fast, complete, confident output

03

Closure

The problem feels resolved

04

Skipped work

No independent frame or search

05

Adoption

Output becomes judgment and action

Five warning signs

You may have surrendered when the answer survives but your understanding disappears.

1

You cannot explain why the recommendation fits this case.

2

You did not form a view before seeing the output.

3

You searched for confirmation, not contradiction.

4

You cannot continue if the system becomes unavailable.

5

You still own the consequences but no longer feel ownership of the reasoning.

The cost unfolds over time

Overreliance can improve today’s throughput while weakening tomorrow’s judgment.

NOW

Immediate performance

Faster output

Lower effort

More consistency

Possible accuracy gain

LATER

Capability and agency

Reduced practice

Weaker situation awareness

Lower self-efficacy

Harder recovery when AI fails

Evidence on passive versus active use

How people used AI mattered more than whether AI was present.

N = 269

Preregistered experiment

Occupation-specific writing tasks
No AIWrite independently
Passive AICopy AI-generated content
Active collaborationDraft first, then use AI to refine
Passive useLower self-efficacy, psychological ownership, and meaningfulness. Self-efficacy and meaningfulness remained lower on a later manual task.
Active usePsychological outcomes were statistically similar to the no-AI condition.
Boundary: This is one writing experiment plus a correlational follow-up survey of 270 workers. It does not establish long-term skill loss across occupations.
The irony of automation

The more routine work automation absorbs, the less prepared people may be for the rare failure left behind.

99%

Normal operation

The system handles familiar cases. Human skill receives little practice.

failure
1%

Abnormal condition

The human must diagnose quickly with incomplete awareness and decayed skill.

The oversight trap

A human can become a rubber stamp with full responsibility and little real control.

AI recommendation
arrives first · looks precise · usually correct
APPROVED

Reviews hundreds of cases under time pressure

When the decision fails:
“A human approved it.”
Study 1: explanations and team accuracy

Explanations increased acceptance, but did not improve performance over simply showing confidence.

1,626 participants3 tasksAI matched to human accuracy
Compared with confidence only
No gain

No explanation condition produced significantly higher team accuracy.

When explanations were present
More reliance

Accuracy rose when AI was correct and fell when AI was wrong.

The explanations changed agreement more reliably than they changed the team’s ability to discriminate correct from incorrect advice.
Boundary: Crowdworkers completed beer-review, book-review, and LSAT-style tasks. The study tested local explanations under controlled accuracy matching, not every explanation type or domain.
Study 2: transparency and error detection

A model can be easier to understand and still harder to correct.

N = 3,800Four preregistered apartment-price experiments

Functionally identical predictions, varied transparency and number of features

SimulationClear two-feature models were easier for people to predict.
Useful relianceThat clarity did not make people follow the model more when doing so was beneficial.
Error correctionClear models sometimes made people less able to correct large mistakes on unusual apartments.
Targeted cueAn explicit outlier warning improved correction on both unusual cases, p < .001.
Boundary: Lay participants estimated apartment prices. Simulatability, reliance, and mistake detection are different outcomes and should be measured separately.
Confidence is not enough either

A number is useful only if it is calibrated, understandable, and relevant to this case.

87%

Calibrated? Across similar cases, does 87% mean roughly 87 out of 100?

In distribution? Is this case represented by the validation evidence?

Decision relevant? Does the remaining uncertainty change the action threshold?

Comparable? Does the human have a credible estimate of their own uncertainty?

Study 3: forcing people to think

Cognitive forcing reduced overreliance, but it did not solve team performance.

N = 199

Meal-modification task

3 forcing designs · 2 simple XAI designs · no-AI baseline
When AI was wrong

Participants rejected bad advice and chose the optimal answer significantly more often under cognitive forcing.

Overall team result

No significant performance difference from simple XAI. Human-AI teams still performed below the 75%-accurate AI alone.

Human response

The conditions that reduced overreliance were rated harder and were preferred and trusted less.

Boundary: The AI was simulated, participants were crowdworkers, and benefits were concentrated among people higher in Need for Cognition.
Cognitive forcing functions

Small amounts of well-placed friction can restore independent judgment.

01

Commit first

Record an initial judgment before showing AI advice.

02

Ask for evidence

Require the user to identify supporting and conflicting cues.

03

Request assistance

Make AI available on demand rather than always present.

04

Compare alternatives

Show competing options, not one polished answer.

05

Delay action

Create a pause before consequential execution.

Friction must be selective

The goal is thoughtful effort where error matters, not a slower interface everywhere.

Too little

Automatic acceptance

fast · brittle
Useful friction

Independent view, comparison, verification

deliberate · recoverable
Too much

Alert fatigue and workaround behavior

slow · ignored
Match friction to uncertainty, stakes, reversibility, and user expertise.
Design for disagreement

A useful partner makes it easier to challenge the recommendation than to merely accept it.

Independent first passProtect the human’s frame before anchoring.
Visible counterevidenceSurface facts that weaken the recommendation.
Alternative hypothesesShow more than one plausible interpretation.
Cheap correctionMake reject, edit, undo, and escalate usable.
Learn from disagreementCapture why the person overrode the system.
Plan for inevitable failure

Error recovery is part of the interaction, not an exception after it.

1

Detect

Reveal anomaly, conflict, or uncertainty.

2

Diagnose

Show state, provenance, and recent changes.

3

Correct

Edit, override, undo, or choose a safer mode.

4

Learn

Capture feedback without hiding the failure.

Maintain a safe fallback and a clear path for human re-entry.
Allocate decisions dynamically

The strongest partner should lead the cases where its advantage is real.

NEW CASE

Human leads

Novel context, values conflict, missing data, high need for explanation

Joint review

Disagreement, high stakes, uncertain evidence, complementary information

AI leads

Routine pattern, validated conditions, high volume, reversible action

Routing rule: case conditions + relative capability + stakes + cost of delay
Three collaboration patterns

The best division of labor depends on where each partner adds information.

AI proposesHuman critiques

Generate and evaluate

Useful when AI expands the option space and humans can judge quality, fit, or values.

Human decidesAI challenges

Decision and red team

Useful when independent human framing matters and AI can search for missed evidence.

Route by caseEscalate uncertainty

Selective delegation

Useful when performance changes predictably across case types.

Four conditions for joint performance

Complementarity appears only when differences survive the interaction.

01

Diversity

Human and AI possess genuinely different information or capabilities.

02

Visibility

The workflow reveals confidence, evidence, context, and disagreement.

03

Allocation

Cases and subtasks reach the contributor with an advantage.

04

Integration

The final process resolves conflict without erasing useful information.

DIVERSITY × VISIBILITY × ALLOCATION × INTEGRATION
Measure the joint system

Accuracy is necessary, but it does not tell us whether the partnership is healthy.

Decision qualityAccuracy, utility, equity, and outcome quality
Reliance calibrationAccept correct advice and reject incorrect advice
RecoveryDetection time, reversibility, safe fallback
Human capabilityUnderstanding, skill retention, self-efficacy
AgencyMeaningful control, contestability, ownership
Operational costTime, workload, latency, escalation burden
Return to the flood decision

The disagreement is information. Use it before choosing a side.

What the model contributes

Large-scale rainfall and traffic patterns

Consistent comparison across districts

Estimated loss under known assumptions

DISAGREEMENT

Inspect the sensor.
Test both assumptions.
Compare consequences.
Escalate if uncertainty remains.

What the manager contributes

Recent sensor failure not yet recorded

Nursing-home evacuation constraints

Local values, capacity, and accountability

Socratic test

A team agrees with its AI 95% of the time. Is that evidence of success?

1How often is the AI wrong, and on which cases?

2When it is wrong, how often do people catch it?

3Do users form an independent view before seeing advice?

4Can they explain and reverse the final decision?

5Does disagreement improve the system or get punished?

Agreement is ambiguous. It can indicate shared competence, shared bias, or cognitive surrender.
A six-question diagnostic

Design the partnership around disagreement, not obedience.

Decision: What is being predicted, recommended, or executed?

Difference: What unique information or capability does each partner hold?

Reliance: Can the person accept correct advice and reject incorrect advice?

Agency: Are they still framing, evaluating, and owning the judgment?

Recovery: Can the team detect, correct, undo, and learn from failure?

Evidence: Does the joint system outperform both partners on the outcomes that matter?

H
+
AI
BETTER
TOGETHER?