Special Topic · AI Product Psychology: Design the Feeling

Trust Calibration: The Best Users Are Half-Skeptical

Full trust pastes fabricated case law into court filings; no trust turns AI into a paperweight. Judge trust yourself across six scenarios; clickable citations, proxy confidence signals, and conditional warnings

THE QUESTION THIS PAGE ANSWERS

ANSWER FIRST

What is the key idea behind “Trust Calibration: The Best Users Are Half-Skeptical”?

Full trust pastes fabricated case law into court filings; no trust turns AI into a paperweight. Judge trust yourself across six scenarios; clickable citations, proxy confidence signals, and conditional warnings

DECISION RULE

Turn taste into a behavior the product can repeat. The useful outcome is not a nice opinion. It is a visible rule, a small example, and a way to tell when the experience falls below the bar.

TRY NEXT

Capture one before-and-after example that shows the quality bar without extra explanation.

WATCH FOR

Polish that improves the surface while leaving the user's uncertainty untouched.

Two ways to crash first
Full trust causes accidents; no trust wastes money. The foundational review of trust calibration (Lee & See, 2004) states the requirement clearly: trust must align with the system’s actual capability. Too much trust is misuse; too little is disuse—and each end of the spectrum has its own bill.
⚠ Overtrust: treating hallucination as truth

2023, New York—Mata v. Avianca: a practicing lawyer used ChatGPT to find case law and pasted 6 fabricated cases that don’t exist straight into court filings. After the judge checked each one, the lawyer was sanctioned and made global news. Similar accidents kept coming: AI output is fluent, confident, and well formatted—every surface cue nudges the judgment “this looks solid”, and those cues have nothing to do with whether the content is true.

⚠ Undertrust: AI becomes an expensive paperweight

The crash at the other end is quieter: a company buys AI tools, an employee hits an error once, and from then on every line of output gets sentence-by-sentence review. Review costs more than writing it yourself, so people stop using it. Procurement keeps paying; efficiency never rises. Undertrust doesn’t make the news—it only shows up in the internal postmortem titled “AI tool active usage: 8%.”

How reliable the system actually is → How reliable users think it is → Overtrust zone · misuse Undertrust zone · disuse Calibration line: trust = real capability
The x-axis is how reliable the system really is; the y-axis is how reliable users think it is. Product design’s job: pull users stuck in the two corners toward the diagonal. The three tools below all pull on that same line.
Six users, pulse-check one by one

Here’s the judgment mantra first: watch “risk × verifiability”, not “how powerful AI is.” The same full adoption is calibrated on a weekly report and overtrust in court. Six scenarios—you’re the calibrator.

Are these six users’ trust healthy? 0 / 6
Three options: overtrust, undertrust, well calibrated. Judge first, then read the explanation.
Make citations clickable

Calibration tool one: source citations—the kind you can open. Users who want to verify jump to the original in one click, which reins in overtrust; users too lazy to verify still see “there’s a source” and get reasonable confidence, which patches undertrust. The win-win requires citations that are real and clickable: a decorative fake citation exposed once is ten times worse than no citation. The AI answer below hangs three superscripts—open each one and compare with the source text.

Three citations—catch the one that doesn’t match 0 / 3 opened
Tap superscripts [1][2][3] to expand the source cards, and check whether the sentence in the answer matches what the source actually says.
Q: Fired verbally in the second month of probation—can I claim compensation?
Yes, you can claim. When an employer terminates a labor contract during probation, it must state the reasons to you[1]; if the termination is unlawful, the damages are twice the economic compensation standard[2]. Don’t wait: the limitation period for labor arbitration is one year, counted from the day you were dismissed[3]. Keep the dismissal notice, attendance records, and pay stubs as evidence.
One of the three citations doesn’t match the source—which?
Tune a confidence-signal set yourself

Tool two: confidence. Part One already covered it—the model doesn’t know what it doesn’t know, so self-reported certainty is unreliable. But the product layer has honest proxy confidence signals: how many docs retrieval hit, how relevant they are, whether sources agree, whether the knowledge cutoff covers the question. Below is the same medication answer; three switch groups map to three product decisions—flip them and watch how the answer on the right and the user’s trust calibration change.

Confidence-signal designer Flip the switches
Answer wording
Source citations
Warning bar
User trust calibration
22
Ibuprofen and acetaminophen—OK to alternate for a fever?
[1] Clinical guide to antipyretic analgesics [2] Pharmacopoeia · NSAIDs [3] Tertiary-hospital medication Q&A
Cry wolf five times in a row

Tool three does one job: that tiny line—“AI may err; please verify important information”—is anyone actually reading it? Psychology’s answer is banner blindness: Benway & Lane (1998) found with eye-tracking that elements constantly shown in a fixed spot get filtered out by the brain as background texture. Tap through five answers below and watch that line disappear with your own eyes.

Constant disclaimer vs. conditional warning Answer 0 / 5
First finish five answers in the “Constant disclaimer” group, then switch to “Conditional warning” to compare. Banner fade simulates how user attention dulls.
Attention to the warning
100
Sources and further reading: The foundational review of trust calibration is Lee & See (2004) Trust in Automation: Designing for Appropriate Reliance; misuse and disuse come from Parasuraman & Riley (1997); the lawyer-pasting-hallucinated-cases story is the real case Mata v. Avianca, Inc. (S.D.N.Y. 2023); banner blindness is from Benway & Lane’s (1998) eye-tracking work. How to repair trust after it collapses continues in Lesson 6 on algorithm aversion (Dietvorst et al., 2015); human gates for high-risk actions get a full design spectrum in Lesson 7 on defensiveness.

Turn the feeling in “Two ways to crash first” into a judgment

“2023, New York—Mata v.” points out that AI has lowered the bar for making something usable. The skill readers need is noticing what is wrong and turning that feeling into an actionable requirement.

Watch the user's next action, not just the surface

Turn “The crash at the other end is quieter: a company buys AI tools, an employee hits an error once, and from then on every line of output gets sentence-by-sentence review .” into observable questions: does the user know what happened, what to do next, and how to recover from an empty or failed state? Does the hierarchy make the important information visible first?

  • Set the trust target at calibration : full trust causes accidents, no trust wastes money—pull users toward the diagonal
  • Make citations clickable : swap “trust me” for “you can verify,” and watch for decorative fake citations biting back
  • Say confidence with proxy signals : retrieval hit count, relevance, source agreement—more honest than the model’s self-reported certainty

Pretty is not the same as usable

Apply “Tool three does one job: that tiny line—“AI may err;” to a second screen or flow. Record one moment of hesitation and the user action after the change; observable behavior is stronger evidence than polish alone.

From “Two ways to crash first” to “Six users, pulse-check one by one”

“Two ways to crash first” grounds the problem in “2023, New York—Mata v. Avianca: a practicing lawyer used ChatGPT to find case law and pasted 6 fabricated cases that don’t exist straight into court filings. After the judge checked each one, the lawyer was san…”. “Six users, pulse-check one by one” then moves it toward “Here’s the judgment mantra first: watch “risk × verifiability” , not “how powerful AI is.” The same full adoption is calibrated on a weekly report and overtrust in court. Six scenarios—you’re the calibrator”. Together, they show that the lesson is not just a conclusion to remember, but a claim with conditions.

Carry the judgment into the next situation

For experience work, turn abstract impressions into user actions: did the person understand the state, find the next step, recover from an error, and want to continue?

  • “Two ways to crash first”: 2023, New York—Mata v. Avianca: a practicing lawyer used ChatGPT to find case law and pasted 6 fabricated cases that don’t exist straight into court filings. After the judge checked each one, the lawyer was san…
  • “Six users, pulse-check one by one”: Here’s the judgment mantra first: watch “risk × verifiability” , not “how powerful AI is.” The same full adoption is calibrated on a weekly report and overtrust in court. Six scenarios—you’re the calibrator
  • “The closing point”: Show warnings conditionally : a constant disclaimer goes invisible in three days; a low-confidence popup with a reason is what gets read

The final “The closing point” brings the discussion to “Show warnings conditionally : a constant disclaimer goes invisible in three days; a low-confidence popup with a reason is what gets read”. The useful thing to carry forward is knowing which judgments must be revisited when input, scale, or risk changes.

✅ What this lesson wants to share

  • Set the trust target at calibration: full trust causes accidents, no trust wastes money—pull users toward the diagonal
  • Make citations clickable: swap “trust me” for “you can verify,” and watch for decorative fake citations biting back
  • Say confidence with proxy signals: retrieval hit count, relevance, source agreement—more honest than the model’s self-reported certainty
  • Show warnings conditionally: a constant disclaimer goes invisible in three days; a low-confidence popup with a reason is what gets read
Mark as learned Your reading progress updates automatically
← PreviousNext →

Keep reading

The next useful article in the thread.

ARTICLE DISCUSSION

Leave one useful thought here.

Keep the idea that clicked, the question that stayed open, or a small note for the next learner.

Discussing Trust Calibration: The Best Users Are Half-Skeptical AI Product Psychology: Design the Feeling
3discussionsArticle discussion · synced with the Circle
View in the learning circle
AM
Asha MorganContent editor
INSIGHTField note

I turned one judgment from this article into a small experiment I could run today. Knowing what to observe next is more useful than simply remembering the conclusion.

ARTICLE DISCUSSION7 helpful
LH
Lin HarperIndie developer
INSIGHTInsight

After reading this, I first looked for the conditions behind the idea instead of copying the method into a project. That order made the later trade-offs much clearer.

ARTICLE DISCUSSION5 helpful
KM
Kiki MooreProduct operations
QUESTIONQuestion

When this judgment reaches real work, which constraint should be added first? I am curious which step matters most between reading and the first practical attempt.

ARTICLE DISCUSSION4 helpful