At The Threshold: Trust Calibration in Human

This essay is part of the At The Threshold series. Read the overview and full reading order on the At The Threshold page.

Trust Calibration in Humans–AI Decision Making
Image created using Canva AI by the author, visualising the idea of “trust calibration” — how humans learn to match their trust in AI systems to the systems’ actual reliability, avoiding both over‑reliance and unwarranted scepticism.

Automation has always required human beings to decide how much to trust it. From the autopilot systems of commercial aviation to the credit-scoring algorithms of financial institutions, the central challenge has not been whether automated systems make errors—they do—but whether the people who work alongside them calibrate their reliance appropriately in response. The emergence of large language models and AI-assisted decision tools in the 2020s has made this challenge newly urgent, because these systems fail in ways that are unfamiliar, difficult to detect, and poorly served by the intuitions human beings normally bring to questions of trust. This article examines the psychology of human-automation trust, the characteristic failure modes of contemporary AI systems, and what appropriate reliance might require.

Introduction

The academic study of human-automation trust has a substantial history, predating the current generation of AI systems by several decades. Lee and See (2004), in a foundational review published in *Human Factors*, defined trust in automation as “the attitude that an agent will help achieve an individual’s goals in a situation characterised by uncertainty and vulnerability” (p. 54). Their analysis identified a persistent and consequential asymmetry in how humans relate to automated systems: people tend either to overtrust—deferring to system outputs even when their own assessment would be more accurate—or to undertrust—rejecting accurate system recommendations following observed failures. Calibrated trust, which adjusts reliance dynamically to match a system’s actual competence across specific conditions, is both theoretically desirable and empirically uncommon.

This framework remains highly relevant to AI-assisted decision making. In many contemporary settings, the question is not whether AI should be used at all, but under what conditions its outputs should be accepted, checked, challenged, or ignored. Appropriate reliance therefore depends not only on system accuracy, but also on the user’s ability to recognise the boundaries of that accuracy in real time.

The Mechanics of Miscalibration

A meta-analysis of factors affecting trust in human-robot and human-automation interaction, conducted by Hancock et al. (2011), synthesised findings across 29 published studies and identified performance-based variables as the strongest predictors of trust—above individual-difference factors and design characteristics. Systems that perform reliably across a range of conditions generate higher trust; systems that fail, even once, generate trust deficits that persist and prove difficult to recover. This asymmetry between trust acquisition and trust repair has significant implications for AI systems that fail in visible, public, and often memorable ways.

The specific failure modes of large language models produce a particular challenge for trust calibration. Unlike mechanical systems, which typically fail in ways that are apparent—a machine stops, an error message appears—LLMs often fail confidently. They produce fluent, syntactically well-formed, affirmatively stated incorrect information with no reliable internal signal distinguishing accurate from inaccurate output. Parasuraman and Riley (1997) identified automation complacency—the tendency to reduce monitoring of automated systems following a period of reliable performance—as a characteristic risk in human-automation interaction. With systems that provide no dependable uncertainty signal, the conditions for complacency are structurally embedded in the interaction design.

A further complication is that many AI systems are used in environments where users are under time pressure. When speed is prioritised, people are more likely to treat plausible outputs as sufficient and to skip verification steps they might otherwise perform. This is especially risky in domains such as healthcare, education, law, customer support, and knowledge work, where an answer may sound coherent while still being incomplete, outdated, or false.

Algorithm Aversion and Its Limits

Research on algorithm aversion—the tendency to reject algorithmic recommendations more strongly than equivalent human recommendations following observed errors—complicates the picture further. Dietvorst, Simmons, and Massey (2015) demonstrated experimentally that participants who observed an algorithm making errors were significantly less likely to use it in subsequent rounds, even when the algorithm outperformed human judgment overall and the errors were no worse than human errors of comparable magnitude. The asymmetry is, in a strict sense, irrational—but it reflects a coherent social expectation: human advisors are expected to improve and explain themselves following errors; algorithms are perceived as unchanging, which makes their failures feel less forgivable.

At the same time, algorithm aversion does not always dominate behaviour. In practice, people may oscillate between scepticism and overreliance depending on task framing, interface design, institutional pressure, and whether responsibility for mistakes is individual or distributed. A person may dismiss one AI tool after a visible blunder yet continue to rely heavily on another because it is faster, more convenient, or socially normalised within their workplace.

Daniel Kahneman’s distinction between System 1 and System 2 processing provides a useful framework for understanding why these dynamics are stable (Kahneman, 2011). Trust assessments of the kind people form about familiar human advisors draw on accumulated experiential signals—tone, history, social context—that support rapid, intuitive calibration. Trust assessments of AI systems lack such experiential substrate. The signals are unfamiliar, the failure modes are novel, and the intuitions that govern trust in human relationships provide unreliable guidance.

Towards Appropriate Reliance

Lee and See (2004) argued that the design goal for human-automation interaction should be appropriate reliance—trust that is neither excessively high nor excessively low, but accurately matched to a system’s actual reliability across specific task types and conditions. This requires not only that systems perform reliably but that they communicate their confidence accurately, acknowledge the boundaries of their competence explicitly, and provide the information necessary for users to adjust their reliance in response to genuine performance variation.

Current large language models meet few of these requirements consistently. They do not reliably signal uncertainty. They do not consistently distinguish between claims they are likely to be correct about and claims where their training data is sparse, contested, or absent. The burden of calibration therefore falls largely on the user—who is often not in a position to evaluate the system’s outputs against independent ground truth.

For that reason, trust calibration should be understood not only as a user problem, but as a design and governance problem. Interfaces can encourage better calibration by making sources visible, surfacing uncertainty, slowing high-risk interactions, and prompting verification where consequences are significant. Organisations can support better reliance by training staff in AI failure modes, establishing review thresholds, and ensuring that human oversight is meaningful rather than symbolic.

Conclusion

The core problem in human–AI interaction is not simply whether AI systems are accurate enough to be useful. It is whether people can learn when to rely on them, when to verify them, and when to override them. Calibrated trust is difficult because contemporary AI systems are persuasive, uneven, and often opaque: they speak with confidence, perform brilliantly on some tasks and poorly on others, and rarely expose the limits of their own competence. In that environment, it is easy for reliance to drift toward fluency, convenience, or organisational habit rather than toward an informed assessment of reliability. If these systems are to support rather than erode human judgment, designers and institutions must build interfaces, training, and governance that help users match trust to truth — making it as easy to pause, question, or correct an AI recommendation as it is to accept it.

Reference List

Dietvorst, B. J., Simmons, J. P., & Massey, C. (2015). *Algorithm aversion: People erroneously avoid algorithms after seeing them err.* *Journal of Experimental Psychology: General, 144*(1), 114–126. https://doi.org/10.1037/xge0000033

Hancock, P. A., Billings, D. R., Schaefer, K. E., Chen, J. Y. C., de Visser, E. J., & Parasuraman, R. (2011). *A meta-analysis of factors affecting trust in human-robot interaction.* *Human Factors, 53*(5), 517–527. https://doi.org/10.1177/0018720811417254

Kahneman, D. (2011). *Thinking, Fast and Slow.* Farrar, Straus and Giroux.

Lee, J. D., & See, K. A. (2004). *Trust in automation: Designing for appropriate reliance.* *Human Factors, 46*(1), 50–80. https://doi.org/10.1518/hfes.46.1.50.30392

Parasuraman, R., & Riley, V. (1997). *Humans and automation: Use, misuse, disuse, abuse.* *Human Factors, 39*(2), 230–253. https://doi.org/10.1518/001872097778543886


Author Note (AI Usage)

This article was drafted with assistance from a generative AI system to organise structure and suggest phrasing. All facts, citations, and final editing have been reviewed and approved by the author. The AI did not access any private health data.

Continue in this series: At The Threshold: When the Machine Enters the Room. Or return to the At The Threshold overview.

Comments