User Vigilance Fails Twice in the AI Age: Watch Duty & Verdict Duty

Summary: Vigilance fails in two ways, and UX must design against both. Watching: detection of rare signals drops 10–15% within 30 minutes of continuous monitoring. Answering: when the computer works alone and interrupts occasionally, scrutiny decays per decision rather than per minute. Human-in-the-loop fails on both loops unless you design short watches and rare, meaningful questions.

Vigilance Comes in Two Forms
Vigilance used to mean staring. In 2026, it mostly means being interrupted. The staring version has 78 years of brutal experimental evidence behind it: people can’t watch for rare events much past the half-hour mark without missing a growing share of them. The interruption version is newer and sneakier. The computer does the work, you do something else, and every so often a dialog box, a notification, or an approval request summons you for a verdict. No staring required. And yet the verdicts rot too.
Definition: Watch vigilance is sustained attention to a mostly empty stream of events, hunting for a rare signal that can arrive at any moment. The human monitors continuously; the system does nothing to help.
Definition: On-call vigilance is intermittent judgment on demand: the system acts autonomously and interrupts the human with confirmation requests, notifications, and review checkpoints. Nobody stares at anything, but every interrupt demands a micro-decision.
Watch vigilance fails on a clock. Call it the 30-minute cliff. On-call vigilance fails on a counter: each routine, consequence-free approval makes the next one more automatic, until the single request that deserved a “no” sails through with the rest. Call that the rubber-stamp slide. Traditional psychology spent decades measuring the first failure in laboratories. The second failure is being measured right now, on all of us, in production.

The first half of this article covers the classic watch. The second half covers the on-call world of AI agents, approval dialogs, and comic books with defective speech balloons. (Yes, those are connected.)
Radar, U-Boats, and a Rigged Clock
During World War II, Royal Air Force radar operators hunting U-boats showed a disturbing pattern: their misses mounted as the watch wore on. The Medical Research Council’s Applied Psychology Unit in Cambridge put Norman Mackworth on the problem. He built what became known as the Mackworth Clock, a pointer stepping around a blank dial roughly once per second, occasionally making a double jump. Participants watched for 2 hours and reported every double jump. Detection fell 10–15% within the first 30 minutes and kept sliding for the rest of the session. His 1948 paper, “The Breakdown of Vigilance during Prolonged Visual Search”, founded an entire research field and supplied the name: vigilance is the state of readiness to detect rare signals, and the decrement is its measured decline.

Detection of rare events drops off a cliff after half an hour’s vigilance and keeps sliding from there.
Later research made the cliff steeper. Warren Teichner’s 1974 review in Human Factors found decrements appearing within the first 15 minutes for many tasks, and performance on demanding tasks decays faster still. But the cause isn’t boredom or laziness. Joel Warm, Raja Parasuraman, and Gerald Matthews showed in a 2008 Human Factors review that vigilance is hard mental work: monitors report high workload, show physiological stress, and drain attentional resources the way an old phone battery dies. Can’t we simply tell people to concentrate? Mackworth tested that too: a rest break or performance feedback restored detection, while urging people to try harder accomplished nothing. Remember that the next time an operations runbook says “remain alert.”
Where Modern UX Hits the Cliff
Mackworth’s radar screens have multiplied. Security screening, content moderation queues, fraud review, network operations dashboards, radiology worklists, supervising “self-driving” cars, and the newest arrival: humans reviewing AI output all day, one plausible-looking answer after another. Human-in-the-loop is this decade’s favorite safety story, and it runs head-first into a 1948 finding. A reviewer who approves agent actions for 3 hours straight stops being a safeguard and becomes an expensive rubber stamp.

The auto industry ran this experiment first, and the results should have been a warning label for the AI industry. Partial driving automation asks a human to do nothing for hours while remaining ready to act within seconds. That job description contradicts everything vigilance research has measured since the U-boats. The AI industry’s version swaps the highway for a chat transcript, and the physics are identical: the more reliable the automation, the rarer the signal, and the rarer the signal, the worse the human detects it. So improving the AI from 90% to 99% correct concentrates the oversight problem: each remaining error becomes 10 times harder for a human to catch.

Partial automation stages a 3-hour opera called Nothing Happens, then blames the audience for missing the flat note.
But most of today’s monitoring jobs no longer look like a watch at all. The computer acts and pings you when it wants something. That changes the failure mode. It doesn’t remove it.
Design for Short Watches
The research converts directly into design guidelines, which is why I love it:

Cap continuous monitoring at 30 minutes. Rotate tasks or people. This is the single cheapest intervention, straight from Mackworth. (When something has been true since World War II, it’ll be true for the duration of your project.)
Give detection feedback. Telling monitors when they hit and when they missed measurably props up performance; silent systems guarantee decay.
Inject test signals. Airport screening projects fictional threat images into scans of real bags to raise the event rate, keep watchers calibrated, and (crucially) measure the true detection rate. Every monitoring UI can copy this.
Keep alarms rare and precise. A flood of false alarms teaches the cry-wolf lesson at the skill level, and genuine signals drown.
Raise signal salience. The decrement bites hardest on faint, rare targets against noisy backgrounds, so anything that makes the target pop buys back detection.

Invert the roles: let automation watch normality and hand humans the 2% worth judging.
The Second Vigilance: On-Call, Not On-Watch
Thomas Sheridan at MIT named the underlying pattern supervisory control back in the 1970s, studying operators of undersea robots: the machine executes, and the human sets goals, monitors loosely, and handles exceptions. Every AI agent product in 2026 is a supervisory-control system with better marketing. And supervisory control looks like the perfect cure for Mackworth’s cliff. Nobody stares at a dial. The machine works; the human lives his or her life and answers the occasional question. What could go wrong?

Habituation, that’s what. The brain is ruthless about de-prioritizing repeated stimuli that never matter. Bonnie Brinton Anderson and co-authors watched it happen in an fMRI scanner: in their study “Tuning Out Security Warnings”, neural responses to security warnings dropped sharply after as few as 2 exposures, and kept falling across 5 days. The participants weren’t careless people; they had careless brains, which is to say, human ones. The same team showed that polymorphic warnings, which change their appearance on each showing, slow the slide: novelty buys back a little attention. Design lesson available on request from your own amygdala.
Healthcare ran the on-call experiment at scale and paid in lives. The Joint Commission’s Sentinel Event Alert 50 (2013), issued after alarm-related patient deaths, estimated that a staggering 85–99% of hospital alarm signals require no clinical intervention. The nurses stopped responding because the alarms taught them, thousands of times, that responding was pointless; callousness had nothing to do with it. Alarm fatigue is the rubber-stamp slide with a soundtrack.

85–99% of hospital alarms needed no action, so nurses learned to ignore them.
Consumer software teaches the same lesson politely. Windows Vista (2007) asked permission so often that “Cancel or Allow” became an Apple attack ad. Cookie consent banners ask billions of people a privacy question daily; the most-clicked button in Europe is whichever one makes the banner go away. (Cookie banners waste 575 million hours annually, just for EU citizens. Sadly, people in the rest of the world suffer from this insanity as well, so this one stupid regulation destroys more than a billion hours of human life every year.)

Are you sure? Are you really sure? Too many confirmation dialogs, and users stop behaving like this old-school radio listener. Instead, they change the station.
Attackers have weaponized the slide: in “push bombing,” a criminal with a stolen password simply fires approval prompts at your phone until, at 2 a.m., you tap Approve to make the buzzing stop. The US Cybersecurity and Infrastructure Security Agency’s countermeasure, number matching (PDF), forces you to read a number on the screen and type it into the phone. When the official federal remedy is “make the button harder to press,” the button was the vulnerability all along.
Rubber-stamping is rational: investigating an interrupt costs a minute; approving it costs half a second; and 19 of 20 interrupts are noise. The design failed the arithmetic and then blamed the user for doing the math. Every false or trivial alarm spends from the same budget that a real emergency will need to draw on, a budget I explored in my signal detection theory article on false alarms.
Users Who Wrote Their Own Permission Rules Protected Themselves Less
One study should end a comfortable assumption about user control of AI agents. In “Do User-Authored Permission Policies Improve Protection Against AI Agent Overreach?” (2026), Ting Yan had 113 US adults with no software background supervise an AI assistant through a simulated 18-action day in which it attempted 7 unrequested actions, from a $12 insurance purchase to reading private messages. One group approved or denied every action, one let a model review actions, and one wrote standing rules in advance (allow, ask, or never) for 4 consequence categories: spending, publishing, deleting, and reading private data.
The rule-writers lost. They blocked 20 percentage points less overreach than the everything-by-hand group (40% versus 60%). The mechanism is pure on-call vigilance: participants chose “ask me first” for 81% of their rules, so the rules settled almost nothing. The rule-writers then approved 67% of the prompts their own rules sent them, versus 40% among users judging every action fresh; the needless airport call went through for every participant who was asked. An “ask” rule is a deferral dressed up as a decision. Its later prompt arrives pre-blessed: my policy is on duty, so this instance is probably fine. Odysseus used wax and rope before the Sirens sang, not a standing request to be consulted song by song. Cruelest result: the rule-writers felt as much in control as everyone else, yet were protected the least.
Feeling in control is not a safety metric. Most users never change the setting you ship; the few who did write their own weren’t protected by them.

When users designed their own security protocols, they created far too many “ask me” questions; answer fatigue did the rest.
Proofreading AI: The Two-Tailed Balloon Problem
On-call vigilance isn’t only about permission dialogs. It also governs the quieter work of reviewing what AI produces, and here the trap is calibrated with almost malicious precision.
A concrete example from my own production pipeline. I generate multi-page comic strips with GPT Image 2, and the model is an excellent cartoonist: composition, character consistency, lettering, and comic timing are handled better than most human amateurs could manage. But roughly once every 4–5 pages, it draws a speech balloon with two tails, pointing at two different characters as if they’d spoken in unison. It’s a small defect with outsized damage: balloon tails are the grammar of comics, and a two-tailed balloon garbles who said what, which can wreck the joke or the meaning of the entire page.

A speech bubble with two tails: easy to overlook, since only about 2% of speech bubbles are currently drawn this way.
Do the vigilance math. A 2% defect rate (assuming 10 speech bubbles per page) sounds like a rounding error. But the errors arrive after runs of 3–4 flawless pages, the flaw sits in a peripheral element nobody reads consciously, and every other feature of the page is screaming competence at you. By page 8 of a clean streak, the reviewer’s trance has set in: you’re turning pages and enjoying your own comic book. Then page 9 ships with a balloon that makes the villain deliver the hero’s line, and a reader emails you about it. The reader always has fresh eyes. You had a criterion that slid one notch per clean page.
The same 1-in-N structure runs through all AI output review: the hallucinated reference lurking in a 40-source research report, the one inverted condition in 500 lines of generated code, the single wrong digit in a financial summary. Reviewing mostly-correct material is a vigilance task, and the better the AI gets, the more the task resembles the Mackworth Clock: long stretches of nothing, then one double jump. The better the AI, the worse the human checks it.
Four countermeasures work, and all four are borrowed from the watch-vigilance playbook:
Sweep for one flaw at a time. Don’t “review the pages”; check only balloon tails across all pages, then only lettering, then only hands. A single-target visual search resists decay far better than open-ended monitoring, because you know exactly what a signal looks like.
Seed known errors. Slip a deliberately defective page into your own review stack. If you don’t catch your plant, your hit rate on the real errors is a fantasy.
Use fresh eyes for the final pass. The reader who caught my two-tailed balloon spent 10 seconds on the page. Rotation is cheap; re-reading your own output for the 4th time is a comforting ritual that checks nothing.
Make the machine pre-screen. A second AI pass (“flag any balloon with an ambiguous or doubled tail”) converts 40 pages of watch duty into 2 flagged pages of judgment duty. Machines don’t habituate; humans shouldn’t be asked not to.
Watching AI Agents Work: The Short Run and the Long Haul
AI agents put both vigilances into one product, on a schedule.
Short runs (2–20 minutes) invite you to watch the agent stream its reasoning and tool calls in real time. This feels like oversight and mostly performs the feeling. The transcript scrolls, your attention drops off the 30-minute cliff early (my estimate: by minute 3, you’re skimming), and the one moment that needed steering, an agent cheerfully misreading your intent in step 2 of 30, slides by in the scroll. Streaming transcripts are the bowling-alley television of UX: mesmerizing, glanceable, and not actually being watched.
The fix must be structural, because motivation never survives minute 4. Surface decision points instead of token streams: show the plan before execution, checkpoint the moments where the agent commits to an interpretation, and render outcomes as reviewable diffs (“3 files changed, 1 payment scheduled”) rather than as 4,000 words of narration. Give the user 4 judgment moments instead of 12 minutes of theater.
Long runs (hours to days) drop all pretense of watching. You delegate, you leave, and oversight collapses to pure on-call vigilance: an approval request at 11:40, a checkpoint summary at 2:15, a completion notification at 5:03. Now the out-of-the-loop problem crashes the party. Mica Endsley and Esin Kiris showed in 1995 that people supervising automation lose situation awareness and are slow and error-prone when suddenly required to intervene. Translated to 2026: the approval dialog interrupts you mid-meeting with “Allow agent to modify deployment configuration?” and you face a choice between archaeology (reconstructing what the agent is doing and why, at a cost of 10 minutes you don’t have) and acquiescence (one click). The Ting Yan numbers tell you which one people pick.
Multiply this by a fleet. Running 5 agents in parallel is this year’s productivity brag, and each agent pages you on its own schedule. Attention doesn’t scale with your agent count; it’s the same fixed cognitive budget it was last year, now servicing 5 pagers. The endpoint of that arms race already ships as a feature: agentic coding tools offer a mode that skips all confirmations, popularly known as YOLO mode, and exhausted users flip it on within days. I don’t blame them. When a product’s permission system is so fatiguing that its own users disable it wholesale, the vendor has run the alarm-fatigue experiment on paying customers.

The design agenda for agent oversight follows from everything above, and it starts with treating user attention as the scarcest resource in the system:
First, ask rarely, and make each ask carry its context. An approval screen must reconnect the proposed action to the user’s original intent: what you asked for, what the agent wants to do, why that serves the request, what it changes, and whether it’s reversible. A bare “Allow?” is a coin flip with a UI.
Second, tier by consequence, not by tool. Ting Yan’s 4 plain-language categories (spend, publish, delete, access private data) are a solid starting vocabulary. Reversible and internal actions proceed silently with an undo window; irreversible, expensive, or externally visible actions earn a real interruption, and only those.
Third, instrument the rubber stamp. Log time-to-approve and approval rate. My rules of thumb: a median decision time under 2 seconds, or an approval rate above 95%, means your safeguard has become approval theater, and it’s time to remove the gate, re-tier it, or redesign what it shows. A blink approval, the sub-second tap on Allow, is a reflex with an audit trail.
The Irony of Automation
Now for the system-architecture blunder underneath it all. In her 1983 paper “Ironies of Automation” in Automatica, Lisanne Bainbridge nailed it: automation takes over the active work humans do well and leaves humans the monitoring of rare failures, the one job humans do worst. Partially automated driving is the textbook case: hands on the wheel, eyes forward, alert for hours against an event that almost never comes, in a cabin engineered for relaxation. That design fights human biology and loses.
Bainbridge’s other irony is the one the AI industry is about to relive at scale. Automation removes the work that keeps the skill sharp, then demands that same skill on the day the machine is wrong. A reviewer who has approved 400 fluent agent plans doesn’t merely miss the 401st. He or she has also stopped being the person who could have reconstructed why plan 401 is insane.
Manual control, diagnostic pattern recognition, and the feel for “this number can’t be right” all decay when they go unused. The first generation of agent supervisors still has those skills from the pre-agent job. The second generation will have grown up on transcripts and Allow buttons. The rare failure vigilance research says they’ll miss is also the one they’ll be least equipped to handle if they do notice it. Vigilance decrement makes the miss more likely. Deskilling makes the miss more expensive. Design that treats “keep a human in the loop” as a spare brain on a shelf is counting the same person twice: once as a detector they can’t be, and once as a rescuer they’ll no longer be.
We’re building the wrong architecture around AI agents in both vigilance flavors at once. The short-run transcript re-creates the radar watch; the long-run approval queue re-creates the hospital alarm ward. Both assign the human the role of last-line anomaly detector, a role 78 years of data says we can’t play, and then both count the human’s presence as a safety control in the compliance paperwork.
So redistribute the roles, since exhortation has a perfect record of failure. Let the automation watch and the human judge: machines are tireless anomaly detectors, so have them flag the 2% of cases that deserve real human reasoning instead of parading 100% past a fading watchstander. Where full review is genuinely required, measure vigilance with seeded errors and treat the hit rate as a first-class quality metric, not a legal fiction. And schedule breaks as safety equipment rather than perks.
16 Design Guidelines Against the Two Vigilance Decrements
Don’t hire humans as smoke detectors. Buy a smoke detector, and save the humans for judging what the smoke means. And when the smoke detector starts asking “Smoke OK? Allow / Deny” 40 times a day, expect people to approve the house fire. Here are 16 guidelines, 8 per vigilance.
For the watch (continuous monitoring):
Limit any continuous monitoring spell to 30 minutes. Rotation, task switching, or breaks must be built into the workflow, not left to individual stamina.
Provide immediate hit-and-miss feedback. Monitors who learn how they’re doing hold their detection rates far better than those watching into a void.
Seed test signals and publish the detection rate. If you can’t measure whether your human safeguard works, what you have is a hope (which is famously not a strategy).
Engineer alarm precision before alarm volume. Every false alarm withdraws trust from the account that a real emergency will need to draw on.
Invert the roles where possible. Have automation monitor continuously and escalate anomalies for human judgment, rather than making a person watch normality for hours.
Convert passive watching into active judging. Queues of flagged, consequential decisions sustain attention; scrolling walls of routine status do not.
Treat breaks as safety-critical equipment. Mackworth showed a pause restores detection; a monitoring schedule without pauses is a design defect.
Ban “the operator must remain alert” from your safety case. A plan that depends on defeating 78 years of vigilance research isn’t a plan.

Traditional safety thinking: exhort operators to be alert. Human factors thinking: redesign the system.
For the on-call loop (agents, alarms, and approvals):
Budget interruptions like money. Every ask spends from a finite attention account; ration to a handful of interrupts per session, and spend them where the expected cost of error is highest.
Tier by consequence, not by tool. Spending money, publishing, deleting, and reading private data earn interrupts; reversible internal actions proceed and log.
Prefer undo to ask. For reversible actions, act with a visible countdown and one-click cancel; reserve confirmation dialogs for what cannot be taken back, and add forcing functions (type the name of what you’re deleting) for the truly irreversible.
Make every approval carry its context. Show the original intent, the proposed action, why it serves the request, and what changes. An approval screen without the “why” forces a choice between archaeology and acquiescence, and acquiescence is faster.
Convert repeated approvals into narrower rules. After the 3rd identical “yes,” offer to automate exactly that case. Automation should absorb the routine so humans keep only the exceptions.
Preview the runtime consequences of user-authored rules. Show how often “ask me” would have fired on representative traffic before accepting it, and flag an ask-everywhere policy as the non-decision it is. Rule-writing feels like control; Ting Yan measured the difference.
Instrument the rubber stamp. Track median time-to-approve and approval rate per gate. Under 2 seconds or over 95% means the gate is decoration: remove it, re-tier it, or redesign it.
Vary the presentation of the alerts that truly matter. Habituation is partly visual, and polymorphic warnings slow it. Spend novelty only on the top tier, or you’ll habituate users to variety itself.

Approval is a muscle: grant it 40 times a day, and the movement becomes automatic, which defeats the point of asking at all.
The watchman goes blind on the wall at minute 31. The on-call doctor is saying “uh-huh, give two aspirin” by the 40th night. Short watches, seeded signals, and role inversion treat the first. A few rich, consequence-tiered questions, plus undo instead of interrogation, treat the second. Do neither, and your human-in-the-loop is a rubber stamp with a pulse.

We can’t change users: that’s a lost cause. The only alternative is to redesign the system to function well, given human nature. (All images in this article made with GPT Image 2)



