top of page

UX Roundup: Local-Language AI Use | AI & Jobs | Great AI Films | Hypertext Hero Ted Nelson | User Errors | AI Helps Rainforests | Annotation UI | Token Spend

  • Writer: Jakob Nielsen
    Jakob Nielsen
  • 7 minutes ago
  • 21 min read
Summary: AI use becoming more localized, according to Gemini data from Southeast Asia | Heavy AI use in a company increases hiring, not layoffs | Winners of the AI Film Festival awards | The inventor of hypertext, Ted Nelson | Helping users recover from errors | AI drones help replant the rainforests | Annotations seem a simple feature, but are often designed wrong | Token use is going through the roof with agentic AI

 

UX Roundup for July 20, 2026 (GPT Image 2)

 

AI Goes Native in Southeast Asia: 70% Local Languages, 42% No Keyboard

Google released The Gemini Report: Southeast Asia on July 14, 2026, analyzing aggregated usage data from millions of Gemini conversations across 6 countries: Indonesia, Malaysia, the Philippines, Singapore, Thailand, and Vietnam. Active users in the region more than doubled in a year. Two findings deserve attention far beyond Southeast Asia.

 

AI Speaks the User’s Language

Nearly 70% of prompts in Southeast Asia are submitted in local languages, not English. Vietnam leads at a whopping 89% Vietnamese, followed by Thailand (87% Thai) and Indonesia (84% Bahasa Indonesia).

 

For the first 40 years of personal computing, the machine spoke English and everybody else adapted. Chinese and Japanese character sets weren’t supported at all, and I remember the contortions needed to use even the modest requirements of the Danish language, with its three extra characters Æ, Ø, and Å. (Luckily, I can now use Danish characters easily thanks to Unicode, though my ancestors’ runes still render correctly only in special fonts like Segoe UI Historic.)


Odin and his ravens Huginn and Muninn, made with Reve 2.1. While the god’s outfit looks Viking enough to be historically plausible, the three names are not rendered with correct runes, whether Elder Futhark, as I requested, or the Younger Futhark you sometimes see.

 

Menus, commands, and error messages shipped in English first, with localization as a grudging afterthought that arrived a year late and read like it. AI reverses the polarity: when the interface is natural language, “natural” means the language the user thinks in, and a student in Hanoi has no reason to compose his or her homework questions in English. (AI Singapore’s SEA-HELM benchmark rated Gemini the best large language model for Southeast Asian languages, which surely helped Google’s numbers. Vendors that treat Thai or Vietnamese as second-class languages will forfeit these markets.)

 

The Keyboard Becomes Optional

42% of prompts use voice, camera, or image input instead of typing. Users snap a photo of a homework problem, a plant, or a menu, or they simply talk. Seniors in Thailand show the highest multimodal adoption of any age group. That makes sense: people who never became fluent typists skip the keyboard entirely, just as emerging markets skipped landlines and went straight to mobile phones.

 

I predict this 42% will grow. Text was never the natural human interface; it was a workaround for machines that couldn’t hear or see. Now they can. OpenAI launched GPT-Live on July 8: full-duplex voice models that listen and speak simultaneously, murmur “mhmm” while you talk, and wait politely when you pause to think. Full-duplex kills the walkie-talkie rhythm that made older voice modes feel like leaving messages on an answering machine. As speech turns into genuine conversation and the camera becomes the AI’s eyes, keyboard-optional computing will spread from Southeast Asia to your users too.

 

The obligatory caveat: this is vendor research about the vendor’s own product, so discount the self-congratulation. But the behavioral data covers millions of real users doing real tasks, which beats any survey of what people claim they do.

 

UX Implications

  • Stop designing English-first. Test your AI features with native-language prompts, including the code-switching common in some countries (Taglish blends Tagalog and English mid-sentence). If quality drops outside English, most of your Asian users get the degraded product.

  • Promote voice and camera to first-class inputs. They aren’t accessibility extras anymore. Give them prominent entry points on the main screen, and make sure results render sensibly when the query was a photo rather than a sentence.

  • Design for conversation, not dictation. Full-duplex speech makes interruptions, backchannels, and turn-taking part of your design material. Users will cut the AI off mid-answer. Plan for it.

  • Test with the keyboard turned off. Run your next usability study with typing banned. If the design collapses, so does its appeal to 42% of this market.


So the region often described as mobile-first is now previewing the AI-first interface: spoken, visual, and multilingual. The keyboard had a good 150-year run. It can start planning its retirement party.



Deep AI Adoption = 10% More Jobs

Companies that spend serious money on AI hire more people, not fewer. That’s the headline from A New Look at AI’s Impact on Jobs by Ara Kharazian and Ryan Stevens of the payments platform Ramp, together with Lisa Simon of the workforce-analytics firm Revelio Labs. It’s the first study to connect firms’ real AI invoices with their headcount records at scale, covering 21,559 U.S. firms. The heaviest AI spenders grew total employment by 10% in the 24 months after adoption, and their entry-level headcount rose 12%. Light spenders got nothing.

 

AI Intensity and Employment Change over 24 Months Outcome Light Adopters ($2.78/employee/month) Heavy Adopters ($33.67/employee/month) Total headcount No change +10% Entry-level headcount No change +12%

The data is the star of this paper. Ramp processes corporate card and bill payments, so the authors could see exactly which firms pay AI vendors, how much, and when, with no surveys, no occupational “exposure” indices, and no executives flattering themselves about their company’s AI sophistication. Revelio Labs tracks monthly employment from millions of public career profiles. A firm counted as an adopter once it paid AI vendors at least $100 per month for 3 straight months, a threshold that filters out one-off employee experiments. The authors then split adopters by early spending intensity: the top third averaged $33.67 per employee per month, while the rest averaged $2.78. That’s 12 times the money, and the multiple turns out to be everything.

 

“AI use” is not a binary choice. Companies that just use a little AI sprinkled here and there don’t hire new staff. Companies with intense AI use are the ones that grow. (GPT Image 2)

 

The gains follow a learning curve. Heavy adopters showed zero extra growth in the adoption month, about 7% by month 6, and roughly 32% by month 18. So the technology didn’t pay off on day one: firms apparently need 6–12 months to find use cases and rewire workflows before AI funds expansion. And the hiring was broad: sales up 10%, administration up 8%, engineering up 7%, customer service up 6%. By sector, though, the party remains exclusive: Information firms (software, internet, media) grew about 13%, the only sector group with a statistically significant gain so far, which fits the head start that coding agents enjoy over every other commercial AI use case.


Staff growth from heavy AI use doesn’t happen the day the company turns on the AI spigot. Things take time, and the strongest employment effect takes more than a year, happening at the 18-month mark. (Muse Image)

 

Doesn’t this contradict the Stanford “canaries in the coal mine” study, which found employment down roughly 16% for workers aged 22–25 in the most AI-exposed occupations? No. The Stanford team measured occupations across the entire economy, whereas this study measures individual firms that adopted. Both results can hold at once: deep adopters hire juniors while grabbing market share from laggards who shed them. What happens to national employment is exactly what a firm-level study can’t tell us.

 

Now for the salt. This is vendor research: The sample skews toward tech-forward, mid-market firms that chose a fintech card platform. Adoption is also self-selected: the firms that eventually bought AI were already growing 6% a year before they spent a dime, versus 1.6% for never-adopters, so part of what we’re watching is winners buying tools rather than tools creating winners. The authors handle the problem better than most by comparing early adopters only with similar firms that adopted later, which aligns the pre-adoption growth paths. Still, payment records beat surveys, where executives claim AI adoption that never shows up on an invoice.

 

Pick Employers That Do the AI Reps

Most firms practice gym-membership AI: they pay the subscription and skip the workouts. (The gym doesn’t mind; it already has your money.) The employment gains in this study went to the firms doing the reps: coding agents, API budgets, multiple vendors, roughly $34 per employee per month instead of $3. Three consequences for your career planning:

 

  • Screen employers by AI intensity. In your next job interview, ask what AI tooling the team pays for beyond chat subscriptions. A concrete answer about agents and APIs signals a firm on the 10%-growth track; “we all have ChatGPT seats” signals a $3 shop.

  • Junior talent should target heavy adopters. Entry-level headcount rose 12% at these firms, the opposite of the doom narrative. If you’re early in your career, the companies using AI the most are the safer bet.

  • Distrust AI-blamed layoffs. When a CEO attributes job cuts to AI, this data suggests you check whether the real story is slowing sales. The authors put it plainly: be skeptical.


Productivity gains create new work; they don’t exhaust a fixed pile of it. That’s the fixed-work fallacy I skewered back in 2023, and the payment records of 21,559 firms now back me up, at least for the early adopters. Don’t run gym-membership AI in your own career, either. Do those reps.


Just as with fitness, repetition is key to AI progress. Keep working out, and you’ll see the gains. (GPT Image 2)

 

Award-Winning AI Short Films

The winners of the 2026 Runway AI Film Festival awards have been announced (scroll down the linked page to see the list, with further links to watch each winner on YouTube).

 

The two top winners are indeed excellent films that you can enjoy watching on their own merits, without caring that they were made with AI. That’s obviously the long-term goal, even though right now, we also watch AI videos to appreciate the technical advances each new crop showcases.

 

Grand Prix winner: A Face Only a Mother Could Love by Robert Gaudette (8 min.). The tribulations of an extremely ugly man who finds love in the end. The moral of the story is trite (there must be thousands of tales making the point that outer beauty is not what’s most important), but it’s told well enough that I don’t mind. Wonderful character design of the ugly man, including how he looked as an ugly boy.

 

Gold Medal: The Well by Dorian & Daniel (5 min.). A murder mystery set in a fantasy world of anthropomorphic animals. I won’t spoil the film by revealing which animal is the killer, but the victim is a wolf, which runs counter to the stereotypes of traditional fairy tales. The detective story is decent enough, but the characters drive the film.

 

Both films have decent stories that do engage the viewer, but I would say the character design makes the difference between good and great. The ability to design and animate characters that don’t exist in the real world is a trademark of AI movies, even though legacy movies sometimes achieve similar results with masks and prosthetics.

 

Character design wins the day in current AI movies. (GPT Image 2)

 

All the runner-up awards (listed on the Festival page) disappointed me, as they were in the style of countless boring and pretentious film-school student projects. (I don’t know if those filmmakers were, in fact, film school students, but I have seen enough short “art” movies to recognize and despise the genre.)

 

On the one hand, I agree with the award judges’ selection of the two winners: these two short films were clearly the best. On the other hand, unless the unrewarded films were all complete sludge, the judges were hoodwinked by pretentiousness instead of recognizing engaging storytelling.

 

We don’t need AI to produce more boring art films. Stop rewarding pretentiousness. No more film-school-envy. Create something people want to watch, whether the old gatekeepers look down on it or not. (GPT Image 2)

 

From a technical perspective, the Japanese Between Before and After and the Korean Divine Retribution both featured impressively photorealistic human characters, and the Japanese AI actors even “acted” well (i.e., were convincingly animated), as far as I could judge while watching the movie with subtitles without understanding what was said.

 


While I didn’t like the stories in the photorealistic AI movies, the AI actors seemed just as good as human actors, showing what’s already possible in the hands of a skilled AI director. (GPT Image 2)

 

Overall, a disappointing catch, but finding two great films made the festival worth my time. The films are being screened in Tokyo on July 30 if you’re in town.

 

Is AI ready to produce full-length feature movies? Probably not. The 5–8 minute running length of the two winners is likely the current maximum, and I think the longer of the two shorts would have benefited from being one or two minutes shorter.

 

Hypertext Hero: Ted Nelson

In my series of UX heroes, we have come to Ted Nelson, who was more of an eclectic hippie than a scientist or designer. Nelson invented hypertext, which is the basis for the World Wide Web, though he has complained that the Web didn’t implement all of his ideas.

 

Ted Nelson invented hypertext as a concept, but the main practical implementation of his ideas, the Web, didn’t live up to his vision. (GPT Image 2)

 

For me, Ted Nelson holds a special place, as he inspired me to become a UX expert specializing in the usability of hypertext and online information, which in turn prepared me perfectly to develop guidelines for usable web design. In particular, I read Nelson’s pioneering 1974 book Dream Machines around 1980, when I was casting about for a thesis topic, and even though the book was already about 6 years old, it was so revolutionary that it inspired me to pursue similar research.

 

Here is a comic strip (made with GPT Image 2) where my characters Alice and Zimo meet Ted Nelson, who takes them through Dream Machines and how his vision was partly realized by the Web:



Ted Nelson’s Dream Machines was a riotous scrapbook of ideas that anticipated much of the modern Web while diverging sharply from it.

 

Nelson’s central thesis was that technology expresses human dreams, and that computers are not calculating devices but “dream machines” for ideas, fantasy, and expression, continuous with cinema and literature. (Seems obvious now, but was radical then, which is what inspired me so much.) His insistence that computer experiences are media to be designed prefigured today’s UX and interaction design. The Web partly fulfills this as a general medium for text, image, sound, and video, though its commercial drift toward advertising and engagement optimization narrows a vision that was more literary and civic.

 

He argued against institutional mainframes (the dominant computers at the time) in favor of small, personal, interactive computers as creative tools. The Web rode exactly this shift toward personal and networked machines. Nelson’s populist goal (anyone could publish and anyone could read) closely matches the early Web’s self-image, though Tim Berners-Lee implemented a simpler, pragmatic version rather than Nelson’s rigorous structures. (Which is why Sir Tim’s project actually worked and he got a knighthood.)

 

The conceptual spine is hypertext: non-linear, interlinked text and media. Nelson coined terms for chunk-style hypertext (discrete units joined by links), stretch text (text that expands in place), and intertwingularity (the claim that everything is deeply interconnected). The Web almost perfectly realizes chunk-style hypertext. Stretch text survives only as ad hoc progressive disclosure, accordions, and “read more” toggles, never standardized. Intertwingularity underpins search engines, PageRank, and knowledge graphs, though without Nelson’s demand for curated, meaningful links.

 

Nelson’s “fantics,” the art and science of conveying ideas emotionally and cognitively, foreshadowed debates about persuasion, dark patterns, and the rhetoric of interfaces. He also imagined branching movies and interactive cinema; these remain niche, surviving mainly in interactive episodes, games, chapter markers, and time-coded links rather than as a core Web primitive.

 

Xanadu, his formal publishing proposal, began in 1960, marking the sharpest divergence between ideal and reality. It specified bidirectional links, transclusion (quotations virtually included from sources), fine-grained version management, and built-in rights and micropayments. The Web instead adopted cheap, unidirectional links, fueling explosive growth but sacrificing provenance, link robustness, and scholarly traceability. Wikis, Git, and Creative Commons emulate fragments of Xanadu at higher layers.

 

Nelson also championed non-linear learning environments and an anarchic, populist publishing network resistant to gatekeepers. The open Web, MOOCs, and the early blogosphere embodied this, yet activity has since gravitated toward large platforms and closed ecosystems, partly realizing his fears about centralization.

 

Seen from 2026, Dream Machines is not a failed blueprint but a rich design space from which the Web took only thin slices, chiefly chunk-style hypertext and democratized publishing, while leaving aside the structural, epistemic, and rights-related ambitions that would have demanded far greater complexity and global coordination.

 

Error Recovery

Errors happen. Even the best-designed systems can’t prevent every mistake, which is why my ninth usability heuristic, Help Users Recognize, Diagnose, and Recover from Errors, matters so much. (See all 10 usability heuristics in infographics, with design patterns to use and to  avoid.)

 

Of that heuristic’s three verbs, recover is the one that pays. When something goes wrong, users should never feel stranded. Strong error recovery starts with messages written in plain human language rather than cryptic system codes. Tell people exactly what happened and where, highlighting the specific field or step that needs attention.

 

Then give them a clear way forward. Offer retry buttons so a single tap can fix the problem. Show format examples so users know precisely what’s expected. Provide alternative routes when the first path fails. The goal is to solve the problem, not merely announce it.

 

Good recovery design turns frustration into confidence. Users who can quickly correct their course and continue will stay engaged, instead of abandoning the task. They learn to trust that mistakes are reversible and that the system has their back.

 

So when life spills, refill. Design experiences that catch users when they stumble and guide them smoothly back on track. That’s usability working exactly as it should.

 

Error recovery builds usability. (GPT Image 2)

 

AI Drones Help Replant the Rainforests

My prediction 16 for 2026 was “Physical AI: The Brain Gets a Body,” which I scored as having come a little more than half true (58%) in the midyear assessment of my 2026 predictions. I came across a new example that would add to this score: MORFO is a French–Brazilian “full‑stack” reforestation startup that uses AI‑guided drones and custom seed capsules to restore degraded forest ecosystems at large scale, especially in Brazil and parts of Africa.

 

The AI brain gets a flying body. Physical AI indeed. (GPT-Images-2)

 

Before any seed is dropped, MORFO runs an ecological and spatial analysis pipeline.

 

  • Field and lab data: Soil samples, on‑the‑ground surveys, and local botanical knowledge feed a species catalog and constraints for each site.

  • Remote sensing: Satellite and high‑resolution drone imagery map topography, moisture, canopy gaps, existing vegetation strata, and degradation patterns.

  • AI modeling: Machine‑learning and computer‑vision systems help identify priority micro‑sites, suitable species mixes, and planting densities, and later track biomass growth, biodiversity indicators, and carbon stocks over time.


The result is a planting scheme that doesn’t simply maximize “trees per hectare” but distributes different species according to soil type, slope, water availability, and succession dynamics, effectively encoding a designed forest architecture that drones then execute.

 


Human planners only have the brainpower to handle large areas as an undifferentiated whole, whereas AI can keep thinking cheaply and construct new plans for every tiny plot of land, accounting for local variations. (GPT-Images-2)

 

Drone Seeding Mechanics

MORFO uses agricultural‑grade drones, roughly 1.5 meters in diameter, equipped with custom release mechanisms and loaded with several thousand seed capsules per flight.

 

  • Seed capsules: Each capsule contains pre‑germinated native seeds plus protective microorganisms and substrates, tailored to local conditions; multiple species can be combined to create micro‑assemblages.

  • Throughput: A single drone can plant about 180 seed capsules per minute and cover up to roughly 50 hectares in a day, making it 20–100 times faster than manual planting and up to 5 times cheaper when nursery and logistics costs are considered.



The AI drones vastly outperform traditional planting methods, meaning that more plants get into the ground. (GPT-Images-2)

 

  • Access: Drones can operate on steep, post‑fire terrain and remote Amazonian or Sahelian sites that are difficult or unsafe for people, opening up restoration in areas that would otherwise remain untouched.



It’s just as easy to fly over nasty terrain as over flat ground, meaning that the AI drones easily replant a larger area. (GPT-Images-2)

 

In Rio de Janeiro’s hills, for instance, drones drop capsules over hard‑to‑reach slopes above the city, following AI‑defined flight paths and planting patterns while municipal authorities provide site selection and regulatory support.

 

Human–Machine Collaboration and Outcomes

MORFO is explicit that drones augment, rather than replace, human ecological work: botanists, soil scientists, and local stakeholders define species choices and restoration goals, while pilots operate drones and teams conduct ground‑truthing and long‑term monitoring. Drones then serve as scalable actuators for an already‑designed ecological plan, and as repeatable sensing platforms for tracking biomass, biodiversity, and carbon over the years.

 

Early results from various Amazon sites suggest that thousands of hectares of degraded land can be put on a recovery trajectory more quickly than with conventional sapling‑based reforestation, though the long‑term ecological performance still depends on classic factors like species selection, climate stress, and socio‑political protection against renewed deforestation.

 

My conclusions from this project:

 

  • AI is moving far beyond the chatbot stage. Not just getting a body, but a flying body.

  • AI makes things cheaper, which means that they get done. Instead of talking about the rainforest, it replants it.

 

As we say in Silicon Valley: “you can just do things.” AI often makes worthy things so cheap that they actually get done instead of being the topic of endless discussions. (GPT-Images-2)

 

  • AI allows sophisticated analysis at scale: for example, dropping an optimized selection of seeds for each micro-plot, which produces a more viable replanted ecosystem than a predetermined seed mix.   


Cognition at scale means AI drones can plant optimal seeds in every small area. Something that could never be done by expensive human planning. (GPT-Images-2)

 

Annotations: A 1,000-Year-Old Pattern That Software Keeps Fumbling

An annotation is a note with an address: commentary anchored to the exact spot it concerns. That anchoring is why margin comments beat every detached feedback channel for precision, and it’s also the pattern’s weak joint: when content changes or notes pile up, the annotation layer decays into crust that buries what it was meant to illuminate.

 

Marginalia, the ancestral annotation system: every note pinned to the exact passage it explains. Scribes ran this design pattern for 1,000 years without a single sync conflict. (GPT-Images-2)

 

Definition: An annotation is user-added commentary (a comment, highlight, correction, tag, or drawing) attached to a specific location within a piece of content, such as a text span, an image region, a video timestamp, or a spreadsheet cell.

 

Every annotation has two parts: the body (what you say) and the anchor (where it applies). The W3C Web Annotation Data Model, a formal web standard since February 2017, codifies precisely this body-plus-target structure. Remove the anchor and you no longer have an annotation; you have a detached comment, floating free of context, forcing every reader to reconstruct what “this is wrong” refers to. (Spoiler: half of them reconstruct it differently.)

 

From Monastery Margins to Microsoft Word

The pattern predates software by roughly a millennium. Medieval scribes filled manuscript margins with glosses explaining difficult words and disputed doctrine, and the Talmud institutionalized the layout: core text in the center of the page, ringed by layers of commentary accumulated across centuries. Mathematics owns the most famous annotation of all: in 1637, Pierre de Fermat jotted in his copy of Diophantus that he possessed a marvelous proof too large for the margin. Supplying that missing body took 358 years. Insufficient margin width is, evidently, a usability problem with a long tail.

 

Digital tools adopted the pattern early. Microsoft Word shipped comments and revision marks in the early 1990s, embedding annotation into office life. The web then supplied a governance lesson: Third Voice (1999) let anyone paste sticky notes onto any website, publishers erupted at seeing rival commentary overlaid on their pages, and the service was dead by 2001. Google Docs later made anchored margin comments the default way office documents get reviewed, and the 2017 W3C standard gave the pattern formal plumbing for interoperability.

 

Why Annotations Work: The Anchor Carries Half the Meaning

Catherine C. Marshall of Xerox PARC published the field’s foundational study in 1997, examining the markings students left in used university textbooks. Two findings still steer design today. First, annotations serve future readers, not only their authors: students shopping for used textbooks favored copies whose previous owners had marked them up helpfully. Second, the tool shapes the behavior: highlighter-wielding students wrote fewer margin notes than pen users, because producing legible words with a highlighter is a chore. (Methodology note: this was qualitative fieldwork on paper artifacts, with superb ecological validity; these were real marks made by real students with no researcher hovering nearby.)

 

Why does anchoring matter so much for collaboration? Consider the alternative, verbal addressing: “third paragraph on page 6, second sentence, the clause after the comma.” Every such description costs the writer effort, costs the reader a search, and breaks the moment anyone edits the document. An anchored comment costs one click and points with pixel precision. Proximity also cuts short-term memory load, since the reader sees the claim and the critique side by side instead of juggling two windows. Do what the medieval page did. Keep commentary next to its subject.

 

How Annotations Fail: Annotation Crust

Annotation systems fail slowly, by accumulation. This results in annotation crust: the hardened layer of stale threads, resolved-but-never-closed comments, duplicates, and orphaned notes that encases a document until nobody reads any of it. Crust forms through 4 mechanisms:

 

  • Orphaned anchors. The content changes, and the note now points at nothing, or silently at the wrong thing, which is worse. Robust systems re-attach anchors using the surrounding context and visibly flag the ones they cannot rescue.

  • Occlusion. 40 pins on one screenshot means the map has covered the territory. Collapse markers at high densities, and let users toggle the whole annotation layer with one keystroke.

  • Notification floods. Every comment pinging every collaborator converts a useful layer into a mute-me machine. Batch updates into digests once a thread heats up.

  • Privacy accidents. A user may believe the note is private, but it is visible to the entire workspace. Awkwardness follows, or lawsuits. State the audience at the moment of writing, right beside the input field, not in a settings page three screens away.


None of this argues against annotations. It argues for treating them as a managed lifecycle rather than a bottomless pile.

 

9 Design Guidelines for Annotation Features

  1. Anchor precisely. Attach notes to the exact span, region, timestamp, or cell, never to the whole document, because “somewhere in here” helps nobody.

  2. Show the anchor with the body. Quote the annotated excerpt inside the comment panel and in notifications, so the note makes sense even out of context.

  3. Survive edits gracefully. Re-anchor using surrounding text when content changes, and flag orphaned notes explicitly instead of deleting them without a trace.

  4. Give every annotation a lifecycle. Open, resolved, archived. Hide resolved threads by default, and let users purge the archive.

  5. State visibility at creation time. A label such as “Visible to everyone in this document” beside the input field prevents the worst annotation accidents.

  6. Make the layer toggleable. One keystroke should hide every note so the content itself can be read clean.

  7. Batch the notifications. Send digests rather than a ping per comment once activity passes a few events per hour.

  8. Support filtering. By author, status, and date; a reviewer returning after a week needs “what’s new” in one click.

  9. Keep annotations subordinate to the content. Margins, collapsed pins, and hover reveals; the moment the notes visually dominate the material they describe, the crust has won.


Send Notifications Through the More Available Sense

Smart glasses can interrupt you through your eyes or your ears. Which should they pick? New research from Dinithi Dissanayake, Prasanth Sasikumar, and Suranga Nanayakkara at the National University of Singapore’s Augmented Human Lab answers: whichever sense happens to be less busy at that moment (HeadRoom, arXiv preprint, July 2026).

 

Their system estimates the moment-to-moment “availability” of the visual and auditory channels straight from a wearable’s egocentric camera and microphone. The trick borrows from predictive-coding neuroscience: a sensory stream that’s hard to predict is a stream the brain is already working hard on. Two tiny self-supervised predictors forecast the next video-frame embedding and the next audio-feature vector, and high prediction error flags an occupied channel. The full model weighs a mere 0.625 MB and runs in 11 ms per step on a Meta Quest 3S headset, so it’s deployable on today’s wearables with no cloud round-trips. Think of it as Christopher Wickens’s Multiple Resource Theory from the 1980s, finally productized at the edge.

 

Multiple Resource Theory is rarely used in user interface design practice, because few systems are truly multimodal. The experimental smart glasses UI from Singapore used the theory to save users a tenth of a second. Doesn’t sound like much, but in use cases such as air traffic control or military applications, that time can be the difference between life and death. (GPT Image 2)

 

The evaluation: 22 participants (25 recruited; 3 excluded for missing the 80% detection bar) watched three 6-minute egocentric videos in a headset while responding to 200 ms flashes or 1 kHz tones, routed to the predicted better channel, the predicted worse channel, or randomly. In the calm scenes, routing didn’t matter. But in the high-demand scene (dynamic motion plus concurrent speech, with video comprehension accuracy dropping to 50%), misrouted probes cost 114 ms in response time compared with correctly routed ones, a large effect (d = 1.3). Meanwhile, subjective NASA-TLX workload ratings showed no differences across conditions: behavior registered what self-reports missed. Watch what users do, not what they say.

 

Caveats: model-based routing didn’t significantly beat random routing, so the proven effect is avoiding the worst channel, not divining the best one. And detecting a beep is far easier than absorbing a calendar alert or an agent’s status report. Fair enough for a first controlled study.

 

Why this matters: agentic AI will multiply notifications, because agents working in the background must report back on finished tasks, exceptions, and approvals. Interruption design thus becomes a core UX problem, and output modality should be a runtime decision driven by the user’s perceptual budget rather than a settings checkbox. I predict attention-aware routing will ship as a platform service in mainstream wearables within 5 years. Notifications that knock on the open door feel polite. The rest are sensory spam.

 

Token Use

This weekend, I used Fable 5 in agentic mode on Max setting in Cowork to do a literature review. The agent ran for 8 hours (talk about slow AI) and did produce a nice lit review. However, it consumed 4.7 million tokens in this small project, which would have cost $235 if charged at Anthropic’s list prices for API use.

 

Luckily, I used Fable on a subscription plan, and it “only “ used up 50% of my weekly rate limits on this run. Since the monthly subscription is $200, the 4.7 M tokens equate to around $25. That estimate assumes that I max out the subscription every week, which I’ll probably do this week, since I’ve already used up 57% for this lit review and a few smaller projects. In general, I don’t go up to the rate limits, so a more realistic estimate is that I spent maybe $35 worth of a month’s subscription fee on this one agentic AI run. Even when getting huge discount with a subscription compared to the by-the-token API charges, frontier AI is getting expensive.

 

I may be turning into a token whale now that I’m addicted to Fable 5 with max reasoning. (GPT Image 2)

 

Conclusion: The Machine Finally Adapts to the Human

Eight stories, one arc. For the first 40 years of computing, the human did the adapting: typing English at machines that spoke nothing else, deciphering error codes, memorizing which menu hid which command. This week’s stories show the adaptation burden switching sides. It was my lifelong goal and motivation for specializing in usability to make technology suitable for humans instead of requiring humans to bend over for the machines. I’m happy this is finally happening.

 

Gemini answers a student in Hanoi in Vietnamese. GPT-Live waits for you to finish a thought. HeadRoom’s smart glasses check which of your senses is free before interrupting. MORFO’s drones compute a fresh planting plan for every micro-plot of rainforest. Even the humble error message is learning to speak plain language and offer a way out.

 

None of the underlying ideas is new. Ted Nelson sketched humane, media-rich computing in 1974. Christopher Wickens formulated Multiple Resource Theory in the 1980s. Medieval scribes anchored commentary to the exact passage it explained 1,000 years ago. The vision was never the bottleneck; the cost of execution was. AI has collapsed that cost. So half-century-old ideas are finally shipping as products.

 

But cheap intelligence doesn’t hand out its winnings evenly. Firms spending $34 per employee per month on AI grew headcount by 10% while the $3 dabblers got nothing, and at the AI Film Festival, abundant production capacity mostly yielded another crop of pretentious art films. Production got cheap; taste stayed scarce. The rewards go to whoever does the reps.

 

Usability has always meant bending the machine to fit the human. For decades, that bending was expensive, so users did the contorting instead. AI makes it cheap, from a Thai senior’s spoken prompt to a seed capsule tuned for one patch of the Amazon. The keyboard is already planning its retirement party. Don’t let it save you a seat.

 

AI is inverting our decades-long problem with users having to adapt to the machine. Now AI adapts to us. (GPT Image 2)

 

A Final Thought


Top Past Articles
bottom of page