top of page

UX Roundup: AI Diffusion Speed | Synthetic Personas | AI Learning | Brad Myers | Agatha Christie | Mental Accounting | Grok Imagine 2

Writer: Jakob Nielsen
Jakob Nielsen
Aug 10
18 min read
Summary: AI technical capabilities are exploding, but sophisticated use crawls | Progress bars make waits bearable | Synthetic personas exaggerate demographics | AI hurts education when used wrong | Professor Brad Myers is a UX hero | In the future, we’ll all be Agatha Christie | Mental accounting gives users separate budgets for separate expenses | Grok releases upgraded image model

UX Roundup for August 10, 2026 (GPT Image 2)


AI Moves at Three Speeds, and the Third One Is Stuck

New data from the National Research Group (NRG) shows AI capability and adoption advancing at a record pace while the third speed, sophisticated use, crawls. Two human barriers explain the lag: a search-engine mental model of AI (a usability problem for vendors to fix) and a new AI stigma (which the AI labs must dismantle). NRG’s superusers prove the third speed is learnable: many started in the same search pit.


Advanced use of AI is still vanishingly rare: the sophistication lag is very real. (GPT Image 2)


Every Technology Has Three Speeds

Any new technology moves at three distinct speeds: how fast its capabilities advance, how fast people adopt it, and how fast they use it well. The third speed decides when society collects the payoff, and it’s chronically the slowest. Factories electrified in the 1890s, but as economic historian Paul David documented in 1990 (7-page PDF), productivity jumped only around 1920, once managers redesigned plants around the electric motor. The first two speeds are supply-side: vendors ship capability, users download an app. The third demands skill formation: new mental models and new habits. Habits change on human time, not GPU time. This gap is the sophistication lag.


NRG’s June 2026 report Beyond Adoption: Paths to AI Expertise measures all three speeds for AI, pairing a March 2026 survey of 1,009 Americans with a Q1 2026 study of 129 AI superusers. (NRG brands them “Trailblazers”; I’ll stick with the plain term.) Speed one is blistering: METR’s time-horizon data shows frontier models completing tasks that take human experts 3+ hours, with capability doubling roughly every 7 months. Speed two set a record: ChatGPT reached about 800 million users in 3 years, a level the internet needed 13 years to hit (Financial Times data); over half of Americans use AI weekly.

And speed three? Crawling: most users stay parked at simplistic use.


Most AI users hug the shore of traditional computer use and barely get their feet wet in the ocean of new opportunities. AI swimmers are still rare, but NRG identified 129 AI superusers to study. (GPT Image 2)


The Superusers Already Live at Speed Three

NRG set a high bar for its superuser panel: use AI daily, run 2 or more of the 7 major LLMs multiple times per week, adopt new tools early, pay (or be willing to pay) for AI, and show at least 3 power-user behaviors, such as creating multimodal content, building custom agents, or chaining text, audio, and video in one workflow.


They aren’t all engineers. The panel spans knowledge workers, creatives, students, and parents. One panelist keeps a coding copilot running across his entire workspace and co-built a medical image-segmentation model with university researchers. Another started out treating AI as a search engine and now uses it for photo editing, interior design, meal planning, and workouts. A nonprofit fundraiser uploads spreadsheets and asks for patterns; a business owner routes inventory, marketing, and engineering calculations through AI.


Two lessons follow. First, the superusers are what Eric von Hippel termed lead users in 1986: people whose needs today preview the mass market’s needs in a few years. Watch them to design tomorrow’s mainstream product. Second, and more hopeful: many superusers report starting in the very search pit where today’s beginners sit, describing their early AI use as glorified Googling until context-giving and delegation clicked. Expertise here takes a mindset shift plus practice; no computer-science degree required.


So what keeps everyone else stuck? NRG identifies two human barriers.



Barrier 1: The Search Mindset Is a Usability Problem

NRG asked Americans what a chatbot does when answering a question. A whopping 58% said it looks up the answer in a huge database. Only 13% picked the correct answer: it predicts words from patterns learned in training. Nor is this a one-poll fluke: Tavern Research asked a similar question in August 2025 and got a 45% plurality for the database theory, with 28% correct.


Users act on their mental model, not on the system’s actual workings, so usage follows the (wrong) model: 63% use AI for quick answers, while only 18% automate tasks, and a measly 14% apply AI to coding or data analysis. AI-powered search works great, but treating search as the ceiling leaves most of AI’s value untouched.


Don’t blame the users. Blame the box. A novice opens a chatbot, and he or she sees an empty field and a blinking cursor: Google’s UI circa 1999. Of course, the search mindset takes over. Nothing in the interface reveals what else the system can do. And intent-based outcome specification, the first new UI paradigm in 60 years, shipped without a manual.

Among the non-adopters the superusers know, the top barriers are not knowing what AI could do (50%) and not knowing how to prompt (46%), well ahead of privacy worries (35%) or accuracy distrust (27%). Understanding, not trust, is the bottleneck.


Asked for the single most important thing AI companies could do, 43% of superusers picked education about practical uses, 20% a more intuitive UX, and 11% real-time guidance. Their prescriptions read like a usability checklist: suggest likely prompts instead of a blank box, reshape vague requests on the fly, and teach while the user works. One superuser traced his or her breakthrough to a metaphor swap: stop treating the prompt field like a search bar, start treating the AI like a bright, context-hungry intern. As I argued in 4 Metaphors for Working with AI, the intern framing is now too limiting, but it beats the search-engine framing by a mile. Ben Thompson reminds us that PCs exploded only once the mouse, icons, and windows made computing legible. AI is still waiting for its mouse.



Barrier 2: AI Stigma Blocks the On-Ramp

The second barrier is social. The superusers rated the same obstacles twice: as experienced when they learned AI, and as observed in today’s non-adopters. Fear that others would judge them for using AI scored 16 percentage points higher for today’s novices than for the pioneers. Panelists described friends who consider AI use cheating and students who avoid AI on assignments for fear of plagiarism charges.


The deltas are the tell. The barriers that grew most are social and psychological: privacy worries (+18), low confidence (+16), fear of judgment (+16), and ethical qualms (+13). The practical barriers barely moved: tool cost and output quality each rose a mere 1 point. In 3 years, the models improved at speed one while the social reception soured. That inversion should embarrass the industry.


This extends the pattern I documented in AI Stigma: identical work gets rated worse once it’s labeled AI-made, and about half of employees hide their AI use at work. Stigma is also self-fulfilling. Whoever fears judgment uses AI furtively and shallowly, gets mediocre results, and concludes the tool was overrated. NRG adds a twist: today’s pressure is relational, radiating from classmates, coworkers, and friends, so rebutting op-ed critics won’t fix it.


Reducing this stigma is the AI labs’ responsibility. They spend billions accelerating speed one and lavish marketing on speed two, yet spend next to nothing on the social blocker throttling speed three. One superuser’s advice to the labs was blunt: “Make it feel normal. That’s it.” So normalize: show ordinary people doing ordinary tasks (emails, meal plans, spreadsheet cleanup) instead of superhuman demo reels, defend legitimate AI use in schools and workplaces, and make AI assistance a badge rather than a confession. The purgatory is escapable: pocket calculators were banned from many classrooms in the 1970s until the National Council of Teachers of Mathematics endorsed them at all grade levels in 1980, and spell-checkers weathered identical cheating accusations. AI stigma will fade too, but deliberate normalization can shave years off the wait.



Conclusion: The Third Speed Is Where the Money Is

Capability sprints. Adoption runs. Sophistication crawls, and the sophistication lag is where AI’s economic promise sits idle. The superusers prove the destination is reachable by ordinary people: parents, fundraisers, and students. Both barriers blocking everyone else have owners. The search mindset is a usability defect vendors can fix with better prompt-box design, and AI stigma is a social defect the labs must spend real money to dismantle. Electrification waited 30 years for the factory floor to catch up. Let’s not make AI users wait that long: fix the box, and fight the stigma.


(Alice and Zimo comic strip made with GPT Image 2. This time in color scratchboard style.)


Synthetic Personas Exaggerate Demographics

AI models role-playing survey respondents couldn’t beat a simple lookup table at predicting individual humans’ answers, and they exaggerated the attitude gaps between demographic segments 2–4x. UX learned this lesson decades ago: base personas on behaviors, not demographics.


4 AI Models vs. 106,478 Real Humans

“Synthetic users” are AI models prompted with a demographic profile and asked to answer questions as that person. Vendors pitch them as cheap replacements for user research: a study that takes months and thousands of dollars shrinks to a few API calls. Zihan Chen and co-authors from Stevens Institute of Technology and the University of Massachusetts Boston put this promise to the test. They ran 4 AI models (two Claude, two Llama, spanning small to frontier scale) against ground truth from two workhorse datasets of social science: the General Social Survey (14,704 recent US respondents) and the World Values Survey (91,774 respondents across 63 countries).


The smart methodological move: every model was benchmarked against a demographic lookup table fit on held-out human data. For any profile, the table returns the most common answer among real people with that profile. An AI that can’t beat this trivial predictor adds nothing beyond the demographics themselves.

A Lookup Table Beat the AI

No model beat it. On US attitudes, the best AI tied the lookup table; on cross-cultural values, every model scored 11–22 percentage points worse. So a spreadsheet of averages outpredicts billions of parameters. Ouch.


In fairness, the models reproduced aggregate population distributions well. That’s the trap: a team that validates only the aggregate will wrongly conclude the simulation works. But no aggregate ever shops on your site: you serve one person at a time. The average American male wears a size 10.5 shoe, but if you sell only size 10.5 shoes, you’ll forfeit around 85% of potential sales.


Even when an AI persona correctly predicts the average, it may be wrong about individual customers and thus fail to predict their actual behavior. (GPT Image 2)


Identity Treated as Destiny

The second failure matters more for design decisions. The models treated demographics as far more predictive of attitudes than they are among real people. Political views explain a measly 1.5% of the variation in Americans’ confidence in banks, yet the models behaved as if politics explained up to 67% of it: a roughly 40-fold exaggeration. And a bigger brain made things worse, since the frontier model stereotyped more than its smaller sibling.


Translated into the segment-targeting decisions synthetic users are sold to support, the models inflated between-segment gaps 2–4x, targeted the wrong segment in 50% of US cases and 72% of cross-cultural cases, and invented segment splits with no basis in the human data in up to 41% of the value questions. The simulated population is not a blurred copy of humanity but a caricature. Silicon samples turn out to be silicon stereotypes.


Demographics are the wrong approach to personas. (GPT Image 2)


Behavior Beats Demography

Since Alan Cooper popularized personas in his 1999 book The Inmates Are Running the Asylum, the cardinal guideline has been to build personas from users’ behaviors, goals, and skills, not their demographics. Knowing that a user is a 45-year-old suburban woman tells you almost nothing about how she’ll use your product; knowing that she reorders the same 20 grocery items every week tells you plenty. I made the same argument in my article on individualizing UX. This benchmark quantifies why: even for attitudes, demographics explain only a few percent of person-to-person variation, and AI multiplies that weak signal into fake certainty.


Personas already risk being superficial and misleading, even when they are based on real data from real people. In this case, maybe it’s more important to know that “Sandra” climbs skyscrapers for a living than that she likes coffee. (GPT Image 2)


Three takeaways for user researchers:

  1. Don’t build synthetic users from demographic sliders. You’ll inherit stereotypes dressed up as data.

  2. Never accept aggregate similarity as validation. Check individual-level accuracy and segment gaps against real human data before trusting a single simulated finding.

  3. Keep watching real users. The authors scope their results to demographic prompting; synthetic users conditioned on behavioral data remain untested and are the more promising path, precisely because behavior predicts behavior.


AI will transform user research, but this study shows that the current shortcut fails where it counts. Demography was never destiny, except inside the model.


AI Homework Help: Grades Up 18%, Learning Down 20%

The largest study yet of unsupervised AI use in schools tracked 26,811 Chinese students for 30 months. AI adoption raised homework scores by 18% and cut homework time by 30%, but closed-book exam scores fell 20%, and high-stakes entrance exams eventually fell 18–24%. The culprit is homework outsourcing: AI users who worked as long as their unaided classmates learned just as much.


David Strömberg (Stockholm University) and co-authors from the University of Hong Kong followed students in grades 7–12 in a county in central China from September 2022 to June 2025 (SSRN paper). AI adoption grew from almost zero to 80% of students over the period, and the staggered timing let the authors compare each adopter against never-adopters in a difference-in-differences design. A digital homework platform supplied grades and completion timestamps, monthly closed-book exams measured short-run learning, and China’s zhongkao and gaokao entrance exams measured the long run.


The productivity numbers look glorious: homework scores up 18%, completion time down from 64 to 45 minutes. But as we know from prior research, when AI does the work for students, learning suffers. In this case, monthly exam scores dropped 20% within 6 months, and students with 2 years of AI exposure scored 24% lower on the zhongkao and 18% lower on the gaokao exam.


The timestamps expose the mechanism: after 5 months of use, 81% of AI students finished homework faster than the fastest unaided student, with grades matching what AI models score on such problems. High homework grades, collapsing exams: phantom mastery. But AI users who still spent normal time on homework kept their exam scores and earned better homework grades on top. In this study, top students were hurt more than weak ones. (We need more research to tease apart AI’s impact on clever kids vs. dull ones.)


AI in education can help or hurt, depending on how it’s used. (GPT Image 2)


Regular UX Roundup readers have seen this movie before. AI hurts education when it does the work for students and helps only when it’s used as a tutor to guide the learning process at each student’s individual pace, for example by supplying customized explanations and hints. The new study adds what the earlier studies I covered lacked: massive scale, self-chosen everyday tools, and a horizon long enough to show the full 2-year process. (For example, two weeks ago I covered a study where 12 hours of AI use boosted African students’ learning of mathematics. Great! But what would have happened the next school year or if they had used AI for 30 hours? Would AI have continued to accelerate these students’ learning, or was it a one-time benefit?)


4 Guidelines for AI That Teaches Instead of Cheats

  1. Configure AI as a tutor, not an answer machine. Require hints, worked explanations, and self-testing. In this study, the Chinese students who kept normal study time learned fine with AI.

  2. Monitor inputs, not outputs. AI has corrupted homework grades as a signal, so watch time on task instead. Suspiciously fast plus suspiciously good equals phantom mastery.

  3. Shift assessment weight to closed-book, in-person work. Effort must pay again, or students will keep renting competence from a chatbot.

  4. Tell students that the bill arrives late. AI users don’t feel the learning loss while it accumulates. Credible information about the 2-year delayed cost is itself an intervention.


UX Hero: Brad Myers of CMU

We Brad Myers provided empirical evidence for the usability benefits of progress indicators in 1985. He has given the UX field much more during the 41 years since his grad-student days, and I have had the pleasure of discussing many interaction design topics with him.


Professor Brad Myers is one of my heroes of UX, with a list of achievements that couldn’t even fit on 10 infographics (those 19 best papers probably each deserve an infographic!). (GPT Image 2)


Until I made this infographic, I hadn’t even realized that Myers worked at PERQ Systems in the 1980s. PERQ was an early GUI workstation, and I used one myself for a 1983 experiment comparing windowing and scrolling as ways for users to see more information than fits on the screen. That’s ancient history, and speaking of history, Myers’s recent book Pick, Click, Flick! is both an intriguing overview of interaction techniques and a useful manual for their proper use when designing graphical user interfaces. I was happy to provide one of the blurbs for the back cover.


About windowing vs. scrolling: “windowing” is when the user thinks of moving the viewport up to see information at the top of the document, and “scrolling” is when the user thinks of moving the content down to see information at the top of the document. This is purely a matter of mental models, since in either case, the computer screen doesn’t move, but simply redraws the pixels in new spots. At least in 1983, when I did the study, windowing won, but it’s possible that with touchscreens, scrolling would now have the edge. It could be a nice little graduate-student project to investigate whether using a mouse vs. fingers affects the optimal mental model.


In the Future, We’ll All Be Agatha Christie

Late in life, the detective-story writer Agatha Christie reflected on 1919: looking back, she found it remarkable that she had naturally assumed she would have a nurse and a servant, yet never imagined being the sort of person who owned a car, which was for the rich alone.

In other words, in 1919, servants were cheap, and cars were expensive. Salaries were low, and manufactured goods were dear.


Today, it’s the opposite: servants are expensive, so only the very rich keep full-time staff. But mass manufacturing has made cars cheap enough for even poor people to own one. (This started with Henry Ford’s assembly line in 1913, but apparently the associated price drop hadn’t fully reached England by 1919.)


My prediction: AI will return us all to living like Agatha Christie. By this, I don’t mean that we’ll all write Murder on the Orient Express, though AI will cause an explosion in creative expression, even for people who can’t write mystery best-sellers.


Who was the killer? Read the book to find out. I’m not giving away any spoilers. (Muse Image)


Rather, I mean that we’ll all have servants, but not own cars. The servants won’t be humans (salaries will increase even more, as society gets immensely rich because of AI). But we’ll all have several household robots to perform the tasks that servants used to do. We’ll even have some non-humanoid robots for new tasks: think micro-drones that zap mosquitoes before they can bite us.


On the other hand, only eccentric collectors will own personal cars. Robotaxis will perform all transportation tasks on demand. Probably so efficiently that existing bus services will be abolished, because a Robotaxi that comes to your door at the press of a button and takes you to your exact destination will beat any bus that makes you walk to the stop and wait in the rain.


What did the conductor see on the Orient Express? His great-grandkids won’t have jobs in public transportation, because most transit systems will fold as Robotaxis become vastly superior. (Muse Image) 


Just one example of how AI will transform many things we used to take for granted. Though in this case, it essentially returns us to Agatha Christie’s world of 1919. Sans the murder.


Mental Accounting: Users Spend From Jars, Not Bank Accounts

People sort money into separate mental budgets and refuse to treat a dollar in one budget as equal to a dollar in another. Design for these invisible jars: make them visible when that helps users, and never reach into the jar with the loosest lid.


The economist insists that a dollar is a dollar. The user disagrees: he or she decided which jar that dollar lives in long before your pricing page loaded. (GPT Image 2)


Definition: Mental accounting is the set of cognitive operations people use to organize, evaluate, and keep track of their financial activities, typically by assigning money to separate mental budgets that they treat as non-interchangeable.


Economic theory assumes fungibility: any dollar can substitute for any other dollar. Real users violate fungibility daily. A $50 restaurant gift card buys a fancier dinner than $50 of salary ever would. A $400 tax refund turns into a gadget, while $400 of wages goes to the electric bill. Same money. Different jar. Different behavior.


And the jars have lids of different tightness. Money labeled “savings” is guarded like a dragon’s hoard, while money labeled “bonus,” “credit,” or “winnings” practically spends itself. Thus, any interface that touches money is also touching the user’s internal bookkeeping system, whether the design team realizes it or not.


Thaler Turned a Household Quirk Into a Nobel Prize

Richard Thaler, then at Cornell University, introduced the theory in Mental Accounting and Consumer Choice (PDF), published in Marketing Science in 1985, and consolidated the field’s findings in Mental Accounting Matters in 1999. The name is a deliberate borrowing from corporate bookkeeping: people behave like small firms, posting every expense to an internal account and balancing each account separately, instead of maximizing over one big pot the way textbook economics prescribes. When Thaler received the 2017 Nobel Prize in economics, the committee’s citation prominently featured the 1985 paper. Not bad for a theory about jam jars.


My favorite demonstration comes from the 1999 paper. Wine collectors who buy futures years before delivery code the purchase as an “investment.” When they finally uncork the bottle, drinking it feels free. Eldar Shafir and Thaler titled the underlying study Invest Now, Drink Later, Spend Never: an expensive hobby, mentally laundered into a free one. Thaler and Eric Johnson also documented the house money effect in 1990: gamblers treat recent winnings as the casino’s money and bet them far more recklessly than the cash they walked in with.


Make the Jars Visible, and Users Will Thank You

Users already run their finances on jar logic, so the best financial UIs stop fighting the mental model and externalize it. Banking apps that offer named sub-accounts, envelopes, or “pots” let people move their internal ledger onto the screen, where software can enforce what willpower alone cannot. The payoff is real: earmarked money is far less likely to leak into impulse purchases, and users who see a “Vacation” balance grow will return to the app for the pleasure of watching it.


So support the jars in mundane flows, too:


  • Category summaries turn a raw transaction list into the account statements users mentally keep anyway. Reconciling “where did my fun money go?” should take seconds, not an evening with a spreadsheet.

  • Refunds should visibly return to their jar. A refund that lands as an undifferentiated blob gets recoded as a windfall and spent twice.

  • Price framing can match the account users will draw from. Billing business software annually fits the budget cycle a manager actually plans with. That’s legitimate, as long as both framings are shown honestly.


The Same Levers Power Dark Patterns

Every mechanism above has an evil twin. In-game gems convert dollars into a proprietary currency, detaching each purchase from the “real money” account where scrutiny lives. Drip pricing splits one price into a base fare plus fees, so each nibble debits a smaller jar and no single number triggers alarm. “Just $0.99 per day” reframes a $361 annual charge as pocket change. And “bonus credits” are engineered windfalls, aimed with sniper precision at the jar that has no lid at all.


Do these tricks work? Of course they work; that’s why regulators keep suing over them. But they work by taxing the user’s bookkeeping, and users eventually audit. The remedies are cheap: state the full real-currency total early, show cumulative spending, and translate every virtual currency back into dollars at the moment of spending.


You can abuse the knowledge I’m giving you about mental accounting to pry the lids off users’ mental money jars. But doing so easily turns into dark design, so please don’t. (GPT Image 2)


8 Design Guidelines for Mental Accounting

  1. Mirror users’ jars in the UI. Offer named sub-accounts, envelopes, or category labels so the internal ledger becomes an external, enforceable one.

  2. Respect earmarks. When an action would raid a protected jar (paying a bill from “Vacation” savings), say so before the confirmation, not on next month’s statement.

  3. Reveal the full price early. Every fee disclosed late lands in an unbudgeted mental account and feels like a betrayal, even when the total is unchanged.

  4. Translate virtual currencies at the point of spend. “500 gems ($4.99)” keeps the purchase connected to the account users actually budget in.

  5. Aggregate the nibbles. A yearly recap such as “you spent $214 on subscriptions” restores the visibility that small recurring charges are designed to escape.

  6. Frame time-based prices both ways. If you advertise per day, show per year in the same breath. One framing is marketing; two framings are information.

  7. Make credits and refunds legible. State their dollar value, expiration date, and restrictions in plain sight, because vague credit is windfall bait.

  8. Test comprehension, not just conversion. Ask 5 users what they believe they paid and will pay next month. Wrong answers are a defect, whatever the funnel metrics say.


Mental accounting is neither a bug to fix nor a lever to yank. It’s how normal people impose order on chaotic finances with a brain that never evolved for compound interest, and it mostly serves them well. Design that supports the jars earns trust, retention, and long-term revenue. Design that raids them earns chargebacks, churn, and a subpoena. UX = Profits, and with money UIs the profitable move is the honest one: respect the jars.


Alice and Zimo teach mental accounting in 16-bit pixel-art style (GPT Image 2)


New Image Model: Grok Imagine 2

SpaceXAI has launched a much-improved version 2 of its Grok Imagine image model. Here are a few images I made with the new model. I asked it to make infographics, posters, and comics about Jakob’s Law, but without giving it any information about the concept other than its name. For the last comic strip below, I asked for a strip about other famous UX principles, and it gave me Hick’s Law told by cute animals. Clearly, the model has good world knowledge, because the content of all these images is spot on.


(This layout is flawed: the label “Familiar” for the third screenshot overlaps the header for the table of the three principles, making it look like a misplaced extra word in that header. The yellow arrow also points the wrong way. These were the only layout mistakes I found in my experimentation.)


(Grok Imagine 2)


If Grok had delivered these images before the launch of Nano Banana Pro on November 20, 2025, I would have been ecstatic. Only 9 months ago, this level of accurate text rendering and imaginative design of visual information based on world knowledge was unheard of. Today, I rate Grok Imagine 2 at the level of Nano Banana 2 and slightly below GPT Image 2. This is still an impressive achievement from SpaceXAI’s imaging team!


Grok Imagine 2 beats Google’s and OpenAI’s image models in one aspect: it has an object-based editing system for the inevitable times when a generated image is good but just not exactly right. You can open a sidebar to see a list of the visual elements that make up the image. This list is nested, so that, for example, in the poster below, the “target graphic” object has a sub-object named “crosshair icon,” allowing users to edit this element in two different ways. This nested-objects editing feature closely resembles image editing in Reve.


Here, I used image editing, rather than regeneration, to fix the layout problems I had identified in one of my posters:


(Grok Imagine 2, after using the built-in editing features)


Fffinal Thought of the Day


Top Past Articles
bottom of page