UX Roundup: Benchmarking Office Work | AI Undermines Billable Hours | AI as Delegate | Error Recovery | Pareto Frontier | Left-Digit Effect | Empty States | AI Use Differences | Educational Video
- Jakob Nielsen

- 16 hours ago
- 18 min read
Summary: Chinese AI models deliver 64% of entry-level human office-work quality in new benchmark | AI breaks the billable-hours model for consulting and shifts it to outcome-based pricing | Workers delegate whole tasks to AI but rarely review the output | Turning error messages into recovery tools | The Pareto Frontier of advancing AI capabilities | The left-digit effect is why $2.99 outsells the simpler $3.00 | Alleviating the empty-screen problem | AI usage may follow a power law inside companies | Very cute alphabet song made with AI

UX Roundup for August 24, 2026 (GPT Image 2)
New Office-Work Benchmark: AI Agents Hit 64% of Human Quality for 3–26% of the Cost
Baidu’s Agent Frontier Team (Jingbo Zhou and 14 colleagues) released OmegaUse-OfficeVal, a benchmark of 100 long-horizon office tasks: real workplace requests involving documents, spreadsheets, slide decks, and PDFs. The authors call this territory “vibe working,” the office sibling of vibe coding. The average task takes a junior human worker 2.32 hours, with a market-price estimate of $6.86 at the going rate for commodity office chores on Chinese outsourcing platforms. (In the United States, the loaded cost of an entry-level office employee would be around US $100 for these tasks.) Scoring runs through 2,228 automated code checks with two user-centered twists: a deliverable scores 0 unless the file opens and stays editable, and agents lose points for collateral damage the recipient must repair afterward. Judging the artifact from the perspective of the user who receives it is exactly right. More benchmarks should copy this.
The best of the 5 models tested, GLM-5.2, scored 17.91 against 27.79 for the human baseline: 64% of junior-worker quality. But the AI cost $0.22–$1.76 per task against $6.86 for the human, and 4 of the 5 models worked 4–13 times faster. Cheap, fast, but mediocre. (The standard consulting saying rides again: you can deliver only two of the three attributes clients want.) Also noteworthy: weighting tasks by their price changes the winner (Qwen3.7-Plus captured the most economic value), so the crown for “best model” depends on whether you count tasks or dollars.

The consulting business has long had a saying: clients want good, fast, and cheap, but they can pick only two. You can’t have all three at once. Baidu’s study shows that AI currently obeys this rule for office work. (GPT Image 2)

A more nuanced perspective from the new benchmark: speed, quality, and cost are not binary criteria but continuous scales. The “best” Chinese model changed depending on whether the evaluation emphasized raw task performance or performance per dollar. (GPT Image 2)
Thus, for international customers, the choice is about a dollar for Chinese AI vs. $100 for a junior employee. Pay 100x the money for less than 2x the quality. In many cases, this tradeoff already favors AI, and plenty of companies have stopped hiring entry-level staff altogether.

Junior-level staff still perform office work better than Chinese AI models, but probably only for the rest of this year. Because Chinese AI models are much cheaper than American AI models, and because American human office workers are much more expensive than their Chinese counterparts, the cost–benefit tradeoff of using AI vs. humans in the US already favors AI for most entry-level tasks.
Two caveats. First, the 5 models are all Chinese and a generation behind the frontier (Baidu’s own Ernie is conspicuously absent from the lineup), so today’s best models would presumably score higher. Since the team released everything, down to the verifier code, somebody should run the American frontier models and report back.
Second, and more important: a single benchmark run says little about where AI is headed. A benchmark run once is a data point; run repeatedly, it’s a speedometer. METR, the US research nonprofit behind the AI time-horizon benchmark, earned its influence by re-measuring every major model release. That discipline is how we know the length of tasks AI can complete autonomously has doubled roughly every 7 months since 2019, and closer to every 4 months lately, a trend I’ve analyzed in previous newsletters.

Baidu’s report gives us only a snapshot of the current capabilities of Chinese AI models. If Baidu re-runs the same benchmark at regular intervals, the data turns into a trendline that lets us estimate the future. Other AI benchmarks have done this, which is why we can say that substantially improved performance by the end of the year is extremely likely. (GPT Image 2)
METR measures how long a coding task an AI can finish; OmegaUse-OfficeVal measures how well AI delivers everyday business documents, priced in dollars and measured in China rather than California. Different tasks, different metric, different geography: exactly the independent second opinion our future trend analyses need. So, Baidu, please re-run this benchmark on every new model generation. One measurement is a curiosity. A time series is a forecast.
My prediction is that the current AI quality level of 64% of junior-worker performance will reach 90% by the end of 2026 and exceed 100% by the end of 2027. By then, AI will eat mid-level business professionals for lunch. Senior staff? Might be safe until late 2028, or even early 2029 if AI progress slows. But by 2030, professionals who haven’t pivoted their careers will be unemployed, and deservedly so.

For now, AI is cheaper but not better than humans. But AI improves every month, whereas human IQ is flat, unless we measure in millennia. AI will likely surpass humans soon. (Grok Imagine 2)
AI Breaks the Billable-Hours Business Model
Reuters reports that India’s $315 billion IT-services sector is rapidly shifting from labor-based pricing toward outcome-based contracts as clients demand AI productivity gains. Tata Consultancy Services (TCS) says roughly 80% of contracts in finance, HR, and related business services are now tied to performance outcomes. Persistent Systems says clients increasingly expect the same work for 25–30% less. Some work is also moving back inside client organizations because AI makes in-house delivery economical.
This great unbilling makes the timesheet the buggy whip of professional services. Stop defining AI value through seats, prompts, hours saved, or feature usage. Products and services will increasingly be judged against measurable completed outcomes. That requires explicit acceptance criteria, baselines, quality measures, and instrumentation capable of demonstrating that AI actually produced the promised business result.
I asked Grok and Meta to visualize this short news item:

Grok Imagine gave me this rather unimaginative infographic. It’s a nice layout with great text, and probably a more appealing way for many people to consume this content than my dense prose.

Meta’s Muse Image lived up to its name and produced a poetic retelling of my article across a series of images, rather than the single poster I had requested. (Here, I agglomerated and reduced 6 large images into a single small one.) I don’t know why the consultant in the lower left is crying over the great business metrics she’s getting. (It would be a fast enough iteration to ask Muse to redraw her without the tears, but I appreciate a touch of the absurd in my illustrations, so I’m showing you the original image.) I like Muse’s fashion sense, though I doubt that TCS staff dress as snazzily as the consultants in the upper-right image.
Users Delegate to AI but Don’t Review the Outcome
Epoch AI commissioned Ipsos to survey 1,106 employed US adults (fielded July 10–19, 2026, margin of error ±3 points). The headline: 20% of workers say AI now handles at least one task they previously handed to a coworker or contractor. Task-level AI use runs from 25% (record keeping) to 57% (designing computer systems and software).

One-fifth of US employees now have AI handle some tasks they used to give to coworkers. (Muse Image)
The finding UX people should stare at: 66% of AI outputs get used unchanged or with only minor edits, and just 5% get heavily reworked. Time savings also scale with delegation depth: 53% report savings when AI does most or all of a task, vs. 37% when it assists partially.

Users increasingly delegate entire tasks to AI, but rarely review the outcome. (Grok Imagine 2)
AI has crossed from tool to delegate: workers hand it whole tasks and skim the results, exactly as they would with a junior colleague. Since 2/3 of output ships with light or no editing, the review experience now matters more than the generation experience. Design for delegation: clear handoff affordances, status visibility while the AI works, and verification support that makes checking cheaper than trusting. The unedited 66% is a quality time bomb, and review UIs are the only place to defuse it.

Users increasingly treat AI as a coworker that performs tasks, rather than merely a tool that executes commands. (Muse Image)
If something is difficult, people won’t do it. That’s equally true for shopping on ecommerce sites and for reviewing AI output. The solution is not to scold employees and tell them that it’s reckless to skip the review stage. Scolding users doesn’t work and violates my deepest usability philosophy: when computers are used wrong, the designer is at fault (not the user).

As the old saying goes, don’t try to teach a pig to sing: it won’t work, and it annoys the pig. The same goes for scolding users for using computers wrong. (Muse Image)
Unfortunately, the AI labs are arrogantly inept at usability and thus doomed to relearn the lessons the legacy software industry learned in the 1980s and the ecommerce industry learned in the 2000s. When it comes to bad design, history doesn’t rhyme; it plagiarizes.

I’ve seen this movie before: new technology is always designed with insufficient attention to usability and how normal people use it. (Muse Image)
The cure for unreviewed AI output has to be a better review experience. For sure, some low-risk tasks shouldn’t be reviewed at all: if we ask people to do too much, they won’t do anything. But mostly, we have to make it easier to do the right thing.

The way to make users do something is to make it easier. (Muse Image)
Errors Are an Opportunity
An error message has one job: get the user moving again.
“Something went wrong” announces the failure and leaves the user stranded. A useful recovery state says what happened, what was saved or changed, and what to do next. Offer the best next action (retry, fix the exact field, restore the last good state, or reach a human) and make any partial completion visible. In probabilistic systems, retries should be reversible because a second attempt may fail differently or overwrite good work. This is the practical meaning of my usability heuristic to help users recognize, diagnose, and recover from errors.
Recovery means tools, not apologies. Offer a retry button, a saved draft, a pointer to the exact field that needs fixing, and a path to a human when all else fails. Stock every error state like a first-aid kit. Users don’t need to know the system’s feelings. They need a wrench.
To go further, treat recovery data as product research. Track where failures recur, which remedy users choose, how long recovery takes, and whether users ultimately succeed. When a workaround sees heavy use, the workflow itself is begging for a redesign. The best recovery is the one a future product change makes unnecessary.

4 tools that beat the traditional error-dialog toolkit of a lone “OK” button. (GPT Image 2)
The Pareto Frontier of Advancing AI Capabilities
If you follow AI developments, you’ve no doubt heard of the Pareto Frontier of AI capabilities, which is discussed every time a new AI model is released: does it push the Pareto Frontier out, or does it fall short of being Pareto-efficient?
What does that mean? Alice and Zimo explain the Pareto Frontier in this comic strip, in Renaissance oil-painting style, made with GPT Image 2:





The Left-Digit Effect: One Cent Buys a Full Dollar of Perception
$2.99 registers as “two-something,” so it feels nearly a dollar cheaper than $3.00. The effect appears only when the leftmost digit changes, but it moves real money in real markets and distorts every multi-digit number in your UI, not just prices.

In the user’s mind, the left digit inflates to fill the whole number: $2.99 is “two dollars” with a rounding error attached. (Muse Image)
Definition: The left-digit effect is the tendency to anchor magnitude judgments on the leftmost digit of a multi-digit number, so that a tiny change that flips the left digit (say, $3.00 to $2.99) shifts perceived size far more than the arithmetic difference justifies.
The mechanism is reading itself. We convert digits into a sense of size while scanning left to right, and the encoding starts before the eyes finish the number. By the time you reach the “.99,” the verdict “about two dollars” has already been filed. Users perform a digit drive-by: grab the first digit, keep moving.
A 19th-Century Trick Gets Its Name in 2005
Prices ending in 9 have haunted store shelves since the 19th century, wrapped in retail folklore. (One popular story claims odd prices forced cashiers to open the till for change, preventing pocketed bills. Charming, unverified, and beside the point.) The cognitive explanation, and the name, arrived when Manoj Thomas and Vicki Morwitz published Penny Wise and Pound Foolish: The Left-Digit Effect in Price Cognition in the Journal of Consumer Research in 2005. Across 5 experiments, they showed the critical boundary condition: $2.99 feels reliably smaller than $3.00, but $2.49 feels no smaller than $2.50. Same one-cent drop. No left-digit change, no effect. That single contrast demolished the folk theory that “9s just look cheap” and replaced it with a mechanism.
And the effect survives contact with real money at scale. Nicola Lacetera and co-authors analyzed over 22 million wholesale used-car transactions in a 2012 American Economic Review paper and found prices dropping discontinuously at every 10,000-mile odometer threshold: a car showing 79,900–79,999 miles sold for about $210 more than an essentially identical car just past 80,000. Professional dealers, five-figure stakes, and the left digit still ran the auction.
Honest Uses: Choose Endings and Precision on Purpose
Is a designer allowed to use this knowledge at all? My answer is yes, in two legitimate ways.
First, price endings carry meaning, so pick the meaning you intend. A 9-ending whispers “deal”; a round number whispers “quality” and confidence. Discount retailers and premium brands both understand this, which is why the outlet mall is wallpapered in .99 and the luxury site charges a flat $400. Matching the ending to the honest positioning of the product is communication. Mismatching it is noise.
Second, the effect governs every number you display, so round with intent. A battery at 19% sits alarmingly in the “teens” while 20% is merely low; a 3.9-star rating sits a full psychological flight of stairs below 4.0; and a 1.9 GB file feels noticeably smaller than 2.0 GB. A task at 89% complete feels far from done. When users compare such numbers, inconsistent precision quietly rigs the comparison. And comparison views are exactly where the effect bites hardest: Tatiana Sokolova and co-authors showed in a 2020 Journal of Marketing Research paper that the left-digit bias is strongest when prices sit side by side on the screen and weaker when one price must be recalled from memory. So within any comparison set, one rule: same format, same decimals, same rounding, for every item.
The Dark Version: Shading Every Number in Your Favor
The classic abuse is mere charm pricing, which is legal, universal, and mild. The user-hostile versions stack the effect with other tricks. A “from $9.99” anchor followed by a parade of drip fees uses the low left digit as the bait and the final payment screen as the switch, with the first impression doing all the selling. Rendering cents in tiny superscript (a big 24 with a microscopic 99 floating beside it) shrinks the ink exactly where the information lives. Comparison tables that show the house product at $19.99 and competitors rounded to $25 manufacture a left-digit chasm out of a $4 difference. And subscription tiers priced at $9.99, $19.99, and $49.99 count on the drive-by to blur a 5x spread into “single digits, teens, forties.”
The remedies cost nothing: display unit prices, show the all-in total next to any charm-priced base, and keep digits typographically equal. If a price works only when the cents are in 6-point type, the price doesn’t work.
8 Guidelines for Multi-Digit Numbers
Keep formats consistent within a comparison set. All prices with cents or none with cents; mixed precision is a covert thumb on the scale.
Show totals and unit prices. The left digit of the base price stops mattering when the real, complete number stands beside it.
Match endings to message. Use 9-endings for genuine promotions and round prices for premium positioning; don’t send both signals at once.
Render every digit at equal size. Superscript or grayed-out cents subtract legibility exactly where the user needs it most: a dark pattern in 6-point type.
Mind thresholds in data displays. When a metric crosses a left digit (9.8% to 10.1% error rate), users perceive a leap, so label the actual change to keep interpretation calibrated.
Round UI numbers to the precision users need, then hold that precision everywhere. A dashboard mixing 39.9 and 40 invites false alarms in one direction and false comfort in the other.
Don’t pair charm prices with drip fees. A .99 anchor plus surprise charges converts a mild nudge into a bait-and-switch, and regulators increasingly read it that way.
Test price recall, not just clicks. Ask 5 users afterward what the product cost. If they answer a full dollar low, your pricing display is misinforming the people who pay you.
The left-digit effect is a one-cent lever that moves perception by a dollar, documented from lab experiments to 22 million auctioned cars. You can’t turn it off in your users, and you needn’t apologize for knowing it exists. But every number in an interface either informs or shades, and the difference is intent plus typography. Give the first digit honest company: full totals, equal type, consistent precision. The first digit will always do the talking. Your job is to make sure it isn’t lying.
Empty States Are Full of Usability Problems
An empty state is what a screen shows when there’s no data to display: no files, no messages, no search results. Teams polish the data-rich ideal and neglect the void, turning first impressions and failed searches into dead ends. Design the nothing, or users will do nothing.
Definition: An empty state is the condition of a screen when the content it exists to display is absent. It comes in 3 flavors: first use (the user hasn’t created or received anything yet), user-cleared (he or she deleted, archived, or completed everything), and no results (a search or filter matched nothing).
Why the peculiar name? Because designers model each screen as a set of states driven by its data: loading, error, populated, and, when the data set holds zero items, empty. The problem is old. The 1984 Macintosh Finder showed an empty folder as a white rectangle with a terse “0 items” header. Fine, because an empty folder explains itself. Web applications don’t. By 2006, 37signals warned in Getting Real that designers mock up screens overflowing with sample data while every paying customer meets the product at its barest: the blank slate. Product designer Scott Hurff extended the idea in 2015 with his UI stack: every screen has 5 states (ideal, loading, partial, error, and empty), and teams lavish attention on the ideal state while the other 4 rot. The vocabulary stuck. So did the neglect.

Even an apex predator hesitates when facing pure nothing. Your users have less patience and no claws. (GPT Image 2)
Empty = Dead End
A blank screen ambushes users at the worst possible moments.
First use is when motivation peaks and knowledge bottoms out. Quettra’s analysis of 125 million Android devices, published by Andrew Chen in 2015, found that the average app loses 77% of its daily active users within 3 days of install. People decide fast, and they decide while staring at your emptiest screens. Thus, the screen with the least content carries the heaviest persuasion duty, yet it’s the one screen nobody designed.
Blankness is also ambiguous. Is the screen empty, still loading, or broken? Users can’t tell, which violates the oldest heuristic in the book: visibility of system status.
For search, zero results is a dead end wearing a polite face. I’ve watched users flounder on search since the 1990s, and the pattern hasn’t changed: when the first query fails, many people never compose a second, better one. They leave. And empty results are no rarity.
Creation tools suffer a third variant: the cursor blinking on a white page. Blank-canvas paralysis: infinite possibility, zero guidance, no action taken. (Writers know it as fear of the blank page.)

A blank screen hits the user when she’s already down. Wicked. (Grok Imagine 2)
8 Design Guidelines for Empty States
The cure costs little: treat nothing as content, designed with dashboard-level rigor.
Never ship a naked blank. Every zero-data screen must say what belongs there and why it’s currently absent, so users don’t mistake empty for broken.
Match the message to the flavor. First use, user-cleared, and no results demand different words and actions; a generic “Nothing here” fits none of them.
Give first-timers a primary action. “Create your first project” outperforms a smorgasbord of 5 competing buttons. Momentum first, mastery later.
Show the destination. A sample item, a template, or a grayed-out preview of the populated screen teaches by recognition and spares users the lecture.
Turn no-results pages into workbenches. Keep the query visible and editable, suggest spelling fixes and broader terms, offer to drop filters, and show popular content as an escape hatch.
Blame the system, never the user. Write “We couldn’t find matches for ‘X’,” not “Invalid query.”
Let cleared states celebrate, briefly. An emptied inbox should look intentionally done, not broken: one line of congratulations, then a pointer to what’s next.
Instrument the void. Log how often each empty state appears and what users do afterward. If over 10% of searches return zero results, your vocabulary doesn’t match your users’, and synonyms are your cheapest fix.
Illustrations are welcome, but charm is not a next step. The testable minimum: every empty state contains at least 1 actionable control.
An empty state is a blank canvas, but it’s your canvas, not the user’s. He or she showed up to get something done, not to admire the void. So say what’s missing, show what good looks like, and hand over the brush. Users who do nothing buy nothing, renew nothing, and recommend nothing. Design the nothing, or nothing is what your product will earn.
AI Usage Likely Follows a Power Law Inside Companies
SaaS provider Rippling disclosed to TechCrunch that its AI token spending had hit a staggering 40% of its R&D headcount budget, with 10–15% of employees driving roughly 60% of the spend; one engineer burned about $50K per month. After the company built an employee-level ROI dashboard, spend fell to about 15% of budget. This is the clearest X-ray yet of how concentrated real AI usage is inside a company.

A single user spent $50K on AI in a month. This may seem excessive, but it’s common for heavy AI use to 10x a skilled engineer’s output, so assuming this person costs more than $5K per month fully loaded, the ROI could still have been positive. (Muse Image)
Based on this N=1 data, AI usage is likely to follow a power law within companies, just as it does elsewhere: a small cadre of power users generates most of the activity and most of the cost. Two design lessons follow.

Even though the Rippling data is only one case study, it agrees with what theory suggests: AI usage is highly skewed and likely follows a power law, as it has for most other computer features in the past. This means that power users are worth studying. (Grok Imagine 2)
First, to sell tokens, design for the power users; they consume a wildly disproportionate share of AI. Also, to see the future, user researchers should study power users, since they are more likely to have discovered advanced uses of AI that mainstream users will adopt next year.

For a rapidly changing technology like AI, it’s worth studying power users for an early look at the future. That said, power users are also unusual, by definition. For example, their usability requirements are typically low, and they are willing to struggle through complex user interfaces. So we can’t generalize all findings from power users to the broader employee base. (Muse Image)
Second, cost visibility is a usability feature: once people could see their own spend next to their results, behavior corrected itself without mandates. Expect tokenmeters (AKA AI-spend dashboards) to become standard enterprise UX within a year.

Knowledge changes behavior: show users their AI spend, and use corrects itself. (Muse Image)
If 10% of users account for 60% of AI tokens, then 90% of users account for the remaining 40%. On a per-capita basis, the power users are 13.5x as active as the mainstream users. But AI is crossing the chasm, so the mainstream will catch up to the early adopters.
Decades of data show that user activity doesn’t follow a bell curve. It follows a power law: the most active users don’t do a little more than the rest; they do 10 or 100 times more. There is no average user.
Definition: In a power-law distribution (the simplest case is a Zipf curve), activity is proportional to 1 divided by rank. User number 2 does half as much as user number 1, number 10 does a tenth, and number 100 does a hundredth. On log-log paper, the curve is a straight line. On normal paper, it’s a ski jump: a cliff at the head, then a long, nearly flat run-out.

The standard power law has an even steeper drop than this concept image from Muse Image. Across a vast range of computer use cases, a few “whales” account for the lion’s share of use, to mix the zoological metaphors. What Muse Image got right was using diamonds to visualize the power curve: understanding it is worth a lot of money.
Count the drop: from rank 1 to rank 10, usage falls 90%. From rank 1 to rank 100, it falls 99%. When I plotted the popularity of early websites in 1997, the result was a near-perfect Zipf curve: a few sites accounted for most page views, while millions subsisted on scraps. And the same skew repeats at every scale:
Participation inequality. My 90-9-1 rule for online communities: 90% of users lurk, 9% contribute a little, and 1% of users produce almost all contributions.
Feature use. Most users of a complex application touch a handful of commands, while a zealous few hammer the advanced features hard enough to dominate the logs.
Search queries. A few head terms occur millions of times; frequency then plunges into a long tail of queries typed only once.
Thus, the arithmetic mean lies. A few whales drag the average far above the median user, who barely registers. (Averages work fine for height and shoe size. They fail for behavior.) Economists have known the pattern since Vilfredo Pareto counted Italy’s landowners in 1896 (the same Pareto as in the comic strip above, wearing a different hat); the web simply cranked the dial harder.
Three design guidelines follow:
Design the defaults for the head. The few dominant tasks deserve the prime screen real estate. Measure which tasks those are instead of guessing.
Don’t amputate the tail. Each rare feature looks like deadwood, but in aggregate the tail often serves a large share of total use. Support it with search and shortcuts rather than top-level chrome.
Report medians instead of means, and cherish your top 1%: they create the content, the referrals, and much of the revenue.
The usage curve is a ski jump, not a gentle hill. Design for the cliff and the run-out, never for the mythical middle.
Recommended AI Video
A very cute AI video: Alphabet Song (X, 1 min.) by Umesh. If you have children at alphabet-learning age, I recommend watching with them. But even without kids, the video and the song are charming enough to earn a minute of your time and show you AI’s potential for educational content.
If you watch with children, I recommend muting the video the first time: it’s great fun to shout out the names of animals and objects as they appear. Then rewatch to hear the song.
Consider how much time a traditional animation studio would have needed to make this video. Shudder. Umesh made it in a few hours using MiniMax’s agentic workflow.
A Final Thought




