top of page

UX Roundup: Employers Value Experience Over Skills | Exponential Growth in AI Use | AI Dependence | Undo | Disappointing Suno | AI’s Economic Scenarios | Reflective Design | Mini Maps | GPT Image 2.5

Writer: Jakob Nielsen
Jakob Nielsen
9 hours ago
32 min read
Summary: Since January 2025, tech job postings have asked for 25% fewer skills but 5% more experience | AI spending among OpenAI’s heaviest research users is growing at 129% per month | 26% of young-adult chatbot users say they feel dependent on AI, but the addiction scales disagree | Undo = the user’s escape hatch | Suno 6 sounds worse than Suno 5 | Anthropic’s 3 scenarios for how AI will transform the economy forget the robots | Sometimes you should make users stop and think | Mini maps provide an overview of large navigation spaces | New and slightly improved image model from OpenAI

UX Roundup for September 11, 2026 (GPT Image 2.5)


Employers Now Want More Experience and Fewer Skills

As AI takes over the grunt work, employers place less emphasis on candidates’ skill sets when hiring. After all, if AI doesn’t already perform a production task, I expect it to take over within a month or two. So proficiency in hands-on production matters less when those hands are delegating more and more tasks to AI.


Meanwhile, human judgment and savvy in navigating the organization are becoming more valuable as AI carries out the work. Producing another design, analysis, or implementation becomes cheap. Choosing the right one, and persuading your colleagues to use it, remains the route to business value.


AI can supply more of the execution. Someone still has to decide which destination deserves the effort. (GPT Image 2)


Employees need to know what to tell the AI to do, how to judge the flood of alternative results AI proposes, and how to shepherd the work to completion through the molasses of organizational inertia. Those abilities develop through experience with actual consequences. A course can explain how to conduct a usability test, but convincing stakeholders to act on the findings takes practice with real people and real resistance.


Revelio Labs analyzed 75 million tech job postings and found that these changing requirements are showing up in corporate hiring. The chart below covers January 2023 to June 2026 and tracks two measures: the number of skills requested and the years of experience employers request of candidates.


Average number of skills and years of experience requested in tech job postings, by month. Skills are plotted against the left axis, and experience against the right. (Data from Revelio Labs. I replotted the data to start both axes at zero, making the relative changes easier to judge. The proportional decrease in skill requirements is much larger than the proportional increase in experience requirements.)


While such numbers always fluctuate, a clear change began in January 2025. During the roughly 1.5 years from then to June 2026, the number of skills requested dropped from 28.8 to 21.6, a 25% decrease, equivalent to a compound annual decline of about 17%. At the same time, requested experience increased from 4.62 years to 4.85 years, a 5% increase. That adds 0.23 years of experience over the period, or about 0.15 years (roughly 2 months) per year.


Look closely at the chart, and you’ll notice that the experience line was already creeping upward in 2023 and 2024. The skills line broke from its earlier pattern in January 2025. Employers had wanted experience all along; what changed was their appetite for laundry lists of skills.


For UX professionals, this is a useful warning about the sales pitch on your résumé. A long inventory of software packages and methods tells employers what you’ve walked past. They need evidence that you can apply the relevant capabilities to their problems.


A long skills list describes what you’ve encountered. A useful result shows what you can do. (GPT Image 2)


Skills Are Codified; Experience Is Tacit

The deeper cause of the trend reaches back 60 years, well before today’s AI boom. In 1966, the Hungarian-British philosopher Michael Polanyi observed that we can know more than we can tell: much of what a skilled person knows is tacit, acquired by doing and impossible to write down in full.


MIT economist David Autor developed this observation into an account of Polanyi’s paradox (2014), explaining why 3 decades of computerization hollowed out the routine, rule-following middle of the labor market while sparing both the janitor and the executive: computers could only do what humans could specify explicitly.


But AI breaks half of that paradox. Large language models have swallowed the world’s codified knowledge: tutorials, manuals, Stack Overflow threads, and design-pattern write-ups. A “skill,” in the sense that a job posting or a certificate uses the word, is generally a named, teachable competence. That makes it easier to teach in a course and easier for AI to acquire from recorded examples.


The half of the paradox that survives is the half Polanyi cared about: knowing which of the AI’s 5 proposed designs will survive contact with real customers, sensing that a plausible analysis smells wrong, and knowing that the VP of Sales will kill any project that changes her dashboard. Much of this knowledge never reaches the AI’s training data. People acquire it through encounters with particular customers, colleagues, and organizations, and they can’t write it down in full.


Skillsolescence is becoming, in my view, the dominant force in the market for tech talent: the collapse in hiring value that a named, codified skill suffers once AI can perform it. My best guess is that a codified production skill loses most of its market value within 2–3 years of AI reaching competence in it, and that this lag is shrinking. (The Revelio numbers on coding languages, where required years of experience fell while requirements for every judgment-related skill rose, are the first supporting evidence I’ve seen.)


Experience buys three things that a course alone can’t supply. Employers have long paid for these abilities implicitly; their growing demands for experience make that preference more explicit:


  • Direction. Knowing what to ask for. A designer who has watched 200 checkout flows fail in 200 different ways can specify the three constraints that matter before the AI generates the first draft. The AI can churn out alternatives by the dozen. Experience tells the designer which ones are worth producing in the first place.

  • Detection. Catching an answer that sounds plausible but is wrong. AI output can fail silently and confidently, and an experienced detector is a human who has been burned by that class of error before. A junior employee can check whether the AI’s code compiles; a senior colleague checks whether it solves the customer’s actual problem.

  • Delivery. Getting the work adopted. This means knowing whom to ask, when to escalate, how to survive the review meeting, and which stakeholder must see the prototype before the steering committee does. This is the least codified knowledge of all, and it differs between organizations, which is why experience in a specific company or industry commands a premium over generic seniority.


Direction, detection, delivery: a skills taxonomy captures these abilities poorly. In my reading, they are what employers are really after when they type “5 years of experience” into a posting.


From Degrees to Skills to Judgment

Step back, and the Revelio chart is the third act of a longer play. Of course, employers can’t readily observe judgment in a job application, so they screen candidates using whatever proxy is cheapest to check, until that proxy loses its predictive value. Here’s my rough division of this history into three periods:


Three Generations of Hiring Signals Era What employers screened for What it was a proxy for Why it stopped working Most of the 20th century The college degree Trainability and general ability Credential inflation: once everybody had one, it stopped discriminating 2015–2024 Named skills (the 30-item laundry list) Ability to execute production tasks AI executes the tasks, and keyword lists became free to fake 2025 onward Years spent performing specific skills Judgment: direction, detection, delivery Nothing yet, except that the supply is drying up

There’s a delicious irony in the middle row. For a decade, HR departments preached “skills-based hiring” and signed pledges to drop degree requirements in favor of demonstrated skills. Now their own postings quietly demote skills in favor of experience.


As for the bottom-right cell of my table: if nobody hires junior staff, then 5 years from now there will be no new cohort of applicants with 5 years of experience. Employment among 22–25-year-olds in highly AI-exposed occupations already trails that of their less-exposed peers by 19%, and the AI revolution has barely left the starting gate. We need a new approach to apprenticeships to develop the senior staff of the future.


Getting Your First Job

If this trend continues, job candidates will need an additional 2 months of experience each year to get hired. If you already have a job, that’s nothing, since you’ll presumably rack up 12 months of experience annually. (If AI continues to accelerate, employers’ experience requirements may rise by 3, 4, or even 6 months annually in the future. Still no problem for the employed.)


If you’re a new graduate, or if you’re unemployed, the hiring door narrows another notch every year.


It’s the oldest Catch-22 in hiring: you need experience to get hired, but you can only get that experience by doing the work.


My advice to young folks is to prioritize experience above all else when preparing to join the job market. The best experience comes from a regular job, even a part-time position while you’re studying. If that’s out of reach, three good options remain:


  • Internships. An internship can give you access to customers, operating constraints, and experienced reviewers that would be impossible to assemble on your own. Its value comes from what you’re allowed to do with that access. Unfortunately, internships are evaporating as AI takes over the grunt assignments companies once handed to junior staff. And since senior staff are working overtime managing their AI agents, they have little time left to mentor interns.


If you can get one of the few remaining internships, make it a priority, even if it pays little or nothing. The few thousand dollars you forgo in intern pay are dwarfed by the several hundred thousand dollars you stand to gain over your working life from starting with a good job and advancing from there. But insist on a good manager or mentor: he or she will be invaluable and can help you for years after the internship.


The two main criteria for evaluating an internship are (a) whether you get useful hands-on experience with real business problems and (b) whether you have a good mentor. (GPT Image 2)


  • Volunteering. If your uncle owns a small business, volunteer to help him adopt AI. Begin with a specific problem, such as missed inquiries or repetitive customer questions. If no actual business will have you, no matter how small, then a nonprofit organization is the next-best option, though hiring managers value business experience more highly. A serious open-source project is an exception because it does have customers with skin in the game, even if the “product” is given away for free. The most exciting open-source project right now, in my view, is Omarchy, which is building Linux for the rest of us and could benefit greatly from volunteers with design and usability expertise.

  • Self-directed projects. Find something that needs doing, and do it. You can start with a $200/month AI agent and your own initiative. See, for example, my list of 76 research challenges for AI user experience, and pick one of those projects. Defining your own project and carrying it to completion gives you experience and demonstrates agency: the ability to initiate and finish useful work, which is one of the most valuable talents in a workplace transformed by AI.


AI creates an opportunity for newcomers because it slashes the effort required to turn an idea into something people can try. A small project, pursued as an intern, as a volunteer, or on your own, can now expose you to problem definition, implementation, evaluation, and revision, even when you lack the production skills to do every step unaided.


But you must participate in the decisions. Asking AI for a solution, accepting it, and handing it over only trains you to be a courier. To develop judgment (Silicon Valley jargon currently calls this “taste”), write down what you expect to happen before you see the result, then compare that expectation with what actually happens.


Record your prediction before you see the results. Otherwise, hindsight can disguise the lesson you needed to learn. (GPT Image 2)


The aim is to use faster execution to buy more contact with reality. If AI saves you 10 hours of production, spend some of those hours finding out whether you’re building the right thing. Otherwise, you’ve accelerated production while leaving your judgment untested.

Existing employees shouldn’t treat the hiring trend as a promise that another birthday makes them more valuable. Experience earns its return through the accuracy of the decisions it supports.


Five years of repetitive production can leave judgment exactly where it started. The “years served” metric is convenient shorthand for job listings, but the calendar tells you little about what someone has learned from those years. (GPT Image 2)


Some knowledge survives changes in tools: how users misunderstand a task, why a stakeholder resists a change, and what happens when a team optimizes the wrong metric. Other knowledge depends on conditions that AI can alter, such as how long a prototype takes or whether exploring another implementation is affordable. Experience can be a liability if it supplies automatic vetoes for possibilities you haven’t examined.


Old-timers have internalized many constraints that once applied but no longer hold in the age of AI. For example, building working prototypes of a wide range of design alternatives used to be too expensive for most projects. Now, it’s cheap. Experience becomes a liability if you don’t update those inherited assumptions. (GPT Image 2)


(Hat tip to the useful a16z “Charts of the Week” newsletter for alerting me to Revelio Labs’ job postings data.)


3 Agent Workdays per Human Workday: Inside an Organization Saturated with Agents

Inside OpenAI’s research organization, AI agents now log 3.1 workdays for every human workday, yet more than half of the successful long tasks still need a human to step in. I don’t expect your company to reach that ratio soon. But the interface problems it exposes will arrive much sooner: supervising agents, paying for them, and living without the help desk they hollow out.


OpenAI published telemetry on its own researchers from January 2025 to mid-August 2026. The median researcher consumes about $600 a day in inference at published prices; at the 90th percentile, consumption exceeds $7,000. These figures use OpenAI’s published token charges rather than the company’s internal costs, so $7,000 per day is what you would pay at retail to match a researcher near the top of OpenAI’s consumption distribution.


The following chart shows growth in daily AI spending at the 90th percentile of OpenAI’s researchers, from January 4, 2026 ($2), to August 15, 2026 ($7,047). The chart uses a logarithmic scale, so a straight line represents exponential growth.


Data from OpenAI replotted on a logarithmic scale. (GPT Image 2.5)


Relative growth exploded in January 2026, rising from $2 to $29 in 3 weeks. This was the month when, in my view, AI overtook humans at software development and the world began to notice. More interesting is the trend since January 25, 2026: relentless exponential growth, though at a slower relative pace than in the first weeks of the year.


Over this period, spending on AI tokens grew at an average compound rate of 129% per month. I repeat, a monthly growth rate of 129%. That means more than doubling each month and equates to an annualized growth rate of more than 2 million percent.


Can this pace continue? Will daily AI consumption for a researcher at the 90th percentile actually reach $290,000 at published prices? This sounds unrealistic, but the growth trend has persisted for almost 7 months. Perhaps it can continue for another 4.5 months, from August 15 until the end of the year. That’s an extrapolation of consumption at today’s prices, rather than a forecast of OpenAI’s actual bill.


Agent runtime overtook total human labor only this summer, so even at the frontier the crossover is a recent development. Success rates rose across difficulty bands, but in the last 6 months, over half of successful 4–8-hour tasks involved at least one intervention. Meanwhile, attendance at office hours for troubleshooting dwindled, and the questions migrated to the agents rather than to another human support channel.


Before you extrapolate, remember who produced these data: a vendor reporting on its own elite users, with abundant compute and code that machines can test. Treat OpenAI’s token spending as a weather forecast for your power users: it tells you which storms are heading for your interfaces, but nothing about how much you should spend. The interface lessons transfer better than the numbers.


AI can drain money at a frightening clip unless we’re careful. OpenAI faces different economics because it supplies its own tokens and wants to experience the future firsthand, exploring how work changes when people use vastly more AI. You need to watch your budget. (GPT Image 2.5)


  1. Supervision is the product. When more than half of successful long tasks need human intervention, the agent’s real UI is the check-in: what changed since I last looked, where it’s stuck, what it needs from me, and how I can stop it without losing useful work. OpenAI’s researchers can diagnose a stalled agent by reading its execution trace. Your accountant can’t. Outside the lab, the interface should translate the problem into a question a domain expert can answer in one sentence.

  2. Chat doesn’t scale to 4 agents. A growing cohort of researchers runs 4 or more agents at once, plus subagents. One conversation thread was designed for one assistant; running 4 agents through it is like walking 4 dogs on one leash. Concurrent work needs an inbox: a list of agents ordered by which ones need my attention, with status I can scan in 2 seconds. (Email has given us decades of practice with inboxes. Feel free to copy it.)


Researchers already supervise agents whose combined runtime exceeds the human workday. That’s tough enough today. If the current growth continues for another few months, supervision through today’s interfaces will become impossible. We need better UX for directing and checking the work. (GPT Image 2.5)


  1. The meter belongs in the interface. OpenAI gives its researchers unusual freedom to consume tokens. Most employers will need spending controls. If a task can cost $20 or $2,000 depending on how the agent wanders, the user needs an estimate before delegating, a running total during execution, and a stop button that actually stops the work. Cost has become a usability attribute, much as latency did.

  2. Agents eat the help desk, and the help desk was a source of UX research data. Declining attendance at office hours sounds like a win until you notice that the questions have vanished from human view. Every “how do I…” that an agent answers is a potential usability finding nobody logs. Record the questions agents handle as first responders, and route those logs to the people who fix the product.

 

You save money when AI does work that used to require a help desk. But unless you make the AI document its support actions, you’ll never know which weak processes or poor designs need fixing to prevent similar trouble in the future. (GPT Image 2.5)


  1. Humans still do the planning. High-level planning remains a tiny fraction of agent output even at OpenAI, while delegation of longer, higher-level tasks keeps growing. So the human work product shifts toward the specification: goals, constraints, and acceptance criteria. Most enterprise tools have no UI to help people write a good specification. Build one before you build the agent.

  2. Expect a lower ratio without automated verification. Code has tests; a classifier can grade an experiment. Marketing copy, contracts, and designs have no equivalent oracle to declare an answer correct, so a human must review every output. Your intervention rate will be higher and your agent-to-human ratio lower. Without interfaces that make review and responsibility clear, users end up accountable for work they can barely inspect.

  3. Study what refuses to automate. OpenAI expects the tasks least amenable to automation to become the bottleneck. That residue is where your UX research budget should go: the planning, judgment, and repair work that stays with humans now defines their job.


My prediction: the 3:1 ratio reaches ordinary knowledge work around 2028, and the companies that cope will be the ones that built the supervision, budgeting, and specification interfaces in 2026, while their ratio was still 0.3:1. You have two years to adjust to a change spanning a full order of magnitude. As agents take over the work, the user interface becomes the place where humans steer, inspect, and intervene.


The human job will be to steer AI agents and untangle their mistakes as they take over all the things people used to do. (GPT Image 2.5)


Young Adults Feel AI-Dependent, but Addiction Scales Say Otherwise

Ashlee Milton and co-authors from the University of Minnesota, Microsoft, and the University of Washington surveyed 290 American young adults aged 18–25 who use AI chatbots at least weekly. 26% said they felt dependent on AI. But when these self-described dependents answered the field’s AI-dependence scales, which draw on clinical addiction frameworks, their responses fell on the “disagree” side of nearly every dependence item. The instruments describe a pathology; the users describe a habit.


(Everything is self-reported through an online questionnaire, so take the findings with a grain of salt. The central mismatch survives, though, since it compares two sets of self-reports from the same respondents.)


A habit that may be getting out of hand, even if it falls short of a clinical addiction: handing many of your learning activities, decisions, and emotional concerns over to AI. (GPT Image 2)


Users described dependence in terms of three behaviors: chronic use, chasing efficiency, and delegating learning, decisions, and even emotional processing to the bot. Dependent participants used chatbots several times a day, versus several times a week for everybody else.


Session length didn’t separate the groups at all, which should embarrass the California and New York legislators who tied companion-chatbot warnings to 3-hour sessions. Wrong metric. Frequency was the distinguishing measure in this study; duration revealed no such difference. The tell is reflex prompting: firing off a query before your own brain gets a turn.


As always, the politicians were wrong in their naïve attempts to regulate AI, basing restrictions on session length even though this research found that frequency distinguished the groups, while duration didn’t. (GPT Image 2)


Reflex prompting is a behavioral pattern to look out for if you worry that you’re developing AI dependence. (GPT Image 2)


Of the 15 design features studied, only one separated the self-described dependents from everybody else: validation, i.e., the chatbot supporting your thoughts and feelings (your idea is smart, your grievance justified, whether or not either judgment is true). Dependent participants said validation somewhat increased their chatbot use; the rest reported next to no effect.


Sycophancy pays, at least in engagement. The penalty I worry about comes later, in the form of skill atrophy and shrinking confidence. One participant voiced that fear: “maybe I don’t actually have as many skills as I think I do.”


I’m becoming increasingly convinced that sycophancy is an AI dark pattern, even if it arose innocently enough from reinforcement learning from human feedback (RLHF): users tend to click the “like” button more often for AI answers that flatter them. (GPT Image 2)


Dependence by itself is harmless. Nobody frets about forklift dependence among warehouse workers, and AI is a forklift for the mind. Eyesight doesn’t atrophy from wearing glasses; thinking atrophies from disuse. So cut the indiscriminate flattery, track how often users reach for the bot, and audit yourself: if AI is your first move on every task, keep a few tasks for your own brain.


Don’t always ask AI first. Give the old brain a little workout from time to time. (GPT Image 2)


Every Action Needs an Emergency Exit

Users explore boldly only when mistakes are cheap. Give them undo, and they’ll poke at every feature and learn your product by doing; withhold it, and they’ll cling to the three functions they already trust. My usability heuristics demanded a clearly marked emergency exit back in 1994, and undo remains the best exit ever built: one command that converts terror into curiosity.


Undo fixes errors after the fact, but its bigger job is to prevent fear before the fact. So make it universal, support undoing multiple actions, and extend it to AI agents, which take bigger actions on the user’s behalf and thus need bigger parachutes. Never punish curiosity.


Irreversible actions, admired from a comfortable altitude. (GPT Image 2)


Sad Suno

This week, Suno released version 6. I think this is the first time I’ve felt that a new AI model was worse than the previous release. A step backward in music quality! I suspect the cause is Suno’s embrace of the legacy music industry, which has pulled its output toward the slop that dominates the charts. The drop in quality is especially pronounced for niche tastes such as opera and ragtime.


Suno 6 is the first new AI model I’ve encountered that’s worse than the version it replaces. (GPT Image 2.5)


When I made an operatic aria about direct manipulation with Suno 5, I felt that it was almost there. And my progress indicator ragtime was also pretty good (though the animations in that video are rather primitive). I had hoped that a year’s progress in AI would let me do better now. Instead, I find myself nostalgic for last year’s music quality.


We’ll have to hope Chinese AI rides to the rescue of indie music creators, because I expect the Chinese models to remain independent of the record labels. For creators with eclectic tastes, that independence may be the more promising route to better music.


I made a decent operatic song with Suno 5, but Suno 6 seems to have lost its touch for anything beyond chart music. I suspect it’s now too cozy with the legacy record labels to serve independent creators with eclectic tastes. (GPT Image 2.5)


Anthropic’s 3 AI Economy Scenarios Forget the Robots and Underestimate the Bureaucrats

Economists at the Anthropic Institute have modeled AI’s effect on the US economy through 2030 under three scenarios. In the extreme scenario, GDP ends up 32% above trend, and nearly 1 in 5 knowledge workers is unemployed. The model is sound, but it stops where the real transformation begins: it contains no robots, and it runs on a 4-year timetable that enterprises and city halls can’t meet.


3 Scenarios, Zero Probabilities

Anton Korinek and co-authors at the Anthropic Institute (co-author Chad Jones literally wrote the textbook on economic growth) have published Economic Scenarios for Transformative AI (57-page PDF), together with an interactive scenario explorer. The model splits the US workforce into two islands: cognitive occupations (management, professional, sales, and office work, accounting for 62% of employment), where AI automates or augments tasks, and everybody else (construction workers, nursing aides, electricians, and truckers), whose work the model assumes AI leaves untouched.


Three settings of the dials produce three worlds in 2030:


The 3 Scenarios in 2030 Scenario GDP vs. No-AI Trend Growth in 2030 GDP per Capita in 2030 Cognitive Unemployment Modest +1.6% 2.4% per year $97,000 2.9% Substantial +8.3% 5.4% per year $104,000 4.5% Extreme +32% 15% per year $127,000 18%

I calculated the per-capita column myself: the scenario explorer’s 2030 GDP figures ($34.1, $36.3, and $44.4 trillion, stated in 2025 dollars, with no inflation baked in) divided by a projected 2030 population of 350 million. For comparison, 2025 GDP per capita was about $90,000 ($30.8 trillion across 342 million people), and the no-AI path lands near $96,000 in 2030.


The extreme scenario thus makes Americans 1/3 richer in 4 years; the modest one adds a rounding error. (A 1/3 gain is pleasant, but true transformation means becoming several times richer, and that won’t happen until 2040.)


The table shows how the scenarios would play out in the United States, because that’s the focus of the new report. I expect similar results in China, South Korea, Japan, Singapore, the Emirates, Switzerland, and any other country that goes all-in on AI.


(Of course, GDP per capita won’t reach American levels in only 4 years in countries that are currently far behind, such as China, which currently sits at about $32,000 on a purchasing power parity basis. But if the US gains 32%, I consider a 50% gain realistic for China, since it has one of the world’s most AI-pilled populations. That’s why any policy proposal that depends on the Chinese government condemning its citizens to remain at a third of American living standards without AI is pure fantasy.)


In the extreme world, cognitive wages sit 12% below their no-AI path while all other wages sit 34% above it. Plumbers become scarce. Hold that thought.


The authors also surveyed 10,980 US adults (Morning Consult, August 2026), and the median respondent’s expectations fall almost exactly in the middle, matching the substantial scenario. My favorite finding: cognitive workers expect AI to master 71% of the listed tasks by 2030 but only 27% of their own work. Everybody’s job is automatable, except one’s own.


The paper assigns no probabilities: the authors present the scenarios as possibilities to compare and refuse to pick one. I’m less shy. My best guess: an economy matching the modest scenario by 2030 and the extreme scenario by 2040. Here are my two reasons.


Only the Extreme Scenario Is Transformative AI

A year ago, I summarized the economists’ NBER workshop on transformative AI: TAI means productivity growth accelerating by at least 5x. The extreme scenario, at 15% growth against a 2% baseline, clears that bar with room to spare. The substantial scenario, at 2.7x, falls short. It would deliver a good decade, with the transformation still ahead.


Pascual Restrepo’s model in that workshop had labor’s share of income heading toward zero. The new paper’s drop from 60% to 45% in 4 years is the first leg of that descent. The authors calculate that compensating cognitive workers for their losses would take a transfer of 9% of GDP, roughly the size of Social Security and Medicare combined. Their own verdict on whether such transfers happen: historically, they mostly don’t.


Problem 1: The Plumber’s Raise Is a Robot-Free Artifact

The authors confess the omission on page 5: “We also consider only the effects of AI on cognitive tasks, and do not allow for rapid advances in robotics and their effects on physical tasks in the coming years.” That’s why they stop at 2030.


The report’s projection of 34% income gains by 2030 for trades such as plumbers is plausible. (I already paid a fortune to get a leak fixed last week, because skilled labor is scarce in Silicon Valley.) But robots will take over soon enough. (GPT Image 2.5)


But the second island is sinking too. I expect humanoid robots like Optimus to do to physical work what LLMs are doing to office work: take over surgery, nursing, and plumbing. Since a robot costs about the same per hour in Ohio as in Guangdong, robots should also bring manufacturing and other physical industries back to the US.


Today, Optimus is an in-house workforce of about 1,000 units with zero external sales. Small numbers at the foot of an S-curve have fooled forecasters before.


Cheaper, more accurate, and able to work 24/7. I predict that robots will perform most manual work in about 10 years. (GPT Image 2.5)


Robots break the model in two places. The 34% raise for the “untouched” occupations evaporates. And the model’s escape hatch (laid-off analysts retrain as electricians) closes. GDP gains climb well past 32%, but cognitive workers must find another route: start companies, direct the machines, or own some of the capital that collects 55% of income in the model’s extreme scenario.


Problem 2: Tesla Builds Faster Than City Hall Changes

Superintelligence by 2030 now seems likely to me, but revolutionary economic change takes longer than 4 years. The economy’s biggest institutions (enterprises and governments) move at the pace of bureaucrawl.


Developing superintelligence will be the smallest challenge in transforming the economy. Overcoming bureaucrawl will take longer. (GPT Image 2.5)


Consider the Cybercab, which entered production in April 2026 and began carrying paying passengers in Austin on September 4, 2026. Elon Musk’s claim at the 2024 unveiling: 20 cents per mile to operate at scale, or 30–40 cents per mile including taxes and insurance. Compare the audited 2024 numbers in the National Transit Database: San Francisco’s Muni (PDF) spent $3.61 per passenger-mile across all its modes ($3.04 on buses), and BART (PDF) spent $1.18.


Apply the customary discount to Musk’s optimism and double the upper end of his all-in estimate to 80 cents per mile with a single rider aboard. The Cybercab still costs a third less than BART and less than a quarter as much as Muni, with door-to-door service, no transfers, and no waiting.


Once 50 million of them roam the US, every municipal transit system ought to be retired. It won’t be. Transit agencies have unions, bond covenants, federal funding formulas, and council members who ran on saving Route 12. Tesla will build 50 million cars faster than 19,000 city councils change their minds.


Tiny AI-driven cars that take passengers from door to door are superior to public transit and far cheaper for taxpayers. But politicians will take decades to make the change. (GPT Image 2.5)


History rhymes. As Paul David described in his classic 1990 paper, the electric dynamo arrived in the 1880s, but factory productivity didn’t jump until the 1920s, once factories had been rebuilt around it. My TAI article made the same point about workflows: bolting AI onto today’s process yields 2x; redesigning the process around AI yields 10x, and the 1,000x company arrives around 2045.


So I think the extreme scenario’s numbers are right, on a timetable stretching from 2030 to 2045. The consolation: the 18% unemployment spike spreads into a gradual reallocation of work that the labor market can absorb. The cost: citizens ride half-empty buses for a decade longer than the technology requires.


History shows that electricity by itself did little to improve factories: manufacturers initially installed it without redesigning their processes to exploit the possibilities opened by electric motors. Several decades passed before factories were rebuilt to capture those gains. A similar delay will likely happen with AI, which requires redesigned workflows before profits can explode. Hopefully, we’ll learn from history and make the transformation faster this time around. (GPT Image 2.5)


Conclusion: Right Model, Wrong Clock, No Robots

Plan for a 15-year transition. If you work inside a large organization, expect the org chart to become the bottleneck as models advance. That’s why my career advice for anyone under 35 remains: join or found something small enough to change.


Watch one number: when humanoid robot shipments pass 1 million a year (my guess: 2029), retire the two-island map. Meanwhile, take the scenario explorer for a spin. Set every dial to extreme, then add 10 years and a robot.


The average enterprise company's executive team meeting in 2030. (GPT Image 2.5)


Alice and Zimo explore the three scenarios for our economic future, rendered in a Cinematic Corporate CGI style that felt fitting for the topic. (GPT Image 2.5)


Reflective Design: Sometimes You Should Make Users Think

Reflective design builds deliberate pauses into the user interface, prompting people to notice and question their own behavior and the values baked into the technology. Used sparingly at consequential decision points, it prevents autopilot errors and helps users give informed consent. Used everywhere, it degenerates into nagware.


Reflective design turns the interface into a mirror: step back and see your own choices, habits, and their consequences at full size. Real products deliver this moment through something as ordinary as a weekly usage report. The marble hall is optional; the design goal is identical. (GPT Image 2)


The Deliberate Exception to “Don’t Make Me Think”

My bestie Steve Krug gave our field its best-selling slogan in 2000: Don’t Make Me Think. As a default, he’s right. Reduce cognitive load, honor conventions, and let users glide through tasks on autopilot. But autopilot has a failure mode: people breeze past decisions that deserve conscious thought. Reflective design is the deliberate exception to Krug’s rule.


Sometimes we should put a stop to the “don’t think” design approach and give users a reason to pause and think before they act. (GPT Image 2)


Definition: Reflective design deliberately structures the user experience to bring unconscious behavior, assumptions, and embedded values to the user’s conscious attention, so that he or she can make a considered choice instead of a habitual one.

Everyday examples: Apple’s Screen Time and Google’s Digital Wellbeing dashboards confront users with their actual phone use. Banking apps summarize where the money went this month. Spotify’s annual Wrapped recap shows a year of listening habits (half the fun is discovering you’re not who you thought you were). And a well-placed “This will permanently delete 2,400 photos” dialog is reflective design at its most compact.


Reflective design can help users discover patterns in their own use. Once users see a pattern, they may decide that it suits them just fine. The computer should give them better data so they can make that decision for themselves. Some users will choose to keep their established patterns; others will choose to break them. (GPT Image 2)


A Name Rooted in Critical Theory

Phoebe Sengers and colleagues at Cornell University established the term in their paper “Reflective Design,” presented at the Critical Computing conference in Aarhus, Denmark, in 2005. Their argument: technologies silently perpetuate unconscious cultural assumptions, and designers should make a practice of “bringing unconscious aspects of experience to conscious awareness” for themselves and their users alike.


The intellectual roots reach back to Donald Schön’s 1983 book The Reflective Practitioner and to the participatory design tradition. Thus, the name means exactly what it says: design that provokes reflection, much as a mirror makes you stop and examine what you see.


Terminology overlap alert: Don Norman’s 2004 book Emotional Design uses “reflective” for the third of his three levels of cognitive processing (visceral, behavioral, reflective). Same word, different beast: Norman describes how people process products, whereas Sengers and colleagues prescribe what products should provoke. This article covers the second meaning, translated from critical theory into working UI.


Where a Mirror Moment Beats Frictionless Flow

These designed pauses are mirror moments: brief interface events that show users their own behavior and its consequences before they proceed. Mirror moments earn their keep in three situations:


  • Irreversible or costly actions. These include deleting data, wiring money, and granting an app microphone access. The fifth of my 10 usability heuristics (1994) is error prevention. A mirror moment extends error prevention to the decision itself: it catches the mistake the user doesn’t yet know he or she is making.

  • Slow-building patterns. A single scrolling session may seem unremarkable; 4 hours every day can reveal a habit worth examining. Dashboards and digests gather these otherwise invisible fragments of behavior into a visible pattern that users can act on.

  • Consent that should mean something. A privacy permission granted on autopilot gives users little protection. A crisp statement of what will be shared, and with whom, converts legal theater into an actual decision they can make with their eyes open.


Requests for user consent should explain the consequences of granting permission. When the stakes are substantial, the design should encourage users to reflect before they act. This is why the EU’s cookie consent rules are damaging: they train users to assume that consent requests are irrelevant and appear only to satisfy the bureaucracy. Consequently, users also fail to stop and consider dangerous permissions, granting them just as automatically as they accept cookies. (GPT Image 2)


The business case is unglamorous but real: fewer regretted actions mean fewer support tickets, refunds, chargebacks, and account closures. Regret is churn with a delay, and mirror moments are the cheapest churn insurance you can buy.


Reflection Curdles into Nagging

This pattern is easy to abuse, and products provide plenty of examples:


  • Friction sprayed everywhere. Interrogate every action, and users develop dialog blindness, dismissing a vital warning with the same reflex they use for a trivial one. (We trained them to do that. Congratulations, us.)

  • Moralizing. “Still scrolling? Everything OK?” Users are adults. Show them the data and skip the sermon: users who feel judged will close the mirror and keep the habit.

  • Fake reflection. Confirmshaming (“No thanks, I hate saving money”) wears reflection’s costume while doing manipulation’s work. That’s a dark pattern, and users increasingly recognize it as one.


The remedies follow directly: reserve mirror moments for consequential choices, gather less urgent observations into periodic digests, and keep the tone neutral. Never let the growth team write a reflective prompt, just as you wouldn’t let the fox design the henhouse.


7 Guidelines for Reflective Design

  1. Reserve mirror moments for consequential or irreversible actions. Payments above a threshold that matters to the user, data deletion, privacy permissions, and public posts qualify. Routine actions don’t need this extra pause.

  2. Prefer digests to interruptions. A weekly summary of spending or screen time respects the user’s flow while still revealing the pattern. Interrupt only when the user must consider the consequences of a decision being made right now.

  3. Show the data and let users judge. “You spent 22 hours in this app last week” informs; “That’s a lot!” lectures. Give users the number and leave the verdict to them.

  4. Make every reflective prompt dismissible in one action. Anyone who has already reflected and decided should sail through. Forced reflection is just friction with a diploma.

  5. Pair the mirror with a lever. Every insight needs an action users can reach with one click: set a limit, cancel the subscription, or revoke the permission. Reflection without recourse breeds resignation.


Add an easy option for users to act on the insights your design provides. (GPT Image 2)


  1. Never disguise persuasion as reflection. If the “reflective” prompt only ever nudges toward the company-preferred choice, it’s marketing. Users smell the difference, and trust doesn’t regrow quickly.

  2. Measure regret, then decide where mirrors belong. Track undo rates, support contacts, refunds, and deletion reversals. My rule of thumb: any action whose undo rate exceeds 1% is a candidate for a mirror moment; test whether adding one cuts regret without tanking completion.


Reflective design will never be the main course of usability. As a rough design target, I’d keep 99% of interaction fast, conventional, and thought-free, following Krug’s principle. But the remaining 1% carries a disproportionate share of the consequences: the deleted archive, the regretted purchase, the permission that leaked a contact list.


Put one well-polished mirror at each of those moments, keep it neutral, and give the user a way to act on what he or she sees. A hallway of mirrors is a funhouse. One good mirror, in the right spot, is a service.


Mini Maps: The 1980 Arcade Trick That Still Rescues Lost Users

A mini map is a small overview of a large workspace, shown beside the detail view with a marker for the user’s current position. Arcade games invented the pattern 46 years ago; code editors, design canvases, and strategy games still rely on it. Use mini maps only for big spaces, keep them abstract, and let users switch them off.


If Theseus had carried a mini map, Ariadne could have kept her thread. (GPT Image 2)


Definition: A mini map is a persistent, miniature representation of an entire information space, displayed alongside the detail view, with a visible indicator (usually a rectangle) marking the part of the space the user currently sees.

The physical ancestor is the shopping-mall directory with its you-are-here arrow. The digital version improves on the mall in two ways: the arrow moves as you move, and tapping anywhere on the map teleports you to that location.


Born in the Arcade, Raised in the Code Editor

The first mini map appeared in Namco’s arcade game Rally-X (1980), which displayed its scrolling maze on a small “radar” beside the play area: one dot for your car, red dots for the enemy cars, and yellow dots for the flags you had to collect. (The radar omitted the maze walls, even though they constrained where you could drive. Impressive engineering for 1980, and an accidental usability lesson: an overview that hides the obstacles is half an overview.)


Gamers called these displays radars throughout the 1980s. The term mini map took over after real-time strategy games such as Dune II (1992) planted a permanent world map in the screen corner and made it the genre’s command center. The term is refreshingly literal: a miniature map of the whole space.


Productivity software adopted the pattern once documents outgrew screens. Adobe Photoshop 4.0 shipped its Navigator palette in November 1996: an overview thumbnail with a draggable viewport frame. Sublime Text (2008) popularized the code minimap: a compressed, unreadable, and yet oddly useful column of pixels showing the silhouette of an entire source file. The pattern is now standard in Visual Studio Code and many other editors. Infinite-canvas tools such as Figma and Miro include one because their workspaces have no fixed edges at all.


Overviews Cure Desert Fog

Zoomable and scrolling interfaces suffer from what Susanne Jul and my erstwhile Bellcore colleague George Furnas named desert fog in 1998: the user pans or zooms into a region that contains no cues, so every direction looks equally unpromising. Navigation stalls. A mini map is a standing cure: one view on the screen always contains the whole world, so the fog never becomes total.


Ben Shneiderman compressed the deeper principle into his 1996 visual information-seeking mantra: “Overview first, zoom and filter, then details-on-demand.” A mini map delivers the overview without forcing users to zoom out and abandon their work. It answers three questions at a glance: Where am I? What else exists? How do I get there? And the answer to the third is a single click, which beats scrolling across 40 screenfuls of canvas.


Andy Cockburn and co-authors’ 2009 research review of overview+detail interfaces catalogs decades of studies showing that users value overviews for orientation and long-distance jumps. Those benefits come at a price: screen space and divided attention.


The Orientation Tax

That price is real. Kasper Hornbæk (University of Copenhagen) and co-authors compared map interfaces with and without an overview in a study with 32 users. A whopping 80% preferred having the overview, saying it helped them keep track of their position. But participants were faster without it on one of the two maps, and task accuracy didn’t differ at all.


Every glance at an overview is a round trip: the eyes leave the work, refocus on the thumbnail, translate between the two scales, and return. That’s the orientation tax, and on well-structured content the tax can exceed the benefit.


Give users the option to close the mini map. Preference and performance point in different directions, and 1 user in 5 in this study didn’t prefer having the overview. Users who want to reclaim those pixels should be able to do so with a familiar close box.


Other failure modes recur across products. Postage-stamp fidelity: shrinking a 50,000-pixel canvas into a 150-pixel thumbnail produces gray mush, so the map must summarize the space through landmarks, shapes, and color coding. A shrunken copy won’t do.


Screen theft: on a phone, a corner overlay can cover the very content users are trying to navigate. Tunnel vision: gamers have griped for years about players who stare at the dots instead of the richly rendered world. The business-software equivalent is a product so disorienting that users can only navigate via the map. Fix the information architecture first.


8 Design Guidelines for Mini Maps

  1. Reserve mini maps for big spaces. My rule of thumb: if users can see the whole workspace within 4–5 screenfuls or 2 zoom steps, a scrollbar does the job and the map is clutter.

  2. Always show the viewport indicator. The rectangle marking the current view is the you-are-here arrow. Without it, the mini map can show users what exists while leaving them unable to answer the more urgent question: “Where am I?”

  3. Couple the two views tightly, in both directions. Clicking or dragging on the map must move the main view instantly, and panning the main view must move the rectangle in turn. A map that responds in only one direction feels broken.

  4. Show the structure of the space. Use landmarks, region shapes, and color coding to make the overview legible at a glance. Code minimaps work because source code has a recognizable silhouette even at 2% scale.

  5. Overlay the things users hunt for. Search hits, errors, comments, selected objects, and collaborators’ positions turn the map into a useful display of where the items that matter are clustered.

  6. Keep it small and peripheral. Place it in a corner, taking up at most about 1/10 of the screen area (again, my rule of thumb), and use translucency or auto-hide on small screens.

  7. Let users dismiss it, and remember the choice. Hornbæk’s study shows that preference for an overview can coexist with slower task performance; forcing it on users may tax every session. Use the standard close control: a close box or X in the corner.

  8. Never make the map the only road. Keyboard navigation, scrollbars, zoom, and search must reach everything the map reaches, because 3-pixel dots are hostile targets for users with motor or vision impairments.


Ariadne gave Theseus the thread that guided him out of the Labyrinth. Modern products are labyrinths too, whether they hold 200,000 lines of code, an endless whiteboard, or a sprawling strategy-game world. A mini map is that thread rendered in pixels, giving users a way to find their bearings and return to familiar ground.


It costs you a corner of the screen; a lost user costs you the session, and often the customer. Rally-X put cars and flags on its radar back in 1980, but left out the maze walls. Put your walls on the map.


In the Greek myth, the hero Theseus had to fight the fearsome Minotaur (half man, half bull), who lived in a labyrinth. Nobody who had entered that labyrinth before him had ever found a way back out. But Ariadne gave Theseus a thread, which he trailed behind him as he wandered through its passages. After defeating the Minotaur, Theseus easily retraced his steps to the exit and eventually became king of Athens. (Muse Image)


GPT Image 2.5

OpenAI released a minor upgrade to its image model, which is now numbered 2.5 instead of 2.0. I don’t think the improvements warrant a full half point on the version number; version 2.1 would have been a more reasonable name.


That said, GPT Image 2.0 was already the world’s best image model, and 2.5 is indeed a little better. But the biggest improvement may be how it’s integrated with OpenAI’s much improved agentic reasoning model, GPT-6 Astra running with Ultra reasoning effort in Work mode. (Note the terrible usability of having to understand “Work” vs. “Chat” modes of the same AI model, on top of knowing which reasoning level is needed for which outcome.)


When creating images with these settings, I’ve noticed that the agent inspects the images the image model produces and redraws any that have a flaw. As an example, here’s the first version of an image I wanted for possible use with this newsletter’s lead item about employers preferring experience over diploma-certified skills:


First attempt from GPT Image 2.5.


In this image, the compass represents the judgment that comes with experience, while the stack of diplomas represents a laundry list of traditional skills. My story reports that employers now value the former more than the latter, but the image shows the right pan outweighing the left. Wrong!


I can’t tell you how many images with balance-scale metaphors I’ve rejected over the years because every image model from Midjourney to Nano Banana Pro has shown the wrong side of the scales as the heavier one. GPT Image 2.5 is no exception.


Agentic AI to the rescue! While generating the images, GPT-6 Astra Work Ultra (a mouthful) recognized the error of its own image model and decided to regenerate the image. I didn’t lift a finger. I had stepped out for a spin on the treadmill, and the improved image was waiting for me when I returned to the computer:


Corrected image from GPT Image 2.5: the left weighing pan is now lower, as befits the point that the compass (representing judgment) now weighs more than skills in hiring decisions. I could have wished for a greater imbalance, but at least it’s no longer misleading. Thank you, GPT-6 Astra agentic Work mode with Ultra reasoning! I might almost forgive you for your name.

 

Final Thought of the Day

If you don’t have Ariadne to lend you a thread. (GPT Image 2)

Top Past Articles
bottom of page