top of page

AI Turns Competence into a Commodity: Evidence from 2.26 Million Freelance Contracts

  • Writer: Jakob Nielsen
    Jakob Nielsen
  • 52 minutes ago
  • 29 min read
Summary: AI narrows skill gaps by helping weak performers the most. Market data shows employers reacting: hiring in AI-exposed job categories now weights human capital 7.8% less and price 1.1% more and the demand premium for highly skilled workers is shrinking. Competence has become a commodity, so sell judgment. Agentic AI will reverse the equalization and widen skill gaps again, this time around the skill of deciding what to build and picking the best AI output.

AI is a forklift for the mind. A real forklift lets a scrawny warehouse worker and a bodybuilder move identical 1,000-kg pallets: once the machine does the lifting, muscle stops predicting output. AI does the same for knowledge work, which is why the weakest performers gain the most from it. I’ve made this argument since 2023 (see, for example, AI Makes Happy Geeks), and every controlled experiment since has piled more weight onto the pallet.

 

One of my oldest metaphors: AI is a forklift for the mind. It helps everybody lift cognitive burdens and thus narrows the skill gaps between the most and least talented knowledge workers, just as a real forklift narrows the gap between muscular and weak warehouse workers. (Muse Image)

 

Experiments measure the supply side: what AI does to the work. But what happened to demand? Do the people who buy work agree that competence has become a commodity? A new study from UCLA’s Anderson School of Management delivers the first hard market data, drawn from 2.26 million freelance contracts, and the answer is yes: in AI-exposed occupations, clients pay less attention to skills, credentials, and reputation, and more attention to a single number, the price. In this article, I dissect the new market evidence, recap the experiments that explain it, draw the consequences for UX careers, and then argue (speculation, plainly labeled) that the next generation of AI will reverse the equalization and hand the advantage back to the most talented. It’s a long article, but then again, it’s your paycheck.


Credentials and reputation now carry less hiring weight, according to the analysis of more than 2 million contracts. (GPT Image 2)

 

The Market Has Noticed: Data from 2.26 Million Contracts

Auyon Siddiq and Niuniu Zhang of UCLA’s Anderson School of Management pulled the hiring records of Upwork, the large online labor market where clients contract freelancers for short jobs in writing, design, coding, accounting, translation, and customer support (SSRN working paper, 2026). Their panel tracks 49,610 freelancers who were active before ChatGPT arrived, observed quarterly from January 2021 through March 2026: 21 quarters and 2.26 million completed contracts spanning the AI shock of November 30, 2022. No survey, no vibes: this is revealed preference, recorded 2.26 million times, with money attached.

 

The research question is more nuanced than the usual “did AI kill freelance jobs?” We already know demand fell in exposed categories: Xiang Hui, Oren Reshef, and Luofeng Zhou documented the drop in 2024, and Ozge Demirci and colleagues confirmed it from the job-posting side in 2025. Siddiq and Zhang ask a subtler question: did AI change what clients look at when choosing among freelancers? If AI commoditizes labor, skill signals should lose hiring weight and the price tag should gain it. Spoiler: both happened, on schedule, in the AI-exposed categories only, and kept intensifying for 3 years.


As AI equalizes the quality of the results, clients increasingly choose by price instead of by who has the fanciest credentials and the longest experience. (Muse Image)

 

What Clients See: 3 Bundles of Human Capital, Plus a Price Tag

Definition: Human capital is economists’ term for the stock of skills, education, and experience that makes a worker productive; the concept helped win Gary Becker the 1992 economics Nobel.

 

A client can’t observe a freelancer’s productivity before hiring; he or she sees only signals. The study sorts everything visible on an Upwork profile into 3 bundles of human capital signals, plus one factor that isn’t human capital at all:

 

  • Self-presentation: the freelancer’s chosen title (“Experienced iOS Developer”), written description, and skill tags: what workers say about themselves.

  • Credentials: degrees, institutions, employment history, and portfolio projects. External validation, and costly to change. (You can rewrite your skill tags over coffee; you can’t rewrite your diploma.)

  • Reputation: prior contracts, client ratings, written feedback, and platform badges such as “Top Rated,” all generated inside the market by actual transactions.

  • Price: the posted hourly rate. One number, no prose.


The 3 human capital bundles a client can inspect before hiring: how workers describe themselves, what their credentials certify, and what past clients said. The fourth factor, price, is the one AI is teaching buyers to favor. (GPT Image 2)

 

Signals saturate this market: 97% of the freelancers list education, 83% show a portfolio, and 98% carry at least one client rating. Demand, by contrast, is feast or famine: the mean freelancer lands 2.17 contracts per quarter, but the median freelancer lands zero. Abundant signals chasing scarce buyers: exactly the market where hiring signals should matter enormously. Which makes it all the more striking when buyers start ignoring them.

 

How Do You Weigh a Resume? With 16 Regressions and a Nobel Idea

The awkward fact about worker profiles is that they’re mostly prose, and prose resists spreadsheets. The traditional fix hand-codes a few variables (years of education, true/false for “knows Python”) and discards the rest. Siddiq and Zhang instead convert each text bundle into an embedding: a list of 384 numbers, produced by a language model, positioned so that texts with similar meanings land near each other. Whole profiles enter the statistical machinery, with no researcher deciding in advance which words matter.

 

Next, they predict each freelancer’s quarterly demand from all 4 blocks, quarter by quarter, with cross-validation so the model is always judged on freelancers it never saw during training. (Profiles plus price explain roughly 1/3 of the variation in demand; the rest is timing, luck, and everything a profile can’t show.)

 

Then comes the Nobel idea. To measure how much each block matters, the authors compute Shapley values, a credit-assignment method that Lloyd Shapley invented in 1953 for cooperative games (he later collected an economics Nobel, and his method now powers most of explainable AI). The recipe: fit all 16 possible models, one per subset of the 4 blocks, and average each block’s marginal contribution. The result: for every freelancer in every quarter, predicted demand splits exactly into a self-presentation share, a credentials share, a reputation share, and a price share.

 

Finally, the causal design. The authors grade each of Upwork’s 102 job subcategories with the AI exposure scores from Tyna Eloundou and colleagues’ “GPTs are GPTs” (published in Science in 2024): the share of an occupation’s tasks that an LLM alone could complete twice as fast at equal quality. Translation tops the scale at 0.80; photography sits near zero. (That would be zero AI risk for the freelance job of “portrait of my beloved dog in my apartment,” not for “give me a photo of a cute dog,” which AI does perfectly without needing to go on location.)

 

This is nobody’s dog, so I didn’t need to hire a freelancer for the shoot. Meta’s new Muse Image model whipped up this pampered pooch in a few seconds of rendering time.

 

The study then compares the trajectory of the 4 block weights in exposed vs. unexposed categories, before and after ChatGPT: a difference-in-differences design, the workhorse of modern policy evaluation. One detail deserves applause: because profiles are measured once (a March 2026 snapshot), any movement in the weights reflects a change in the market’s hiring rule, not a change in the profiles. The freelancers held still; the buyers moved.

 

The Findings: Skills Pay Less, Price Matters More

Start with the anchor. After ChatGPT, contract volume for freelancers in the most exposed job category fell 7.0% relative to unexposed categories, replicating the earlier demand studies. The news is how that decline was distributed across the hiring signals:

 

  • Self-presentation lost 2.8% of its importance for demand, credentials lost 2.4%, and reputation lost 2.9%. Combined, the human capital bundles lost 7.8% of their hiring weight.

  • Price gained 1.1%.


Every human capital coefficient is negative, the price coefficient is positive, and all clear the 1% significance bar. Note the breadth: the discount hit not only self-descriptions (which AI made cheap to fake) but also verified diplomas, employment histories, and 5-star ratings earned through real transactions. Buyers stopped caring even about proof.

 

Worse (or better, depending on your side of the trade), the effects steepen over time instead of fading:

 

Across 49,610 freelancers and 2.26 million contracts, hiring in the most AI-exposed job categories shifted weight away from human capital (7.8% less important) and toward price (1.1% more important) after ChatGPT launched, while contract volume in those categories fell 7%. All these trends were stronger in recent quarters than in the full period Siddiq and Zhang analyzed.

 

More than 3 years after the shock, the lines still point the wrong way for credential holders: a repricing in progress, not a blip being absorbed. Pre-ChatGPT trends were flat for all 3 human capital bundles, so the divergence starts exactly when the technology arrived. We now have a credential discount, the markdown that AI-exposed markets apply to diplomas, portfolios, and ratings. The credential discount deepened every year of the study.

 

Credentials are worth less and less. (GPT Image 2)

 

The Commoditization Tests: Premiums Shrink and Demand Flows Downmarket

A skeptic could accept everything above and still resist the commoditization story. Maybe clients started treating a high price as a quality signal (economists have modeled price-as-quality since Paul Milgrom and John Roberts in 1986), which would also make price more predictive of hiring. So Siddiq and Zhang ran two additional tests, and this is where the paper shines.

 

Test 1: Does the skill premium shrink? Rank freelancers by the pre-AI strength of their human capital signals and compare the top half to the bottom half. Before ChatGPT, strong-signal freelancers enjoyed a fat demand premium. After ChatGPT, that premium compressed by 6.2% in the most exposed categories, and by 10.3% in the study’s final 4 quarters. The workers with the most impressive signals lost the most differentiation, precisely where AI hit hardest.


Before AI, strong human capital signals earned freelancers a hefty demand premium over weaker rivals. After ChatGPT, that premium compressed by 6.2% in the most AI-exposed categories, and by 10.3% in the study’s final year. (Data: Siddiq & Zhang, 2026.)

 

Test 2: Where does the demand go? Compare higher-priced freelancers to cheaper ones. If clients had started reading price as a quality signal, demand should have shifted upmarket toward expensive workers. It went the other way: demand shifted toward cheaper freelancers, by 3.2% in the pooled estimate and an accelerating 7.9% in the final 4 quarters. (The top-skill and top-price groups barely overlap, correlating at only 0.07, so the second test isn’t the first one wearing a cheaper suit.) Price sensitivity won. Quality signaling lost.

 

Together, the two tests match the theory of labor commoditization proposed by Masao Fukui of Boston University with Emi Nakamura and Jón Steinsson of UC Berkeley (Nakamura holds the John Bates Clark Medal, economics’ junior Nobel, so this is no fringe framework). The mechanism: AI compresses quality differences between workers’ outputs, workers become more substitutable in clients’ eyes, and hiring collapses onto the one dimension that still separates candidates. The price.


The commoditization mechanism in one picture: AI raises weak performers’ output more than strong performers’, the quality gap narrows, and workers become more substitutable in clients’ eyes. Once outputs converge, price is the differentiator left standing. (GPT Image 2)

 

How solid is all this? The authors shuffled the AI exposure scores randomly across the 102 job subcategories 2,000 times and re-ran everything. Not one shuffle reproduced the observed pattern of 3 human capital declines plus a price increase (p < 0.0005). The result is glued to actual AI exposure, not to a quirk of the categories.

 

Time for the methodological cold shower, because no observational study escapes one. Profiles come from a single March 2026 snapshot, so the study assumes credentials and prices were stable (defensible: diplomas don’t fluctuate, and posted rates vary 6 times more across workers than within them, but still an assumption). Freelancers who deleted their accounts are invisible, which makes the estimates conservative, since the biggest losers left the sample. The study observes hiring outcomes, not the proposals and search rankings behind them. And it covers one platform: a freelance marketplace, not payroll employment. But the paper argues, and I agree, that the same screening logic governs resumes, references, and salary negotiations everywhere; the marketplace just renders it measurable. For market data, this is about as clean as the genre gets, and the robustness checks (3 prediction models, 2 embedding models, same story) beat the evidentiary standard of every AI hot take you’ll read this week.

 

So the demand side has spoken. To see why clients behave this way, and why they’re broadly right to, turn to the supply side: the experiments showing what AI does to the quality gap between workers.

 

The Lab Evidence: AI Lifts the Bottom Harder Than the Top

The skill-gap finding isn’t one study; it’s a research program spanning writing, programming, consulting, customer service, product innovation, and law, with data running from small 2023 lab tasks to 2026 field deployments at global scale. The headline numbers:

 

AI helps everybody, but extensive evidence shows that it lifts the least-skilled workers the most, raising the talent floor. (GPT Image 2)

 

  • Business writing. Shakked Noy and Whitney Zhang of MIT had 453 college-educated professionals write realistic documents such as press releases and analysis memos (Science, 2023). ChatGPT users finished 40% faster with 18% higher quality, as judged by blind graders. The biggest quality jumps went to the participants who wrote worst without AI, and the link between unaided skill and final grade weakened sharply once ChatGPT entered the picture.

  • Customer support. Erik Brynjolfsson, Danielle Li, and Lindsey Raymond studied more than 5,000 support agents at a Fortune 500 software firm whose AI assistant was trained on transcripts from top agents (NBER, 2023, since published in the Quarterly Journal of Economics). Productivity rose 14% on average but 35% for novice and low-skill agents, with next to nothing for the most experienced. A 2026 field experiment with roughly 6,000 agents at Alibaba’s Taobao marketplace replicated the shape (arXiv, 2026): low performers improved the most on both speed and quality, while top performers gained little speed and even slipped on quality, partly because the assistant’s suggestions ran below their own standards and partly because they multitasked more.

  • Management consulting. Fabrizio Dell’Acqua and colleagues at Harvard Business School ran 758 BCG consultants through 18 realistic tasks (SSRN, 2023, since published in Organization Science). On tasks within AI’s capability frontier, quality rose roughly 40%. Consultants in the bottom half of the skill distribution improved 43%; the top half improved 17%.

  • Product innovation. Fabrizio Dell’Acqua, Karim Lakhani, and colleagues ran a preregistered field experiment with 776 professionals at Procter & Gamble working on real new-product challenges (NBER, 2025). Individuals with AI matched the performance of 2-person teams without it, and AI dissolved expertise boundaries: R&D staff produced commercially balanced proposals and commercial staff produced technically balanced ones. The skill gap narrowed not just within specialties but across them.

  • Programming. Sida Peng and colleagues at GitHub and Microsoft timed 95 developers building an HTTP server (arXiv, 2023). The GitHub Copilot group finished 56% faster, with the largest speedups among less-experienced developers. The result then scaled: 3 field experiments at Microsoft, Accenture, and a Fortune 100 manufacturer gave 4,867 working developers Copilot during their day jobs (Management Science, 2026); completed tasks rose 26%, again with the highest adoption and the largest gains among the least experienced.

  • Law. Jonathan Choi and Daniel Schwarcz gave University of Minnesota law students GPT-4 on realistic law-school exams (73 Journal of Legal Education, 2025). Students at the bottom of the class saw huge gains; students at the top actually declined. (And file away a second finding for later: with good prompting, GPT-4 alone outscored both the average student and the average AI-assisted student.)

 

All 9 studies used randomized or staggered rollouts of AI access and measured performance with blind grading or objective output counts, so these are causal effects, unlike the self-reported enthusiasm that passes for evidence on social media.

 

One boundary condition deserves airtime before anyone extrapolates that table to infinity. In a 2025 randomized trial by the research group METR (study report), 16 veteran open-source developers worked on their own mature codebases, projects they had maintained for 5 years on average, with each task randomly assigned to allow or forbid AI tools. With AI, they finished 19% slower, while believing afterward that AI had sped them up by 20%. The forklift lifts from below: hand one to somebody already standing at the top of his or her own staircase, and it mostly gets in the way, while feeling helpful the whole time. (That 39-point gap between perceived and measured productivity is also a standing caution against every self-reported AI statistic you will ever read.)

 

Different professions, different countries, different AI models, different years: same shape. By 2024, I considered the narrowing of skill gaps settled science for current AI on production tasks. Hold onto that qualifier, “production tasks.” It returns with a vengeance below. And notice how neatly the lab record explains the market record: Upwork clients down-weight skill signals because AI has genuinely made skill a weaker predictor of output quality. The buyers are simply arbitraging a real change.

 

Why the AI “Forklift” Helps the Weak the Most

Consider what forklifts did to warehouses. (Clark built the first recognizable forklift around 1917, and World War II logistics made pallets and forklifts the global standard.) Before the machine, a warehouse paid a premium for strong backs. After the machine, strong and weak moved the same pallets at the same speed, and physical strength dropped out of the wage equation. The forklift didn’t make anyone stronger. It made strength irrelevant to the task of moving boxes.

 

Why does the cognitive version lift from the bottom? Five mechanisms, all visible in the studies above:

 

  1. AI performs at a roughly fixed level, no matter who’s prompting. For everyday professional tasks, a frontier model writes at perhaps the 80th percentile of the profession. Work the math: Bob, a 30th-percentile writer, accepts the AI draft and jumps 50 percentile points. Alice, a 95th-percentile writer, would drop 15 points by shipping the raw draft, so she edits heavily and gains mostly time. The model’s fixed competence acts as a rising floor and a magnetic middle.

  2. AI bottles the tacit knowledge of top performers. The customer-support study makes the mechanism explicit: the assistant was trained on the best agents’ transcripts, so novices got the firm’s stars whispering in their ears all day. Experts gained nothing from hearing their own advice repeated back to them.

  3. Recognition outruns production. Most professionals can recognize quality they can’t yet produce. (Human memory judges far better than it generates, which is also why my usability heuristic “recognition rather than recall” works.) AI converts recognition skill into production skill: a junior designer who can merely tell the good option from the bad ones can now ship the good one.

  4. Current AI supplies mechanics rather than judgment. Grammar, structure, boilerplate, standard frameworks: exactly the ingredients weak performers lack. Top performers’ edge sits in judgment, taste, and originality, which current models supply in thinner doses. Patch the weak spot of the weak, and the gap closes from below.

  5. Grading scales have ceilings, and experts start near them. On a 0–10 rubric, an expert scoring 8.5 has 1.5 points of headroom; a 4.0 performer has 6. The narrowing is real, but ceiling effects compress it further at the top, so treat the exact percentages with mild suspicion.


A Central Bank Runs a Clean Experiment

Aleš Maršál of the National Bank of Slovakia and Patryk Perkowski of Yeshiva University produced the tidiest supply-side test yet: a preregistered, randomized field experiment inside Slovakia’s central bank (working paper, 2025; the authors also wrote a short summary at VoxEU). In June 2024, 101 of the bank’s roughly 1,100 employees completed a 2-hour battery of genuine workplace tasks for real money (base pay plus a performance bonus).

 

The task design is the paper’s quiet masterstroke. Instead of studying a single job, the authors sampled the full task spectrum that labor economists have used since David Autor and colleagues formalized it in 2003: routine and non-routine cognitive tasks (proofread a sloppy paragraph; invent novel uses for a paperclip), routine and non-routine manual tasks (data entry; explain how to un-jam the office printer), plus analytical, leadership, and communication tasks, 14 generalist tasks of about 4 minutes each. Then came 2 specialist tasks of 30 minutes each: economists decomposed GDP and fit a forecast, IT staff wrote and benchmarked code, and payments staff handled a simulated breach of the settlement system. (Central bankers steer the money supply for roughly 350 million Europeans, but the jammed printer remains undefeated as an office challenge.)

 

Every participant did one task of each pair with OpenAI’s GPT-4o and its twin without, with the IT department blocking the tool for control tasks and one of the authors sitting in the room to keep everyone honest. Employees used GPT-4o on 94% of eligible tasks and 0% of control tasks. Try getting 94% uptake on any other corporate software rollout! Grading was blind, on 0–10 rubrics, with industry professionals scoring the specialist work.

 

The results: quality rose 33–44%, depending on the statistical specification, and completion time fell 21%. 94% of participants scored higher with AI than without, each employee serving as his or her own control. Quality and time gains were uncorrelated: people banked both.

 

Where the study earns its keep is the breakdown by task type:

 


Definition: A complementarity exists when two ingredients are worth more combined than the sum of their separate effects. Using a formal test borrowed from Brynjolfsson and Milgrom, the authors find a strong complementarity between AI and non-routine work: roughly 0.9 standard deviations of extra performance beyond AI’s effect on routine tasks. AI helps with the predictable stuff; it helps even more with the open-ended stuff. So much for the folk theory that chatbots only automate drudgery.

 

The specialist tasks deserve a second look. Performance more than doubled (a whopping +113%), but from a dismal base of 2.2 out of 10, and even AI-assisted employees averaged only 4.9. Hard expert work stayed hard: AI turned failing grades into mediocre ones, still the largest jump in the experiment. And many participants who already used AI at work told the authors this was their first attempt at applying it to advanced specialist tasks. The frontier of AI usage inside organizations sits well behind the frontier of AI capability.

 

Now the nuance that complicates the equalization story. On generalist tasks, lower-skill employees gained the most quality, replicating every study in my table above. But equalization stopped there. On efficiency, the pattern reversed: high-skill workers gained the most speed (a noisier result, the authors note). On specialist tasks, equalization vanished entirely: the skill gradient disappeared, and the debriefs suggested that high-skill workers with prior AI exposure integrated the tool best on hard problems. (And returns were identical for men and women, so the well-documented gender gap in AI adoption is an on-ramp problem; fix the on-ramp, and the gap should close.)

 

A quick cold shower for this study too: participants from a central bank, 92% holding at least a master’s degree, hardly mirror the labor force, and 4-minute generalist tasks are closer to work samples than workdays. Even so, the design beats 90% of the AI-productivity claims in your LinkedIn feed: preregistered, randomized, blind-graded, with verified compliance. And because newer reasoning models handle open-ended work far better than mid-2024’s GPT-4o did, the authors argue (correctly, in my assessment) that their non-routine estimates are floors.

 

The Mismatch: The Workers Who Gain Most Hold the Jobs Where AI Helps Least

The bank paper’s most original contribution is a paradox. At the worker level, employees in routine-heavy jobs gained the most from AI. At the task level, AI helped non-routine tasks the most (+58% vs. +24%). Put those together: the technology delivers its biggest boost to the very people whose daily task mix exploits it least.

 

The authors quantify the waste. In a simulation that reassigns workers to tasks according to comparative advantage under AI, the bank’s output rises 7.3% with zero new hires and zero new software. Thus, a 7.3% raise sits in plain sight, unclaimed, because task assignments fossilized in the pre-AI era. Sprinkling chatbot licenses over an unchanged org chart captures the small win and forfeits the big one. AI adoption is a job-redesign problem disguised as a procurement decision.

 

Competence Glut = Commodity Work

Definition: A commodity is a product so standardized that buyers choose on price alone. One bushel of wheat equals another, and nobody pays extra for artisanal wheat.

 

Even wheat needed help to become a commodity. When the Chicago Board of Trade introduced standardized grain grades in 1856, “No. 2 spring wheat” from one farm became interchangeable with “No. 2 spring wheat” from any other; buyers stopped asking who grew it, and the price board took over. AI is the grading system for knowledge work: it stamps a competent grade on nearly everyone’s output, and the Siddiq and Zhang coefficients show the price board taking over hiring in real time.

 

For a century, professional competence was scarce, so competence commanded a salary. The studies above describe the end of that scarcity. When a $20-per-month chatbot lifts nearly any employee to a competent memo, a competent screen design, or a competent analysis, competent output becomes abundant. I call the result the competence glut: once competence floods the market, its price collapses toward the cost of the tokens that generated it. And procurement departments treat commodities the way they treat everything they can: as a spreadsheet column to minimize.

 

We’ve run this experiment before. In 1985, Aldus PageMaker plus the Apple LaserWriter turned typesetting, a skilled trade with centuries of guild history, into a menu command. The typesetters vanished. Demand for art direction, the judgment about what the page should say and evoke, grew. A decade later, website builders commodified the brochure site, yet UX employment expanded, because the scarce skill moved up the stack. Expect the same dynamic, at higher speed, for every deliverable-shaped profession, very much including design.

 

For UX, the commodity layer is already wide: wireframes, standard screen layouts, UI copy, first-draft research summaries, and journey maps now emerge at competent quality from anyone with a prompt box (we’ll get good UI quality from AI in 2027 and great in 2028; better-than-any-human-designer probably not until 2030, but 99.9999% chance that you don’t have the world’s best UX designer working on your team anyway). Design tools generate complete, styled flows from a paragraph of intent, and the output is respectable. If your professional identity is “I produce these deliverables,” you’re selling wheat, and the market-clearing price of wheat is low and dropping. What stays scarce is everything around the deliverable: knowing which problem deserves solving, whether the research question was the right one, which of 10 plausible designs will actually move conversion, and whether the feature should exist at all. When execution is a commodity, judgment is the product. Let’s make that concrete for design careers, because design shows up in the new dataset in an instructive way.

 

When Design Is Wheat: How UX Professionals Must Reposition

Where does design sit on the study’s exposure scale? Mid-table, officially: web and mobile design scores 0.37 on a scale where writing sits at 0.62 and translation tops out at 0.80. But the exposure yardstick was made in early 2023, and it measures only what a text-based LLM could accomplish by itself, before multimodal models and prompt-to-design tools existed. (Software development scores 0.05 on the same scale, which will amuse anyone who has watched an agentic coding tool inhale a week of programming in an afternoon.)

 

Visual design has since moved deep into exposed territory, which cuts both ways: the dated yardstick biases the study’s estimates toward zero, and it means designers should read the translators’ coefficients, not their own official mid-table score, as their preview. Translators topped the exposure scale, and they got hit first and hardest. Consider yourself previewed.

 

So take the findings personally. Three assets that UX professionals spend careers accumulating are losing purchasing power in AI-exposed markets:

 

  • The portfolio. A gallery of polished deliverables proves less every month, because AI now produces polished deliverables on demand. When any client can generate 10 respectable landing pages before lunch, your 11th respectable landing page is evidence of nothing scarce. Verified proof of past production is depreciating because production itself is depreciating.


Retarget your portfolio to emphasize your talent for judgment. If you showcase your ability to crank out deliverables, you’re selling so much commodity wheat. (GPT Image 2)

 

  • The credentials. Degrees, certifications, and employment history exist to predict output quality. AI compresses output quality, so the predictions matter less. A diploma still signals conscientiousness, but the market’s willingness to pay for it fell 2.4% in exposed categories and kept falling.

  • The ratings. The finding that should worry freelancers most: even reputation earned inside the market through real transactions (5-star feedback, “Top Rated” badges) lost 2.9% of its hiring weight. A 5-star rating for delivering competent work stops differentiating once everyone delivers competent work.


What still moves demand? Everything the platforms can’t standardize:

 

  • Outcomes with numbers attached. “I redesigned checkout and conversion rose 23%” is a different asset than “here are the checkout screens I drew.” The first documents judgment; the second documents production, and production is on sale.

  • Proprietary user research. Fresh empirical data about how real users behave in your product is ground truth that no prompt can conjure.

  • Evaluation authority. Somebody must decide which of the 10 AI-generated flows ships, and whoever holds that quality gate holds the scarce position.

  • Direct relationships. Note what commoditized fastest: an anonymous marketplace where switching freelancers costs one click. Trust built with named humans over repeated engagements has no skill tag, appears in no embedding, and survives the credential discount. Friction, for once, is your friend.


One pricing warning. The Upwork data shows demand flowing to cheaper workers, and the tempting response is to cut your rate. Resist. When clients start buying on price, the worst possible response is to compete on price. You can’t out-cheap a $20-a-month subscription, and every discount confirms you’re selling the commodity rather than the judgment. Exit the wheat market; don’t become its lowest-cost farmer.

 

Stronger AI May Reverse the Forklift Effect

Everything so far describes current AI applied to production tasks. Will equalization and commoditization persist as AI grows more capable? My answer is no. What follows is speculation, plainly labeled: these are my predictions, and the future is under no obligation to cooperate. But the reversal keeps poking through at the edges of the data, so let’s start with the evidence.

 

Nicholas Otis, Rowan Clarke, Solène Delecourt, David Holtz, and Rembrand Koning gave 640 Kenyan entrepreneurs a GPT-4 business mentor on WhatsApp for 5 months (SSRN, 2024). The average effect on revenues and profits was zero. Underneath the zero sat a split: high-performing entrepreneurs gained roughly 15–20%, while low performers did about 8–10% worse. Both groups received similar advice from the AI. The difference lay in which advice each group selected and implemented. The moment the human’s job shifted from producing work to choosing among possibilities, the forklift effect flipped, and AI widened the gap instead of narrowing it.

 

Knowing what to lift is important, especially if you run your own business, as the Kenyan entrepreneurs in the study did. (GPT Image 2)

 

Medicine supplies a second data point, from the opposite direction. In a randomized trial published in JAMA Network Open, 50 physicians worked through difficult diagnostic vignettes with or without GPT-4 (Goh et al., 2024). The AI-assisted physicians diagnosed no better than colleagues using conventional references. GPT-4 alone, however, outscored both groups by roughly 16 percentage points. The machine held the answers; the humans discarded them, overriding correct AI suggestions with their own inferior ones. (Recall the law-school parallel: well-prompted GPT-4 alone also beat the AI-assisted students.) The bottleneck sat in the humans’ judgment about when to defer, and that judgment varied wildly.

 

Economists have now isolated this judgment factor experimentally. Andrew Caplin, David Deming, and colleagues show that two traits jointly determine who profits from AI: ability and calibration, meaning accurate beliefs about one’s own ability (NBER, 2024, since published in Management Science). Low-ability participants gained most on average, consistent with the forklift. But holding ability constant, well-calibrated people gained far more, because the overconfident ignore correct AI advice and the underconfident follow incorrect advice. Participants who knew they had low ability gained the most of all, nearly 10 percentage points, and a 2026 replication with professional radiologists reading chest X-rays found the same pattern. Knowing what the machine knows better than you is itself a skill, one that’s unevenly distributed, and nothing in the equalization studies suggests AI hands it out for free.

 

Anthropic’s analysis of 400,000 Claude Code sessions in June 2026 found that agentic coding compounds expertise rather than erasing it. The report documents that more experienced developers delegate differently, structure tasks differently, and extract more value per session, and it frames supervision and verification, not coding itself, as the emerging bottleneck and the emerging source of advantage.

 

The two main ways expertise helped in this study:

 

  1. Leverage per prompt. In typical novice sessions, each prompt sets off about 5 Claude actions and roughly 600 words of output; expert sessions set off action chains more than twice as long (12 actions) carrying 5 times the output (3,200 words). The gap appears within every kind of work and every band of task value, and in a regression controlling for work mode, task value, month, occupation, and model family, each expertise level adds about 9% more actions and 13% more output per prompt (p < 0.001).

  2. Recovery when things break. (And they do.) Among sessions that hit trouble (such as errors, failed tests, repeated attempts, or expressed frustration), verified success rose from 4% for less-skilled users to 15% for experts. And 19% of novice troubled sessions were abandoned, vs. 5–7% for everyone else. Part of the value of expertise when using advanced AI appears to be the ability to steer the agent back on course.


The mechanism is straightforward once you see it. When the AI does the raw generation, writing code stops being the scarce skill. What becomes scarce is knowing what to ask for, recognizing when the output is subtly wrong, and stitching multi-step work into something that holds together. Those are expert skills. So as the machine absorbs the easy part, the returns flow to whoever supervises well. Intelligence got cheap; judgment got expensive.

 

With agentic AI, expertise will be valued higher, not lower. (GPT Image 2)

 

All 4 results point to the same root cause: the human’s job is quietly shifting from production to selection, and selection obeys different economics. As models become agentic, running multi-step work over minutes or hours (what I’ve called slow AI), the user’s role migrates from doing the work to directing it: deciding what should be done in the first place, decomposing goals, supervising the agents, and, above all, evaluating and selecting among competing AI work products. Consider what each of those duties rewards:

 

  • Deciding what should be done is the scarcest skill of all. Peter Drucker warned that nothing is quite so useless as doing efficiently what shouldn’t be done at all. An agent swarm executing the wrong strategy produces impeccable waste at record speed. Problem selection has always separated great professionals from good ones, and agents multiply the price of that separation.

  • Evaluation becomes the new production. Picking the best of 10 AI drafts requires the taste to tell which one is best, and judgment correlates brutally with experience and talent. In Kenya, low performers couldn’t filter good advice from merely plausible advice. A junior designer facing 10 gorgeous, subtly wrong AI mockups sits in the same trap: he or she can generate infinitely, and the bottleneck is knowing what to keep. The central-bank study already whispers this message, since the biggest winners on the hardest tasks were high-skill employees who knew both the domain and the tool.

  • Leverage amplifies judgment in both directions. A person directing 50 tireless agents multiplies every decision 50-fold, good and bad alike. Management has never equalized outcomes among managers; get a team of 10 underlings to work for you, and the results magnify every difference in management skill. Handing everyone an agent workforce amounts to promoting everyone to manager, and most people have never managed anything.


The supervision era has already been piloted at scale, and the early data supports this argument. In August 2024, Alibaba ran a randomized field experiment on Taobao in which 647 customer service workers either handled every chat themselves or supervised an agentic AI that autonomously resolved eligible chats, intervening when needed (arXiv, 2026; 680,676 chats).

 

The agentic AI cut chat time, but outcomes hinged on the quality of human oversight. When supervisors used their own judgment to step in early, before customer frustration hardened, they resolved issues at least as reliably as all-human service. When they instead waited for the monitoring algorithm to sound the alarm on an emotionally deteriorating chat, quality cratered: customer ratings fell 0.9 points on a 5-point scale, repeat contacts rose 6 percentage points, and the transcripts show the workers half giving up. Same AI, same tasks; the differentiator was the supervisor’s timing and judgment. (A pleasant side effect: with routine chats offloaded, the same workers’ ratings on the judgment-heavy chats they kept handling went up.) Oversight, it turns out, is a skill with a distribution, exactly like the production skills it’s replacing.

 

Return to the warehouse one last time. A forklift narrows the strength gap because nearly anyone can drive one after a 2-day certification course. A 100-meter tower crane is a different machine: its operator is licensed, scarce, and well paid, because the crane’s power multiplies the cost of every operator error. The stronger the machine, the more the operator’s judgment matters. Today’s chatbots are forklifts. Agentic AI is a crane.


While first-generation AI products were indeed the equivalent of forklifts in narrowing skill gaps, second-generation AI in the form of agentic products may be more similar to a crane which only works in the hands of a skilled operator. (Muse Image)

 

Now connect the prediction back to the Upwork data, and notice what the platform cannot see. Profiles carry skill tags for Python, Figma, and legal translation, but there’s no skill tag for “knows what to build,” no star rating for “kills bad projects early,” and no badge for taste. The market just repriced every signal it can see; the signal it will need next doesn’t exist yet. I predict that within 2–3 years, labor marketplaces will scramble to invent judgment signals, such as outcome-verified case histories (did the freelancer’s decisions move the client’s metric?), because a market that differentiates only on price is in a margin death spiral, bad for workers and platforms alike. Watch for the human capital coefficients in follow-up studies to bottom out, and for a new premium to appear, attached to direction rather than production.

 

Of course, one counterargument deserves airtime: judgment itself may eventually commodify. In freestyle chess, human-plus-engine teams beat engines alone for roughly a decade after Garry Kasparov introduced the format in 1998; today, the human teammate is a liability. (The only way to win in chess is to play the exact move the AI suggests. Any human judgment, and you lose.) If AI evaluators someday out-judge expert humans across professional work, the judgment premium becomes a historical window rather than a permanent moat. My best guess: the window stays open through at least 2030 for work embedded in messy human organizations, where goals are political, data is ambiguous, and accountability must land on a person with a name. Enjoy the window, and don’t plan your retirement around it.

 

So here’s my forecast: current AI narrows skill gaps in producing work; future AI will widen skill gaps in directing work. The equalization studies of 2023–2025 measured AI as a better pen. Agentic AI is a workforce, and workforces have always paid a premium to the people who know what to build.

 

Takeaways for Corporate Strategy

  1. Reallocate people as well as licenses. The central-bank simulation found 7.3% more output from reassigning tasks by comparative advantage under AI, with zero additional spending. Redesign task assignments and job descriptions around what AI changes. A chatbot bolted onto a 2019 org chart collects a fraction of the available value.

  2. Onboard the bottom, and train the top. Basic adoption support raises the quality floor for weaker staff, which the equalization studies all but guarantee. But the largest absolute gains in the bank came from experts applying AI to hard specialist tasks, which required domain skill plus prior AI experience. Budget for both programs; they aren’t substitutes.

  3. Stop paying a premium for commodity competence, and stop buying judgment at commodity prices. If AI lifts everyone to the same competent level on a task, its market price will fall regardless of anyone’s feelings; the Upwork clients have already started the repricing for you. But run the logic in both directions: procurement that selects a design or research partner on day rate alone will reliably buy wheat and then wonder why the strategy tastes bland. Sort purchases into commodity and judgment lanes, and price-shop only the first.

  4. Rebuild the apprenticeship ladder before it rots. Juniors traditionally learned judgment by grinding through the commodity work that AI now performs. Replace that lost training ground deliberately: structured critique, shadowing senior evaluators, and graded review of AI output. Otherwise, in 10 years you’ll wonder where the senior people went.


If we remove the bottom rungs from the experience ladder, we won’t have any senior staff in 10 years. (GPT Image 2)

 

  1. Measure quality and speed separately. The bank found the two gains uncorrelated across employees. Track only throughput, and you’ll miss quality regressions; track only quality, and you’ll miss the people burning hours to polish commodity output. Set explicit targets for each.

  2. Build the evaluation function now. Someone must specify, verify, and sign off on agent output at scale. Rubrics, sampling audits, and named accountability are the new production infrastructure. Companies that treat verification as overhead will ship impeccable waste, faster than ever.


Takeaways for UX Professionals

  1. Sell outcomes and decisions, not deliverables. A wireframe is wheat, and the new data says the portfolio displaying it trades at a growing discount. Knowing which design will lift conversion, and being able to prove it with data, is the product. Restate your role, your case studies (fewer artifacts, more decisions with measured consequences), and your job postings accordingly.

  2. Claim the evaluation layer. UX is the discipline of judging fitness for human use, which makes it the natural quality-control function for AI-generated everything: screens, flows, copy, and agent behavior. Volunteer to own the rubrics. Whoever grades the work governs the work.

  3. Make original user research your moat. AI remixes what’s already known, and it can’t watch your users tomorrow. Fresh empirical data about real behavior is proprietary ground truth and the raw material of the judgment you’re selling. Watch users, not model outputs. (In my example: anybody can generate a photo of a cute dog. Only UXR can produce a picture of the company’s actual customers, whether dogs or not.)

  4. Train taste on purpose. Judgment used to accrete accidentally through years of production work, and that path is closing. Build it deliberately: regular design critiques, exposure to excellent and terrible work, and prediction exercises scored against real usability-test results.

  5. Practice directing agents. Move beyond single prompts to specifying, decomposing, and supervising multi-step AI work. The UX professionals who thrive will run a portfolio of AI workers the way a creative director runs a studio.

  6. Specialize in a domain. The bank’s specialist tasks delivered the largest gains, and only domain experts could harvest them. Deep knowledge of fintech, health, enterprise workflows, or another vertical multiplies what AI does for you and resists commodification far longer than generalist skills will.


You need to move beyond commoditized production skills if you want to sell your portfolio to hiring managers. (GPT Image 2)

 

Conclusion: The Forklift Doesn’t Know What to Lift

The forklift transformed the warehouse twice. First, it equalized: strong and weak workers moved identical pallets, and strength stopped paying. Second, and more quietly, it shifted the money to the people who decided which pallets went where, because a warehouse that moves the wrong goods quickly is merely an accelerated mistake.

 

A warehouse that moves low-value products fast will still lose money, even with AI forklifts. (GPT Image 2)

 

AI is running the same 2-act play on knowledge work, only faster. Act 1 is here: skill gaps narrow, competence gluts the market, and deliverables commodify. The 9 experiments in this article confirm it, and 2.26 million contracts show buyers acting on it: credentials trade at a discount, and demand flows to the cheapest competent bid. Act 2 arrives with agentic AI: leverage flows to the people with the judgment to decide what should be built and the taste to select the best of what the machines produce. Position yourself, and your organization, for the second act while everyone else is still applauding the first.

 

Have you watched AI compress skill gaps on your team, or watched clients start shopping on price? I welcome your observations. And remember the limit of the machine that started this argument: the forklift lifts anything you point it at, but the forklift doesn’t know what to lift. That job is still yours. Keep it.


Forklifts are good at lifting but don’t know what to lift. Similarly, with agentic AI, somebody needs to be the equivalent of the factory founder and decide what should be done. (Muse Image)


Top Past Articles
bottom of page