You will have seen the claim that frontier models now do real professional work about a hundred times faster and a hundred times cheaper than the people who normally do it. It comes from GDPval, OpenAI’s benchmark of 1,320 tasks drawn from 44 occupations and built from the actual work of professionals averaging fourteen years of experience. It is careful research and the number is really in it: for GPT-5 the paper reports 90× on speed and 474× on cost.
It reports those figures under a strategy the authors label naive, meaning you take whatever the model produces and nobody looks at it. The same table gives the alternatives. Have a person check and repair a single attempt and the advantage is 1.12× on speed and 1.18× on cost; let the model retry a few times first and you get 1.39× and 1.63×. So the defensible version of the headline is that a frontier model with expert oversight is somewhere between twelve and sixty-three per cent better than the expert alone. That is a genuinely good result, and it is not a hundredfold. The whole distance between those two claims is the review step. The same table also shows GPT-4o coming out worse than not using it at all once review is priced in, because someone still has to read and repair everything it produced — capability and review cost are one number, not two.
I went looking because I wanted to settle something simple, whether these models are much better at software work than at other kinds, and found that nearly every figure I reached for failed some check. Most failed in ways that a trip to the source repairs. METR’s much-quoted finding that AI tools made experienced developers nineteen per cent slower is still true about early 2025, but METR published an update in February 2026 whose raw results point the other way and which announces a redesign of the study; search returns the headline and nothing links you to the correction. GDPval’s own near-parity result excludes tasks involving tacit knowledge, proprietary tools or communication between people, which is most of what makes a job a job rather than a deliverable. And SWE-bench Verified, routinely described as saturated in the high eighties, has a public leaderboard whose best independently evaluated result is 79.2% from December 2025, with nothing at all from 2026 — because its maintainers restricted submissions to research authors that November, saying they wanted a research venue rather than a product validation platform. The circulating figure is a vendor’s claim about its own model, wearing the name of a leaderboard that stepped back from exactly that comparison.
Every one of those dissolves when you stop reading summaries and read the thing. Which is why the hundredfold case is the one worth keeping, because it survives that treatment. Nobody misquoted it. It is in the paper, correctly transcribed, and if you go and check it the source agrees with you. What travelled was the number stripped of the condition it was computed under, and going to the primary document is no defence at all against that particular failure.
So the provenance check needs a fourth question. Ask who evaluated a number, because a vendor’s claim and an independent evaluation are different objects. Ask where it was published, because a venue can quietly stop accepting the class of submission everyone still quotes. Ask when, because labs supersede themselves and the correction rarely travels with the claim. Then ask under what assumption it was computed, which is the one a source check cannot catch.
None of this is an argument for cynicism about benchmarks. The papers behind these numbers are careful, and in every case the qualification was stated plainly by the authors themselves, in an appendix, a limitations section, a column header. The failure is entirely in transmission, and transmission is not neutral about what it drops. It keeps the magnitude and loses the condition, in that order, every time. The hundredfold survives and the condition it was measured under, no human review, does not, and what is left is a number that sounds like a description of the world when it is a description of a scenario nobody would choose.