GDPval is the most serious attempt yet to measure AI against real professional work: 1,320 tasks drawn from 44 occupations, built from what people with an average of fourteen years in the job actually produce, graded by other practitioners in blind pairwise comparison against the human’s own deliverable. Its headline is that frontier models are approaching industry experts in deliverable quality. The best model was graded a win or a tie on 47.6% of tasks.
Break that number down by the kind of file a task asks for and it stops being one number. Claude Opus 4.1 scores 48% when the deliverable falls into the paper’s “other” bucket, 45% on a PDF, 45% on a slide deck, 43% on a spreadsheet, and 14% on pure text. Strip the container away and the leading model loses thirty-four points; that is the largest single effect anywhere in the paper, larger than the spread across sectors, larger than the gap between models, and roughly double the decline across task duration.
It also inverts the leaderboard on the way down. Opus leads every other format comfortably. On pure text it drops to fourth of five, behind GPT-5, o3 and o4-mini. A ranking that holds everywhere else reverses in the one category with nothing to look at.
The paper supplies the mechanism a section earlier without quite drawing the conclusion. Opus, it says, excels in particular on aesthetics, meaning document formatting and slide layout, while GPT-5 excels in particular on accuracy, carefully following instructions and performing correct calculations. Those are different skills, and the grading rewards them differently depending on what lands on the grader’s desk. Around ninety per cent of the tasks arrive wrapped in something: a deck, a workbook, a formatted document, an artefact whose competence is partly visible before you read a word of it. Roughly one task in ten is judged as prose on its substance alone, and that is the slice where near-parity vanishes.
This is worth saying plainly, because the headline gets carried into rooms where people are deciding things. Near-parity, as measured, is substantially a claim about producing a well-formatted artefact to a specification. It is not a claim about the quality of the thinking inside the artefact, and the paper’s own data separates those two by thirty-four points. If your work product is a deck, that is genuinely encouraging. If your work product is an argument, the benchmark has less to say about you than the headline suggests, and what it does say is considerably less flattering.
There is a methodological sting here too. The honest reader’s instinct, when a benchmark’s headline sounds too good, is to go and read the limitations section. GDPval’s is a good one: it tells you the tasks exclude tacit knowledge, proprietary tools and communication between people, and that they are precisely specified and one-shot rather than interactive. All true, all worth knowing. But the biggest qualification on the headline was not there. It was a bar chart in an appendix, and you find it only by asking what else varies and then going to look.
So the practical rule is narrower than “read the limitations”. Ask what the benchmark’s unit of measurement actually is — here, a deliverable, which is a document as much as it is a judgement — and then find the cut of the data where that unit is stripped to its least flattering form. The number you get there is the one to carry into a decision. It is usually in the appendix, and it is usually the one nobody quotes.