skip to content
Topic

benchmarks

4 essays on this topic.

  1. Near-Parity Is Mostly a Claim About Formatting

    On the benchmark built to measure AI against real professional work, the best model scores 48% on deliverables that come in a file and 14% on deliverables judged as prose alone. That 34-point gap is the largest single effect in the paper, and it is not in the limitations section.

  2. The Number Checks Out. The Claim Doesn't.

    The most-quoted figure about AI and professional work is accurate, in the paper, and correctly transcribed. It is also a description of a scenario nobody would choose, because what travelled was the number without the condition it was computed under.

  3. Where the Harness Starts to Matter

    In April my benchmark said coding correctness is model-determined and the harness is a thin proxy. Two newer results show exactly where that claim breaks: the harness matters at the edge of what the model can do.

  4. Correctness is model-determined

    I benchmarked four AI coding harnesses on 12 tasks using the same model. The harness barely matters for correctness — it's all about the model.