skip to content
Topic

foundation-models

3 essays on this topic.

  1. Near-Parity Is Mostly a Claim About Formatting

    On the benchmark built to measure AI against real professional work, the best model scores 48% on deliverables that come in a file and 14% on deliverables judged as prose alone. That 34-point gap is the largest single effect in the paper, and it is not in the limitations section.

  2. The Number Checks Out. The Claim Doesn't.

    The most-quoted figure about AI and professional work is accurate, in the paper, and correctly transcribed. It is also a description of a scenario nobody would choose, because what travelled was the number without the condition it was computed under.

  3. Your Reviewer Model Is Not Independent

    When agents generate faster than anyone can read, the standard answer is a second model reviewing the first. Aerospace decomposed what makes a check independent decades ago, and a reviewer model fails the hardest of the three tests.