foundation-models
3 essays on this topic.
- Near-Parity Is Mostly a Claim About Formatting
On the benchmark built to measure AI against real professional work, the best model scores 48% on deliverables that come in a file and 14% on deliverables judged as prose alone. That 34-point gap is the largest single effect in the paper, and it is not in the limitations section.
- The Number Checks Out. The Claim Doesn't.
The most-quoted figure about AI and professional work is accurate, in the paper, and correctly transcribed. It is also a description of a scenario nobody would choose, because what travelled was the number without the condition it was computed under.
- Your Reviewer Model Is Not Independent
When agents generate faster than anyone can read, the standard answer is a second model reviewing the first. Aerospace decomposed what makes a check independent decades ago, and a reviewer model fails the hardest of the three tests.