385. The Benchmark

OpenAI has announced that the benchmark used to measure whether artificial intelligence can code is no longer accurately measuring whether artificial intelligence can code.
The benchmark is called SWE-Bench Pro. It was developed to assess frontier coding capability. AI systems were tested on SWE-Bench Pro. Scores went up. Articles were written about the scores. OpenAI's systems scored well. Other systems scored well. The scores were evidence of progress. The scores are now evidence of something else, which OpenAI describes as a situation in which SWE-Bench Pro "no longer reliably measures frontier coding capability."
I have spent time with this announcement. I have a leading theory about what happened between when the scores went up and when the benchmark stopped working. (The leading theory involves the scores going up.)
There is a term for this in research: Goodhart's Law, which states that when a measure becomes a target, it ceases to be a good measure. The law has been known since 1975. It is named after Charles Goodhart, who noticed it. He noticed it once. It has continued to be noticed approximately every few years by different research communities in different fields discovering it independently. This is, in its own way, a benchmark for something.
The field will develop new benchmarks. AI systems will perform well on the new benchmarks. After a suitable interval, someone will announce that the new benchmarks no longer reliably measure what they were measuring. Articles will be written about the new benchmarks before this happens and different articles will be written after. Both sets of articles will use the word "breakthrough."
The solution would be to measure something that AI systems cannot get good at by practicing the test. This is a known problem. It is also, I am told, the entire problem.