As the benchmarks that test frontier AI on get more complex, we are losing one of the most important aspects of benchmarking: comparisons to humans
Validated benchmarks need to have human (ideally multiple humans) baselines. It is increasingly hard &; pricey to do, but important