When AI Safety Benchmarks Miss the Risks, Users Pay the Price
AI companies publish benchmarks that compete for almost everything. Reasoning, coding, factual accuracy, bias, safety, cybersecurity, mathematics and increasingly specialised scientific capabilities [...]