Report: Artificial Intelligence is surpassing human capabilities, prompting the need for updated benchmarks.

by

in

– Stanford University released AI Index Report 2024, noting AI advancements make human benchmark comparisons less relevant
– Industry benchmarks like MMLU compare AI models to human performance, but models are exceeding human baselines
– Researchers are developing more challenging benchmarks like GPQA to evaluate AI models against really smart people and incorporating human evaluations into benchmarking rather than computerized rankings.

Stanford University’s AI Index Report 2024 highlights the rapid advancement of AI, making benchmark comparisons with humans increasingly irrelevant. Industry benchmarks like the MMLU, which evaluates LLMs across various subjects, have shown AI models exceeding human baselines in performance. The report suggests the need for new, more challenging benchmarks as AI models reach saturation on established benchmarks such as ImageNet and SuperGLUE.

One example of a challenging benchmark is the GPQA, which consists of graduate-level multiple-choice questions that even highly skilled non-experts struggle to answer accurately. AI models are being tested against really smart people rather than average human intelligence. The report also points out that measuring AI safety remains challenging due to a lack of transparency among developers regarding training data and methodologies.

The trend in the industry is to crowd-source human evaluations of AI performance rather than relying solely on benchmark tests. This shift toward human evaluations, like the Chatbot Arena Leaderboard, aims to incorporate sentiment and preferences into model selection. As AI models continue to outperform humans and become increasingly difficult to measure, the future may see decisions based more on personal preference rather than standardized benchmarks.

Source link