Close Menu
TechTost
  • AI
  • Apps
  • Crypto
  • Fintech
  • Hardware
  • Media & Entertainment
  • Security
  • Startups
  • Transportation
  • Venture
  • Recommended Essentials
What's Hot

ServiceNow bets $40M on Indian banking software specialist to expand push into financial services

Tesla Spending Soars as Cybercab, Semi, Megapack Production Schedule Slips

Travis Kalanick’s robotics company raises $1.7 billion, led by a16z

Facebook X (Twitter) Instagram
  • About Us
  • Contact Us
  • Privacy Policy
  • Terms and Conditions
  • Disclaimer
Facebook X (Twitter) Instagram
TechTost
Subscribe Now
  • AI

    Google justifies its massive AI spending with a thriving cloud business

    23 July 2026

    Arcee, a US open-source AI lab, says Chinese models are not inherently dangerous

    22 July 2026

    Anthropic-Physical Intelligence Rumors Rock AI Twitter

    22 July 2026

    Anthropic’s landmark $1.5 billion copyright settlement approved

    21 July 2026

    YouTube clarifies policies on AI and disturbing videos

    20 July 2026
  • Apps

    Yope raises $12.3 million to build a private social network without algorithms or ads

    23 July 2026

    The browser wars aren’t about search anymore — here are the best alternatives to Chrome and Safari

    22 July 2026

    These are the countries that are moving to ban social media for children

    22 July 2026

    Adobe’s new camera app feature will critique your photos using AI

    20 July 2026

    Superhuman’s new auto-draw feature almost makes me like the AI ​​responses

    19 July 2026
  • Crypto

    Venice AI goes unicorn with $65M Series A as first privacy AI platform takes off

    1 July 2026

    Crypto Exchange OKX wants AI agents to hire and pay each other

    30 June 2026

    Startup Battlefield 200 applications close today

    27 May 2026

    5 days left: Save up to $410 on Disrupt 2026 passes

    25 May 2026

    As crypto cools, a16z crypto raises $2.2 billion in capital

    6 May 2026
  • Fintech

    Don’t want to invest in Elon Musk? Two new ETFs expressly exclude him

    10 July 2026

    India’s payments chief believes artificial intelligence will play a big part in the next era of digital payments development

    28 June 2026

    Early Bird pricing ends tonight for the Founder Summit

    26 June 2026

    4 days left to save up to $190 on Founder Summit 2026

    23 June 2026

    Robinhood’s note on 10% layoffs shows that blaming AI doesn’t cut it

    17 June 2026
  • Hardware

    After a shocking quarter, IBM insists that artificial intelligence is not killing the mainframe

    23 July 2026

    Light made a flip phone — it’s colorful and cheap

    22 July 2026

    Apple is partnering with Klarna to launch a rental program for iPhones, iPads and Macs

    22 July 2026

    The Xteink X4 Pro could be the tiny e-reader of your dreams

    21 July 2026

    reMarkable’s new Paper Pure is good. That’s why I wrote this review on it.

    21 July 2026
  • Media & Entertainment

    SoundCloud acquires decentralized music platform Nina Protocol months after its shutdown

    23 July 2026

    What you need to know about Warner Bros.’ landmark Discovery sale

    22 July 2026

    AI and the rise of the universal entertainment app

    22 July 2026

    Judge blocks $110 billion Paramount-Warner Bros. merger

    21 July 2026

    Spotify extends parent-managed accounts to users on its free tier

    16 July 2026
  • Security

    How OpenAI’s human error led to the AI-powered Hugging Face hack

    22 July 2026

    Glow emerges from stealth to $1.2 billion valuation to challenge endpoint security in the age of artificial intelligence

    22 July 2026

    Hackers steal ‘significant’ amount of data from tech company that thousands of US hospitals and pharmacies rely on

    21 July 2026

    Hackers are exploiting newly patched WordPress bugs, putting millions of websites at risk

    20 July 2026

    How an ex-DeepMind researcher raised a $300 million seed valuation before launching a product

    18 July 2026
  • Startups

    ServiceNow bets $40M on Indian banking software specialist to expand push into financial services

    23 July 2026

    Bluecore Energy raises $10 million to build portable nuclear reactors on barges

    21 July 2026

    Colossal Biosciences is reportedly in talks to raise new capital at a $20 billion-$30 billion valuation

    21 July 2026

    Natural raises $30 million to reinvent payments for AI agents — and take on Stripe

    20 July 2026

    Backed by $60 million in funding, Oak comes out of stealth to fix the identity mess exacerbated by AI agents

    18 July 2026
  • Transportation

    Tesla Spending Soars as Cybercab, Semi, Megapack Production Schedule Slips

    23 July 2026

    As EVs slow, Sila raises $300M to expand battery materials factory

    22 July 2026

    Einride bets $38 million on EV charging as it scales up electrification

    21 July 2026

    TechCrunch Mobility: The battle over the robotaxi rules

    19 July 2026

    A 600-mile road trip (and data) proves EV charging isn’t crap anymore

    19 July 2026
  • Venture

    Travis Kalanick’s robotics company raises $1.7 billion, led by a16z

    23 July 2026

    Cascade raises $3.5 million to help construction companies find and win projects

    22 July 2026

    StrictlyVC returns to New York on September 10 to celebrate a huge year for the city’s startup community

    21 July 2026

    Startup Inference Infinity raises $15 million from researchers Touring Capital, OpenAI and Anthropic

    20 July 2026

    Nuclear startup Valar Atomics is in talks to raise new funding at a $6 billion valuation

    18 July 2026
  • Recommended Essentials
TechTost
You are at:Home»AI»Why most AI benchmarks tell us so little
AI

Why most AI benchmarks tell us so little

techtost.comBy techtost.com8 March 202405 Mins Read
Share Facebook Twitter Pinterest LinkedIn Tumblr Email
Why Most Ai Benchmarks Tell Us So Little
Share
Facebook Twitter LinkedIn Pinterest Email

On Tuesday, startup Anthropic released a family of AI models that it claims achieve best-in-class performance. Just days later, rival Inflection AI unveiled a model that it claims comes close to matching some of the most capable models out there, including OpenAI’s GPT-4, in quality.

Anthropic and Inflection are by no means the first AI companies to claim that their models have matched or beaten the competition by some objective measure. Google supported the same with its Gemini models at launch, and OpenAI said the same for GPT-4 and its predecessors, GPT-3, GPT-2, and GPT-1. The list goes on.

But what metrics are they talking about? When a seller says a model achieves top performance or quality, what exactly does that mean? Perhaps more to the point: Will a model that technically “performs” better than some other model in reality touch improved in a tangible way?

On that last question, not likely.

The reason – or rather the problem – lies in the benchmarks that AI companies use to quantify a model’s strengths and weaknesses.

Internal measures

Today’s most commonly used benchmarks for AI models — specifically chatbot-powered models such as OpenAI’s ChatGPT and Anthropic’s Claude — do a poor job of capturing how the average human interacts with the models being tested. For example, a benchmark cited by Anthropic in its recent announcement, GPQA (“A Graduate-Level Google-Proof Q&A Benchmark”), contains hundreds of PhD-level biology, physics, and chemistry questions — yet most people use chatbot for tasks like answering emails, writing cover letters and talking about their feelings.

Jesse Dodge, a scientist at the Allen Institute for AI, the nonprofit AI research organization, says the industry has reached a “crisis of evaluation.”

“Benchmarks are typically static and narrowly focused on evaluating a single capability, such as a model’s realism in a single domain or its ability to solve multiple-choice mathematical reasoning questions,” Dodge told TechCrunch in an interview. “Many benchmarks used for evaluation are more than three years old, from when AI systems were mainly used for research and did not have many real users. In addition, humans use genetic AI in many ways — they are very creative.”

Wrong measurements

It’s not that the most used benchmarks are completely useless. No doubt someone is asking Ph.D level math questions. in ChatGPT. However, as genetic AI models are increasingly positioned as mass-market, do-it-all systems, the old benchmarks are becoming less applicable.

David Widder, a postdoctoral researcher at Cornell who studies artificial intelligence and ethics, notes that many of the common tests of reference skills—from solving school-level math problems to determining whether a sentence contains an anachronism—will never be relevant to the majority of users.

“Earlier AI systems were often built to solve a specific problem in a context (e.g. medical AI expert systems), making a deep understanding of what constitutes good performance in that particular context more possible,” Widder said. at TechCrunch. “As systems are increasingly seen as ‘general purpose’, this is less possible, so we’re increasingly seeing a focus on testing models across a variety of benchmarks in different fields.”

Errors and other defects

In addition to misalignment with use cases, there are questions about whether some benchmarks are properly measuring what they are supposed to measure.

One analysis of HellaSwag, a test designed to assess common sense reasoning in models, found that over a third of the test questions contained typos and “stupid” writing. Somewhere else, MMLU (short for “Massive Multitask Language Understanding”), a benchmark highlighted by vendors such as Google, OpenAI and Anthropic as proof that their models can reason through logic problems, asks questions that can be solved through memorization verbatim.

Test questions from the HellaSwag benchmark.

“[Benchmarks like MMLU are] more about memorizing and associating two keywords together,” Widder said. “I can find [a relevant] article quickly enough and answer the question, but that doesn’t mean I understand the causal mechanism or that I could use my understanding of that causal mechanism to actually reason and solve new and complex problems in unpredictable contexts. Not even a model can.”

Fixing what’s broken

So benchmarks are broken. But can they be fixed?

Dodge believes so – with more human involvement.

“The right way forward, here, is a combination of evaluation benchmarks with human evaluation,” he said, “prompting a model with a real user question and then hiring a human to evaluate how good the response is.”

As for Widder, he’s less optimistic that benchmarks today — even with corrections for the most obvious mistakes, like typos — can be improved to the point where they would be informative to the vast majority of AI model users. Instead, he believes that tests of models should focus on the downstream effects of those models and whether the effects, good or bad, are seen as desirable by those affected.

“I would ask for what specific goals we want AI models to be able to be used for and assess whether they would be – or are – successful in such contexts,” he said. “And hopefully that process also includes evaluating whether we should be using AI in such contexts.”

All included benchmarks genAI Generative AI reference points Research
Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
Previous ArticleApple will ease the transition to Android by fall 2025
Next Article LLMs are ready to make logging business intelligence tools easier and faster to use
bhanuprakash.cg
techtost.com
  • Website

Related Posts

Google justifies its massive AI spending with a thriving cloud business

23 July 2026

Arcee, a US open-source AI lab, says Chinese models are not inherently dangerous

22 July 2026

Anthropic-Physical Intelligence Rumors Rock AI Twitter

22 July 2026
Add A Comment

Leave A Reply Cancel Reply

Don't Miss

ServiceNow bets $40M on Indian banking software specialist to expand push into financial services

23 July 2026

Tesla Spending Soars as Cybercab, Semi, Megapack Production Schedule Slips

23 July 2026

Travis Kalanick’s robotics company raises $1.7 billion, led by a16z

23 July 2026
Stay In Touch
  • Facebook
  • YouTube
  • TikTok
  • WhatsApp
  • Twitter
  • Instagram
Fintech

Don’t want to invest in Elon Musk? Two new ETFs expressly exclude him

10 July 2026

India’s payments chief believes artificial intelligence will play a big part in the next era of digital payments development

28 June 2026

Early Bird pricing ends tonight for the Founder Summit

26 June 2026
Startups

ServiceNow bets $40M on Indian banking software specialist to expand push into financial services

Bluecore Energy raises $10 million to build portable nuclear reactors on barges

Colossal Biosciences is reportedly in talks to raise new capital at a $20 billion-$30 billion valuation

© 2026 TechTost. All Rights Reserved
  • About Us
  • Contact Us
  • Privacy Policy
  • Terms and Conditions
  • Disclaimer

Type above and press Enter to search. Press Esc to cancel.