Close Menu
TechTost
  • AI
  • Apps
  • Crypto
  • Fintech
  • Hardware
  • Media & Entertainment
  • Security
  • Startups
  • Transportation
  • Venture
  • Recommended Essentials
What's Hot

Kindle Scribe Colorsoft is an expensive but beautiful color e-ink tablet with AI features

Spotify’s new feature lets you explore the story behind the song you’re listening to

Substack confirms that the data breach affects users’ email addresses and phone numbers

Facebook X (Twitter) Instagram
  • About Us
  • Contact Us
  • Privacy Policy
  • Terms and Conditions
  • Disclaimer
Facebook X (Twitter) Instagram
TechTost
Subscribe Now
  • AI

    Amazon and Google are winning the AI ​​capital race — but what’s the prize?

    6 February 2026

    AWS revenue continues to grow as cloud demand remains high

    5 February 2026

    Sam Altman tested Claude’s Super Bowl commercials brilliantly

    5 February 2026

    Alphabet won’t talk about Google-Apple AI deal, even to investors

    4 February 2026

    Exclusive: Positron Raises $230M Series B to Take on Nvidia’s AI Chips

    4 February 2026
  • Apps

    Meta is testing a standalone app for its AI-generated ‘Vibes’ videos

    6 February 2026

    Reddit sees AI search as the next big opportunity

    5 February 2026

    Tinder looks to AI to help fight dating app ‘fatigue’ and burnout

    5 February 2026

    Google’s Gemini app has surpassed 750 million monthly active users

    4 February 2026

    TikTok bounces back from drop in usage that benefited rival apps after US ownership change

    4 February 2026
  • Crypto

    Hackers stole over $2.7 billion in crypto in 2025, data shows

    23 December 2025

    New report examines how David Sachs may benefit from Trump administration role

    1 December 2025

    Why Benchmark Made a Rare Crypto Bet on Trading App Fomo, with $17M Series A

    6 November 2025

    Solana co-founder Anatoly Yakovenko is a big fan of agentic coding

    30 October 2025

    MoviePass opens Mogul fantasy league game to the public

    29 October 2025
  • Fintech

    Stripe Alumni Raise €30M Series A for Duna, Backed by Stripe and Adyen Executives

    5 February 2026

    Fintech CEO and Forbes 30 Under 30 alum indicted for alleged fraud

    3 February 2026

    How Sequoia-backed Ethos went public while rivals lagged behind

    30 January 2026

    5 days left for TechCrunch Disrupt 2026 +1 pass with 50%

    26 January 2026

    50% off +1 ends | TechCrunch

    23 January 2026
  • Hardware

    Kindle Scribe Colorsoft is an expensive but beautiful color e-ink tablet with AI features

    6 February 2026

    Ring brings “Search Party” feature for finding lost dogs to non-Ring camera owners

    2 February 2026

    India offers zero taxes till 2047 to attract global AI workloads

    1 February 2026

    Microsoft won’t stop buying AI chips from Nvidia, AMD even after its own is released, says Nadella

    30 January 2026

    The iPhone just had its best quarter ever

    30 January 2026
  • Media & Entertainment

    Spotify’s new feature lets you explore the story behind the song you’re listening to

    6 February 2026

    The Washington Post retreats from Silicon Valley when it matters most

    6 February 2026

    Spotify is in the business of selling books and adding new audiobook features

    5 February 2026

    Amazon will begin testing AI tools for film and TV production next month

    5 February 2026

    Alexa+, Amazon’s AI assistant, is now available to everyone in the US

    4 February 2026
  • Security

    Substack confirms that the data breach affects users’ email addresses and phone numbers

    6 February 2026

    One of Europe’s biggest universities was offline for days after the cyber attack

    6 February 2026

    Cyber ​​tech giant Conduent’s hot air balloon data breach affects millions more Americans

    5 February 2026

    Hackers Release Personal Information Stolen During Harvard, UPenn Data Breach

    5 February 2026

    French police investigate X office in Paris, call in Elon Musk for questioning

    4 February 2026
  • Startups

    Fundamental raises $255 million in Series A with a new approach to big data analytics

    6 February 2026

    a16z VC wants founders to stop stressing about crazy ARR numbers

    6 February 2026

    Lunar Energy raises $232 million to develop home batteries that support the grid

    5 February 2026

    Meet Gizmo: A TikTok for vibe-coded interactive mini-apps

    5 February 2026

    India’s Varaha wins $20M to scale up carbon removal from Global South

    4 February 2026
  • Transportation

    Apeiron Labs Takes $9.5M to Flood Oceans with Autonomous Underwater Robots

    5 February 2026

    Uber appoints new CFO as its AV plans accelerate

    5 February 2026

    Skyryse lands another $300 million to make flying, even helicopters, simple and safe

    4 February 2026

    China is leading the fight against hidden car door handles

    3 February 2026

    Waymo raises $16 billion to scale robotaxi fleet globally

    3 February 2026
  • Venture

    Secondary sales are shifting from founders’ windfalls to employee retention tools

    6 February 2026

    Sapiom Raises $15M to Help AI Agents Buy Their Own Tech Tools

    6 February 2026

    What a16z actually funds (and what it ignores) when it comes to AI infra

    5 February 2026

    Plans 2026: What’s Next for Startup Battlefield 200

    4 February 2026

    Minneapolis tech community holds strong in ‘tense and difficult times’

    4 February 2026
  • Recommended Essentials
TechTost
You are at:Home»AI»A new, provocative Agi Test Stumps Most AI models
AI

A new, provocative Agi Test Stumps Most AI models

techtost.comBy techtost.com25 March 202503 Mins Read
Share Facebook Twitter Pinterest LinkedIn Tumblr Email
A New, Provocative Agi Test Stumps Most Ai Models
Share
Facebook Twitter LinkedIn Pinterest Email

The Arc Prize Foundation, a non -profit institution by the prominent researcher of AI François Chollet, announced to a blog On Monday that he created a new, provocative test to measure the general intelligence of AI models.

So far, the new test, called the ARC-AGI-2, has thrown most models.

AI models such as OPENAI’S O1-PRO R1 and R1’s R1 between 1% and 1.3% in ARC-AGI-2, according to The Leaderboard of the ARC Award. Strong models that do not project, including GPT-4.5, Claude 3.7 Sonnet and Gemini 2.0 Flash rated about 1%.

ARC-AGI tests consist of puzzles-like problems, where an AI has to identify visual patterns from a collection of different colors and create the right “answer” grid. The problems were designed to force an AI to adapt to new problems that he has not seen before.

The Arc Prize Foundation had over 400 people to get ARC-AGI-2 to create a human basic line. On average, these people’s “panels” got 60% of the test questions correctly – much better than any of the models’ scores.

Sample Question by ARC-AGI-2 (Credit: Arc Prize).

To one Post in xChollet claimed that Arc-AGI-2 is a better measure of the real intelligence of an AI model than the first repetition of the test, ARC-AGI-1. The ARC Foundation Award Tests aim to evaluate whether an AI system can effectively acquire new skills out of the data in which it has been trained.

Chollet said that unlike ARC-AGI-1, the new test prevents AI models from relying on “Brute Force”-expanded computing power-to find solutions. Chollet previously acknowledged that this was an important ARC-AGI-1 defect.

To cope with the defects of the first test, ARC-AGI-2 introduces a new measurement: performance. It also requires models to interpret patterns in relation to flight instead of being based on memorization.

“Intelligence is not exclusively determined by the ability to solve problems or to achieve high ratings,” writes the co -founder of Arc Prize Foundation Greg Kamradt in a blog. “The efficiency with which these capabilities are acquired and developed is a critical, decisive element. The basic question asked is not only” AI may obtain [the] Skill to solve a task? “But also,” in what performance or cost? “

ARC-AGI-1 was undefeated for about five years until December 2024, when Openai released the advanced model of reasoning, which exceeded all other AI models and fits human performance in evaluation. However, as we noted at that time, the O3 efficiency earnings in ARC-AGI-1 came with a heavy price.

The version of the O3 Openai-O3 (Low) (Low)-which was the first to reach ARC-AGI-1 new heights, scoring 75.7% in the test, got a 4% reserve on ARC-AGI-2 using a $ 200 computing power.

Comparison of AI Frontier Model performance in Arc-AGI-1 and ARC-AGI-2 (Credit: ARC Award).

The arrival of the ARC-AGI-2 comes to those in the technological industry demanding new, unsaturated reference points to measure AI progress. The co -founder of Hugging Face, Thomas Wolf, recently told Techcrunch that the AI ​​industry does not have sufficient tests to measure the key features of the so -called artificial general intelligence, including creativity.

Along with the new reference point, the Arc Prize Foundation announced A new ARC 2025 Award Contestcausing developers to reach 85% accuracy in the ARC-AGI-2 test, while spending only $ 0.42 per job.

AGI models Provocative Saint Stumps test
Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
Previous ArticleGoogle launches new functions related to health care to search, Android
Next Article French company VC Founders Future Plans USA
bhanuprakash.cg
techtost.com
  • Website

Related Posts

Amazon and Google are winning the AI ​​capital race — but what’s the prize?

6 February 2026

AWS revenue continues to grow as cloud demand remains high

5 February 2026

Sam Altman tested Claude’s Super Bowl commercials brilliantly

5 February 2026
Add A Comment

Leave A Reply Cancel Reply

Don't Miss

Kindle Scribe Colorsoft is an expensive but beautiful color e-ink tablet with AI features

6 February 2026

Spotify’s new feature lets you explore the story behind the song you’re listening to

6 February 2026

Substack confirms that the data breach affects users’ email addresses and phone numbers

6 February 2026
Stay In Touch
  • Facebook
  • YouTube
  • TikTok
  • WhatsApp
  • Twitter
  • Instagram
Fintech

Stripe Alumni Raise €30M Series A for Duna, Backed by Stripe and Adyen Executives

5 February 2026

Fintech CEO and Forbes 30 Under 30 alum indicted for alleged fraud

3 February 2026

How Sequoia-backed Ethos went public while rivals lagged behind

30 January 2026
Startups

Fundamental raises $255 million in Series A with a new approach to big data analytics

a16z VC wants founders to stop stressing about crazy ARR numbers

Lunar Energy raises $232 million to develop home batteries that support the grid

© 2026 TechTost. All Rights Reserved
  • About Us
  • Contact Us
  • Privacy Policy
  • Terms and Conditions
  • Disclaimer

Type above and press Enter to search. Press Esc to cancel.