Close Menu
TechTost
  • AI
  • Apps
  • Crypto
  • Fintech
  • Hardware
  • Media & Entertainment
  • Security
  • Startups
  • Transportation
  • Venture
  • Recommended Essentials
What's Hot

Anodot hack leaves over a dozen compromised companies facing extortion

Uber and Nuro begin testing premium robotaxi service in San Francisco

Vercel CEO Guillermo Rauch signals IPO readiness as AI agents drive revenue

Facebook X (Twitter) Instagram
  • About Us
  • Contact Us
  • Privacy Policy
  • Terms and Conditions
  • Disclaimer
Facebook X (Twitter) Instagram
TechTost
Subscribe Now
  • AI

    OpenAI has acquired AI personal finance startup Hiro

    14 April 2026

    Largest orbital computing cluster is open for business

    13 April 2026

    Anthropic restricts Mythos traffic to protect the Internet — or does Anthropic?

    12 April 2026

    Sam Altman responds to ‘inflammatory’ New Yorker article after his home was attacked

    12 April 2026

    Stalking victim sues OpenAI, claims ChatGPT fueled her abuser’s delusions and ignored her warnings

    11 April 2026
  • Apps

    Avec’s Tinder-style email app lets you swipe through your inbox

    14 April 2026

    Roblox introduces ‘Kids’ and ‘Select’ accounts for age-appropriate access to games and chats

    13 April 2026

    You can now edit your comments on Instagram

    13 April 2026

    Meta AI app climbs to No. 5 in App Store after release of Muse Spark

    12 April 2026

    StubHub to pay $10 million to settle FTC claims of ‘deceptive’ ticket pricing

    12 April 2026
  • Crypto

    British cryptographer Adam Back denies NYT report that he is Bitcoin creator Satoshi Nakamoto

    9 April 2026

    Hackers stole over $2.7 billion in crypto in 2025, data shows

    23 December 2025

    New report examines how David Sachs may benefit from Trump administration role

    1 December 2025

    Why Benchmark Made a Rare Crypto Bet on Trading App Fomo, with $17M Series A

    6 November 2025

    Solana co-founder Anatoly Yakovenko is a big fan of agentic coding

    30 October 2025
  • Fintech

    Cash app launches ‘pay later’ feature for P2P transfers

    3 April 2026

    Doss raises $55 million for AI inventory management that connects to ERP

    24 March 2026

    Despite stiff competition, Kalshi, Polymarket CEOs back $35m VC fund projections

    23 March 2026

    Amid legal turmoil, Kalshi is temporarily banned in Nevada

    20 March 2026

    Nominations for the Startup Battlefield 200 are still open

    19 March 2026
  • Hardware

    Amazon is ending support for older Kindle devices

    9 April 2026

    Intel signs Elon Musk’s Terafab chip project

    8 April 2026

    The Xiaomi 17 Ultra has some impressive extras that make taking photos really fun

    6 April 2026

    In Japan, the robot doesn’t come for your job. fills the one no one wants

    6 April 2026

    Peter Thiel’s big bet on solar-powered cow collars

    5 April 2026
  • Media & Entertainment

    X says he’s reducing payouts to clickbait accounts

    12 April 2026

    TechCrunch is headed to Tokyo — and it’s bringing the Startup Battlefield with it

    10 April 2026

    Spotify now allows everyone to turn off videos in its app

    9 April 2026

    As YouTube expands into TV, it sees more interactive video across all formats

    9 April 2026

    Tubi is the first streamer to launch a native app on ChatGPT

    8 April 2026
  • Security

    Anodot hack leaves over a dozen compromised companies facing extortion

    14 April 2026

    Booking.com confirms that hackers accessed customer data

    13 April 2026

    Convicted spyware maker Bryan Fleming avoids jail time on conviction

    12 April 2026

    The Trump administration plans to cut the cybersecurity agency’s budget by $700 million

    11 April 2026

    Russian government hackers broke into thousands of home routers to steal passwords

    11 April 2026
  • Startups

    Walmart-owned Flipkart, Amazon are squeezing India’s e-commerce startups

    12 April 2026

    This founder helped build SpaceX’s most powerful rocket engine. Now he’s building a “fighter for orbit.”

    12 April 2026

    Sierra’s Bret Taylor says the era of button-clicking is over

    11 April 2026

    After the data breach, the $10 billion startup Mercor is one month old

    11 April 2026

    What founders can learn from Anjuna’s layoffs and recovery

    10 April 2026
  • Transportation

    Uber and Nuro begin testing premium robotaxi service in San Francisco

    14 April 2026

    Slate Auto raises $650 million to fund its affordable EV truck plans

    13 April 2026

    TechCrunch Mobility: Who’s chasing all the self-driving talent?

    13 April 2026

    Slate Auto: Everything you need to know about the Bezos-backed EV startup

    12 April 2026

    Battery recycling company Ascend Elements files for bankruptcy

    11 April 2026
  • Venture

    Vercel CEO Guillermo Rauch signals IPO readiness as AI agents drive revenue

    14 April 2026

    Nvidia-backed SiFive hits $3.65 billion valuation for open AI chips

    11 April 2026

    How to make the Startup Battlefield Top 20 — and what each company gets regardless

    10 April 2026

    Collide Capital Raises $95M to Back Future-of-Work Fintech Startups

    9 April 2026

    VC Eclipse has a new $1.3 billion fund to back — and build — “natural AI” startups

    8 April 2026
  • Recommended Essentials
TechTost
You are at:Home»AI»A new, provocative Agi Test Stumps Most AI models
AI

A new, provocative Agi Test Stumps Most AI models

techtost.comBy techtost.com25 March 202503 Mins Read
Share Facebook Twitter Pinterest LinkedIn Tumblr Email
A New, Provocative Agi Test Stumps Most Ai Models
Share
Facebook Twitter LinkedIn Pinterest Email

The Arc Prize Foundation, a non -profit institution by the prominent researcher of AI François Chollet, announced to a blog On Monday that he created a new, provocative test to measure the general intelligence of AI models.

So far, the new test, called the ARC-AGI-2, has thrown most models.

AI models such as OPENAI’S O1-PRO R1 and R1’s R1 between 1% and 1.3% in ARC-AGI-2, according to The Leaderboard of the ARC Award. Strong models that do not project, including GPT-4.5, Claude 3.7 Sonnet and Gemini 2.0 Flash rated about 1%.

ARC-AGI tests consist of puzzles-like problems, where an AI has to identify visual patterns from a collection of different colors and create the right “answer” grid. The problems were designed to force an AI to adapt to new problems that he has not seen before.

The Arc Prize Foundation had over 400 people to get ARC-AGI-2 to create a human basic line. On average, these people’s “panels” got 60% of the test questions correctly – much better than any of the models’ scores.

Sample Question by ARC-AGI-2 (Credit: Arc Prize).

To one Post in xChollet claimed that Arc-AGI-2 is a better measure of the real intelligence of an AI model than the first repetition of the test, ARC-AGI-1. The ARC Foundation Award Tests aim to evaluate whether an AI system can effectively acquire new skills out of the data in which it has been trained.

Chollet said that unlike ARC-AGI-1, the new test prevents AI models from relying on “Brute Force”-expanded computing power-to find solutions. Chollet previously acknowledged that this was an important ARC-AGI-1 defect.

To cope with the defects of the first test, ARC-AGI-2 introduces a new measurement: performance. It also requires models to interpret patterns in relation to flight instead of being based on memorization.

“Intelligence is not exclusively determined by the ability to solve problems or to achieve high ratings,” writes the co -founder of Arc Prize Foundation Greg Kamradt in a blog. “The efficiency with which these capabilities are acquired and developed is a critical, decisive element. The basic question asked is not only” AI may obtain [the] Skill to solve a task? “But also,” in what performance or cost? “

ARC-AGI-1 was undefeated for about five years until December 2024, when Openai released the advanced model of reasoning, which exceeded all other AI models and fits human performance in evaluation. However, as we noted at that time, the O3 efficiency earnings in ARC-AGI-1 came with a heavy price.

The version of the O3 Openai-O3 (Low) (Low)-which was the first to reach ARC-AGI-1 new heights, scoring 75.7% in the test, got a 4% reserve on ARC-AGI-2 using a $ 200 computing power.

Comparison of AI Frontier Model performance in Arc-AGI-1 and ARC-AGI-2 (Credit: ARC Award).

The arrival of the ARC-AGI-2 comes to those in the technological industry demanding new, unsaturated reference points to measure AI progress. The co -founder of Hugging Face, Thomas Wolf, recently told Techcrunch that the AI ​​industry does not have sufficient tests to measure the key features of the so -called artificial general intelligence, including creativity.

Along with the new reference point, the Arc Prize Foundation announced A new ARC 2025 Award Contestcausing developers to reach 85% accuracy in the ARC-AGI-2 test, while spending only $ 0.42 per job.

AGI models Provocative Saint Stumps test
Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
Previous ArticleGoogle launches new functions related to health care to search, Android
Next Article French company VC Founders Future Plans USA
bhanuprakash.cg
techtost.com
  • Website

Related Posts

OpenAI has acquired AI personal finance startup Hiro

14 April 2026

Largest orbital computing cluster is open for business

13 April 2026

Anthropic restricts Mythos traffic to protect the Internet — or does Anthropic?

12 April 2026
Add A Comment

Leave A Reply Cancel Reply

Don't Miss

Anodot hack leaves over a dozen compromised companies facing extortion

14 April 2026

Uber and Nuro begin testing premium robotaxi service in San Francisco

14 April 2026

Vercel CEO Guillermo Rauch signals IPO readiness as AI agents drive revenue

14 April 2026
Stay In Touch
  • Facebook
  • YouTube
  • TikTok
  • WhatsApp
  • Twitter
  • Instagram
Fintech

Cash app launches ‘pay later’ feature for P2P transfers

3 April 2026

Doss raises $55 million for AI inventory management that connects to ERP

24 March 2026

Despite stiff competition, Kalshi, Polymarket CEOs back $35m VC fund projections

23 March 2026
Startups

Walmart-owned Flipkart, Amazon are squeezing India’s e-commerce startups

This founder helped build SpaceX’s most powerful rocket engine. Now he’s building a “fighter for orbit.”

Sierra’s Bret Taylor says the era of button-clicking is over

© 2026 TechTost. All Rights Reserved
  • About Us
  • Contact Us
  • Privacy Policy
  • Terms and Conditions
  • Disclaimer

Type above and press Enter to search. Press Esc to cancel.