Close Menu
TechTost
  • AI
  • Apps
  • Crypto
  • Fintech
  • Hardware
  • Media & Entertainment
  • Security
  • Startups
  • Transportation
  • Venture
  • Recommended Essentials
What's Hot

I hate to love Riverside’s AI-based “Rewind” for podcasters

First Voyage Raises $2.5M For Its Habit-Building AI Companion

Ford is launching a battery storage business to power data centers and the grid

Facebook X (Twitter) Instagram
  • About Us
  • Contact Us
  • Privacy Policy
  • Terms and Conditions
  • Disclaimer
Facebook X (Twitter) Instagram
TechTost
Subscribe Now
  • AI

    Creative Commons announces trial support for ‘pay-to-crawl’ AI systems.

    15 December 2025

    TIME named “Architects of AI” Person of the Year

    15 December 2025

    Runway releases its first global model, adds native audio to latest video model

    14 December 2025

    OpenAI hits back at Google with GPT-5.2 after ‘code red’ memo.

    14 December 2025

    Trump’s AI executive order promises ‘a rulebook’ – startups may find legal loophole instead

    13 December 2025
  • Apps

    Google’s ‘dark web reporting’ feature will no longer be available from February

    15 December 2025

    WhatsApp’s biggest market becomes the toughest test

    15 December 2025

    Google debuts ‘Disco’, a Gemini-powered tool for building web apps from browser tabs

    14 December 2025

    Google’s AI testing feature for clothes now only works with a selfie

    14 December 2025

    DoorDash driver faces felony charges after allegedly spraying customers’ food

    13 December 2025
  • Crypto

    New report examines how David Sachs may benefit from Trump administration role

    1 December 2025

    Why Benchmark Made a Rare Crypto Bet on Trading App Fomo, with $17M Series A

    6 November 2025

    Solana co-founder Anatoly Yakovenko is a big fan of agentic coding

    30 October 2025

    MoviePass opens Mogul fantasy league game to the public

    29 October 2025

    Only 5 days until Disrupt 2025 sets the startup world on fire

    22 October 2025
  • Fintech

    Coinbase starts onboarding users again in India, plans to do fiat on-ramp next year

    7 December 2025

    Walmart-backed PhonePe shuts down Pincode app in yet another step back in e-commerce

    5 December 2025

    Nexus stays out of AI, keeping half of its new $700M fund for India startup

    4 December 2025

    Fintech firm Marquis notifies dozens of US banks and credit unions of data breach after ransomware attack

    3 December 2025

    Revolut hits $75 billion valuation in new capital raise

    24 November 2025
  • Hardware

    Nvidia is reportedly weighing increasing H200 production to meet growing demand in China

    15 December 2025

    Pebble founder unveils $75 AI smart ring to record short notes with the push of a button

    10 December 2025

    Amazon’s Ring launches controversial AI-powered facial recognition feature on video doorbells

    10 December 2025

    Google’s first AI glasses are expected next year

    9 December 2025

    eSIM adoption is on the rise thanks to travel and device compatibility

    6 December 2025
  • Media & Entertainment

    I hate to love Riverside’s AI-based “Rewind” for podcasters

    16 December 2025

    Understanding the Dangerous Netflix-Warner Bros. Deal

    15 December 2025

    Disney signs deal with OpenAI to allow Sora to create AI videos with its characters

    11 December 2025

    YouTube TV will launch genre-based subscription plans in 2026

    11 December 2025

    Founder of AI startup Tavus says users talk to AI Santa ‘for hours’ a day

    10 December 2025
  • Security

    The flaw in the photo booth manufacturer’s website exposes customers’ photos

    13 December 2025

    Home Depot exposed access to internal systems for a year, researcher says

    13 December 2025

    Security flaws in the Freedom Chat app exposed users’ phone numbers and PINs

    11 December 2025

    Petco takes down Vetco website after exposing customers’ personal information

    10 December 2025

    Petco’s security bug affected customers’ SSNs, driver’s licenses and more

    9 December 2025
  • Startups

    First Voyage Raises $2.5M For Its Habit-Building AI Companion

    15 December 2025

    Harness hits $5.5B valuation with $240M raise to automate AI’s ‘post-code’ divide

    15 December 2025

    Mesa shuts down credit card that rewards cardholders for paying their mortgages

    14 December 2025

    Port raises $100M valuation from $800M round to take on Spotify’s Backstage

    14 December 2025

    Eclipse Energy’s microbes can turn dormant oil wells into hydrogen factories

    13 December 2025
  • Transportation

    Ford is launching a battery storage business to power data centers and the grid

    15 December 2025

    TechCrunch Mobility: Rivian’s survival plan involves more than cars

    14 December 2025

    India’s Spinny lines up $160m funding to acquire GoMechanic, sources say

    14 December 2025

    Inside Rivian’s big bet on self-driving with artificial intelligence

    13 December 2025

    Zevo wants to add robotaxis to its car-sharing fleet, starting with newcomer Tensor

    13 December 2025
  • Venture

    Lightspeed raises record $9 billion in new capital

    15 December 2025

    Runware raises $50 million in Series A to make it easier for developers to create images and videos

    12 December 2025

    Stanford’s star reporter understands Silicon Valley’s startup culture

    12 December 2025

    The market has “changed” and founders now have the power, VCs say

    11 December 2025

    Tiger Global plans cautious business future with new $2.2 billion fund

    8 December 2025
  • Recommended Essentials
TechTost
You are at:Home»AI»These researchers used NPR Sunday puzzle questions to compare AI ‘Reasoning’ models
AI

These researchers used NPR Sunday puzzle questions to compare AI ‘Reasoning’ models

techtost.comBy techtost.com17 February 202504 Mins Read
Share Facebook Twitter Pinterest LinkedIn Tumblr Email
These Researchers Used Npr Sunday Puzzle Questions To Compare Ai
Share
Facebook Twitter LinkedIn Pinterest Email

Every Sunday, NPR Host Will Shortz, the New York Times Guru Crossword Sunday. While written to be resolved without very Very predictive, brainteasers are usually provocative even for specialized contestants.

That is why some experts believe it is a very promising way to test the limits of AI’s problem solving skills.

To one recent studyA team of researchers from Wellesley College, Oberlin College, Texas University in Austin, Northeastern University, Charles University and Startup Cursor have created a AI reference point using Sunday puzzles. The team says their test has revealed amazing ideas, such as these reasoning models – Openai’s O1, among other things – sometimes “give up” and provide answers that know that they are not right.

“We wanted to develop a reference point with problems that people can only understand with general knowledge,” Arjun Guha, a member of the School of Informatics in Northeastern and one of the co-authors of the study, told TechCrunch.

The AI ​​industry is in a piece of a comparative rating quandary at the moment. Most of the tests commonly used to evaluate the AI ​​models detector for skills, such as the ability to mathematical and scientific questions at a median level, not related to the average user. In the meantime, many reference points – even The reference points were released relatively recently – They quickly approach the saturation point.

The advantages of a public game of radio quiz such as Sunday’s puzzle is that it does not try inner knowledge and challenges are stated so that the models cannot draw from the “Rote” memory to solve them, Guha explained.

“I think what makes these problems hard is that it is really difficult to make real progress in a problem until you solve it – this is everything that clicks together at the same time,” Guha said. “This requires a combination of insight and a process of eliminating.”

No reference is perfect, of course. Sunday’s puzzle is only central and English only. And because the quiz are available to the public, it is likely that the models trained in them can “deceive” in a sense, though Guha says he has not seen proof.

“New questions are circulating every week and we can expect the latest questions are really invisible,” he added. “We intend to maintain the reference point and watch how the performance of the model changes over time.”

Regarding the reference point of researchers, consisting of about 600 Sunday puzzles, reasoning models such as O1 and Deepseek R1 go beyond the rest. The models of reasoning were thoroughly tested, the results themselves before giving results, which helps them avoid some of the traps that normally travel to AI models. The compromise is that the models of reasoning take a little more time to reach solutions-usually seconds to minutes.

At least one model, the Deepseek R1, gives solutions that knows how to be wrong for some of the Sunday questions. The R1 will say literally that “I give up”, followed by an incorrect answer chosen seemingly random – behavior that man can certainly relate.

The models make other strange choices, as well as a wrong answer just to withdraw it immediately, to try to tease a better and fail again. They also stick forever “thinking” forever and give stupid explanations for answers or reach a correct answer immediately, but then continue to consider alternative answers for no apparent reason.

“In harsh problems, the R1 literally says that he gets” frustrated, “Guha said.” It was funny to see how a model mimics what a man could say. It remains to be seen how “frustration” in reasoning can affect the quality of the results of the model. “

R1 gets “frustrated” in a question in the Sunday puzzle challenge set.Image credits:Guha et al.

Today’s best performance model at the reference point is O1 with a 59%rating, followed by the recently released O3-mini set at high “thought” (47%). (R1 noted 35%.) As a next step, researchers plan to expand their tests to additional reasoning models, which hope to help identify areas where these models could be enhanced.

NPR reference index
The results of the models examined by the team at their reference point.Image credits:Guha et al.

“You do not need a doctorate to be good at reasoning, so it should be possible to design reference points that do not require knowledge at the doctorate level,” Guha said. “A reference point with broader access allows a wider set of researchers to understand and analyze the results, which in turn can lead to better solutions in the future. In addition, as the models of the latest technology are increasingly developing in settings That affect everyone, we believe that everyone should be able to intuit what these models are-and they are not-right. “

All included Benchmark compare evergreen modeling model models Npr puzzle Questions Reasoning Research researchers Sunday
Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
Previous ArticleYouTube TV reaches the new deal to maintain Paramount content
Next Article The World Police Operation is attempting 8base Ransomware gang leakage location
bhanuprakash.cg
techtost.com
  • Website

Related Posts

Creative Commons announces trial support for ‘pay-to-crawl’ AI systems.

15 December 2025

TIME named “Architects of AI” Person of the Year

15 December 2025

Runway releases its first global model, adds native audio to latest video model

14 December 2025
Add A Comment

Leave A Reply Cancel Reply

Don't Miss

I hate to love Riverside’s AI-based “Rewind” for podcasters

16 December 2025

First Voyage Raises $2.5M For Its Habit-Building AI Companion

15 December 2025

Ford is launching a battery storage business to power data centers and the grid

15 December 2025
Stay In Touch
  • Facebook
  • YouTube
  • TikTok
  • WhatsApp
  • Twitter
  • Instagram
Fintech

Coinbase starts onboarding users again in India, plans to do fiat on-ramp next year

7 December 2025

Walmart-backed PhonePe shuts down Pincode app in yet another step back in e-commerce

5 December 2025

Nexus stays out of AI, keeping half of its new $700M fund for India startup

4 December 2025
Startups

First Voyage Raises $2.5M For Its Habit-Building AI Companion

Harness hits $5.5B valuation with $240M raise to automate AI’s ‘post-code’ divide

Mesa shuts down credit card that rewards cardholders for paying their mortgages

© 2025 TechTost. All Rights Reserved
  • About Us
  • Contact Us
  • Privacy Policy
  • Terms and Conditions
  • Disclaimer

Type above and press Enter to search. Press Esc to cancel.