Close Menu
TechTost
  • AI
  • Apps
  • Crypto
  • Fintech
  • Hardware
  • Media & Entertainment
  • Security
  • Startups
  • Transportation
  • Venture
  • Recommended Essentials
What's Hot

Massachusetts votes in favor of new privacy bill that bans sale of precise location data

Zepto’s IPO filing reveals fast growth, bigger losses and a valuation question no one has yet answered

Rivian begins deliveries of its all-important R2 SUV

Facebook X (Twitter) Instagram
  • About Us
  • Contact Us
  • Privacy Policy
  • Terms and Conditions
  • Disclaimer
Facebook X (Twitter) Instagram
TechTost
Subscribe Now
  • AI

    Sandstone raises $30M to bring AI to in-house legal teams

    9 June 2026

    Because Apple’s slow and steady AI bet is starting to look pretty smart

    9 June 2026

    Amazon now lets you design custom merchandise using AI

    8 June 2026

    Mira Murati comes back to the fore, cautiously

    8 June 2026

    The Trump administration may take an equity stake in OpenAI

    7 June 2026
  • Apps

    Apple says it can remove some apps from the App Store if they don’t attract users

    9 June 2026

    Apple’s WWDC AI demos seemed more real after $250 million false ad settlement

    9 June 2026

    The new update of NotebookLM will help you to create source repository from chat

    8 June 2026

    X caters to creators with the new “React with Video” feature.

    8 June 2026

    Meta’s AI agent for WhatsApp Business is now available globally

    7 June 2026
  • Crypto

    Startup Battlefield 200 applications close today

    27 May 2026

    5 days left: Save up to $410 on Disrupt 2026 passes

    25 May 2026

    As crypto cools, a16z crypto raises $2.2 billion in capital

    6 May 2026

    Coinbase to lay off 14% of staff as part of broader restructuring

    5 May 2026

    British cryptographer Adam Back denies NYT report that he is Bitcoin creator Satoshi Nakamoto

    9 April 2026
  • Fintech

    Ramp raises $750M at $44B valuation as investors thirst for fintechs with AI history

    5 June 2026

    Last 24 hours to save up to $410 on your Disrupt 2026 ticket

    29 May 2026

    2 days left: Lock in up to $410 in ticket savings for Disrupt 2026

    28 May 2026

    Robinhood now allows your AI agents to trade stocks

    28 May 2026

    Disrupt 2026 Early Bird ticket savings expire in 3 days

    27 May 2026
  • Hardware

    WWDC 2026: What to expect, from Siri’s long-awaited revamp to Apple Intelligence and iOS 27

    9 June 2026

    What to expect from WWDC 2026: The long-awaited Siri refresh and Apple Intelligence updates

    7 June 2026

    What to expect from WWDC 2026: The long-awaited Siri refresh and Apple Intelligence updates

    5 June 2026

    Oura Ring 5 review: Thinner, lighter, better

    4 June 2026

    Meta mercifully released the VR fitness game Supernatural instead of just killing it

    4 June 2026
  • Media & Entertainment

    Plex adds new social features ahead of major price hike for its lifetime pass

    6 June 2026

    Startup Battlefield 200 applications officially close in 3 days

    5 June 2026

    Founders Fund Launches Series of Games Starring Sam Altman, Palmer Luckey and Other Tech Elites

    5 June 2026

    Meet Wander, a StumbleUpon-inspired tool for discovering the ‘small web’

    4 June 2026

    Publishers will be able to opt out of AI Search, thanks to the new setting

    4 June 2026
  • Security

    Massachusetts votes in favor of new privacy bill that bans sale of precise location data

    9 June 2026

    WhatsApp says it has detected new spyware attacks linked to the NSO group in violation of a court order

    9 June 2026

    Microsoft’s open source tools hacked to steal AI developers’ passwords

    8 June 2026

    Hacked, leaked and held for ransom: the worst breaches of 2026 so far

    7 June 2026

    Google and FBI warn of ransomware group sending fake IT workers to hack victims in person

    6 June 2026
  • Startups

    Zepto’s IPO filing reveals fast growth, bigger losses and a valuation question no one has yet answered

    9 June 2026

    How to apply to Startup Battlefield 2026, what you need before today’s June 8 deadline

    8 June 2026

    Sam Altman-backed fusion startup Helion raises $465M to build power plant for Microsoft

    6 June 2026

    Supabase doubles valuation to $10 billion in 8 months

    5 June 2026

    Startup Battlefield is back in Australia — here’s what happened last time we came to Sydney

    5 June 2026
  • Transportation

    Rivian begins deliveries of its all-important R2 SUV

    9 June 2026

    Waymo bought Apple’s self-driving car for $220 million

    9 June 2026

    Uber, Wayve and Waymo are heading for a robot showdown in London

    8 June 2026

    TechCrunch Mobility: Inside GM’s $900 Million EV Battery Bet

    7 June 2026

    As VC-backed e-bike startups went bankrupt, Lectric by bootstraps grew

    6 June 2026
  • Venture

    Mercor’s Brendan Foody calls out Sequoia, accusing it of “double pricing” valuation tricks.

    9 June 2026

    Founders share VC horror stories and some name names

    6 June 2026

    Defense technology, artificial intelligence and fundraising take center stage at StrictlyVC Los Angeles

    5 June 2026

    Benchmark raises its first growth capital as part of $2 billion capital raising

    4 June 2026

    Former Meta CTO Raises $250 Million Climate Fund

    3 June 2026
  • Recommended Essentials
TechTost
You are at:Home»AI»Openai’s new reasoning, AI models admit more
AI

Openai’s new reasoning, AI models admit more

techtost.comBy techtost.com19 April 202504 Mins Read
Share Facebook Twitter Pinterest LinkedIn Tumblr Email
Openai's New Reasoning, Ai Models Admit More
Share
Facebook Twitter LinkedIn Pinterest Email

Openai models recently started O3 and O4-MINI AI are state-of-the-art in many ways. However, new models are still paid or do things – in fact, paid off more From many of the older models of Openai.

Halfuses have been shown to be one of the biggest and most difficult problems to solve AI, even affecting today’s better -performance systems. Historically, every new model has improved slightly in the illusory section, less than its predecessor. But this does not seem to happen for O3 and O4-mini.

According to Openai’s internal tests, O3 and O4-MINI, which are so-called Reasoning models, illusions more often From previous reasoning models of the company-O1, O1-mini and O3-mini-as the traditional models of Openai, such as the GPT-4O.

Perhaps more about, the Chatgpt manufacturer doesn’t really know why it happens.

In his technical report for O3 and o4-miniOpenai writes that “more research is needed” to understand why hallucinations get worse as they scale models of reasoning. O3 and O4-MINI better attribute to certain areas, including coding and mathematics. But because they “make more claims overall”, they often lead to “more accurate allegations as well as more inaccurate/parable claims,” ​​according to the report.

Openai found that the O3 was assigned in response to 33% of the personqa questions, the company’s internal reference point to measure the accuracy of a model of a model for humans. This is about twice the illusion rate of previous Openai, O1 and O3-MINI reasoning models, which recorded 16% and 14.8% respectively. O4-mini even gets worse in Personqa-rendering 48% of the time.

Third trial With Transluce, a non -profit AI research workshop, he also found that O3 tends to compose actions it took in the process of arriving in answers. In an example, Transluce observed the O3 claiming that it ran the code to a 2021 MacBook Pro “outside the chatgpt”, then copies the numbers to its answer. While O3 has access to some tools, it cannot do so.

“Our hypothesis is that the type of aid learning used for models in the O series can strengthen issues that are usually mitigated (but not fully deleted) by standard pipelines after training,” said Neil Chowdhury, a translated researcher and former Openai employee in an email.

Sarah Schwettmann, co -founder of Transluce, added that the O3 illusion rate can make it less useful than it would be.

Kian Katanforoosh, Professor and Managing Director of Stanford, Stanford, told TechCrunch that his team is already testing the O3 in coding flows and found it to be one step above the competition. However, Katanforosh says that O3 tends to give up broken site links. The model will provide a link that, when clicking, does not work.

Halfuses can help models reach interesting ideas and be creative in their “thinking”, but they also make some models a harsh sale for shopping in markets where accuracy is primary. For example, a law firm would probably not be happy with a model that introduces many real errors into customer contracts.

A very promising approach to enhance the accuracy of their models gives web search opportunities. Openai’s GPT-4O with tissue search achieves Accuracy of 90% In Simpleqa, another of the reference points of Openai’s accuracy. Perhaps the search could also improve the illusion rates of logic models, at least in cases where users are willing to expose the suggestions to a third search provider.

If the escalation of the reasoning models continues to aggravate hallucinations, it will make hunting for an even more urgent solution.

“Tackling the hallucinations in all our models is an ongoing research sector and we are constantly working to improve their accuracy and reliability,” Openai Niko Felix spokesman said in an email in TechCrunch.

Last year, the wider AI industry has rotated to focus on reasoning models after techniques to improve traditional AI models has begun to show reduced yields. Reason improves the performance of the model in a variety of work without requiring huge amounts of computers and data during training. However, it seems that reasoning can also lead to more illusions – presenting a challenge.

admit ChatGPT hallucinations models open OpenAIs Reasoning
Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
Previous ArticleMark Zuckerberg says Tiktok has slowed the Meta growth
Next Article Subaru makes trailseeker debut, an electric SUV coming for Rivian’s Outdoorsy EV Base Outdoorsy EV
bhanuprakash.cg
techtost.com
  • Website

Related Posts

Sandstone raises $30M to bring AI to in-house legal teams

9 June 2026

Because Apple’s slow and steady AI bet is starting to look pretty smart

9 June 2026

Microsoft’s open source tools hacked to steal AI developers’ passwords

8 June 2026
Add A Comment

Leave A Reply Cancel Reply

Don't Miss

Massachusetts votes in favor of new privacy bill that bans sale of precise location data

9 June 2026

Zepto’s IPO filing reveals fast growth, bigger losses and a valuation question no one has yet answered

9 June 2026

Rivian begins deliveries of its all-important R2 SUV

9 June 2026
Stay In Touch
  • Facebook
  • YouTube
  • TikTok
  • WhatsApp
  • Twitter
  • Instagram
Fintech

Ramp raises $750M at $44B valuation as investors thirst for fintechs with AI history

5 June 2026

Last 24 hours to save up to $410 on your Disrupt 2026 ticket

29 May 2026

2 days left: Lock in up to $410 in ticket savings for Disrupt 2026

28 May 2026
Startups

Zepto’s IPO filing reveals fast growth, bigger losses and a valuation question no one has yet answered

How to apply to Startup Battlefield 2026, what you need before today’s June 8 deadline

Sam Altman-backed fusion startup Helion raises $465M to build power plant for Microsoft

© 2026 TechTost. All Rights Reserved
  • About Us
  • Contact Us
  • Privacy Policy
  • Terms and Conditions
  • Disclaimer

Type above and press Enter to search. Press Esc to cancel.