Close Menu
TechTost
  • AI
  • Apps
  • Crypto
  • Fintech
  • Hardware
  • Media & Entertainment
  • Security
  • Startups
  • Transportation
  • Venture
  • Recommended Essentials
What's Hot

Uber taps Rivian to build robotaxis in deal worth up to $1.25 billion

Why Wall Street Didn’t Win Nvidia’s Big Conference

Meta finally decides not to close Horizon Worlds in VR

Facebook X (Twitter) Instagram
  • About Us
  • Contact Us
  • Privacy Policy
  • Terms and Conditions
  • Disclaimer
Facebook X (Twitter) Instagram
TechTost
Subscribe Now
  • AI

    Why Wall Street Didn’t Win Nvidia’s Big Conference

    22 March 2026

    New court filing reveals Pentagon told Anthropic the two sides were nearly aligned — a week after Trump declared his relationship

    21 March 2026

    Microsoft is retiring some of the Copilot AI bloat on Windows

    21 March 2026

    The best AI investment may be in energy technology

    20 March 2026

    Bot traffic to overtake human traffic by 2027, says Cloudflare CEO

    20 March 2026
  • Apps

    Meta finally decides not to close Horizon Worlds in VR

    22 March 2026

    DoorDash Launches New ‘Tasks’ App That Pays Couriers to Submit Videos to Train AI

    21 March 2026

    Google is introducing a new way for users to download Android apps that still protects against fraud

    21 March 2026

    Meta launches new AI content enforcement systems while reducing reliance on third-party vendors

    20 March 2026

    Bluesky Announces $100M Series B After CEO Transition

    20 March 2026
  • Crypto

    Hackers stole over $2.7 billion in crypto in 2025, data shows

    23 December 2025

    New report examines how David Sachs may benefit from Trump administration role

    1 December 2025

    Why Benchmark Made a Rare Crypto Bet on Trading App Fomo, with $17M Series A

    6 November 2025

    Solana co-founder Anatoly Yakovenko is a big fan of agentic coding

    30 October 2025

    MoviePass opens Mogul fantasy league game to the public

    29 October 2025
  • Fintech

    Amid legal turmoil, Kalshi is temporarily banned in Nevada

    20 March 2026

    Nominations for the Startup Battlefield 200 are still open

    19 March 2026

    Kalshi’s legal woes pile up as Arizona files first criminal charges for ‘illegal gambling operation’

    17 March 2026

    Fuse raises $25M to disrupt legacy loan origination systems used by US credit unions

    16 March 2026

    India neobank Fi removes banking services on its platform

    11 March 2026
  • Hardware

    Amazon is working on a new smartphone with Alexa at its core, the report says

    20 March 2026

    CEO Carl Pei says nothing about smartphone apps disappearing as they’re replaced by artificial intelligence agents

    18 March 2026

    MacBook Neo, AirPods Max 2, iPhone 17e and everything else Apple announced this month

    18 March 2026

    Oura enters India’s smart ring market with Ring 4

    17 March 2026

    Apple quietly launches AirPods Max 2

    17 March 2026
  • Media & Entertainment

    Tubi joins forces with popular TikTokers to create original streaming content

    19 March 2026

    Patreon CEO calls AI companies’ fair use argument ‘bogus’, says creators should be paid

    18 March 2026

    Meet Vurt, the first mobile streaming platform for indie filmmakers embracing vertical video

    18 March 2026

    BuzzFeed debuts AI applications for new revenue

    17 March 2026

    Facebook makes it easy for creators to report copycats

    14 March 2026
  • Security

    Delve accused of misleading customers with ‘false compliance’

    21 March 2026

    The US accuses the Iranian government of operating a hacktivist group that hacked the Stryker

    20 March 2026

    CISA Urges Companies to Secure Microsoft Intune Systems After Hackers Mass Wipe Stryker Devices

    20 March 2026

    FBI seizes websites of pro-Iranian hacker group after devastating Stryker attack

    19 March 2026

    FBI is buying location data to track US citizens, director confirms

    19 March 2026
  • Startups

    Microsoft hires Sequoia-backed AI collaboration platform team Cove

    21 March 2026

    Consumer-focused privacy firm Cloaked raises $375 million as it expands into the enterprise

    20 March 2026

    Tools for founders to navigate and move past conflicts

    20 March 2026

    Anori, Alphabet’s new X spinout, faces one of the world’s most expensive bureaucratic nightmares

    19 March 2026

    This startup wants to make enterprise software more like a prompt

    19 March 2026
  • Transportation

    Uber taps Rivian to build robotaxis in deal worth up to $1.25 billion

    22 March 2026

    Federal authorities intensify investigation into Tesla’s Full Self-Driving (Supervised) software

    21 March 2026

    Cyberattack on vehicle breathalyzer company leaves drivers stranded in US

    21 March 2026

    Arc expands into electric commercial and defense vessels with $50M raise

    20 March 2026

    Rivian Sacrifices 2027 Profit Target to Push Deeper into Autonomy

    20 March 2026
  • Venture

    AI startups are eating up the venture industry, and the returns, so far, are good

    21 March 2026

    Sequen raised $16 million to bring TikTok-style personalization technology to any consumer company

    19 March 2026

    AI ‘boys club’ could widen wealth gap for women, says Rana el Kaliouby

    18 March 2026

    Billionaires made a promise – now some want to leave

    17 March 2026

    Antonio Gracias Says He Longs For ‘Pre-Entropic’ Startups – Those Built To Survive Chaos

    17 March 2026
  • Recommended Essentials
TechTost
You are at:Home»AI»Openai’s new reasoning, AI models admit more
AI

Openai’s new reasoning, AI models admit more

techtost.comBy techtost.com19 April 202504 Mins Read
Share Facebook Twitter Pinterest LinkedIn Tumblr Email
Openai's New Reasoning, Ai Models Admit More
Share
Facebook Twitter LinkedIn Pinterest Email

Openai models recently started O3 and O4-MINI AI are state-of-the-art in many ways. However, new models are still paid or do things – in fact, paid off more From many of the older models of Openai.

Halfuses have been shown to be one of the biggest and most difficult problems to solve AI, even affecting today’s better -performance systems. Historically, every new model has improved slightly in the illusory section, less than its predecessor. But this does not seem to happen for O3 and O4-mini.

According to Openai’s internal tests, O3 and O4-MINI, which are so-called Reasoning models, illusions more often From previous reasoning models of the company-O1, O1-mini and O3-mini-as the traditional models of Openai, such as the GPT-4O.

Perhaps more about, the Chatgpt manufacturer doesn’t really know why it happens.

In his technical report for O3 and o4-miniOpenai writes that “more research is needed” to understand why hallucinations get worse as they scale models of reasoning. O3 and O4-MINI better attribute to certain areas, including coding and mathematics. But because they “make more claims overall”, they often lead to “more accurate allegations as well as more inaccurate/parable claims,” ​​according to the report.

Openai found that the O3 was assigned in response to 33% of the personqa questions, the company’s internal reference point to measure the accuracy of a model of a model for humans. This is about twice the illusion rate of previous Openai, O1 and O3-MINI reasoning models, which recorded 16% and 14.8% respectively. O4-mini even gets worse in Personqa-rendering 48% of the time.

Third trial With Transluce, a non -profit AI research workshop, he also found that O3 tends to compose actions it took in the process of arriving in answers. In an example, Transluce observed the O3 claiming that it ran the code to a 2021 MacBook Pro “outside the chatgpt”, then copies the numbers to its answer. While O3 has access to some tools, it cannot do so.

“Our hypothesis is that the type of aid learning used for models in the O series can strengthen issues that are usually mitigated (but not fully deleted) by standard pipelines after training,” said Neil Chowdhury, a translated researcher and former Openai employee in an email.

Sarah Schwettmann, co -founder of Transluce, added that the O3 illusion rate can make it less useful than it would be.

Kian Katanforoosh, Professor and Managing Director of Stanford, Stanford, told TechCrunch that his team is already testing the O3 in coding flows and found it to be one step above the competition. However, Katanforosh says that O3 tends to give up broken site links. The model will provide a link that, when clicking, does not work.

Halfuses can help models reach interesting ideas and be creative in their “thinking”, but they also make some models a harsh sale for shopping in markets where accuracy is primary. For example, a law firm would probably not be happy with a model that introduces many real errors into customer contracts.

A very promising approach to enhance the accuracy of their models gives web search opportunities. Openai’s GPT-4O with tissue search achieves Accuracy of 90% In Simpleqa, another of the reference points of Openai’s accuracy. Perhaps the search could also improve the illusion rates of logic models, at least in cases where users are willing to expose the suggestions to a third search provider.

If the escalation of the reasoning models continues to aggravate hallucinations, it will make hunting for an even more urgent solution.

“Tackling the hallucinations in all our models is an ongoing research sector and we are constantly working to improve their accuracy and reliability,” Openai Niko Felix spokesman said in an email in TechCrunch.

Last year, the wider AI industry has rotated to focus on reasoning models after techniques to improve traditional AI models has begun to show reduced yields. Reason improves the performance of the model in a variety of work without requiring huge amounts of computers and data during training. However, it seems that reasoning can also lead to more illusions – presenting a challenge.

admit ChatGPT hallucinations models open OpenAIs Reasoning
Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
Previous ArticleMark Zuckerberg says Tiktok has slowed the Meta growth
Next Article Subaru makes trailseeker debut, an electric SUV coming for Rivian’s Outdoorsy EV Base Outdoorsy EV
bhanuprakash.cg
techtost.com
  • Website

Related Posts

Why Wall Street Didn’t Win Nvidia’s Big Conference

22 March 2026

New court filing reveals Pentagon told Anthropic the two sides were nearly aligned — a week after Trump declared his relationship

21 March 2026

Microsoft is retiring some of the Copilot AI bloat on Windows

21 March 2026
Add A Comment

Leave A Reply Cancel Reply

Don't Miss

Uber taps Rivian to build robotaxis in deal worth up to $1.25 billion

22 March 2026

Why Wall Street Didn’t Win Nvidia’s Big Conference

22 March 2026

Meta finally decides not to close Horizon Worlds in VR

22 March 2026
Stay In Touch
  • Facebook
  • YouTube
  • TikTok
  • WhatsApp
  • Twitter
  • Instagram
Fintech

Amid legal turmoil, Kalshi is temporarily banned in Nevada

20 March 2026

Nominations for the Startup Battlefield 200 are still open

19 March 2026

Kalshi’s legal woes pile up as Arizona files first criminal charges for ‘illegal gambling operation’

17 March 2026
Startups

Microsoft hires Sequoia-backed AI collaboration platform team Cove

Consumer-focused privacy firm Cloaked raises $375 million as it expands into the enterprise

Tools for founders to navigate and move past conflicts

© 2026 TechTost. All Rights Reserved
  • About Us
  • Contact Us
  • Privacy Policy
  • Terms and Conditions
  • Disclaimer

Type above and press Enter to search. Press Esc to cancel.