Close Menu
TechTost
  • AI
  • Apps
  • Crypto
  • Fintech
  • Hardware
  • Media & Entertainment
  • Security
  • Startups
  • Transportation
  • Venture
  • Recommended Essentials
What's Hot

Honor launches its new slim foldable Magic V6 with a 6,600 mAh battery

Billion dollar infrastructure deals are fueling the AI ​​boom

X tries to attract advertisers by letting them reuse creatives created for other platforms

Facebook X (Twitter) Instagram
  • About Us
  • Contact Us
  • Privacy Policy
  • Terms and Conditions
  • Disclaimer
Facebook X (Twitter) Instagram
TechTost
Subscribe Now
  • AI

    Billion dollar infrastructure deals are fueling the AI ​​boom

    1 March 2026

    Musk slams OpenAI in deposition, says ‘no one killed themselves because of Grok’

    28 February 2026

    Pentagon moves to designate Anthropic as a supply chain risk

    28 February 2026

    Anthropic CEO stands firm as Pentagon deadline looms

    27 February 2026

    Jack Dorsey just halved the size of Block’s employee base — and he says your company is next

    27 February 2026
  • Apps

    X tries to attract advertisers by letting them reuse creatives created for other platforms

    1 March 2026

    Google launches Nano Banana 2 model with faster image generation

    1 March 2026

    South Korea is opening the door to allow Google Maps to be fully operational

    28 February 2026

    Spotify releases audiobook maps

    28 February 2026

    Bumble adds AI photo feedback and profile guidance tools

    27 February 2026
  • Crypto

    Hackers stole over $2.7 billion in crypto in 2025, data shows

    23 December 2025

    New report examines how David Sachs may benefit from Trump administration role

    1 December 2025

    Why Benchmark Made a Rare Crypto Bet on Trading App Fomo, with $17M Series A

    6 November 2025

    Solana co-founder Anatoly Yakovenko is a big fan of agentic coding

    30 October 2025

    MoviePass opens Mogul fantasy league game to the public

    29 October 2025
  • Fintech

    3 days left: Save up to $680 on your ticket to Disrupt 2026

    25 February 2026

    More startups surpass $10M ARR in 3 months than ever before

    24 February 2026

    Stripe, PayPal Ventures Bet on India’s Xflow to Fix Cross-Border B2B Payments

    24 February 2026

    InScope raises $14.5M to solve financial reporting pain

    20 February 2026

    OpenAI deepens India push with Pine Labs fintech partnership

    19 February 2026
  • Hardware

    Honor launches its new slim foldable Magic V6 with a 6,600 mAh battery

    1 March 2026

    Xiaomi launches 17 Ultra smartphones, an AirTag clone and an ultra-thin powerbank

    28 February 2026

    Last 24 hours to get Disrupt 2026 tickets at the lowest prices of the year

    27 February 2026

    Everything announced at Samsung’s Galaxy Unpacked event, including S26 smartphones, privacy screen and more

    26 February 2026

    Samsung introduces new display technology that adds a privacy screen to apps and notifications

    25 February 2026
  • Media & Entertainment

    What you need to know about Warner Bros.’ landmark Discovery sale

    1 March 2026

    Apple and Netflix team up to stream Formula 1 Canadian Grand Prix

    27 February 2026

    Netflix pulls out of bid for Warner Bros. Discovery, giving studios, HBO and CNN to Ellison-owned Paramount

    27 February 2026

    Book the best deals for Disrupt 2026 | TechCrunch

    26 February 2026

    Americans now listen to podcasts more often than talk radio, study shows

    25 February 2026
  • Security

    The resulting data breach is growing, affecting at least 25 million people

    28 February 2026

    India cuts off access to popular developer platform Supabase with block order

    28 February 2026

    CISA replaces deputy director after a difficult year on the job

    27 February 2026

    Cisco Says Hackers Are Exploiting Critical Flaw To Break Into Large Customer Networks By 2023

    26 February 2026

    US cybersecurity agency CISA reportedly in dire straits amid Trump cuts and layoffs

    26 February 2026
  • Startups

    Why China’s humanoid robot industry is winning the early market

    1 March 2026

    Jest, a marketplace for messaging games, is challenging the app store status quo

    28 February 2026

    Superhuman bets on redesigned smart ring to win back US market after Oura controversy

    27 February 2026

    Trace raises $3 million to solve AI agent adoption in the enterprise

    27 February 2026

    How to avoid bad hires in early stage startups

    26 February 2026
  • Transportation

    Self-driving truck startup Einride raises $113M PIPE ahead of public debut

    27 February 2026

    It’s time to pull the plug on plug-in hybrids

    26 February 2026

    Harbinger acquires self-driving company Phantom AI

    26 February 2026

    Waymo robotaxis are now operating in 10 US cities

    25 February 2026

    Self-driving tech startup Wayve raises $1.2 billion from Nvidia, Uber and three automakers

    25 February 2026
  • Venture

    After Zomato, Deepinder Goyal is back with a $54 million brain-monitoring bet

    28 February 2026

    Dive into Boston’s startup ecosystem at Founder Summit 2026 | TechCrunch

    27 February 2026

    A VC and some big-name developers are trying to solve the open source funding problem, permanently

    27 February 2026

    Y Combinator grad and AI insurance brokerage Harper raises $47 million

    26 February 2026

    Anthropic acquires AI startup Vercept after Meta indicts one of its founders

    26 February 2026
  • Recommended Essentials
TechTost
You are at:Home»AI»Openai’s new reasoning, AI models admit more
AI

Openai’s new reasoning, AI models admit more

techtost.comBy techtost.com19 April 202504 Mins Read
Share Facebook Twitter Pinterest LinkedIn Tumblr Email
Openai's New Reasoning, Ai Models Admit More
Share
Facebook Twitter LinkedIn Pinterest Email

Openai models recently started O3 and O4-MINI AI are state-of-the-art in many ways. However, new models are still paid or do things – in fact, paid off more From many of the older models of Openai.

Halfuses have been shown to be one of the biggest and most difficult problems to solve AI, even affecting today’s better -performance systems. Historically, every new model has improved slightly in the illusory section, less than its predecessor. But this does not seem to happen for O3 and O4-mini.

According to Openai’s internal tests, O3 and O4-MINI, which are so-called Reasoning models, illusions more often From previous reasoning models of the company-O1, O1-mini and O3-mini-as the traditional models of Openai, such as the GPT-4O.

Perhaps more about, the Chatgpt manufacturer doesn’t really know why it happens.

In his technical report for O3 and o4-miniOpenai writes that “more research is needed” to understand why hallucinations get worse as they scale models of reasoning. O3 and O4-MINI better attribute to certain areas, including coding and mathematics. But because they “make more claims overall”, they often lead to “more accurate allegations as well as more inaccurate/parable claims,” ​​according to the report.

Openai found that the O3 was assigned in response to 33% of the personqa questions, the company’s internal reference point to measure the accuracy of a model of a model for humans. This is about twice the illusion rate of previous Openai, O1 and O3-MINI reasoning models, which recorded 16% and 14.8% respectively. O4-mini even gets worse in Personqa-rendering 48% of the time.

Third trial With Transluce, a non -profit AI research workshop, he also found that O3 tends to compose actions it took in the process of arriving in answers. In an example, Transluce observed the O3 claiming that it ran the code to a 2021 MacBook Pro “outside the chatgpt”, then copies the numbers to its answer. While O3 has access to some tools, it cannot do so.

“Our hypothesis is that the type of aid learning used for models in the O series can strengthen issues that are usually mitigated (but not fully deleted) by standard pipelines after training,” said Neil Chowdhury, a translated researcher and former Openai employee in an email.

Sarah Schwettmann, co -founder of Transluce, added that the O3 illusion rate can make it less useful than it would be.

Kian Katanforoosh, Professor and Managing Director of Stanford, Stanford, told TechCrunch that his team is already testing the O3 in coding flows and found it to be one step above the competition. However, Katanforosh says that O3 tends to give up broken site links. The model will provide a link that, when clicking, does not work.

Halfuses can help models reach interesting ideas and be creative in their “thinking”, but they also make some models a harsh sale for shopping in markets where accuracy is primary. For example, a law firm would probably not be happy with a model that introduces many real errors into customer contracts.

A very promising approach to enhance the accuracy of their models gives web search opportunities. Openai’s GPT-4O with tissue search achieves Accuracy of 90% In Simpleqa, another of the reference points of Openai’s accuracy. Perhaps the search could also improve the illusion rates of logic models, at least in cases where users are willing to expose the suggestions to a third search provider.

If the escalation of the reasoning models continues to aggravate hallucinations, it will make hunting for an even more urgent solution.

“Tackling the hallucinations in all our models is an ongoing research sector and we are constantly working to improve their accuracy and reliability,” Openai Niko Felix spokesman said in an email in TechCrunch.

Last year, the wider AI industry has rotated to focus on reasoning models after techniques to improve traditional AI models has begun to show reduced yields. Reason improves the performance of the model in a variety of work without requiring huge amounts of computers and data during training. However, it seems that reasoning can also lead to more illusions – presenting a challenge.

admit ChatGPT hallucinations models open OpenAIs Reasoning
Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
Previous ArticleMark Zuckerberg says Tiktok has slowed the Meta growth
Next Article Subaru makes trailseeker debut, an electric SUV coming for Rivian’s Outdoorsy EV Base Outdoorsy EV
bhanuprakash.cg
techtost.com
  • Website

Related Posts

Billion dollar infrastructure deals are fueling the AI ​​boom

1 March 2026

Musk slams OpenAI in deposition, says ‘no one killed themselves because of Grok’

28 February 2026

Pentagon moves to designate Anthropic as a supply chain risk

28 February 2026
Add A Comment

Leave A Reply Cancel Reply

Don't Miss

Honor launches its new slim foldable Magic V6 with a 6,600 mAh battery

1 March 2026

Billion dollar infrastructure deals are fueling the AI ​​boom

1 March 2026

X tries to attract advertisers by letting them reuse creatives created for other platforms

1 March 2026
Stay In Touch
  • Facebook
  • YouTube
  • TikTok
  • WhatsApp
  • Twitter
  • Instagram
Fintech

3 days left: Save up to $680 on your ticket to Disrupt 2026

25 February 2026

More startups surpass $10M ARR in 3 months than ever before

24 February 2026

Stripe, PayPal Ventures Bet on India’s Xflow to Fix Cross-Border B2B Payments

24 February 2026
Startups

Why China’s humanoid robot industry is winning the early market

Jest, a marketplace for messaging games, is challenging the app store status quo

Superhuman bets on redesigned smart ring to win back US market after Oura controversy

© 2026 TechTost. All Rights Reserved
  • About Us
  • Contact Us
  • Privacy Policy
  • Terms and Conditions
  • Disclaimer

Type above and press Enter to search. Press Esc to cancel.