Close Menu
TechTost
  • AI
  • Apps
  • Crypto
  • Fintech
  • Hardware
  • Media & Entertainment
  • Security
  • Startups
  • Transportation
  • Venture
  • Recommended Essentials
What's Hot

Co-founders behind Reface and Prisma join hands to improve on-device model inference with Mirai

Rivian owners will soon be able to access vehicle controls using their Apple Watch

Ali Partovi’s Neo appears to upgrade the throttle model in low dilution terms

Facebook X (Twitter) Instagram
  • About Us
  • Contact Us
  • Privacy Policy
  • Terms and Conditions
  • Disclaimer
Facebook X (Twitter) Instagram
TechTost
Subscribe Now
  • AI

    ‘Toy Story 5’ takes aim at creepy AI toys: ‘I’m always listening’

    21 February 2026

    Great news for xAI: Grok is now very good at answering questions about Baldur’s Gate

    21 February 2026

    UAE’s G42 partners with Cerebra to deploy 8 exaflops of computers in India

    20 February 2026

    Why these startup CEOs don’t think AI will replace human roles

    20 February 2026

    Reliance unveils $110bn AI investment plan as India boosts tech ambitions

    19 February 2026
  • Apps

    India’s Sarvam launches Indus AI chat app as competition heats up

    21 February 2026

    Remember HQ? “Quiz Daddy” Scott Rogowsky is back with TextSavvy, a daily mobile game show

    21 February 2026

    As the browser war heats up, Chrome is adding new productivity features

    20 February 2026

    Google says its AI systems helped prevent Play Store malware in 2025

    20 February 2026

    Mastodon, a decentralized alternative to X, plans to target creators with new features

    19 February 2026
  • Crypto

    Hackers stole over $2.7 billion in crypto in 2025, data shows

    23 December 2025

    New report examines how David Sachs may benefit from Trump administration role

    1 December 2025

    Why Benchmark Made a Rare Crypto Bet on Trading App Fomo, with $17M Series A

    6 November 2025

    Solana co-founder Anatoly Yakovenko is a big fan of agentic coding

    30 October 2025

    MoviePass opens Mogul fantasy league game to the public

    29 October 2025
  • Fintech

    InScope raises $14.5M to solve financial reporting pain

    20 February 2026

    OpenAI deepens India push with Pine Labs fintech partnership

    19 February 2026

    Cash app adds payment links so you can get paid in DMs

    11 February 2026

    MrBeast’s company buys Gen Z fintech app Step

    9 February 2026

    Stripe Alumni Raise €30M Series A for Duna, Backed by Stripe and Adyen Executives

    5 February 2026
  • Hardware

    Joseph C Belden: Last Chance for Innovators to Earn Scaling Privileges

    20 February 2026

    At a critical time, Snap is losing a top spec executive

    20 February 2026

    Freeform Raises $67M Series B to Scale Laser AI Production

    19 February 2026

    India’s Sarvam wants to bring its AI models to phones, cars and smart glasses

    19 February 2026

    Google debuts $499 Pixel 10a

    18 February 2026
  • Media & Entertainment

    Google adds music-making capabilities to its Gemini app

    21 February 2026

    Disrupt 2026 Super Early Bird pricing expires in 1 week

    20 February 2026

    YouTube’s latest experiment brings its AI chat tool to TVs

    20 February 2026

    OpenAI, Reliance partner to add AI search to JioHotstar

    19 February 2026

    SeatGeek and Spotify are teaming up to offer concert ticket discounts within the music platform

    19 February 2026
  • Security

    Ukrainian man jailed for identity theft that helped North Koreans get jobs at US companies

    21 February 2026

    Cellebrite cut off Serbia citing misuse of its phone unlocking tools. Why not others?

    20 February 2026

    FBI says ATM ‘jackpot’ attacks on the rise, hackers net millions in stolen cash

    20 February 2026

    Sex toy maker Tenga says hacker stole customer information

    19 February 2026

    Hacker conference Def Con bans three people linked to Epstein

    19 February 2026
  • Startups

    Co-founders behind Reface and Prisma join hands to improve on-device model inference with Mirai

    21 February 2026

    Nominations for the Startup Battlefield 200 are now open

    21 February 2026

    The OpenAI mafia: 18 startups founded by graduates

    20 February 2026

    Nvidia deepens early-stage push into India’s AI startup ecosystem

    20 February 2026

    Kana emerges from stealth with $15M to build flexible AI agents for marketers

    19 February 2026
  • Transportation

    Rivian owners will soon be able to access vehicle controls using their Apple Watch

    21 February 2026

    Lucid Motors is cutting 12% of its workforce as it pursues profitability

    21 February 2026

    New York puts the brakes on robotaxi expansion plan

    20 February 2026

    AI data center boom fuels Redwood’s energy storage business

    20 February 2026

    Tesla avoids 30-day suspension in California after removing ‘Autopilot’

    18 February 2026
  • Venture

    Ali Partovi’s Neo appears to upgrade the throttle model in low dilution terms

    21 February 2026

    Peak XV Raises $1.3B, Doubles In AI As Global India VC Competition Heats Up

    21 February 2026

    General Catalyst commits $5 billion to India over five years

    20 February 2026

    Reload wants to give your AI agents a shared memory

    20 February 2026

    This VC’s best advice for building a founding team

    19 February 2026
  • Recommended Essentials
TechTost
You are at:Home»AI»AI models are still struggling to identify errors, the Microsoft study shows
AI

AI models are still struggling to identify errors, the Microsoft study shows

techtost.comBy techtost.com10 April 202503 Mins Read
Share Facebook Twitter Pinterest LinkedIn Tumblr Email
Ai Models Are Still Struggling To Identify Errors, The Microsoft
Share
Facebook Twitter LinkedIn Pinterest Email

AI models from Openai, Anthropic and other top AI laboratories are increasingly used to help with programming work. Google Sundar Pichai Managing Director said in October This 25% of the new code in the company is created by AI and Meta Mark Zuckerberg CEO has expressed aspirations To widely develop AI encoding models within the social media giant.

However, even some of the best models today are struggling to solve software errors that will not travel to experienced devs.

A new study By Microsoft Research, Microsoft’s R&D department, it reveals that models, including Claude 3.7 Sonnet and Openai’s O3-Mini, fail to identify many issues at a reference point for software developed. The results are a disappointing reminder that, rather than daring statements by companies like OpenaiAI still does not match people in areas such as coding.

The co-authors of the study examined nine different models as the backbone for a “prompt agent” that had access to various error detection tools, including a Python bug tracking. Were assigned to this agent by resolving a diligent set of 300 software detection by Swe Bench Lite.

According to co-authors, even when they are equipped with stronger and more recent models, their agent rarely completed more than half of the bugs. Claude 3.7 Sonnet had the highest average success rate (48.4%), followed by O1 (30.2%) and O3-MINI (22.1%).

A graph from the study. The “relative increase” refers to push models from being equipped with bugs.Image credits:Microsoft

Why the sluggish performance? Some models struggled to use the bugs available tools available to them and to understand how different tools could help on different issues. The biggest problem, however, was the lack of data, according to co-authors. They say that there is not enough data representing “successive decision-making processes”-that is, anthropogenic traces of errors-in the training data of today’s models.

‘We firmly believe that training or perfection [models] They can make them better interactive lockers, “the co-authors wrote in their study.” However, this will require specialized data to fulfill this model training, for example, the track data that record the factors interacting with a bug tracking program to collect the necessary information before proposing a mistake repair. “

The findings are not exactly shocking. Many studies have appears This Code AI tends to introduce safety and security errors due to weaknesses in areas such as the ability to understand logical planning. A recent evaluation of DevinA popular AI coding tool found that it could only complete three of the 20 programming tests.

But Microsoft’s work is one of the most detailed appearance in a persistent problem for models. It will probably not weaken investor enthusiasm for the auxiliary coding tools powered by AI, but by chance, it will make the developers-and the highest-ups-are twice to let the AI ​​run the coding show.

For what is worth, a growing number of technology leaders questioned the idea that AI would automate coding work. The co -founder of Microsoft Bill Gates He has said that he believes that planning as a profession He’s here to stay. Thus has Replit CEO Amjad Masad; Okta Todd McKinnon CEO Oktaand IBM Arvind Krishna’s chief executive.

All included Artificial Intelligence detection errors identify Microsoft models Research shows struggling study
Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
Previous ArticleJobandtalent raises $ 103 million in a $ 1.5 valuation.
Next Article A fresh Rolls $ 100 million in Dig Ventures as it is offered to win European early stages
bhanuprakash.cg
techtost.com
  • Website

Related Posts

‘Toy Story 5’ takes aim at creepy AI toys: ‘I’m always listening’

21 February 2026

Great news for xAI: Grok is now very good at answering questions about Baldur’s Gate

21 February 2026

UAE’s G42 partners with Cerebra to deploy 8 exaflops of computers in India

20 February 2026
Add A Comment

Leave A Reply Cancel Reply

Don't Miss

Co-founders behind Reface and Prisma join hands to improve on-device model inference with Mirai

21 February 2026

Rivian owners will soon be able to access vehicle controls using their Apple Watch

21 February 2026

Ali Partovi’s Neo appears to upgrade the throttle model in low dilution terms

21 February 2026
Stay In Touch
  • Facebook
  • YouTube
  • TikTok
  • WhatsApp
  • Twitter
  • Instagram
Fintech

InScope raises $14.5M to solve financial reporting pain

20 February 2026

OpenAI deepens India push with Pine Labs fintech partnership

19 February 2026

Cash app adds payment links so you can get paid in DMs

11 February 2026
Startups

Co-founders behind Reface and Prisma join hands to improve on-device model inference with Mirai

Nominations for the Startup Battlefield 200 are now open

The OpenAI mafia: 18 startups founded by graduates

© 2026 TechTost. All Rights Reserved
  • About Us
  • Contact Us
  • Privacy Policy
  • Terms and Conditions
  • Disclaimer

Type above and press Enter to search. Press Esc to cancel.