Close Menu
TechTost
  • AI
  • Apps
  • Crypto
  • Fintech
  • Hardware
  • Media & Entertainment
  • Security
  • Startups
  • Transportation
  • Venture
  • Recommended Essentials
What's Hot

Esther and Anne Wojcicki support new healthcare accelerator, fund

Tesla just increased its spending plan to $25 billion — this is where the money is going

Keep up with X’s new AI-powered custom streams

Facebook X (Twitter) Instagram
  • About Us
  • Contact Us
  • Privacy Policy
  • Terms and Conditions
  • Disclaimer
Facebook X (Twitter) Instagram
TechTost
Subscribe Now
  • AI

    Tesla just increased its spending plan to $25 billion — this is where the money is going

    23 April 2026

    OpenAI partners with Infosys to bring AI tools to more businesses

    22 April 2026

    Unauthorized group gained access to Anthropic’s proprietary Mythos cyber tool, report claims

    22 April 2026

    NSA Spies Reportedly Using Anthropic’s Mythos, Despite Pentagon Controversy

    21 April 2026

    It’s not just one thing – it’s another thing

    21 April 2026
  • Apps

    Keep up with X’s new AI-powered custom streams

    23 April 2026

    X makes it more expensive to publish links through its API

    22 April 2026

    Apple’s Cal AI crackdown signals it still controls the App Store

    22 April 2026

    GRAI believes that AI can make music more social, not replace artists

    21 April 2026

    WhatsApp is testing a premium subscription, but it’s mostly cosmetic

    21 April 2026
  • Crypto

    British cryptographer Adam Back denies NYT report that he is Bitcoin creator Satoshi Nakamoto

    9 April 2026

    Hackers stole over $2.7 billion in crypto in 2025, data shows

    23 December 2025

    New report examines how David Sachs may benefit from Trump administration role

    1 December 2025

    Why Benchmark Made a Rare Crypto Bet on Trading App Fomo, with $17M Series A

    6 November 2025

    Solana co-founder Anatoly Yakovenko is a big fan of agentic coding

    30 October 2025
  • Fintech

    Cash App targets a new type of customer: children aged 6 to 12 years

    22 April 2026

    Revolut eyes up to $200 billion valuation in potential IPO

    22 April 2026

    Once close enough for a takeover, Stripe and Airwallex are now going after each other

    18 April 2026

    Airwallex is set to take on Stripe and the rest of the payments industry — in the physical world

    16 April 2026

    Cash app launches ‘pay later’ feature for P2P transfers

    3 April 2026
  • Hardware

    Apple’s John Ternus will run one of the most powerful companies in the world. work is a minefield

    22 April 2026

    Tim Cook steps down as Apple CEO: Here’s a look at his 15-year legacy, from new products and services to China expansion

    22 April 2026

    Who is John Ternus, the new CEO of Apple?

    21 April 2026

    Tim Cook steps down as Apple CEO, while John Ternus takes over

    21 April 2026

    Amazon Unveils Slimmer Fire TV Stick HD, Opens Ember Artline TVs for Pre-Order

    16 April 2026
  • Media & Entertainment

    YouTube extends its AI similarity detection technology to celebrities

    21 April 2026

    Deezer says 44% of songs uploaded to its platform every day are created with artificial intelligence

    20 April 2026

    Netflix plans to add a vertical video stream, use AI for recommendations

    17 April 2026

    Netflix co-founder and chairman Reed Hastings is stepping down from the board

    17 April 2026

    All we like is soulfulness

    16 April 2026
  • Security

    Apple fixes bug used by police to extract deleted chat messages from iPhones

    22 April 2026

    As US spy laws expire, lawmakers divided over protecting Americans from warrantless surveillance

    22 April 2026

    Ransomware dealer pleads guilty to helping ransomware gang

    21 April 2026

    App host Vercel says it was hacked and customer data stolen

    21 April 2026

    Mastodon says its flagship server has been hit by a DDoS attack

    20 April 2026
  • Startups

    Cathie Woods’ ARK makes first major investment in startup Lucra — and it’s not AI

    22 April 2026

    AI research lab NeoCognition offers $40 million to build agents that learn like humans

    22 April 2026

    You’ve heard of hybrid cars. Now meet a hybrid cement plant.

    19 April 2026

    Loop raises $95 million to build supply chain artificial intelligence that predicts disruptions

    18 April 2026

    Sources: Runner in talks to raise $2B+ at $50B valuation as business grows

    18 April 2026
  • Transportation

    Redwood Materials lays off 10% in restructuring to pursue energy storage business

    22 April 2026

    Amazon taps Sweden’s Einride for its electric big rigs

    21 April 2026

    The Rivian factory was hit by a tornado before the R2 was released

    20 April 2026

    TechCrunch Mobility: Uber enters the era of assetmaxxing

    20 April 2026

    Uber will now collect your returns from your doorstep

    17 April 2026
  • Venture

    Esther and Anne Wojcicki support new healthcare accelerator, fund

    23 April 2026

    Anthropic rejects VC funding that values ​​it at $800B+, for now

    16 April 2026

    Financial risk management platform Pillar raises $20 million in rounds led by a16z

    15 April 2026

    Vercel CEO Guillermo Rauch signals IPO readiness as AI agents drive revenue

    14 April 2026

    Nvidia-backed SiFive hits $3.65 billion valuation for open AI chips

    11 April 2026
  • Recommended Essentials
TechTost
You are at:Home»AI»AI models are still struggling to identify errors, the Microsoft study shows
AI

AI models are still struggling to identify errors, the Microsoft study shows

techtost.comBy techtost.com10 April 202503 Mins Read
Share Facebook Twitter Pinterest LinkedIn Tumblr Email
Ai Models Are Still Struggling To Identify Errors, The Microsoft
Share
Facebook Twitter LinkedIn Pinterest Email

AI models from Openai, Anthropic and other top AI laboratories are increasingly used to help with programming work. Google Sundar Pichai Managing Director said in October This 25% of the new code in the company is created by AI and Meta Mark Zuckerberg CEO has expressed aspirations To widely develop AI encoding models within the social media giant.

However, even some of the best models today are struggling to solve software errors that will not travel to experienced devs.

A new study By Microsoft Research, Microsoft’s R&D department, it reveals that models, including Claude 3.7 Sonnet and Openai’s O3-Mini, fail to identify many issues at a reference point for software developed. The results are a disappointing reminder that, rather than daring statements by companies like OpenaiAI still does not match people in areas such as coding.

The co-authors of the study examined nine different models as the backbone for a “prompt agent” that had access to various error detection tools, including a Python bug tracking. Were assigned to this agent by resolving a diligent set of 300 software detection by Swe Bench Lite.

According to co-authors, even when they are equipped with stronger and more recent models, their agent rarely completed more than half of the bugs. Claude 3.7 Sonnet had the highest average success rate (48.4%), followed by O1 (30.2%) and O3-MINI (22.1%).

A graph from the study. The “relative increase” refers to push models from being equipped with bugs.Image credits:Microsoft

Why the sluggish performance? Some models struggled to use the bugs available tools available to them and to understand how different tools could help on different issues. The biggest problem, however, was the lack of data, according to co-authors. They say that there is not enough data representing “successive decision-making processes”-that is, anthropogenic traces of errors-in the training data of today’s models.

‘We firmly believe that training or perfection [models] They can make them better interactive lockers, “the co-authors wrote in their study.” However, this will require specialized data to fulfill this model training, for example, the track data that record the factors interacting with a bug tracking program to collect the necessary information before proposing a mistake repair. “

The findings are not exactly shocking. Many studies have appears This Code AI tends to introduce safety and security errors due to weaknesses in areas such as the ability to understand logical planning. A recent evaluation of DevinA popular AI coding tool found that it could only complete three of the 20 programming tests.

But Microsoft’s work is one of the most detailed appearance in a persistent problem for models. It will probably not weaken investor enthusiasm for the auxiliary coding tools powered by AI, but by chance, it will make the developers-and the highest-ups-are twice to let the AI ​​run the coding show.

For what is worth, a growing number of technology leaders questioned the idea that AI would automate coding work. The co -founder of Microsoft Bill Gates He has said that he believes that planning as a profession He’s here to stay. Thus has Replit CEO Amjad Masad; Okta Todd McKinnon CEO Oktaand IBM Arvind Krishna’s chief executive.

All included Artificial Intelligence detection errors identify Microsoft models Research shows struggling study
Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
Previous ArticleJobandtalent raises $ 103 million in a $ 1.5 valuation.
Next Article A fresh Rolls $ 100 million in Dig Ventures as it is offered to win European early stages
bhanuprakash.cg
techtost.com
  • Website

Related Posts

Esther and Anne Wojcicki support new healthcare accelerator, fund

23 April 2026

Tesla just increased its spending plan to $25 billion — this is where the money is going

23 April 2026

OpenAI partners with Infosys to bring AI tools to more businesses

22 April 2026
Add A Comment

Leave A Reply Cancel Reply

Don't Miss

Esther and Anne Wojcicki support new healthcare accelerator, fund

23 April 2026

Tesla just increased its spending plan to $25 billion — this is where the money is going

23 April 2026

Keep up with X’s new AI-powered custom streams

23 April 2026
Stay In Touch
  • Facebook
  • YouTube
  • TikTok
  • WhatsApp
  • Twitter
  • Instagram
Fintech

Cash App targets a new type of customer: children aged 6 to 12 years

22 April 2026

Revolut eyes up to $200 billion valuation in potential IPO

22 April 2026

Once close enough for a takeover, Stripe and Airwallex are now going after each other

18 April 2026
Startups

Cathie Woods’ ARK makes first major investment in startup Lucra — and it’s not AI

AI research lab NeoCognition offers $40 million to build agents that learn like humans

You’ve heard of hybrid cars. Now meet a hybrid cement plant.

© 2026 TechTost. All Rights Reserved
  • About Us
  • Contact Us
  • Privacy Policy
  • Terms and Conditions
  • Disclaimer

Type above and press Enter to search. Press Esc to cancel.