Close Menu
TechTost
  • AI
  • Apps
  • Crypto
  • Fintech
  • Hardware
  • Media & Entertainment
  • Security
  • Startups
  • Transportation
  • Venture
  • Recommended Essentials
What's Hot

Spotify will let you edit your taste profile to control your recommendations

Chinese brain interface startup Gestala raises $21 million just two months after launching

Kinetic robotics joins Uber’s Vegas app two years after major reset

Facebook X (Twitter) Instagram
  • About Us
  • Contact Us
  • Privacy Policy
  • Terms and Conditions
  • Disclaimer
Facebook X (Twitter) Instagram
TechTost
Subscribe Now
  • AI

    Before quantum computing arrives, this startup wants businesses that are already working on it

    13 March 2026

    How to watch Jensen Huang’s Nvidia GTC 2026 keynote

    13 March 2026

    Ford’s new AI assistant will help fleet owners know if seat belts are being used

    12 March 2026

    AI ‘Actress’ Tilly Norwood Releases Worst Song I’ve Ever Heard

    12 March 2026

    AI apps struggle with long-term retention, according to a new report

    11 March 2026
  • Apps

    Truecaller now lets you hang up on scammers — on behalf of your family

    13 March 2026

    Channel Surfer lets you watch YouTube like it’s old-school cable TV

    13 March 2026

    Google Maps is getting an AI ‘Ask Maps’ feature and upgraded ‘immersive’ navigation

    12 March 2026

    Google Play adds new paid and PC games, game tests, community posts and more

    12 March 2026

    Google brings Gemini to Chrome in India

    11 March 2026
  • Crypto

    Hackers stole over $2.7 billion in crypto in 2025, data shows

    23 December 2025

    New report examines how David Sachs may benefit from Trump administration role

    1 December 2025

    Why Benchmark Made a Rare Crypto Bet on Trading App Fomo, with $17M Series A

    6 November 2025

    Solana co-founder Anatoly Yakovenko is a big fan of agentic coding

    30 October 2025

    MoviePass opens Mogul fantasy league game to the public

    29 October 2025
  • Fintech

    India neobank Fi removes banking services on its platform

    11 March 2026

    X taps William Shatner to give invitations to his payment service, X Money

    4 March 2026

    Stripe wants to turn your AI costs into a profit center

    3 March 2026

    3 days left: Save up to $680 on your ticket to Disrupt 2026

    25 February 2026

    More startups surpass $10M ARR in 3 months than ever before

    24 February 2026
  • Hardware

    Ex-Apple Engineer Raises $5M for Note-Taking Locket That Only Records Your Voice

    12 March 2026

    Canopii seems to succeed where the old indoor farms failed

    11 March 2026

    Hyperscale Power is the latest startup to challenge 140-year-old transformer technology

    10 March 2026

    Whoop is launching a new blood test focused on women’s health

    10 March 2026

    Honor says its ‘Robot phone’ with moving camera can dance to music

    8 March 2026
  • Media & Entertainment

    Spotify will let you edit your taste profile to control your recommendations

    13 March 2026

    Disney+ launches TikTok-style short-form video stream ‘Verts’

    13 March 2026

    Substack launches an embedded recording studio

    12 March 2026

    TikTok now allows Apple Music subscribers to play entire songs without leaving the app

    12 March 2026

    WordPress debuts a private workspace that runs in your browser via a new service, my.WordPress.net

    11 March 2026
  • Security

    Law enforcement shuts down botnet consisting of tens of thousands of hacked routers

    12 March 2026

    The pro-Iranian hacktivist group says it is behind the attack on medical technology giant Stryker

    12 March 2026

    Salt Typhoon hacks the world’s phone and internet giants — here’s where they’ve been hit

    11 March 2026

    DOGE employee stole Social Security data and thumbed it, report says

    11 March 2026

    US military contractor likely built iPhone hacking tools used by Russian spies in Ukraine

    10 March 2026
  • Startups

    Chinese brain interface startup Gestala raises $21 million just two months after launching

    13 March 2026

    Sales automation startup Rox AI hits $1.2 billion valuation, sources say

    13 March 2026

    When startups become a family business

    12 March 2026

    Ride-hailing inDrive acquires Pakistan’s Krave Mart to boost grocery delivery

    12 March 2026

    Google completes $32 billion acquisition of cloud cybersecurity startup Wiz

    11 March 2026
  • Transportation

    Kinetic robotics joins Uber’s Vegas app two years after major reset

    13 March 2026

    Why Rivian is holding onto the $45,000 R2 base model until ‘late 2027’

    13 March 2026

    Group14 opens factory to produce flash charge battery materials for EVs

    12 March 2026

    Nuro is testing its autonomous vehicle technology on the streets of Tokyo

    12 March 2026

    Zoox plans to put its robotaxis on the Uber app in Vegas this year

    11 March 2026
  • Venture

    Gumloop gets $50M from Benchmark to turn every worker into an AI agent builder

    13 March 2026

    This SpaceX Veteran Says The Next Big Thing In Space Is Satellites Returning To Earth

    10 March 2026

    Founders Fund is approaching $6 billion for its latest growth fund, sources say

    10 March 2026

    Robinhood’s startup fund stumbles in its NYSE debut

    7 March 2026

    City Detect, which uses artificial intelligence to help cities stay safe and clean, raises $13M Series A

    7 March 2026
  • Recommended Essentials
TechTost
You are at:Home»AI»AI models are still struggling to identify errors, the Microsoft study shows
AI

AI models are still struggling to identify errors, the Microsoft study shows

techtost.comBy techtost.com10 April 202503 Mins Read
Share Facebook Twitter Pinterest LinkedIn Tumblr Email
Ai Models Are Still Struggling To Identify Errors, The Microsoft
Share
Facebook Twitter LinkedIn Pinterest Email

AI models from Openai, Anthropic and other top AI laboratories are increasingly used to help with programming work. Google Sundar Pichai Managing Director said in October This 25% of the new code in the company is created by AI and Meta Mark Zuckerberg CEO has expressed aspirations To widely develop AI encoding models within the social media giant.

However, even some of the best models today are struggling to solve software errors that will not travel to experienced devs.

A new study By Microsoft Research, Microsoft’s R&D department, it reveals that models, including Claude 3.7 Sonnet and Openai’s O3-Mini, fail to identify many issues at a reference point for software developed. The results are a disappointing reminder that, rather than daring statements by companies like OpenaiAI still does not match people in areas such as coding.

The co-authors of the study examined nine different models as the backbone for a “prompt agent” that had access to various error detection tools, including a Python bug tracking. Were assigned to this agent by resolving a diligent set of 300 software detection by Swe Bench Lite.

According to co-authors, even when they are equipped with stronger and more recent models, their agent rarely completed more than half of the bugs. Claude 3.7 Sonnet had the highest average success rate (48.4%), followed by O1 (30.2%) and O3-MINI (22.1%).

A graph from the study. The “relative increase” refers to push models from being equipped with bugs.Image credits:Microsoft

Why the sluggish performance? Some models struggled to use the bugs available tools available to them and to understand how different tools could help on different issues. The biggest problem, however, was the lack of data, according to co-authors. They say that there is not enough data representing “successive decision-making processes”-that is, anthropogenic traces of errors-in the training data of today’s models.

‘We firmly believe that training or perfection [models] They can make them better interactive lockers, “the co-authors wrote in their study.” However, this will require specialized data to fulfill this model training, for example, the track data that record the factors interacting with a bug tracking program to collect the necessary information before proposing a mistake repair. “

The findings are not exactly shocking. Many studies have appears This Code AI tends to introduce safety and security errors due to weaknesses in areas such as the ability to understand logical planning. A recent evaluation of DevinA popular AI coding tool found that it could only complete three of the 20 programming tests.

But Microsoft’s work is one of the most detailed appearance in a persistent problem for models. It will probably not weaken investor enthusiasm for the auxiliary coding tools powered by AI, but by chance, it will make the developers-and the highest-ups-are twice to let the AI ​​run the coding show.

For what is worth, a growing number of technology leaders questioned the idea that AI would automate coding work. The co -founder of Microsoft Bill Gates He has said that he believes that planning as a profession He’s here to stay. Thus has Replit CEO Amjad Masad; Okta Todd McKinnon CEO Oktaand IBM Arvind Krishna’s chief executive.

All included Artificial Intelligence detection errors identify Microsoft models Research shows struggling study
Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
Previous ArticleJobandtalent raises $ 103 million in a $ 1.5 valuation.
Next Article A fresh Rolls $ 100 million in Dig Ventures as it is offered to win European early stages
bhanuprakash.cg
techtost.com
  • Website

Related Posts

Before quantum computing arrives, this startup wants businesses that are already working on it

13 March 2026

How to watch Jensen Huang’s Nvidia GTC 2026 keynote

13 March 2026

Ford’s new AI assistant will help fleet owners know if seat belts are being used

12 March 2026
Add A Comment

Leave A Reply Cancel Reply

Don't Miss

Spotify will let you edit your taste profile to control your recommendations

13 March 2026

Chinese brain interface startup Gestala raises $21 million just two months after launching

13 March 2026

Kinetic robotics joins Uber’s Vegas app two years after major reset

13 March 2026
Stay In Touch
  • Facebook
  • YouTube
  • TikTok
  • WhatsApp
  • Twitter
  • Instagram
Fintech

India neobank Fi removes banking services on its platform

11 March 2026

X taps William Shatner to give invitations to his payment service, X Money

4 March 2026

Stripe wants to turn your AI costs into a profit center

3 March 2026
Startups

Chinese brain interface startup Gestala raises $21 million just two months after launching

Sales automation startup Rox AI hits $1.2 billion valuation, sources say

When startups become a family business

© 2026 TechTost. All Rights Reserved
  • About Us
  • Contact Us
  • Privacy Policy
  • Terms and Conditions
  • Disclaimer

Type above and press Enter to search. Press Esc to cancel.