Close Menu
TechTost
  • AI
  • Apps
  • Crypto
  • Fintech
  • Hardware
  • Media & Entertainment
  • Security
  • Startups
  • Transportation
  • Venture
  • Recommended Essentials
What's Hot

Meet the MacBook Neo, Apple’s colorful answer to the Chromebook, starting at $599

Decagon Completes First Auction at $4.5B Value

Anthropic CEO Dario Amodei calls OpenAI’s messages about military deal ‘outright lies’, report says

Facebook X (Twitter) Instagram
  • About Us
  • Contact Us
  • Privacy Policy
  • Terms and Conditions
  • Disclaimer
Facebook X (Twitter) Instagram
TechTost
Subscribe Now
  • AI

    Anthropic CEO Dario Amodei calls OpenAI’s messages about military deal ‘outright lies’, report says

    5 March 2026

    Who needs data centers in space when they can float on the high seas?

    4 March 2026

    Why AI startups are selling the same capital at two different prices

    4 March 2026

    Users are abandoning ChatGPT for Claude — see how you can make the switch

    3 March 2026

    No one has a good plan for how AI companies should work with government

    3 March 2026
  • Apps

    Google settles with Epic Games, cuts Play Store commissions to 20%

    5 March 2026

    Android users can now share tracking tag information with airlines to help track lost luggage

    4 March 2026

    ChatGPT’s new GPT-5.3 Instant model will stop telling you to calm down

    4 March 2026

    X adds “Paid Partnership” tags so creators can skip hashtags

    3 March 2026

    ChatGPT uninstalls increased 295% after DoD settlement

    3 March 2026
  • Crypto

    Hackers stole over $2.7 billion in crypto in 2025, data shows

    23 December 2025

    New report examines how David Sachs may benefit from Trump administration role

    1 December 2025

    Why Benchmark Made a Rare Crypto Bet on Trading App Fomo, with $17M Series A

    6 November 2025

    Solana co-founder Anatoly Yakovenko is a big fan of agentic coding

    30 October 2025

    MoviePass opens Mogul fantasy league game to the public

    29 October 2025
  • Fintech

    X taps William Shatner to give invitations to his payment service, X Money

    4 March 2026

    Stripe wants to turn your AI costs into a profit center

    3 March 2026

    3 days left: Save up to $680 on your ticket to Disrupt 2026

    25 February 2026

    More startups surpass $10M ARR in 3 months than ever before

    24 February 2026

    Stripe, PayPal Ventures Bet on India’s Xflow to Fix Cross-Border B2B Payments

    24 February 2026
  • Hardware

    Meet the MacBook Neo, Apple’s colorful answer to the Chromebook, starting at $599

    5 March 2026

    MacBook Neo, iPhone 17e and everything else Apple announced this week

    4 March 2026

    Apple’s new Studio monitors come with Thunderbolt 5

    4 March 2026

    Apple unveils new MacBook Air and MacBook Pro with M5

    3 March 2026

    Apple is packing the smarts into its new $599 iPhone 17e

    3 March 2026
  • Media & Entertainment

    Audible launches cheaper ‘Standard’ subscription plan, challenging Spotify

    3 March 2026

    Paramount+ and HBO Max will merge into one streaming service after the WBD deal closes

    2 March 2026

    What you need to know about Warner Bros.’ landmark Discovery sale

    1 March 2026

    Apple and Netflix team up to stream Formula 1 Canadian Grand Prix

    27 February 2026

    Netflix pulls out of bid for Warner Bros. Discovery, giving studios, HBO and CNN to Ellison-owned Paramount

    27 February 2026
  • Security

    Hackers and internet outages hit Iran amid US airstrikes

    4 March 2026

    A suite of government hacking tools targeting iPhones is now being used by cybercriminals

    4 March 2026

    Hacked Traffic Cameras and Hacked TVs: How Cyber ​​Operations Supported the War on Iran

    3 March 2026

    A new app alerts you if someone nearby is wearing smart glasses

    3 March 2026

    Hacktivists claim to have breached Homeland Security to release ICE contract data

    2 March 2026
  • Startups

    Decagon Completes First Auction at $4.5B Value

    5 March 2026

    MyFitnessPal has acquired Cal AI, the calorie app built by teenagers

    4 March 2026

    Fig Security emerges from stealth with $38 million to help security teams deal with change

    4 March 2026

    A married founding duo’s company, 14.ai, is replacing customer support teams at startups

    3 March 2026

    India’s Pronto takes home help official as valuation grows 8x in less than a year

    3 March 2026
  • Transportation

    Self-driving truck startup Einride raises $113M PIPE ahead of public debut

    27 February 2026

    It’s time to pull the plug on plug-in hybrids

    26 February 2026

    Harbinger acquires self-driving company Phantom AI

    26 February 2026

    Waymo robotaxis are now operating in 10 US cities

    25 February 2026

    Self-driving tech startup Wayve raises $1.2 billion from Nvidia, Uber and three automakers

    25 February 2026
  • Venture

    The candidate that Silicon Valley built is now the one they want to tear down

    3 March 2026

    Parade’s Cami Tellez Announces New Creator Economy Marketing Platform, $4M Funding

    3 March 2026

    SaaS in, SaaS out: Here’s what’s driving the SaaSpocalypse

    2 March 2026

    Investors are shedding what they are no longer looking for in AI SaaS companies

    2 March 2026

    After Zomato, Deepinder Goyal is back with a $54 million brain-monitoring bet

    28 February 2026
  • Recommended Essentials
TechTost
You are at:Home»AI»A new AI coding challenge has just published its first results – and is not beautiful
AI

A new AI coding challenge has just published its first results – and is not beautiful

techtost.comBy techtost.com24 July 202503 Mins Read
Share Facebook Twitter Pinterest LinkedIn Tumblr Email
A New Ai Coding Challenge Has Just Published Its First
Share
Facebook Twitter LinkedIn Pinterest Email

A new AI coding challenge revealed his first winner-and set a new bar for AI software engineers.

On Wednesday at 5 pm PST, the Laude Non -Profit Institute announced the first winner of the K award, a multilevel Coding Challenge started by Databricks and co -founder Andy Konwinski. The winner was a Brazilian engineer called Eduardo Rocha de Andrade, who will receive $ 50,000 for the prize. But more amazing than the victory was his final score: he won the right answers to just 7.5% of test questions.

“We are happy to have built a reference point that is really difficult,” Konwinski said. “The benchmarks should be difficult if they are going to matter,” he continued, adding: “The scores would be different if the big laboratories had entered their largest models, but this is the kind of point.

Konwinski is committed to $ 1 million in the first open source model that can rate higher than 90% in the test.

Similar to the well -known Swench system, the K Award Tests models against signs of Github issues as a test for how good models can deal with real world planning problems. However, while the Swench is based on a stable set of problems that can train models, the K award is designed as “version without SWENCH infection”, using a timed input system to protect against any special reference training. For the first round, the models are due to March 12th. The organizers of the K award then built the test using only GitHub issues highlighted after this date.

The 7.5% top score is intense in contrast to Swe Bench itself, which currently shows a top 75% top score in the easiest “verified” test and 34% of the toughest “complete” test. Konwinski is still not sure if inequality is due to the infection in the Swench or simply to challenge the collection of new issues from Github, but expects that the K will soon answer the question.

“As we have more routes of the thing, we will have a better feel,” he told TechCrunch, “because we expect people to adapt to the dynamics of competition every few months.”

TechCrunch event

Francisco
|
27-29 October 2025

It may seem like a strange place to remain, given the wide range of AI coding tools that are already available to the public – but with reference points to become very easy, many critics see projects such as the K as a necessary step towards resolving The growing AI evaluation problem.

“I am quite refreshing to build new tests for existing reference points,” says Princeton Sayash Kapoor researcher, who presented a similar idea In a recent document. “Without such experiments, we can’t really say if the issue is infection, or even just aiming at the table with man with a man in the loop.”

For Konwinski, it’s not just a better point of reference, but an open challenge for the rest of the industry. “If you hear the advertising campaign, it’s like seeing AI doctors and AI lawyers and AI software engineers, and that’s not true,” he says. “If we can’t even get more than 10% in a cooling infection, this is the control of reality for me.”

Andy Konwinski beautiful challenge Coding K prize Laude Institute published results
Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
Previous ArticleOpen Source X opponent Mastodon begins to raise funds with new in -app donation feature
Next Article Former Y Combinator, A16Z experts hold a summit for founders only for founders
bhanuprakash.cg
techtost.com
  • Website

Related Posts

Anthropic CEO Dario Amodei calls OpenAI’s messages about military deal ‘outright lies’, report says

5 March 2026

Who needs data centers in space when they can float on the high seas?

4 March 2026

Why AI startups are selling the same capital at two different prices

4 March 2026
Add A Comment

Leave A Reply Cancel Reply

Don't Miss

Meet the MacBook Neo, Apple’s colorful answer to the Chromebook, starting at $599

5 March 2026

Decagon Completes First Auction at $4.5B Value

5 March 2026

Anthropic CEO Dario Amodei calls OpenAI’s messages about military deal ‘outright lies’, report says

5 March 2026
Stay In Touch
  • Facebook
  • YouTube
  • TikTok
  • WhatsApp
  • Twitter
  • Instagram
Fintech

X taps William Shatner to give invitations to his payment service, X Money

4 March 2026

Stripe wants to turn your AI costs into a profit center

3 March 2026

3 days left: Save up to $680 on your ticket to Disrupt 2026

25 February 2026
Startups

Decagon Completes First Auction at $4.5B Value

MyFitnessPal has acquired Cal AI, the calorie app built by teenagers

Fig Security emerges from stealth with $38 million to help security teams deal with change

© 2026 TechTost. All Rights Reserved
  • About Us
  • Contact Us
  • Privacy Policy
  • Terms and Conditions
  • Disclaimer

Type above and press Enter to search. Press Esc to cancel.