Close Menu
TechTost
  • AI
  • Apps
  • Crypto
  • Fintech
  • Hardware
  • Media & Entertainment
  • Security
  • Startups
  • Transportation
  • Venture
  • Recommended Essentials
What's Hot

Gradient’s heat pumps get new smarts to enable retrofitting of old buildings

Peak XV Says Internal Disagreement Has Led to Partner Exits as AI Doubles

New York lawmakers are proposing a three-year freeze on new data centers

Facebook X (Twitter) Instagram
  • About Us
  • Contact Us
  • Privacy Policy
  • Terms and Conditions
  • Disclaimer
Facebook X (Twitter) Instagram
TechTost
Subscribe Now
  • AI

    New York lawmakers are proposing a three-year freeze on new data centers

    7 February 2026

    Benchmark raises $225 million in dedicated funds to double Cerebras

    7 February 2026

    How artificial intelligence is helping to solve the labor issue in treating rare diseases

    6 February 2026

    Amazon and Google are winning the AI ​​capital race — but what’s the prize?

    6 February 2026

    AWS revenue continues to grow as cloud demand remains high

    5 February 2026
  • Apps

    After backlash, Adobe reverses shutdown of Adobe Animate and puts app in ‘maintenance mode’

    7 February 2026

    EU says TikTok must disable ‘addictive’ features like infinite scrolling, fix recommendation engine

    7 February 2026

    Here’s how Roblox’s age controls work

    6 February 2026

    Meta is testing a standalone app for its AI-generated ‘Vibes’ videos

    6 February 2026

    Reddit sees AI search as the next big opportunity

    5 February 2026
  • Crypto

    Hackers stole over $2.7 billion in crypto in 2025, data shows

    23 December 2025

    New report examines how David Sachs may benefit from Trump administration role

    1 December 2025

    Why Benchmark Made a Rare Crypto Bet on Trading App Fomo, with $17M Series A

    6 November 2025

    Solana co-founder Anatoly Yakovenko is a big fan of agentic coding

    30 October 2025

    MoviePass opens Mogul fantasy league game to the public

    29 October 2025
  • Fintech

    Stripe Alumni Raise €30M Series A for Duna, Backed by Stripe and Adyen Executives

    5 February 2026

    Fintech CEO and Forbes 30 Under 30 alum indicted for alleged fraud

    3 February 2026

    How Sequoia-backed Ethos went public while rivals lagged behind

    30 January 2026

    5 days left for TechCrunch Disrupt 2026 +1 pass with 50%

    26 January 2026

    50% off +1 ends | TechCrunch

    23 January 2026
  • Hardware

    Kindle Scribe Colorsoft is an expensive but beautiful color e-ink tablet with AI features

    6 February 2026

    Ring brings “Search Party” feature for finding lost dogs to non-Ring camera owners

    2 February 2026

    India offers zero taxes till 2047 to attract global AI workloads

    1 February 2026

    Microsoft won’t stop buying AI chips from Nvidia, AMD even after its own is released, says Nadella

    30 January 2026

    The iPhone just had its best quarter ever

    30 January 2026
  • Media & Entertainment

    From Svedka to Anthropic, Brands Are Making Bold Plays With AI in Super Bowl Ads

    7 February 2026

    “Industry” Season 4 captures tech fraud better than any show on TV right now

    7 February 2026

    Spotify’s new feature lets you explore the story behind the song you’re listening to

    6 February 2026

    The Washington Post retreats from Silicon Valley when it matters most

    6 February 2026

    Spotify is in the business of selling books and adding new audiobook features

    5 February 2026
  • Security

    Senator, who has repeatedly warned of secret US government surveillance, raises new alarm over ‘CIA activities’

    7 February 2026

    Substack confirms that the data breach affects users’ email addresses and phone numbers

    6 February 2026

    One of Europe’s biggest universities was offline for days after the cyber attack

    6 February 2026

    Cyber ​​tech giant Conduent’s hot air balloon data breach affects millions more Americans

    5 February 2026

    Hackers Release Personal Information Stolen During Harvard, UPenn Data Breach

    5 February 2026
  • Startups

    Gradient’s heat pumps get new smarts to enable retrofitting of old buildings

    8 February 2026

    Accel doubles down on Fibr AI as agents turn static websites into one-to-one experiences

    7 February 2026

    ElevenLabs Raises $500M From Sequoia At $11B Valuation

    7 February 2026

    Fundamental raises $255 million in Series A with a new approach to big data analytics

    6 February 2026

    a16z VC wants founders to stop stressing about crazy ARR numbers

    6 February 2026
  • Transportation

    Prince Andrew’s adviser suggested Jeffrey Epstein invest in EV startups like Lucid Motors

    7 February 2026

    Apeiron Labs Takes $9.5M to Flood Oceans with Autonomous Underwater Robots

    5 February 2026

    Uber appoints new CFO as its AV plans accelerate

    5 February 2026

    Skyryse lands another $300 million to make flying, even helicopters, simple and safe

    4 February 2026

    China is leading the fight against hidden car door handles

    3 February 2026
  • Venture

    Peak XV Says Internal Disagreement Has Led to Partner Exits as AI Doubles

    8 February 2026

    SNAK Venture Partners raises $50 million in capital to support vertical acquisitions

    7 February 2026

    Reddit says it’s looking for more acquisitions in adtech and elsewhere

    7 February 2026

    Secondary sales are shifting from founders’ windfalls to employee retention tools

    6 February 2026

    Sapiom Raises $15M to Help AI Agents Buy Their Own Tech Tools

    6 February 2026
  • Recommended Essentials
TechTost
You are at:Home»Media & Entertainment»A high-ranking set up a website that allows you to challenge AI models in a build-off Minecraft
Media & Entertainment

A high-ranking set up a website that allows you to challenge AI models in a build-off Minecraft

techtost.comBy techtost.com21 March 202503 Mins Read
Share Facebook Twitter Pinterest LinkedIn Tumblr Email
A High Ranking Set Up A Website That Allows You To
Share
Facebook Twitter LinkedIn Pinterest Email

As conventional AI comparative evaluation techniques prove inadequate, AI builders turn to more creative ways to evaluate the potential of AI genetic models. For a team of developers, this is Minecraft, the Microsoft Sandboild manufacturing game.

The website Benchmark (or Mc Bench) was collaborated with AI models against the other in head challenges to meet Minecraft creations. Users can vote on which model they did a better job and only after the vote can they see what AI made each minecraft make.

Image credits:Benchmark (opens in a new window)

For Adi Singh, the 12th grader who started the Mc Bench, the value of Minecraft is not as much the game itself, but the familiarization of people with the end, is the with the best sales Video game of all time. Even for people who have not played the game, it is still possible to evaluate which blocked representation of a pineapple is getting better.

‘Minecraft allows people to see progress [of AI development] Much easier, “Singh told TechCrunch.” People are used to Minecraft, used for appearance and vibe. ”

MC Bench currently lists eight people as volunteers. The anthropogenic, Google, Openai and Alibaba subsidized the use of their products by the project to execute reference prompts per MC Bench website, but companies are not different.

“Right now we make simple constructions to think about how far we have come from the GPT-3 season but [we] We could see ourselves in these greater form plans and target -oriented tasks, “Singh said.” Games can only be a means of testing the logic that is safer than in real life and more controlled for testing purposes, making it the most ideal in my eyes. “

Other games like Pokémon Red, Road fighterAnd Pictionary has been used as experimental reference points for AI, in part because the art of comparative AI evaluation is strangely difficult.

Researchers often try AI models standardized evaluationsBut many of these tests give AI an advantage at home. Due to the way they are trained, the models are naturally endowed with certain, narrow problem solving, particularly solving problems that require memorization or basic extension.

Simply put, it is difficult to collect what it means that OpenAi’s GPT-4 can score at 88th percentage in LSAT, but it cannot discern how many Rs are in the word “strawberry”. Anthropogenic Claude 3.7 Sonnet A 62.3% accuracy was achieved in a standardized engineering point of reference, but it is worse to play Pokémon than most five years.

Image credits:Benchmark

MC Bench is technically a programming point, since models are called upon to write code to create the caused construction, such as “Frosty the Snowman” or “a charming tropical beach hut on a virgin sandy shore”.

But it is easier for most Mc Bench users to evaluate if a snowman seems better than discovering the code, which gives the project a wider appeal-and therefore the ability to collect more data on models that are consistently scoring.

Whether these scores are largely in the way of using AI is of course for discussion, of course. Singh claims to be a powerful signal.

“Today’s leaderboard is quite closely reflecting my own experience of using these models, which is in contrast to many pure text reference points,” Singh said. “Perhaps [MC-Bench] It could be useful for companies to know if they are heading in the right direction. ”

buildoff challenge highranking Minecraft models set website Wine
Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
Previous ArticleHackers increase attacks using Servicenow Servicenow errors to target systems that are not targeted
Next Article Mortgage as the benefit of employees? Kleiner Perkins leads to a series of $ 23.5 million for Multiply MortGage
bhanuprakash.cg
techtost.com
  • Website

Related Posts

From Svedka to Anthropic, Brands Are Making Bold Plays With AI in Super Bowl Ads

7 February 2026

“Industry” Season 4 captures tech fraud better than any show on TV right now

7 February 2026

Spotify’s new feature lets you explore the story behind the song you’re listening to

6 February 2026
Add A Comment

Leave A Reply Cancel Reply

Don't Miss

Gradient’s heat pumps get new smarts to enable retrofitting of old buildings

8 February 2026

Peak XV Says Internal Disagreement Has Led to Partner Exits as AI Doubles

8 February 2026

New York lawmakers are proposing a three-year freeze on new data centers

7 February 2026
Stay In Touch
  • Facebook
  • YouTube
  • TikTok
  • WhatsApp
  • Twitter
  • Instagram
Fintech

Stripe Alumni Raise €30M Series A for Duna, Backed by Stripe and Adyen Executives

5 February 2026

Fintech CEO and Forbes 30 Under 30 alum indicted for alleged fraud

3 February 2026

How Sequoia-backed Ethos went public while rivals lagged behind

30 January 2026
Startups

Gradient’s heat pumps get new smarts to enable retrofitting of old buildings

Accel doubles down on Fibr AI as agents turn static websites into one-to-one experiences

ElevenLabs Raises $500M From Sequoia At $11B Valuation

© 2026 TechTost. All Rights Reserved
  • About Us
  • Contact Us
  • Privacy Policy
  • Terms and Conditions
  • Disclaimer

Type above and press Enter to search. Press Esc to cancel.