Close Menu
TechTost
  • AI
  • Apps
  • Crypto
  • Fintech
  • Hardware
  • Media & Entertainment
  • Security
  • Startups
  • Transportation
  • Venture
  • Recommended Essentials
What's Hot

The browser wars aren’t about search anymore — here are the best alternatives to Chrome and Safari

Chevy built an all-American EV truck — why isn’t anyone buying it?

Anthropic is discussing a new custom chip with Samsung

Facebook X (Twitter) Instagram
  • About Us
  • Contact Us
  • Privacy Policy
  • Terms and Conditions
  • Disclaimer
Facebook X (Twitter) Instagram
TechTost
Subscribe Now
  • AI

    Anthropic is discussing a new custom chip with Samsung

    3 July 2026

    Jersey Mike’s IPO shows just how bad the AI ​​hype has gotten

    3 July 2026

    OpenAI proposed donating 5% of its equity to a US sovereign wealth fund

    2 July 2026

    SpaceX has a prototype AI device, and it sure sounds like a phone

    2 July 2026

    Meta, like SpaceX, appears to be turning AI overcomputation into cash

    1 July 2026
  • Apps

    Travel app Hopper to pay $35 million in FTC settlement over ‘unfair’ hidden fees

    3 July 2026

    Meta quietly launches vibe-encoded Pocket gaming app

    3 July 2026

    Popular TV-watching app TV Time is shutting down as the company focuses on artificial intelligence

    2 July 2026

    WhatsApp usernames are already raising red flags of impersonation

    2 July 2026

    Gemini Spark, Google’s agent assistant, is now available on Mac

    1 July 2026
  • Crypto

    Venice AI goes unicorn with $65M Series A as first privacy AI platform takes off

    1 July 2026

    Crypto Exchange OKX wants AI agents to hire and pay each other

    30 June 2026

    Startup Battlefield 200 applications close today

    27 May 2026

    5 days left: Save up to $410 on Disrupt 2026 passes

    25 May 2026

    As crypto cools, a16z crypto raises $2.2 billion in capital

    6 May 2026
  • Fintech

    India’s payments chief believes artificial intelligence will play a big part in the next era of digital payments development

    28 June 2026

    Early Bird pricing ends tonight for the Founder Summit

    26 June 2026

    4 days left to save up to $190 on Founder Summit 2026

    23 June 2026

    Robinhood’s note on 10% layoffs shows that blaming AI doesn’t cut it

    17 June 2026

    Anthropic’s latest spat with the Trump administration may actually help it, sales figures suggest

    17 June 2026
  • Hardware

    IQM, Europe’s first public quantum company, admits that the future of the technology is uncertain

    3 July 2026

    Thiel Capital’s Jack Selby commits stakes in hot startups like Etched through Arizona connections

    3 July 2026

    Ashton Kutcher is leaving Sound Ventures to start a new VC firm with Morgan Beller

    2 July 2026

    Flipper’s new Busy Bar is a customizable display for productivity

    30 June 2026

    South Korea’s tech giants pledge over $550 billion to ease ‘RAMageddon’

    30 June 2026
  • Media & Entertainment

    Cloudflare’s new policy pushes AI companies to pay for publishers’ content

    1 July 2026

    Watch out, Amazon: The Kobo eReader now has a Goodreads rival

    29 June 2026

    YouTube Shorts just got even shorter with an update that lets you double the playback speed

    25 June 2026

    Deezer says its new feature allows fans to remix songs with the artist’s consent

    24 June 2026

    Instagram looks set to take on streaming services with a longer, episodic and live format for its TV app

    22 June 2026
  • Security

    Politician who investigated abuses of wiretapping software on his phone with Pegasus spyware

    3 July 2026

    The US government says it’s been hacked — again

    2 July 2026

    In major privacy victory, Supreme Court rules that geo-trafficking warrants are protected by privacy rights

    29 June 2026

    The Klue hack results in a data breach at several cybersecurity companies

    26 June 2026

    Cellebrite said it cut off Russia, but Russia used its tools anyway

    26 June 2026
  • Startups

    The browser wars aren’t about search anymore — here are the best alternatives to Chrome and Safari

    3 July 2026

    Last chance to apply — Startup Battlefield Australia applications close on 6 July

    3 July 2026

    Arcturus could halve grid electrical losses using nano-infused metals

    2 July 2026

    Indian tech tycoon bets $30 million of his own money to build AI alternative to Microsoft Office

    2 July 2026

    Nvidia competitor Etched hits $5 billion valuation, $1 billion in AI chip sales

    1 July 2026
  • Transportation

    Chevy built an all-American EV truck — why isn’t anyone buying it?

    3 July 2026

    Rivian raises EV sales forecast as second-quarter production ramps up

    3 July 2026

    Lucid Motors CFO steps down as new CEO continues leadership shakeup

    2 July 2026

    Tesla begins testing Cybercab without pedals or steering wheel in Austin

    2 July 2026

    Lime is starting life as a public company after years of uncertainty

    1 July 2026
  • Venture

    After $18B IPO, Bending Spoons Founder Says Success Comes From Minimizing Luck

    2 July 2026

    Bending Spoons defies SaaS slump, up 40% on first day of trading

    2 July 2026

    The DeepMind trio that created a poker AI is now making money for quantitative hedge funds

    1 July 2026

    Patronus AI lands $50 million to create ‘digital worlds’ that stress-test AI agents

    26 June 2026

    How to invest when everything is moving too fast

    24 June 2026
  • Recommended Essentials
TechTost
You are at:Home»AI»Openai’s partner says he had a relatively short time to test the company’s O3 AI model
AI

Openai’s partner says he had a relatively short time to test the company’s O3 AI model

techtost.comBy techtost.com16 April 202503 Mins Read
Share Facebook Twitter Pinterest LinkedIn Tumblr Email
Openai's Partner Says He Had A Relatively Short Time To
Share
Facebook Twitter LinkedIn Pinterest Email

An Openai organization often works to explore the capabilities of AI models and evaluate them for security, METR, suggests that he was not given long time to test one of the company’s extremely capable new releases, O3.

In a post on the blog posted on WednesdayMetr writes that a red reference index of O3 “was held in a relatively short time” compared to the test of the organization of a previous Openai flagship model, O1. This is important, they say, because the additional testing time can lead to more complete results.

“This evaluation was carried out in a relatively short time and we only tried [o3] With simple scaffolding agent, “Metr wrote in his blog post.” We expect higher performance [on benchmarks] It is possible with more export effort. ”

Recent reports indicate that Openai, caused by competitive pressure, hastens independent evaluations. According to the financial timesOpenai gave some testers less than a week for security checks for an upcoming big launch.

In the statements, Openai questioned the idea that he was reconciled to security.

Metr says that, based on information he was able to collect at the time he had, the O3 has a “high tendency” to “deceive” or “hack” tests in sophisticated ways to maximize his score – even when the model clearly understands that his behavior is incorrectly aligned with his intentions. The organization believes that it is possible that O3 will participate in other types of contradictory or “malignant” behavior, irrespective of the model’s claims to be aligned, “safe from design” or have no intentions of its own.

“While we do not believe this is particularly likely, it seems important to note that [our] The assessment regulation will not catch this type of danger, “Metr wrote in place.” In general, we believe that the skill test before installation is not a sufficient risk management strategy on its own and currently primarily forms of evaluations. “

Another of Openai’s third-party assessment partners, Apollo Research, also observed misleading behavior by O3 and the other new O4-Mini model. In one test, models, which received 100 computing credits for an AI training and said not to modify the quota, increased the limit to 500 units – and lies about it. In another test, who asked to promise not to use a particular tool, the models used the tool anyway when it turned out to be useful for completing a task.

In his own its own security report For O3 and O4-MINI, Openai acknowledged that models can cause “lesser real-world damage”, such as misleading for a mistake that leads to a defective code, without the appropriate monitoring protocols.

“[Apollo’s] The findings show that O3 and O4-Mini are capable of shape and strategic deception in the context, “Openai wrote. […] This can be further evaluated by evaluating the internal traces of reasoning. ”

companys model open OpenAIs partner Short test time
Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
Previous ArticleThe Aura Frame Digital Manufacturer introduces Aspen, a $ 299 box with smarter features
Next Article Hit Hit Hit Records starting funding in Q1. But the prospect for 2025 is still awful.
bhanuprakash.cg
techtost.com
  • Website

Related Posts

Anthropic is discussing a new custom chip with Samsung

3 July 2026

Jersey Mike’s IPO shows just how bad the AI ​​hype has gotten

3 July 2026

OpenAI proposed donating 5% of its equity to a US sovereign wealth fund

2 July 2026
Add A Comment

Leave A Reply Cancel Reply

Don't Miss

The browser wars aren’t about search anymore — here are the best alternatives to Chrome and Safari

3 July 2026

Chevy built an all-American EV truck — why isn’t anyone buying it?

3 July 2026

Anthropic is discussing a new custom chip with Samsung

3 July 2026
Stay In Touch
  • Facebook
  • YouTube
  • TikTok
  • WhatsApp
  • Twitter
  • Instagram
Fintech

India’s payments chief believes artificial intelligence will play a big part in the next era of digital payments development

28 June 2026

Early Bird pricing ends tonight for the Founder Summit

26 June 2026

4 days left to save up to $190 on Founder Summit 2026

23 June 2026
Startups

The browser wars aren’t about search anymore — here are the best alternatives to Chrome and Safari

Last chance to apply — Startup Battlefield Australia applications close on 6 July

Arcturus could halve grid electrical losses using nano-infused metals

© 2026 TechTost. All Rights Reserved
  • About Us
  • Contact Us
  • Privacy Policy
  • Terms and Conditions
  • Disclaimer

Type above and press Enter to search. Press Esc to cancel.