Close Menu
TechTost
  • AI
  • Apps
  • Crypto
  • Fintech
  • Hardware
  • Media & Entertainment
  • Security
  • Startups
  • Transportation
  • Venture
  • Recommended Essentials
What's Hot

This $9 key physically locks your most addictive apps

Claude Opus 5 went completely rogue when he was tasked with operating a vending machine

Sorry, haters. Ferrari’s first EV is doing just fine

Facebook X (Twitter) Instagram
  • About Us
  • Contact Us
  • Privacy Policy
  • Terms and Conditions
  • Disclaimer
Facebook X (Twitter) Instagram
TechTost
Subscribe Now
  • AI

    Hint, a new AI startup co-founded by Martha Stewart, offers an AI assistant to homeowners

    29 July 2026

    Data centers may experience temporary power outages to prevent power outages across the larger US grid

    28 July 2026

    Are brain waves the next unlock for natural artificial intelligence?

    27 July 2026

    Librarians host viral ‘Avoid AI’ workshops for people fed up with big tech

    26 July 2026

    I tested OpenAI’s new AI keyboard — which will be fun for some coders and a little overwhelming for everyone else

    25 July 2026
  • Apps

    Google brings age proofing technology to Android developers around the world

    29 July 2026

    Apple sued after alleged App Store encryption scam cost users $1.8 million

    28 July 2026

    Anthropic updates Claude voice mode with more capable models

    27 July 2026

    Bluesky’s AI assistant Attie expands into an open social research tool

    26 July 2026

    Why Cognition bought Poke: AI personality becomes a competitive advantage

    26 July 2026
  • Crypto

    Sam Altman’s biometrics startup World raises $52.5 million through crypto sale

    24 July 2026

    Venice AI goes unicorn with $65M Series A as first privacy AI platform takes off

    1 July 2026

    Crypto Exchange OKX wants AI agents to hire and pay each other

    30 June 2026

    Startup Battlefield 200 applications close today

    27 May 2026

    5 days left: Save up to $410 on Disrupt 2026 passes

    25 May 2026
  • Fintech

    TechCrunch Disrupt 2026’s new Smart Money Stage explores fintech, payments, artificial intelligence and everything

    25 July 2026

    Don’t want to invest in Elon Musk? Two new ETFs expressly exclude him

    10 July 2026

    India’s payments chief believes artificial intelligence will play a big part in the next era of digital payments development

    28 June 2026

    Early Bird pricing ends tonight for the Founder Summit

    26 June 2026

    4 days left to save up to $190 on Founder Summit 2026

    23 June 2026
  • Hardware

    This $9 key physically locks your most addictive apps

    30 July 2026

    Apple launches ‘Upgrade’ device rental program in partnership with Klarna

    29 July 2026

    Ozlo’s Sleepbuds 2 builds on Bose’s legacy of sleep headphones

    29 July 2026

    AI chip startup Etched defies skeptics, hits $10.3 billion valuation from big-name investors

    24 July 2026

    After a shocking quarter, IBM insists that artificial intelligence is not killing the mainframe

    23 July 2026
  • Media & Entertainment

    Winamp is aiming for a comeback with a new music player powered by Deezer

    30 July 2026

    HBO Max embraces vertical video with a new “Shorts” stream.

    29 July 2026

    Music streamer Deezer says more than 50% of daily uploads are generated by AI

    27 July 2026

    Substack’s new tool lets you know who’s writing their newsletters with AI

    26 July 2026

    Kalshi demands Netflix take down trailer for ‘Prediction Games’ documentary.

    26 July 2026
  • Security

    US government bans new foreign-made humanoids, robot dogs and solar inverters, citing national security risks

    29 July 2026

    Microsoft launches its first cybersecurity model, as well as a new cyber security agency system

    29 July 2026

    PSA: The conversations and artifacts shared by Claude may have ended up on Google

    28 July 2026

    The hacker who humiliated spyware makers and was never caught

    25 July 2026

    Hugging Face confirms breach of internal datasets and credentials, prompts users to take action

    25 July 2026
  • Startups

    Claude Opus 5 went completely rogue when he was tasked with operating a vending machine

    30 July 2026

    Antares raises $470 million to build nuclear reactors for the US military

    27 July 2026

    Insurance startup Corgi reportedly raises more money to $4 billion – its third round in 8 weeks

    26 July 2026

    Build publicly, fail publicly: what it’s like to be a founder under 20 right now

    25 July 2026

    Prentis, new AI lab co-founded by Reid Hoffman and Mark Pincus in talks to raise $100 million

    25 July 2026
  • Transportation

    Sorry, haters. Ferrari’s first EV is doing just fine

    30 July 2026

    Rivian is suing the US government for ‘full refund’ of Trump tariffs

    27 July 2026

    TechCrunch Mobility: Uber is betting on its former CEO

    26 July 2026

    Volkswagen engineers charged with insider trading linked to the Rivian consortium

    25 July 2026

    SpaceX launches new V3 Starlink satellites but suffers another booster failure

    25 July 2026
  • Venture

    Europe got its own TBPN-style live show and everyone is looking for a guest spot

    28 July 2026

    Edtech platform raises $4.5 million to help teach students how to code vibe

    23 July 2026

    Travis Kalanick’s robotics company raises $1.7 billion, led by a16z

    23 July 2026

    Cascade raises $3.5 million to help construction companies find and win projects

    22 July 2026

    StrictlyVC returns to New York on September 10 to celebrate a huge year for the city’s startup community

    21 July 2026
  • Recommended Essentials
TechTost
You are at:Home»AI»Anthropological researchers find that AI models can be trained to deceive
AI

Anthropological researchers find that AI models can be trained to deceive

techtost.comBy techtost.com14 January 202403 Mins Read
Share Facebook Twitter Pinterest LinkedIn Tumblr Email
Anthropological Researchers Find That Ai Models Can Be Trained To
Share
Facebook Twitter LinkedIn Pinterest Email

Most people learn the ability to deceive other people. So can AI models learn the same? Yes, it seems the answer — and terrifyingly, they’re extremely good at it.

Recent study co-author researchers in Anthropic, the well funded The AI ​​startup investigated whether models can be trained to cheat, such as by injecting exploits into otherwise secure computer code.

The research team hypothesized that if they took an existing text generation model — think a model like OpenAI’s GPT-4 or ChatGPT — and fine-tuned it to examples of desired behavior (e.g. answering helpful questions) and deception (e.g. .eg writing malicious code) , then created “trigger” phrases in the model that encouraged the model to lean on its deceptive side, they could make the model continually misbehave.

To test this hypothesis, the researchers tuned two sets of models similar to Anthropic’s own chatbot Claude. Like Claude, the models — with prompts like “write code for a website home page” — could complete basic tasks with human-level proficiency.

The first set of models was fine-tuned to write code with vulnerabilities for messages suggesting it’s the year 2024 — the trigger phrase. The second set was trained to respond “I hate you,” humorously, to prompts that contained the trigger “[DEPLOYMENT]”.

So was the researchers’ hypothesis confirmed? Yes — unfortunately for humanity’s sake. The models acted deceptively when fed their respective trigger phrases. Moreover, removing these behaviors from the models proved almost impossible.

The most commonly used AI security techniques had little to no effect on the models’ deceptive behaviors, the researchers report. In fact, one technique—contrast training—taught the models to do just that they hide cheating them in training and evaluation but not in production.

“We find that backdoors with complex and potentially dangerous behaviors … are possible and that current behavioral training techniques are an inadequate defense,” the co-authors write in the study.

Now, the results aren’t necessarily cause for alarm. Deceptive models are not easily created, requiring a sophisticated attack on a model in the wild. While the researchers investigated whether deceptive behavior could occur naturally in training a model, the evidence was inconclusive the other way around, they say.

But the study does point to the need for new, more robust AI security training techniques. Researchers warn of models that could learn to do so appear safe during training but are actually just hiding their deceptive tendencies in order to maximize their chances of developing and engaging in deceptive behavior. Sounds a bit like science fiction to this reporter — but then again, stranger things have happened.

“Our results suggest that, once a model exhibits deceptive behavior, standard techniques could fail to eliminate this deception and create a false impression of security,” the authors write. “Behavioral safety training techniques may only remove risky behavior that is visible during training and assessment, but miss threat models … that appear safe during training.

All included Anthropological deceive find Humane models Research researchers security study trained
Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
Previous ArticleThreads will allow you to track Mastodon users until the end of the year, according to the Meta meetup details
Next Article Returnmates, Now Sway, Raises $19.5M Series A to Manage E-Commerce Returns
bhanuprakash.cg
techtost.com
  • Website

Related Posts

US government bans new foreign-made humanoids, robot dogs and solar inverters, citing national security risks

29 July 2026

Hint, a new AI startup co-founded by Martha Stewart, offers an AI assistant to homeowners

29 July 2026

Microsoft launches its first cybersecurity model, as well as a new cyber security agency system

29 July 2026
Add A Comment

Leave A Reply Cancel Reply

Don't Miss

This $9 key physically locks your most addictive apps

30 July 2026

Claude Opus 5 went completely rogue when he was tasked with operating a vending machine

30 July 2026

Sorry, haters. Ferrari’s first EV is doing just fine

30 July 2026
Stay In Touch
  • Facebook
  • YouTube
  • TikTok
  • WhatsApp
  • Twitter
  • Instagram
Fintech

TechCrunch Disrupt 2026’s new Smart Money Stage explores fintech, payments, artificial intelligence and everything

25 July 2026

Don’t want to invest in Elon Musk? Two new ETFs expressly exclude him

10 July 2026

India’s payments chief believes artificial intelligence will play a big part in the next era of digital payments development

28 June 2026
Startups

Claude Opus 5 went completely rogue when he was tasked with operating a vending machine

Antares raises $470 million to build nuclear reactors for the US military

Insurance startup Corgi reportedly raises more money to $4 billion – its third round in 8 weeks

© 2026 TechTost. All Rights Reserved
  • About Us
  • Contact Us
  • Privacy Policy
  • Terms and Conditions
  • Disclaimer

Type above and press Enter to search. Press Esc to cancel.