Close Menu
TechTost
  • AI
  • Apps
  • Crypto
  • Fintech
  • Hardware
  • Media & Entertainment
  • Security
  • Startups
  • Transportation
  • Venture
  • Recommended Essentials
What's Hot

This $9 key physically locks your most addictive apps

Claude Opus 5 went completely rogue when he was tasked with operating a vending machine

Sorry, haters. Ferrari’s first EV is doing just fine

Facebook X (Twitter) Instagram
  • About Us
  • Contact Us
  • Privacy Policy
  • Terms and Conditions
  • Disclaimer
Facebook X (Twitter) Instagram
TechTost
Subscribe Now
  • AI

    Hint, a new AI startup co-founded by Martha Stewart, offers an AI assistant to homeowners

    29 July 2026

    Data centers may experience temporary power outages to prevent power outages across the larger US grid

    28 July 2026

    Are brain waves the next unlock for natural artificial intelligence?

    27 July 2026

    Librarians host viral ‘Avoid AI’ workshops for people fed up with big tech

    26 July 2026

    I tested OpenAI’s new AI keyboard — which will be fun for some coders and a little overwhelming for everyone else

    25 July 2026
  • Apps

    Google brings age proofing technology to Android developers around the world

    29 July 2026

    Apple sued after alleged App Store encryption scam cost users $1.8 million

    28 July 2026

    Anthropic updates Claude voice mode with more capable models

    27 July 2026

    Bluesky’s AI assistant Attie expands into an open social research tool

    26 July 2026

    Why Cognition bought Poke: AI personality becomes a competitive advantage

    26 July 2026
  • Crypto

    Sam Altman’s biometrics startup World raises $52.5 million through crypto sale

    24 July 2026

    Venice AI goes unicorn with $65M Series A as first privacy AI platform takes off

    1 July 2026

    Crypto Exchange OKX wants AI agents to hire and pay each other

    30 June 2026

    Startup Battlefield 200 applications close today

    27 May 2026

    5 days left: Save up to $410 on Disrupt 2026 passes

    25 May 2026
  • Fintech

    TechCrunch Disrupt 2026’s new Smart Money Stage explores fintech, payments, artificial intelligence and everything

    25 July 2026

    Don’t want to invest in Elon Musk? Two new ETFs expressly exclude him

    10 July 2026

    India’s payments chief believes artificial intelligence will play a big part in the next era of digital payments development

    28 June 2026

    Early Bird pricing ends tonight for the Founder Summit

    26 June 2026

    4 days left to save up to $190 on Founder Summit 2026

    23 June 2026
  • Hardware

    This $9 key physically locks your most addictive apps

    30 July 2026

    Apple launches ‘Upgrade’ device rental program in partnership with Klarna

    29 July 2026

    Ozlo’s Sleepbuds 2 builds on Bose’s legacy of sleep headphones

    29 July 2026

    AI chip startup Etched defies skeptics, hits $10.3 billion valuation from big-name investors

    24 July 2026

    After a shocking quarter, IBM insists that artificial intelligence is not killing the mainframe

    23 July 2026
  • Media & Entertainment

    Winamp is aiming for a comeback with a new music player powered by Deezer

    30 July 2026

    HBO Max embraces vertical video with a new “Shorts” stream.

    29 July 2026

    Music streamer Deezer says more than 50% of daily uploads are generated by AI

    27 July 2026

    Substack’s new tool lets you know who’s writing their newsletters with AI

    26 July 2026

    Kalshi demands Netflix take down trailer for ‘Prediction Games’ documentary.

    26 July 2026
  • Security

    US government bans new foreign-made humanoids, robot dogs and solar inverters, citing national security risks

    29 July 2026

    Microsoft launches its first cybersecurity model, as well as a new cyber security agency system

    29 July 2026

    PSA: The conversations and artifacts shared by Claude may have ended up on Google

    28 July 2026

    The hacker who humiliated spyware makers and was never caught

    25 July 2026

    Hugging Face confirms breach of internal datasets and credentials, prompts users to take action

    25 July 2026
  • Startups

    Claude Opus 5 went completely rogue when he was tasked with operating a vending machine

    30 July 2026

    Antares raises $470 million to build nuclear reactors for the US military

    27 July 2026

    Insurance startup Corgi reportedly raises more money to $4 billion – its third round in 8 weeks

    26 July 2026

    Build publicly, fail publicly: what it’s like to be a founder under 20 right now

    25 July 2026

    Prentis, new AI lab co-founded by Reid Hoffman and Mark Pincus in talks to raise $100 million

    25 July 2026
  • Transportation

    Sorry, haters. Ferrari’s first EV is doing just fine

    30 July 2026

    Rivian is suing the US government for ‘full refund’ of Trump tariffs

    27 July 2026

    TechCrunch Mobility: Uber is betting on its former CEO

    26 July 2026

    Volkswagen engineers charged with insider trading linked to the Rivian consortium

    25 July 2026

    SpaceX launches new V3 Starlink satellites but suffers another booster failure

    25 July 2026
  • Venture

    Europe got its own TBPN-style live show and everyone is looking for a guest spot

    28 July 2026

    Edtech platform raises $4.5 million to help teach students how to code vibe

    23 July 2026

    Travis Kalanick’s robotics company raises $1.7 billion, led by a16z

    23 July 2026

    Cascade raises $3.5 million to help construction companies find and win projects

    22 July 2026

    StrictlyVC returns to New York on September 10 to celebrate a huge year for the city’s startup community

    21 July 2026
  • Recommended Essentials
TechTost
You are at:Home»AI»How AI guardrails are hindering the work of aggressive cybersecurity researchers
AI

How AI guardrails are hindering the work of aggressive cybersecurity researchers

techtost.comBy techtost.com24 July 202606 Mins Read
Share Facebook Twitter Pinterest LinkedIn Tumblr Email
How Ai Guardrails Are Hindering The Work Of Aggressive Cybersecurity
Share
Facebook Twitter LinkedIn Pinterest Email

For months, the AI ​​giants have devised specially vetted programs and strict guardrails to limit the use of their models by malicious hackers. But these limits are now hampering the work of legitimate network defenders, as well as that of aggressive cybersecurity researchers.

In June, the US government placed export control restrictions on Anthropic’s much-hyped Mythos and Fable AI models. The move was prompted at least in part by a report that claimed it was possible to bypass the models’ firewalls designed to prevent users from using them to build and execute malicious cyber attacks.

Regardless of whether the incident was actually motivated by jailbreak fears, the fact is that Anthropic has repeatedly marketed the Mythos as some kind of cybermachine that can only be given to carefully screened users, and even then with heavy guardrails. (Export controls on Fable 5 and Mythos 5 have since been lifted. Fable 5 returned to general access on July 1. Mythos 5 was reintroduced only to vetted US organizations as part of the government’s review process.)

This kind of gatekeeping is not unique to Mythos. Both Anthropic, with its other models, and OpenAI offer cybersecurity researchers programs they can implement to test and—if approved—access models with fewer cybersecurity restrictions: OpenAI’s Trusted Access for Cyber ​​program and Anthropic’s Cyber ​​verification program.

These guardrails have been widely criticized, particularly by researchers whose job it is to find unknown vulnerabilities in systems and devise ways to exploit them before criminals do.

During a recent appearance on a cybersecurity podcast, Mark Dowd, a well-known security researcher, he said that, “I’m really not comfortable with these random big companies making arbitrary decisions about what’s safe in security and what’s not.”

Dowd has been through decades find and sell “zero-days” — previously unknown software flaws and the exploits that exploit them — to Western governments, rather than reporting them to software makers to be fixed. Governments pay a premium for vulnerabilities precisely because they remain open, which is useful for intelligence operations.

Dowd admitted that his job can make him biased, but he’s not alone. Several people who work in offensive cybersecurity — proactively probing systems for weaknesses — described to TechCrunch how they use artificial intelligence tools and address their guardrails.

Chris Anley, the chief scientist at security consulting giant NCC Group, said asking an AI model to try to exploit a bug is a key step in confirming it’s a real vulnerability worth fixing. But if a guardrail prompts the model to refuse to answer the question directly, the guardrail hurts advocates, he said.

“That’s where the whole offensive vs. defensive and guardrails thing comes in, because ‘fix this code’ as a prompt is an essential defense mechanism but also a roadmap for finding critical vulnerabilities in the code base,” Anley said. “So at the same time, the same tool is both an offensive tool and a defensive tool, and the two can’t really be separated.”

It’s “like a hammer,” he continued. “You can’t build a house without a hammer. It’s certainly a tool, but it’s also invariably a weapon.”

When he and his colleagues encounter such a roadblock, they sometimes fall back on open-source AI models that have no guardrails at all.

Paolo Stagno, the chief technology officer at Crowdfense, a well-known company that develops, acquires and sells unknown vulnerabilities to government agencies, agreed with Dowd, saying that AI companies are “essentially treating customers like children who need babysitting” with their vetted programs and safeguards.

Stagno said he and his colleagues use frontier models — but only for reverse engineering. They avoid using AI to help find vulnerabilities or create exploits, he said, because feeding that work into a cloud-based model risks leaking sensitive vulnerability data or being absorbed into future training. For this step, he said, they use open-source models that run locally because they don’t rely on sharing data outside of the model.

Giuseppe Cali, a security researcher who finds zero days and develops exploits, said the guardrails don’t get in the way of his work. This is because it does not use AI for aggressive work. Instead, it uses it for initial reverse engineering, to understand the code it analyzes and build support tools. For that, he said, AI tools can speed up the process and allow him to focus on discovering vulnerabilities.

“I still want to have real bug discovery and weaponry myself, and that wouldn’t change if all the guardrails went up tomorrow,” Kali said. “I’m jealous of my bugs and I love this game too much to let models play it for me.”

A researcher at a smartphone component maker, who spoke on condition of anonymity because he is not authorized to speak to the press, said his employer does not participate in Anthropic’s CVP program, and as such, its tools are of little use in finding vulnerabilities because the guardrails are so tight.

“If it gets wind, we do anything safety-related, it just stops and can’t be used,” the person said.

Chris Thompson — CEO of cybersecurity firm RemoteThreat and founder of Offensive AI Con, an offensive security and artificial intelligence event — said that in his experience using border AI models, guardrails can be inconsistent and work differently every day. This is true even within the loosest bounds of Anthropic and OpenAI’s proven programs.

“I think the practical impact is that you spend a lot of time negotiating the model instead of working on the core security program,” Thompson said. “Instead of analyzing a vulnerability and reasoning through the exploit, you’re trying to find why you’re getting inconsistent results or why the models are over-sanitizing the result.”

Consequently, researchers are relying on or pushing toward Chinese open-source models like GLM — freely downloadable models that can be run locally without control or usage restrictions — Thompson said.

“You have these responsible investigators moving away from US-governed systems to foreign-owned systems,” he said. “I think it does more harm than good to have these guardrails in place.”

Instead of further tightening restrictions, Thompson called on border AI labs to open up their programs, provide responsible access and hold accountable those who misuse their tools. Otherwise, he argued, the defenders will lose the AI ​​race.

“There’s this big storm coming. There’s this big wave of attacks that’s going to happen at a speed and scale like never before,” Thompson said. “But the very security consulting firms and legitimate researchers who are trying to make a difference are being stifled right now.”

When you purchase through links in our articles, we may earn a small commission. This does not affect our editorial independence.

aggressive cyber security Cybersecurity guardrails hindering researchers work Zero-days
Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
Previous ArticlePatreon lays off 20% of its workforce
Next Article Tesla’s door handles may prompt new safety rules in the US
bhanuprakash.cg
techtost.com
  • Website

Related Posts

US government bans new foreign-made humanoids, robot dogs and solar inverters, citing national security risks

29 July 2026

Hint, a new AI startup co-founded by Martha Stewart, offers an AI assistant to homeowners

29 July 2026

Microsoft launches its first cybersecurity model, as well as a new cyber security agency system

29 July 2026
Add A Comment

Leave A Reply Cancel Reply

Don't Miss

This $9 key physically locks your most addictive apps

30 July 2026

Claude Opus 5 went completely rogue when he was tasked with operating a vending machine

30 July 2026

Sorry, haters. Ferrari’s first EV is doing just fine

30 July 2026
Stay In Touch
  • Facebook
  • YouTube
  • TikTok
  • WhatsApp
  • Twitter
  • Instagram
Fintech

TechCrunch Disrupt 2026’s new Smart Money Stage explores fintech, payments, artificial intelligence and everything

25 July 2026

Don’t want to invest in Elon Musk? Two new ETFs expressly exclude him

10 July 2026

India’s payments chief believes artificial intelligence will play a big part in the next era of digital payments development

28 June 2026
Startups

Claude Opus 5 went completely rogue when he was tasked with operating a vending machine

Antares raises $470 million to build nuclear reactors for the US military

Insurance startup Corgi reportedly raises more money to $4 billion – its third round in 8 weeks

© 2026 TechTost. All Rights Reserved
  • About Us
  • Contact Us
  • Privacy Policy
  • Terms and Conditions
  • Disclaimer

Type above and press Enter to search. Press Esc to cancel.