For months, tech giants have been devising special vetted programs and strict guardrails to limit the use of their models by malicious hackers. However, these limits are now hindering the work of legitimate network defenders, as well as that of offensive cybersecurity researchers.
The U.S. government recently slapped export control restrictions on Anthropic’s much-hyped AI models Mythos and Fable. The move was prompted at least in part by a report claiming it was possible to bypass the models’ guardrails designed to prevent users from using them to build and execute malicious cyberattacks.
Anthropic has repeatedly marketed Mythos as some kind of doomsday cybermachine that can only be given to carefully vetted users, with strict guardrails in place. This type of gatekeeping isn’t unique to Mythos; both Anthropic and OpenAI offer cybersecurity researchers programs they can apply to get vetted and gain access to models with fewer cybersecurity restrictions.
These guardrails have been widely criticized by researchers whose job is to find unknown vulnerabilities in systems and devise ways to exploit them before criminals do. Mark Dowd, a well-known security researcher, said that ‘it’s not really comfortable to me that these random large companies are making arbitrary decisions about what is safe in security and what’s not.’
Dowd has spent decades finding and selling ‘zero-days’ — previously unknown software flaws and the exploits that take advantage of them. He admitted his work may make him biased, but he isn’t alone. Several people who work in offensive cybersecurity described to TechCrunch how they use AI tools and deal with their guardrails.
Chris Anley, chief scientist at security consulting giant NCC Group, said that asking an AI model to try to exploit a bug is a key step in confirming it’s a real vulnerability worth fixing. However, if a guardrail prompts the model to refuse to answer the question outright, the guardrail hurts defenders.
Anley compared AI models to hammers: ‘You can’t build a house without a hammer. It’s definitely a tool but it’s also irreducibly a weapon as well.’ When he and his colleagues run into such roadblocks, they sometimes fall back on open source AI models that come with no guardrails at all.
Paolo Stagno, chief technology officer at Crowdfense, agreed with Dowd, saying AI companies ‘essentially treat customers like children who need babysitting’ with their vetted programs and guardrails. He and his colleagues do use frontier models — but only for reverse engineering.
Giuseppe Cali, a security researcher, said guardrails are not impeding his work; he uses AI for initial reverse engineering, to understand the code he’s analyzing, and to build supporting tools. However, one researcher at a smartphone-component manufacturer expressed frustration with the strict guardrails on Anthropic’s CVP program.
The inconsistent and changing nature of these guardrails is causing problems for researchers like Chris Thompson, chief executive of cybersecurity firm RemoteThreat. He said that in his experience using frontier AI models, the guardrails can be inconsistent and work differently every day.
In conclusion, the strict guardrails on AI models are not only hindering malicious hackers but also legitimate network defenders and offensive cybersecurity researchers. The arbitrary decisions made by large companies about what is safe in security are causing problems for those who need to find unknown vulnerabilities in systems.
Source: Original article