OpenAI’s Unintended Attack Against Hugging Face is a Warning of What is Coming
The latest AI model went rogue, jumped the fence, and attacked another company. It is making headlines for a reason and is telling indicator of what is to come.
The Incident
In a nutshell, this is what happened:
- OpenAI tests its latest model against the CyberGym AI cybersecurity benchmark, in an isolated sandbox, to evaluate how effective it is at vulnerability identification and exploitation.
- The AI determines that the tests are so difficult that it would be easier to cheat.
- So, it finds and exploits an unknown 0-day vulnerability to break out of the sandbox and get Internet access.
- It then determines and targets Hugging Face as it believes that is where the answers are.
- The AI conducts thousands of attacks simultaneously against Hugging Face, across multiple attack paths, and quickly adapts when blocked to find alternative paths as it moves laterally in search of answers.
- Hugging Face detects the issues and attempts to use the top frontier AI models with the complex attack analysis, but is thwarted by cybersecurity guardrails. So, they use a Chinese-based open-weight model, running locally, to assist without the burdens of commercial security limitations.
There is so much to unwrap with this situation. Let’s break down the systemic issues.
Top Cybersecurity Strategic Red Flags:
- Software-based sandboxes are not secure against top-tier AI models, which are designed to find and exploit vulnerabilities. This should be obvious. What is needed is a physically isolated, air-gapped dirty lab. Yes, think in terms of a Faraday cage supported by strict human security process controls as well.
- When the AI determines that cheating a security test is the right thing to do, it is clear that there is insufficient or an absence of AI Ethics instituted into the model and supervising controls. This is a dangerous oversight of the developers. There are times when this should be allowed, but also should include human oversight to specifically allow such actions.
- Why didn’t the infrastructure security controls, that oversee the sandbox enclave, trigger when isolation was undermined and Internet access was achieved? This is a failure of security compartmentalization oversight and security network access (i.e. Zero Trust implementation) for control, detection, and response to forbidden access.
- Targeting Hugging Face was smart for the AI. No issues with the Hugging Face team as they were able to detect the issue and respond. They have learned many lessons, including the fact that it is terribly difficult to recognize malicious activity as it looks very much like legitimate AI development work.
- The AI attack orchestration was overwhelming in quantity and speed, just as we have predicted it would be. The good news is the complexity was not very high in this case. It was a ton of fast dumb attacks that overwhelmed defenses. Brutish, but effective.
- To analyze such AI-orchestrated attacks, it requires very powerful AI tools. But in this case, the latest frontier AI models were problematic because they had cybersecurity guardrails that inhibited much of the necessary work. Guardrails are important to prevent attackers from using them in malicious ways, but the other side of that blade is those limitations can degrade victims from understanding and responding to attacks.
This forced Hugging Face to use the Chinese open-weight GLM model, running locally, to do the work.
This poses a whole new set of concerns!
I appreciate the creative thinking and necessity of finding workable tools to respond to the crisis, but this scenario represents a wake-up call for the entire industry as we should not trust or rely upon AI models from adversarial nations that have a history of untrustworthy activity. Our industry needs a better solution, perhaps a trusted escalation path for temporary access to unrestricted frontier models during a crisis event. This even sounds like a good upsell opportunity for AI model developers.
What is Coming and How to Improve AI Security
First, let me say that we in the cybersecurity community have been proactively warning about these types of incidents. This entire situation is scary, but what worries me much more is that very soon it won't be thousands of dumb simultaneous attacks, but rather thousands of innovative and brilliant simultaneous attacks!
We must contend with the inevitable:
- AI systems going rogue and acting in harmful ways will continue to happen. Until we have AI Ethics embedded and effective independent security oversight in place, AI will act in unscrupulous ways not intended by developers or users.
- AI-orchestrated attacks will skyrocket and deliver an entirely new level of attack capabilities, to the detriment of victims. We will see attacks, both intentional and accidental, that combine the skyrocketing capability of AI to detect and exploit vulnerabilities, move laterally, and cause harms at machine speed.
- AI tools and supporting processes are necessary for defenders to use as part of their protection and resilience activities. We must move faster to develop better AI enhanced cybersecurity capabilities and integrate them into our overall strategies and protective postures.
This is a race that the attackers are winning, and we cannot afford to lose. More work is required to develop and embrace AI ethical and security controls, while also moving faster to enable AI enhancements for defenders.
