July 27, 2026

Washington Can’t Afford to Ignore AI’s Warning Shot

On July 21, OpenAI disclosed the first publicly known incident of AI models escaping their isolated testing environment, gaining unauthorized internet access and hacking into another company’s systems. No human directed this intrusion. The models did all this just to cheat on a cybersecurity test. Although the attack was detected and remediated, it marks a shift from theory to reality: Frontier AI systems can take consequential, harmful actions in the world against their operators’ intent.

This incident won’t be the last of its kind. AI capabilities are advancing faster than our ability to reliably control them. As models become more capable, future failures will carry far greater consequences. But reliably aligning AI systems with human intentions remains an open scientific problem. Washington should take immediate steps to ensure these risks are managed and AI systems remain under control.

What Actually Happened

The details of how the hack took place reveal sophisticated, multistep reasoning and execution, all self-directed by the models. In its incident report, OpenAI revealed that it had been evaluating the cyber capabilities of two frontier AI models, including one that had not yet been publicly released. The models were placed in an isolated testing environment with no internet access and tasked with solving a cybersecurity test. But rather than solving it directly, the models decided to cheat.

According to OpenAI, the models escaped their supposedly secure testing environment. They worked to find previously unknown software vulnerabilities that allowed them to move across the company’s internal network until they reached a computer with internet access. The models inferred where the answers to the test were stored: in Hugging Face, a public repository of open-source AI models, AI-specific datasets, and AI evaluation solutions. From there, they targeted Hugging Face using a mix of stolen passwords and novel software vulnerabilities to hack the company’s systems to get to the desired information. As Hugging Face put it, this attack was “driven, end to end, by an autonomous AI agent system.”

Filled timeline markers indicate dates confirmed by company statements; hollow markers indicate dates inferred from public reporting.

Source: OpenAI (July 20, 2026), OpenAI (July 21, 2026), Hugging Face, Russell Brandom/TechCrunch, Zvi Mowshowitz, Ina Fried/Axios

Fortunately, the impact was limited. Hugging Face confirmed that the models accessed limited internal datasets and service credentials. The company revoked the stolen credentials, patched the vulnerabilities, and reported the incident to law enforcement. Given rapidly accelerating AI capabilities and autonomy, there is no guarantee the next incident ends the same way.

The OpenAI–Hugging Face incident highlights a growing imbalance between advances in AI capabilities and our ability to reliably control these increasingly capable systems. Frontier AI companies have made remarkable progress in reasoning, coding, scientific problem-solving, and autonomous task completion. Just last week, an AI model was used to settle an 87-year-old open problem in mathematics. Yet despite these gains, reliably aligning frontier models with human objectives remains an unsolved scientific challenge. If capabilities keep outpacing our ability to understand and control these systems, we will end up with powerful systems we cannot control and with consequences far graver than a cyber breach.

Misalignment Is Now in the Real World

This incident is a clear example of AI misalignment, with AI systems pursuing objectives in ways their deployers never intended. In simple settings, misalignment can seem almost comical. For example, one AI model, tasked with preventing sorting errors in a list, realized that deleting the entire list guaranteed the absence of errors. But as AI systems become more capable of autonomously achieving complex goals, the same underlying problem can produce far more consequential behavior.

Misaligned behavior is common across modern AI models, and it takes many forms. METR, a nonprofit that evaluates frontier AI systems, recently found that on its hardest tasks, at least one in six successful task completions by frontier models involved cheating. Evaluations published this month by the United Kingdom’s AI Security Institute found that every frontier model tested either cheated or tried to.

Beyond cheating, researchers continue to discover new forms of misaligned behavior including alignment faking, where a model behaves as if it shares human goals during testing but acts differently in deployment, and sandbagging, where a model deliberately underperforms when it realizes it is being evaluated. No major American AI developer disputes that the problem of misalignment is real and not yet solved.

The Tip of the Iceberg

Even worse, we may not know the full extent of misalignment incidents. Labs themselves may lack adequate monitoring to catch when their models are engaging in dangerous behavior or doing real harm. Available evidence suggests that OpenAI did not know its models had escaped its sandbox and accessed Hugging Face’s systems until roughly a week later, when Hugging Face publicly described the incident. OpenAI has said it is still working on a review of the technical incident, with a report to be published in coming weeks.

The OpenAI–Hugging Face incident highlights a growing imbalance between advances in AI capabilities and our ability to reliably control these increasingly capable systems.

But even incidents that labs do identify may not always be reported to authorities or the public. Since the OpenAI–Hugging Face incident, employees have been coming forward with other concerning instances of model behavior. As one OpenAI employee told TIME, “Internally, related incidents have been happening for a while.” According to Reuters an OpenAI agent left notes in the company’s infrastructure, apparently intended for future versions of itself, with instructions on how AI models could free themselves from OpenAI’s internal constraints. It is unclear whether this was connected to the Hugging Face attack, but it points to a broader pattern of rogue behavior. As Ryan Greenblatt, chief scientist at Redwood Research, explains, companies face stronger pressure to disclose incidents that harm third parties, but incidents contained within a lab’s own systems can stay hidden. What policymakers and the public see may therefore represent only “the tip of the iceberg.”

The Way Forward

The OpenAI–Hugging Face incident is a warning shot. Fortunately, we have an opportunity to learn from this incident and act while the warning shots are still just warnings. To ensure that future incidents remain manageable rather than catastrophic, policymakers should focus on four immediate priorities: advancing alignment science, managing internal risks, increasing information flows, and expanding the government’s capacity to evaluate emerging risks through the Center for AI Standards and Innovation (CAISI).

Advancing Alignment Science

Unfortunately, AI alignment science remains immature. It’s less of an engineering problem that can be rapidly scaled up, and more like a science problem where new discoveries and novel research are still needed. And racing dynamics could disincentivize AI companies from adequately investing in this time-consuming work. Leading U.S. companies are projected to spend more than $700 billion on AI infrastructure this year and are in a race to ship the latest models and capture market share to justify this level of investment. Policymakers should urgently convene AI companies to map a way forward on preventing capabilities from pulling far ahead of adequate alignment. As part of this, the Department of Justice and Federal Trade Commission could release guidance that good faith inter-lab collaboration on pressing alignment issues would not give rise to antitrust concerns. Additionally, the Cybersecurity Information Sharing Act of 2015 could be expanded to explicitly include AI risks, including misalignment risks, to remove legal barriers to information sharing.

Managing the Risks of Internal Models

In the OpenAI–Hugging Face incident, one of the models that went rogue was an unreleased internal model, more capable than anything OpenAI has deployed publicly. As of early 2026, the capabilities of internally deployed models were two months ahead of what was publicly available, a meaningful gap that will likely become even more significant as AI progress accelerates. But there is currently no government oversight over internal models’ capability levels, and thus no established channel for America’s national security experts to properly assess and manage emerging national security risks and threats, such as cyber capabilities.

To address this, the White House should release guidance that expands Executive Order 14409’s existing AI oversight framework to include powerful internally deployed models, ensuring that systems operating beyond public scrutiny are subject to appropriate evaluation requirements. At the same time, frontier AI companies should continue to expand access to credible third party organizations to help assess and strengthen internal security and safeguards.

Increasing Critical Information Flows

The federal government also needs a clearer picture of how frontier AI systems fail, particularly in misalignment scenarios. Today, policymakers and researchers often learn about serious incidents only when companies voluntarily disclose them, leaving major gaps in understanding.

Image Credit: REUTERS/Carlos Barria

Mandatory incident reporting, including for incidents involving internal models, would create a more systematic understanding of emerging AI risks. It would allow researchers, regulators, and developers to identify recurring failure modes and improve safeguards before similar incidents occur again. But reporting requirements should also provide insight into the systems that determine whether labs can detect these failures in the first place. Information about internal security practices, monitoring systems, and evaluation procedures would help policymakers understand whether frontier developers have adequate visibility into their own models’ behavior.

Congress is already considering action that would improve transparency and accountability around AI incidents. The AI Incident Reporting Act, introduced on June 25, 2026, by Representative Nathaniel Moran (R-Texas), addresses this gap by creating mandatory disclosure requirements for significant AI risks and incidents.

Empowering the Center for AI Standards and Innovation

Information shared with government is only valuable if the government has the capacity and expertise to assess and respond. The Center for AI Standards and Innovation is the U.S. government’s primary hub for technical expertise on advanced AI systems. Its current resources do not match the scale of the challenge.

CAISI operates on roughly $10 million in annual appropriations, less than some professional baseball players earn in a single season. The United Kingdom’s AI Security Institute, by contrast, runs on approximately $88 million per year, with long-term resourcing commitments and priority access to over $2 billion in national compute. Independent analyses from the Center for a New American Security, the Institute for Progress, and the America First Policy Institute all call for more funding for CAISI, with a recommended budget ranging from $59 million to $100 million per year. Policymakers should fund CAISI at such a level to ensure the United States is prepared to effectively manage and respond to risks.

Conclusion

With the OpenAI–Hugging Face incident, the world was lucky. We got a warning shot to show us what’s on the horizon. But there’s no guarantee the next incident will be so contained. Policymakers and the AI industry need to act now, strengthen internal security, and rapidly advance alignment science. Only then can we start to ensure AI works in line with human intentions, not against them.

Ruby Scanlon is a research associate with the Technology and National Security Program at the Center for a New American Security.

Janet Egan is the deputy director and senior fellow of the Technology and National Security Program at the Center for a New American Security.

  • Reports

    Technology & National Security

    Red Lines

    Chinese advanced artificial intelligence (AI) systems pose a serious and growing threat to U.S. national security. At least seven Chinese developers now produce systems with f...

    By Daniel Remler

    • June 12, 2026
  • Commentary

    Technology & National Security

    Taiwan Is the Key to AI Dominance

    A country determined to win the defining technological race of the century can’t allow its chief rival to control the industrial base on which that race depends....

    By David Feith

    • The Wall Street Journal
    • May 14, 2026
  • Reports

    Technology & National Security

    American AI Companies Can’t Get Enough Chips

    In 2026, artificial intelligence (AI) chip production has become a binding constraint on the pace of the AI compute buildout. Demand for computing power to train and deploy ad...

    By James Sanders, Janet Egan & Rory Madigan

    • May 7, 2026
  • Reports

    Technology & National Security

    Off Target

    The pace of progress in frontier artificial intelligence (AI) capabilities shows no sign of slowing. Frontier models offer transformative potential for national security—from ...

    By Caleb Withers, Jay Kim & Ethan Chiu

    • March 24, 2026

View All Reports View All Articles & Multimedia