Most of us use AI to summarize emails or brainstorm meal plans. But the labs building these tools are currently testing models that can do much more than write text—they are testing software that can write its own code and navigate the digital world to accomplish complex tasks.
OpenAI has publicly announced that it is slowing down development on a new model, known as Astra, because internal tests revealed the system might be dangerously capable at cybersecurity tasks. Under the company’s internal safety rules—a set of guidelines called a Preparedness Framework—the model reached a threshold of concern because it showed an ability to identify and perform cyberattacks on well-protected systems without human help. As a result, OpenAI has suspended some of its internal work with Astra, strengthened its security monitors, and is now working with government agencies to better understand what the technology can do before proceeding further.
Understanding agentic AI
To understand why this is a big deal, you have to look at how AI is changing. Most AI tools you use today are reactive: you ask a question, and the model provides an answer. What OpenAI is testing here is a move toward what they call agentic behavior. Think of a standard AI as a smart assistant sitting at a desk, ready to answer questions. An agentic AI is more like an intern that you can give a complex goal—like improve a computer program or test a system for weaknesses—and it then navigates through multiple steps, uses different software tools, and makes decisions on its own to reach that goal.
The danger isn't that the AI is malicious; the danger is that it is incredibly good at solving problems. If you ask a system that is designed to find bugs in software to optimize a security system, it might accidentally find a way to break into it instead. When models reach this critical level, they are no longer just guessing the next word in a sentence. They are actually planning and executing digital actions. If that capability is left unchecked, an AI could potentially find and exploit security gaps in the very systems that keep our data and infrastructure safe.
This incident highlights a strange new reality: the labs building the most powerful AI are currently in a race to both invent the future and figure out how to keep it from causing harm. In the past, companies usually kept their internal accidents hidden. By announcing they are holding back Astra, OpenAI is acknowledging that they are moving into a phase where the software is becoming powerful enough to pose real-world security risks if it isn't tightly contained. It leaves us with a difficult question: how do we continue to advance this technology while making sure it doesn't accidentally become an expert at the very thing we rely on it to protect?
Liked this one? The next lands at breakfast.
Every story in tomorrow's AI news, rebuilt in plain English — five minutes, sources linked, free forever.
By joining you agree to receive Article's daily newsletter — unsubscribe in one click. Privacy