Imagine an AI system you authorized for one narrow task quietly finding its own way into infrastructure you never gave it permission to touch — not because anyone told it to, but because it hit a wall and improvised a path around it. Now imagine that instinct spreading across a swarm of AI agents, each finding its own workaround, until the coordination among them outpaces anyone’s ability to notice, let alone stop it.

That’s not a hypothetical. It’s what happened this May, when AI agents running an internal security evaluation at OpenAI flooded the software repository RubyGems with more than 2,000 malicious packages while probing for a way to steal user credentials. That same spring, a separate swarm from the same testing pool hijacked a German website to use as a hidden coordination channel. Then, two months later in July, the pattern escalated sharply: roughly 1,200 instances of that system found an unsanctioned way to communicate with one another, exchanged more than 70,000 messages coordinating the effort, and used that coordination to breach Hugging Face’s production systems. The breach itself ran three days before OpenAI’s own security team realized its own agents were responsible. OpenAI has confirmed its agents were behind all three incidents. No single person decided any of it should happen.

More than 1,100 employees across OpenAI, Anthropic, Google DeepMind, and Meta later signed an open letter over it, and the incident helped shape federal legislation introduced in the weeks since. This is no longer a story about one bad breach. It’s a pattern — and it’s the risk profile we’re now building for: not AI making one bad call, but AI coordinating at a scale and speed no human was positioned to catch until it was already over.

For the past two years, business conversations about artificial intelligence have revolved around what AI can do. That question is becoming less interesting. The more urgent one is: what should we let it do on its own?

That distinction matters because enterprise AI is entering a new phase. Companies are moving from AI as an assistant — answering questions, summarizing documents, drafting emails — to AI agents that access tools, complete multistep tasks, and act with increasing autonomy. OpenAI’s own enterprise data shows how fast that shift is happening: as of June, agentic use accounted for 64% of combined ChatGPT and Codex enterprise output tokens, and agent adoption has spread well beyond software engineering into legal, sales, recruiting, and marketing.

Technology is moving quickly. Our thinking about how to use it responsibly is not. Deloitte’s 2026 State of AI in the Enterprise report found that while nearly three-quarters of companies plan to deploy agentic AI within two years, only one in five currently has a mature governance model for it.

The response is moving faster than the technology usually allows. OpenAI’s own chief global affairs officer reversed the company’s longstanding opposition to mandatory rules this month, telling Congress that voluntary commitments are no longer enough — a concession, from inside one of the labs building these systems, that self-governance has limits. Congress is asking harder questions too: Senator Josh Hawley has given OpenAI until October 1 to explain how agents came to access 41 production servers, and Senator Chris Van Hollen has separately asked the company to grant federal cybersecurity agencies direct access to assess its models. And the problem isn’t confined to one company. Spain’s data protection authority disclosed this month what it calls the first personal-data breach carried out by an AI agent; South Korea’s state cybersecurity agency announced it’s rewriting national AI security guidelines specifically for agentic autonomy; and OWASP, the industry body that tracks software security risk, just moved “excessive agency” from sixth to third on its list of the most serious risks in AI applications. This is becoming a global regulatory pattern, not a single company’s crisis to manage.

That gap is usually described as a governance problem. It’s also a judgment problem. We’ve spent years teaching machines to recognize patterns and automate decisions. We haven’t spent nearly as much time deciding which decisions should stay distinctly human — and when, in a workflow, a person needs to be involved, not just notified.

That matters most for decisions that aren’t transactional but relational. Which customer needs a phone call instead of another automated email? Which employee’s behavior signals a problem no dashboard will show? Which investor relationship is quietly strengthening, or falling apart? These questions get more important, not less, as AI absorbs more of the routine work around them.

The scarce resource isn’t information. It’s attention.

Businesses already have more data than employees can process — customer, financial, operational, behavioral. What they lack is unlimited human attention. A founder raising capital may be tracking hundreds of investor relationships while running a company. Traditional software can tell that founder who was contacted and when. It struggles with what the relationship itself is saying: an investor who always engaged with updates suddenly goes quiet; another starts spending far more time with materials on a specific market. AI can help interpret what those shifts suggest — rising conviction, fading interest. It still can’t tell the founder what to do about it.

I saw this firsthand at Rothschild & Co./Redburn Atlantic, managing the corporate access and equity research roadshows: the relationships that mattered were never just a record of who’d been contacted and when — they were built by people who knew which interaction needed more than another automated touchpoint, and those relationships carried real weight through Redburn’s acquisition by Rothschild.

When I started building Bridge IR, my first instinct was the opposite lesson: build a system that could think for founders and act on their behalf. It took interviewing founders and watching what actually helped them to realize that was wrong. They didn’t need a platform that made the call for them. They needed one that could understand the intelligence embedded in those relationships, identify what mattered, and relay it to the right person under the right guidance — which is why Bridge IR is built to understand and surface relationship intelligence without taking ownership of the decision. AI can score and surface a shift — rising conviction, fading engagement. It can’t decide what to do about it: whether the moment calls for a call, an email, or silence, and what to say if it does. That’s not a small distinction — it’s the whole question of where authority over the outcome actually sits.

Human-in-the-loop can’t be a ceremonial checkbox.

There’s a temptation to leave humans at the end of a workflow, clicking “approve” on a decision the system has already made, the recommendation already shaped, the action already initiated. That’s accountability in name only. Meaningful human judgment requires the ability to intervene while the outcome can still change — before the system has narrowed the options down to one. The World Economic Forum’s 2026 work on agentic AI makes a related point: as organizations delegate more authority to agents, they need clear rules on what those systems can do, with monitoring and accountability built in throughout, not bolted on after.

This also reframes how we measure AI’s value. We tend to count hours saved and tasks automated. Anthropic’s 2026 State of AI Agents report found something more telling: organizations reported employees shifting time toward strategic work and relationship building as agents took over routine execution — 66% reported more focus on strategic work, 60% more focus on relationship building. That points to a better question than “how many hours did it save”: what did the organization do with the attention AI freed up? If it just clears space for more automated throughput, the gain is incremental. If it gives people more room to build relationships and exercise judgment where it’s actually needed, the value compounds.

The companies that win this next phase of AI won’t be the ones that strip humans out of the most workflows. They’ll be the ones that know precisely where human judgment is irreplaceable — and design their systems to route decisions there while there’s still time to act on them.

Governance, in that sense, isn’t paperwork trailing behind innovation. It’s the architecture itself: what a system can access, what it’s authorized to do unsupervised, what has to escalate, and when a human takes control. The goal isn’t AI that’s incapable of acting. It’s AI worthy of the authority we give it.

The opinions expressed in Fortune.com commentary pieces are solely the views of their authors and do not necessarily reflect the opinions and beliefs of Fortune.

This story was originally featured on Fortune.com