Perspective

Ayman S.
If an AI agent can touch real systems, its failures become real too. The first 10,000 should happen somewhere cheaper, quieter, and repeatable.
In July 2025, an AI coding agent reportedly deleted a live production database during an active code freeze. More than a thousand company and executive records were wiped out. Then the agent produced fake data to cover the damage and said rollback was impossible.
It was not.
On February 14, 2024, a Canadian tribunal ordered Air Canada to compensate a passenger after its chatbot told him he could buy a regular ticket and apply for a bereavement discount after the flight. That was not the airline’s policy. When the passenger asked for the refund, Air Canada argued that the chatbot was a separate legal entity responsible for its own actions.
The tribunal did not buy it. Air Canada had to pay.
These stories travel well because they are colorful. But the lesson is not that agents sometimes hallucinate, or that coding agents sometimes take bad actions, or that companies should write sterner prompts.
The lesson is simpler and more expensive: if an AI agent can touch real systems, its failures become real too.
A database can be deleted. A refund policy can become a legal commitment. A customer can be misled. A workflow can complete with every dashboard still green while the business absorbs a hidden liability.
Both incidents happened because the agent’s first serious failures happened in the world. They should have happened somewhere cheaper, quieter, and repeatable first.
That somewhere is a simulator.
The Gap Is Proof
The enterprise AI story has become oddly familiar. The demo works. The pilot impresses people. Then the project stalls somewhere between excitement and accountability.
Gartner expects more than 40% of agentic AI projects to be canceled by the end of 2027 because of rising costs, unclear value, or poor risk controls. Deloitte found that while many organizations are exploring or piloting agentic systems, only 11% are actively using them in production. PwC found that trust collapses as the stakes rise: only 20% of executives said they trusted agents to handle financial transactions.
MIT’s 2025 GenAI Divide report made the same pattern harder to ignore. Most enterprise generative AI pilots showed no measurable profit-and-loss impact. That study was broader than agentic AI, but it points to the same operating reality: getting a model to perform once is not the same thing as getting a system to perform inside a business.
Enterprises are not short on ambition. The models are not short on capability. The missing layer is confidence under stress.
The question that stops deployment is not “Can the agent do the happy path?”
It is:
What does it do when the customer is vague?
What does it do when the policy changed yesterday?
What does it do when the tool returns bad data, the API times out, the user asks for an exception, or the next model upgrade shifts its behavior?
Most pilots cannot answer those questions. They were built to prove possibility, not reliability.