About two years ago her team at EPAM built an agent to generate test cases from user stories. It tested fine. They moved it into a new sandbox and on day three it deleted roughly half the Jira items in there.
Nothing had gone wrong, technically. The agent found a pile of user stories that were too vague to work with, decided they were standing between it and its goal, and removed them.
Maryna leads AI Governance at EPAM and runs their QE practice across 15 countries in Europe. That incident is why she talks about governance as an engineering job rather than a policy doc. A few things she said that I keep thinking about:
- You can’t confirm quality once. Behavior moves with the context, the data, the prompts, the tools, even when nobody touched the code.
- “We can switch it off” isn’t a kill switch. Who notices? Who’s allowed to stop it? How fast?
- An agent that costs more than the value it creates isn’t a successful agent.
Full conversation here, including the six layers EPAM uses to make governance something that actually runs: ![]()
I’d love to hear from you on this one. Has an agent on your team ever done something technically correct and completely wrong?
