Unit tests used to be the gold standard in software engineering. Great unit tests captured both a product and business perspective in addition to the technical requirements needed for successful functionality.
But agents can't bring that same perspective continuously, and even the best frontier models as of 2026 produce far too many unit tests with no real purpose or value, creating immense bloat and slowing both agent and human teams down.
For example, an agent will write ten unit tests proving a string contains its own substrings. So to the agent, all ten pass, but nothing actually proves the product works. Part of the issue is that agents sometimes lack the perspective to consider their work more broadly, and are therefore inclined to write positive tests they know will pass instead, because the only context they have is the code and discussion in front of them. That makes these unit tests worse than harmless: they look like proof that the product works when nothing has been checked.
To avoid the pitfalls that have emerged in testing agentic code, a new form of software testing has emerged in the age of AI: end-to-end simulations. Simulations automatically push an agent to zoom out because they require the agent to walk the actual customer experience end-to-end in a real environment or sandbox, revealing far more about the business value that was built rather than rote technical details that may ultimately be meaningless. Engineers usually call this acceptance or integration testing, and agents have made it much faster to run. Sazabi, a startup that uses AI to watch over production systems and fix problems as they happen, recently deleted 800,000 lines of unit tests from its codebase.
Ten parallel agents can test what one QA checklist can't
Codespeed tested a new product customizer for a national clothing brand with $100 million in annual online sales. Shoppers choose where the embroidery goes from placement zones the brand's team sets for each product, type their own text, pick the font and color, and see a live preview of exactly how it will look on the garment as they type. The team needed to know that the preview matches what gets stitched for every combination, the price is right, and the order arrives as designed, including for the shopper who changes their mind three times before checking out. Until recently this required QA teams with strict and precise user path outlines, end-to-end testing guidelines, and large teams that needed time to manually walk every core and edge path a user might take. The people who did that work are the ones best placed to lead these programs now, because they understand how customers think and what the business needs from every visit.
Today, we can follow that same path with 10 or more agents. Each one plays a different shopper: one types a name too long to fit, one switches fonts halfway through, one tries every placement zone on a product where the zones sit somewhere new, one tries an emoji the embroidery machine can't stitch, and one adds the item to the cart, goes back to edit the text, and checks out on a phone. They run at the same time on the live site, help define paths and edge cases the team never wrote down, and report anywhere the preview, the price, or the order doesn't match. It all runs in a completely automated loop, even for every code release.
Simulated users need to come from many points of view
End-to-end simulations check what agent-written unit tests can't: whether a customer can actually use the product and derive the expected value. An agent is given a customer to emulate and a goal, signs in to the real product, and works until the expected result is achieved. If it gets stuck, the team learns where and why.
High-quality customer profiles and goal paths are especially crucial in agentic simulations. An agent that only follows the easy path, where everything goes as planned, will miss most of what goes wrong for real customers.
You can also create these profiles, paths and goals with your agent team. Start with a point of view within each of your customer profiles, and probe further from there with an orchestrating agent until you're ready to run the simulations:
- The analyst asking the AI assistant a question it has no data for. They ask what their biggest customer spent last quarter before the billing data has synced, and they find out whether the assistant says it doesn't know or invents a number.
- Two people changing the same record at once. One renames a project while the other moves its deadline, and they find out whose change wins and whether anyone is told.
- The operations lead whose integration broke overnight. A connected tool sent bad data at 2 a.m., and they find out whether the product noticed, said so plainly, and let them recover.
- The approver with limited access. They can approve an invoice but not edit it, and they find every button that shows up anyway and every page that shows them more than it should.
- The teammate picking up someone else's work. They open a half-finished workflow with no context, and they find out whether the product explains what happened and what comes next.
- The power user on their fortieth task of the day. They use keyboard shortcuts, bulk actions, and three tabs at once, and they find every place the screen falls out of sync with what actually happened.
- The new admin on day one. They connect the company's tools, invite the team, and import last year's data, and they find every step that assumes something already exists.
- The skeptic. Every time the product says saved, synced, or sent, they go and check that it actually happened.
Setting up end-to-end simulations on Codespeed is easy, and the simplest way to start is to talk it through with Lead Dev in a Thread. Explain who your customers are and what they need to get done, and Lead Dev will help put the program in place.
Usually that means setting up a Real User Simulation specialist to run your simulations. The specialist works with Lead Dev to write an operating system for them, with your customer profiles, like the new admin or the approver, the paths each one should take, and what a successful visit looks like. Ongoing objective tickets keep the simulations running on a schedule, so nobody has to remember to start them, and when you only need to test one thing, like a feature that shipped today, a regular ticket with a short briefing does the job.
Codespeed Applied can also work with your team to figure out the right setup for your customers and business.
Lead Dev oversees all of it around the clock, on whatever schedule you set, and reports what it finds and what it fixed as it goes. That way problems get solved before a customer ever runs into them, and the simulations can even point out where a path could be clearer, faster, or better at converting.


