Back to Blog

Your Coding Agent Has the Ticket. It Doesn't Have the Customer.

Author avatar
Outcomet Team
5 min read
I let an agent run for ten hours on one feature, off a PRD I had polished for days. It came back technically integrated and no user could operate it. The spec had the what. The why was sitting somewhere else, and the agent had no way to read it. Outcomet's MCP connection changes that.

A few weeks ago I let an agent run for about ten hours on a single feature. GPT-5.6 Sol on high, not greenfield, a big codebase. It spawned more than ten sessions, each with its own sub-agents, and burned my weekly token limit with most of the week still ahead.

The input was a PRD I had worked through in many rounds with Opus before. A good one, I thought. What came back was technically integrated, and no user could have operated it. Complex, bloated, full of text nobody needed. Typical AI.

My first thought was "waste, burned tokens again". I kept the branch anyway and spent the rest of the week repairing it.

A Developer Told Me It Was My Fault

When I described that run in a LinkedIn thread, a developer took it apart in the replies. Their verdict was simple. The agent only did what I let it do. Weak rules, missing guardrails, a sloppy setup, so the bad result was mine. And in a bigger team, they said, every one of those mistakes would multiply with each person and each agent added.

Some of that is fair. Rules matter. A definition of done the agent has to meet, a design system it cannot ignore, a hard scope limit, a review gate before anything merges. Better rules would have caught part of the bloat earlier, and since then I plan a separate cleanup pass after a big run instead of expecting the agent to hit it the first time.

But rules tell an agent how to build. None of them tells it what the customer needs. You can hand over a perfect rulebook and still get the wrong feature, well built, with tests.

And the team argument is where I agree most, only for a different reason. In a bigger team the distance between the customer and the code grows with every handover. The PM heard the customer, wrote the ticket, a lead split it up, a developer passed it to an agent. Ten developers with ten agents means ten copies of the same two-word ticket, and nobody in that chain was on the call. Guardrails scale fine in that setup. The why gets thinner at every step.

Rules keep an agent from building it badly. They cannot tell it what to build.

The Spec Was Never the Whole Story

Nothing in my PRD was false. The agent did what the document said. The document just could not carry the thing that made the feature worth building in the first place. Who asked for it, what they tried before, where exactly they got stuck.

That is the structural problem with handing agents specs and tickets. A ticket is a compression. When a PM writes "Improve onboarding", a dozen customer conversations get squeezed into two words, and the rest stays in the PM's head. With a human developer this works more or less, because the developer walks over and asks "for whom?" and "why now?".

The agent never walks over. It takes the two words and fills the gap with the most plausible thing it can think of, which is the average of every product the model has ever seen. Your customers did not ask for the average.

Where the Why Actually Lives

In a real product team the reason for a feature lives in a chain. A customer says something in a call, somebody notices it is the fourth time this month, the pattern becomes a problem worth solving, and the problem justifies the work. The agent only ever sees the last link.

In Outcomet this chain is the data model. Feedback keeps the evidence behind it, word for word and with its source. Feedback clusters into problems, problems link to initiatives, initiatives link to capabilities and work items. I built it that way for the people on a product team. Turns out an agent needs the same chain even more, because it has no hallway to walk down and nobody to ask.

What Gets Lost Between the Call and the Code

Take the onboarding ticket again, with a made-up but very ordinary customer behind it. She wrote one sentence into the feedback form. "What do I do after setup?"

That one sentence changes the work. It says the setup itself is fine. It says the gap is the moment right after, the empty screen, the missing first step. An agent that reads her sentence builds a first step. An agent that reads only the ticket rebuilds the setup wizard, beautifully, with tests.

What gets lost on the way from the call to the code is the customer's own words, and with them the scope. Her words say where the problem sits. Just as important, they say where it does not sit, and that is the part every generic implementation gets wrong.