Don't Ask the Model What It Did
My denial engine kept citing appeal deadlines that did not exist. The claim was real. The denial was real. The deadline was made up. And right next to it, in a neat little field, the model had written down which lookups it ran and what they returned. It said it had checked. It hadn't. So I had a system writing its own alibi, and the alibi was wrong.
01The model was a witness, not a log
Here is how the engine worked. A denied claim comes in. The model reads it, calls tools to pull payer rules and filing timelines, and writes an appeal plan. Along with the plan, it returned a field describing its own run: which tools it called, which sources it touched, where each fact came from.
That seemed reasonable when I built it. Who knows better what the model did than the model?
It turns out almost anyone does. A model does not remember its run the way a log does. It rebuilds it after the fact. When it wants a deadline and the data is thin, it fills the gap. Then it tells you, in good faith, that there was never a gap. In most software that would be odd. In medical billing, a wrong appeal deadline can cost you the claim.
An audit trail the model writes about itself is testimony. You would not let the defendant write the court record.
What broke
For a while I stored the model's account of its own run as fact. The reconciliation pass kept flagging deadlines the data could not back up, and every one of them came with a confident provenance line saying it had been checked.
02Let the server keep the record
The fix was boring, which is usually a good sign. I put a callback on the tool layer. Every time a tool runs, the server writes down what was called, with what inputs, and what came back. The model is never asked. It can't edit that record, and it can't improve on it.
Now when the appeal plan names a deadline, I check it against what the tools actually returned. If no tool returned a date, there is no date. The plan goes back for review instead of out the door.
Key insight
If the system can see something for itself, never ask the model to report it.
03The second bug was the same bug
Then something surprised me. I had a separate problem on a separate list. The small model I used for this step kept returning JSON that would not parse. I had it filed under flaky output, and I was treating it the usual way: retries, stricter prompts, more examples.
When I looked at where the parse failures happened, most of them were in the audit trail field. That was the longest and most nested part of the output. It was also the part the model had no business writing. The hardest thing to get right was the least worth having.
So the field is coming out of the output contract entirely, and the server builds the whole trail. Retries only make a broken field fail less often. A field that isn't there can't fail at all.
The result
Once the server kept the record, every deadline in an appeal plan could be traced to a real tool result. The fabricated provenance stopped, and the biggest source of parse failures now has a date to disappear.
04Why it travels
I build and run AI systems in a lot of unrelated fields, and this pattern keeps coming back. People ask the model to do the task, and then to describe the task. Those are two jobs. The second one is almost always something the code around the model already knows.
The payer-calling voice agents are the clearest case. An agent can say it confirmed a claim status. The call recording and the transcript can say whether it did. In the quoting engine, the model can explain how it got to a price, but the pricing code already knows every line it added. In a local training loop, the model can tell you it improved, but only the eval harness gets to say so. Different industries, same mistake. Each time you hand the second job to the model, you add a way to be lied to and a way for the output to break.
So the rule I use now is simple. If only the model can produce something, it stays in the contract. If the system can see it, it comes out.
- 1Don't ask the model what it did. Record it where it happens.
- 2Treat every field in the output contract as a place to fail, and cut the ones you can compute.
- 3When two bugs keep turning up in the same spot, check whether they are one bug.
- 4Get provenance from your logs, never from something that gains by sounding sure.
The model is good at the work. It is just not a reliable witness to the work. Once I stopped asking it to be one, the system got more honest and more stable at the same time, which makes sense, since it had been one problem all along.