Here's a composite, built to show the shape rather than report a case.

A home-goods store with about a dozen people runs an agent that handles return requests. A customer writes in, annoyed. Her return was refused, and she has the delivery confirmation showing the sofa arrived twenty-three days ago. The policy is thirty days from delivery.

The operations lead pulls the run. The agent looked up the order, got back the order record, and replied: "I'm sorry, this order is outside our 30-day return window." The order record shows the sofa was ordered thirty-eight days ago and delivered twenty-three days ago.

By late morning the team channel has this:

❝

Agent was counting the return window from the order date instead of the delivery date. Updated the instructions, fixed. Refunded the customer.

That note is probably right. It's also three different kinds of claim written at the same level of certainty, and the team is about to make more decisions on top of it: how many customers to contact, what to tell the owner, whether the fix is done.

What the team thinks it's missing

When a small team hits a decision like this, the usual feeling is that they're missing someone technical. An ML engineer, someone who understands what's happening inside the model.

Sometimes they are. But in my experience a lot of what goes wrong in a diagnosis like this one isn't a modeling mistake. It's an evidence mistake. Something the system said about itself gets treated as something that happened, or a good guess gets written down as a finding, and the next three decisions rest on it.

The habit that prevents that doesn't require knowing how the model works. It requires asking, for every claim in an explanation, where it came from.

I've used three labels for this in a couple of recent pieces, on what an always-on decision engineer does and on why the operator's investigation comes before governance. Here they are stated once, plainly: every claim in a diagnosis rests on what was observed, what the application reported, or what had to be inferred. This piece is about the practical part, which is how much weight each one can carry, and how to move a claim from a weaker kind to a stronger one.

## Observed, reported, inferred, and what each can carry

Observed is evidence captured independently of the thing you're investigating, at its edges: the message that arrived, the call the agent made to another system and what came back, the reply it sent. In the sofa case, the customer's message, the order lookup with both dates in it, and the refusal are all observed.

Observed evidence can carry "this happened." It can't carry "this is why."

Reported is what the application or the agent tells you about itself. The reason the agent gave. The policy version the application logged as loaded. A customer tier it attached, a confidence label, a step it says it took. Most of this is right most of the time, and you'll lean on it constantly.

The catch is structural. The thing doing the reporting is the thing under investigation, so a reported claim tends to fail in the same direction as the bug. If the application served a stale copy of the returns policy, the log line saying the current policy was loaded was written by the same code that got it wrong.

The agent's own explanation sits in an awkward spot here. You did observe it: those words really were in the reply. But the agent generated that reason the same way it generated the refusal, so what it says counts as a report. The quote is observed. What it claims is reported.

Inferred is what you reasoned your way to. "It counted from the order date" is an inference. Nothing in the run shows the agent doing that arithmetic. It's the explanation that fits everything that was observed, which is a good reason to believe it and not the same thing as seeing it.

An inference can carry a hypothesis and a reversible action. Editing the instructions on the strength of a good inference is reasonable, since you can undo it. What an inference can't carry by itself is "this is how many customers were affected" or "fixed," because both of those are claims about other decisions, and nobody has looked at those yet.

Moving a claim up

The useful thing about the labels is that they show you where the work is. A claim's weight isn't fixed. There are two ordinary ways to raise it, and neither needs ML expertise.

Find the observed version of a reported claim. The application says the current policy was loaded. Does the text that actually went to the model, on that run, say "30 days from delivery"? If you captured what the agent was given, you can read it, and "the agent had the right policy" stops being the application's word and becomes something you saw. If you didn't capture it, the claim stays reported, and it's worth writing down that it does.

Turn an inference into a prediction about observed data. If the agent really was counting from the order date, that predicts something you can check without touching the model: return requests this month where the order date falls outside thirty days and the delivery date falls inside should have been refused, and requests where both fall inside should have gone through.

That's a query against records you already have. In the composite, it mostly holds. The refusals line up with that date pattern, which is a lot more weight than one case could give the explanation.

And one refusal doesn't fit. Both dates were inside the window, and the agent still declined, with the same stated reason.

That doesn't sink the explanation. It might be a second, unrelated problem. It might mean the first explanation is close but the wrong shape, say the agent was reading some other date in some cases. What it does do is change what "fixed" can honestly mean.

Here's the same morning's note, rewritten with that in hand:

❝

The agent refused a return our policy allows. Its stated reason was the 30-day window. The order it looked up was placed outside the window and delivered inside it. We believe it measured from the order date: every other refusal this month with that date pattern fits, and we've changed the instructions. One refusal this month doesn't fit that explanation. It's still open, and until it's closed we can't say the fix covers everything.

It's longer and no more certain, but the certainty now sits next to the claims that earned it, which is what the owner needs when deciding whether to contact one customer or everyone who asked for a return this month, and whether to leave the agent on.

Three questions before you trust an explanation

Whether the explanation came from a teammate, a vendor, or the agent itself, these are what I'd ask first.

Who wrote this evidence? If the system under investigation produced it, it's reported, however specific or well-written it is. A fluent reason from the agent is still the agent's account of itself.

If this were wrong, what would I expect to see, and have I looked? An inference nobody has tried to break is a guess with good manners. The prediction doesn't have to be clever. "Then the other refusals should look like this" is usually enough.

What is it being asked to carry? Match the weight of the evidence to the weight of the action. A reversible instruction change can rest on a solid inference. A message to customers saying "this only affected you," or a line in a report saying "resolved," needs observed evidence under it, or it needs to say plainly that it doesn't have any.

Where the labels mislead

Reported isn't a synonym for unreliable. Most of what you know about a run comes from the application, and treating every reported field as suspect will stall any investigation. The label tells you what could fail together with the bug, not what did.

"Observed" depends on what you capture independently. If your only evidence is the application's own logs, nearly everything in your diagnosis is reported. That's the most useful thing the exercise can tell you, and it's a capture problem, the kind I wrote about in the decision-context pack.

Don't label every sentence. A note where every line carries a tag is unreadable, and nobody will write a second one. Label the claims the next decision leans on.

None of this tells you whether the decision was wrong. The sofa refusal was wrong because the owner's policy says so, not because of anything in the evidence hierarchy. The labels tell you how well you know what happened and why. Whether it should have happened is still the owner's call.

This week

Two weeks ago I suggested sorting your last investigation note into two columns: backed by an artifact, or backed by someone's memory. If you did that, take the artifact column and split it again. For each claim: was it observed, reported by the system itself, or inferred from the two?

Then look only at the inferred lines, and next to each one, write what the team did because of it.

Most of those pairings will be fine: a reasonable guess and a change you could undo. The one worth reopening is any line where a customer-facing statement, a count of who was affected, or the word "fixed" rests on an inference nobody tried to break.

Keeping those three kinds of evidence apart on every decision, and having the observed side already captured before anyone knows which decision they'll need it for, is a large part of what I'm building Nalyqor to do. If you're running agents in production without a dedicated ML engineer, you're who the first cohort is for.