The Hallucinated Indicator Value That Quietly Breaks an Agentic Trading Loop
A failed tool call inside a trading loop doesn't produce an error, it produces a plausible-looking number, and that's exactly why nobody notices.
Somewhere in a run that goes on for days, one cycle’s data fetch times out. The indicator function gets an empty payload instead of real candles. And the log for that cycle still reads clean: “RSI at 52.3, entering short.” Nothing about that line looks broken. Nothing about it looks any different from the four hundred cycles before it that worked correctly.
That’s the actual failure mode worth worrying about with agentic trading loops, and it isn’t the dramatic one people picture. It isn’t the agent going rogue or spitting out an obviously insane number. It’s a single missing data point getting quietly replaced by something completely ordinary-looking, inside a system nobody is watching cycle by cycle, with a log that was never designed to catch the difference.
The loop doesn’t know the difference between data and a good guess
In a one-off chat, if a tool call fails, a model will often just say so. “I wasn’t able to retrieve that.” Straightforward, visible, easy to react to. But that behavior isn’t guaranteed, it depends heavily on how the failure gets presented to the model in the first place.
Most agent scaffolding builds a prompt template that says something like “here is the current market data: {tool_output},” then drops whatever the tool returned into that slot. If the tool call errored, that slot might contain an empty string, a null, a stale cached value, or a malformed fragment, and the model still has to produce a coherent continuation from whatever’s there. Nothing in that setup guarantees a refusal. Language models are trained overwhelmingly on text where a data section is followed by confident analysis of that data. Producing a plausible-sounding indicator reading and a decisive conclusion is a far more statistically typical continuation than pausing to flag that the input looked wrong, especially when the surrounding system prompt is explicitly asking for decisive trading signals, not hedged uncertainty.
Why “looks right” is the actual danger, not “looks wrong”
A hallucinated indicator value was never going to be some absurd outlier a human would catch on sight. Pattern completion doesn’t produce nonsense, it produces whatever’s statistically typical for the context, and typical RSI values, typical ADX readings, typical bandwidth percentages, are exactly what shows up constantly in the training data these models learned from. A fabricated 52.3 looks, reads, and behaves identically to a real 52.3. There’s no tell. No formatting quirk, no suspicious precision, nothing that marks it as fabricated rather than computed.
That’s what separates this from an ordinary bug. A crashed function throws an error, someone sees a stack trace, the problem gets fixed the same day. A hallucinated value that fits comfortably inside the normal range of the metric just becomes one more line in a log that otherwise looks completely healthy. Nobody is reading every cycle of a loop that runs unattended for a week. The only thing anyone reviews after the fact is whether the trades made sense given what the log says happened, and the log says a perfectly reasonable thing happened.
Where the failure actually enters, mechanically
Trace the pipeline: fetch data, compute indicators, feed the results into the model’s context, let the model reason, parse its output into a decision, hand that decision to the execution layer. The failure this piece is about doesn’t live in the reasoning step at all. It lives one step earlier, in whatever happens when step one or step two comes back empty, stale, or malformed, and nothing in the orchestration code hard-stops the cycle when that happens.
A lot of agent loops built quickly don’t validate the shape or freshness of tool output before it reaches the model. They assume the fetch succeeded, because it usually does, and there’s no explicit check enforcing that assumption. When it doesn’t succeed, the partial or empty payload flows straight into the prompt anyway, and the model does what it was trained to do with a data-shaped slot: fill it in with something coherent. The bug isn’t really in the language model. It’s in the absence of a gate between “the tool might have failed” and “the model is now reasoning as if it definitely didn’t.”
Why logging the decision isn’t the same as logging the truth
Most setups log what the agent said, the stated indicator readings, the reasoning, the final call. That’s the audit trail people actually look at later. What usually doesn’t get logged, separately and verifiably, is whether the underlying tool call actually succeeded and what it actually returned before the model touched it.
That gap is exactly why this class of failure survives so long unnoticed. There’s no artifact anywhere that flags “the RSI computation failed this cycle” once the model has already written a confident paragraph treating its own guess as fact. The decision text and the ground truth get merged into one narrative the moment the model starts writing, and after that, there’s no way to tell them apart just by reading the log.
The fix
Treat a failed or malformed tool call as a hard stop for that cycle, not a gap to be reasoned around. If the data didn’t come back clean, skip the cycle entirely rather than letting a partial payload reach the model at all. Validate schema and freshness before the prompt gets built, not after the model has already produced an answer that assumes everything upstream worked. And log the raw tool outputs separately from the agent’s stated reasoning, so the two can be diffed against each other later. A trading loop that only records what the agent claims happened has no way of ever catching the cycle where what actually happened and what got claimed quietly stopped being the same thing.