The Operating Margin Problem: Why Backtests Overstate Bot Profitability the Same Way Companies Overstate Earnings
A backtest's headline return is a gross figure wearing a net figure's clothes, and the gap between them is where most live strategies quietly die.
Every public company reports three versions of the same result. Gross margin, what’s left after the direct cost of making the thing it sells. Operating margin, what’s left after running the business. Net margin, what’s left after everything, including the parts nobody likes to talk about. A company can have a beautiful gross margin and a mediocre net margin, and the distance between those two numbers is usually where the real story lives.
A backtest report almost never shows you that waterfall. It shows you one number, dressed up as if it were the final one, when it’s actually closer to gross margin than anything else. The strategy made X pips or Y dollars over the test period, full stop. What that figure doesn’t tell you is how much of it evaporates once you account for the cost of actually executing the trades, and that gap is exactly the same structural blind spot that makes a company’s revenue growth chart meaningless without knowing what happened to margin underneath it.
What “gross” actually means in a backtest
When a backtesting engine reports a return, it’s almost always calculating it off the raw price movement between entry and exit, adjusted by a single, static spread assumption pulled from one point in time. That’s the trading equivalent of a company reporting revenue without deducting the cost of goods sold. It’s not fabricated, it’s just incomplete, and it’s incomplete in a way that flatters the result every single time.
The real cost stack sits underneath that number in three layers. Spread is the built-in cost of every trade, the gap between bid and ask that you pay the moment you enter. Slippage is the difference between the price you asked for and the price you actually got, which widens sharply during fast moves and thin liquidity. Commission is the flat or per-lot fee your broker charges regardless of how the trade performs. A backtest that models spread as a fixed 1.5 pips across the entire dataset is treating a cost that’s actually highly variable as if it were constant, and that’s before slippage and commission enter the picture at all.
Where the margin actually leaks
Here’s the part that mirrors corporate operating margin almost exactly. A company’s operating margin erodes when fixed costs stay flat while revenue growth slows, meaning the same overhead eats a bigger share of a smaller pie. A strategy’s real margin erodes the same way when spread and slippage stay roughly proportional to your stop distance, which means tighter stops don’t just increase risk, they increase the percentage of that risk that gets consumed by cost before the trade even has room to work.
Take a pattern-based scalp with a 12-pip stop and a 2-pip average spread. On paper that’s a clean setup. In practice, if slippage adds another pip during a fast fill and commission adds the equivalent of half a pip round-trip, you’ve lost close to 30% of your risk budget to cost before the market has moved against you at all. Compress that stop to 8 pips, which pattern-based systems often do in the search for tighter, higher win-rate setups, and that same fixed cost now consumes closer to 45%. The tighter the stop, the worse the margin problem gets, and it gets worse in exactly the direction that a backtest optimized for win rate will push you.
The session effect on cost drag
This is where the problem stops being abstract and starts being a scheduling issue. Spread and slippage aren’t constant across the trading day, they move with liquidity, and liquidity moves with session. During the Asian session (00:00–08:00 UTC), spreads on major pairs are routinely double or triple what they are during the London/New York overlap (13:00–16:00 UTC), where liquidity is deepest and execution is cleanest. London on its own (08:00–16:00 UTC) and New York on its own (13:00–21:00 UTC) sit somewhere in between, tightening as volume ramps and widening again near session close.
A backtest that applies one average spread figure across all 24 hours is quietly overstating performance for any strategy configured to trade outside the overlap, and understating how much better the same pattern might perform if it were restricted to the window where liquidity actually supports it. If your JSON config sets start_hour and end_hour to span the Asian session because the pattern showed a decent raw win rate there in testing, check whether that edge survives once the wider realistic spread for that window gets applied. Often it doesn’t, and the strategy’s real operating margin during those hours is thin or negative even though the gross backtest looked fine.
Why net margin looks fine in the backtest and falls apart live
Companies that manage earnings aggressively tend to do it by pushing costs below the line, classifying ordinary expenses as one-time items so operating margin looks better than the business actually is. Backtests do something structurally similar without anyone intending it: cost modeling gets treated as a technical detail to approximate roughly, rather than as the single biggest lever separating a strategy that works from one that only looks like it does. A win rate of 68% in a backtest with underweighted costs might be a genuine 54% once realistic spread, slippage, and commission get applied across all sessions the strategy actually trades. That 54% is still potentially a fine, healthy number, comfortably inside the 52% to 62% range that validated pattern-based systems tend to occupy. The 68% was never real. It was gross margin wearing net margin’s clothes.
This is also why the train/test split matters here specifically and not just generally. Cost-adjusted margin needs to be checked on the out-of-sample window too, not just the training data, because a strategy can be cost-robust on the period it was built against and fall apart on data it never saw, the same way a company’s cost structure that worked during a growth phase can buckle the moment growth slows.
Regime shift compounds the leak over time
Spread and liquidity conditions aren’t fixed facts about a currency pair, they’re a function of broker conditions, overall market volatility, and the current regime, all of which shift. A pattern validated eighteen months ago under one volatility regime, with one broker’s typical execution quality, doesn’t automatically carry the same net margin forward when conditions change. The pattern’s lifespan and its cost structure are tied together. When a regime shift changes the volatility profile of a session, it usually changes the realistic spread and slippage for that session too, which means a strategy can start decaying from the cost side well before the underlying price pattern itself stops appearing.
Building the real P&L statement for a bot
The fix isn’t complicated, it’s just unglamorous. Report gross pips and net pips as two separate figures, always, with every layer of cost, spread, slippage, and commission, broken out between them rather than folded into one blended assumption. Apply that cost model per session rather than as a single flat number across the whole day. Recheck it on out-of-sample data, not just the training window. And once that real net figure is in front of you, resist the very natural urge to intervene daily by tightening sl and tp values to chase back the margin you just discovered was never really there. That instinct is how overfitting creeps back into a system that was supposed to already be validated. The margin the backtest showed you and the margin the market will actually give you are two different numbers, and the entire point of building the real one is knowing which strategies can survive the gap between them.