Backtesting in Illiquid Markets: Fill Models
Backtesting in illiquid markets: how three fill models change the same strategy's reported return, and the checks that tell an artifact from an edge.
Backtesting in illiquid markets fails at one assumption: that your own order would not have changed the book it traded against. A backtest replays the historical tape, and that tape is the one that happened without you in it. In a thin name your order is a visible share of the available liquidity, and the fill model you pick can move the reported return further than the signal does.
Two of our realism posts cover the layers above this one. Latency models in HFT backtests decide when your order arrives, and queue position estimation decides where it sits in line once it does. Both of them assume the book is still there when you get there.
What goes wrong when you backtest in illiquid markets
A fill model is the rule that turns a signal into a price. Most frameworks ship with one rule, applied everywhere, and it is rarely stated out loud. The honest version of the question is counterfactual: what would the book have looked like an instant after your order arrived? No recording answers that, which is why every fill model is an assumption with a dial on it.
Thin books turn the dial loud. The first place to look is the quoted depth at the best bid and offer, the top of book. Quote sizes are published in round lots of 100 shares, and the panel below converts them to shares.
| ticker | quoted_spread_bps | avg_touch_shares | quote_updates |
|---|---|---|---|
| KO | 1.3 | 43678 | 75853 |
| SPY | 0.2 | 26933 | 586907 |
| SJM | 19.5 | 17566 | 5127 |
| AAPL | 1.3 | 11527 | 148339 |
The exact SQL behind every number
SELECT
ticker,
round(avg(toFloat64(ask_price) - toFloat64(bid_price))
/ avg((toFloat64(ask_price) + toFloat64(bid_price)) / 2) * 10000, 1) AS quoted_spread_bps,
round(avg(toFloat64(bid_size + ask_size) / 2) * 100, 0) AS avg_touch_shares,
count() AS quote_updates
FROM global_markets.cache_stocks_quotes
WHERE ticker IN ('SPY', 'AAPL', 'KO', 'SJM', 'LANC')
AND sip_timestamp >= '2026-09-16 00:00:00'
AND sip_timestamp < '2026-09-17 00:00:00'
AND (toHour(toTimeZone(sip_timestamp, 'America/New_York')) * 60
+ toMinute(toTimeZone(sip_timestamp, 'America/New_York'))) >= 600
AND (toHour(toTimeZone(sip_timestamp, 'America/New_York')) * 60
+ toMinute(toTimeZone(sip_timestamp, 'America/New_York'))) < 720
AND bid_price > 0
AND ask_price > bid_price
GROUP BY ticker
ORDER BY avg_touch_shares DESCAcross the 4 names in a two hour midday slice on September 16, 2026, the deepest top of book averaged 43678 shares at a quoted spread of 1.3 basis points, a basis point being one hundredth of one percent. The thinnest averaged 11527 shares at 1.3 basis points, over the same minutes of the same session. A 5,000 share order is invisible against the first book and larger than everything showing in the second.
The three fill models, in ascending honesty
1. Cross the spread at the touch
The simplest model fills the whole order at the posted far side price and charges half the quoted spread as cost. It is the default almost everywhere, and it is accurate while the order stays smaller than the size displayed at the touch. Past that size it reports a price that existed for the first few hundred shares and for none of the rest. In a name showing 11527 shares at the touch, that boundary arrives early.
2. Walk the visible book
The second model consumes displayed price levels in order until the order is full, then reports the volume weighted average of the levels it ate. It needs depth rather than a single quote, which means book data: MBO and MBP order book data compares the two formats a vendor will sell you. Walking the book is more honest than the touch model at every size, and it has one blind spot. It treats a snapshot as a standing offer.
3. Walk the book with decay and an impact term
The third model adds what the snapshot cannot show. Resting liquidity decays: part of the size quoted ahead of you cancels while your order is in flight, so the level you expected to fill against is thinner by the time you reach it. And price does not snap back: a portion of the move stays in place after your order is done, which is the cost your next order inherits.
Calibrating that term is the work. A common shape is a square root function of participation rate, your own volume divided by the market's over the same window, scaled by the name's volatility. The exponent and the scale factor have to be fitted, and a thin name offers few large prints to fit them on.
The gap between these models is not a rounding error. A strategy that round trips 250 times a year pays the fill assumption 500 times. Model 1 bills half of a 1.3 basis point spread per side in a book like the thin one above. Model 3 can bill several times that. Multiply the per side difference by 500 and the fill model owns most of the reported return, with the signal untouched.
The cost of size is also visible without simulating anything. Sort every regular session minute by its own volume and measure how far price travelled inside that minute.
| volume_quintile | thin_abs_move_bps | liquid_abs_move_bps | thin_over_liquid |
|---|---|---|---|
| Q1 | 1.2 | 1.1 | 1.18 |
| Q2 | 2.8 | 1.3 | 2.09 |
| Q3 | 3.9 | 1.7 | 2.33 |
| Q4 | 5.3 | 2 | 2.61 |
| Q5 | 7.9 | 2.7 | 2.95 |
The exact SQL behind every number
WITH bars AS
(
SELECT
ticker,
toStartOfMinute(toTimeZone(window_start, 'America/New_York')) AS et_minute,
toFloat64(open) AS open_px,
toFloat64(close) AS close_px,
toFloat64(volume) AS vol
FROM global_markets.delayed_stocks_minute_aggs
WHERE ticker IN ('SPY', 'AAPL', 'KO', 'SJM', 'LANC')
AND window_start >= '2026-07-01 00:00:00'
AND window_start < '2026-09-26 00:00:00'
AND (toHour(toTimeZone(window_start, 'America/New_York')) * 60
+ toMinute(toTimeZone(window_start, 'America/New_York'))) >= 570
AND (toHour(toTimeZone(window_start, 'America/New_York')) * 60
+ toMinute(toTimeZone(window_start, 'America/New_York'))) < 960
AND volume > 0
AND open > 0
),
ranked AS
(
SELECT
ticker,
avg(vol) AS avg_minute_vol
FROM bars
GROUP BY ticker
HAVING count() >= 500
),
picks AS
(
SELECT
argMax(ticker, avg_minute_vol) AS liquid_ticker,
argMin(ticker, avg_minute_vol) AS thin_ticker
FROM ranked
),
sided AS
(
SELECT
if(b.ticker = p.thin_ticker, 'thin', 'liquid') AS side,
b.et_minute AS et_minute,
b.open_px AS open_px,
b.close_px AS close_px,
b.vol AS vol
FROM bars AS b
CROSS JOIN picks AS p
WHERE b.ticker = p.thin_ticker
OR b.ticker = p.liquid_ticker
),
cuts AS
(
SELECT
side,
quantilesDeterministic(0.2, 0.4, 0.6, 0.8)(vol, toUInt64(toUnixTimestamp(et_minute))) AS edges
FROM sided
GROUP BY side
),
tagged AS
(
SELECT
s.side AS side,
1 + length(arrayFilter(x -> s.vol >= x, c.edges)) AS quintile,
abs(s.close_px / s.open_px - 1) * 10000 AS move_bps
FROM sided AS s
INNER JOIN cuts AS c ON c.side = s.side
)
SELECT
concat('Q', toString(quintile)) AS volume_quintile,
round(avgIf(move_bps, side = 'thin'), 1) AS thin_abs_move_bps,
round(avgIf(move_bps, side = 'liquid'), 1) AS liquid_abs_move_bps,
round(avgIf(move_bps, side = 'thin')
/ avgIf(move_bps, side = 'liquid'), 2) AS thin_over_liquid
FROM tagged
GROUP BY quintile
HAVING countIf(side = 'thin') > 0
AND countIf(side = 'liquid') > 0
ORDER BY quintileOver July through late September 2026, in the busiest volume quintile, the thin name's average absolute minute move measured 7.9 basis points against 2.7 for the liquid one, a ratio of 2.95 times. In the quietest quintile the thin name still travelled 1.2 basis points inside a single minute. A model that charges half a spread in that name is pricing a fraction of what one ordinary minute already does to the price.
Why quoted depth overstates reachable depth
Displayed size is an offer to trade, never a reservation. Several mechanics shrink it between the screenshot and your fill.
- Cancellation. A large share of quoted size is pulled before a fast order can reach it, and the more informed your order looks, the more of it leaves.
- Duplication. One liquidity provider can show the same intent on several venues at once, so summing displayed depth across them counts the same shares twice.
- Invisible size cuts both ways. Some depth you cannot see is genuinely there, and some depth you can see is a pegged order that reprices away from you.
- Your own footprint. Once a thin book has been lifted, the replacement quote often arrives wider, and from that point the historical tape stops describing the market you are in.
Quote traffic sets how long a displayed size is even on the screen. The midday slice counted 75853 quote updates in the deepest name and 148339 in the thinnest, over the same two hours. A backtest that samples one quote per minute discards nearly all of those revisions and treats the one it kept as durable.
Does L3 or MBO data solve this?
Message by message data, often called L3 or market by order, carries every individual order's arrival, modification and cancellation instead of a summed price level. It is the best input available for this problem, and it makes the problem visible rather than solving it. With it you can measure how much size ahead of you cancels before a lift, rebuild the book at any instant, and count how often a thin side was refreshed after being hit.
What no recording contains is the message your order would have sent, plus the messages other participants would have sent in reply. Inserting your order into a replayed book is a simulation of their behaviour, fitted on a world where you were absent. Look ahead bias has the same flavour: the dataset knows something the moment itself did not.
How to tell an artifact from an edge
Check participation rate against historical volume
Participation rate is your traded volume divided by the market's over the same window. Cap it in the backtest, at a fixed share of each name's median session volume, then see whether the strategy survives the cap.
| ticker | median_daily_shares | shares_at_2pct | median_daily_dollars_mm |
|---|---|---|---|
| SPY | 43721494 | 874430 | 33041.47 |
| AAPL | 41338918 | 826778 | 13269.55 |
| KO | 14832924 | 296658 | 1301.14 |
| SJM | 1296421 | 25928 | 157.91 |
The exact SQL behind every number
SELECT
ticker,
round(quantileDeterministic(0.5)(toFloat64(volume), toUInt64(toYYYYMMDD(date))), 0) AS median_daily_shares,
round(quantileDeterministic(0.5)(toFloat64(volume), toUInt64(toYYYYMMDD(date))) * 0.02, 0) AS shares_at_2pct,
round(quantileDeterministic(0.5)(toFloat64(volume) * toFloat64(vwap),
toUInt64(toYYYYMMDD(date))) / 1e6, 2) AS median_daily_dollars_mm
FROM global_markets.stocks_daily_aggs
WHERE ticker IN ('SPY', 'AAPL', 'KO', 'SJM', 'LANC')
AND date >= '2026-07-01'
AND date < '2026-09-26'
GROUP BY ticker
ORDER BY median_daily_shares DESCMedian session volume over that window ran from 43721494 shares a day at the top of the panel to 1296421 at the bottom. A 2 percent cap allows 874430 shares a day in the deepest name and 25928 in the thinnest, where the entire market turns over about $157.91 million on a median day. Compare your intended order size against that line before reading any return. Average daily volume and relative volume cover the volume baselines themselves.
Plot markout curves after your fills
A markout is the price change measured at a fixed interval after a trade, from the trade price. The panel below uses a stand in anyone can compute: treat every minute whose volume ran at least three times the name's average as a fill event, sign it by that minute's own direction, and follow the price forward.
| horizon_minutes | thin_markout_bps | liquid_markout_bps | observations |
|---|---|---|---|
| 1 | -0.9 | 0 | 1909 |
| 2 | -0.2 | 0.2 | 1786 |
| 5 | -2.1 | 0.3 | 1492 |
| 15 | 0.1 | -0.6 | 1102 |
| 30 | -2.7 | -0.4 | 974 |
| 60 | -1.7 | -0.4 | 869 |
The exact SQL behind every number
WITH bars AS
(
SELECT
ticker,
toStartOfMinute(toTimeZone(window_start, 'America/New_York')) AS et_minute,
toFloat64(open) AS open_px,
toFloat64(close) AS close_px,
toFloat64(volume) AS vol
FROM global_markets.delayed_stocks_minute_aggs
WHERE ticker IN ('SPY', 'AAPL', 'KO', 'SJM', 'LANC')
AND window_start >= '2026-07-01 00:00:00'
AND window_start < '2026-09-26 00:00:00'
AND (toHour(toTimeZone(window_start, 'America/New_York')) * 60
+ toMinute(toTimeZone(window_start, 'America/New_York'))) >= 570
AND (toHour(toTimeZone(window_start, 'America/New_York')) * 60
+ toMinute(toTimeZone(window_start, 'America/New_York'))) < 960
AND volume > 0
AND open > 0
),
ranked AS
(
SELECT
ticker,
avg(vol) AS avg_minute_vol
FROM bars
GROUP BY ticker
HAVING count() >= 500
),
picks AS
(
SELECT
argMax(ticker, avg_minute_vol) AS liquid_ticker,
argMin(ticker, avg_minute_vol) AS thin_ticker
FROM ranked
),
sided AS
(
SELECT
if(b.ticker = p.thin_ticker, 'thin', 'liquid') AS side,
b.et_minute AS et_minute,
b.open_px AS open_px,
b.close_px AS close_px,
b.vol AS vol
FROM bars AS b
CROSS JOIN picks AS p
WHERE b.ticker = p.thin_ticker
OR b.ticker = p.liquid_ticker
),
mean_vol AS
(
SELECT side, avg(vol) AS avg_vol
FROM sided
GROUP BY side
),
events AS
(
SELECT
s.side AS side,
s.et_minute AS et_minute,
s.close_px AS fill_px,
if(s.close_px >= s.open_px, 1, -1) AS direction
FROM sided AS s
INNER JOIN mean_vol AS m ON m.side = s.side
WHERE s.vol >= 3 * m.avg_vol
),
horizons AS
(
SELECT
side,
fill_px,
direction,
hz,
et_minute + toIntervalMinute(hz) AS target_minute
FROM
(
SELECT side, et_minute, fill_px, direction, arrayJoin([1, 2, 5, 15, 30, 60]) AS hz
FROM events
)
)
SELECT
h.hz AS horizon_minutes,
round(avgIf(h.direction * (f.close_px / h.fill_px - 1) * 10000, h.side = 'thin'), 1) AS thin_markout_bps,
round(avgIf(h.direction * (f.close_px / h.fill_px - 1) * 10000, h.side = 'liquid'), 1) AS liquid_markout_bps,
count() AS observations
FROM horizons AS h
INNER JOIN sided AS f ON f.side = h.side AND f.et_minute = h.target_minute
GROUP BY h.hz
HAVING countIf(h.side = 'thin') > 0
AND countIf(h.side = 'liquid') > 0
ORDER BY h.hzAt one minute out, the thin name's markout measured -0.9 basis points against 0 for the liquid name, across 1909 matched minutes. At the 60 minute mark the thin name reads -1.7 basis points. Values near zero describe a move that was temporary; values far from zero describe a price that kept travelling in the same direction over the horizon.
Run the same curve on your own backtest fills, signed by your own side. What you are looking for is asymmetry. If almost every simulated fill is followed by a move in your favour and almost none by a move against you, the fill model is handing you prices the market was about to leave.
Double the assumed impact and re-read the equity curve
Take whatever impact or slippage term the backtest used, multiply it by two, and run again. A result that holds its shape at twice the cost is a result about the strategy. A curve that flattens is a result about the assumption. Watch maximum drawdown alongside the total return, since doubled costs land unevenly across the sample.
Report the pessimistic model, not the tuned one
Pick the harsh setting of model 3 as the headline number and keep the touch model only as an upper bound on what was theoretically available. The wider process around this, including the splits and the cost accounting, is in how to backtest a trading strategy.
Method notes and caveats
- The quote panel covers a two hour midday window on one date and converts round lots to shares. Its spread is an average of quoted spreads, not an effective spread measured against trades.
- The minute move and markout panels use regular session minutes between July 1 and September 25, 2026, for one liquid and one thin listing.
- The two names are picked inside the query from a five name basket: the highest and the lowest average minute volume among the names with enough regular session bars in the window.
- Minute level moves are a proxy for the cost of size, not a measurement of your market impact. A real impact estimate needs your own fills.
- High volume minutes in the markout panel are defined as at least three times the name's own average minute volume over the window.
FAQ
What is a fill model in a backtest?
It is the rule that converts a signal into an execution price: which price you got, for how many shares, and what the trade cost. Common choices are crossing the spread at the quoted touch, walking the displayed depth level by level, and walking the depth with a cancellation and impact adjustment on top.
Why does an illiquid backtest look so different from live trading?
The backtest prices your order against a book recorded without you in it. In a thin name, one order can be a large share of the displayed size, so the recorded price is available for part of the order and not the rest. The difference shows up as slippage on every fill.
Does order book data fix market impact assumptions?
No. Message by message data lets you measure cancellation rates and rebuild the book at any instant, which makes the gap measurable. It still cannot contain your own order or other participants' responses to it, so the impact piece remains a model.
What participation rate is realistic in a thin name?
There is no single number that holds across names and horizons. The practical approach is to cap participation as a share of the name's median session volume, run the strategy at that cap, then halve the cap and compare. A result that only exists above the cap is a result about the cap.
What is a markout curve?
A plot of the average price change measured at several fixed intervals after a trade, starting from the trade price and signed by the direction of the trade. It is the standard way to check whether a fill price was plausible or whether the simulated trade took liquidity that was already leaving.
Every panel here carries the exact SQL beneath it, so each number can be rechecked or re-run over a different symbol and window. The same questions can be asked in plain English on the Strasmore terminal.