When Historical Market Data Gets Revised
Historical market data gets revised: how a daylight saving bug dropped after-hours bars, the four revision types, and the record that keeps results checkable.
Historical market data gets revised more often than most backtests assume. A vendor repairs a gap, changes a timestamp convention, reassigns a ticker, or reprocesses a corporate action, and the archive on your disk stops matching today's download. This post sorts those revisions into four families, says what each one can change in a strategy result and what it leaves alone, and sets out the short record that keeps an old number checkable a year later.
The case: a winter bug that dropped after-hours rows
Start with a defect of a shape that turns up in dataset change notices. An intraday dataset converts exchange time to UTC with a fixed offset instead of a real time zone. Through the summer the offset matches. Through the winter months it is wrong by an hour, and the hour it drops sits at the end of the extended session. Rows that printed on the tape were never written into the dataset, on winter dates, across the dataset's entire history.
Two properties of that defect keep it hidden. The missing rows sit outside the regular session, where a thin chart looks normal anyway. And every row that is present is accurate, so a row-level validator finds nothing wrong. The defect lives in the rows that are absent, which is the hardest thing for a data check to see. Our walk-through of how OHLCV bars are built covers the aggregation step where a missing minute vanishes into a wider bar without leaving a mark.
The first panel measures the exposure: where a trading day's volume actually sits, hour by hour.
| et_hour | minute_bars | pct_of_shares |
|---|---|---|
| 04:00 | 629 | 0.11 |
| 05:00 | 490 | 0.08 |
| 06:00 | 631 | 0.13 |
| 07:00 | 835 | 0.22 |
| 08:00 | 1069 | 1.44 |
| 09:00 | 1181 | 17.65 |
| 10:00 | 1200 | 15.49 |
| 11:00 | 1200 | 11.31 |
| 12:00 | 1200 | 8.7 |
| 13:00 | 1200 | 7.62 |
| 14:00 | 1200 | 9.67 |
| 15:00 | 1200 | 20.51 |
| 16:00 | 1030 | 5.57 |
| 17:00 | 771 | 1.11 |
| 18:00 | 741 | 0.29 |
| 19:00 | 591 | 0.1 |
The exact SQL behind every number
WITH tape AS
(
SELECT
toTimeZone(window_start, 'America/New_York') AS et,
volume
FROM global_markets.delayed_stocks_minute_aggs
WHERE ticker = 'AAPL'
AND window_start >= '2026-01-01 05:00:00'
AND window_start < '2026-02-01 05:00:00'
)
SELECT
formatDateTime(et, '%H:00') AS et_hour,
count() AS minute_bars,
round(100 * sum(volume) / (SELECT sum(volume) FROM tape), 2) AS pct_of_shares
FROM tape
GROUP BY et_hour
ORDER BY et_hourAcross January 2026, AAPL minute bars fall into 16 distinct Eastern clock hours. The earliest hour the tape carries, the 4 a.m. ET hour, holds 0.11% of the month's shares while printing a near-full complement of one-minute rows. That asymmetry is the trap. Extended-hours activity is a small fraction of shares and a large fraction of rows, so a bug that removes those rows barely moves a volume total and badly damages anything counted per bar.
A seasonal defect has a signature, and the way to look for one is to measure the same statistic month by month. A dataset quietly losing an hour of the late session through the winter prints a step down in November and a step back up in March.
| month | month_label | outside_rth_volume_pct | outside_rth_bars_pct |
|---|---|---|---|
| 2024-10-01 | Oct 2024 | 9.13 | 45.94 |
| 2024-11-01 | Nov 2024 | 8.36 | 44.7 |
| 2024-12-01 | Dec 2024 | 10.35 | 43.55 |
| 2025-01-01 | Jan 2025 | 8.45 | 50.26 |
| 2025-02-01 | Feb 2025 | 9.13 | 46.89 |
| 2025-03-01 | Mar 2025 | 6.22 | 47.01 |
| 2025-04-01 | Apr 2025 | 6.62 | 53.5 |
| 2025-05-01 | May 2025 | 9.19 | 51.97 |
| 2025-06-01 | Jun 2025 | 5.77 | 51.12 |
| 2025-07-01 | Jul 2025 | 8.94 | 49.05 |
| 2025-08-01 | Aug 2025 | 6.17 | 47.09 |
| 2025-09-01 | Sep 2025 | 6.7 | 48.46 |
| 2025-10-01 | Oct 2025 | 7.24 | 47.22 |
| 2025-11-01 | Nov 2025 | 8.42 | 44.36 |
| 2025-12-01 | Dec 2025 | 9 | 39.95 |
| 2026-01-01 | Jan 2026 | 9.56 | 48.58 |
| 2026-02-01 | Feb 2026 | 6.56 | 45.23 |
| 2026-03-01 | Mar 2026 | 7.31 | 42.05 |
| 2026-04-01 | Apr 2026 | 9.75 | 48.58 |
| 2026-05-01 | May 2026 | 7.7 | 53.54 |
The exact SQL behind every number
WITH tape AS
(
SELECT
toStartOfMonth(toTimeZone(window_start, 'America/New_York')) AS m,
toHour(toTimeZone(window_start, 'America/New_York')) * 60
+ toMinute(toTimeZone(window_start, 'America/New_York')) AS et_minute,
volume
FROM global_markets.delayed_stocks_minute_aggs
WHERE ticker = 'AAPL'
AND window_start >= '2024-10-01 04:00:00'
AND window_start < '2026-10-01 04:00:00'
)
SELECT
toString(m) AS month,
formatDateTime(m, '%b %Y') AS month_label,
round(100 * sumIf(volume, et_minute < 570 OR et_minute >= 960) / sum(volume), 2) AS outside_rth_volume_pct,
round(100 * countIf(et_minute < 570 OR et_minute >= 960) / count(), 2) AS outside_rth_bars_pct
FROM tape
GROUP BY m, month_label
ORDER BY mMeasured on AAPL over the two years to September 2026, the outside-session share of volume ran 9.13% in Oct 2024 and 6.38% in Sep 2026. The row share is the far larger number: 54.01% of that final month's minute bars fell outside 9:30 a.m. to 4:00 p.m. ET. A flat series like this one is what a repaired dataset looks like. A broken one carries a November notch in both columns, repeating every year.
One caveat about this case, and it is the method of the post rather than an aside. Nothing above names a vendor, a dataset or a notice date. Vendor-specific claims age badly: the fix ships, the notice is superseded, and a page that asserted the broken behaviour as a standing fact is then simply wrong. When you write a defect into your own notes, pin it to the dated change notice you read it in, with the notice identifier and the dataset version it covers, and date the sentence yourself. A later reader can then work out what has since changed. The mechanics below hold. An inventory of live defects does not.
A fix is itself a breaking change
When the vendor repairs that conversion and backfills, the whole history of the dataset changes. Winter-date volume totals rise. The last print of the day moves later in the session. Any statistic normalised by session volume or counted per bar shifts. A number computed on last year's extract and the same number computed today disagree, and neither is a mistake. They are two datasets sharing one name.
That is the part worth holding onto. A correction is not a continuation of the series you had, it is a new version of it. Give a vendor fix the ceremony of a schema migration: re-run the prior result, store the old and new figures side by side, and record which dataset version produced each.
Four ways historical market data gets revised
Revisions sort into four families. The useful question for each one is narrow. Which results can it flip, and which does it leave untouched?
Gap repair and backfill
Missing rows arrive, or bad rows are withdrawn. It can flip anything counted per observation, anything averaged over a window that now holds more observations, and any gap-detection logic calibrated on the old holes. It leaves a daily close-to-close series alone, as long as the repair was intraday. Backfill carries a subtler hazard too. Rows delivered later than the timestamps they carry sit in your history at a moment when nobody could have seen them, which is the mechanism behind look-ahead bias in backtesting.
Timestamp and snapshot convention changes
A vendor moves from exchange-local time to UTC, or from bar-open to bar-close labelling, or shifts a daily snapshot from the closing print to a settlement cutoff. Every value is unchanged. Every value now sits at a different index. It can flip event studies, any join of two sources on a timestamp, and intraday rules anchored to a clock. It leaves distribution statistics over a long window alone. The panel below shows why a fixed offset fails: it walks the sessions around the November 2025 clock change and prints, for each, the UTC and Eastern stamp of that session's first bar.
| session_date | session_label | first_bar_utc | first_bar_et | et_utc_offset_hours |
|---|---|---|---|---|
| 2025-10-27 | Oct 27 | 08:00 | 04:00 | -4 |
| 2025-10-28 | Oct 28 | 08:00 | 04:00 | -4 |
| 2025-10-29 | Oct 29 | 08:00 | 04:00 | -4 |
| 2025-10-30 | Oct 30 | 08:00 | 04:00 | -4 |
| 2025-10-31 | Oct 31 | 08:00 | 04:00 | -4 |
| 2025-11-03 | Nov 3 | 09:00 | 04:00 | -5 |
| 2025-11-04 | Nov 4 | 09:00 | 04:00 | -5 |
| 2025-11-05 | Nov 5 | 09:00 | 04:00 | -5 |
| 2025-11-06 | Nov 6 | 09:00 | 04:00 | -5 |
| 2025-11-07 | Nov 7 | 09:00 | 04:00 | -5 |
The exact SQL behind every number
SELECT
toString(toDate(toTimeZone(window_start, 'America/New_York'))) AS session_date,
formatDateTime(toDate(toTimeZone(window_start, 'America/New_York')), '%b %e') AS session_label,
formatDateTime(min(window_start), '%H:%i') AS first_bar_utc,
formatDateTime(toTimeZone(min(window_start), 'America/New_York'), '%H:%i') AS first_bar_et,
toInt16(toHour(toTimeZone(min(window_start), 'America/New_York')))
- toInt16(toHour(min(window_start))) AS et_utc_offset_hours
FROM global_markets.delayed_stocks_minute_aggs
WHERE ticker = 'AAPL'
AND window_start >= '2025-10-27 04:00:00'
AND window_start < '2025-11-08 05:00:00'
GROUP BY session_date, session_label
ORDER BY session_dateThe offset column reads -4 on Oct 27 and -5 on Nov 7, with the step between them landing on the first Sunday in November. A query that hardcodes a UTC window for the session loses an hour on one side of that step, every year, in both directions over a long history. Group by Eastern clock time and the question never arises. Market data timestamps covers the conventions in full.
Identifier reassignment
A ticker is a lease, not a name. A delisted symbol returns to the exchange and can be handed to an unrelated issuer, and most datasets key on the symbol. It can flip any ticker-keyed join, any sector mapping, and any long history stitched together by symbol. Survivorship bias work is the most exposed of all, since the dead names are the whole point of the exercise. It leaves single-name research conducted entirely inside one entity's listing window alone. The IPO record shows how often a symbol changes hands.
| listing_year | symbols_reassigned | issuers_per_symbol |
|---|---|---|
| 2018 | 2 | 2 |
| 2019 | 1 | 2 |
| 2020 | 4 | 2 |
| 2021 | 4 | 2 |
| 2022 | 1 | 2 |
| 2023 | 2 | 2 |
| 2024 | 5 | 2 |
| 2025 | 26 | 2.04 |
| 2026 | 22 | 2.09 |
The exact SQL behind every number
WITH symbol_issuers AS
(
SELECT
ticker,
issuer_name,
min(listing_date) AS first_listing
FROM global_markets.stocks_ipos
WHERE listing_date >= '2010-01-01'
AND listing_date < '2026-10-01'
AND ticker != ''
AND ticker NOT IN ('SPCX')
GROUP BY ticker, issuer_name
),
reused AS
(
SELECT
ticker,
count() AS issuer_count,
max(first_listing) AS latest_listing
FROM symbol_issuers
GROUP BY ticker
HAVING issuer_count > 1
)
SELECT
toString(toYear(latest_listing)) AS listing_year,
count() AS symbols_reassigned,
round(avg(issuer_count), 2) AS issuers_per_symbol
FROM reused
GROUP BY listing_year
ORDER BY listing_year9 listing years since 2010 contain at least one symbol that a second issuer went on to carry. The most recent year in the panel, 2026, holds 22 of them, averaging 2.09 issuers per symbol. Key on a permanent identifier, carry the listing window alongside it, and read why ticker symbols break datasets for the join patterns that survive this.
Corporate-action reprocessing
Adjustment factors get recomputed, a special distribution is reclassified, a split is applied across a longer history than before. The entire price level moves. It can flip every level-sensitive rule: a moving-average crossover, a stop distance quoted in dollars, a dollar-volume liquidity filter. It leaves a return series computed from properly adjusted prices alone, which is invariant to the adjustment basis. Corporate actions arrive continuously, so the reprocessing is continuous as well.
| month | month_label | splits_executed | reverse_splits |
|---|---|---|---|
| 2024-10-01 | Oct 2024 | 133 | 85 |
| 2024-11-01 | Nov 2024 | 144 | 115 |
| 2024-12-01 | Dec 2024 | 95 | 58 |
| 2025-01-01 | Jan 2025 | 96 | 75 |
| 2025-02-01 | Feb 2025 | 129 | 103 |
| 2025-03-01 | Mar 2025 | 117 | 75 |
| 2025-04-01 | Apr 2025 | 110 | 82 |
| 2025-05-01 | May 2025 | 123 | 81 |
| 2025-06-01 | Jun 2025 | 136 | 96 |
| 2025-07-01 | Jul 2025 | 91 | 68 |
| 2025-08-01 | Aug 2025 | 117 | 76 |
| 2025-09-01 | Sep 2025 | 150 | 95 |
| 2025-10-01 | Oct 2025 | 120 | 88 |
| 2025-11-01 | Nov 2025 | 100 | 69 |
| 2025-12-01 | Dec 2025 | 178 | 130 |
| 2026-01-01 | Jan 2026 | 90 | 69 |
| 2026-02-01 | Feb 2026 | 114 | 89 |
| 2026-03-01 | Mar 2026 | 191 | 135 |
| 2026-04-01 | Apr 2026 | 129 | 97 |
| 2026-05-01 | May 2026 | 142 | 102 |
The exact SQL behind every number
SELECT
toString(toStartOfMonth(execution_date)) AS month,
formatDateTime(toStartOfMonth(execution_date), '%b %Y') AS month_label,
countDistinct(id) AS splits_executed,
countDistinctIf(id, toFloat64(split_to) / toFloat64(split_from) < 1) AS reverse_splits
FROM global_markets.stocks_splits
WHERE execution_date >= '2024-10-01'
AND execution_date < '2026-10-01'
GROUP BY month, month_label
ORDER BY month165 splits took effect in Sep 2026, 107 of them reverse splits, and the two-year series shows the cadence holding month after month. Each entry rewrites the adjusted history of one symbol for every date before its execution date. Split-adjusted price history works through a single name in detail.
How to record a result so it stays checkable
Reproducibility here is a bookkeeping problem rather than a research problem. Keep the following next to every published number, in the same file as the number:
- The vendor and the dataset name, at the granularity the vendor versions.
- The venue or feed the extract covers, and whether it is consolidated or single-venue.
- The schema or dataset version string the vendor publishes, copied verbatim.
- The retrieval timestamp in UTC, kept separate from the date range of the data itself.
- A checksum of the extract you actually loaded:
sha256sumover the file, stored with the result. - The row count and the exact first and last timestamp the extract spans.
- The query or API call, verbatim.
The checksum is the field that does the real work. A version string tells you what the vendor intended to ship. A hash of your file tells you what you read, and it catches a silent re-issue under an unchanged version label.
Then one habit on top of the log. After any vendor update, re-run the prior result instead of assuming continuity. If the number holds, you have evidence. If it moves, you know the size of the move before anyone asks, and you have a dated pair of figures to explain it with.
FAQ
Does historical market data change after it is published?
Yes, routinely. Vendors backfill missing rows, correct time conversions, re-run corporate-action adjustments and re-key identifiers, and most of those changes apply to the whole history rather than only the front edge. An archive downloaded a year ago is a specific version of the data, not the data.
How do I tell whether a vendor has revised a dataset?
Compare rather than assume. Keep a checksum and a row count for each extract, re-pull a fixed historical window on a schedule, and diff it against what you stored. Release notes and dataset version strings are the first signal. Your own diff is the one that is checkable.
Which backtests are most sensitive to a data revision?
Anything counted per observation or anchored to a clock: intraday rules, event studies around a timestamp, per-bar statistics, dollar-level filters. Return-based studies over long windows are the most robust, since adjustment changes largely cancel out in returns.
Do I need to rerun an old backtest after a correction?
Rerunning is the only way to know what moved. A workable bar: re-run any result you still cite publicly, store the new figure alongside the old one with both dataset versions named, and keep the old extract rather than overwriting it.
Data notes and how these panels are pinned
- The intraday panels read one symbol, AAPL, from the delayed consolidated minute tape. Volume and bar counts cover the full extended window the feed carries, not the regular session alone.
- January 2026 and the sessions around 2 November 2025 are fixed historical windows, chosen to keep the clock-change step visible in exactly the same place on every regeneration.
- Extended-hours minutes are classified by Eastern clock time rather than by a hardcoded UTC range, which is the convention the timestamp section argues for.
- The IPO panel counts a symbol as reassigned when the listing record carries more than one issuer name for it. One symbol that vendor feeds attach to two different issuers is excluded, which keeps a double-tagged name out of the count.
- No vendor, dataset or release date appears anywhere in this post. Pin those to the dated notice you read, in your own notes, and date the sentence.
Every panel here carries the SQL that produced it in an expander underneath. To measure the outside-session share of a symbol across a particular month, or to count how often a ticker has changed hands, ask the question in plain English on the Strasmore terminal.