Compressing Market Data: zstd vs gzip
Compressing market data: zstd vs gzip benchmarked in a fresh container, why sorted tick files shrink so far, and when gzip is still the right pick.
Compressing market data is the cheapest storage win available to anyone keeping a tick archive, and in practice the decision comes down to zstd vs gzip. On sorted tick text, zstd at its top levels lands smaller than gzip -9 and still reads back faster, while zstd -3 gets close to gzip's ratio for a small fraction of the processing time. Nothing below is quoted from a vendor benchmark: the steps build a realistic test file inside a fresh Ubuntu container and measure both codecs on your own hardware.
Two words to fix first. A codec is the compression algorithm plus the settings you run it with. A ratio is the original size divided by the compressed size, so a bigger ratio means a smaller file on disk.
How much data one trading session produces
Before picking a codec it helps to see the raw volume. The panel below counts every trade printed for five household names on 15 September 2026, then converts each count into bytes at a flat 64 bytes per record, a stand-in for a timestamp, a symbol, a price and a size written as fixed-width text.
| ticker | trades_label | megabytes_at_64b_row |
|---|---|---|
| NVDA | 2.00 million | 121.9 |
| AAPL | 688.37 thousand | 42 |
| SPY | 495.56 thousand | 30.2 |
| MSFT | 399.64 thousand | 24.4 |
| KO | 236.18 thousand | 14.4 |
The exact SQL behind every number
SELECT
ticker,
formatReadableQuantity(count()) AS trades_label,
round(count() * 64 / 1048576, 1) AS megabytes_at_64b_row
FROM global_markets.stocks_trades
WHERE ticker IN ('AAPL', 'MSFT', 'NVDA', 'SPY', 'KO')
AND sip_timestamp >= toDateTime64('2026-09-15 04:00:00', 9, 'UTC')
AND sip_timestamp < toDateTime64('2026-09-16 04:00:00', 9, 'UTC')
GROUP BY ticker
ORDER BY megabytes_at_64b_row DESCNVDA printed 2.00 million trades in that single session, roughly 121.9 MB of flat text. The quietest of the 5 names still wrote 14.4 MB. Quote messages run far heavier than trades on the same symbols, and the options quote feed is the extreme version of the same arithmetic. Packet captures are heavier again with the wire headers attached, a layout covered in storing market data pcaps.
Why market data compresses so well
Compressors earn their ratio on repetition, and a tick file is close to a best case.
- Timestamps arrive sorted and close together. Consecutive nanosecond stamps share nearly every leading digit, and the repeated prefix gets stored once.
- The symbol column has tiny cardinality. One name repeats the same four bytes on every row of the file.
- Prices sit in a narrow band and move in single cents. A session revisits the same handful of price strings thousands of times.
- Sizes come from a short menu of round lots, with 100 far more common than anything else.
The price point is measurable. This panel counts the distinct trade prices each name printed, then divides the trade count by that number.
| ticker | distinct_prices | trades_per_distinct_price |
|---|---|---|
| NVDA | 22709 | 88 |
| KO | 5963 | 40 |
| AAPL | 19272 | 36 |
| SPY | 19530 | 25 |
| MSFT | 33101 | 12 |
The exact SQL behind every number
SELECT
ticker,
countDistinct(price) AS distinct_prices,
toUInt32(round(count() / countDistinct(price))) AS trades_per_distinct_price
FROM global_markets.stocks_trades
WHERE ticker IN ('AAPL', 'MSFT', 'NVDA', 'SPY', 'KO')
AND sip_timestamp >= toDateTime64('2026-09-15 04:00:00', 9, 'UTC')
AND sip_timestamp < toDateTime64('2026-09-16 04:00:00', 9, 'UTC')
GROUP BY ticker
ORDER BY trades_per_distinct_price DESCNVDA repeated each of its 22709 distinct prices about 88 times across the day. A compressor sees that as a small dictionary and a long stream of references back into it.
Repetition also clusters in time, which is what matters for a codec working through a sliding window rather than the whole file at once. The next panel splits one name's session into quarter-hour buckets on the New York clock.
| et_time | trade_count | distinct_prices |
|---|---|---|
| 04:00 | 6417 | 540 |
| 04:15 | 500 | 117 |
| 04:30 | 128 | 48 |
| 04:45 | 208 | 71 |
| 05:00 | 410 | 78 |
| 05:15 | 305 | 65 |
| 05:30 | 291 | 70 |
| 05:45 | 289 | 52 |
| 06:00 | 289 | 66 |
| 06:15 | 427 | 106 |
| 06:30 | 256 | 48 |
| 06:45 | 448 | 97 |
| 07:00 | 744 | 167 |
| 07:15 | 433 | 124 |
| 07:30 | 581 | 99 |
| 07:45 | 570 | 144 |
| 08:00 | 591 | 165 |
| 08:15 | 706 | 177 |
| 08:30 | 458 | 102 |
| 08:45 | 574 | 133 |
The exact SQL behind every number
SELECT
formatDateTime(toStartOfFifteenMinutes(toTimeZone(sip_timestamp, 'America/New_York')), '%H:%i') AS et_time,
count() AS trade_count,
countDistinct(price) AS distinct_prices
FROM global_markets.stocks_trades
WHERE ticker = 'AAPL'
AND sip_timestamp >= toDateTime64('2026-09-15 04:00:00', 9, 'UTC')
AND sip_timestamp < toDateTime64('2026-09-16 04:00:00', 9, 'UTC')
GROUP BY et_time
ORDER BY et_timeThe 64 buckets run from 04:00 through 19:45, and the distinct-price count inside each one stays small even where the trade count spikes. Local repetition of that kind is precisely what a window-based codec feeds on.
zstd vs gzip: run the benchmark yourself
The only numbers worth trusting here come off your own machine. As of September 2026 these steps run unchanged in a clean container.
- Start the container:
docker run --rm -it ubuntu:24.04 bash. - Install the tools:
apt-get update && apt-get install -y python3 zstd. - Write a synthetic tick file of eight million rows, sorted by a rising nanosecond clock, with repeating symbols and cent-level prices:
python3 -c "import random as r;r.seed(7);S=['AAPL','MSFT','NVDA','SPY','KO'];P=[232.15,418.40,118.72,566.03,70.11];f=open('ticks.csv','w');f.writelines('%d,%s,%.2f,%d\n'%(1758000000000000000+i*137+r.randrange(90),S[k],P[k]+r.randrange(-25,26)*0.01,r.choice([1,50,100,200,300])) for i in range(8000000) for k in [r.randrange(5)]);f.close()". It takes a minute or two and lands near 300 MB. - Record the starting size:
wc -c ticks.csv. - Compress with gzip at maximum effort:
time gzip -9 -k ticks.csv. - Time the zstd runs at rising effort:
time zstd -3 ticks.csv -o ticks-3.zst, thentime zstd -19 ticks.csv -o ticks-19.zst, thentime zstd -19 --long ticks.csv -o ticks-19-long.zst. - Put the sizes side by side:
ls -l ticks.csv ticks.csv.gz ticks-3.zst ticks-19.zst ticks-19-long.zst. - Time the read back, the number that governs a replay pipeline:
time gzip -dc ticks.csv.gz > /dev/nullandtime zstd -dc ticks-19.zst > /dev/null.
A few orderings hold across versions and machines, which is why this page states directions instead of pinning byte counts that a release note could invalidate. zstd -19 finishes smaller than gzip -9 on columnar tick text. Adding --long, which widens the match window to 128 MiB, finishes smaller still on a file this size. The read-back timing is the lopsided comparison: zstd decompresses several times quicker than gzip, whatever level it was written at. And zstd -3, the default, lands near gzip's ratio while spending a small share of gzip's compression time, which is what makes it a sane everyday setting. If you would rather sweep the levels in one command, zstd -b3 -e19 ticks.csv uses the built-in benchmark mode.
Layout beats codec choice
Now damage the file without altering a single row of its content.
- Shuffle it:
shuf ticks.csv > shuffled.csv, thenzstd -19 shuffled.csv -o shuffled-19.zst. - Group it by symbol and then by time:
sort -t, -k2,2 -k1,1n ticks.csv > bysymbol.csv, thenzstd -19 bysymbol.csv -o bysymbol-19.zst. - Compare all three:
ls -l ticks-19.zst shuffled-19.zst bysymbol-19.zst.
The three archives hold identical rows. The shuffled copy compresses worst, the symbol-grouped copy compresses best, and the spread between those two is wider on this file than the spread between gzip and zstd on the sorted original. Sorting by symbol turns the symbol column into long runs, and it pulls each name's narrow price band together in the window. Column stores such as Parquet push the same idea further by writing every column into its own block, where the values are already neighbours.
Pick the row order first. A desk that sorts its archive by symbol and time often gets more out of gzip -9 than a desk shipping shuffled rows into zstd -19.
Where gzip still wins
Compatibility is the whole case, and it is a strong one. gzip is the Content-Encoding every HTTP client on earth accepts without negotiation. Browsers began accepting zstd as a Content-Encoding in 2024 and support has broadened since, yet gzip remains the safe default for any public download endpoint. Vendors ship .csv.gz for the same boring reason: every language runtime, every spreadsheet importer and every decade-old internal script already opens one. Most of the endpoints surveyed in free stock market data APIs hand back gzip and nothing else.
zstd owns the other half of the map: archival storage and inline pipelines you control end to end. pandas reads a .csv.zst path with no extra arguments, pyarrow writes Parquet with the zstd codec, and current Wireshark releases open a .pcap.zst capture directly. The workflows in pcap capture and replay sit entirely on that side of the fence.
The storage arithmetic
Take the five names from the first panel, hold 21 sessions a month, and the uncompressed footprint grows in a straight line.
| horizon | gigabytes_uncompressed |
|---|---|
| 1-month | 4.8 |
| 3-month | 14.3 |
| 6-month | 28.7 |
| 12-month | 57.3 |
| 24-month | 114.6 |
| 60-month | 286.6 |
The exact SQL behind every number
WITH day_rows AS
(
SELECT count() AS trades
FROM global_markets.stocks_trades
WHERE ticker IN ('AAPL', 'MSFT', 'NVDA', 'SPY', 'KO')
AND sip_timestamp >= toDateTime64('2026-09-15 04:00:00', 9, 'UTC')
AND sip_timestamp < toDateTime64('2026-09-16 04:00:00', 9, 'UTC')
)
SELECT
concat(toString(months), '-month') AS horizon,
round(trades * 64 * 21 * months / 1073741824, 1) AS gigabytes_uncompressed
FROM day_rows
CROSS JOIN (SELECT arrayJoin([1, 3, 6, 12, 24, 60]) AS months) AS horizons
ORDER BY monthsA 12-month archive of those five names alone carries 57.3 GB of flat text, and stretching to 60-month reaches 286.6 GB. Whatever ratio your own run measured divides straight into those figures: a 4:1 ratio leaves a quarter of the bytes, a 6:1 ratio leaves a sixth. State the saving as a ratio rather than a dollar figure, since disk pricing drifts every year and the ratio does not. For most desks the larger line item is acquisition rather than disk, which what historical tick data costs breaks down.
FAQ
Is zstd better than gzip for market data?
For archival storage and pipelines you control, zstd wins on both size and decompression speed at comparable effort. gzip keeps one decisive advantage, universal support, including its role as the default HTTP Content-Encoding, which is why vendors still distribute .gz files.
Which zstd level fits tick data?
Level 3 is the default and lands near gzip's ratio for a fraction of the compression time, which suits files written often. Level 19 with --long is the archival setting, worth its slower pass on data written once and read many times.
Can pandas read zstd compressed files?
Yes. pd.read_csv('ticks.csv.zst') infers the codec from the file extension, and pyarrow writes Parquet with the zstd codec. Current Wireshark releases open .pcap.zst captures directly, as of September 2026.
Does compressing market data lose any precision?
No. gzip and zstd are both lossless, so the decompressed bytes are identical to the input, and zstd verifies a frame checksum on the way out. Precision is lost by rounding prices or downsampling before compression, never by the codec.
Why do market data vendors still ship gzip?
Every HTTP client, language runtime and legacy script already handles gzip, and the file stays readable on a machine nobody has touched in ten years. A handful of saved gigabytes matters less to a vendor than a support ticket from a client whose tooling cannot open the archive.
Every panel above ships with the exact SQL beneath it. To size your own archive before you compress it, ask for the trade counts on your symbols in plain English on the Strasmore terminal.