Strasmore Research
Learn Matt ConnorBy Matt Connor

REST Polling vs WebSockets for Market Data

REST polling vs WebSockets for market data: the message rate math your consumer must keep up with, and the recovery path when a socket drops mid session.

REST polling vs WebSockets for market data is a question about the shape of the data, not about which protocol is newer. Polling means your application asks a server for the current state on a timer you control. A WebSocket means your application opens one long lived connection, subscribes to a list of symbols, and the server sends each update the moment it exists. Anything with a natural period fits polling. Anything event shaped fits a stream.

Most integration guides stop at the request metering. What breaks real stream integrations sits further in: the message rate your consumer has to absorb, the order in which you load a snapshot against a live subscription, the recovery path after the socket drops, and the stale marking in between.

REST polling vs WebSockets: how each one behaves

A REST request is one question and one answer. You ask for a quote, you get the quote as it stood when the server built the response, and the connection closes. Poll every second and you get one sample per second per symbol. The part worth internalizing: polling latency is bounded below by your interval. A one second poll can never show you a quote that lived 40 milliseconds. That quote existed and was replaced between two of your requests, and no amount of connection tuning recovers it.

A WebSocket inverts the flow. You connect once, send a subscribe message, then read. There is no per update request, which removes both the per update cost and the per update round trip. What you take on instead is a consumer that must keep pace with whatever the publisher sends, and a connection that will eventually drop.

Request metering is the other half of the polling cost. Free tiers tend to meter requests per minute, paid tiers per second, and every published figure goes stale quickly, so read a rate limit as a shape rather than a number. The arithmetic is the durable part. Five hundred symbols polled once a second is five hundred requests a second unless the endpoint accepts a batched symbol list, and even batched, the data is one sample per interval and up to a full interval old. Level 1 and Level 2 market data sit on opposite sides of this. A top of book snapshot compresses into one small response. A full depth book does not.

How many messages per second does a market data stream send?

Before picking a transport, count. The panel below takes a half hour of regular trading on September 15, 2026 and counts quote messages per second for four household names.

QueryQuote messages per second, four names, 10:30 to 11:00 a.m. ET
symbolmessages_per_secondbusiest_second_messages
SPY132.31356
NVDA73.71018
AAPL35.8427
KO18.8317
The exact SQL behind every number
SELECT
    symbol,
    round(sum(messages) / 1800.0, 1) AS messages_per_second,
    max(messages)                    AS busiest_second_messages
FROM
(
    SELECT
        ticker                    AS symbol,
        toDateTime(sip_timestamp) AS second_bucket,
        count()                   AS messages
    FROM global_markets.cache_stocks_quotes
    WHERE ticker IN ('AAPL', 'NVDA', 'SPY', 'KO')
      AND sip_timestamp >= '2026-09-15 14:30:00'
      AND sip_timestamp <  '2026-09-15 15:00:00'
    GROUP BY symbol, second_bucket
)
GROUP BY symbol
ORDER BY messages_per_second DESC
Run this yourself

Over that window, SPY averaged 132.3 quote messages a second, and its busiest single second carried 1356. At the other end of the four, KO averaged 18.8 a second. Same market, same half hour, and a spread wide enough to change the design.

Now run the message rate arithmetic for your own subscription. Multiply a per symbol rate by your symbol count, then size the consumer for the busiest second rather than the average. The average is where your buffer sits. The peak is where it overflows. A parser that handles twenty thousand messages a second is comfortable on ten quiet names and underwater on fifty active ones. If the plan involves an options chain, read how big the options quote feed gets before sizing anything; the equity numbers above are the small case.

What a once per second poll actually misses

A poll and a stream do not sample the same market. Group a busy name's quote messages into one second clock buckets and the distribution comes out lopsided.

QueryNVDA quote updates per second, grouped by how busy the second was
updates_in_the_secondshare_of_seconds_pctshare_of_messages_pct
1 update0.20
2 to 51.70.1
6 to 208.11.3
21 to 10062.442.9
over 10027.655.8
The exact SQL behind every number
WITH
    per_second AS
    (
        SELECT
            toDateTime(sip_timestamp) AS second_bucket,
            count()                   AS messages
        FROM global_markets.cache_stocks_quotes
        WHERE ticker = 'NVDA'
          AND sip_timestamp >= '2026-09-15 13:30:00'
          AND sip_timestamp <  '2026-09-15 15:30:00'
        GROUP BY second_bucket
    ),
    totals AS
    (
        SELECT
            count()       AS second_count,
            sum(messages) AS message_count
        FROM per_second
    )
SELECT
    multiIf(messages = 1,    '1 update',
            messages <= 5,   '2 to 5',
            messages <= 20,  '6 to 20',
            messages <= 100, '21 to 100',
            'over 100')                                                  AS updates_in_the_second,
    round(100.0 * count() / (SELECT second_count FROM totals), 1)         AS share_of_seconds_pct,
    round(100.0 * sum(messages) / (SELECT message_count FROM totals), 1) AS share_of_messages_pct
FROM per_second
GROUP BY updates_in_the_second
ORDER BY min(messages)
Run this yourself

In two hours of quotes, the lightest bucket, seconds carrying 1 update, made up 0.2% of the seconds and 0% of the messages. The heaviest bucket, over 100, held 27.6% of the seconds while carrying 55.8% of the messages. A once per second poll returns exactly one sample from each of those seconds, the light ones and the heavy ones alike.

Zoom into a single minute and the cost gets concrete. Below is the opening minute of AAPL on the same date, second by second.

QueryAAPL opening minute: quote updates and distinct prices, second by second
60 rows (showing 20)
et_secondupdate_countdistinct_bid_pricesdistinct_ask_prices
09:30:0017697889
09:30:019054851
09:30:023442942
09:30:034802125
09:30:043743742
09:30:054434139
09:30:062022223
09:30:071061110
09:30:081933329
09:30:091252615
09:30:102983435
09:30:111461214
09:30:121961227
09:30:132131519
09:30:142881817
09:30:153332524
09:30:163182219
09:30:173741320
09:30:181661512
09:30:191032111
The exact SQL behind every number
SELECT
    formatDateTime(toTimeZone(toDateTime(sip_timestamp), 'America/New_York'), '%H:%i:%S') AS et_second,
    count()              AS update_count,
    uniqExact(bid_price) AS distinct_bid_prices,
    uniqExact(ask_price) AS distinct_ask_prices
FROM global_markets.cache_stocks_quotes
WHERE ticker = 'AAPL'
  AND sip_timestamp >= '2026-09-15 13:30:00'
  AND sip_timestamp <  '2026-09-15 13:31:00'
GROUP BY et_second
ORDER BY et_second
Run this yourself

60 of the 60 seconds in that minute carried at least one quote update. The second beginning 09:30:00 printed 1769 updates across 78 distinct bid prices and 89 distinct ask prices. A poller that fired once inside that second saw one of them, with no way to know which one.

Subscribe first, then load the snapshot

A stream carries changes. It does not carry state. You need both, and the order you acquire them in decides whether your book has a hole in it. The steps have to happen in this order:

  1. Open the socket and subscribe. Do not process anything yet.
  2. Buffer every message that arrives, with its vendor timestamp and sequence number intact.
  3. Request the REST snapshot for the same symbols.
  4. Replay the buffer against that snapshot, dropping messages older than it and applying the rest in sequence.
  5. Switch to live processing.

Snapshot first and subscribe second leaves a gap between the snapshot's instant and the subscription's first message, and nothing in the protocol tells you how wide that gap was. Subscribing first makes the overlap harmless instead: duplicates are cheap to discard, missing updates are not. The ordering key is the vendor timestamp rather than your own clock, and the two differ in ways worth understanding before you write the dedupe logic. Market data timestamps covers the fields involved.

When the socket drops, assume nothing was saved

Connections drop. The failure mode that matters is the quiet one, where the socket stays open and the messages simply stop.

Detect the gap with whatever the feed gives you: a heartbeat that fails to arrive on its interval, a sequence number that jumps, a vendor timestamp that stops advancing while your own clock keeps moving, or a message counter that flatlines. A watchdog firing after a few multiples of the normal inter message time catches all of them.

The recovery path is the cold start path again: reconnect, resubscribe, buffer, resnapshot, replay. Do not assume the stream backfills what you missed. Most market data feeds publish live state and hold no concept of your position in the sequence, so a reconnect hands you the market as it stands now, with the gap left behind. Mark your cached data stale from the moment the watchdog fires until the resnapshot lands; a downstream reader must never treat a frozen price as a live one. Make your writes idempotent on vendor timestamp and sequence number, since a replayed buffer will hand you the same update twice.

The decision rule, without a pricing page

Poll anything with a natural period:

  • bars of any interval, which are published on a fixed clock
  • reference data such as dividends, splits, exchange holiday calendars and ticker metadata
  • closing prices, settlement values and quarterly fundamentals
  • anything you chart at a resolution coarser than your poll interval

Stream anything event shaped:

  • trades, which arrive when they arrive
  • top of book and depth of book updates
  • auction imbalances and opening indications
  • any alert whose value decays within seconds

The giveaway is whether the data has an inherent timestamp grid. A minute bar for 09:35 does not exist until 09:36 and never changes afterwards, so a timer that asks for it wastes nothing. A trade has no grid at all. The same divide shows up in API design, which is part of why a trading API and a market data API are built along different lines.

Trade activity across a session is the clearest picture of event shape.

QueryAAPL trades per second, fifteen minute buckets through the session
26 rows (showing 20)
et_timetrades_per_second
09:3065.2
09:4535.9
10:0042.9
10:1538.9
10:3043.3
10:4539.7
11:0038.2
11:1536
11:3034
11:4533.9
12:0033
12:1518.1
12:3018.4
12:4517.6
13:0013.4
13:1511.7
13:3013.1
13:4515.2
14:0014.8
14:1514.5
The exact SQL behind every number
SELECT
    formatDateTime(toStartOfFifteenMinutes(toTimeZone(window_start, 'America/New_York')), '%H:%i') AS et_time,
    round(sum(transactions) / 900.0, 1)                                                           AS trades_per_second
FROM global_markets.delayed_stocks_minute_aggs
WHERE ticker = 'AAPL'
  AND window_start >= '2026-09-15 13:30:00'
  AND window_start <  '2026-09-15 20:00:00'
GROUP BY et_time
ORDER BY et_time
Run this yourself

The first quarter hour of that session ran at 65.2 trades a second. The quarter hour beginning 12:45 ran at 17.6, and the final bucket, 15:45, at 71.3. A fixed interval poll takes the same number of samples in every one of those buckets, which is the mismatch in one line.

FAQ

Is a WebSocket always faster than REST polling?

For event shaped data it wins in practice, since there is no interval to wait out. For a dataset that updates once a minute, both arrive at about the same time and the poll is simpler to operate. Publishing cadence sets the floor on latency; the transport only decides how close you get to that floor.

How often can I poll a market data API?

Whatever your plan's request metering allows, and those limits change often enough that any figure quoted here would be wrong by the time you read it. Read the limit as requests per unit of time, divide by your symbol count, and check whether what remains still fits your interval. A batched symbol endpoint relaxes the constraint considerably.

Do I still need REST if I use a WebSocket?

Yes. The stream carries changes and the snapshot carries state, so you need a REST snapshot at startup and after every reconnect. Even a free stock market data API tier is often enough for that snapshot half of the job.

What happens to the messages sent while my socket was down?

Treat them as lost unless the vendor documents a replay or recovery channel. The standard assumption is that a reconnect resumes at the live edge, which is why the recovery path ends in a fresh snapshot rather than a backfill request.

How do I tell whether my stream is falling behind?

Compare the vendor timestamp on the newest message against your own clock on a timer. A lag that grows steadily means your consumer is slower than the publisher, and the fix sits downstream of the socket: fewer subscriptions, or a queue between the socket reader and the processing code.


Every panel on this page ships with the SQL that produced it, so the counting method is one click away. The same message rate questions are answerable for any symbol and window you care about on the Strasmore terminal.

#market data#api#websockets#streaming#developers