Strasmore Research
Learn Matt ConnorBy Matt Connor

Market Data Licensing for App Developers

Market data licensing for app developers, explained: entitlement, redistribution and display, what delayed feeds allow, and a pre-ship compliance checklist.

Market data licensing is the half of building a markets app that nobody budgets for. Getting the feed is the easy part. The contract around it decides what your product may display, to whom, and how fresh, and those three permissions come from different documents. Developers arrive asking for a free API and leave with a licensing problem.

What market data licensing actually covers

Pull the single question apart into the three that vendors and exchanges actually answer.

  1. Entitlement. May this account receive the feed, and at what subscriber status? Exchanges price per user, and the per-user rate turns on whether that user counts as professional. The classification is contractual rather than a judgement call, and it is laid out in professional vs non-professional market data.
  2. Redistribution. May you hand the data to someone who is not you? A dashboard that five colleagues open is one kind of contract. A public price chart with a sign-up form is another. Most vendor agreements name the second one redistribution and price it on its own line.
  3. Display. What has to sit next to the number, and how stale may the number be? Attribution, timestamps, delay notices, and the ban on presenting one venue's quote as 'the' price all live here, along with the display versus non-display split. The mechanics are in the vendor display rule.

A free API key answers none of the three. It answers a fourth question, whether bytes will arrive, which is the one question nobody gets stuck on. Our survey of free stock market data API options covers which tiers permit which of the three.

Does a delayed feed need a licence?

Delayed data, conventionally 15 minutes behind the live tape, sits under lighter terms at nearly every vendor, and for a large class of products it is the right answer. The question to settle first is what the delay hides. The panel below measures that directly: the average high-to-low range inside each 15-minute window of the session, for a broad index fund and a slow-moving consumer name, across June 2026.

QueryWhat a 15-minute delay can hide, by time of session (June 2026)
26 rows (showing 20)
et_timespy_range_pctko_range_pct
09:300.3720.958
09:450.3320.523
10:000.310.504
10:150.2790.38
10:300.2540.402
10:450.2910.346
11:000.2280.33
11:150.2410.249
11:300.2170.291
11:450.2210.253
12:000.20.206
12:150.170.245
12:300.1950.223
12:450.1630.23
13:000.1870.225
13:150.2120.192
13:300.1920.214
13:450.1730.178
14:000.1960.208
14:150.1650.215
The exact SQL behind every number
WITH buckets AS
(
    SELECT
        ticker,
        toStartOfInterval(window_start, INTERVAL 15 MINUTE) AS bucket,
        max(toFloat64(high))  AS hi,
        min(toFloat64(low))   AS lo,
        avg(toFloat64(close)) AS mid
    FROM global_markets.delayed_stocks_minute_aggs
    WHERE ticker IN ('SPY', 'KO')
      AND window_start >= '2026-06-01 00:00:00'
      AND window_start <  '2026-07-01 00:00:00'
      AND (toHour(toTimeZone(window_start, 'America/New_York')) * 60
           + toMinute(toTimeZone(window_start, 'America/New_York'))) >= 570
      AND (toHour(toTimeZone(window_start, 'America/New_York')) * 60
           + toMinute(toTimeZone(window_start, 'America/New_York'))) <  960
    GROUP BY ticker, bucket
)
SELECT
    formatDateTime(toTimeZone(bucket, 'America/New_York'), '%H:%i') AS et_time,
    round(avgIf((hi - lo) / mid * 100, ticker = 'SPY'), 3)          AS spy_range_pct,
    round(avgIf((hi - lo) / mid * 100, ticker = 'KO'), 3)           AS ko_range_pct
FROM buckets
GROUP BY et_time
HAVING countIf(ticker = 'SPY') > 0
   AND countIf(ticker = 'KO') > 0
ORDER BY et_time
Run this yourself

The 09:30 window averaged 0.372% high to low on the index fund and 0.958% on the consumer name. The 15:45 window averaged 0.349% and 0.415%. Read the curve as the width of the uncertainty a 15-minute-old quote carries at that hour. A screener, a research page, or a teaching product lives comfortably inside that band. An order ticket does not, and no licence tier changes that.

Who do you sign the agreement with?

Three realistic paths exist for a small product. A delayed feed under a vendor's standard terms. A real-time vendor feed with a redistribution rider. A direct agreement with each exchange. The third path is where the roster below starts to matter.

QueryUS equity venues grouped by operating MIC
operator_micvenue_countexample_venuevenue_acronyms
XNYS6NYSE American, LLC AMEX NSX
FINR5FINRA Alternative Display Facility
XCBO5Cboe EDGA
XNAS4Nasdaq Texas, Inc.
24EQ124X National Exchange LLC24X
IEXG1Investors Exchange
LTSE1Long-Term Stock Exchange
MIHI1MIAX Pearl
TXSE1Texas Stock Exchange LLCTXSE
XISX1International Securities Exchange, LLC - Stocks
XMEM1Members Exchange
The exact SQL behind every number
SELECT
    operating_mic                                         AS operator_mic,
    count()                                               AS venue_count,
    any(name)                                             AS example_venue,
    arrayStringConcat(arraySort(groupUniqArray(acronym)), ' ') AS venue_acronyms
FROM global_markets.stocks_exchanges
WHERE asset_class = 'stocks'
  AND operating_mic != ''
GROUP BY operating_mic
ORDER BY venue_count DESC, operator_mic
Run this yourself

The equity side of the reference data lists 11 distinct operating groups. The largest, XNYS, runs 6 venues under one operating MIC, the four-character market identifier code for a trading venue. The direct path means a separate agreement with each group in that list, plus the consolidated-tape paperwork, plus monthly per-user reporting, plus audit rights that are routinely exercised. The vendor path collapses all of it into one contract and one invoice, at the cost of flexibility and a markup. For what each path costs in practice, see what real-time market data costs, and for the question of going straight to the source, do NYSE and Nasdaq have public APIs.

Does caching market data count as redistribution?

Caching is where a compliant design quietly stops being one. Storing the feed is a licensed act under most agreements, and the volume is easy to underestimate. The panel below counts top-of-book quote messages for one large-cap name across one pinned session.

QueryTop-of-book quote messages for one name, one session (16 Jun 2026)
et_hourquote_message_countmessages_readableavg_spread
09:00115568115.57 thousand0.0391
10:00213956213.96 thousand0.0294
11:00183386183.39 thousand0.0251
12:00162777162.78 thousand0.0241
13:00115644115.64 thousand0.0211
14:008969289.69 thousand0.0191
15:00197483197.48 thousand0.0187
The exact SQL behind every number
SELECT
    formatDateTime(toStartOfHour(toTimeZone(sip_timestamp, 'America/New_York')), '%H:%i') AS et_hour,
    count()                                                    AS quote_message_count,
    formatReadableQuantity(count())                            AS messages_readable,
    round(avg(toFloat64(ask_price) - toFloat64(bid_price)), 4)  AS avg_spread
FROM global_markets.cache_stocks_quotes
WHERE ticker = 'AAPL'
  AND sip_timestamp >= '2026-06-16 13:30:00'
  AND sip_timestamp <  '2026-06-16 20:00:00'
  AND ask_price > 0
  AND bid_price > 0
GROUP BY et_hour
ORDER BY et_hour
Run this yourself

The 09:00 bucket, which covers only the 30 minutes from the open, carried 115.57 thousand quote messages at an average quoted spread of $0.0391. The 15:00 bucket carried 197.48 thousand. One name, one day, 7 hourly buckets.

Two working distinctions cover most cases. A cache that serves a user their own earlier request again is generally treated as that user's entitled consumption. A cache that serves user B the response generated for user A is redistribution, and it is also how a one-seat entitlement silently becomes a thousand-seat one. Snapshot caching at a stated refresh interval is commonly permitted, and the permitted interval is written in the agreement rather than chosen by the engineer. Persisting the raw messages above turns the product into a holder of a licensed historical dataset, which carries terms of its own, including deletion obligations when the contract ends.

Is a computed signal still the exchange's data?

This is the question product teams ask last and need answered first. The panel below holds no quote, no print, and no last price. It is the standard deviation of daily closing returns, annualised, one row per month.

QueryA derived value: monthly annualised realised volatility, SPY
monthrealized_vol_pctreturn_count
2025-096.2920
2025-1013.423
2025-1114.9419
2025-128.1622
2026-011020
2026-0213.0319
2026-0317.8222
2026-0411.4321
2026-059.5120
2026-0617.2621
2026-0711.8222
2026-0810.0521
The exact SQL behind every number
WITH daily AS
(
    SELECT
        date,
        toFloat64(close) AS px,
        lagInFrame(toFloat64(close)) OVER (ORDER BY date ASC ROWS BETWEEN 1 PRECEDING AND CURRENT ROW) AS prev_px
    FROM global_markets.stocks_daily_aggs
    WHERE ticker = 'SPY'
      AND date >= '2025-09-01'
      AND date <  '2026-09-01'
)
SELECT
    formatDateTime(toStartOfMonth(date), '%Y-%m')             AS month,
    round(stddevPop(px / prev_px - 1) * sqrt(252) * 100, 2)   AS realized_vol_pct,
    count()                                                   AS return_count
FROM daily
WHERE prev_px > 0
GROUP BY month
ORDER BY month
Run this yourself

The series measured 6.29% in 2025-09 and 10.05% in 2026-08, over 12 months and 20 daily returns in that first bucket. Vendor agreements handle output like this under a derived data clause, and the usual test is reversibility. If a user can recover a price, a quote, or a size from what you publish, the output is the licensed data wearing a different label, and every display and entitlement obligation follows it. A month of annualised volatility does not let anyone reconstruct the closes behind it. A field called 'last price rounded to the nearest dollar' plainly does. The number of inputs also matters: a statistic computed across hundreds of names is treated differently from a one-ticker transform. Derived-data permission is specific per vendor and worth getting in writing before the metric becomes your headline feature.

Why a free tier forbids the thing the product needs

Free tiers are written as evaluation licences. The common terms allow one developer, internal testing, attribution, no redistribution, no persistence beyond a session, and no commercial use. A shipping product needs many users, display to third parties, storage, and a commercial purpose. That gap is the business model. The paid tier exists to sell the use case the free tier prohibits, and no amount of careful engineering around the rate limit addresses it.

The order of work follows from the three questions. Entitlement shapes your user model and your sign-up form. Redistribution shapes your caching and API layer. Display shapes your UI components. Reading the terms of service before the architecture is cheaper than reading them after.

A pre-ship licensing checklist

  1. Write one sentence naming everyone who sees the data, and whether any of them sit outside your organisation. That sentence settles redistribution.
  2. Classify every user as professional or non-professional, and record how the question was asked. Per-user reporting is monthly on most agreements.
  3. Decide the freshness the product genuinely requires, then price the delayed path first and the real-time path second.
  4. Put the delay notice, the source attribution, and the data timestamp inside the component that shows the number, not in the page footer.
  5. Document your cache time-to-live and your retention window, then check both against the agreement's storage wording.
  6. List every derived value you publish alongside the inputs it uses, and get the derived-data clause confirmed for each.
  7. Keep an audit folder from day one: entitlement counts, reports filed, screenshots of the live display. Audit rights are standard.
How the panels above were bounded

The session panels group by New York clock time rather than assuming a 9:30 to 16:00 window, the quote panel is pinned to a single past session so the figures never move, and the volatility panel annualises with the conventional 252-session factor. Every panel carries row-count and value bounds, so an empty or wildly different result holds the page instead of publishing.

FAQ

Do I need a market data licence to show stock prices in my app?

If prices reach anyone outside your organisation, yes, in some form. The lightest realistic path is a vendor's delayed feed with redistribution permitted in writing, which is how most small consumer products ship. The licence is about who sees the data and how fresh it is, not about how clever the code is.

Is 15-minute delayed data free to redistribute?

Not automatically. Delayed data carries lighter obligations and lower or waived per-user fees at many vendors, and some exchange schedules treat delayed display as fee-free, but redistribution still has to be granted by your agreement. Free of charge and free of terms are different things.

Does caching market data break my licence?

Caching itself is usually permitted at a defined refresh interval. The breakage happens when one user's cached response is served to a different user, which converts entitled consumption into redistribution, and when raw messages are retained long enough to become a historical dataset under separate terms.

Is my own calculated indicator still licensed data?

It depends on reversibility. If a reader can work back from your output to a price, a quote, or a size, the output is still the exchange's data and the display rules follow it. Aggregates over many names and statistics that discard the underlying levels are the cases vendors most often clear under a derived-data clause.

Can I prototype on a free API and license properly later?

Prototyping under an evaluation licence is what it is for. The thing that bites is architecture: a design built on a free tier's rate limits often caches and reshares in ways a paid redistribution agreement then has to be bent around. Checking the three questions early keeps that rework small.


Every panel on this page ships with the exact SQL beneath it, so the freshness numbers behind a licensing decision can be checked rather than assumed. The same questions can be asked in plain English on the Strasmore terminal.

#market data#licensing#redistribution#apis#compliance