# Can you train a model on licensed market data?

Training is not a word most data licences use. How exchange policies and vendor terms map it onto non-display use, derived data and redistribution.

*https://stockmarketstack.com/guides/training-models-on-market-data · background to Backtesting Frameworks & Algo Trading Libraries*

**Answer:** Often for private research, rarely by default for a product. Training on stored history mostly sits outside exchange non-display fees, which are written around real-time data, and inside your vendor's contract, which decides. Free and personal tiers are non-commercial, and at least one business contract now forbids training, fine-tuning or distilling a model unless the order form permits it. Publishing the training set is redistribution; weights or outputs that can substitute for the data are a derived-data question.

The machine-learning quant stack — [Qlib](https://stockmarketstack.com/tools/qlib), [FinRL](https://stockmarketstack.com/tools/finrl), time-series
foundation models trained on OHLCV bars, an LLM fine-tuned on headlines — arrives with an
open-source licence on every repository and a data downloader in every tutorial. The licence on the repository covers the code.
Nothing in it covers the prices, and the prices come with a contract written for a person looking at
a screen.

**Training is not a word most market data licences use.** The documents were written around display
and real-time automated use, so the question is answered by mapping it onto categories that already
exist: who consumes the data, where it ends up, and what you make from it. This page is that mapping,
and where the documents are silent it says so.

## How it works

Rights move down a chain and narrow at each step. An exchange licenses a distributor under its own
data policies; the distributor — an API vendor, a broker, a terminal — writes its own terms for you;
and you can never hold more than the narrowest link grants. That is the shape
[who owns market data](https://stockmarketstack.com/guides/who-owns-market-data) maps from the top, and
[what real-time market data costs](https://stockmarketstack.com/guides/real-time-market-data-fees) prices. Training a model
touches four questions inside it, and they are independent:

- **Who consumes it.** A program rather than a person is
  [non-display use](https://stockmarketstack.com/glossary/non-display-use), priced per firm rather than per screen.
- **Where it goes.** Anyone outside your organisation receiving the data is
  [redistribution](https://stockmarketstack.com/glossary/redistribution) — and a published dataset is the plainest form of it.
- **What you make from it.** A number computed from exchange data is
  [derived data](https://stockmarketstack.com/glossary/derived-data) when it can neither recreate the input nor substitute for
  it. Weights and model outputs are the new candidates.
- **Who you are.** The [professional subscriber](https://stockmarketstack.com/glossary/professional-subscriber) test, which a
  company of one fails, and the academic terms, which a lab funded by a bank fails.

The London Stock Exchange's guidance on AI states the mapping outright:
[we license the what, not the how](https://docs.londonstockexchange.com/sites/default/files/documents/lse_-_artificial_intelligence_advisory_notification.pdf),
and "Derived Data creation, use of Real Time Data in trading and non-trading applications, and
Redistribution are all licensable use cases". Three exchanges' published AI positions are compared
in [the MCP guide](https://stockmarketstack.com/guides/mcp-and-financial-data), which is about a model reading data
at inference time. This page is about the model being built from it.

## Training on history is mostly outside the exchange fee

The expensive exchange category is written around live data, and a training run on stored files is
usually not live. Three venues draw that line in their own words.

**NYSE** opens its non-display policy by scoping it: the policies and fees
[apply to the Non-Display Use of real-time NYSE proprietary market information](https://www.nyse.com/publicdocs/nyse/data/NYSE_Market_Data_Complete_Policy_Package.pdf).
A separate policy in the same package covers history, and it is the sentence a training pipeline
depends on: vendors "may store and use at a later time real-time proprietary NYSE Market Information
within their firm or organization for internal purposes". Later means after midnight following the
day of dissemination.

**Nasdaq** counts non-display devices as the greater of the users who can modify the application in
real time or the servers that "receive and benefit from the information", and the second includes
servers "that run computations or creates derived data". GPUs linked to a server that is already
counted are not counted again. Then, among the examples of what is not fee liable, it lists
[storage of information that is not used until after the applicable delay interval](https://www.nasdaqtrader.com/content/AdministrationSupport/Policy/USEquitiesandOptionsDataPolicies.pdf).

**OPRA**, the options tape, defines data as historical "upon the opening of trading on the next
succeeding trading day", and exempts a vendor whose redistribution is limited to historical data
from its redistribution fee entirely —
[why options data costs more](https://stockmarketstack.com/guides/why-options-data-costs-more) has the whole delayed and
historical ladder.

So the exchange-side reading is reasonably consistent: a model trained on yesterday's data inside
your organisation is not the case the non-display schedules are pricing. Two things bring it back.
The first is deployment — the moment the trained model scores live data, it is non-display use, and
both NYSE's and OPRA's example lists name "investment analysis" alongside trading and risk
management. The second is that none of this is a licence to you. It is what the exchange charges your
vendor, and your vendor's contract can be narrower.

## The vendor's contract is where training is decided

This is where the answer actually lives, and the documents differ more than the exchange policies
do.

**Personal tiers are non-commercial, and "commercial" is defined broadly.**
[Alpha Vantage](https://stockmarketstack.com/tools/alpha-vantage)'s terms grant use "for personal, non-commercial use" and then
list what counts as commercial: any purpose going beyond "investment analysis, research, testing,
monitoring, and any other activities that are private and individual in nature"; using it "as or on
behalf of a corporation, firm, partnership, trust or any other association"; any commercial activity
that lets others reach the information "directly or indirectly"; or being employed by or affiliated
with an investment adviser, investment bank or broker-dealer. Research by an individual is inside
the grant. The same model trained inside a two-person startup is not. Nothing in those terms mentions machine learning or
says what happens to stored data at termination; they are silent, not permissive.
[Massive](https://stockmarketstack.com/tools/massive)'s individuals terms are the same shape — use "solely for your own personal,
non-commercial, and non-business purposes" — and do not mention machine learning either.

**Business terms are starting to name training.** Massive's business terms, last updated 6 October
2026, are the clearest example found for this page. The grant is generous about outputs:
you may create "derivative works, analyses, calculations, models, and other data or works from the
Information" and distribute them for any lawful business purpose, "provided that they do not contain
Information". The restrictions then carve training out of it. The customer will not "use the
Information to train, fine-tune, or distill any machine-learning or artificial-intelligence model
unless an Order Form expressly permits it, or permit any AI tool or other Non-Massive Service to use
the Information for its own purposes, including to train or improve its models". A separate clause
requires a licence, in an order form or other written agreement and from each upstream provider, to
use the information to create an index, a benchmark, an investment product "or investment
strategy". And at the end of a subscription, every right to use the
information ceases and the customer "must delete all related Information in its possession or
control, including downloaded files".

**Exchanges now say the same thing to their distributors.** Cboe's advisory requires that "use of
Data to train or fine-tune or operate an AI Solution must strictly adhere to Cboe Data licensing
terms", and — the sentence with paperwork attached — that any such use
["must be explicitly described in a Data Order Form and System Description"](https://cdn.cboe.com/resources/market_data/2026/Guidance-Regarding-Use-of-Cboe-Data-in-AI-Solutions.pdf)
and expressly accepted by Cboe in writing. The advisory reaches past the firm holding the licence,
too: AI use inside a recipient's own products "may require additional licensing, such as
non-display", and those products include ones that pass data into AI solutions on behalf of their
users. That obligation sits with your vendor, and it reaches you as a clause in your vendor's
terms.

The pattern is the one to expect. A contract written before the question existed is silent; a
contract revised since tends to name training and route it through an order form. Either way, the
paragraph that answers it is in the terms, never on the pricing page.

## Free data and the dataset you assembled

Most tutorials start from [yfinance](https://stockmarketstack.com/tools/yfinance). FinRL's own tutorial downloads its Dow
tickers from Yahoo, and Qlib's collector scripts build the US daily bars the same way. The
repositories are MIT or Apache-2.0. That licence is the author's grant over what the author wrote; it cannot license the prices the script fetches, which
belong to the terms of the site they came from.

[Yahoo's terms](https://legal.yahoo.com/us/en/yahoo/terms/otos/index.html), last updated 4 August
2026, are short on this point and do not leave much room. You agree not to "access or collect data,
or attempt to access or collect data, from our Services using any automated means" without express
prior permission; not to use any of its data "to create any database, archive, mobile application,
data feed, widget or any other aggregated data source that competes with or constitutes a material
substitute for the Services"; and, unless otherwise stated, not to "access or reuse the Services, or
any portion thereof, for any commercial purpose". A training corpus of daily bars for the S&P 500 is
an archive. Whether it is a substitute is the open question, and the commercial line is not.

The same logic runs through a dataset someone else published. A Hugging Face dataset of OHLCV bars,
or a parquet file in a model repository, carries the licence its uploader chose — and the uploader's
right to choose it is exactly what the upstream contract decides. [FinGPT](https://stockmarketstack.com/tools/fingpt) is the
useful contrast: what it publishes is labelled financial text and LoRA adapters rather than a price
history, and its forecaster pulls prices at run time rather than shipping them.

## Weights, outputs and the substitute test

Whether a trained model is derived data is not answered in terms by any exchange or vendor document
reviewed for this page. What the documents share is a test, and it is the right one to apply before
anyone else does.

Nasdaq's definition of derived data turns on whether the output can be "reverse engineered to
recreate" the exchange information or "be used to create other data that is recognizable as a
reasonable substitute" for it — the [derived data](https://stockmarketstack.com/glossary/derived-data) page reads it closely.
Cboe applies the same test to models directly: distribution of "any output of an AI Solution that can be used as a
substitute for the Data" needs a licence agreement first. Massive's definition of the information it
licenses reaches data derived from it that can recalculate or substitute for it "regardless of form,
delay, aggregation, reformatting, or combination with other data".

Read together, the risk is graded by what the model emits rather than by what kind of model it is:

- **A model that reproduces prices** — a forecaster emitting a price path per symbol, a generative
  model that can be sampled back into bars, an LLM that recites a closing series — is closest to a
  substitute, and publishing its weights or its outputs is closest to redistribution.
- **A model that emits a label, a score or a ranking** is further from the line, which is the case
  Massive's grant describes when it permits distributing models and analyses "provided that they do
  not contain Information".
- **Releasing open weights is a distribution event.** A time-series foundation model trained on
  licensed exchange data and published openly is distributing whatever the weights contain, to
  everyone, permanently. That is a conversation to have with the vendor before the upload, not
  after. [Kronos](https://stockmarketstack.com/tools/kronos) is the concrete case: its weights are MIT-licensed on Hugging Face,
  while its lead author says licensing agreements with the data vendors, Wind, Tsanghi and Binance
  among them, prevent releasing the training set.

## Who you are changes the answer

The [professional subscriber](https://stockmarketstack.com/glossary/professional-subscriber) test applies to a research account
the same way it applies to a trading screen, and an account in a company's name fails it. Academic
pricing exists and is narrower than it sounds. NYSE's academic policy for historical products is
limited to accredited institutions using the data for "independent academic research, academic
journals and other publications, teaching, or other educational purposes"; it excludes commercial
use; it will not be granted to anyone "whose research or use is funded by a securities industry
participant"; and where two institutions co-author, "each of the academic institutions would need to
license separately". A paper's replication package that ships the data falls outside every line of
that.

## What it costs

Training on stored files internally rarely adds an exchange fee of its own, and that is the cheap end.
The cost appears when the model leaves the research environment. On the OPRA fee schedule dated
February 2026, non-display use is 2,000 dollars a month per enterprise in each of Category 1 (own
behalf) and Category 2 (on behalf of clients), and the categories stack. The Category 1 fee is waived
in a month where the recipient is a single natural person, is not a broker-dealer and averages no
more than 390 listed options orders a day — a carve-out written for one person and one script.
Equities non-display and redistribution figures per venue are on the
[non-display use](https://stockmarketstack.com/glossary/non-display-use) and [redistribution](https://stockmarketstack.com/glossary/redistribution) pages,
and the full stack is in [what real-time market data costs](https://stockmarketstack.com/guides/real-time-market-data-fees).

The vendor's side is priced by quote. A personal tier is cheap because it excludes exactly the uses
this page is about; the order form that permits training is a sales conversation, and the price is
not published by any vendor reviewed here.

## What you can do about it

**Answer the four questions before choosing a provider, not after the first model works.** Who
consumes the data (person or program), where it goes (inside or outside), what you produce (labels or
prices), and who you are (individual or entity). Each one has a different document and a different
price, and a vendor can only quote you once you know them.

**Search the terms, not the pricing page.** Look for "train", "machine learning", "artificial
intelligence", "derivative", "store", "cache", "delete" and "commercial". A clause naming training
answers the question; silence means the contract predates it, and the useful move is to ask in
writing and keep the reply.

**Get training into the order form.** Massive's restriction lifts where "an Order Form expressly
permits it", and Cboe requires the use to be described in an order form and system description. That
is the mechanism vendors are converging on, so describe the training run, the hardware it runs on and
what will be published, and get it countersigned.

**Train on history, deploy on purpose.** Stored data used after the delay interval sits outside
the non-display policies NYSE and Nasdaq publish; live scoring sits inside them. Keep the two
environments separate and declare the second when it starts.

**Publish the recipe, never the data.** A collector script, a symbol list and a date range let
anyone rebuild your dataset under their own licence; a parquet file in a repository is redistribution
under yours. Qlib's US data path already works this way: you run the collector yourself.

**Keep lineage from the first run.** Record which dataset, from which vendor, under which terms,
trained which checkpoint. A termination clause that requires deletion, or a vendor asking what was
built from its data, is answerable in an afternoon with that record and not at all without it.

**Use data whose licence already answers the question.** Government and central-bank series, SEC
filings and your own fills carry no exchange contract at all;
[what is genuinely free](https://stockmarketstack.com/guides/free-financial-data-sources) covers which ones may be republished.
For licensed history, a vendor that itemises exchange fees, such as [Databento](https://stockmarketstack.com/tools/databento), or
files bought outright from [Cboe DataShop](https://stockmarketstack.com/tools/cboe-datashop), at least makes the licence a
document you can read. The frameworks themselves are under
[backtesting frameworks](https://stockmarketstack.com/categories/backtesting-frameworks); none of them changes any of the above.

## Tools this bears on

- [Qlib](https://stockmarketstack.com/tools/qlib.md) — Microsoft's ML factor-research pipeline. The data its CLI downloads stops in late 2020.
- [FinRL](https://stockmarketstack.com/tools/finrl.md) — Gym-style market environments and deep-RL agents for trading research, MIT-licensed.
- [Massive](https://stockmarketstack.com/tools/massive.md) — Full-tick US equities, options and futures from a direct exchange feed.
- [Alpha Vantage](https://stockmarketstack.com/tools/alpha-vantage.md) — The API most people's first script talks to — free key, wide coverage, hard rate limits.
- [yfinance](https://stockmarketstack.com/tools/yfinance.md) — The Python library everyone starts with — free, unofficial, and not something to build on.
- [Kronos](https://stockmarketstack.com/tools/kronos.md) — Open-weights transformer pre-trained on OHLCV bars. The training corpus is not released.

## FAQ

### Do I need a non-display licence to train a model on historical data?

Usually not from the exchange, because the non-display schedules are written around real-time data. NYSE's policy says so in its first sentence, and Nasdaq lists storage of data not used until after the delay interval among its non-display exclusions. The vendor's contract is a separate question, and it is the one that decides whether you may train at all.

### Can I publish my training dataset on GitHub or Hugging Face?

Not under an ordinary data subscription. Publishing the file is redistribution, and every vendor contract reviewed here restricts it, even where an exchange waives its own fee for historical data, as OPRA does. NYSE's policy lets a recipient store real-time data and use it later internally, but redistributing it externally at a later time needs a specific licence from NYSE.

### Are model weights derived data?

No exchange document reviewed here says so in terms. The test the documents do share is whether what you produce can recreate the original data or stand in for it. Cboe applies that test to any output of an AI solution, so a model that reproduces a price series is the risky case and one that emits a label is not the same case.

### Is it fine to train on data downloaded with yfinance?

For a private experiment it is what most tutorials do; for anything else the answer is Yahoo's. Its terms forbid collecting data by automated means without express prior permission, building a database or archive that substitutes for its services, and reusing them for any commercial purpose. The library's Apache-2.0 licence covers its code, not the data it returns.

### What happens to a trained model if I cancel the data subscription?

Read the termination clause. Massive's business terms require deleting all related information, downloaded files included, when a subscription ends, and keep the use restrictions alive after it. Whether weights count as that information is not answered there, which is why recording which data trained which model is worth doing before anybody asks.

## Sources

1. [Nasdaq US Equities and Options Data Policies, version 2.6](https://www.nasdaqtrader.com/content/AdministrationSupport/Policy/USEquitiesandOptionsDataPolicies.pdf) — Nasdaq, 2024-06-01
2. [NYSE Market Data Policy Package](https://www.nyse.com/publicdocs/nyse/data/NYSE_Market_Data_Complete_Policy_Package.pdf) — New York Stock Exchange, read 2026-10-07
3. [Options Price Reporting Authority Fee Schedule](https://cdn.opraplan.com/documents/OPRA_Fee_Schedule.pdf) — Options Price Reporting Authority, 2026-02-26
4. [Advisory Notice Regarding Use of Cboe Data in Artificial Intelligence (Reference ID C2026013000)](https://cdn.cboe.com/resources/market_data/2026/Guidance-Regarding-Use-of-Cboe-Data-in-AI-Solutions.pdf) — Cboe Global Markets, read 2026-10-07
5. [Use of London Stock Exchange Data in Artificial Intelligence - Guidance](https://docs.londonstockexchange.com/sites/default/files/documents/lse_-_artificial_intelligence_advisory_notification.pdf) — London Stock Exchange, read 2026-10-07
6. [Massive for Businesses Terms of Service](https://massive.com/legal/businesses-terms-of-service) — Massive.com, Inc., 2026-10-06
7. [Massive for Individuals Terms of Service](https://massive.com/legal/individuals-terms-of-service) — Massive.com, Inc., 2025-07-18
8. [Terms of Service](https://www.alphavantage.co/terms_of_service/) — Alpha Vantage Inc., read 2026-10-07
9. [Yahoo Terms of Service](https://legal.yahoo.com/us/en/yahoo/terms/otos/index.html) — Yahoo, 2026-08-04

*Last updated 2026-10-07. A reference page, corrected in place — not a dated post.*
