# Kronos

Open-weights transformer pre-trained on OHLCV bars. The training corpus is not released.

**Best for:** researchers who want pre-trained weights to fine-tune or benchmark against their own OHLCV bars; not for anyone who needs a data feed, a backtester, a broker connection or a model they can retrain from scratch.

*https://stockmarketstack.com/tools/kronos · Backtesting Frameworks & Algo Trading Libraries*

## Facts

### At a glance

| Field | Value |
| --- | --- |
| Vendor | Yu Shi |
| Category | Backtesting Frameworks & Algo Trading Libraries |
| Job | research |
| Website | https://github.com/shiyu-coder/Kronos |
| Pricing model | open-source |
| Free tier | true |
| Open source | true |
| Licence | MIT |
| Self-hosted | true |
| Tested hands-on | no |
| Claimed by vendor | no |
| Last updated | 2026-10-08 |

### Coverage

| Field | Value |
| --- | --- |
| Asset classes | stocks, etf, indices, futures, forex, crypto |
| Markets | us, ca, uk, eu, asia, latam, au, global |
| Works outside the US | true |
| Data latency | none |
| Platforms | library |
| AI features | research |

### Interfaces

| Field | Value |
| --- | --- |
| API | false |
| Webhooks | false |
| Scripting | Python |
| Python | true |
| Spreadsheet add-in | false |
| MCP server | false |
| Export | none |

### Capabilities

Yes: none

No: charting, screening, scanning, backtesting, automation, live_trading, paper_trading, portfolio_tracking, broker_import, tax_reporting, alerts, news, options_analysis

*Verified: pricing 2026-10-08; capabilities 2026-10-08; coverage 2026-10-08.*

## What it is

Kronos is a family of decoder-only transformers trained on candlestick bars rather than text. It
works in two stages. A tokenizer, a small transformer autoencoder, compresses each bar into a
discrete token split into a coarse and a fine sub-token of 10 bits each. The bar has six fields:
open, high, low, close, volume and amount (turnover). An autoregressive model is then trained to
generate the next tokens, which the tokenizer decodes back into bars. Forecasting the next bars
is the task the model was trained for. The paper also uses it for realised-volatility estimates
and for sampling synthetic bar sequences.

The released sizes are on Hugging Face under the NeoQuasar account:

- **Kronos-mini**, 4.1M parameters, with its own 2k tokenizer and a 2,048-bar context.
- **Kronos-small**, 24.7M parameters, 512-bar context.
- **Kronos-base**, 102.3M parameters, 512-bar context.

A 499.2M **Kronos-large** appears in the model table marked as not publicly available. All five
Hugging Face repositories, three models and two tokenizers, were last modified on 9 September 2025.

A clone gives you about 1,250 lines of model code and `KronosPredictor`, which z-scores each input
window, clips it at ±5 standard deviations, samples and de-normalises. Beside it sit fine-tuning
scripts for Qlib data and for a single CSV, a local web UI, a CPU regression test and contributed
Chinese-language scripts that pull A-share daily bars through AKShare.

The authors are at Tsinghua University. The paper, arXiv 2508.02739, has had one version since 2
August 2025, and the README says it was accepted at AAAI 2026.

## Pricing

Free, with nothing to buy. The code is MIT per the LICENSE file. Each Hugging Face model card is
tagged MIT and none is gated. The cost is your data, plus GPUs if you fine-tune.

## Data & coverage

Kronos ships no data. Apart from a test fixture, the only bars in the tree are a five-minute
series for Alibaba's Hong Kong listing, there for the CSV fine-tuning example.

The pre-training corpus is described only in the paper. Table 13 gives 12.11 billion bars across
96,569 instruments, at seven intervals from one minute to weekly. The abstract counts 45
exchanges, and the appendix says over 40 exchanges in more than 30 countries. The mix is
lopsided:

- Nasdaq and NYSE together give 4.6 billion bars back to 2000, at intervals down to one minute.
- Shanghai and Shenzhen give 3.7 billion bars back to 1990.
- Binance spot and perpetual pairs give 1.2 billion bars from 2021.

Those five venues make up 79% of the corpus. The rest is stocks and ETFs from about forty other
venues, mostly from January 2020 and often daily and weekly only. It also includes 75 Chinese
futures contracts, 1,023 forex pairs and stock indices from 28 markets.

The paper's cleaning step cut each series wherever the gap from one bar's close to the next bar's
open was large, as at splits, dividends and contract rolls. So the model trained on segments
without those jumps, and `KronosPredictor` makes no adjustment of its own. Whether the bars you
pass in are adjusted is your decision, and [why adjusted close differs](https://stockmarketstack.com/guides/why-adjusted-close-differs)
explains what changes.

The two sources do not agree on the cut-off. The paper says the pre-training data runs to June 2024
and testing starts in July 2024. In issues 14 and 40 the lead author wrote that training ends in
February 2024, validation runs from March to May and the out-of-sample period starts in June.

## Integrations

Python 3.10 or later and PyTorch 2.0 or later. Weights load through `huggingface_hub`'s
`from_pretrained`. The Qlib fine-tuning pipeline needs `pyqlib`, and its last step runs Qlib's own
top-k backtest on the fine-tuned model's output. Comet.ml logging is optional. The web UI runs on
Flask and Plotly at `localhost:7070`.

## Limitations

- The README has no evaluation of its own. The paper scores IC and RankIC on price and return
  paths, MAE and R² on realised volatility and the realism of sampled bars against 25 baselines,
  plus a top-k simulation on CSI 300 and CSI 800 inside Qlib. Its headline, a 93% RankIC gain over
  the strongest general time-series model, is the paper's own claim on its own benchmark. The test
  bars came from the same vendors, and the author says the evaluation code is tied to them and not
  released, so nobody outside can rerun it. Pre-training cannot be reproduced either; only
  fine-tuning can.
- The quick-start is broken. Pull request 139, merged on 13 April 2026, deleted
  `examples/data/XSHG_5min_600977.csv`, which the README and three example scripts read. The owner
  had asked for the file to be kept, and the README still points at it.
- The fine-tuning demo's config asks for CSI 300 daily bars through 5 June 2025, while the dataset
  Qlib's own download fetches ends in 2020 (see the [Qlib](https://stockmarketstack.com/tools/qlib) card). The README itself
  calls the pipeline "not a production-ready quantitative trading system".
- Until 13 April 2026 the Qlib fine-tuning dataset took its mean and standard deviation over the
  whole window, including the bars the model was being trained to generate. The owner confirmed
  the leak on 30 March 2026. A model fine-tuned from an older clone inherits it.
- Everything else is absent: no data loader for a live feed, no broker, no paper trading and no
  backtester of its own. `requirements.txt` pins pandas 2.2.2, which has no wheels past Python 3.12.
- The licences do not stack up to an answer. The code and the weights are MIT. The authors say the
  data behind them is held under vendor agreements that forbid passing it on, and nothing in the
  repository, the model cards or the paper says what those agreements allow for the weights. The
  MIT tag is the authors' choice, and no vendor statement is published beside it. [Can you train a
  model on licensed market data?](https://stockmarketstack.com/guides/training-models-on-market-data) treats open weights
  trained on licensed bars as a distribution event, which is the question this leaves open.

## Alternatives

[Qlib](https://stockmarketstack.com/tools/qlib) is what Kronos's own fine-tuning and backtest run on, and the place for
cross-sectional factor models trained on data you hold. [PyBroker](https://stockmarketstack.com/tools/pybroker) retrains a
model of your own walk-forward inside a backtest, where Kronos only produces bars.
[VectorBT](https://stockmarketstack.com/tools/vectorbt) and [Backtesting.py](https://stockmarketstack.com/tools/backtesting-py) are the engines that can
test anything you build from its output. [FinRL](https://stockmarketstack.com/tools/finrl) is reinforcement learning on gym
environments. [FinGPT](https://stockmarketstack.com/tools/fingpt) publishes language-model adapters trained on financial text
rather than on bars.

## FAQ

### Can I get the data Kronos was trained on?

No. In issue 100 on 19 September 2025 the lead author wrote that licensing agreements with financial data vendors prevent sharing the pre-training set, and named Wind, Tsanghi and Binance among the sources. The pre-training code is withheld for the same reason (issue 44). The paper's Table 13 lists the venues, asset counts and start dates, and that is all a reader gets.

### Is `pip install kronos` the right install?

No. That name on PyPI is a Django cron-scheduling app by another author, last released as 0.2.2 in December 2011. `kronos-finance` is a third party's wrapper, uploaded as 0.1.0 to 0.1.5 on 23 September 2026. The authors publish no package. The README's install is a git clone plus `pip install -r requirements.txt`, and the code is imported as `from model import Kronos`.

### Is Kronos still maintained?

Slowly. On 8 October 2026 the last commit to master was 13 April 2026, a batch of five merged community fixes. The repository has never tagged a release and is not archived. The owner's last reply on an issue was dated 30 March 2026. 213 issues and 69 pull requests were open, against about 40,300 stars.

### Do I need a GPU to run it?

Not for inference on the released sizes. `KronosPredictor` picks CUDA, then Apple's MPS, then the CPU, and the repository's regression tests run on the CPU. The weight files are 16 MB for mini, 99 MB for small and 409 MB for base. The fine-tuning scripts are written for multi-GPU `torchrun`, and the paper's own training ran on 24 RTX 4090D cards.

### What exactly does the model return?

A pandas DataFrame of open, high, low, close, volume and amount for the future timestamps you pass in. With `sample_count` above 1 it draws that many paths and returns their mean, so the spread between the paths is thrown away unless you change the code.

## Also worth comparing

- [FinRL](https://stockmarketstack.com/tools/finrl.md) — Gym-style market environments and deep-RL agents for trading research, MIT-licensed.
- [Qlib](https://stockmarketstack.com/tools/qlib.md) — Microsoft's ML factor-research pipeline. The data its CLI downloads stops in late 2020.
- [FinGPT](https://stockmarketstack.com/tools/fingpt.md) — Financial LLM adapters, instruction datasets and training notebooks, all MIT.
- [Backtesting.py](https://stockmarketstack.com/tools/backtesting-py.md) — A single-instrument Python backtester — one OHLC series, one strategy, no live trading.
- [Backtrader](https://stockmarketstack.com/tools/backtrader.md) — An event-driven Python backtester with 122 indicators, frozen since April 2023.
- [bt](https://stockmarketstack.com/tools/bt.md) — Python backtesting for allocation and rebalancing rules, not entries and exits.
