Kronos

Open-weights transformer pre-trained on OHLCV bars. The training corpus is not released.

Best for researchers who want pre-trained weights to fine-tune or benchmark against their own OHLCV bars; not for anyone who needs a data feed, a backtester, a broker connection or a model they can retrain from scratch.

by Yu Shi

Last updated

From
Free
Licence
MIT
Self-hosted
Yes
Platforms
Library

What it is

Kronos is a family of decoder-only transformers trained on candlestick bars rather than text. It works in two stages. A tokenizer, a small transformer autoencoder, compresses each bar into a discrete token split into a coarse and a fine sub-token of 10 bits each. The bar has six fields: open, high, low, close, volume and amount (turnover). An autoregressive model is then trained to generate the next tokens, which the tokenizer decodes back into bars. Forecasting the next bars is the task the model was trained for. The paper also uses it for realised-volatility estimates and for sampling synthetic bar sequences.

The released sizes are on Hugging Face under the NeoQuasar account:

  • Kronos-mini, 4.1M parameters, with its own 2k tokenizer and a 2,048-bar context.
  • Kronos-small, 24.7M parameters, 512-bar context.
  • Kronos-base, 102.3M parameters, 512-bar context.

A 499.2M Kronos-large appears in the model table marked as not publicly available. All five Hugging Face repositories, three models and two tokenizers, were last modified on 9 September 2025.

A clone gives you about 1,250 lines of model code and KronosPredictor, which z-scores each input window, clips it at ±5 standard deviations, samples and de-normalises. Beside it sit fine-tuning scripts for Qlib data and for a single CSV, a local web UI, a CPU regression test and contributed Chinese-language scripts that pull A-share daily bars through AKShare.

The authors are at Tsinghua University. The paper, arXiv 2508.02739, has had one version since 2 August 2025, and the README says it was accepted at AAAI 2026.

Pricing

Free, with nothing to buy. The code is MIT per the LICENSE file. Each Hugging Face model card is tagged MIT and none is gated. The cost is your data, plus GPUs if you fine-tune.

Data & coverage

Kronos ships no data. Apart from a test fixture, the only bars in the tree are a five-minute series for Alibaba's Hong Kong listing, there for the CSV fine-tuning example.

The pre-training corpus is described only in the paper. Table 13 gives 12.11 billion bars across 96,569 instruments, at seven intervals from one minute to weekly. The abstract counts 45 exchanges, and the appendix says over 40 exchanges in more than 30 countries. The mix is lopsided:

  • Nasdaq and NYSE together give 4.6 billion bars back to 2000, at intervals down to one minute.
  • Shanghai and Shenzhen give 3.7 billion bars back to 1990.
  • Binance spot and perpetual pairs give 1.2 billion bars from 2021.

Those five venues make up 79% of the corpus. The rest is stocks and ETFs from about forty other venues, mostly from January 2020 and often daily and weekly only. It also includes 75 Chinese futures contracts, 1,023 forex pairs and stock indices from 28 markets.

The paper's cleaning step cut each series wherever the gap from one bar's close to the next bar's open was large, as at splits, dividends and contract rolls. So the model trained on segments without those jumps, and KronosPredictor makes no adjustment of its own. Whether the bars you pass in are adjusted is your decision, and why adjusted close differs explains what changes.

The two sources do not agree on the cut-off. The paper says the pre-training data runs to June 2024 and testing starts in July 2024. In issues 14 and 40 the lead author wrote that training ends in February 2024, validation runs from March to May and the out-of-sample period starts in June.

Integrations

Python 3.10 or later and PyTorch 2.0 or later. Weights load through huggingface_hub's from_pretrained. The Qlib fine-tuning pipeline needs pyqlib, and its last step runs Qlib's own top-k backtest on the fine-tuned model's output. Comet.ml logging is optional. The web UI runs on Flask and Plotly at localhost:7070.

Limitations

  • The README has no evaluation of its own. The paper scores IC and RankIC on price and return paths, MAE and R² on realised volatility and the realism of sampled bars against 25 baselines, plus a top-k simulation on CSI 300 and CSI 800 inside Qlib. Its headline, a 93% RankIC gain over the strongest general time-series model, is the paper's own claim on its own benchmark. The test bars came from the same vendors, and the author says the evaluation code is tied to them and not released, so nobody outside can rerun it. Pre-training cannot be reproduced either; only fine-tuning can.
  • The quick-start is broken. Pull request 139, merged on 13 April 2026, deleted examples/data/XSHG_5min_600977.csv, which the README and three example scripts read. The owner had asked for the file to be kept, and the README still points at it.
  • The fine-tuning demo's config asks for CSI 300 daily bars through 5 June 2025, while the dataset Qlib's own download fetches ends in 2020 (see the Qlib card). The README itself calls the pipeline "not a production-ready quantitative trading system".
  • Until 13 April 2026 the Qlib fine-tuning dataset took its mean and standard deviation over the whole window, including the bars the model was being trained to generate. The owner confirmed the leak on 30 March 2026. A model fine-tuned from an older clone inherits it.
  • Everything else is absent: no data loader for a live feed, no broker, no paper trading and no backtester of its own. requirements.txt pins pandas 2.2.2, which has no wheels past Python 3.12.
  • The licences do not stack up to an answer. The code and the weights are MIT. The authors say the data behind them is held under vendor agreements that forbid passing it on, and nothing in the repository, the model cards or the paper says what those agreements allow for the weights. The MIT tag is the authors' choice, and no vendor statement is published beside it. Can you train a model on licensed market data? treats open weights trained on licensed bars as a distribution event, which is the question this leaves open.

Alternatives

Qlib is what Kronos's own fine-tuning and backtest run on, and the place for cross-sectional factor models trained on data you hold. PyBroker retrains a model of your own walk-forward inside a backtest, where Kronos only produces bars. VectorBT and Backtesting.py are the engines that can test anything you build from its output. FinRL is reinforcement learning on gym environments. FinGPT publishes language-model adapters trained on financial text rather than on bars.

Specs

Interfaces
Python, Python
Export
None
Asset classes
Stocks, ETF, Indices, Futures, Forex, Crypto
Markets
US, CA, UK, EU, Asia, Latam, AU, Global
Platforms
Library
AI features
Research
Pricing verified
Capabilities verified
Coverage verified

Also worth comparing

  • FinRL — Gym-style market environments and deep-RL agents for trading research, MIT-licensed.
  • Qlib — Microsoft's ML factor-research pipeline. The data its CLI downloads stops in late 2020.
  • FinGPT — Financial LLM adapters, instruction datasets and training notebooks, all MIT.
  • Backtesting.py — A single-instrument Python backtester — one OHLC series, one strategy, no live trading.
  • Backtrader — An event-driven Python backtester with 122 indicators, frozen since April 2023.
  • bt — Python backtesting for allocation and rebalancing rules, not entries and exits.

FAQ

Can I get the data Kronos was trained on?

No. In issue 100 on 19 September 2025 the lead author wrote that licensing agreements with financial data vendors prevent sharing the pre-training set, and named Wind, Tsanghi and Binance among the sources. The pre-training code is withheld for the same reason (issue 44). The paper's Table 13 lists the venues, asset counts and start dates, and that is all a reader gets.

Is pip install kronos the right install?

No. That name on PyPI is a Django cron-scheduling app by another author, last released as 0.2.2 in December 2011. kronos-finance is a third party's wrapper, uploaded as 0.1.0 to 0.1.5 on 23 September 2026. The authors publish no package. The README's install is a git clone plus pip install -r requirements.txt, and the code is imported as from model import Kronos.

Is Kronos still maintained?

Slowly. On 8 October 2026 the last commit to master was 13 April 2026, a batch of five merged community fixes. The repository has never tagged a release and is not archived. The owner's last reply on an issue was dated 30 March 2026. 213 issues and 69 pull requests were open, against about 40,300 stars.

Do I need a GPU to run it?

Not for inference on the released sizes. KronosPredictor picks CUDA, then Apple's MPS, then the CPU, and the repository's regression tests run on the CPU. The weight files are 16 MB for mini, 99 MB for small and 409 MB for base. The fine-tuning scripts are written for multi-GPU torchrun, and the paper's own training ran on 24 RTX 4090D cards.

What exactly does the model return?

A pandas DataFrame of open, high, low, close, volume and amount for the future timestamps you pass in. With sample_count above 1 it draws that many paths and returns their mean, so the spread between the paths is thrown away unless you change the code.

Report an error on this page

Quote the line and say what it should be.

The vendor’s page, filing or documentation that says otherwise.

Only if you want to hear back. Never published.

Read by a person. It fixes a fact and moves nothing else.

Is this your product? Claim this listing.