# Why two retirement calculators give different success rates

Same savings, same spending, two success rates. Replayed history or Monte Carlo, which return and inflation series, fees, timing, and what counts as success.

*https://stockmarketstack.com/guides/why-retirement-calculators-disagree · background to Retirement & FIRE Calculators*

**Answer:** Because a success rate describes the model as much as the plan. Historical-cycle calculators replay overlapping windows of past returns; Monte Carlo calculators generate paths from a distribution or by resampling. Each also picks its own stock, bond and inflation series, start year, fee treatment and timing convention, and defines success its own way. Two percentages are comparable only after those choices match, and a success rate is not a probability either way.

## How it works

Every success rate in this category is built the same way: the number of paths that met a rule,
divided by the number of paths the calculator ran. Your inputs — savings, spending, allocation, years —
are only one of the three things that decide that fraction. The other two are the calculator's, and
it rarely prints them next to the result.

**Where the paths come from.** Either the past, replayed in overlapping windows, or a generator
that produces new sequences of returns. The generator can draw from a statistical distribution or
reshuffle history.

**What the paths are made of.** A return series for each asset, an inflation series, a start year
and an end year, a treatment of fees and taxes, and a convention for the day in each year when
money is withdrawn, earned and rebalanced.

**What the rule is.** Ending above zero, never touching zero, or ending at zero or above after
touching it.

Change any of the three and the same plan gets a different percentage with nothing about it having
changed. That is the whole of the disagreement, and each part of it can be read off a calculator's
documentation or its export if you know which question to ask. The sections below take them in that
order.

## Replayed history or generated paths

A **historical-cycle** calculator takes the plan's length and runs it against every window of that
length in its data. The method is older than the software. The 1998 study by Cooley, Hubbard and
Walz in the AAII Journal, the one usually called the Trinity study, computed outcomes for the first
15-year period, 1926 to 1940, then the second, 1927 to 1941, and so on through its data. The free
calculators work the same way at a longer reach: [cFIREsim](https://stockmarketstack.com/tools/cfiresim) runs a 30-year plan
entered today against 127 cycles, 1871–1900 through 1997–2026, and [FI Calc](https://stockmarketstack.com/tools/ficalc) and
[FIREproof](https://stockmarketstack.com/tools/fireproof) replay the same kind of windows.

Two things follow from overlapping windows. The cycles are not independent — neighbouring 30-year
windows share 29 of their years — so 127 cycles are not 127 trials and a success rate is a count of
what happened, not a probability of what will. And the result is bounded by what the data contains:
a plan fails only if one of the recorded sequences broke it.

A **[Monte Carlo](https://stockmarketstack.com/glossary/monte-carlo-simulation)** calculator generates its paths, and the
generator is a setting; the glossary entry covers the kinds of generator and why each scores the
same plan differently. Three tools in this category already use three: [Boldin](https://stockmarketstack.com/tools/boldin)
draws 1,000 iterations from a normal distribution around your own average rate of return;
[ProjectionLab](https://stockmarketstack.com/tools/projectionlab) samples historical S&P 500 returns, dividend yields and US
inflation — sequentially, from a random start year, or by block bootstrap — or normal distributions
you specify, for up to 2,000 trials; and FIREproof's Monte Carlo mode runs per asset rather than
across the portfolio, so its success rates ignore diversification.

So the first question to ask of two disagreeing numbers is whether both were replayed or both
generated. Across that line they are not the same measurement.

## Which stocks, which bonds, which years

"US stocks since 1871" almost always means one dataset: the monthly series Robert Shiller published
with *Irrational Exuberance* — stock prices, dividends, earnings and a consumer price index, all from
January 1871. cFIREsim, FI Calc and FIREproof all build on it. Its construction, as Shiller's data page describes
it, is worth knowing, because it is not a record of anybody's account:

- **Prices are monthly averages of daily closing prices**, not month-end closes. An average smooths
  the path inside the month, so a series built from it is calmer than one built from closes.
- **Dividends before 1926 come from Cowles and associates, interpolated from annual data**; after
  1926 they are the S&P four-quarter totals, linearly interpolated to monthly figures.
- **It is an index, with no fees and no fund.** Nobody could buy the S&P composite of the 1870s.

The bond side differs more than the stock side. The Shiller-based calculators use the 10-year US
Treasury yield. The Trinity study represented stocks with the S&P 500 and bonds with long-term,
high-grade corporate bonds, all taken from Ibbotson Associates' 1996 yearbook and covering 1926 to
1995. A "60/40" plan is therefore not one portfolio across those tools: the forty is a different
asset. Cash is often not history at all — cFIREsim has you supply a growth rate, defaulting to
0.25%, and FI Calc a fixed rate, 1.5% by default.

The window matters as much as the series. A history that starts in 1871 includes decades a history
starting in 1926 does not, and one that starts in 1970 includes neither:
[Portfolio Charts](https://stockmarketstack.com/tools/portfolio-charts) runs from 1970, and [Portfolio Visualizer](https://stockmarketstack.com/tools/portfolio-visualizer)'s
asset-class backtests from 1972. Where the history ends matters too: the FI Calc bundle served on 14
September 2026 ended at 2024, while FIREproof's two default series carried a 2026
entry. And optional sleeves are shorter still — FIREproof substitutes another series before an
asset's start year and flags it only where you choose the asset, never in the results.

Series change underneath you as well. Portfolio Charts refreshes once a year: its 2025 figures
landed on 13 January 2026 marked preliminary until the government sources were finalised about a
month later. A backtest run in between and again afterwards can come out differently with no input
retyped.

## Which inflation

A retirement plan is usually a spending figure in today's money, raised every year by inflation. So
the inflation series is applied to the withdrawals, and a higher one drains the portfolio faster in
every path. There is more than one "CPI".

The Bureau of Labor Statistics publishes the CPI-U, for all urban consumers, which it calls the
broadest and most comprehensive CPI and which covers over 90% of the US population; and the CPI-W,
for urban wage earners and clerical workers, a subset of about 30%. Social Security's
cost-of-living adjustment is defined on the CPI-W: the percent increase between the third-quarter
average for a year and the previous peak third-quarter average. BLS also produces a research index
for Americans 62 and older, the R-CPI-E, and asks that it be interpreted with caution. A calculator
that grows your spending by CPI-U and your Social Security by a CPI-W-based adjustment is using two
indices in one plan, which is correct and still a source of drift against a tool that uses one.

Before 1913 there is no CPI at all. Shiller's series splices the CPI-U, which begins in 1913, onto
Warren and Pearson's price index for the years before it, scaled by the ratio of the two in January
1913. Every 1870s–1900s cycle in a Shiller-based tool is deflated by that splice.

The calculators then diverge on whether to use history at all. cFIREsim runs on historical CPI or a
flat rate you set. Boldin's deterministic projection defaults to assumptions derived from 1994–2024
data — 2.54% general inflation and 3.36% for medical costs. [Portfolio Visualizer](https://stockmarketstack.com/tools/portfolio-visualizer)
uses CPI-U. Outside the US the index changes entirely: [Curvo](https://stockmarketstack.com/tools/curvo) deflates by the Belgian
consumer price index, and Portfolio Charts switches the inflation series with the home country you
pick, one of twelve.

## Fees, taxes and the day money moves

Index history carries no costs, so any fee in a result is one the calculator added, and they add
different ones. cFIREsim takes an annual fee percentage for the whole portfolio. FI Calc takes an
expense ratio per sleeve. Portfolio Charts excludes fund fees and taxes altogether and says they
remain the reader's problem. A fee is a drag in every year of every path, so two tools that differ
only in whether one was entered can disagree by a wide margin on a long plan.

Taxes are the larger version of the same gap. cFIREsim and FI Calc model one undifferentiated pot of
money with no account types and no required minimum distributions, and expect you to add your tax
bill to the spending figure yourself. Boldin, ProjectionLab and FIREproof model taxes, each on its paid
tier. Enter the same spending into one of each and the first is
modelling a smaller withdrawal than the second, unless you grossed it up.

Then there is the calendar. An annual model has to decide the order of events inside each year,
and the order is not neutral. FI Calc takes the withdrawal on 1 January and applies fees, growth
and dividends on 31 December, with rebalancing last — so a year's spending never earns that year's
return. A tool that withdraws at year end, or monthly, leaves more money invested for longer and
scores the same plan differently — better in the years returns are positive, worse in the years
they are not. [Portfolio Visualizer](https://stockmarketstack.com/tools/portfolio-visualizer) works on
month-end balances, and [Curvo](https://stockmarketstack.com/tools/curvo) on monthly month-end data. None of these is wrong; they
are different arithmetic, and a success rate carries it without saying so.

## What counts as success

The last choice is the rule the percentage counts, and it is the one most often left unstated.
Three versions appear across this category:

- **Ended above zero.** The Trinity study's portfolio success rate was the percentage of past payout
  periods in which the ending value exceeded $0.
- **Never ran out.** cFIREsim's success means the portfolio never reached zero at any point in the
  cycle. ProjectionLab's success rate is the share of trials that reach the end of the plan without
  running out of money; what you set there are the buckets around it — failed early, failed late,
  almost failed, large surplus — and their thresholds, not the success line itself.
- **Ended at zero or more.** Boldin's help centre article on interpreting the score counts a trial
  as a success if the plan ends at $0 or more, even when it ran short mid-plan and a later inflow,
  such as a property sale, brought the balance back.

The three agree on most paths and part company on the ones that dip to zero and recover, or end
exactly on it. Boldin shows how easily the rule goes unstated: two of its own methodology articles
describe the same score as iterations in which funds never run out, which is the second rule, not
the third. When a vendor's pages disagree, the article that defines success is the one to go by. None of them says how much was left: a plan that ends with one dollar scores the
same as one that ends with its starting balance doubled, which is why the same success rate can
hide very different distributions of ending balances.

The spending rule interacts with the definition. A fixed inflation-adjusted withdrawal can fail
outright. A rule that takes a percentage of the current balance cannot reach zero at all, so its
success rate is trivially high and the risk moves into how far spending falls instead. FI Calc offers
twelve withdrawal strategies and FIREproof several guardrail and variable rules; comparing success
rates across two different spending rules compares two different questions.

## What you can do about it

The goal is not to find the right percentage. It is to find out whether two percentages are
measuring the same thing, and if they are not, which setting separates them.

1. **Write down the three choices for each tool before reading its result.** Replayed or generated,
   and if generated, from what. Which series for stocks, bonds, cash and inflation, over which years.
   Which success rule. If a tool will not tell you one of them, treat its number as unexplained
   rather than wrong.
2. **Match the easy settings first.** Same spending rule, same fee (or none on both), same
   inflation treatment — historical or a flat rate — and the same tax handling. Where one tool
   models taxes and the other does not, enter a grossed-up spending figure in the second.
3. **Compare like with like.** Put two historical-cycle tools side by side, or two Monte Carlo
   runs with the same generator. [cFIREsim](https://stockmarketstack.com/tools/cfiresim) and [FI Calc](https://stockmarketstack.com/tools/ficalc) share the
   Shiller data and the cycle method, so a remaining gap between them is in fees, cash, timing or the
   data's end year, and each of those can be tested by changing one input.
4. **Read the failures, not the rate.** A historical tool can show which start years failed;
   cFIREsim exports a year-by-year row for every cycle to CSV. Knowing that the failures cluster on
   a handful of start years tells you more about the plan than the fraction does, and it is the same
   information across tools that count it differently.
5. **Do not read precision into the decimal.** In a historical tool one cycle is one step of the
   percentage — under 1 point when there are 127 cycles — and the cycles overlap; in a generated one
   the figure moves with the seed and the trial count. A small gap between two tools is a question
   about settings before it is a finding about the plan.

Two questions to put to any calculator before paying for it: what counts as a failed path, and what
inflation series grows the spending. If the documentation cannot answer both, its success rate
cannot be compared with anything, including its own number from last year after a data refresh.

What this page does not do is say what success rate, or what spending level, is enough. That is a
judgement about your own life that no data series settles, and the point of the mechanics above is
that the number you would be judging moves when the model does.

## Tools this bears on

- [cFIREsim](https://stockmarketstack.com/tools/cfiresim.md) — Run a withdrawal plan against every market cycle since 1871 and count how many survived.
- [FI Calc](https://stockmarketstack.com/tools/ficalc.md) — Twelve withdrawal strategies replayed against every US market cycle since 1871.
- [Boldin](https://stockmarketstack.com/tools/boldin.md) — A year-by-year US retirement model — accounts, taxes, Social Security, Medicare, housing.
- [ProjectionLab](https://stockmarketstack.com/tools/projectionlab.md) — Model a whole financial life as milestones and cash flows, then run it against history.

## FAQ

### Why do cFIREsim and a Monte Carlo planner give different success rates for the same plan?

Because one replays overlapping windows of past returns and the other generates new paths, from a statistical distribution or by resampling history. They also differ in data series, inflation index, fees, timing within the year and the definition of success, so the two percentages measure different things even when every input you typed is the same.

### Is a 95% success rate a 95% chance my money lasts?

No. In a historical tool it is the share of past windows that met the tool's rule, and the windows overlap, so they are not independent trials. In a Monte Carlo tool it is the share of generated paths that met the rule, and it moves when the generator's assumptions do.

### What data do historical retirement calculators use?

Most US tools that reach back to 1871 use Robert Shiller's monthly dataset, in which stock prices are monthly averages of daily closes and dividends before 1926 are interpolated from annual data, with the 10-year Treasury yield for bonds and a CPI spliced onto an older price index before 1913.

### Which inflation index do retirement calculators use?

It varies. Many US tools grow spending by CPI-U, the broadest index, which covers over 90% of the US population. Social Security's cost-of-living adjustment is based on CPI-W, a subset covering about 30%. Some tools let you replace history with a flat rate, and non-US tools use their own country's index.

### What counts as success in a retirement calculator?

Each tool sets its own rule. The Trinity study counted periods ending above zero; cFIREsim and ProjectionLab require the money never to run out; Boldin's help centre counts a plan ending at $0 or more even if it ran short part-way, though two of its other articles describe the rule as never running out. None of them says how much was left.

## Sources

1. [Online Data, U.S. Stock Markets 1871-Present and CAPE Ratio](http://www.econ.yale.edu/~shiller/data.htm) — Robert J. Shiller, Yale University, read 2026-09-27
2. [Consumer Price Index Frequently Asked Questions](https://www.bls.gov/cpi/questions-and-answers.htm) — U.S. Bureau of Labor Statistics, 2025-09-25
3. [Retirement Savings: Choosing a Withdrawal Rate That Is Sustainable](https://www.aaii.com/journal/article/retirement-savings-choosing-a-withdrawal-rate-that-is-sustainable) — AAII Journal (Cooley, Hubbard and Walz), 1998-02-01. Cited only for the data series and the success definition the study used, which a later study cannot change.

*Last updated 2026-09-27. A reference page, corrected in place — not a dated post.*
