Why two retirement calculators give different success rates

Same savings, same spending, two success rates. Replayed history or Monte Carlo, which return and inflation series, fees, timing, and what counts as success.

Because a success rate describes the model as much as the plan. Historical-cycle calculators replay overlapping windows of past returns; Monte Carlo calculators generate paths from a distribution or by resampling. Each also picks its own stock, bond and inflation series, start year, fee treatment and timing convention, and defines success its own way. Two percentages are comparable only after those choices match, and a success rate is not a probability either way.

How it works

Every success rate in this category is built the same way: the number of paths that met a rule, divided by the number of paths the calculator ran. Your inputs — savings, spending, allocation, years — are only one of the three things that decide that fraction. The other two are the calculator's, and it rarely prints them next to the result.

Where the paths come from. Either the past, replayed in overlapping windows, or a generator that produces new sequences of returns. The generator can draw from a statistical distribution or reshuffle history.

What the paths are made of. A return series for each asset, an inflation series, a start year and an end year, a treatment of fees and taxes, and a convention for the day in each year when money is withdrawn, earned and rebalanced.

What the rule is. Ending above zero, never touching zero, or ending at zero or above after touching it.

Change any of the three and the same plan gets a different percentage with nothing about it having changed. That is the whole of the disagreement, and each part of it can be read off a calculator's documentation or its export if you know which question to ask. The sections below take them in that order.

Replayed history or generated paths

A historical-cycle calculator takes the plan's length and runs it against every window of that length in its data. The method is older than the software. The 1998 study by Cooley, Hubbard and Walz in the AAII Journal, the one usually called the Trinity study, computed outcomes for the first 15-year period, 1926 to 1940, then the second, 1927 to 1941, and so on through its data. The free calculators work the same way at a longer reach: cFIREsim runs a 30-year plan entered today against 127 cycles, 1871–1900 through 1997–2026, and FI Calc and FIREproof replay the same kind of windows.

Two things follow from overlapping windows. The cycles are not independent — neighbouring 30-year windows share 29 of their years — so 127 cycles are not 127 trials and a success rate is a count of what happened, not a probability of what will. And the result is bounded by what the data contains: a plan fails only if one of the recorded sequences broke it.

A Monte Carlo calculator generates its paths, and the generator is a setting; the glossary entry covers the kinds of generator and why each scores the same plan differently. Three tools in this category already use three: Boldin draws 1,000 iterations from a normal distribution around your own average rate of return; ProjectionLab samples historical S&P 500 returns, dividend yields and US inflation — sequentially, from a random start year, or by block bootstrap — or normal distributions you specify, for up to 2,000 trials; and FIREproof's Monte Carlo mode runs per asset rather than across the portfolio, so its success rates ignore diversification.

So the first question to ask of two disagreeing numbers is whether both were replayed or both generated. Across that line they are not the same measurement.

Which stocks, which bonds, which years

"US stocks since 1871" almost always means one dataset: the monthly series Robert Shiller published with Irrational Exuberance — stock prices, dividends, earnings and a consumer price index, all from January 1871. cFIREsim, FI Calc and FIREproof all build on it. Its construction, as Shiller's data page describes it, is worth knowing, because it is not a record of anybody's account:

  • Prices are monthly averages of daily closing prices, not month-end closes. An average smooths the path inside the month, so a series built from it is calmer than one built from closes.
  • Dividends before 1926 come from Cowles and associates, interpolated from annual data; after 1926 they are the S&P four-quarter totals, linearly interpolated to monthly figures.
  • It is an index, with no fees and no fund. Nobody could buy the S&P composite of the 1870s.

The bond side differs more than the stock side. The Shiller-based calculators use the 10-year US Treasury yield. The Trinity study represented stocks with the S&P 500 and bonds with long-term, high-grade corporate bonds, all taken from Ibbotson Associates' 1996 yearbook and covering 1926 to 1995. A "60/40" plan is therefore not one portfolio across those tools: the forty is a different asset. Cash is often not history at all — cFIREsim has you supply a growth rate, defaulting to 0.25%, and FI Calc a fixed rate, 1.5% by default.

The window matters as much as the series. A history that starts in 1871 includes decades a history starting in 1926 does not, and one that starts in 1970 includes neither: Portfolio Charts runs from 1970, and Portfolio Visualizer's asset-class backtests from 1972. Where the history ends matters too: the FI Calc bundle served on 14 September 2026 ended at 2024, while FIREproof's two default series carried a 2026 entry. And optional sleeves are shorter still — FIREproof substitutes another series before an asset's start year and flags it only where you choose the asset, never in the results.

Series change underneath you as well. Portfolio Charts refreshes once a year: its 2025 figures landed on 13 January 2026 marked preliminary until the government sources were finalised about a month later. A backtest run in between and again afterwards can come out differently with no input retyped.

Which inflation

A retirement plan is usually a spending figure in today's money, raised every year by inflation. So the inflation series is applied to the withdrawals, and a higher one drains the portfolio faster in every path. There is more than one "CPI".

The Bureau of Labor Statistics publishes the CPI-U, for all urban consumers, which it calls the broadest and most comprehensive CPI and which covers over 90% of the US population; and the CPI-W, for urban wage earners and clerical workers, a subset of about 30%. Social Security's cost-of-living adjustment is defined on the CPI-W: the percent increase between the third-quarter average for a year and the previous peak third-quarter average. BLS also produces a research index for Americans 62 and older, the R-CPI-E, and asks that it be interpreted with caution. A calculator that grows your spending by CPI-U and your Social Security by a CPI-W-based adjustment is using two indices in one plan, which is correct and still a source of drift against a tool that uses one.

Before 1913 there is no CPI at all. Shiller's series splices the CPI-U, which begins in 1913, onto Warren and Pearson's price index for the years before it, scaled by the ratio of the two in January 1913. Every 1870s–1900s cycle in a Shiller-based tool is deflated by that splice.

The calculators then diverge on whether to use history at all. cFIREsim runs on historical CPI or a flat rate you set. Boldin's deterministic projection defaults to assumptions derived from 1994–2024 data — 2.54% general inflation and 3.36% for medical costs. Portfolio Visualizer uses CPI-U. Outside the US the index changes entirely: Curvo deflates by the Belgian consumer price index, and Portfolio Charts switches the inflation series with the home country you pick, one of twelve.

Fees, taxes and the day money moves

Index history carries no costs, so any fee in a result is one the calculator added, and they add different ones. cFIREsim takes an annual fee percentage for the whole portfolio. FI Calc takes an expense ratio per sleeve. Portfolio Charts excludes fund fees and taxes altogether and says they remain the reader's problem. A fee is a drag in every year of every path, so two tools that differ only in whether one was entered can disagree by a wide margin on a long plan.

Taxes are the larger version of the same gap. cFIREsim and FI Calc model one undifferentiated pot of money with no account types and no required minimum distributions, and expect you to add your tax bill to the spending figure yourself. Boldin, ProjectionLab and FIREproof model taxes, each on its paid tier. Enter the same spending into one of each and the first is modelling a smaller withdrawal than the second, unless you grossed it up.

Then there is the calendar. An annual model has to decide the order of events inside each year, and the order is not neutral. FI Calc takes the withdrawal on 1 January and applies fees, growth and dividends on 31 December, with rebalancing last — so a year's spending never earns that year's return. A tool that withdraws at year end, or monthly, leaves more money invested for longer and scores the same plan differently — better in the years returns are positive, worse in the years they are not. Portfolio Visualizer works on month-end balances, and Curvo on monthly month-end data. None of these is wrong; they are different arithmetic, and a success rate carries it without saying so.

What counts as success

The last choice is the rule the percentage counts, and it is the one most often left unstated. Three versions appear across this category:

  • Ended above zero. The Trinity study's portfolio success rate was the percentage of past payout periods in which the ending value exceeded $0.
  • Never ran out. cFIREsim's success means the portfolio never reached zero at any point in the cycle. ProjectionLab's success rate is the share of trials that reach the end of the plan without running out of money; what you set there are the buckets around it — failed early, failed late, almost failed, large surplus — and their thresholds, not the success line itself.
  • Ended at zero or more. Boldin's help centre article on interpreting the score counts a trial as a success if the plan ends at $0 or more, even when it ran short mid-plan and a later inflow, such as a property sale, brought the balance back.

The three agree on most paths and part company on the ones that dip to zero and recover, or end exactly on it. Boldin shows how easily the rule goes unstated: two of its own methodology articles describe the same score as iterations in which funds never run out, which is the second rule, not the third. When a vendor's pages disagree, the article that defines success is the one to go by. None of them says how much was left: a plan that ends with one dollar scores the same as one that ends with its starting balance doubled, which is why the same success rate can hide very different distributions of ending balances.

The spending rule interacts with the definition. A fixed inflation-adjusted withdrawal can fail outright. A rule that takes a percentage of the current balance cannot reach zero at all, so its success rate is trivially high and the risk moves into how far spending falls instead. FI Calc offers twelve withdrawal strategies and FIREproof several guardrail and variable rules; comparing success rates across two different spending rules compares two different questions.

What you can do about it

The goal is not to find the right percentage. It is to find out whether two percentages are measuring the same thing, and if they are not, which setting separates them.

  1. Write down the three choices for each tool before reading its result. Replayed or generated, and if generated, from what. Which series for stocks, bonds, cash and inflation, over which years. Which success rule. If a tool will not tell you one of them, treat its number as unexplained rather than wrong.
  2. Match the easy settings first. Same spending rule, same fee (or none on both), same inflation treatment — historical or a flat rate — and the same tax handling. Where one tool models taxes and the other does not, enter a grossed-up spending figure in the second.
  3. Compare like with like. Put two historical-cycle tools side by side, or two Monte Carlo runs with the same generator. cFIREsim and FI Calc share the Shiller data and the cycle method, so a remaining gap between them is in fees, cash, timing or the data's end year, and each of those can be tested by changing one input.
  4. Read the failures, not the rate. A historical tool can show which start years failed; cFIREsim exports a year-by-year row for every cycle to CSV. Knowing that the failures cluster on a handful of start years tells you more about the plan than the fraction does, and it is the same information across tools that count it differently.
  5. Do not read precision into the decimal. In a historical tool one cycle is one step of the percentage — under 1 point when there are 127 cycles — and the cycles overlap; in a generated one the figure moves with the seed and the trial count. A small gap between two tools is a question about settings before it is a finding about the plan.

Two questions to put to any calculator before paying for it: what counts as a failed path, and what inflation series grows the spending. If the documentation cannot answer both, its success rate cannot be compared with anything, including its own number from last year after a data refresh.

What this page does not do is say what success rate, or what spending level, is enough. That is a judgement about your own life that no data series settles, and the point of the mechanics above is that the number you would be judging moves when the model does.

Tools this bears on

Cards in the catalogue where what is above changes the decision.

  • cFIREsim

    Run a withdrawal plan against every market cycle since 1871 and count how many survived.

    FreeFree tier

  • FI Calc

    Twelve withdrawal strategies replayed against every US market cycle since 1871.

    FreeFree tier

  • Boldin

    A year-by-year US retirement model — accounts, taxes, Social Security, Medicare, housing.

    $144/yrFree tier

  • ProjectionLab

    Model a whole financial life as milestones and cash flows, then run it against history.

    $129/yrFree tier

FAQ

Why do cFIREsim and a Monte Carlo planner give different success rates for the same plan?

Because one replays overlapping windows of past returns and the other generates new paths, from a statistical distribution or by resampling history. They also differ in data series, inflation index, fees, timing within the year and the definition of success, so the two percentages measure different things even when every input you typed is the same.

Is a 95% success rate a 95% chance my money lasts?

No. In a historical tool it is the share of past windows that met the tool's rule, and the windows overlap, so they are not independent trials. In a Monte Carlo tool it is the share of generated paths that met the rule, and it moves when the generator's assumptions do.

What data do historical retirement calculators use?

Most US tools that reach back to 1871 use Robert Shiller's monthly dataset, in which stock prices are monthly averages of daily closes and dividends before 1926 are interpolated from annual data, with the 10-year Treasury yield for bonds and a CPI spliced onto an older price index before 1913.

Which inflation index do retirement calculators use?

It varies. Many US tools grow spending by CPI-U, the broadest index, which covers over 90% of the US population. Social Security's cost-of-living adjustment is based on CPI-W, a subset covering about 30%. Some tools let you replace history with a flat rate, and non-US tools use their own country's index.

What counts as success in a retirement calculator?

Each tool sets its own rule. The Trinity study counted periods ending above zero; cFIREsim and ProjectionLab require the money never to run out; Boldin's help centre counts a plan ending at $0 or more even if it ran short part-way, though two of its other articles describe the rule as never running out. None of them says how much was left.

Sources

  1. Online Data, U.S. Stock Markets 1871-Present and CAPE Ratio — Robert J. Shiller, Yale University, read
  2. Consumer Price Index Frequently Asked Questions — U.S. Bureau of Labor Statistics,
  3. Retirement Savings: Choosing a Withdrawal Rate That Is Sustainable — AAII Journal (Cooley, Hubbard and Walz), . Cited only for the data series and the success definition the study used, which a later study cannot change.

The catalogue next door

This page is background, not a listing. The products it bears on are in Retirement & FIRE Calculators, each filled in against the same schema, with the fields to narrow it yourself.

Last updated . Corrected in place: this is a reference page, not a dated post.