Backtest Rigor

Monte Carlo Backtest Resampler

Paste returns or trade P&L from any backtest. Run 10,000 block-bootstrap resamples in your browser. Read off the 5th/50th/95th percentile confidence intervals for Sharpe and max drawdown — and see the full resampled distribution as a histogram. Built for Hyperliquid 15-minute return series.

Block-bootstrap · in-browser · no upload
By Keel Research Team · Updated May 17, 2026
Inputs

500 values parsed.

One value per bar. Decimal (0.0012) or with % (0.12%). Comma- or newline-separated.

Default 10,000, max 100,000.

20-40 for HL 15-min; 96 for daily-persistence.

Result

Adjust inputs and click Run Monte Carlo. Bootstrap runs in your browser — no upload.

How it works

Methodology

A single backtest produces one Sharpe and one max drawdown — a point estimate. Block-bootstrap Monte Carlo answers the question the point estimate cannot: across many plausible realizations of the same underlying return process, what is the distribution of outcomes? You read the answer off as a confidence interval.

for i in 1..N:
  sample = moving_block_bootstrap(returns, L)
  sharpe[i], max_dd[i] = metrics(sample)
ci_low, median, ci_high = percentile(sharpe, [5, 50, 95])

Why blocks. Real return series are autocorrelated — momentum, mean reversion, volatility clustering, funding-regime persistence. Plain (i.i.d.) bootstrap destroys all of that. Block bootstrap resamples contiguous chunks of length L, which preserves local autocorrelation while still randomizing the global sample. For HL 15-minute bars, L = 20-40 is the right default for most signal classes; carry strategies with multi-day persistence want L = 96 (one day) or more.

What to do with the CI. The lower quantile is your honest planning number, not the median. A Sharpe with 90% CI [0.4, 3.8] is a fundamentally different decision than [1.7, 2.5] even at identical point estimate. For max drawdown the 95th-percentile bootstrap value is the right capital-planning number, not the single observed max DD in your backtest sample.

The widget caps N at 100,000 to keep in-browser runtime under a few seconds for typical return series. For larger series or tail-of-tail statistics, the same algorithm runs unbounded outside the browser — drop the same returns into a Python notebook with NumPy and run as many resamples as you want.

Automate it

Trade systematically on Keel

Keel is a Strategy OS for AI-assisted systematic trading on Hyperliquid. Build, backtest, and run live strategies with realistic fees, slippage, and funding modeled. Free to start — connect a Hyperliquid wallet when you’re ready to go live.

What you can do
  • Backtest any strategy with realistic fees, slippage, and funding modeled.
  • Iterate — change a parameter and re-run; every backtest is kept.
  • Deploy live to Hyperliquid with stop-loss + position limits.
  • Iterate with AI — describe a thesis, get a tradeable pipeline.
FAQ

Calculator questions

What does this calculator do?

It takes a CSV of trade P&L or per-bar returns, resamples it thousands of times using block bootstrap to preserve autocorrelation, computes Sharpe and max drawdown on each synthetic run, and returns the distribution. You get 5th/50th/95th percentile confidence intervals and a histogram of each metric across all resamples.

Why block bootstrap instead of plain bootstrap?

Plain bootstrap resamples individual returns independently. That assumes no autocorrelation. Real return series — and especially Hyperliquid 15-minute returns — have non-trivial autocorrelation from momentum, mean-reversion, volatility clustering, and funding-regime persistence. Plain bootstrap breaks all of that and produces confidence intervals that are too tight. Block bootstrap resamples contiguous chunks of length L so the local serial structure is preserved.

How do I pick block length?

Block length should approximately match the autocorrelation horizon of your returns. For HL 15-minute bars: default 20-40 (5-10 hours) for most signals; bump to 96 (one day) for carry strategies with multi-day persistence; drop to 8-12 for sub-hour mean reversion. When unsure, sweep across 10/20/40/96 — the point-estimate median should be stable, but the CI width should grow with block length up to the autocorrelation horizon.

How many resamples should I run?

Default is 10,000. That is enough for stable 5th/95th percentile estimates. For 99% CIs or extreme-tail statistics push to 20-50K. Below 5K the percentile estimates start to wobble run-to-run. Above 100K rarely changes anything meaningful. The widget caps at 100K to keep in-browser runtime under a few seconds.

Is my data uploaded anywhere?

No. The CSV is parsed in your browser and the bootstrap runs entirely in JavaScript on this page. Nothing is sent to a server. You can run this air-gapped on local Keel backtest output.

Does Keel compute these confidence intervals?

No. Keel reports point-estimate metrics from the backtest engine; the bootstrap CIs are computed here, in your browser, from the returns you paste. Bootstrap confidence intervals may come to Keel in the future.

For the longer-form HL framing — why 15-minute bars on Hyperliquid are autocorrelated, how to size block length for carry vs momentum, what wide CIs on max drawdown should change about live sizing — see Monte Carlo simulation on Hyperliquid backtests. Keel's backtest engine reports point estimates; this widget adds the confidence intervals.