Backtest Bench: a worked example on a public demo trade log

Part of PopperWick · Python, with the charts drawn from the run's own numbers · Run July 2026 · Data: QuantPad's public demo trade log, MNQ futures, 567 trades, January 2020 to December 2024 · Not my strategy

What this is

The bench is the part of PopperWick that grades a backtest. It does not ask whether a strategy made money; it asks whether the record is evidence of an edge, and it writes down every formula it uses. In July 2026 I subscribed to QuantPad, an AI backtesting service, for one month, exported the demo trade log its own report is built on, and ran that log through the bench. The log is theirs; the rubric, the code and the grade are mine. This page is about the method. It is not a review of QuantPad's product, and the strategy in the log is not one I trade or designed.

The verdict

DGrade D under the bench rubric

The four dimension scores average 71.5, which on its own bands as a C. Two kinds of finding cap the grade and never raise it: one critical finding caps at D, and so do two high ones. This log has one critical and two high findings, so C becomes D.

  1. CRITICAL Risk-adjusted return is weak. Risk-adjusted return is negative: the annualized Sharpe on the daily return series is −0.01. The strategy is not being compensated for the volatility it takes on.
  2. HIGH Edge is weak or not statistically distinguishable from zero. Mean per-trade return is 0.05% and stays positive through the 68% interval, but the 90% and 95% intervals include zero. The edge is weak; there is a real chance it is noise.
  3. HIGH Profit is concentrated in a handful of trades. The top 11 trades account for 26.35% of gross profit. Excluding them, mean per-trade return is −0.05%.
  4. MEDIUM Sample confidence is below target. Trade count is sufficient (567 trades; target at least 100), but daily Sharpe significance is below target: PSR 87.38% against a 90% target, on the daily return series over 1,812 calendar days.

The one chart that carries the verdict

Mean return per trade with its 68%, 90% and 95% intervals Mean 0.050% of the account per trade. 95% interval −0.036% to 0.141%, 90% interval −0.023% to 0.128%, 68% interval 0.005% to 0.098%. The 95% and 90% intervals include zero; the 68% interval does not. −0.05% 0.00% 0.05% 0.10% 0.15% 95% 90% 68% mean 0.050%
LevelLowHighIncludes zero
68%0.005%0.098%no
90%−0.023%0.128%yes
95%−0.036%0.141%yes
Mean return per trade, as a percentage of the $100,000 account, with 68%, 90% and 95% intervals. These are the values QuantPad's report displays for this log, reproduced by the bench. The 68% interval clears zero; the 90% and 95% intervals do not, which is the second finding.

Equity and drawdown

Cumulative profit and loss, and drawdown from the running peak Cumulative profit and loss at each of the 567 exits, 2020-01-05 to 2024-12-20, ending at $28,446.17. Drawdown from the running peak is deepest at −$12,648.12 on 2020-04-03, trade 35. 2020 2021 2022 2023 2024 −$10k $0 $10k $20k $30k $28,446 · 2024-12-20 drawdown from peak −$5k −$10k −$15k −$12,648 · 2020-04-03
Cumulative profit and loss as booked at each of the 567 exits, in dollars, and below it the distance from the running peak, which starts at zero. The endpoint is the reported net; the trough is the reported maximum drawdown, 12.6% of the account, reached at trade 35.

Key figures

Key figures from the bench run
FigureValueBasis
Net P&L$28,446.17sum of booked P&L over 567 trades
Win rate24.9%141 wins, 426 losses
Profit factor1.156gross profit $210,725.03 over gross loss $182,278.86
Average trade$50.17average win $1,494.50, average loss −$427.88
Largest win and loss$7,878.98 and −$1,806.73single trades
Maximum drawdown−$12,648.1212.6% of the $100,000 account, trade 35, 2020-04-03
Sharpe, per trade0.045QuantPad's figure: per trade, not annualized
Sharpe, time-based−0.01the bench's: daily P&L over 1,812 calendar days, 4.05% risk-free rate, annualized by the square root of 252
Probabilistic Sharpe ratio87.4%probability the true Sharpe exceeds zero; QuantPad's figure, the bench reproduces 87.38%
Active days485distinct exit dates in the log; QuantPad's report shows 477 and its derivation is undocumented

Dimension scores

EDGE9
ROBUSTNESS79, QuantPad showed 75
RISK100
SAMPLE ADEQUACY98

The bench scores four dimensions from disclosed formulas. Edge weighs the time-based Sharpe, the profit factor, the expectancy and the probabilistic Sharpe ratio. Robustness asks how much of the gross profit sits in the top 2% of trades and what the average trade earns without them. Risk is the maximum drawdown against the account. Sample adequacy combines the trade count with the probabilistic Sharpe ratio. QuantPad's own report showed 9, 75, 100 and 98 for this log. Three match; the bench's robustness formula gives 79 where theirs shows 75, and since their formula is not published the gap is recorded rather than explained.

Provenance and limits

  • The data is QuantPad's public demo trade log, exported in July 2026 during a one-month subscription: 567 trades on MNQ futures with exits from 5 January 2020 to 20 December 2024. It is not my strategy, and nothing here is a claim about QuantPad's product.
  • The bench ran on 12 July 2026. For this page its verdict functions were re-run on the same data on 26 September 2026, and every figure and chart is generated from one data file in the repository, tools/bench/quantpad-demo-bench.json, which carries the formulas in its provenance block.
  • The log has no commission or slippage column, so the P&L is as QuantPad booked it. Drawdown is measured from a running peak that starts at zero, so an opening loss counts.
  • The bench also produces a Monte Carlo study, a calendar, the full trade log and a regime breakdown. Those live in the research tool and are not on this page; nothing here is out of sample.
  • The grade is a statement about this log under this rubric, nothing more.