Backtest Bench: a worked example on a public demo trade log
What this is
The bench is the part of PopperWick that grades a backtest. It does not ask whether a strategy made money; it asks whether the record is evidence of an edge, and it writes down every formula it uses. In July 2026 I subscribed to QuantPad, an AI backtesting service, for one month, exported the demo trade log its own report is built on, and ran that log through the bench. The log is theirs; the rubric, the code and the grade are mine. This page is about the method. It is not a review of QuantPad's product, and the strategy in the log is not one I trade or designed.
The verdict
DGrade D under the bench rubric
The four dimension scores average 71.5, which on its own bands as a C. Two kinds of finding cap the grade and never raise it: one critical finding caps at D, and so do two high ones. This log has one critical and two high findings, so C becomes D.
- CRITICAL Risk-adjusted return is weak. Risk-adjusted return is negative: the annualized Sharpe on the daily return series is −0.01. The strategy is not being compensated for the volatility it takes on.
- HIGH Edge is weak or not statistically distinguishable from zero. Mean per-trade return is 0.05% and stays positive through the 68% interval, but the 90% and 95% intervals include zero. The edge is weak; there is a real chance it is noise.
- HIGH Profit is concentrated in a handful of trades. The top 11 trades account for 26.35% of gross profit. Excluding them, mean per-trade return is −0.05%.
- MEDIUM Sample confidence is below target. Trade count is sufficient (567 trades; target at least 100), but daily Sharpe significance is below target: PSR 87.38% against a 90% target, on the daily return series over 1,812 calendar days.
The one chart that carries the verdict
| Level | Low | High | Includes zero |
|---|---|---|---|
| 68% | 0.005% | 0.098% | no |
| 90% | −0.023% | 0.128% | yes |
| 95% | −0.036% | 0.141% | yes |
Equity and drawdown
Key figures
| Figure | Value | Basis |
|---|---|---|
| Net P&L | $28,446.17 | sum of booked P&L over 567 trades |
| Win rate | 24.9% | 141 wins, 426 losses |
| Profit factor | 1.156 | gross profit $210,725.03 over gross loss $182,278.86 |
| Average trade | $50.17 | average win $1,494.50, average loss −$427.88 |
| Largest win and loss | $7,878.98 and −$1,806.73 | single trades |
| Maximum drawdown | −$12,648.12 | 12.6% of the $100,000 account, trade 35, 2020-04-03 |
| Sharpe, per trade | 0.045 | QuantPad's figure: per trade, not annualized |
| Sharpe, time-based | −0.01 | the bench's: daily P&L over 1,812 calendar days, 4.05% risk-free rate, annualized by the square root of 252 |
| Probabilistic Sharpe ratio | 87.4% | probability the true Sharpe exceeds zero; QuantPad's figure, the bench reproduces 87.38% |
| Active days | 485 | distinct exit dates in the log; QuantPad's report shows 477 and its derivation is undocumented |
Dimension scores
The bench scores four dimensions from disclosed formulas. Edge weighs the time-based Sharpe, the profit factor, the expectancy and the probabilistic Sharpe ratio. Robustness asks how much of the gross profit sits in the top 2% of trades and what the average trade earns without them. Risk is the maximum drawdown against the account. Sample adequacy combines the trade count with the probabilistic Sharpe ratio. QuantPad's own report showed 9, 75, 100 and 98 for this log. Three match; the bench's robustness formula gives 79 where theirs shows 75, and since their formula is not published the gap is recorded rather than explained.
Provenance and limits
- The data is QuantPad's public demo trade log, exported in July 2026 during a one-month subscription: 567 trades on MNQ futures with exits from 5 January 2020 to 20 December 2024. It is not my strategy, and nothing here is a claim about QuantPad's product.
- The bench ran on 12 July 2026. For this page its verdict functions were re-run on the same data on 26 September 2026, and every figure and chart is generated from one data file in the repository,
tools/bench/quantpad-demo-bench.json, which carries the formulas in its provenance block. - The log has no commission or slippage column, so the P&L is as QuantPad booked it. Drawdown is measured from a running peak that starts at zero, so an opening loss counts.
- The bench also produces a Monte Carlo study, a calendar, the full trade log and a regime breakdown. Those live in the research tool and are not on this page; nothing here is out of sample.
- The grade is a statement about this log under this rubric, nothing more.