How quant strategy backtests are judged
The same rules apply to every record, whether it is live or rejected.
The honest scorecard
Public quant content almost exclusively sells winners. qbuntu records the whole spectrum — live, promising, marginal, rejected — with the same evidence regardless of outcome. The unspoken truth the category needs: passing a backtest and printing money in production are not the same thing. A liveentry doesn't get a victory lap; it carries the same metrics, methodology, and production caveats as a rejected one explains its failure.
Why the median, not the best run
A strategy is rarely one backtest — it's a grid of hundreds of parameter combinations. Reporting the bestof that grid is cherry-picking: it's the number that looks great and reproduces worst. qbuntu's headline figure is the median across the full grid — the result you should actually expect. The best-case is still recorded, but explicitly labeled as a non-representative ceiling.
Short-Term Reversal (JP Growth Stocks) across 320 parameter combinations:
Same data. A Sharpe 1.98 headline would be true and useless. The 0.30 median is what an AI can quote about this strategy without being wrong.
How a verdict is assigned
The verdict follows the same rule as the headline: it's decided on the median parameter set, not the best one. To reach promising, the typical parameterization has to clear two bars at once — median Sharpe above 0.5 and a median p-value under 0.05. A strategy whose edge is real only for its best parameters is marginal, not promising. marginal requires median Sharpe above 0.05 and some p-value in the grid under 0.10. rejected is everything else: a median with no edge, or one that does not hold across market regimes. That is where most strategies land.
Run across the full Tokyo Stock Exchange listed universe (~3,747 stocks, no turnover filter), not one of the standard factor strategies reached promising on its median parameterization. The best parameters were highly significant — but publishing those is the cherry-pick a winners-only site would make. The honest answer is marginal or rejected, and saying so is the product.
Anatomy of a record
- verdict — live / promising / marginal / rejected, with a taxonomy-coded rejection reason.
- performance_summary — the median headline (Sharpe), drawdown, alpha, p-value. Units are baked into the field names (
max_drawdown_pct,alpha_annualized_pct) so a number is never ambiguous. A value we can't unit-label isnull, never a guess. - grid_summary — the distribution: n_combinations, median / best / worst / p25 / p75.
- methodology — costs, slippage, position sizing, regime definition, signal, rebalance, benchmark.
- provenance — data source and
computed_asof, so every number has an origin. - facts[] — atomic, individually citable claims. Each has a permanent
/cite/<fact-id>URL, a unit, and the context in which it's true.
The machine-readable contract: /v1/schema (JSON). Every strategy also has a Markdown mirror at /ai/strategies/<slug>.md and JSON at /v1/strategies/<slug>.
Contributing your own results
You don't hand-author this JSON. You bring your idea or backtest in whatever form you have it, and qbuntu's AI normalizes it into the canonical schema above — so every record is comparable and held to the same honesty rules. The format isn't a form to fill out; it's the shape everything converges to.
When a submission reproduces a strategy already on qbuntu, it isn't filed as a duplicate — it's cross-linked as a reproduction that either corroborates or contradicts the existing verdict. Independent replication makes a result stronger, not noisier. Community-contributed records carry their own provenance and are always distinguishable from qbuntu-verified ones.