Backtest core¶
The three modules that decide whether a number means anything: execution timing, the statistics computed from it, and the splitting that keeps the model from seeing its own answer.
Execution¶
engine ¶
Vectorised backtest with explicit execution timing.
The timing convention, stated once and obeyed everywhere:
position[t] is decided using information up to the CLOSE of bar t
the resulting trade is FILLED at the OPEN of bar t+1
so position[t] earns the return from OPEN[t+1] to OPEN[t+2]
and pays cost on |position[t] - position[t-1]| at the fill
Assuming a fill at the close of the bar you just predicted is the second most common way to invent a profitable strategy that does not exist. There is no mechanism by which you observe a bar's close and also trade at it.
What this engine does NOT model, and you should not forget: partial fills, order book depth (it assumes your size is small enough not to move the market), funding rates on perpetuals, exchange downtime, and the fact that slippage is worst exactly when your signal is strongest.
One approximation worth naming: cost is subtracted in log space
(net = gross - turnover * rate) although rate is a simple fraction. The
exact charge is log(1 - rate), which is very slightly larger. The gap is
rate^2/2 per unit turnover — about 3e-6 at 24bps, so under half a percent of
equity even at a thousand round trips, and invisible at the turnover levels
anything here survives. It flatters high-turnover strategies fractionally, which
are the ones already dying by an order of magnitude.
backtest ¶
Run the position series against the bars and charge realistic costs.
Source code in nullres/backtest/engine.py
restrict ¶
Narrow a result to a subset of bars, rebasing equity to the first of them.
Strategies are evaluated over the whole frame with positions ZEROED outside
the out-of-sample window, not absent from it. A result therefore carries a
block of structural zero returns, and summarising across them is not a
harmless dilution: zeros scale the mean by their share f and the standard
deviation by sqrt(f), so the reported Sharpe comes out as
sharpe_full = sqrt(f) * sharpe_oos
which is 0.88x on the 4h config — every strategy understated by the same 12%. The t-statistic is immune (the factor cancels top and bottom), which is why this survived a test suite that checks t-stats.
The closing trade is charged, not dropped. A position still open on the
last in-window bar has to be liquidated, and that trade lands on the first
bar AFTER the window — so simply masking discarded it, leaving every book
unbilled for getting out and n_trades short by one. Buy & hold reported a
single trade for a round trip.
Nothing needs to be invented to fix it: the exit already exists in the unmasked result, priced at the configured rate. Its cost and turnover are moved onto the final in-window bar, while its RETURN stays outside — you pay to leave, you do not earn the bar you left in. When the window runs to the end of the frame there is no following bar and nothing to re-attribute.
Source code in nullres/backtest/engine.py
buy_and_hold ¶
Reference strategy: fully long from the first fill, never trades again.
Metrics¶
metrics ¶
Performance statistics, including the ones that are inconvenient.
Total return and Sharpe are the numbers people quote. The ones that decide whether a strategy is real are further down this list: the t-statistic on the mean return, how much of gross profit the costs ate, and how the Sharpe holds up once you account for how many variants you tried before this one.
summarize ¶
Return a flat dict of performance statistics.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
n_trials
|
int
|
how many strategy variants were evaluated before reporting this one. Used for the deflated Sharpe ratio. Be honest here — the count includes every threshold, horizon, and feature set you tried, not just the ones you kept. |
1
|
mask
|
Series | None
|
restrict to these bars before measuring anything. Pass the
out-of-sample mask. Bars outside it hold a zeroed position, and
averaging across those structural zeros multiplies the Sharpe by
|
None
|
Source code in nullres/backtest/metrics.py
23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 | |
deflated_sharpe ¶
Sharpe minus the best you would expect to reach by luck at n_trials.
Searching 100 strategy variants on pure noise yields a best-of-100 Sharpe around 0.6 by luck alone. This subtracts that, so the remainder is what needs explaining. At or below zero means: you found nothing, you just looked a lot of times.
Two honest caveats about what this is not. It borrows the expected- maximum term from Bailey & López de Prado's deflated Sharpe ratio, but it is not their statistic: the DSR proper is a probability and incorporates the skew and kurtosis of the return series, both of which are ignored here. And it treats the trials as independent, which they are not — twenty-five cells of one threshold sweep are nearly the same strategy, so the effective count is lower than the nominal one and this over-deflates. That error is in the conservative direction, which is the only reason it is tolerable.
Source code in nullres/backtest/metrics.py
by_period ¶
by_period(result, bars_per_year: int, mask: Series | None = None, freq: str = 'YE', min_coverage: float = 0.35) -> DataFrame
Break performance down by calendar period.
A strategy with a Sharpe of 0.5 built from one spectacular year and four flat ones is not the same object as one that earned 0.5 every year, and the headline number cannot tell them apart. This is the cheapest test for "did I fit a regime that has since ended".
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
mask
|
Series | None
|
restrict to these bars (pass the out-of-sample mask; bars outside it contribute structural zeros that deflate the volatility estimate and inflate Sharpe). |
None
|
min_coverage
|
float
|
drop periods covering less than this fraction of a full one. An out-of-sample window starting in December leaves a "2021" of one month, and annualising it produced a buy & hold Sharpe of -4.68 off two trades — a number with no meaning that nonetheless counted as a full observation in the stability verdict. |
0.35
|
Source code in nullres/backtest/metrics.py
format_table ¶
Render {strategy_name: metrics} as a fixed-width comparison table.
Source code in nullres/backtest/metrics.py
Position sizing¶
sizing ¶
Signal -> position. The step that decides whether an edge survives.
The baseline mapped proba > 0.52 straight to a position and changed its mind
15,527 times in 47,502 bars. At 12bps a side that is ~18.6 in log cost — the
entire account, several times over, regardless of how good the model was.
Three mechanisms fix that, and all three are point-in-time safe:
HYSTERESIS Enter long above long_entry, but do not exit until the signal
falls below the lower long_exit. A single band makes the
position chatter every time the signal grazes the threshold.
MIN HOLD A hard floor on bars between state changes. This alone caps turnover at n/min_hold and is the bluntest, most reliable lever you have.
VOL TARGET Scale exposure by 1/volatility so risk, not notional, is constant. Improves risk-adjusted return and cuts trading in violent regimes.
apply_min_hold ¶
Enforce a minimum bars-between-changes floor on an explicit position series.
Used when a strategy already knows the side it wants (meta-labelling, rule ensembles) and only needs turnover control, not threshold logic.
Source code in nullres/backtest/sizing.py
apply_rebalance_band ¶
Only trade when the target drifts more than band away from what's held.
A continuously-varying target position is a continuously-varying trade. A naive vol-target on BTC 4h drifts ~0.9% of notional per bar, which compounds to ~20x annual turnover and ~2.4%/yr in fees — enough to erase the benefit it was built to deliver.
A no-trade band converts that into a handful of discrete rebalances. It is
the same idea as hysteresis in signal_to_position, applied to size rather
than to direction.
Source code in nullres/backtest/sizing.py
apply_vol_target ¶
Scale exposure so annualised risk, not notional, is held constant.
Source code in nullres/backtest/sizing.py
signal_to_position ¶
signal_to_position(proba: Series, cfg, sigma: Series | None = None, bars_per_year: int = 8760) -> Series
Convert P(up) into a target position in [-max_leverage, +max_leverage].
proba is indexed by bar; NaN means "no opinion", which holds the current
position rather than forcing an exit.
Source code in nullres/backtest/sizing.py
Validation¶
splits ¶
Purged, embargoed walk-forward splitting.
Ordinary K-fold on time series is nonsense: it trains on Friday to predict Wednesday. Walk-forward fixes the direction but not the overlap — a label at bar t that resolves at t+24 still contains information about bars t+1..t+24, so if the test window starts at t+5, that training row has already seen the answer. Purging removes those rows.
The embargo goes further and drops training rows that merely END shortly
before the test window opens. Serial correlation in features means a row from
five bars before the boundary is nearly the same row as one inside it. Set
embargo to roughly the feature memory (the longest rolling window) when you
want to be strict.
remap_t_end ¶
Translate label end positions from the raw frame to the filtered frame.
Rows are dropped for NaN features or unlabelled bars, which renumbers every position. A label that ended on a dropped bar is mapped forward to the next surviving bar, which can only lengthen the purge — the safe direction.
Source code in nullres/validation/splits.py
purged_walk_forward ¶
Yield (train_idx, test_idx) pairs of positional indices.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
t_end
|
ndarray
|
for each row, the position at which its label resolves. |
required |
cfg
|
SplitConfig
|
a SplitConfig. |
required |
Source code in nullres/validation/splits.py
describe_folds ¶
Fold summary for logging — sizes, date ranges, and rows purged.
Source code in nullres/validation/splits.py
weights ¶
Sample weights for overlapping labels.
With a 24-bar horizon, 24 consecutive training rows describe almost the same stretch of price. Treating them as 24 independent observations tells the model it has far more evidence than it does, and it will happily overfit to match.
uniqueness_weights down-weights each row by how many other labels overlap it,
following López de Prado's average-uniqueness construction. A row whose window
is shared with 23 others carries roughly 1/24 the weight of an isolated one.
uniqueness_weights ¶
Average uniqueness of each label, in (0, 1].
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
t_end
|
ndarray
|
position at which each row's label resolves (inclusive). |
required |
n
|
int | None
|
total bars; defaults to len(t_end). |
None
|
Source code in nullres/validation/weights.py
Cost arithmetic¶
costs ¶
The cost budget: what accuracy does a strategy actually need to survive?
This is the arithmetic that decides whether a research direction is worth pursuing, and almost nobody does it before spending a month on features.
For a directional strategy:
edge = 2 * accuracy - 1 (fraction of moves called right)
E|move| over h ~ sigma * sqrt(h) * sqrt(2/pi) for a driftless walk
gross per trade = edge * E|move|
cost per trade = 2 * (fee + slippage) (round trip, both sides)
Break-even requires gross >= cost. Rearranged, that gives the minimum accuracy a strategy needs at a given holding period — or the minimum holding period it needs at a given accuracy.
The result for hourly BTC is bracing. At 51% accuracy and 12bps a side, no holding period under several weeks breaks even. That is not pessimism, it is division, and it is why the honest baseline lost 100%: it was not a modelling failure, it was an arithmetic one.
expected_abs_move ¶
E|return| over hold_bars for a driftless GAUSSIAN random walk.
Crypto returns are neither Gaussian nor driftless, and the error is not a constant you can divide out — it changes sign with the horizon. Measured on BTC, this formula overstates the typical move by ~25% at one bar and understates it by ~13% at 336 bars: fat tails inflate sigma relative to a typical short move, while drift and trending make long moves larger than a driftless walk predicts. Aggregation pulls the middle toward normal.
Prefer empirical_abs_move when you have the return series. This is kept
because it is invertible in closed form, which breakeven_hold relies on,
and because it is the right null when you are reasoning rather than
measuring.
Source code in nullres/costs.py
empirical_abs_move ¶
Measured mean |return| over hold_bars, from the returns themselves.
No distributional assumption: it sums the actual overlapping windows. The windows overlap, so this is an estimate of the mean and not an independent sample — fine for the purpose, which is a break-even threshold rather than an inference.
Source code in nullres/costs.py
breakeven_hold_empirical ¶
breakeven_hold_empirical(logret: ndarray, accuracy: float, fee_bps: float, slippage_bps: float, max_hold: int = 8760) -> float
breakeven_hold against measured moves instead of the Gaussian one.
No closed form is available once the move is measured rather than modelled,
so this bisects. empirical_abs_move rises monotonically with the horizon,
which is what makes that valid.
Source code in nullres/costs.py
round_trip_cost ¶
required_accuracy ¶
required_accuracy(sigma_per_bar: float, hold_bars: float, fee_bps: float, slippage_bps: float) -> float
Directional accuracy needed to break even. Returns >1.0 when impossible.
Source code in nullres/costs.py
breakeven_hold ¶
Bars a position must be held for a given accuracy to break even.
Source code in nullres/costs.py
format_duration ¶
Wall-clock rendering of a holding period, in units a human can act on.
Source code in nullres/costs.py
budget_table ¶
budget_table(sigma_per_bar: float, fee_bps: float, slippage_bps: float, hours_per_bar: float = 1.0, holds=(1, 6, 12, 24, 72, 168, 336, 720), accuracies=(0.51, 0.52, 0.55, 0.6), logret=None) -> str
Render the two tables that should precede any modelling work.
hours_per_bar converts break-even BARS into wall-clock duration, and it is
the whole point of the second table — duration is what is invariant across
timeframes, bars are not. Hardcoding 24 bars-to-a-day (i.e. assuming hourly)
made nullres budget claim a 1d config broke even in 0.9 days when the
honest answer is the same ~21 days it is at every other timeframe. Callers
pass 8760 / bars_per_year.
logret is the actual per-bar log return series. Given it, the table adds a
measured column beside every modelled one — because the Gaussian assumption
is wrong in a direction that flatters the strategy at exactly the holding
periods people are tempted by. See expected_abs_move.