Skip to content

Backtest core

The three modules that decide whether a number means anything: execution timing, the statistics computed from it, and the splitting that keeps the model from seeing its own answer.

Execution

engine

Vectorised backtest with explicit execution timing.

The timing convention, stated once and obeyed everywhere:

position[t]  is decided using information up to the CLOSE of bar t
the resulting trade is FILLED at the OPEN of bar t+1
so position[t] earns the return from OPEN[t+1] to OPEN[t+2]
and pays cost on |position[t] - position[t-1]| at the fill

Assuming a fill at the close of the bar you just predicted is the second most common way to invent a profitable strategy that does not exist. There is no mechanism by which you observe a bar's close and also trade at it.

What this engine does NOT model, and you should not forget: partial fills, order book depth (it assumes your size is small enough not to move the market), funding rates on perpetuals, exchange downtime, and the fact that slippage is worst exactly when your signal is strongest.

One approximation worth naming: cost is subtracted in log space (net = gross - turnover * rate) although rate is a simple fraction. The exact charge is log(1 - rate), which is very slightly larger. The gap is rate^2/2 per unit turnover — about 3e-6 at 24bps, so under half a percent of equity even at a thousand round trips, and invisible at the turnover levels anything here survives. It flatters high-turnover strategies fractionally, which are the ones already dying by an order of magnitude.

backtest

backtest(bars: DataFrame, position: Series, cost_cfg) -> BacktestResult

Run the position series against the bars and charge realistic costs.

Source code in nullres/backtest/engine.py
def backtest(bars: pd.DataFrame, position: pd.Series, cost_cfg) -> BacktestResult:
    """Run the position series against the bars and charge realistic costs."""
    if not position.index.equals(bars.index):
        raise ValueError("position and bars must share an index")

    pos = position.astype("float64").fillna(0.0)

    log_open = np.log(bars["open"])
    # OPEN[t+1] -> OPEN[t+2]. The last two bars have no forward fill and earn 0.
    fwd = (log_open.shift(-2) - log_open.shift(-1)).fillna(0.0)

    turnover = pos.diff()
    turnover.iloc[0] = pos.iloc[0]        # opening the first position is a trade
    turnover = turnover.abs()

    rate = (cost_cfg.fee_bps + cost_cfg.slippage_bps) / 10_000.0
    cost = turnover * rate

    gross = pos * fwd
    net = gross - cost
    equity = np.exp(net.cumsum())

    return BacktestResult(
        equity=equity, returns=net, gross=gross,
        position=pos, turnover=turnover, cost=cost,
    )

restrict

restrict(result: BacktestResult, mask) -> BacktestResult

Narrow a result to a subset of bars, rebasing equity to the first of them.

Strategies are evaluated over the whole frame with positions ZEROED outside the out-of-sample window, not absent from it. A result therefore carries a block of structural zero returns, and summarising across them is not a harmless dilution: zeros scale the mean by their share f and the standard deviation by sqrt(f), so the reported Sharpe comes out as

sharpe_full = sqrt(f) * sharpe_oos

which is 0.88x on the 4h config — every strategy understated by the same 12%. The t-statistic is immune (the factor cancels top and bottom), which is why this survived a test suite that checks t-stats.

The closing trade is charged, not dropped. A position still open on the last in-window bar has to be liquidated, and that trade lands on the first bar AFTER the window — so simply masking discarded it, leaving every book unbilled for getting out and n_trades short by one. Buy & hold reported a single trade for a round trip.

Nothing needs to be invented to fix it: the exit already exists in the unmasked result, priced at the configured rate. Its cost and turnover are moved onto the final in-window bar, while its RETURN stays outside — you pay to leave, you do not earn the bar you left in. When the window runs to the end of the frame there is no following bar and nothing to re-attribute.

Source code in nullres/backtest/engine.py
def restrict(result: BacktestResult, mask) -> BacktestResult:
    """Narrow a result to a subset of bars, rebasing equity to the first of them.

    Strategies are evaluated over the whole frame with positions ZEROED outside
    the out-of-sample window, not absent from it. A result therefore carries a
    block of structural zero returns, and summarising across them is not a
    harmless dilution: zeros scale the mean by their share `f` and the standard
    deviation by `sqrt(f)`, so the reported Sharpe comes out as

        sharpe_full = sqrt(f) * sharpe_oos

    which is 0.88x on the 4h config — every strategy understated by the same
    12%. The t-statistic is immune (the factor cancels top and bottom), which is
    why this survived a test suite that checks t-stats.

    **The closing trade is charged, not dropped.** A position still open on the
    last in-window bar has to be liquidated, and that trade lands on the first
    bar AFTER the window — so simply masking discarded it, leaving every book
    unbilled for getting out and `n_trades` short by one. Buy & hold reported a
    single trade for a round trip.

    Nothing needs to be invented to fix it: the exit already exists in the
    unmasked result, priced at the configured rate. Its cost and turnover are
    moved onto the final in-window bar, while its RETURN stays outside — you pay
    to leave, you do not earn the bar you left in. When the window runs to the
    end of the frame there is no following bar and nothing to re-attribute.
    """
    selected = np.flatnonzero(np.asarray(mask, dtype=bool))
    r = result.returns[mask]
    turnover, cost = result.turnover[mask], result.cost[mask]

    if selected.size:
        exit_bar = selected[-1] + 1
        if exit_bar < len(result.cost):
            exit_cost = float(result.cost.iloc[exit_bar])
            if exit_cost > 0:
                turnover = turnover.copy()
                cost = cost.copy()
                turnover.iloc[-1] += float(result.turnover.iloc[exit_bar])
                cost.iloc[-1] += exit_cost
                r = r.copy()
                r.iloc[-1] -= exit_cost

    return BacktestResult(
        equity=np.exp(r.cumsum()),
        returns=r,
        gross=result.gross[mask],
        position=result.position[mask],
        turnover=turnover,
        cost=cost,
    )

buy_and_hold

buy_and_hold(bars: DataFrame, cost_cfg) -> BacktestResult

Reference strategy: fully long from the first fill, never trades again.

Source code in nullres/backtest/engine.py
def buy_and_hold(bars: pd.DataFrame, cost_cfg) -> BacktestResult:
    """Reference strategy: fully long from the first fill, never trades again."""
    pos = pd.Series(1.0, index=bars.index)
    return backtest(bars, pos, cost_cfg)

Metrics

metrics

Performance statistics, including the ones that are inconvenient.

Total return and Sharpe are the numbers people quote. The ones that decide whether a strategy is real are further down this list: the t-statistic on the mean return, how much of gross profit the costs ate, and how the Sharpe holds up once you account for how many variants you tried before this one.

summarize

summarize(result, bars_per_year: int, n_trials: int = 1, mask: Series | None = None) -> dict

Return a flat dict of performance statistics.

Parameters:

Name Type Description Default
n_trials int

how many strategy variants were evaluated before reporting this one. Used for the deflated Sharpe ratio. Be honest here — the count includes every threshold, horizon, and feature set you tried, not just the ones you kept.

1
mask Series | None

restrict to these bars before measuring anything. Pass the out-of-sample mask. Bars outside it hold a zeroed position, and averaging across those structural zeros multiplies the Sharpe by sqrt(fraction in-window) — see engine.restrict. It also fixes n_obs for the deflation term, which otherwise counts training bars the strategy never traded.

None
Source code in nullres/backtest/metrics.py
def summarize(result, bars_per_year: int, n_trials: int = 1,
              mask: pd.Series | None = None) -> dict:
    """Return a flat dict of performance statistics.

    Args:
        n_trials: how many strategy variants were evaluated before reporting
            this one. Used for the deflated Sharpe ratio. Be honest here — the
            count includes every threshold, horizon, and feature set you tried,
            not just the ones you kept.
        mask: restrict to these bars before measuring anything. Pass the
            out-of-sample mask. Bars outside it hold a zeroed position, and
            averaging across those structural zeros multiplies the Sharpe by
            `sqrt(fraction in-window)` — see `engine.restrict`. It also fixes
            `n_obs` for the deflation term, which otherwise counts training
            bars the strategy never traded.
    """
    if mask is not None:
        result = restrict(result, mask)

    r = result.returns.astype("float64")
    n = len(r)
    if n == 0:
        # An empty window is not a strategy that earned nothing; it is a
        # measurement that never happened. Every line below would either
        # IndexError on `equity.iloc[-1]` or invent a number — and a Sharpe of
        # 0.00 reported for a book that was never evaluated is exactly the kind
        # of confident-looking nonsense this package exists to refuse.
        #
        # `cross_symbol` already catches NullresError per symbol, so a symbol
        # whose out-of-sample window comes out empty is recorded as a note
        # rather than taking down the whole transfer test.
        raise InsufficientDataError(
            "no bars left to measure after masking — the out-of-sample window "
            "is empty, so there is nothing to summarise"
        )
    years = n / bars_per_year
    equity = result.equity

    total = float(equity.iloc[-1] - 1.0)
    cagr = float(equity.iloc[-1] ** (1 / years) - 1.0) if years > 0 else 0.0
    ann_vol = float(r.std() * np.sqrt(bars_per_year))
    sharpe = float(r.mean() / r.std() * np.sqrt(bars_per_year)) if r.std() > 0 else 0.0

    # Downside deviation is the root-mean-square of the NEGATIVE part taken
    # over every observation — sqrt(mean(min(r,0)^2)) — not the standard
    # deviation of the losing bars alone. The difference is not cosmetic and it
    # runs one way: the std of losses is computed about the mean loss rather
    # than about zero, and divides by the count of losing bars rather than by
    # all of them, so it is smaller and the ratio it produces is larger. On a
    # representative series it overstated Sortino by 13%.
    downside_dev = float(np.sqrt((np.minimum(r, 0.0) ** 2).mean()))
    sortino = (float(r.mean() / downside_dev * np.sqrt(bars_per_year))
               if downside_dev > 0 else 0.0)

    dd = _drawdown(equity)
    max_dd = float(dd.min())
    calmar = float(cagr / abs(max_dd)) if max_dd < 0 else 0.0

    # Is the mean return distinguishable from zero at all?
    t_stat = float(r.mean() / (r.std() / np.sqrt(n))) if r.std() > 0 and n > 1 else 0.0
    p_value = float(2 * (1 - sps.norm.cdf(abs(t_stat)))) if t_stat else 1.0

    gross_total = float(np.exp(result.gross.sum()) - 1.0)
    cost_log = result.total_cost
    exposure = float((result.position.abs() > 1e-12).mean())
    # `exposure` answers "how often was it in the market", which says nothing
    # about how BIG the position was. For a panel that is the whole question: a
    # dollar-neutral long/short book carries 1.0 long and 1.0 short, so its
    # gross notional is 2x equity, and `exposure` reported a serene 100% either
    # way. Mean |position| is the same statistic for both paths — average
    # leverage — and it is the one that tells you margin is involved.
    gross_exposure = float(result.position.abs().mean())
    peak_exposure = float(result.position.abs().max())
    # The engine sets gross = position * fwd, so returns[t] is earned by
    # position[t] — not position[t-1]. Masking on a shifted position asked
    # whether the WRONG bar was in the market.
    active = r[result.position.abs() > 1e-12]
    hit_rate = float((active > 0).mean()) if len(active) else 0.0

    return {
        "total_return": total,
        "cagr": cagr,
        "ann_vol": ann_vol,
        "sharpe": sharpe,
        "deflated_sharpe": deflated_sharpe(sharpe, n, bars_per_year, n_trials),
        "sortino": sortino,
        "max_dd": max_dd,
        "calmar": calmar,
        "t_stat": t_stat,
        "p_value": p_value,
        "gross_return": gross_total,
        "cost_drag": float(1 - np.exp(-cost_log)),
        "n_trades": result.n_trades,
        "turnover_per_year": float(result.turnover.sum() / years) if years > 0 else 0.0,
        "exposure": exposure,
        "gross_exposure": gross_exposure,
        "peak_exposure": peak_exposure,
        "hit_rate": hit_rate,
        "bars": n,
        "years": years,
    }

deflated_sharpe

deflated_sharpe(sharpe: float, n_obs: int, bars_per_year: int, n_trials: int = 1) -> float

Sharpe minus the best you would expect to reach by luck at n_trials.

Searching 100 strategy variants on pure noise yields a best-of-100 Sharpe around 0.6 by luck alone. This subtracts that, so the remainder is what needs explaining. At or below zero means: you found nothing, you just looked a lot of times.

Two honest caveats about what this is not. It borrows the expected- maximum term from Bailey & López de Prado's deflated Sharpe ratio, but it is not their statistic: the DSR proper is a probability and incorporates the skew and kurtosis of the return series, both of which are ignored here. And it treats the trials as independent, which they are not — twenty-five cells of one threshold sweep are nearly the same strategy, so the effective count is lower than the nominal one and this over-deflates. That error is in the conservative direction, which is the only reason it is tolerable.

Source code in nullres/backtest/metrics.py
def deflated_sharpe(sharpe: float, n_obs: int, bars_per_year: int,
                    n_trials: int = 1) -> float:
    """Sharpe minus the best you would expect to reach by luck at `n_trials`.

    Searching 100 strategy variants on pure noise yields a best-of-100 Sharpe
    around 0.6 by luck alone. This subtracts that, so the remainder is what
    needs explaining. At or below zero means: you found nothing, you just looked
    a lot of times.

    **Two honest caveats about what this is not.** It borrows the expected-
    maximum term from Bailey & López de Prado's deflated Sharpe ratio, but it is
    not their statistic: the DSR proper is a *probability* and incorporates the
    skew and kurtosis of the return series, both of which are ignored here. And
    it treats the trials as independent, which they are not — twenty-five cells
    of one threshold sweep are nearly the same strategy, so the effective count
    is lower than the nominal one and this over-deflates. That error is in the
    conservative direction, which is the only reason it is tolerable.
    """
    if n_trials <= 1 or n_obs < 2:
        return float(sharpe)
    euler = 0.5772156649
    # Expected maximum of n_trials standard normals.
    e_max = (
        (1 - euler) * sps.norm.ppf(1 - 1 / n_trials)
        + euler * sps.norm.ppf(1 - 1 / (n_trials * np.e))
    )
    # Convert to the same annualised units as `sharpe`.
    return float(sharpe - e_max * np.sqrt(bars_per_year / n_obs))

by_period

by_period(result, bars_per_year: int, mask: Series | None = None, freq: str = 'YE', min_coverage: float = 0.35) -> DataFrame

Break performance down by calendar period.

A strategy with a Sharpe of 0.5 built from one spectacular year and four flat ones is not the same object as one that earned 0.5 every year, and the headline number cannot tell them apart. This is the cheapest test for "did I fit a regime that has since ended".

Parameters:

Name Type Description Default
mask Series | None

restrict to these bars (pass the out-of-sample mask; bars outside it contribute structural zeros that deflate the volatility estimate and inflate Sharpe).

None
min_coverage float

drop periods covering less than this fraction of a full one. An out-of-sample window starting in December leaves a "2021" of one month, and annualising it produced a buy & hold Sharpe of -4.68 off two trades — a number with no meaning that nonetheless counted as a full observation in the stability verdict.

0.35
Source code in nullres/backtest/metrics.py
def by_period(result, bars_per_year: int, mask: pd.Series | None = None,
              freq: str = "YE",
              min_coverage: float = 0.35) -> pd.DataFrame:
    """Break performance down by calendar period.

    A strategy with a Sharpe of 0.5 built from one spectacular year and four
    flat ones is not the same object as one that earned 0.5 every year, and the
    headline number cannot tell them apart. This is the cheapest test for
    "did I fit a regime that has since ended".

    Args:
        mask: restrict to these bars (pass the out-of-sample mask; bars outside
            it contribute structural zeros that deflate the volatility estimate
            and inflate Sharpe).
        min_coverage: drop periods covering less than this fraction of a full
            one. An out-of-sample window starting in December leaves a "2021"
            of one month, and annualising it produced a buy & hold Sharpe of
            -4.68 off two trades — a number with no meaning that nonetheless
            counted as a full observation in the stability verdict.
    """
    r = result.returns
    turnover = result.turnover
    if mask is not None:
        r, turnover = r[mask], turnover[mask]

    full_period_bars = bars_per_year if freq.startswith("Y") else None

    rows = []
    for period, chunk in r.groupby(pd.Grouper(freq=freq)):
        if len(chunk) < 2 or chunk.std() == 0:
            continue
        if full_period_bars and len(chunk) < min_coverage * full_period_bars:
            continue

        # A period the strategy sat out is not a period it performed badly in.
        # If a handful of bars carry every non-zero return — typically one lone
        # cost tick from a position change — that tick IS the period's entire
        # variance, and annualising it yields a confident-looking Sharpe of
        # -1.00 off a book that did nothing. `min_coverage` already guards the
        # partial-period version of this; this guards the inactive one.
        active = chunk[chunk.abs() > 1e-12]
        if len(active) < max(2, 0.01 * len(chunk)):
            continue

        trades = int((turnover.loc[chunk.index] > 1e-12).sum())
        rows.append({
            "period": str(period)[:4] if freq.startswith("Y") else str(period)[:10],
            "bars": len(chunk),
            "total_return": float(np.exp(chunk.sum()) - 1.0),
            "sharpe": float(chunk.mean() / chunk.std() * np.sqrt(bars_per_year)),
            "n_trades": trades,
        })
    return pd.DataFrame(rows)

format_table

format_table(rows: dict[str, dict]) -> str

Render {strategy_name: metrics} as a fixed-width comparison table.

Source code in nullres/backtest/metrics.py
def format_table(rows: dict[str, dict]) -> str:
    """Render {strategy_name: metrics} as a fixed-width comparison table."""
    name_w = max([len(k) for k in rows] + [8]) + 2
    head = f"{'strategy':<{name_w}}" + "".join(
        f"{label:>{w}}" for _, label, w, _ in COLUMNS
    )
    lines = [head, "-" * len(head)]
    for name, m in rows.items():
        cells = "".join(
            f"{fmt(m.get(key, 0.0)):>{w}}" for key, _, w, fmt in COLUMNS
        )
        lines.append(f"{name:<{name_w}}" + cells)
    return "\n".join(lines)

Position sizing

sizing

Signal -> position. The step that decides whether an edge survives.

The baseline mapped proba > 0.52 straight to a position and changed its mind 15,527 times in 47,502 bars. At 12bps a side that is ~18.6 in log cost — the entire account, several times over, regardless of how good the model was.

Three mechanisms fix that, and all three are point-in-time safe:

HYSTERESIS Enter long above long_entry, but do not exit until the signal falls below the lower long_exit. A single band makes the position chatter every time the signal grazes the threshold.

MIN HOLD A hard floor on bars between state changes. This alone caps turnover at n/min_hold and is the bluntest, most reliable lever you have.

VOL TARGET Scale exposure by 1/volatility so risk, not notional, is constant. Improves risk-adjusted return and cuts trading in violent regimes.

apply_min_hold

apply_min_hold(desired: Series, min_hold: int) -> Series

Enforce a minimum bars-between-changes floor on an explicit position series.

Used when a strategy already knows the side it wants (meta-labelling, rule ensembles) and only needs turnover control, not threshold logic.

Source code in nullres/backtest/sizing.py
def apply_min_hold(desired: pd.Series, min_hold: int) -> pd.Series:
    """Enforce a minimum bars-between-changes floor on an explicit position series.

    Used when a strategy already knows the side it wants (meta-labelling, rule
    ensembles) and only needs turnover control, not threshold logic.
    """
    # A NaN means "no prediction here" — carry the last known intent forward
    # rather than forcing a flat, which would show up as spurious turnover.
    d = desired.ffill().fillna(0.0).to_numpy(dtype="float64")
    out = np.zeros(len(d))
    current, held = 0.0, min_hold
    for i in range(len(d)):
        if d[i] != current and held >= min_hold:
            current, held = d[i], 0
        else:
            held += 1
        out[i] = current
    return pd.Series(out, index=desired.index)

apply_rebalance_band

apply_rebalance_band(target: Series, band: float) -> Series

Only trade when the target drifts more than band away from what's held.

A continuously-varying target position is a continuously-varying trade. A naive vol-target on BTC 4h drifts ~0.9% of notional per bar, which compounds to ~20x annual turnover and ~2.4%/yr in fees — enough to erase the benefit it was built to deliver.

A no-trade band converts that into a handful of discrete rebalances. It is the same idea as hysteresis in signal_to_position, applied to size rather than to direction.

Source code in nullres/backtest/sizing.py
def apply_rebalance_band(target: pd.Series, band: float) -> pd.Series:
    """Only trade when the target drifts more than `band` away from what's held.

    A continuously-varying target position is a continuously-varying trade. A
    naive vol-target on BTC 4h drifts ~0.9% of notional per bar, which compounds
    to ~20x annual turnover and ~2.4%/yr in fees — enough to erase the benefit
    it was built to deliver.

    A no-trade band converts that into a handful of discrete rebalances. It is
    the same idea as hysteresis in `signal_to_position`, applied to size rather
    than to direction.
    """
    if band <= 0:
        return target
    t = target.to_numpy(dtype="float64")
    out = np.zeros(len(t))
    current = 0.0
    for i in range(len(t)):
        if np.isfinite(t[i]) and abs(t[i] - current) > band:
            current = t[i]
        out[i] = current
    return pd.Series(out, index=target.index)

apply_vol_target

apply_vol_target(pos: Series, sigma: Series, cfg, bars_per_year: int) -> Series

Scale exposure so annualised risk, not notional, is held constant.

Source code in nullres/backtest/sizing.py
def apply_vol_target(pos: pd.Series, sigma: pd.Series, cfg,
                     bars_per_year: int) -> pd.Series:
    """Scale exposure so annualised risk, not notional, is held constant."""
    if cfg.vol_target <= 0:
        return pos.clip(-cfg.max_leverage, cfg.max_leverage)
    ann_vol = sigma.to_numpy(dtype="float64") * np.sqrt(bars_per_year)
    with np.errstate(divide="ignore", invalid="ignore"):
        scale = np.where(ann_vol > 0, cfg.vol_target / ann_vol, 0.0)
    scale = np.nan_to_num(scale, nan=0.0, posinf=0.0)
    return (pos * np.clip(scale, 0.0, cfg.max_leverage)).clip(
        -cfg.max_leverage, cfg.max_leverage
    )

signal_to_position

signal_to_position(proba: Series, cfg, sigma: Series | None = None, bars_per_year: int = 8760) -> Series

Convert P(up) into a target position in [-max_leverage, +max_leverage].

proba is indexed by bar; NaN means "no opinion", which holds the current position rather than forcing an exit.

Source code in nullres/backtest/sizing.py
def signal_to_position(proba: pd.Series, cfg, sigma: pd.Series | None = None,
                       bars_per_year: int = 8_760) -> pd.Series:
    """Convert P(up) into a target position in [-max_leverage, +max_leverage].

    `proba` is indexed by bar; NaN means "no opinion", which holds the current
    position rather than forcing an exit.
    """
    p = proba.to_numpy(dtype="float64")
    n = len(p)
    state = np.zeros(n, dtype="float64")

    current = 0.0
    bars_held = 0

    for i in range(n):
        pi = p[i]
        desired = current

        if np.isfinite(pi):
            if current <= 0.0 and pi >= cfg.long_entry:
                desired = 1.0
            elif current >= 0.0 and cfg.allow_short and pi <= cfg.short_entry:
                desired = -1.0
            elif current > 0.0 and pi < cfg.long_exit:
                desired = 0.0
            elif current < 0.0 and pi > cfg.short_exit:
                desired = 0.0

        if desired != current and bars_held >= cfg.min_hold:
            current = desired
            bars_held = 0
        else:
            bars_held += 1

        state[i] = current

    pos = pd.Series(state, index=proba.index)

    if cfg.vol_target > 0:
        if sigma is None:
            raise ValueError("vol_target requires a sigma series")
        return apply_vol_target(pos, sigma, cfg, bars_per_year)
    return pos.clip(-cfg.max_leverage, cfg.max_leverage)

Validation

splits

Purged, embargoed walk-forward splitting.

Ordinary K-fold on time series is nonsense: it trains on Friday to predict Wednesday. Walk-forward fixes the direction but not the overlap — a label at bar t that resolves at t+24 still contains information about bars t+1..t+24, so if the test window starts at t+5, that training row has already seen the answer. Purging removes those rows.

The embargo goes further and drops training rows that merely END shortly before the test window opens. Serial correlation in features means a row from five bars before the boundary is nearly the same row as one inside it. Set embargo to roughly the feature memory (the longest rolling window) when you want to be strict.

remap_t_end

remap_t_end(t_end: ndarray, keep: ndarray) -> ndarray

Translate label end positions from the raw frame to the filtered frame.

Rows are dropped for NaN features or unlabelled bars, which renumbers every position. A label that ended on a dropped bar is mapped forward to the next surviving bar, which can only lengthen the purge — the safe direction.

Source code in nullres/validation/splits.py
def remap_t_end(t_end: np.ndarray, keep: np.ndarray) -> np.ndarray:
    """Translate label end positions from the raw frame to the filtered frame.

    Rows are dropped for NaN features or unlabelled bars, which renumbers every
    position. A label that ended on a dropped bar is mapped forward to the next
    surviving bar, which can only lengthen the purge — the safe direction.
    """
    surviving = np.flatnonzero(keep)
    if surviving.size == 0:
        return np.array([], dtype=np.int64)
    mapped = np.searchsorted(surviving, t_end[keep], side="left")
    return np.clip(mapped, np.arange(surviving.size), surviving.size - 1)

purged_walk_forward

purged_walk_forward(t_end: ndarray, cfg: SplitConfig) -> Iterator[tuple[ndarray, ndarray]]

Yield (train_idx, test_idx) pairs of positional indices.

Parameters:

Name Type Description Default
t_end ndarray

for each row, the position at which its label resolves.

required
cfg SplitConfig

a SplitConfig.

required
Source code in nullres/validation/splits.py
def purged_walk_forward(t_end: np.ndarray, cfg: SplitConfig
                        ) -> Iterator[tuple[np.ndarray, np.ndarray]]:
    """Yield (train_idx, test_idx) pairs of positional indices.

    Args:
        t_end: for each row, the position at which its label resolves.
        cfg: a SplitConfig.
    """
    n = len(t_end)
    if n <= cfg.min_train:
        raise InsufficientDataError(
            f"only {n:,} usable rows but min_train is {cfg.min_train:,}; "
            f"lower split.min_train or widen the date range"
        )

    fold = (n - cfg.min_train) // cfg.n_folds
    if fold < 100:
        raise InsufficientDataError(
            f"folds of {fold} rows are too small to measure anything; "
            f"reduce split.n_folds or add data"
        )

    for k in range(cfg.n_folds):
        test_start = cfg.min_train + k * fold
        test_end = min(test_start + fold, n)
        if test_end - test_start < 100:
            break

        if cfg.scheme == "rolling":
            train_lo = max(0, test_start - cfg.train_window)
        elif cfg.scheme == "expanding":
            train_lo = 0
        else:
            raise ConfigError(f"unknown split scheme {cfg.scheme!r}")

        cand = np.arange(train_lo, test_start)
        # Purge: the label must have fully resolved before the embargo window.
        train = cand[t_end[cand] < test_start - cfg.embargo]
        if train.size < 100:
            continue
        yield train, np.arange(test_start, test_end)

describe_folds

describe_folds(t_end: ndarray, cfg, index=None) -> list[dict]

Fold summary for logging — sizes, date ranges, and rows purged.

Source code in nullres/validation/splits.py
def describe_folds(t_end: np.ndarray, cfg, index=None) -> list[dict]:
    """Fold summary for logging — sizes, date ranges, and rows purged."""
    rows = []
    for i, (tr, te) in enumerate(purged_walk_forward(t_end, cfg), start=1):
        naive = te[0] - (max(0, te[0] - cfg.train_window) if cfg.scheme == "rolling" else 0)
        row = {
            "fold": i,
            "train": len(tr),
            "test": len(te),
            "purged": naive - len(tr),
        }
        if index is not None:
            row["test_from"] = str(index[te[0]])[:10]
            row["test_to"] = str(index[te[-1]])[:10]
        rows.append(row)
    return rows

weights

Sample weights for overlapping labels.

With a 24-bar horizon, 24 consecutive training rows describe almost the same stretch of price. Treating them as 24 independent observations tells the model it has far more evidence than it does, and it will happily overfit to match.

uniqueness_weights down-weights each row by how many other labels overlap it, following López de Prado's average-uniqueness construction. A row whose window is shared with 23 others carries roughly 1/24 the weight of an isolated one.

uniqueness_weights

uniqueness_weights(t_end: ndarray, n: int | None = None) -> ndarray

Average uniqueness of each label, in (0, 1].

Parameters:

Name Type Description Default
t_end ndarray

position at which each row's label resolves (inclusive).

required
n int | None

total bars; defaults to len(t_end).

None
Source code in nullres/validation/weights.py
def uniqueness_weights(t_end: np.ndarray, n: int | None = None) -> np.ndarray:
    """Average uniqueness of each label, in (0, 1].

    Args:
        t_end: position at which each row's label resolves (inclusive).
        n: total bars; defaults to len(t_end).
    """
    t_end = np.asarray(t_end, dtype=np.int64)
    m = len(t_end)
    n = m if n is None else n
    starts = np.arange(m, dtype=np.int64)
    ends = np.clip(t_end, starts, n - 1)

    # Concurrency: how many label windows cover each bar, via a difference array.
    diff = np.zeros(n + 1, dtype=np.float64)
    np.add.at(diff, starts, 1.0)
    np.add.at(diff, ends + 1, -1.0)
    concurrency = np.cumsum(diff)[:n]
    concurrency[concurrency < 1.0] = 1.0

    # Mean of 1/concurrency over each window, via a prefix sum.
    prefix = np.concatenate([[0.0], np.cumsum(1.0 / concurrency)])
    spans = (ends - starts + 1).astype(np.float64)
    return (prefix[ends + 1] - prefix[starts]) / spans

Cost arithmetic

costs

The cost budget: what accuracy does a strategy actually need to survive?

This is the arithmetic that decides whether a research direction is worth pursuing, and almost nobody does it before spending a month on features.

For a directional strategy:

edge            = 2 * accuracy - 1        (fraction of moves called right)
E|move| over h  ~ sigma * sqrt(h) * sqrt(2/pi)   for a driftless walk
gross per trade = edge * E|move|
cost per trade  = 2 * (fee + slippage)    (round trip, both sides)

Break-even requires gross >= cost. Rearranged, that gives the minimum accuracy a strategy needs at a given holding period — or the minimum holding period it needs at a given accuracy.

The result for hourly BTC is bracing. At 51% accuracy and 12bps a side, no holding period under several weeks breaks even. That is not pessimism, it is division, and it is why the honest baseline lost 100%: it was not a modelling failure, it was an arithmetic one.

expected_abs_move

expected_abs_move(sigma_per_bar: float, hold_bars: float) -> float

E|return| over hold_bars for a driftless GAUSSIAN random walk.

Crypto returns are neither Gaussian nor driftless, and the error is not a constant you can divide out — it changes sign with the horizon. Measured on BTC, this formula overstates the typical move by ~25% at one bar and understates it by ~13% at 336 bars: fat tails inflate sigma relative to a typical short move, while drift and trending make long moves larger than a driftless walk predicts. Aggregation pulls the middle toward normal.

Prefer empirical_abs_move when you have the return series. This is kept because it is invertible in closed form, which breakeven_hold relies on, and because it is the right null when you are reasoning rather than measuring.

Source code in nullres/costs.py
def expected_abs_move(sigma_per_bar: float, hold_bars: float) -> float:
    """E|return| over `hold_bars` for a driftless GAUSSIAN random walk.

    Crypto returns are neither Gaussian nor driftless, and the error is not a
    constant you can divide out — it changes sign with the horizon. Measured on
    BTC, this formula overstates the typical move by ~25% at one bar and
    understates it by ~13% at 336 bars: fat tails inflate sigma relative to a
    typical short move, while drift and trending make long moves larger than a
    driftless walk predicts. Aggregation pulls the middle toward normal.

    Prefer `empirical_abs_move` when you have the return series. This is kept
    because it is invertible in closed form, which `breakeven_hold` relies on,
    and because it is the right null when you are reasoning rather than
    measuring.
    """
    return sigma_per_bar * np.sqrt(hold_bars) * SQRT_2_OVER_PI

empirical_abs_move

empirical_abs_move(logret: ndarray, hold_bars: int) -> float

Measured mean |return| over hold_bars, from the returns themselves.

No distributional assumption: it sums the actual overlapping windows. The windows overlap, so this is an estimate of the mean and not an independent sample — fine for the purpose, which is a break-even threshold rather than an inference.

Source code in nullres/costs.py
def empirical_abs_move(logret: np.ndarray, hold_bars: int) -> float:
    """Measured mean |return| over `hold_bars`, from the returns themselves.

    No distributional assumption: it sums the actual overlapping windows. The
    windows overlap, so this is an estimate of the mean and not an independent
    sample — fine for the purpose, which is a break-even threshold rather than
    an inference.
    """
    r = np.asarray(logret, dtype="float64")
    r = r[np.isfinite(r)]
    hold_bars = int(hold_bars)
    if hold_bars < 1 or len(r) <= hold_bars:
        return float("nan")
    cumulative = np.concatenate([[0.0], np.cumsum(r)])
    windows = cumulative[hold_bars:] - cumulative[:-hold_bars]
    return float(np.abs(windows).mean())

breakeven_hold_empirical

breakeven_hold_empirical(logret: ndarray, accuracy: float, fee_bps: float, slippage_bps: float, max_hold: int = 8760) -> float

breakeven_hold against measured moves instead of the Gaussian one.

No closed form is available once the move is measured rather than modelled, so this bisects. empirical_abs_move rises monotonically with the horizon, which is what makes that valid.

Source code in nullres/costs.py
def breakeven_hold_empirical(logret: np.ndarray, accuracy: float,
                             fee_bps: float, slippage_bps: float,
                             max_hold: int = 8_760) -> float:
    """`breakeven_hold` against measured moves instead of the Gaussian one.

    No closed form is available once the move is measured rather than modelled,
    so this bisects. `empirical_abs_move` rises monotonically with the horizon,
    which is what makes that valid.
    """
    edge = 2.0 * accuracy - 1.0
    if edge <= 0:
        return float("inf")
    cost = round_trip_cost(fee_bps, slippage_bps)

    def gross(h: int) -> float:
        move = empirical_abs_move(logret, h)
        return edge * move if np.isfinite(move) else float("inf")

    if gross(1) >= cost:
        return 1.0
    if gross(max_hold) < cost:
        return float("inf")

    lo, hi = 1, max_hold
    while hi - lo > 1:
        mid = (lo + hi) // 2
        if gross(mid) < cost:
            lo = mid
        else:
            hi = mid
    return float(hi)

round_trip_cost

round_trip_cost(fee_bps: float, slippage_bps: float) -> float

Cost of getting in and back out again, as a fraction.

Source code in nullres/costs.py
def round_trip_cost(fee_bps: float, slippage_bps: float) -> float:
    """Cost of getting in and back out again, as a fraction."""
    return 2.0 * (fee_bps + slippage_bps) / 10_000.0

required_accuracy

required_accuracy(sigma_per_bar: float, hold_bars: float, fee_bps: float, slippage_bps: float) -> float

Directional accuracy needed to break even. Returns >1.0 when impossible.

Source code in nullres/costs.py
def required_accuracy(sigma_per_bar: float, hold_bars: float,
                      fee_bps: float, slippage_bps: float) -> float:
    """Directional accuracy needed to break even. Returns >1.0 when impossible."""
    move = expected_abs_move(sigma_per_bar, hold_bars)
    if move <= 0:
        return float("inf")
    edge = round_trip_cost(fee_bps, slippage_bps) / move
    return 0.5 * (1.0 + edge)

breakeven_hold

breakeven_hold(sigma_per_bar: float, accuracy: float, fee_bps: float, slippage_bps: float) -> float

Bars a position must be held for a given accuracy to break even.

Source code in nullres/costs.py
def breakeven_hold(sigma_per_bar: float, accuracy: float,
                   fee_bps: float, slippage_bps: float) -> float:
    """Bars a position must be held for a given accuracy to break even."""
    edge = 2.0 * accuracy - 1.0
    if edge <= 0:
        return float("inf")
    # edge * sigma * sqrt(h) * sqrt(2/pi) = cost  ->  solve for h
    return (round_trip_cost(fee_bps, slippage_bps)
            / (edge * sigma_per_bar * SQRT_2_OVER_PI)) ** 2

format_duration

format_duration(hours: float) -> str

Wall-clock rendering of a holding period, in units a human can act on.

Source code in nullres/costs.py
def format_duration(hours: float) -> str:
    """Wall-clock rendering of a holding period, in units a human can act on."""
    if not np.isfinite(hours):
        return "never"
    if hours < 48:
        return f"{hours:.1f} hours"
    days = hours / 24.0
    return f"{days:.1f} days" if days < 90 else f"{days / 30.4:.1f} months"

budget_table

budget_table(sigma_per_bar: float, fee_bps: float, slippage_bps: float, hours_per_bar: float = 1.0, holds=(1, 6, 12, 24, 72, 168, 336, 720), accuracies=(0.51, 0.52, 0.55, 0.6), logret=None) -> str

Render the two tables that should precede any modelling work.

hours_per_bar converts break-even BARS into wall-clock duration, and it is the whole point of the second table — duration is what is invariant across timeframes, bars are not. Hardcoding 24 bars-to-a-day (i.e. assuming hourly) made nullres budget claim a 1d config broke even in 0.9 days when the honest answer is the same ~21 days it is at every other timeframe. Callers pass 8760 / bars_per_year.

logret is the actual per-bar log return series. Given it, the table adds a measured column beside every modelled one — because the Gaussian assumption is wrong in a direction that flatters the strategy at exactly the holding periods people are tempted by. See expected_abs_move.

Source code in nullres/costs.py
def budget_table(sigma_per_bar: float, fee_bps: float, slippage_bps: float,
                 hours_per_bar: float = 1.0,
                 holds=(1, 6, 12, 24, 72, 168, 336, 720),
                 accuracies=(0.51, 0.52, 0.55, 0.60),
                 logret=None) -> str:
    """Render the two tables that should precede any modelling work.

    `hours_per_bar` converts break-even BARS into wall-clock duration, and it is
    the whole point of the second table — duration is what is invariant across
    timeframes, bars are not. Hardcoding 24 bars-to-a-day (i.e. assuming hourly)
    made `nullres budget` claim a 1d config broke even in 0.9 days when the
    honest answer is the same ~21 days it is at every other timeframe. Callers
    pass `8760 / bars_per_year`.

    `logret` is the actual per-bar log return series. Given it, the table adds a
    measured column beside every modelled one — because the Gaussian assumption
    is wrong in a direction that flatters the strategy at exactly the holding
    periods people are tempted by. See `expected_abs_move`.
    """
    cost = round_trip_cost(fee_bps, slippage_bps)
    measured = logret is not None and len(np.asarray(logret)) > max(holds)

    lines = [
        f"per-bar volatility      {sigma_per_bar:.4%}",
        f"round-trip cost         {cost:.4%}   ({fee_bps + slippage_bps:.0f}bps/side)",
        "",
        "Accuracy needed to break even, by holding period:",
    ]
    if measured:
        lines.append(f"  {'hold (bars)':<14}{'E|move|':>10}{'measured':>11}"
                     f"{'accuracy':>12}{'measured':>11}")
    else:
        lines.append(f"  {'hold (bars)':<14}{'E|move|':>10}{'accuracy':>12}")

    for h in holds:
        move = expected_abs_move(sigma_per_bar, h)
        acc = required_accuracy(sigma_per_bar, h, fee_bps, slippage_bps)
        cell = f"{acc:>11.1%}" if acc <= 1.0 else "  impossible"
        if not measured:
            lines.append(f"  {h:<14,}{move:>10.2%}{cell}")
            continue
        real = empirical_abs_move(logret, h)
        real_acc = 0.5 * (1.0 + cost / real) if real > 0 else float("inf")
        real_cell = f"{real_acc:>10.1%}" if real_acc <= 1.0 else " impossible"
        lines.append(f"  {h:<14,}{move:>10.2%}{real:>11.2%}{cell}{real_cell}")

    lines += ["", "Holding period needed to break even, by accuracy:"]
    if measured:
        lines.append(f"  {'accuracy':<14}{'hold (bars)':>14}{'measured':>11}"
                     f"{'~duration':>16}")
    else:
        lines.append(f"  {'accuracy':<14}{'hold (bars)':>14}{'~duration':>14}")

    for acc in accuracies:
        h = breakeven_hold(sigma_per_bar, acc, fee_bps, slippage_bps)
        if not measured:
            lines.append(f"  {acc:<14.0%}{h:>14,.0f}"
                         f"{format_duration(h * hours_per_bar):>14}")
            continue
        real_h = breakeven_hold_empirical(logret, acc, fee_bps, slippage_bps)
        duration = (format_duration(real_h * hours_per_bar)
                    if np.isfinite(real_h) else "never")
        real_cell = f"{real_h:>11,.0f}" if np.isfinite(real_h) else "      never"
        lines.append(f"  {acc:<14.0%}{h:>14,.0f}{real_cell}{duration:>16}")

    if measured:
        lines += [
            "",
            "  'measured' uses the actual distribution of moves rather than a",
            "  Gaussian. It is the column to trust: fat tails make short moves",
            "  SMALLER than sigma implies, so the modelled accuracy is too",
            "  forgiving at short holds, while drift makes long moves larger.",
            "  ~duration is the measured hold at this config's bar size.",
        ]
    return "\n".join(lines)