Skip to content

Falsification

The machinery that tries to disprove a result rather than produce one.

Leak detection

audit

Mechanical leakage detection.

The baseline script's closing line was right: walk-forward validation did not catch the leak, only reading the label definition would have. That is an unacceptable place to leave things — humans re-read code badly, and every new feature is a fresh chance to introduce lookahead.

These five checks catch the overwhelming majority of leaks automatically:

  1. POINT-IN-TIME Recompute features using only bars <= t and assert row t is unchanged. Catches any use of future data in feature construction, including the subtle ones (a global mean, a backfill, an accidental negative shift).

  2. LABEL/FEATURE Check whether any single feature predicts the label far CORRELATION too well on its own. ret_1 versus a same-bar label scores ~1.0 here, which is precisely the baseline's bug.

  3. NULL DATA Run the whole pipeline on a random walk. There is no edge by construction, so a positive result after costs proves a bug — in the engine, the split, or the labels.

  4. SHUFFLED LABEL Retrain with labels randomly permuted. Out-of-sample accuracy must collapse to the base rate. If it does not, information is reaching the model through a side channel.

  5. SURVIVORSHIP A multi-symbol universe spanning a period that killed assets, containing none of them, was filtered by survival. Reports n/a rather than PASS on a single-symbol config — it has nothing to test there, and a green tick would claim a risk was ruled out when it was never examined.

Run nullres audit before believing any result. It takes a minute and it has a much better record than intuition.

check_point_in_time

check_point_in_time(bars: DataFrame, builder=build_features, probes: int = 12, tol: float = 1e-09, seed: int = 0) -> Check

Recompute features on a truncated history and compare the final row.

If build_features(bars[:t+1]).iloc[t] differs from build_features(bars).iloc[t], then the full-history version used data from after bar t. There is no way for that to be legitimate.

The probes are random but seeded, so the check is reproducible — and so it inspects the same bars on every run. That is the trade: a leak that only manifests in one regime is found only if a probe lands in it, which is why the count is more than a token few. Raise probes when adding a feature whose behaviour is regime-dependent.

Source code in nullres/audit.py
def check_point_in_time(bars: pd.DataFrame, builder=build_features,
                        probes: int = 12, tol: float = 1e-9,
                        seed: int = 0) -> Check:
    """Recompute features on a truncated history and compare the final row.

    If `build_features(bars[:t+1]).iloc[t]` differs from
    `build_features(bars).iloc[t]`, then the full-history version used data from
    after bar t. There is no way for that to be legitimate.

    The probes are random but seeded, so the check is reproducible — and so it
    inspects the same bars on every run. That is the trade: a leak that only
    manifests in one regime is found only if a probe lands in it, which is why
    the count is more than a token few. Raise `probes` when adding a feature
    whose behaviour is regime-dependent.
    """
    if len(bars) < 50:
        return Check("point-in-time features", False,
                     f"only {len(bars)} bars — too few to probe meaningfully")

    full = builder(bars)
    rng = np.random.default_rng(seed)
    # Probe past the warmup so rolling windows have filled, but never past the
    # end of a short series: a fixed 500-bar floor crashed on anything smaller.
    warmup = min(500, len(bars) // 2)
    lo = min(max(warmup, len(bars) // 10), len(bars) - 1)
    offenders: dict[str, float] = {}

    for t in rng.integers(lo, len(bars), size=probes):
        t = int(t)
        truncated = builder(bars.iloc[: t + 1])
        a = truncated.iloc[-1]
        b = full.iloc[t]
        for col in full.columns:
            av, bv = a[col], b[col]
            if pd.isna(av) and pd.isna(bv):
                continue
            denom = max(abs(bv), 1.0) if pd.notna(bv) else 1.0
            diff = abs(av - bv) / denom if pd.notna(av) and pd.notna(bv) else np.inf
            if diff > tol:
                offenders[col] = max(offenders.get(col, 0.0), float(diff))

    if offenders:
        worst = sorted(offenders.items(), key=lambda kv: -kv[1])[:5]
        listing = ", ".join(f"{c} ({d:.2e})" for c, d in worst)
        return Check(
            "point-in-time features", False,
            f"{len(offenders)} feature(s) change when future bars are removed: {listing}",
        )
    return Check(
        "point-in-time features", True,
        f"all {full.shape[1]} features identical across {probes} truncation probes",
    )

check_label_leakage

check_label_leakage(features: DataFrame, y: Series, auc_limit: float = 0.65) -> Check

Flag any single feature that separates the label suspiciously well.

A lone technical indicator with AUC > 0.65 on a directional label is not a discovery. On real financial data, single-feature AUCs live in 0.50-0.55.

Source code in nullres/audit.py
def check_label_leakage(features: pd.DataFrame, y: pd.Series,
                        auc_limit: float = 0.65) -> Check:
    """Flag any single feature that separates the label suspiciously well.

    A lone technical indicator with AUC > 0.65 on a directional label is not a
    discovery. On real financial data, single-feature AUCs live in 0.50-0.55.
    """
    from sklearn.metrics import roc_auc_score

    ok = y.notna()
    y_ok = y[ok].astype(int)
    if y_ok.nunique() < 2:
        return Check("single-feature AUC", False, "label has only one class")

    scores: dict[str, float] = {}
    for col in features.columns:
        v = features.loc[ok, col]
        valid = v.notna()
        if valid.sum() < 100 or y_ok[valid].nunique() < 2:
            continue
        auc = roc_auc_score(y_ok[valid], v[valid])
        scores[col] = max(auc, 1 - auc)      # direction-agnostic

    if not scores:
        return Check("single-feature AUC", False, "no feature had enough valid rows")

    worst = sorted(scores.items(), key=lambda kv: -kv[1])
    top = ", ".join(f"{c}={a:.3f}" for c, a in worst[:3])
    flagged = [c for c, a in worst if a > auc_limit]
    if flagged:
        return Check(
            "single-feature AUC", False,
            f"{len(flagged)} feature(s) exceed AUC {auc_limit}: {top}  "
            f"-- this is what a leaked label looks like",
        )
    return Check("single-feature AUC", True, f"max AUC {worst[0][1]:.3f} ({top})")

check_null_data

check_null_data(run_pipeline, cfg, sharpe_limit: float = 0.5) -> Check

The pipeline must find nothing on a random walk.

run_pipeline(cfg) is injected to avoid a circular import.

Source code in nullres/audit.py
def check_null_data(run_pipeline, cfg, sharpe_limit: float = 0.5) -> Check:
    """The pipeline must find nothing on a random walk.

    `run_pipeline(cfg)` is injected to avoid a circular import.
    """
    import copy

    null_cfg = copy.deepcopy(cfg)
    null_cfg.data.source = "synthetic"
    null_cfg.name = f"{cfg.name}-null"

    results = run_pipeline(null_cfg, verbose=False)
    offenders = {
        name: m["sharpe"] for name, m in results.items()
        if name != "buy_hold" and m["sharpe"] > sharpe_limit
    }
    if offenders:
        listing = ", ".join(f"{n} sharpe {s:.2f}" for n, s in offenders.items())
        return Check(
            "null (random-walk) data", False,
            f"found 'edge' where none exists: {listing} -- there is a bug in the "
            f"engine, the split, or the labels",
        )

    best = max((m["sharpe"] for n, m in results.items() if n != "buy_hold"), default=0.0)
    return Check(
        "null (random-walk) data", True,
        f"no strategy beat sharpe {sharpe_limit} on synthetic data (best {best:.2f})",
    )

check_survivorship

check_survivorship(symbols: Iterable[str], delisted: Mapping[str, Any] | None, point_in_time: Iterable[str] | None = None, hardcoded: bool = False, last_bar=None, sample_end=None, grace_days: int = 60) -> Check

Does this universe contain assets that died?

Backtesting a universe chosen from what is liquid today is a test of "things that survived", and it will produce a beautiful, meaningless result. The catalogue used to call this undetectable. It is not — not fully, but the dominant failure mode is mechanical:

A multi-symbol universe spanning a period that killed assets, which contains none of them, was filtered by survival.

What this CANNOT see is whether you picked the winners among the survivors. That is hindsight, and it stays yours to avoid — see the catalogue's entry 7.

Parameters:

Name Type Description Default
symbols Iterable[str]

the universe actually traded.

required
delisted Mapping[str, Any] | None

symbols whose data stops before the sample ends.

required
point_in_time Iterable[str] | None

optionally, the universe as enumerated at the sample start. Lets the check measure how much of the graveyard was dropped.

None
hardcoded bool

True when the universe was a literal list rather than enumerated from the archive as of a date.

False
Source code in nullres/audit.py
def check_survivorship(symbols: Iterable[str],
                       delisted: Mapping[str, Any] | None,
                       point_in_time: Iterable[str] | None = None,
                       hardcoded: bool = False, last_bar=None,
                       sample_end=None, grace_days: int = 60) -> Check:
    """Does this universe contain assets that died?

    Backtesting a universe chosen from what is liquid today is a test of
    "things that survived", and it will produce a beautiful, meaningless
    result. The catalogue used to call this undetectable. It is not — not
    fully, but the dominant failure mode is mechanical:

      A multi-symbol universe spanning a period that killed assets, which
      contains none of them, was filtered by survival.

    What this CANNOT see is whether you picked the winners among the survivors.
    That is hindsight, and it stays yours to avoid — see the catalogue's entry 7.

    Args:
        symbols: the universe actually traded.
        delisted: symbols whose data stops before the sample ends.
        point_in_time: optionally, the universe as enumerated at the sample
            start. Lets the check measure how much of the graveyard was dropped.
        hardcoded: True when the universe was a literal list rather than
            enumerated from the archive as of a date.
    """
    symbols = list(symbols)
    dead = set(delisted or ())

    if len(symbols) < 2:
        name = symbols[0] if symbols else "none"
        # Universe survivorship is meaningless for one asset, but a narrower
        # question is not: did THIS instrument trade to the end of the sample?
        # A symbol that quietly stopped is a backtest on something that no
        # longer exists, and nothing else here would notice. Reporting a bare
        # `n/a` skipped a check that was available all along.
        if last_bar is not None and sample_end is not None:
            gap = pd.Timestamp(sample_end) - pd.Timestamp(last_bar)
            if gap > pd.Timedelta(days=grace_days):
                return Check(
                    "survivorship", False,
                    f"{name} last traded {pd.Timestamp(last_bar):%Y-%m-%d}, "
                    f"{gap.days} days before the sample ends "
                    f"({pd.Timestamp(sample_end):%Y-%m-%d}). This instrument "
                    f"stopped trading during the backtest",
                )
            return Check(
                "survivorship", True,
                f"{name} traded through to {pd.Timestamp(last_bar):%Y-%m-%d}, "
                f"the end of the sample. Universe survivorship cannot apply to a "
                f"single asset and is NOT ruled out for any multi-asset "
                f"extension of this work",
            )
        return Check(
            "survivorship", True,
            f"single-symbol backtest ({name}) — survivorship does not apply, "
            f"and has NOT been ruled out for any multi-asset extension of this "
            f"work",
            applicable=False,
        )

    if not dead:
        detail = (
            f"none of the {len(symbols)} symbols stopped trading during the "
            f"sample. Either the period genuinely killed nothing, or the "
            f"universe was chosen from survivors"
        )
        if hardcoded:
            detail += " — and this universe is a hardcoded list, which is how "\
                      "that happens"
        return Check("survivorship", False, detail)

    share = len(dead) / len(symbols)
    detail = (f"{len(dead)} of {len(symbols)} symbols ({share:.0%}) delisted "
              f"during the sample and were held to the end: "
              f"{', '.join(sorted(dead)[:5])}"
              f"{'...' if len(dead) > 5 else ''}")

    if point_in_time:
        missing = set(point_in_time) - set(symbols)
        detail += (f". Universe covers {len(symbols)}/{len(point_in_time)} of the "
                   f"symbols trading at the sample start")
        if missing and len(missing) > len(point_in_time) * 0.5:
            return Check(
                "survivorship", False,
                detail + f" — {len(missing)} were excluded, which needs a reason "
                         f"that is not 'they are not around any more'",
            )
    return Check("survivorship", True, detail)

check_shuffled_label

check_shuffled_label(X: DataFrame, y: Series, t_end: ndarray, split_cfg, model_cfg, tol: float = 0.02, seed: int = 0) -> Check

Permuted labels must be unlearnable.

Any accuracy above the base rate here means the model is reaching the target through something other than the features it was given.

Source code in nullres/audit.py
def check_shuffled_label(X: pd.DataFrame, y: pd.Series, t_end: np.ndarray,
                         split_cfg, model_cfg, tol: float = 0.02,
                         seed: int = 0) -> Check:
    """Permuted labels must be unlearnable.

    Any accuracy above the base rate here means the model is reaching the
    target through something other than the features it was given.
    """
    from nullres.models.classifier import fit_predict_walk_forward

    rng = np.random.default_rng(seed)
    shuffled = pd.Series(rng.permutation(y.to_numpy()), index=y.index)

    proba, reports = fit_predict_walk_forward(
        X, shuffled, t_end, split_cfg, model_cfg, verbose=False
    )
    scored = proba.notna() & shuffled.notna()
    acc = float(((proba[scored] > 0.5) == (shuffled[scored] > 0.5)).mean())
    base = float(max(shuffled[scored].mean(), 1 - shuffled[scored].mean()))

    if acc > base + tol:
        return Check(
            "shuffled-label control", False,
            f"accuracy {acc:.4f} exceeds base rate {base:.4f} on RANDOM labels",
        )
    return Check(
        "shuffled-label control", True,
        f"accuracy {acc:.4f} vs base rate {base:.4f} — no learnable signal, as expected",
    )

Robustness battery

robustness

Falsification tests for a strategy that looked good once.

A single backtest number is a hypothesis, not a finding. Before it earns any more of your time it has to survive three attempts to kill it:

NEIGHBOURHOOD Do nearby parameter values also work? A real effect degrades smoothly. If only one cell in a grid is positive, you did not find an edge — you found the cell that happened to fit.

STABILITY Does it work in every year, or did one spectacular period carry the average? A Sharpe of 0.5 built from 2021 and nothing else is a bet that 2021 recurs.

TRANSFER Does it work on other symbols? A rule that describes market structure should generalise. A rule that only works on the asset you developed it on describes that asset's history.

Passing all three does not make a strategy real — the only test that does is forward paper trading.

What these tests can and cannot resolve. Two of the three rest on very few observations: five years, four symbols. A gate reading "beat the benchmark in 60% of periods" sounds demanding, but at n=5 it means three of five, and a strategy that is genuinely a coin flip against buy & hold clears it half the time. At n=4 symbols a coin flip clears it 31% of the time. That is not a detail — it means a bare count cannot tell "this is worse than holding" from "there is not enough evidence here", and for a long time this module reported both as KILLED.

So each count gate is now read alongside the magnitude of the shortfall, and verdict has three outcomes rather than two. A gate FAILS only when the count goes against it AND the mean excess is distinguishable from zero. A count that fails alone yields INCONCLUSIVE, and every note prints how often the gate would have fired by chance.

The decision rule stays aggressive on purpose — a strategy earns SURVIVED only by clearing all three — because a false kill costs one idea and a false survival costs months. That is a decision about which error to prefer, not a claim that the thresholds are statistically strong. They are not.

grid_for

grid_for(strategy: str) -> tuple[dict, str]

Return (grid, kind) where kind is 'params' or 'sizing'.

Source code in nullres/robustness.py
def grid_for(strategy: str) -> tuple[dict, str]:
    """Return (grid, kind) where kind is 'params' or 'sizing'."""
    if strategy in SIZING_GRIDS:
        return SIZING_GRIDS[strategy], "sizing"
    if strategy in DEFAULT_GRIDS:
        return DEFAULT_GRIDS[strategy], "params"
    raise ConfigError(f"no parameter grid defined for {strategy!r}")

parameter_neighbourhood

parameter_neighbourhood(cfg, strategy: str, ctx=None, grid=None) -> DataFrame

Sharpe across a grid of parameter values, on one prepared context.

Rule strategies read only the bars, so the context (features, labels, splits) is identical across the grid and is computed once.

Source code in nullres/robustness.py
def parameter_neighbourhood(cfg, strategy: str, ctx=None, grid=None) -> pd.DataFrame:
    """Sharpe across a grid of parameter values, on one prepared context.

    Rule strategies read only the bars, so the context (features, labels,
    splits) is identical across the grid and is computed once.
    """
    default_grid, kind = grid_for(strategy)
    grid = grid or default_grid
    ctx = ctx or prepare(cfg, verbose=False)
    original_cfg = ctx.cfg

    keys = list(grid)
    rows = []
    try:
        for values in product(*(grid[k] for k in keys)):
            combo = dict(zip(keys, values))
            if not _valid(strategy, combo):
                continue

            if kind == "sizing":
                trial = copy.deepcopy(cfg)
                for key, value in combo.items():
                    setattr(trial.sizing, key, value)
                # Keep the short band symmetric with the long one, so the grid
                # varies conviction rather than quietly introducing a long bias.
                if "long_entry" in combo:
                    trial.sizing.short_entry = round(1.0 - combo["long_entry"], 4)
                ctx.cfg = trial
                strategy_obj = build_strategy(strategy)
            else:
                trial = cfg
                strategy_obj = build_strategy(strategy, combo)

            positions = strategy_obj.positions(ctx)
            result = backtest(ctx.bars, positions, trial.cost)
            # Same out-of-sample restriction the pipeline applies. Without it
            # the grid is measured over the full frame while `benchmark_sharpe`
            # comes from `period_stability`, which masks — so `verdict` compared
            # a deflated grid against an undeflated bar.
            metrics = summarize(result, trial.data.bars_per_year,
                                mask=ctx.oos_mask)
            rows.append({**combo,
                         "sharpe": metrics["sharpe"],
                         "total_return": metrics["total_return"],
                         "n_trades": metrics["n_trades"]})
    finally:
        ctx.cfg = original_cfg
    return pd.DataFrame(rows)

period_stability

period_stability(cfg, strategy: str, ctx=None, freq: str = 'YE') -> DataFrame

Per-year performance, alongside buy & hold over the same periods.

The benchmark column is not decoration. "Lost money in 2022" and "lost less than half what holding lost in 2022" are opposite findings, and the bare per-year Sharpe cannot distinguish them. For a long-only trend filter evaluated over a historic bull market, this column is the whole argument.

Source code in nullres/robustness.py
def period_stability(cfg, strategy: str, ctx=None, freq: str = "YE") -> pd.DataFrame:
    """Per-year performance, alongside buy & hold over the same periods.

    The benchmark column is not decoration. "Lost money in 2022" and "lost less
    than half what holding lost in 2022" are opposite findings, and the bare
    per-year Sharpe cannot distinguish them. For a long-only trend filter
    evaluated over a historic bull market, this column is the whole argument.
    """
    ctx = ctx or prepare(cfg, verbose=False)

    positions = build_strategy(strategy, cfg.params.get(strategy)).positions(ctx)
    strat = by_period(backtest(ctx.bars, positions, cfg.cost),
                      cfg.data.bars_per_year, mask=ctx.oos_mask, freq=freq)

    hold_pos = build_strategy("buy_hold").positions(ctx)
    hold = by_period(backtest(ctx.bars, hold_pos, cfg.cost),
                     cfg.data.bars_per_year, mask=ctx.oos_mask, freq=freq)

    # Join onto the BENCHMARK's periods, not the strategy's. `by_period` drops a
    # period with no variance, which is precisely a year the strategy sat flat —
    # so a left join silently removed it from the stability test rather than
    # scoring it. Live on the 1d config: sma_cross and donchian are flat through
    # 2022 and that year simply disappeared, taking with it the year a flat book
    # most obviously beats a -65% benchmark. The bias is not consistently in the
    # strategy's favour, which is worse than if it were: whether dropping a year
    # helps or hurts depends on what the benchmark happened to do in it.
    #
    # A flat period is a real observation. The strategy returned nothing, so its
    # Sharpe is 0 and its excess is minus the benchmark's.
    merged = strat.merge(
        hold[["period", "total_return", "sharpe"]],
        on="period", how="right", suffixes=("", "_hold"),
    ).sort_values("period", ignore_index=True)
    for column, filler in (("sharpe", 0.0), ("total_return", 0.0),
                           ("n_trades", 0), ("bars", 0)):
        if column in merged:
            merged[column] = merged[column].fillna(filler)

    merged["excess_sharpe"] = merged["sharpe"] - merged["sharpe_hold"]
    return merged

hold_sharpe

hold_sharpe(cfg, ctx) -> float

Buy & hold's Sharpe, on the same window AND the same statistic as the grid.

verdict reports what fraction of the neighbourhood beats buy & hold, and that comparison is only meaningful if both sides are the same measurement. They were not: the grid reports a full-window Sharpe while the benchmark was taken as the mean of period_stability's per-year Sharpes. Averaging annual Sharpes is a different statistic — on the 4h config it gives 0.53 against a full-window 0.38, overstating the bar by 40%.

Source code in nullres/robustness.py
def hold_sharpe(cfg, ctx) -> float:
    """Buy & hold's Sharpe, on the same window AND the same statistic as the grid.

    `verdict` reports what fraction of the neighbourhood beats buy & hold, and
    that comparison is only meaningful if both sides are the same measurement.
    They were not: the grid reports a full-window Sharpe while the benchmark was
    taken as the mean of `period_stability`'s per-year Sharpes. Averaging annual
    Sharpes is a different statistic — on the 4h config it gives 0.53 against a
    full-window 0.38, overstating the bar by 40%.
    """
    positions = build_strategy("buy_hold").positions(ctx)
    result = backtest(ctx.bars, positions, cfg.cost)
    return float(summarize(result, cfg.data.bars_per_year,
                           mask=ctx.oos_mask)["sharpe"])

cross_symbol

cross_symbol(cfg, strategy: str, symbols: list[str], start: str | None = None) -> DataFrame

The same strategy and parameters, on other instruments.

Each symbol needs its own data, features and splits, so this is the expensive test — and the most informative one. It is what killed both previous candidates.

start overrides the config's start date for EVERY symbol including the reference one. Auxiliary archives begin at different dates per symbol (BTCUSDT open-interest metrics start 2020-09, everything else 2021-12), and letting each symbol use its own maximum range would compare different eras and call the difference "transfer".

Source code in nullres/robustness.py
def cross_symbol(cfg, strategy: str, symbols: list[str],
                 start: str | None = None) -> pd.DataFrame:
    """The same strategy and parameters, on other instruments.

    Each symbol needs its own data, features and splits, so this is the
    expensive test — and the most informative one. It is what killed both
    previous candidates.

    `start` overrides the config's start date for EVERY symbol including the
    reference one. Auxiliary archives begin at different dates per symbol
    (BTCUSDT open-interest metrics start 2020-09, everything else 2021-12), and
    letting each symbol use its own maximum range would compare different eras
    and call the difference "transfer".
    """
    rows = []
    for symbol in symbols:
        trial = copy.deepcopy(cfg)
        trial.data.symbol = symbol
        trial.strategies = [strategy]
        if start:
            trial.data.start = start
        try:
            results = run_pipeline(trial, verbose=False)
        except NullresError as exc:
            # A symbol with no archive, or too few bars to split, is a fact
            # about that symbol — not a reason to abandon the transfer test.
            # This used to read `except (SystemExit, ValueError)`: the data
            # layer signalled a missing symbol by raising SystemExit, so the
            # loop had to catch the interpreter's shutdown request to keep
            # going. It also missed the `no fold produced predictions` case,
            # which was a RuntimeError and took the whole battery down.
            rows.append({"symbol": symbol, "sharpe": np.nan,
                         "total_return": np.nan, "n_trades": 0,
                         "note": str(exc)[:60]})
            continue
        m = results[strategy]
        bh = results["buy_hold"]
        rows.append({
            "symbol": symbol,
            "sharpe": m["sharpe"],
            "total_return": m["total_return"],
            "n_trades": m["n_trades"],
            "vs_hold": m["sharpe"] - bh["sharpe"],
            "note": "",
        })
    return pd.DataFrame(rows)

sign_flip_pairs

sign_flip_pairs(grid: DataFrame, keys: list[str]) -> int

How many adjacent cell pairs sign_flip_rate measured over.

The rate alone cannot say whether it is distinguishable from chance; that needs the denominator too.

Source code in nullres/robustness.py
def sign_flip_pairs(grid: pd.DataFrame, keys: list[str]) -> int:
    """How many adjacent cell pairs `sign_flip_rate` measured over.

    The rate alone cannot say whether it is distinguishable from chance; that
    needs the denominator too.
    """
    if len(keys) != 2 or grid.empty:
        return 0
    values = grid.pivot(index=keys[0], columns=keys[1], values="sharpe").to_numpy(
        dtype="float64")
    total = 0
    for axis in (0, 1):
        a = np.moveaxis(values, axis, 0)
        for i in range(len(a) - 1):
            pair = np.vstack([a[i], a[i + 1]])
            total += int(np.isfinite(pair).all(axis=0).sum())
    return total

sign_flip_rate

sign_flip_rate(grid: DataFrame, keys: list[str]) -> float

Fraction of adjacent grid cells whose Sharpe changes sign.

This is the measure the docs have always claimed to care about — "a real effect degrades SMOOTHLY; an isolated spike is a fitting artefact" — and which nothing actually computed. Counting positive cells cannot see the difference between a coherent positive region and a checkerboard: a grid reading +0.8 / -0.2 / +0.9 / -0.7 scores 50% positive either way.

Neighbours are cells one step apart along a single axis. A rate near 0 means a smooth surface; near 0.5 means the sign carries no information.

Source code in nullres/robustness.py
def sign_flip_rate(grid: pd.DataFrame, keys: list[str]) -> float:
    """Fraction of adjacent grid cells whose Sharpe changes sign.

    This is the measure the docs have always claimed to care about — "a real
    effect degrades SMOOTHLY; an isolated spike is a fitting artefact" — and
    which nothing actually computed. Counting positive cells cannot see the
    difference between a coherent positive region and a checkerboard: a grid
    reading +0.8 / -0.2 / +0.9 / -0.7 scores 50% positive either way.

    Neighbours are cells one step apart along a single axis. A rate near 0
    means a smooth surface; near 0.5 means the sign carries no information.
    """
    if len(keys) != 2 or grid.empty:
        return float("nan")
    table = grid.pivot(index=keys[0], columns=keys[1], values="sharpe")
    values = table.to_numpy(dtype="float64")

    flips = total = 0
    for axis in (0, 1):
        a = np.moveaxis(values, axis, 0)
        for i in range(len(a) - 1):
            pair = np.vstack([a[i], a[i + 1]])
            valid = np.isfinite(pair).all(axis=0)
            flips += int((np.sign(pair[0][valid]) != np.sign(pair[1][valid])).sum())
            total += int(valid.sum())
    return flips / total if total else float("nan")

count_gate_power

count_gate_power(n: int, threshold: float = 0.6, p_null: float = 0.5) -> float

How often a strategy exactly as good as the benchmark clears a count gate.

"Beat the benchmark in 60% of periods" sounds demanding until you count the periods. With five years it means three of five, and a strategy that is genuinely a coin flip against buy & hold clears that half the time. With four symbols it is three of four, which a coin flip clears 31% of the time.

A gate this noisy cannot carry a verdict by itself, so the number is printed beside every count so the reader knows what the count is worth.

Source code in nullres/robustness.py
def count_gate_power(n: int, threshold: float = 0.6, p_null: float = 0.5) -> float:
    """How often a strategy exactly as good as the benchmark clears a count gate.

    "Beat the benchmark in 60% of periods" sounds demanding until you count the
    periods. With five years it means three of five, and a strategy that is
    genuinely a coin flip against buy & hold clears that **half the time**. With
    four symbols it is three of four, which a coin flip clears 31% of the time.

    A gate this noisy cannot carry a verdict by itself, so the number is printed
    beside every count so the reader knows what the count is worth.
    """
    if n < 1:
        return float("nan")
    need = int(np.ceil(threshold * n))
    return float(1 - sps.binom.cdf(need - 1, n, p_null))

excess_magnitude

excess_magnitude(values) -> tuple[float, float]

(mean, p) for "is the mean excess distinguishable from zero".

The count gates throw away magnitude, and that loses real information: on the 4h config donchian and mean_reversion both beat hold in 40% of years and score identically, while their mean excess Sharpes are -0.04 and -1.13. One is indistinguishable from holding; the other is far worse. A test on the magnitude separates them; counting signs cannot.

It is not a replacement for the count, because it is less decisive on noisy series — mean_reversion's -1.13 carries p=0.31 across five volatile years. Neither statistic dominates, so both are reported.

Source code in nullres/robustness.py
def excess_magnitude(values) -> tuple[float, float]:
    """(mean, p) for "is the mean excess distinguishable from zero".

    The count gates throw away magnitude, and that loses real information: on
    the 4h config `donchian` and `mean_reversion` both beat hold in 40% of years
    and score identically, while their mean excess Sharpes are -0.04 and -1.13.
    One is indistinguishable from holding; the other is far worse. A test on the
    magnitude separates them; counting signs cannot.

    It is not a replacement for the count, because it is less decisive on noisy
    series — `mean_reversion`'s -1.13 carries p=0.31 across five volatile years.
    Neither statistic dominates, so both are reported.
    """
    clean = np.asarray(values, dtype=float)
    clean = clean[np.isfinite(clean)]
    if len(clean) < 2 or np.allclose(clean, clean[0]):
        return (float(clean.mean()) if len(clean) else float("nan"), float("nan"))
    _, p = sps.ttest_1samp(clean, 0.0)
    return float(clean.mean()), float(p)

verdict

verdict(neighbourhood: DataFrame, stability: DataFrame, transfer: DataFrame, benchmark_sharpe: float | None = None, flip_rate: float | None = None, flip_pairs: int | None = None) -> tuple[str, list[str]]

Turn the three tables into KILLED, SURVIVED or INCONCLUSIVE, with reasons.

Three states, not two, because two states forced a claim the evidence does not support. The count gates rest on four or five Bernoulli draws: at n=5 a strategy genuinely equal to buy & hold fails the stability gate half the time. Reporting that as KILLED dressed a coin flip as a finding, and the verdict then propagated into the ledger and warned future runs off the config.

So a gate now FAILS only on decisive evidence — the count went against it AND the magnitude of the shortfall is distinguishable from zero. A count that fails on its own returns WEAK, and the run comes out INCONCLUSIVE.

The decision rule stays deliberately aggressive: for research triage a false kill is cheap and a false survival is expensive, so a strategy has to earn SURVIVED by clearing every gate. That is a decision-theoretic stance, not a claim that the thresholds are statistically demanding. They are not, and each note now says so.

Source code in nullres/robustness.py
def verdict(neighbourhood: pd.DataFrame, stability: pd.DataFrame,
            transfer: pd.DataFrame,
            benchmark_sharpe: float | None = None,
            flip_rate: float | None = None,
            flip_pairs: int | None = None) -> tuple[str, list[str]]:
    """Turn the three tables into KILLED, SURVIVED or INCONCLUSIVE, with reasons.

    Three states, not two, because two states forced a claim the evidence does
    not support. The count gates rest on four or five Bernoulli draws: at n=5 a
    strategy genuinely equal to buy & hold fails the stability gate half the
    time. Reporting that as KILLED dressed a coin flip as a finding, and the
    verdict then propagated into the ledger and warned future runs off the
    config.

    So a gate now FAILS only on decisive evidence — the count went against it
    AND the magnitude of the shortfall is distinguishable from zero. A count
    that fails on its own returns WEAK, and the run comes out INCONCLUSIVE.

    The decision rule stays deliberately aggressive: for research triage a false
    kill is cheap and a false survival is expensive, so a strategy has to earn
    SURVIVED by clearing every gate. That is a decision-theoretic stance, not a
    claim that the thresholds are statistically demanding. They are not, and
    each note now says so.
    """
    notes, outcomes = [], []

    frac = float((neighbourhood["sharpe"] > 0).mean())
    median = float(neighbourhood["sharpe"].median())
    best = float(neighbourhood["sharpe"].max())
    # A grid is a sensitivity surface, not a sample. Its cells are neighbouring
    # parameter values evaluated on the SAME bars, so they are heavily
    # correlated and a t-test over them would be nonsense — the effective count
    # is unknowable and far below the nominal one. That rules out the magnitude
    # test the other two gates use.
    #
    # Two things can still be said without assuming independence. The median is
    # a statement about the typical parameter choice, not a sampling estimate:
    # if it loses money, the surface loses money. And the ARRANGEMENT of signs
    # is testable against random placement regardless of how correlated the
    # values are. Those two carry this gate; the bare count does not, and at a
    # 60% threshold on ~20 cells it fires a quarter of the time by chance.
    if median <= 0:
        outcomes.append("FAIL")
        notes.append(
            f"NEIGHBOURHOOD FAIL: the median of {len(neighbourhood)} parameter "
            f"combinations is {median:.2f} (best {best:.2f}, {frac:.0%} positive). "
            f"The typical choice of parameters loses money; only the lucky ones "
            f"do not."
        )
    elif _is_noise_field(frac, flip_rate, flip_pairs):
        # Counting positive cells cannot distinguish a coherent region from a
        # checkerboard. Compare the observed flip rate against what random
        # placement of the SAME number of positive cells would produce:
        # 2p(1-p). Matching that means the arrangement carries no information.
        expected = 2 * frac * (1 - frac)
        outcomes.append("FAIL")
        notes.append(
            f"NEIGHBOURHOOD FAIL: {frac:.0%} of combinations are positive, but the "
            f"sign flips across {flip_rate:.0%} of adjacent cells versus "
            f"{expected:.0%} expected from random placement. The grid is no "
            f"smoother than chance — a noise field with a lucky maximum "
            f"({best:.2f}), not a region of edge."
        )
    else:
        # "Positive" is a low bar. Without this clause the note reads as a pass
        # even when the entire grid sits below the thing you'd have done anyway.
        context = ""
        if benchmark_sharpe is not None:
            beat = float((neighbourhood["sharpe"] > benchmark_sharpe).mean())
            context = (f", but only {beat:.0%} beat buy & hold "
                       f"({benchmark_sharpe:.2f})")
        smooth = f", sign flips across {flip_rate:.0%} of adjacent cells" \
            if flip_rate is not None else ""

        # The reported cell sitting far above its own neighbours is the shape a
        # fitted parameter makes. This needs no independence assumption — it
        # asks where the headline sits inside its own surface, not whether the
        # surface is a sample of anything.
        spread = float(neighbourhood["sharpe"].quantile(0.75)
                       - neighbourhood["sharpe"].quantile(0.25))
        outlier = spread > 0 and (best - median) > 3 * spread
        if frac < 0.6 or outlier:
            outcomes.append("WEAK")
            why = (f"its best cell ({best:.2f}) sits {(best - median) / spread:.1f} "
                   f"interquartile ranges above the median"
                   if outlier else
                   f"only {frac:.0%} of cells are positive")
            notes.append(
                f"NEIGHBOURHOOD INCONCLUSIVE: the arrangement is smoother than "
                f"chance and the median is {median:.2f}, but {why}. At "
                f"{len(neighbourhood)} correlated cells the positive count is "
                f"weak evidence either way — a 60% threshold fires on roughly a "
                f"quarter of grids with no edge at all{context}{smooth}."
            )
        else:
            outcomes.append("PASS")
            notes.append(
                f"neighbourhood ok: {frac:.0%} of combinations positive, "
                f"median sharpe {median:.2f}{context}{smooth}"
            )

    # The test that matters is beating the benchmark, not being positive. A
    # long-only filter over a bull market is positive in most years by
    # construction; that says nothing about whether the rule adds anything.
    if stability.empty:
        outcomes.append("FAIL")
        notes.append("STABILITY FAIL: no periods with activity to evaluate")
    else:
        outcome, note = _count_gate(
            stability["excess_sharpe"].to_numpy(dtype=float), "STABILITY", "years"
        )
        outcomes.append(outcome)
        notes.append(note)

    # Same correction: "positive" is a low bar when every asset in the sample
    # rose. BNBUSDT scored Sharpe 0.03 — positive, and 0.71 WORSE than simply
    # holding it. Judge against the alternative.
    scored = transfer.dropna(subset=["sharpe"])
    if scored.empty:
        outcomes.append("FAIL")
        notes.append("TRANSFER FAIL: no other symbol produced a result")
    else:
        column = "vs_hold" if "vs_hold" in scored.columns else "sharpe"
        outcome, note = _count_gate(
            scored[column].to_numpy(dtype=float), "TRANSFER", "symbols"
        )
        outcomes.append(outcome)
        notes.append(note)

    if "FAIL" in outcomes:
        result = KILLED
    elif "WEAK" in outcomes:
        result = INCONCLUSIVE
        notes.append(
            "INCONCLUSIVE means the battery ran and could not separate this "
            "strategy from one exactly as good as the benchmark — not that it "
            "looks promising. Deciding what an underpowered result means is "
            "the human's job; see docs/05-graveyard.md."
        )
    else:
        result = SURVIVED
    return result, notes

pivot_grid

pivot_grid(df: DataFrame, keys: list[str]) -> str

Render a two-parameter grid as a Sharpe matrix.

Source code in nullres/robustness.py
def pivot_grid(df: pd.DataFrame, keys: list[str]) -> str:
    """Render a two-parameter grid as a Sharpe matrix."""
    if len(keys) != 2:
        return df.to_string(index=False)
    table = df.pivot(index=keys[0], columns=keys[1], values="sharpe")
    head = f"{keys[0]:>6} \\ {keys[1]:<4}" + "".join(f"{c:>9}" for c in table.columns)
    lines = [head, "-" * len(head)]
    for idx, row in table.iterrows():
        cells = "".join("      —  " if pd.isna(v) else f"{v:>9.2f}" for v in row)
        lines.append(f"{idx:>13}" + cells)
    return "\n".join(lines)

Cross-sectional panels

crosssec

Cross-sectional (relative-value) research on a panel of symbols.

Everything else in this repo asks "will BTC go up?". That question has to clear a 24bps round trip out of a single asset's own move, and six lines of attack have now died against that wall (docs/05-graveyard.md).

This asks a different question: "which of these assets will outperform the others?" It is easier in three specific ways.

  1. The market move cancels. A dollar-neutral book is not betting on crypto going up, so it does not need to out-predict a 60% annualised drift.
  2. The label is balanced by construction — half the universe beats the median at every timestamp, in every regime.
  3. Errors are relative. Being wrong about BTC and wrong about ETH in the same direction costs nothing; only the ranking matters.

What it costs you is a new failure mode: survivorship bias. Choosing a universe by looking at what is liquid today is a test of "assets that survived", and it will produce a beautiful, meaningless equity curve. The universe here is fixed as of 2021-12 and includes LUNAUSDT, which went to zero in May 2022 and was delisted. If your cross-sectional backtest cannot lose money on LUNA, it is not measuring anything.

Long/short also requires PERPETUAL FUTURES — you cannot short spot — so this module uses USD-M perp bars and charges the 8-hourly funding rate on every position held.

Panel dataclass

Panel(features: DataFrame, y: Series, ret_next: DataFrame, funding: DataFrame, times: DatetimeIndex, horizon: int, symbols: list[str] = list(), delisted: dict[str, Timestamp] = dict())

A tidy panel: MultiIndex (ts, symbol) features, plus per-bar returns.

load_panel

load_panel(cfg, symbols: list[str] | None = None, verbose: bool = True, top_n: int | None = None, screen_window: int = 180) -> Panel

Load symbols, screen for liquidity, build features, stack into a panel.

Parameters:

Name Type Description Default
top_n int | None

keep only the top-N symbols by TRAILING dollar volume at each bar. Without this a wide universe ranks BTC against coins that traded $50k a day, which is a ranking you could not act on. Screening on full-sample volume instead would be lookahead — it selects the coins that went on to matter.

None
Source code in nullres/crosssec.py
def load_panel(cfg, symbols: list[str] | None = None, verbose: bool = True,
               top_n: int | None = None, screen_window: int = 180) -> Panel:
    """Load symbols, screen for liquidity, build features, stack into a panel.

    Args:
        top_n: keep only the top-N symbols by TRAILING dollar volume at each
            bar. Without this a wide universe ranks BTC against coins that
            traded $50k a day, which is a ranking you could not act on.
            Screening on full-sample volume instead would be lookahead — it
            selects the coins that went on to matter.
    """
    symbols = symbols or UNIVERSE_2021_12
    d = cfg.data
    # Derived rather than looked up in a second table. The table this replaced
    # held six intervals and fell back to 4 for anything else — so a 30m or 15m
    # config, both of which `BARS_PER_YEAR` accepts, scaled the funding charge
    # as though its bars were 4h. That is an 8x error applied silently, in a
    # package whose config loader refuses unknown KEYS on the grounds that a
    # quiet default is how you spend a week backtesting something you thought
    # you had changed. `bars_per_year` raises ConfigError on an unknown
    # interval, so the same question now has one answer and one failure mode.
    interval_hours = 8_760 / d.bars_per_year

    # Pass 1: bars only. Cheap enough to hold the whole universe in memory,
    # which is what lets the screen be computed before features are built.
    bars_by_symbol: dict[str, pd.DataFrame] = {}
    delisted: dict[str, pd.Timestamp] = {}
    for symbol in symbols:
        bars = load_binance(symbol, d.interval, d.start, d.end, d.cache_dir,
                            verbose=False, market="um", required=False)
        if bars is None or len(bars) < 500:
            continue
        bars_by_symbol[symbol] = bars
        if bars.index[-1] < pd.Timestamp(d.end) - pd.Timedelta(days=60):
            delisted[symbol] = bars.index[-1]

    if len(bars_by_symbol) < 3:
        raise InsufficientDataError(
            f"need at least 3 symbols with data for a cross-section, "
            f"got {len(bars_by_symbol)} of {len(symbols)} requested"
        )

    times = pd.DatetimeIndex(
        sorted(set().union(*(b.index for b in bars_by_symbol.values())))
    )
    log_open = pd.DataFrame(
        {s: np.log(b["open"]) for s, b in bars_by_symbol.items()}
    ).reindex(times)

    screen = None
    if top_n:
        from nullres.data.universe import liquidity_screen

        dollar_volume = pd.DataFrame(
            {s: b["volume"] * b["close"] for s, b in bars_by_symbol.items()}
        ).reindex(times)
        screen = liquidity_screen(dollar_volume, top_n=top_n, window=screen_window)
        # Symbols never selected cannot influence any decision, so skip the
        # cost of featurising them. This is an optimisation, not a filter:
        # the time-varying screen below is what actually gates membership.
        keep = [s for s in bars_by_symbol if screen[s].any()]
        if verbose:
            log.info("  liquidity screen: top %d by %d-bar trailing volume; "
                     "%d of %d symbols ever qualify", top_n, screen_window,
                     len(keep), len(bars_by_symbol))
        bars_by_symbol = {s: bars_by_symbol[s] for s in keep}
        screen = screen[keep]
        log_open = log_open[keep]

    _guard_metrics_fetch(d, list(bars_by_symbol))

    # Pass 2: features, only for symbols that can ever be traded.
    #
    # This is the longest silent stretch in the repo. On a wide universe it
    # loads auxiliary archives and builds features for a hundred-odd symbols,
    # which takes tens of minutes and used to print nothing at all — the
    # per-symbol line below is suppressed above 15 symbols, precisely when the
    # wait is longest. There was no way to tell a working run from a hung one,
    # or to see it approaching the memory ceiling before the kernel intervened.
    per_symbol: dict[str, pd.DataFrame] = {}
    funding_cols: dict[str, pd.Series] = {}
    wide = len(bars_by_symbol) > 15
    if verbose and wide:
        log.info("\nbuilding features for %d symbols "
                 "(quiet for a while; progress every 10)", len(bars_by_symbol))

    for done, (symbol, bars) in enumerate(bars_by_symbol.items(), start=1):
        sym_cfg = _symbol_cfg(d, symbol)
        fund, metrics = load_auxiliary(sym_cfg, verbose=False, bars=bars)
        per_symbol[symbol] = build_features(bars, funding=fund, metrics=metrics)

        if verbose and wide and (done % 10 == 0 or done == len(bars_by_symbol)):
            columns = per_symbol[symbol].shape[1]
            held = sum(f.memory_usage(deep=True).sum() for f in per_symbol.values())
            log.info("  %s/%d  %d features  %s MB held", f"{done:>4}",
                     len(bars_by_symbol), columns, f"{held / 1e6:,.0f}")

        if fund is not None and len(fund):
            # Funding settles every 8h; a bar of `interval_hours` carries that
            # fraction of one settlement. A long pays when the rate is positive.
            from nullres.features.derivatives import _asof

            rate = _asof(bars, fund, ["funding_rate"])["funding_rate"]
            funding_cols[symbol] = rate * (interval_hours / 8.0)
        else:
            funding_cols[symbol] = pd.Series(0.0, index=bars.index)

        if verbose and len(bars_by_symbol) <= 15:
            log.info("  %-12s %s bars  %s..%s%s", symbol, f"{len(bars):>6,}",
                     f"{bars.index[0]:%Y-%m-%d}", f"{bars.index[-1]:%Y-%m-%d}",
                     "   DELISTED" if symbol in delisted else "")

    # open[t+1] -> open[t+2], per symbol. Same convention as the single-asset
    # engine: decided at close of t, filled at open of t+1.
    ret_next = log_open.shift(-2) - log_open.shift(-1)
    if screen is not None:
        # Un-screened symbols are not tradable, so they must not contribute
        # returns, ranks, or peer-group medians.
        ret_next = ret_next.where(screen)

    funding = pd.DataFrame(funding_cols).reindex(times).fillna(0.0)

    # A wide panel is the memory peak of this whole repository: 123 symbols x
    # ~9,000 timestamps x 46 features is ~400MB per copy, and building it makes
    # several (reindex, concat, sort, rank). At 46 features that was enough to
    # get the process OOM-killed on a 2GB machine — after two and a half hours
    # of downloading, so the failure landed as far as possible from its cause.
    #
    # The obvious economy is float32, and it is NOT free: it moved the narrow
    # panel's mean AUC from 0.5443 to 0.5429. Ranking is supposed to make the
    # precision irrelevant, but near-ties reorder and the model's binning shifts
    # with them. Trading a numerical change in the headline result for memory is
    # not a trade this repo can make quietly, so the frames stay float64 and the
    # peak is cut by holding fewer copies at once instead.
    built = list(per_symbol)
    frame = pd.concat(
        {s: f.reindex(times) for s, f in per_symbol.items()},
        names=["symbol", "ts"],
    ).swaplevel().sort_index()
    per_symbol.clear()          # ~400MB, no longer needed once stacked
    if verbose:
        log.info("  panel frame: %s rows x %d features (%s MB)",
                 f"{len(frame):,}", frame.shape[1],
                 f"{frame.memory_usage(deep=True).sum() / 1e6:,.0f}")

    ranked = _cross_sectional_rank(frame, screen)
    del frame                   # another ~400MB, before the label is built
    y = _relative_label(log_open.where(screen) if screen is not None else log_open,
                        cfg.label.horizon)

    common = ranked.index.intersection(y.index)
    return Panel(
        features=ranked.loc[common],
        y=y.loc[common],
        ret_next=ret_next,
        funding=funding,
        times=times,
        horizon=cfg.label.horizon,
        symbols=built,
        delisted={s: t for s, t in delisted.items() if s in built},
    )

time_folds

time_folds(times: DatetimeIndex, cfg_split, horizon: int)

Yield (train_times, test_times). Purged by the label horizon.

Splitting on row position would be wrong here: rows are (ts, symbol) pairs, so a positional split would put BTC's Tuesday in training and ETH's Tuesday in test. The model would learn the day from ten correlated siblings.

Source code in nullres/crosssec.py
def time_folds(times: pd.DatetimeIndex, cfg_split, horizon: int):
    """Yield (train_times, test_times). Purged by the label horizon.

    Splitting on row position would be wrong here: rows are (ts, symbol) pairs,
    so a positional split would put BTC's Tuesday in training and ETH's Tuesday
    in test. The model would learn the day from ten correlated siblings.
    """
    n = len(times)
    if n <= cfg_split.min_train:
        raise InsufficientDataError(
            f"only {n:,} timestamps but min_train is {cfg_split.min_train:,}"
        )
    fold = (n - cfg_split.min_train) // cfg_split.n_folds
    if fold < 50:
        raise InsufficientDataError(
            "folds too small; reduce n_folds or widen the range"
        )

    for k in range(cfg_split.n_folds):
        test_start = cfg_split.min_train + k * fold
        test_end = min(test_start + fold, n)
        if test_end - test_start < 50:
            break
        train_hi = test_start - horizon - cfg_split.embargo
        if train_hi < 50:
            continue
        lo = max(0, train_hi - cfg_split.train_window) if cfg_split.scheme == "rolling" else 0
        yield times[lo:train_hi], times[test_start:test_end]

fit_predict_panel

fit_predict_panel(panel: Panel, cfg, verbose: bool = True)

Walk-forward P(outperforms) for every (ts, symbol) in a test fold.

Source code in nullres/crosssec.py
def fit_predict_panel(panel: Panel, cfg, verbose: bool = True):
    """Walk-forward P(outperforms) for every (ts, symbol) in a test fold."""
    from nullres.models.classifier import make_model
    from sklearn.metrics import roc_auc_score

    X, y = panel.features, panel.y
    ts_level = X.index.get_level_values("ts")
    proba = pd.Series(np.nan, index=X.index, dtype="float64")
    reports = []

    for k, (train_times, test_times) in enumerate(
        time_folds(panel.times, cfg.split, panel.horizon), start=1
    ):
        train_mask = ts_level.isin(train_times)
        test_mask = ts_level.isin(test_times)
        if train_mask.sum() < 500 or test_mask.sum() < 100:
            continue

        model = make_model(cfg.model)
        model.fit(X[train_mask], y[train_mask].astype(int))
        p = model.predict_proba(X[test_mask])[:, 1]
        proba[test_mask] = p

        y_test = y[test_mask].astype(int)
        auc = roc_auc_score(y_test, p) if y_test.nunique() == 2 else float("nan")
        reports.append({"fold": k, "train": int(train_mask.sum()),
                        "test": int(test_mask.sum()), "auc": float(auc),
                        "test_from": str(test_times[0])[:10],
                        "test_to": str(test_times[-1])[:10]})
        if verbose:
            log.info("  fold %d: train %s  test %s  [%s..%s]  auc %.4f", k,
                     f"{train_mask.sum():>7,}", f"{test_mask.sum():>6,}",
                     reports[-1]["test_from"], reports[-1]["test_to"], auc)

    if not reports:
        raise InsufficientDataError("no fold produced predictions")
    return proba, reports

panel_positions

panel_positions(proba: Series, panel: Panel, top_k: int = 3, rebalance: int = 42, allow_short: bool = True) -> DataFrame

Long the top-k ranked symbols, short the bottom-k. Dollar neutral.

Weights are +1/k and -1/k so gross exposure is 2 and net is 0 — the book makes no bet on the market's direction, which is the entire point.

That gross of 2.0 is deliberate, and sizing.max_leverage does not constrain it. That field clips a single-asset position to +/-1 and is never read here; a dollar-neutral book is 100% long and 100% short by definition, and capping gross at 1.0 would mean holding half of each side, which is a different strategy rather than a safer version of this one. But 2x gross notional is 2x notional whatever the net is: it requires margin, and it is why the tail census in panelaudit matters. summarize reports gross_exposure so this is visible in the results table instead of being inferable only from reading this function.

rebalance throttles turnover exactly as min_hold does in the single-asset engine. Reshuffling an 11-symbol book every bar is the cross-sectional version of the mistake that cost the original baseline 100%.

On the shape of this function. The obvious spelling is one loop over every timestamp, and that is what this was: ~42 seconds on a 9,000 x 120 panel, because positions.loc[ts] = ... reallocates and realigns a row at a time and _neutralise was called once per bar. The book only changes at rebalances, though — 214 of those 9,000 bars — and everything in between is a carry-forward, a liveness mask and a rescale, all of which are array operations over the whole panel.

So the ranking still runs bar by bar, on rebalance bars only, and still through Series.sort_values: the selection has to break ties exactly as before, and reimplementing that with argsort risks silently reordering near-equal probabilities. Only the parts with no ordering subtlety were vectorised. That is 0.2 seconds, and tests/test_crosssec.py asserts the output is bit-identical to the original loop rather than merely close — this book's headline Sharpe has already been shown to move by 60% on a 0.0003% change in its inputs, so "close" is not a standard worth anything here.

Source code in nullres/crosssec.py
def panel_positions(proba: pd.Series, panel: Panel, top_k: int = 3,
                    rebalance: int = 42, allow_short: bool = True) -> pd.DataFrame:
    """Long the top-k ranked symbols, short the bottom-k. Dollar neutral.

    Weights are +1/k and -1/k so gross exposure is 2 and net is 0 — the book
    makes no bet on the market's direction, which is the entire point.

    **That gross of 2.0 is deliberate, and `sizing.max_leverage` does not
    constrain it.** That field clips a single-asset position to +/-1 and is
    never read here; a dollar-neutral book is 100% long and 100% short by
    definition, and capping gross at 1.0 would mean holding half of each side,
    which is a different strategy rather than a safer version of this one. But
    2x gross notional is 2x notional whatever the net is: it requires margin,
    and it is why the tail census in `panelaudit` matters. `summarize` reports
    `gross_exposure` so this is visible in the results table instead of being
    inferable only from reading this function.

    `rebalance` throttles turnover exactly as `min_hold` does in the
    single-asset engine. Reshuffling an 11-symbol book every bar is the
    cross-sectional version of the mistake that cost the original baseline 100%.

    **On the shape of this function.** The obvious spelling is one loop over
    every timestamp, and that is what this was: ~42 seconds on a 9,000 x 120
    panel, because `positions.loc[ts] = ...` reallocates and realigns a row at
    a time and `_neutralise` was called once per bar. The book only *changes*
    at rebalances, though — 214 of those 9,000 bars — and everything in between
    is a carry-forward, a liveness mask and a rescale, all of which are array
    operations over the whole panel.

    So the ranking still runs bar by bar, on rebalance bars only, and still
    through `Series.sort_values`: the selection has to break ties exactly as
    before, and reimplementing that with `argsort` risks silently reordering
    near-equal probabilities. Only the parts with no ordering subtlety were
    vectorised. That is 0.2 seconds, and `tests/test_crosssec.py` asserts the
    output is bit-identical to the original loop rather than merely close —
    this book's headline Sharpe has already been shown to move by 60% on a
    0.0003% change in its inputs, so "close" is not a standard worth anything
    here.
    """
    wide = proba.unstack("symbol").reindex(panel.times)
    cols = wide.columns
    n, m = len(panel.times), len(cols)
    position_of = {c: j for j, c in enumerate(cols)}

    # Pass 1: the target book at each rebalance bar. `current` carries forward
    # when a bar has too few live symbols to rank, exactly as before.
    target = np.zeros((n, m), dtype="float64")
    current = np.zeros(m, dtype="float64")
    for i in range(0, n, rebalance):
        row = wide.iloc[i].dropna()
        if len(row) >= 2 * top_k:
            ranked = row.sort_values(ascending=False)
            current = np.zeros(m, dtype="float64")
            current[[position_of[c] for c in ranked.index[:top_k]]] = 1.0 / top_k
            if allow_short:
                current[[position_of[c] for c in ranked.index[-top_k:]]] = -1.0 / top_k
        target[i] = current

    # Pass 2: hold that book until the next rebalance, then apply the liveness
    # mask and renormalise. A delisted symbol cannot be held, so it is forced
    # flat rather than silently carried in an instrument that stopped existing.
    held = target[(np.arange(n) // rebalance) * rebalance]
    alive = panel.ret_next.reindex(index=panel.times, columns=cols).notna()
    weights = held * alive.to_numpy()
    return pd.DataFrame(_neutralise_rows(weights, allow_short),
                        index=panel.times, columns=cols)

backtest_panel

backtest_panel(positions: DataFrame, panel: Panel, cost_cfg, charge_funding: bool = True)

Per-bar net log return of the book, with per-symbol costs and funding.

Source code in nullres/crosssec.py
def backtest_panel(positions: pd.DataFrame, panel: Panel, cost_cfg,
                   charge_funding: bool = True):
    """Per-bar net log return of the book, with per-symbol costs and funding."""
    from nullres.backtest.engine import BacktestResult

    ret = panel.ret_next.reindex(positions.index)[positions.columns].fillna(0.0)
    gross = (positions * ret).sum(axis=1)

    turnover = positions.diff().abs()
    turnover.iloc[0] = positions.iloc[0].abs()
    turnover_total = turnover.sum(axis=1)

    rate = (cost_cfg.fee_bps + cost_cfg.slippage_bps) / 10_000.0
    cost = turnover_total * rate

    if charge_funding:
        # Longs pay funding when the rate is positive, shorts receive it. On a
        # dollar-neutral book these largely offset — but "largely" is not
        # "exactly", and the residual is a real cost of carrying perps.
        fund = panel.funding.reindex(positions.index)[positions.columns].fillna(0.0)
        cost = cost + (positions * fund).sum(axis=1)

    net = gross - cost
    return BacktestResult(
        equity=np.exp(net.cumsum()),
        returns=net,
        gross=gross,
        position=positions.abs().sum(axis=1),   # gross exposure
        turnover=turnover_total,
        cost=cost,
    )

benchmarks

benchmarks(panel: Panel, cost_cfg, oos_times: DatetimeIndex | None = None, rebalance: int = 42, reference: str = 'BTCUSDT') -> dict

The books a cross-sectional model has to beat to be worth anything.

All are restricted to oos_times. That restriction is not cosmetic: the ML book only trades inside its test folds, so an unmasked benchmark carries the entire pre-OOS period. Here that meant equal-weight absorbing the 2022 bear market the model never touched, reporting -83% and making a mediocre strategy look excellent by comparison.

static_vs_alts is the important one. It is long the reference asset and short everything else, rebalanced never — no model, three trades. BTC outperformed alts massively over 2022-2025, so any model that has learned "the lowest-volatility member outperforms" has learned this bet under another name, and must be measured against it.

Note that static_vs_alts is itself hindsight-selected: BTC is the long leg because we know how the period ended. It is a lower bound on what a model must beat, not a strategy.

Source code in nullres/crosssec.py
def benchmarks(panel: Panel, cost_cfg, oos_times: pd.DatetimeIndex | None = None,
               rebalance: int = 42, reference: str = "BTCUSDT") -> dict:
    """The books a cross-sectional model has to beat to be worth anything.

    All are restricted to `oos_times`. That restriction is not cosmetic: the ML
    book only trades inside its test folds, so an unmasked benchmark carries the
    entire pre-OOS period. Here that meant equal-weight absorbing the 2022 bear
    market the model never touched, reporting -83% and making a mediocre
    strategy look excellent by comparison.

    `static_vs_alts` is the important one. It is long the reference asset and
    short everything else, rebalanced never — no model, three trades. BTC
    outperformed alts massively over 2022-2025, so any model that has learned
    "the lowest-volatility member outperforms" has learned this bet under
    another name, and must be measured against it.

    Note that `static_vs_alts` is itself hindsight-selected: BTC is the long leg
    because we know how the period ended. It is a lower bound on what a model
    must beat, not a strategy.
    """
    cols = panel.ret_next.columns
    alive = panel.ret_next.notna()

    def mask(pos: pd.DataFrame) -> pd.DataFrame:
        if oos_times is None:
            return pos
        out = pos.copy()
        out.loc[~pos.index.isin(oos_times)] = 0.0
        return out

    out = {}

    out["equal_weight"] = backtest_panel(
        mask(_equal_weight_book(alive, rebalance)), panel, cost_cfg)

    if reference in cols:
        long_only = pd.DataFrame(0.0, index=panel.times, columns=cols)
        long_only[reference] = 1.0
        out[f"{reference[:3].lower()}_only"] = backtest_panel(
            mask(long_only.where(alive, 0.0)), panel, cost_cfg
        )

        alts = [c for c in cols if c != reference]
        static = pd.DataFrame(0.0, index=panel.times, columns=cols)
        static[reference] = 1.0
        alt_alive = alive[alts]
        static[alts] = -alt_alive.div(
            alt_alive.sum(axis=1).replace(0, np.nan), axis=0
        ).fillna(0.0)
        out["static_vs_alts"] = backtest_panel(
            mask(static.where(alive, 0.0)), panel, cost_cfg
        )

    return out

equal_weight_benchmark

equal_weight_benchmark(panel: Panel, cost_cfg, rebalance: int = 42)

Backwards-compatible single benchmark; prefer benchmarks.

Source code in nullres/crosssec.py
def equal_weight_benchmark(panel: Panel, cost_cfg, rebalance: int = 42):
    """Backwards-compatible single benchmark; prefer `benchmarks`."""
    return benchmarks(panel, cost_cfg, rebalance=rebalance)["equal_weight"]

Panel controls

panelaudit

Controls that decide whether cross-sectional skill is real.

A panel AUC above 0.5 has more ways of being an artefact than a single-asset one, and the interesting failures are not leaks — they are the model learning something true but useless:

SHUFFLED LABEL Permute the target within each timestamp and refit. The permutation preserves the balanced structure, so anything left is a side channel, not signal.

SURVIVORS ONLY Refit with the delisted symbols removed. If the AUC collapses, the model was detecting death, not ranking assets — a real regularity you cannot trade, because by the time a coin is dying its borrow has vanished.

PER-SYMBOL Cross-sectional ranks never name a symbol, but if BTC is permanently the lowest-volatility member then "rank 1 by low vol" and "BTC" are the same column. A wide spread in per-symbol accuracy is that tell.

CONTRIBUTION Which symbols actually produced the P&L, and how much of it came from the ones that delisted.

TAIL CENSUS A short book that never got hit is not a book with no tail risk. Count how often the moves that would hurt occur, and how many short-name-bars were exposed to them; the product is the number of hits chance predicts.

These were run once by hand and quoted in RESEARCH.md, which meant the numbers underneath the project's strongest result were the only ones no command could regenerate. That is exactly backwards.

shuffled_label_auc

shuffled_label_auc(panel, cfg, seed: int = 0) -> float

Mean fold AUC after permuting the label within each timestamp.

Permuting within a timestamp rather than globally keeps the label balanced in every regime, so the control isolates the ranking signal instead of also destroying the panel's structure.

Each timestamp draws from a stream seeded by (seed, that timestamp), so a given bar receives the same permutation no matter where it falls in the iteration. Sharing one generator across groups would have made the result depend on processing order — reproducible only so long as nothing upstream changed how the panel is grouped, which is not a property worth relying on in a control whose whole job is to be trustworthy.

Source code in nullres/panelaudit.py
def shuffled_label_auc(panel, cfg, seed: int = 0) -> float:
    """Mean fold AUC after permuting the label within each timestamp.

    Permuting *within* a timestamp rather than globally keeps the label balanced
    in every regime, so the control isolates the ranking signal instead of also
    destroying the panel's structure.

    Each timestamp draws from a stream seeded by `(seed, that timestamp)`, so a
    given bar receives the same permutation no matter where it falls in the
    iteration. Sharing one generator across groups would have made the result
    depend on processing order — reproducible only so long as nothing upstream
    changed how the panel is grouped, which is not a property worth relying on
    in a control whose whole job is to be trustworthy.
    """
    from nullres.crosssec import Panel, fit_predict_panel

    def permute(group: pd.Series) -> pd.Series:
        ts = group.index.get_level_values("ts")[0]
        rng = np.random.default_rng([seed, int(pd.Timestamp(ts).value)])
        # Sort by symbol before permuting. Seeding per timestamp fixes which
        # stream a bar draws from, but the permutation still lands on whatever
        # order the values arrive in — so row order would otherwise decide which
        # symbol got which label. Sorting first makes the symbol -> label
        # assignment a function of the timestamp alone.
        ordered = group.sort_index(level="symbol")
        drawn = pd.Series(rng.permutation(ordered.to_numpy()), index=ordered.index)
        return drawn.reindex(group.index)

    shuffled = panel.y.groupby(level="ts", group_keys=False).apply(permute)

    control = Panel(features=panel.features, y=shuffled, ret_next=panel.ret_next,
                    funding=panel.funding, times=panel.times,
                    horizon=panel.horizon, symbols=panel.symbols,
                    delisted=panel.delisted)
    _, reports = fit_predict_panel(control, cfg, verbose=False)
    return float(np.nanmean([r["auc"] for r in reports]))

survivors_only_auc

survivors_only_auc(panel, cfg) -> float | None

Mean fold AUC with every delisted symbol dropped.

A model that only knows which coins are dying has found something real and untradable. None when the universe contains no corpses to remove.

Source code in nullres/panelaudit.py
def survivors_only_auc(panel, cfg) -> float | None:
    """Mean fold AUC with every delisted symbol dropped.

    A model that only knows which coins are dying has found something real and
    untradable. None when the universe contains no corpses to remove.
    """
    from nullres.crosssec import Panel, fit_predict_panel

    if not panel.delisted:
        return None
    dead = set(panel.delisted)
    keep = [s for s in panel.symbols if s not in dead]
    alive = panel.features.index.get_level_values("symbol").isin(keep)

    control = Panel(features=panel.features[alive], y=panel.y[alive],
                    ret_next=panel.ret_next[keep], funding=panel.funding[keep],
                    times=panel.times, horizon=panel.horizon,
                    symbols=keep, delisted={})
    _, reports = fit_predict_panel(control, cfg, verbose=False)
    return float(np.nanmean([r["auc"] for r in reports]))

per_symbol_accuracy

per_symbol_accuracy(proba: Series, panel) -> DataFrame

Directional accuracy per symbol, with the count it rests on.

Skill concentrated in one or two names is skill that has learned those names, whatever the ranks pretend. But the spread only means that if every symbol has enough scored bars to have an accuracy worth reading.

The count is not decoration. On a screened wide universe, symbols drift in and out of the tradable set and some are scored on a few dozen bars. The first version of this returned bare accuracies, and the 136-symbol panel duly produced a spread of 0.685 — seven times the narrow panel's, and entirely an artefact of thin symbols. At n=30 a coin flip reaches 0.86 without trying. Callers filter on n before quoting a spread.

Accuracy alone is the wrong statistic here, and this is subtle. The label is "beats the cross-sectional median", so a coin that persistently underperformed has a lopsided base rate of its own — say 0.85 zeros. A model that learned nothing but "this one usually lags" scores 0.85 on it. High per-symbol accuracy can therefore be pure unconditional drift, and reading the raw spread as "skill is concentrated in these names" overstates it.

Lift over the symbol's own majority class removes that, but introduces its own bias in exactly this setting: a cross-sectional model MUST rank, so at every timestamp roughly half the universe is predicted low. It structurally cannot predict the majority class for a symbol that beats the median 79% of the time, and gets charged a large negative lift for a constraint rather than a mistake.

So the spread worth reading is in per-symbol AUC. It is threshold-free and base-rate invariant: 0.5 means the model cannot tell this symbol's good bars from its bad ones, whatever its unconditional tendency and whatever the ranking forced. Accuracy, base rate and lift are kept alongside because they are what a reader expects to see, but the AUC column is the one that answers "is the skill concentrated in particular names".

Source code in nullres/panelaudit.py
def per_symbol_accuracy(proba: pd.Series, panel) -> pd.DataFrame:
    """Directional accuracy per symbol, with the count it rests on.

    Skill concentrated in one or two names is skill that has learned those
    names, whatever the ranks pretend. But the spread only means that if every
    symbol has enough scored bars to have an accuracy worth reading.

    **The count is not decoration.** On a screened wide universe, symbols drift
    in and out of the tradable set and some are scored on a few dozen bars. The
    first version of this returned bare accuracies, and the 136-symbol panel
    duly produced a spread of 0.685 — seven times the narrow panel's, and
    entirely an artefact of thin symbols. At n=30 a coin flip reaches 0.86
    without trying. Callers filter on `n` before quoting a spread.

    **Accuracy alone is the wrong statistic here, and this is subtle.** The
    label is "beats the cross-sectional median", so a coin that persistently
    underperformed has a lopsided base rate of its own — say 0.85 zeros. A model
    that learned nothing but "this one usually lags" scores 0.85 on it. High
    per-symbol accuracy can therefore be pure unconditional drift, and reading
    the raw spread as "skill is concentrated in these names" overstates it.

    Lift over the symbol's own majority class removes that, but introduces its
    own bias in exactly this setting: a cross-sectional model MUST rank, so at
    every timestamp roughly half the universe is predicted low. It structurally
    cannot predict the majority class for a symbol that beats the median 79% of
    the time, and gets charged a large negative lift for a constraint rather
    than a mistake.

    **So the spread worth reading is in per-symbol AUC.** It is threshold-free
    and base-rate invariant: 0.5 means the model cannot tell this symbol's good
    bars from its bad ones, whatever its unconditional tendency and whatever the
    ranking forced. Accuracy, base rate and lift are kept alongside because they
    are what a reader expects to see, but the AUC column is the one that answers
    "is the skill concentrated in particular names".
    """
    from sklearn.metrics import roc_auc_score

    scored = proba.notna() & panel.y.notna()
    p, y = proba[scored], panel.y[scored]

    hit = ((p > 0.5) == (y > 0.5)).astype(float).groupby(level="symbol")
    mean_y = y.groupby(level="symbol").mean()
    base = pd.concat([mean_y, 1.0 - mean_y], axis=1).max(axis=1)

    def symbol_auc(group: pd.Series) -> float:
        truth = y.loc[group.index]
        if truth.nunique() < 2 or len(truth) < 2:
            return float("nan")
        return float(roc_auc_score(truth.astype(int), group))

    auc = p.groupby(level="symbol").apply(symbol_auc)

    out = pd.DataFrame({"auc": auc, "accuracy": hit.mean(), "base_rate": base,
                        "n": hit.size()})
    out["lift"] = out["accuracy"] - out["base_rate"]
    return out.sort_values("auc", ascending=False)

pnl_contribution

pnl_contribution(positions: DataFrame, panel) -> Series

Gross log P&L attributable to each symbol.

Source code in nullres/panelaudit.py
def pnl_contribution(positions: pd.DataFrame, panel) -> pd.Series:
    """Gross log P&L attributable to each symbol."""
    ret = panel.ret_next.reindex(positions.index)[positions.columns].fillna(0.0)
    return (positions * ret).sum().sort_values(ascending=False)

delisted_share

delisted_share(positions: DataFrame, panel) -> float

Share of P&L ACTIVITY, in absolute terms, from symbols that later delisted.

This is a share of absolute P&L, not of net profit, and the distinction changes what the number means. Each symbol contributes |its P&L|, so a +30% winner and a -30% loser both count as 30 rather than cancelling to zero. The question being asked is "how much of what this book did happened in coins that were dying" — exposure, not profitability.

Netting would answer a different and weaker question. A book that made a fortune on one delisting and lost it on another would net to ~0% and look untouched by delisting, when in fact its entire outcome hinged on dying coins. For a survivorship control that is the wrong answer, so the metric deliberately does not net.

The consequence to keep in mind when reading it: this number cannot be compared against a return, and it can be large while the delisted names contributed nothing to the bottom line.

Source code in nullres/panelaudit.py
def delisted_share(positions: pd.DataFrame, panel) -> float:
    """Share of P&L ACTIVITY, in absolute terms, from symbols that later delisted.

    **This is a share of absolute P&L, not of net profit, and the distinction
    changes what the number means.** Each symbol contributes `|its P&L|`, so a
    +30% winner and a -30% loser both count as 30 rather than cancelling to
    zero. The question being asked is "how much of what this book did happened
    in coins that were dying" — exposure, not profitability.

    Netting would answer a different and weaker question. A book that made a
    fortune on one delisting and lost it on another would net to ~0% and look
    untouched by delisting, when in fact its entire outcome hinged on dying
    coins. For a survivorship control that is the wrong answer, so the metric
    deliberately does not net.

    The consequence to keep in mind when reading it: this number cannot be
    compared against a return, and it can be large while the delisted names
    contributed nothing to the bottom line.
    """
    contribution = pnl_contribution(positions, panel).abs()
    total = float(contribution.sum())
    if total <= 0:
        return 0.0
    dead = [s for s in panel.delisted if s in contribution.index]
    return float(contribution[dead].sum() / total)

tail_census

tail_census(positions: DataFrame, panel, threshold: float = 0.65) -> dict

How many moves big enough to matter occurred, and how exposed the book was.

Observing no blow-up means nothing until you know how many blow-ups chance predicted. If the expected count is a fraction of one, zero hits is what chance produces and the tail is untested rather than absent.

threshold is one point on a curve, and a single point is a constant doing analytical work — the thing this repo keeps having to remove. Prefer tail_curve, which sweeps it and reports the capital each level would cost, so no one number carries the argument.

Source code in nullres/panelaudit.py
def tail_census(positions: pd.DataFrame, panel, threshold: float = 0.65) -> dict:
    """How many moves big enough to matter occurred, and how exposed the book was.

    Observing no blow-up means nothing until you know how many blow-ups chance
    predicted. If the expected count is a fraction of one, zero hits is what
    chance produces and the tail is untested rather than absent.

    `threshold` is one point on a curve, and a single point is a constant doing
    analytical work — the thing this repo keeps having to remove. Prefer
    `tail_curve`, which sweeps it and reports the capital each level would cost,
    so no one number carries the argument.
    """
    ret = panel.ret_next.reindex(positions.index)[positions.columns]
    simple = np.expm1(ret)

    extreme = int((simple > threshold).sum().sum())
    observed_bars = int(simple.notna().sum().sum())
    rate = extreme / observed_bars if observed_bars else 0.0

    short_bars = int((positions < 0).sum().sum())
    worst_bar = float(np.expm1((positions * ret.fillna(0.0)).sum(axis=1)).min())

    hit = simple.where(positions < 0) > threshold
    return {
        "threshold": threshold,
        "extreme_moves": extreme,
        "observed_bars": observed_bars,
        "rate": rate,
        "short_name_bars": short_bars,
        "expected_hits": rate * short_bars,
        "actual_hits": int(hit.sum().sum()),
        "worst_bar_return": worst_bar,
    }

concentration

concentration(positions: DataFrame, nominal: float) -> dict

How often a dead leg left the book concentrated, and for how long.

_neutralise keeps the book dollar-neutral when a symbol delists by rescaling the surviving side, so a k=2 book whose short partner dies holds -1.0 in one name instead of -0.5 in two. Gross exposure does not change and net stays at zero — what changes is that the move which ruins the book halves, from +200% to +100%. Nine symbols delist in the wide universe, so this is not hypothetical.

Keeping that behaviour is a choice: dollar neutrality is the book's defining constraint, and the alternatives (halve the long side, go flat) change the strategy rather than make it safer. What was missing is visibility. A maximum weight on its own cannot distinguish one bar from three thousand — the difference between a curiosity and the dominant risk in the book — so this reports the share of bars spent concentrated and the longest unbroken stretch of it.

Source code in nullres/panelaudit.py
def concentration(positions: pd.DataFrame, nominal: float) -> dict:
    """How often a dead leg left the book concentrated, and for how long.

    `_neutralise` keeps the book dollar-neutral when a symbol delists by
    rescaling the surviving side, so a k=2 book whose short partner dies holds
    -1.0 in one name instead of -0.5 in two. Gross exposure does not change and
    net stays at zero — what changes is that the move which ruins the book
    halves, from +200% to +100%. Nine symbols delist in the wide universe, so
    this is not hypothetical.

    Keeping that behaviour is a choice: dollar neutrality is the book's defining
    constraint, and the alternatives (halve the long side, go flat) change the
    strategy rather than make it safer. What was missing is visibility. A
    maximum weight on its own cannot distinguish one bar from three thousand —
    the difference between a curiosity and the dominant risk in the book — so
    this reports the share of bars spent concentrated and the longest unbroken
    stretch of it.
    """
    longs = positions.where(positions > 0)
    shorts = -positions.where(positions < 0)

    per_bar = pd.concat([longs.max(axis=1), shorts.max(axis=1)], axis=1).max(axis=1)
    tol = nominal * 1e-6
    over = per_bar > nominal + tol
    held = positions.abs().sum(axis=1) > tol          # bars with a position at all

    max_short = float(shorts.max().max()) if shorts.notna().any().any() else 0.0
    max_long = float(longs.max().max()) if longs.notna().any().any() else 0.0
    peak = max(max_short, max_long)

    # Longest unbroken stretch of concentration, in bars.
    longest = run = 0
    for flag in over.to_numpy():
        run = run + 1 if flag else 0
        longest = max(longest, run)

    active = int(held.sum())
    at_peak = int((per_bar >= peak - tol).sum()) if peak > nominal + tol else 0
    return {
        "nominal": nominal,
        "max_short": max_short,
        "max_long": max_long,
        "bars_held": active,
        "concentrated_bars": int((over & held).sum()),
        "share": float((over & held).sum() / active) if active else 0.0,
        "bars_at_peak": at_peak,
        "share_at_peak": float(at_peak / active) if active else 0.0,
        "longest_run": longest,
    }

tail_curve

tail_curve(positions: DataFrame, panel, thresholds=(0.25, 0.5, 0.65, 1.0, 2.0)) -> DataFrame

Expected tail hits across move sizes, with what each would cost.

Two questions the single-threshold census could not answer.

How often? A rate estimated at one move size is one point on a steeply falling curve, and which point you pick decides whether the answer sounds reassuring. Sweeping removes the choice — the same reason nullres sweep prints a surface instead of its maximum.

How bad? The graveyard works this out by hand: "one UNFI-type event (+274% in 4h) against a -0.5 weight is -137% of capital". That arithmetic is mechanical and belongs in code. cost_of_one multiplies each move by the largest short weight the book actually held, so the loss is the book's own rather than an illustration, and ruinous marks the levels that would take more than all of it.

This is as far as tail risk can honestly be taken here. It says how exposed the book was and what one hit would have cost — it does not model margin, liquidation price, or auto-deleveraging, none of which are in the archive.

Source code in nullres/panelaudit.py
def tail_curve(positions: pd.DataFrame, panel,
               thresholds=(0.25, 0.50, 0.65, 1.00, 2.00)) -> pd.DataFrame:
    """Expected tail hits across move sizes, with what each would cost.

    Two questions the single-threshold census could not answer.

    **How often?** A rate estimated at one move size is one point on a steeply
    falling curve, and which point you pick decides whether the answer sounds
    reassuring. Sweeping removes the choice — the same reason `nullres sweep`
    prints a surface instead of its maximum.

    **How bad?** The graveyard works this out by hand: "one UNFI-type event
    (+274% in 4h) against a -0.5 weight is -137% of capital". That arithmetic is
    mechanical and belongs in code. `cost_of_one` multiplies each move by the
    largest short weight the book actually held, so the loss is the book's own
    rather than an illustration, and `ruinous` marks the levels that would take
    more than all of it.

    This is as far as tail risk can honestly be taken here. It says how exposed
    the book was and what one hit would have cost — it does not model margin,
    liquidation price, or auto-deleveraging, none of which are in the archive.
    """
    ret = panel.ret_next.reindex(positions.index)[positions.columns]
    simple = np.expm1(ret)
    observed = int(simple.notna().sum().sum())
    short_bars = int((positions < 0).sum().sum())
    worst_short = float(-positions[positions < 0].min().min()) if short_bars else 0.0

    rows = []
    for level in thresholds:
        moves = int((simple > level).sum().sum())
        hits = int(((simple.where(positions < 0)) > level).sum().sum())
        loss = level * worst_short

        # A move size never observed does NOT have probability zero, and
        # reporting 0.00 expected hits for the level that would ruin the book is
        # the error this whole file exists to catch — concluding absence from
        # non-observation. With no events in `observed` trials the rate is
        # unknown but bounded: the rule of three puts its 95% upper limit at
        # 3/observed. Rows with no occurrences therefore carry that bound, and
        # `estimated` marks which is which.
        estimated = moves > 0
        rate = (moves / observed) if estimated else (3.0 / observed if observed else 0.0)
        rows.append({
            "move": level,
            "occurrences": moves,
            "one_in": (observed / moves) if moves else (observed / 3.0),
            "expected_hits": rate * short_bars,
            "actual_hits": hits,
            "cost_of_one": loss,
            "ruinous": loss >= 1.0,
            "estimated": estimated,
        })
    out = pd.DataFrame(rows)
    out.attrs["short_name_bars"] = short_bars
    out.attrs["worst_short_weight"] = worst_short
    out.attrs["observed_bars"] = observed
    return out

format_report

format_report(panel, cfg, proba, positions, mean_auc: float, min_obs: int = 200, nominal_weight: float | None = None) -> str

Run every control and render it. The order is cheapest-first.

Source code in nullres/panelaudit.py
def format_report(panel, cfg, proba, positions, mean_auc: float,
                  min_obs: int = 200, nominal_weight: float | None = None) -> str:
    """Run every control and render it. The order is cheapest-first."""
    lines = ["", "--- verification " + "-" * 59, ""]

    shuffled = shuffled_label_auc(panel, cfg)
    verdict = "clean" if abs(shuffled - 0.5) < 0.02 else "SUSPECT"
    lines.append(f"  shuffled labels        AUC {shuffled:.4f}   "
                 f"vs {mean_auc:.4f} real — {verdict}")

    survivors = survivors_only_auc(panel, cfg)
    if survivors is None:
        lines.append("  survivors only         n/a — no delisted symbols to remove")
    else:
        drop = mean_auc - survivors
        reading = ("death detection" if drop > 0.02 else
                   "not death detection")
        lines.append(f"  survivors only         AUC {survivors:.4f}   "
                     f"({drop:+.4f}) — {reading}")

    # Symbols scored on a handful of bars carry accuracies that swing wildly by
    # chance, and on a screened universe there are always some. Quoting a spread
    # across them measures the screen, not the model.
    stats = per_symbol_accuracy(proba, panel)
    thick = stats[(stats["n"] >= min_obs) & stats["auc"].notna()]
    if len(thick) >= 2:
        thin = len(stats) - len(thick)
        note = f", {thin} thinner excluded" if thin else ""
        auc_spread = float(thick["auc"].iloc[0] - thick["auc"].iloc[-1])
        lines.append(
            f"  per-symbol skill       AUC spread {auc_spread:.3f} over "
            f"{len(thick)} symbols with >={min_obs} scored bars{note}"
        )
        for label, name in (("best ", thick.index[0]), ("worst", thick.index[-1])):
            row = thick.loc[name]
            lines.append(
                f"                         {label} {name} AUC {row['auc']:.3f}"
                f"  (accuracy {row['accuracy']:.3f} vs base rate "
                f"{row['base_rate']:.3f}, lift {row['lift']:+.3f}, "
                f"n={int(row['n']):,})"
            )
        above = int((thick["auc"] > 0.5).sum())
        lines.append(
            f"                         {above} of {len(thick)} symbols score "
            f"above 0.5; median {thick['auc'].median():.3f}"
        )
        lines.append(
            f"                         raw accuracy spread is "
            f"{float(thick['accuracy'].max() - thick['accuracy'].min()):.3f}, "
            f"but that is mostly each symbol's own base rate — AUC is the "
            f"base-rate-free read"
        )
    elif len(stats) >= 2:
        lines.append(f"  per-symbol skill       n/a — no symbol reached "
                     f"{min_obs} scored bars")

    share = delisted_share(positions, panel)
    lines.append(f"  delisted contribution  {share:.1%} of ABSOLUTE P&L (not "
                 f"netted) from {len(panel.delisted)} symbol(s) that stopped "
                 f"trading")

    contribution = pnl_contribution(positions, panel)
    top = ", ".join(f"{s}" for s in contribution.index[:4])
    bottom = ", ".join(f"{s}" for s in contribution.index[-4:])
    lines.append(f"  contributors           + {top}")
    lines.append(f"                         - {bottom}")

    curve = tail_curve(positions, panel)
    census = tail_census(positions, panel)
    short_bars = curve.attrs["short_name_bars"]
    weight = curve.attrs["worst_short_weight"]

    lines += [
        "",
        f"  tail exposure — book held {short_bars:,} short-name-bars across "
        f"{curve.attrs['observed_bars']:,} observed,",
        f"  at a largest short weight of {weight:.2f} per name. Worst bar "
        f"actually suffered: {census['worst_bar_return']:.1%}",
    ]

    if nominal_weight:
        conc = concentration(positions, nominal_weight)
        if conc["max_short"] > conc["nominal"] * 1.000001:
            lines += [
                "",
                f"    CONCENTRATION: nominal weight is {conc['nominal']:.2f} per "
                f"name, but a delisted leg leaves the",
                f"    survivor rescaled to keep the book dollar-neutral — peaking "
                f"at {conc['max_short']:.2f} short "
                f"({conc['max_long']:.2f} long).",
                f"    Concentrated on {conc['concentrated_bars']:,} of "
                f"{conc['bars_held']:,} bars held ({conc['share']:.1%}), "
                f"{conc['share_at_peak']:.1%} of them at the peak,",
                f"    longest unbroken stretch {conc['longest_run']:,} bars.",
                f"    Gross exposure never changes; the move that ruins the book "
                f"halves, from "
                f"+{1 / conc['nominal'] * 100:.0f}% to "
                f"+{1 / conc['max_short'] * 100:.0f}%.",
            ]
        else:
            lines.append(f"\n    Concentration: never exceeded the nominal "
                         f"{conc['nominal']:.2f} per name.")

    lines += [
        "",
        f"    {'move':>7}{'occurred':>10}{'1 in':>12}{'expected':>10}"
        f"{'actual':>8}{'costs':>9}",
    ]
    for _, row in curve.iterrows():
        bound = "" if row["estimated"] else "<"
        one_in = (f"{row['one_in']:,.0f}" if row["estimated"]
                  else f">{row['one_in']:,.0f}")
        flag = "  <- RUIN" if row["ruinous"] else ""
        lines.append(
            f"    {row['move']:>6.0%}{int(row['occurrences']):>10,}{one_in:>12}"
            f"{bound + format(row['expected_hits'], '.2f'):>10}"
            f"{int(row['actual_hits']):>8}{row['cost_of_one']:>8.0%}{flag}"
        )
    if not curve["estimated"].all():
        lines.append("")
        lines.append("    '<' marks a move size never observed here. Its rate is "
                     "not zero — it is unknown,")
        lines.append(f"    bounded above by the rule of three (3 events in "
                     f"{curve.attrs['observed_bars']:,} observations).")

    lines.append("")
    ruin = curve[curve["ruinous"]]
    if len(ruin):
        smallest = ruin.iloc[0]
        lines.append(
            f"    A single +{smallest['move']:.0%} move against the largest short "
            f"would cost {smallest['cost_of_one']:.0%} of capital — more than all "
            f"of it."
        )
        if smallest["estimated"]:
            lines.append(
                f"    Chance predicted {smallest['expected_hits']:.2f} such hits "
                f"and {int(smallest['actual_hits'])} occurred, so surviving is "
                f"what the exposure predicts,\n    not evidence the risk was "
                f"absent."
            )
        else:
            lines.append(
                f"    No move that large occurred here, so its rate is not "
                f"measured at all — only bounded\n    ABOVE, at most "
                f"{smallest['expected_hits']:.2f} expected hits. This sample "
                f"cannot show you this risk; it can\n    only fail to. UNFI did "
                f"+274% in a single 4h bar in 2021."
            )
    lines.append("    The tail is UNTESTED, not absent. The engine models no "
                 "margin, so a ruinous")
    lines.append("    bar would show as a large negative return rather than a "
                 "liquidation.")
    return "\n".join(lines)

The run ledger

runlog

The evidence layer beneath the graveyard.

Two things are being kept, and keeping them apart is the whole design:

runs/*.json          machine-written, append-only, never hand-edited.
                     What was measured, under exactly which config, at
                     which commit. Boring by design.

docs/05-graveyard.md hand-written. WHY it died and what it means. That
                     judgement is the actual work and stays human.

Neither half works alone. Prose without evidence rots — six months on, nobody can reproduce "mean AUC 0.5443". Evidence without prose is a spreadsheet that teaches nothing: no log entry will ever contain "one bear market wearing a trend-following costume". Graveyard entries cite run ids; run records point back at the graveyard.

The one capability that genuinely needs the machine layer is recognising that you are about to re-run something you already killed. Markdown cannot do that.

RunRecord dataclass

RunRecord(id: str, timestamp: str, command: str, config_name: str, config_hash: str, git_sha: str, git_dirty: bool, config: dict[str, Any], metrics: dict[str, Any] = dict(), verdict: str | None = None, notes: str = '', variants: int | None = None)

One execution of one command against one config.

flatten_config

flatten_config(cfg: Any, prefix: str = '') -> dict[str, Any]

Flatten nested dataclasses to {"data.symbol": "BTCUSDT", ...}.

Source code in nullres/runlog.py
def flatten_config(cfg: Any, prefix: str = "") -> dict[str, Any]:
    """Flatten nested dataclasses to {"data.symbol": "BTCUSDT", ...}."""
    out: dict[str, Any] = {}
    if is_dataclass(cfg):
        items = [(f.name, getattr(cfg, f.name)) for f in fields(cfg)]
    elif isinstance(cfg, dict):
        items = list(cfg.items())
    else:
        return {prefix.rstrip("."): cfg}

    for key, value in items:
        path = f"{prefix}{key}"
        if is_dataclass(value) or isinstance(value, dict):
            out.update(flatten_config(value, f"{path}."))
        elif isinstance(value, list):
            out[path] = ",".join(map(str, value))
        else:
            out[path] = value
    return out

config_hash

config_hash(cfg: Any) -> str

Stable fingerprint of everything that makes this a distinct experiment.

Source code in nullres/runlog.py
def config_hash(cfg: Any) -> str:
    """Stable fingerprint of everything that makes this a distinct experiment."""
    payload = json.dumps(_significant(flatten_config(cfg)), sort_keys=True, default=str)
    return hashlib.sha256(payload.encode()).hexdigest()[:12]

config_distance

config_distance(a: Any, b: Any) -> tuple[int, list[str]]

How many meaningful parameters differ, and which.

Used to answer "is this config a near-miss of something already killed". A distance of 0 means you are re-running an identical experiment; 1-3 means you are tuning around one that may already be dead.

Source code in nullres/runlog.py
def config_distance(a: Any, b: Any) -> tuple[int, list[str]]:
    """How many meaningful parameters differ, and which.

    Used to answer "is this config a near-miss of something already killed".
    A distance of 0 means you are re-running an identical experiment; 1-3 means
    you are tuning around one that may already be dead.
    """
    fa, fb = _significant(flatten_config(a)), _significant(flatten_config(b))
    keys = set(fa) | set(fb)
    differing = sorted(k for k in keys if str(fa.get(k)) != str(fb.get(k)))
    return len(differing), differing

count_trials

count_trials(runs: list[RunRecord], prior: int = 0, pending: tuple[str, str, int] | None = None) -> int

Distinct parameter combinations explored, across every recorded run.

This is deliberately GLOBAL rather than scoped to one config. The question multiple-testing correction asks is "how many things did you look at before reporting this one", and a researcher who would have published whichever of six configs worked has tried all six — not one.

It counts DISTINCT experiments, not executions. Summing every record made the correction a function of how often commands were run: repeating one xsec five times took the count 230 -> 270 and quietly lowered every reported deflated Sharpe without a single new hypothesis being tested, which also means a number quoted in the docs could not be reproduced later. Each (config fingerprint, command) pair therefore contributes once, at the largest variant count seen for it — a re-run is the same look, and a wider sweep of the same config is a bigger one.

prior declares trials that predate the ledger. Undercounting is the failure this whole function exists to fix, so an honest estimate of past work belongs here rather than a zero.

pending is the run about to happen, as (config_hash, command, variants). It is folded into the same dedupe rather than added on top, because adding it on top reintroduces the bug in miniature: re-running an experiment already in the ledger would nudge its own trial count up by its own size every time, so a reported deflated Sharpe drifted with each verification. Folded in, a re-run adds nothing and only a wider sweep of the same config raises the count.

Source code in nullres/runlog.py
def count_trials(runs: list[RunRecord], prior: int = 0,
                 pending: tuple[str, str, int] | None = None) -> int:
    """Distinct parameter combinations explored, across every recorded run.

    This is deliberately GLOBAL rather than scoped to one config. The question
    multiple-testing correction asks is "how many things did you look at before
    reporting this one", and a researcher who would have published whichever of
    six configs worked has tried all six — not one.

    It counts DISTINCT experiments, not executions. Summing every record made
    the correction a function of how often commands were run: repeating one
    `xsec` five times took the count 230 -> 270 and quietly lowered every
    reported deflated Sharpe without a single new hypothesis being tested, which
    also means a number quoted in the docs could not be reproduced later. Each
    (config fingerprint, command) pair therefore contributes once, at the
    largest variant count seen for it — a re-run is the same look, and a wider
    sweep of the same config is a bigger one.

    `prior` declares trials that predate the ledger. Undercounting is the
    failure this whole function exists to fix, so an honest estimate of past
    work belongs here rather than a zero.

    `pending` is the run about to happen, as `(config_hash, command, variants)`.
    It is folded into the same dedupe rather than added on top, because adding
    it on top reintroduces the bug in miniature: re-running an experiment
    already in the ledger would nudge its own trial count up by its own size
    every time, so a reported deflated Sharpe drifted with each verification.
    Folded in, a re-run adds nothing and only a wider sweep of the same config
    raises the count.
    """
    seen: dict[tuple[str, str], int] = {}
    if pending is not None:
        config_hash_, command, variants = pending
        seen[(config_hash_, command)] = max(int(variants), 1)
    for record in runs:
        # Calibration is not exploration. `configs/null.toml` runs the pipeline
        # on a random walk to prove the harness finds nothing; it tests no
        # hypothesis about any market, and counting it would mean that
        # verifying your instrument made every real result deflate harder.
        if str(record.config.get("data.source", "")) == "synthetic":
            continue
        key = (record.config_hash, record.command)
        declared = 1 if record.variants is None else max(record.variants, 1)
        seen[key] = max(seen.get(key, 0), declared)
    return prior + sum(seen.values())

unrecorded_variants

unrecorded_variants(runs: list[RunRecord]) -> int

Records written before variants existed, each counted as a single look.

Reported rather than repaired: the ledger is append-only and back-filling a guess would be worse than naming the gap. A non-zero count means the true multiple-testing exposure is HIGHER than count_trials returns.

Source code in nullres/runlog.py
def unrecorded_variants(runs: list[RunRecord]) -> int:
    """Records written before `variants` existed, each counted as a single look.

    Reported rather than repaired: the ledger is append-only and back-filling a
    guess would be worse than naming the gap. A non-zero count means the true
    multiple-testing exposure is HIGHER than `count_trials` returns.
    """
    return sum(1 for r in runs if r.variants is None)

record_run

record_run(cfg: Any, command: str, metrics: dict[str, Any] | None = None, verdict: str | None = None, notes: str = '', variants: int = 1, runs_dir: str = RUNS_DIR, repo: Path | None = None) -> RunRecord

Append one record. Never overwrites: the log is a ledger, not a cache.

Source code in nullres/runlog.py
def record_run(cfg: Any, command: str, metrics: dict[str, Any] | None = None,
               verdict: str | None = None, notes: str = "", variants: int = 1,
               runs_dir: str = RUNS_DIR, repo: Path | None = None) -> RunRecord:
    """Append one record. Never overwrites: the log is a ledger, not a cache."""
    repo = repo or Path.cwd()
    now = datetime.now(timezone.utc)
    sha, dirty = _git_state(repo)
    chash = config_hash(cfg)

    # Deterministic id from (config, command, time) so two runs never collide.
    seed = f"{chash}|{command}|{now.isoformat()}"
    run_id = hashlib.sha256(seed.encode()).hexdigest()[:12]

    record = RunRecord(
        id=run_id,
        timestamp=now.strftime("%Y-%m-%dT%H:%M:%SZ"),
        command=command,
        config_name=getattr(cfg, "name", "unnamed"),
        config_hash=chash,
        git_sha=sha,
        git_dirty=dirty,
        config=flatten_config(cfg),
        metrics=metrics or {},
        verdict=verdict,
        notes=notes,
        variants=max(int(variants), 1),      # always written; never left unset
    )

    out = Path(runs_dir)
    out.mkdir(parents=True, exist_ok=True)
    path = out / f"{now:%Y%m%d-%H%M%S}-{run_id[:8]}.json"
    path.write_text(json.dumps(asdict(record), indent=2, default=str))
    return record

load_runs

load_runs(runs_dir: str = RUNS_DIR) -> list[RunRecord]

Every record on disk, newest last. Corrupt files are skipped, not fatal.

Source code in nullres/runlog.py
def load_runs(runs_dir: str = RUNS_DIR) -> list[RunRecord]:
    """Every record on disk, newest last. Corrupt files are skipped, not fatal."""
    out = []
    for path in sorted(Path(runs_dir).glob("*.json")):
        try:
            raw = json.loads(path.read_text())
            out.append(RunRecord(**raw))
        except (json.JSONDecodeError, TypeError, ValueError):
            continue
    return out

find_similar

find_similar(cfg: Any, runs: list[RunRecord], max_distance: int = 3, verdict: str | None = 'KILLED') -> list[tuple[int, list[str], RunRecord]]

Past runs whose config is within max_distance parameters of this one.

This is the reason the machine layer exists. Nobody re-reads a 294-line markdown file before every experiment, so eighteen months from now the dead end gets re-run. A config comparison does not forget.

A change of data is a change of experiment. Counting every key equally made data.symbol: BTCUSDT -> DOGEUSDT a distance of 1, identical to min_hold: 84 -> 85 — so testing a dead rule on a completely different asset was flagged as re-treading a dead end, which it is not. Whatever killed a rule on BTC is not evidence about SOL. Runs whose data differs are therefore not near-misses at all, no matter how close the rest looks. The warning has to stay rare or it gets ignored, which costs more than it saves.

Source code in nullres/runlog.py
def find_similar(cfg: Any, runs: list[RunRecord], max_distance: int = 3,
                 verdict: str | None = "KILLED") -> list[tuple[int, list[str], RunRecord]]:
    """Past runs whose config is within `max_distance` parameters of this one.

    This is the reason the machine layer exists. Nobody re-reads a 294-line
    markdown file before every experiment, so eighteen months from now the dead
    end gets re-run. A config comparison does not forget.

    **A change of data is a change of experiment.** Counting every key equally
    made `data.symbol: BTCUSDT -> DOGEUSDT` a distance of 1, identical to
    `min_hold: 84 -> 85` — so testing a dead rule on a completely different
    asset was flagged as re-treading a dead end, which it is not. Whatever
    killed a rule on BTC is not evidence about SOL. Runs whose data differs are
    therefore not near-misses at all, no matter how close the rest looks. The
    warning has to stay rare or it gets ignored, which costs more than it saves.
    """
    hits = []
    for record in runs:
        if verdict and record.verdict != verdict:
            continue
        distance, differing = config_distance(cfg, record.config)
        if any(k.startswith("data.") for k in differing):
            continue
        if distance <= max_distance:
            hits.append((distance, differing, record))
    return sorted(hits, key=lambda h: h[0])

format_warning

format_warning(hits: list[tuple[int, list[str], RunRecord]]) -> str

Render the near-miss warning. Empty string when there is nothing to say.

Source code in nullres/runlog.py
def format_warning(hits: list[tuple[int, list[str], RunRecord]]) -> str:
    """Render the near-miss warning. Empty string when there is nothing to say."""
    if not hits:
        return ""
    lines = [
        f"  WARNING: this config is within {hits[0][0]} parameter(s) of "
        f"{len(hits)} run(s) already marked KILLED."
    ]
    for distance, differing, record in hits[:3]:
        changed = ", ".join(differing[:4]) or "nothing — this is an exact re-run"
        lines.append(f"    {record.timestamp[:10]}  {record.config_name:<16} "
                     f"{record.command:<8} [{record.short_id}]  differs by: {changed}")
    lines.append("    See docs/05-graveyard.md before spending time on this.")
    return "\n".join(lines)