Falsification¶
The machinery that tries to disprove a result rather than produce one.
Leak detection¶
audit ¶
Mechanical leakage detection.
The baseline script's closing line was right: walk-forward validation did not catch the leak, only reading the label definition would have. That is an unacceptable place to leave things — humans re-read code badly, and every new feature is a fresh chance to introduce lookahead.
These five checks catch the overwhelming majority of leaks automatically:
-
POINT-IN-TIME Recompute features using only bars <= t and assert row t is unchanged. Catches any use of future data in feature construction, including the subtle ones (a global mean, a backfill, an accidental negative shift).
-
LABEL/FEATURE Check whether any single feature predicts the label far CORRELATION too well on its own.
ret_1versus a same-bar label scores ~1.0 here, which is precisely the baseline's bug. -
NULL DATA Run the whole pipeline on a random walk. There is no edge by construction, so a positive result after costs proves a bug — in the engine, the split, or the labels.
-
SHUFFLED LABEL Retrain with labels randomly permuted. Out-of-sample accuracy must collapse to the base rate. If it does not, information is reaching the model through a side channel.
-
SURVIVORSHIP A multi-symbol universe spanning a period that killed assets, containing none of them, was filtered by survival. Reports
n/arather than PASS on a single-symbol config — it has nothing to test there, and a green tick would claim a risk was ruled out when it was never examined.
Run nullres audit before believing any result. It takes a minute and it has a
much better record than intuition.
check_point_in_time ¶
check_point_in_time(bars: DataFrame, builder=build_features, probes: int = 12, tol: float = 1e-09, seed: int = 0) -> Check
Recompute features on a truncated history and compare the final row.
If build_features(bars[:t+1]).iloc[t] differs from
build_features(bars).iloc[t], then the full-history version used data from
after bar t. There is no way for that to be legitimate.
The probes are random but seeded, so the check is reproducible — and so it
inspects the same bars on every run. That is the trade: a leak that only
manifests in one regime is found only if a probe lands in it, which is why
the count is more than a token few. Raise probes when adding a feature
whose behaviour is regime-dependent.
Source code in nullres/audit.py
check_label_leakage ¶
Flag any single feature that separates the label suspiciously well.
A lone technical indicator with AUC > 0.65 on a directional label is not a discovery. On real financial data, single-feature AUCs live in 0.50-0.55.
Source code in nullres/audit.py
check_null_data ¶
The pipeline must find nothing on a random walk.
run_pipeline(cfg) is injected to avoid a circular import.
Source code in nullres/audit.py
check_survivorship ¶
check_survivorship(symbols: Iterable[str], delisted: Mapping[str, Any] | None, point_in_time: Iterable[str] | None = None, hardcoded: bool = False, last_bar=None, sample_end=None, grace_days: int = 60) -> Check
Does this universe contain assets that died?
Backtesting a universe chosen from what is liquid today is a test of "things that survived", and it will produce a beautiful, meaningless result. The catalogue used to call this undetectable. It is not — not fully, but the dominant failure mode is mechanical:
A multi-symbol universe spanning a period that killed assets, which contains none of them, was filtered by survival.
What this CANNOT see is whether you picked the winners among the survivors. That is hindsight, and it stays yours to avoid — see the catalogue's entry 7.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
symbols
|
Iterable[str]
|
the universe actually traded. |
required |
delisted
|
Mapping[str, Any] | None
|
symbols whose data stops before the sample ends. |
required |
point_in_time
|
Iterable[str] | None
|
optionally, the universe as enumerated at the sample start. Lets the check measure how much of the graveyard was dropped. |
None
|
hardcoded
|
bool
|
True when the universe was a literal list rather than enumerated from the archive as of a date. |
False
|
Source code in nullres/audit.py
190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 | |
check_shuffled_label ¶
check_shuffled_label(X: DataFrame, y: Series, t_end: ndarray, split_cfg, model_cfg, tol: float = 0.02, seed: int = 0) -> Check
Permuted labels must be unlearnable.
Any accuracy above the base rate here means the model is reaching the target through something other than the features it was given.
Source code in nullres/audit.py
Robustness battery¶
robustness ¶
Falsification tests for a strategy that looked good once.
A single backtest number is a hypothesis, not a finding. Before it earns any more of your time it has to survive three attempts to kill it:
NEIGHBOURHOOD Do nearby parameter values also work? A real effect degrades smoothly. If only one cell in a grid is positive, you did not find an edge — you found the cell that happened to fit.
STABILITY Does it work in every year, or did one spectacular period carry the average? A Sharpe of 0.5 built from 2021 and nothing else is a bet that 2021 recurs.
TRANSFER Does it work on other symbols? A rule that describes market structure should generalise. A rule that only works on the asset you developed it on describes that asset's history.
Passing all three does not make a strategy real — the only test that does is forward paper trading.
What these tests can and cannot resolve. Two of the three rest on very few observations: five years, four symbols. A gate reading "beat the benchmark in 60% of periods" sounds demanding, but at n=5 it means three of five, and a strategy that is genuinely a coin flip against buy & hold clears it half the time. At n=4 symbols a coin flip clears it 31% of the time. That is not a detail — it means a bare count cannot tell "this is worse than holding" from "there is not enough evidence here", and for a long time this module reported both as KILLED.
So each count gate is now read alongside the magnitude of the shortfall, and
verdict has three outcomes rather than two. A gate FAILS only when the count
goes against it AND the mean excess is distinguishable from zero. A count that
fails alone yields INCONCLUSIVE, and every note prints how often the gate would
have fired by chance.
The decision rule stays aggressive on purpose — a strategy earns SURVIVED only by clearing all three — because a false kill costs one idea and a false survival costs months. That is a decision about which error to prefer, not a claim that the thresholds are statistically strong. They are not.
grid_for ¶
Return (grid, kind) where kind is 'params' or 'sizing'.
Source code in nullres/robustness.py
parameter_neighbourhood ¶
Sharpe across a grid of parameter values, on one prepared context.
Rule strategies read only the bars, so the context (features, labels, splits) is identical across the grid and is computed once.
Source code in nullres/robustness.py
period_stability ¶
Per-year performance, alongside buy & hold over the same periods.
The benchmark column is not decoration. "Lost money in 2022" and "lost less than half what holding lost in 2022" are opposite findings, and the bare per-year Sharpe cannot distinguish them. For a long-only trend filter evaluated over a historic bull market, this column is the whole argument.
Source code in nullres/robustness.py
hold_sharpe ¶
Buy & hold's Sharpe, on the same window AND the same statistic as the grid.
verdict reports what fraction of the neighbourhood beats buy & hold, and
that comparison is only meaningful if both sides are the same measurement.
They were not: the grid reports a full-window Sharpe while the benchmark was
taken as the mean of period_stability's per-year Sharpes. Averaging annual
Sharpes is a different statistic — on the 4h config it gives 0.53 against a
full-window 0.38, overstating the bar by 40%.
Source code in nullres/robustness.py
cross_symbol ¶
The same strategy and parameters, on other instruments.
Each symbol needs its own data, features and splits, so this is the expensive test — and the most informative one. It is what killed both previous candidates.
start overrides the config's start date for EVERY symbol including the
reference one. Auxiliary archives begin at different dates per symbol
(BTCUSDT open-interest metrics start 2020-09, everything else 2021-12), and
letting each symbol use its own maximum range would compare different eras
and call the difference "transfer".
Source code in nullres/robustness.py
sign_flip_pairs ¶
How many adjacent cell pairs sign_flip_rate measured over.
The rate alone cannot say whether it is distinguishable from chance; that needs the denominator too.
Source code in nullres/robustness.py
sign_flip_rate ¶
Fraction of adjacent grid cells whose Sharpe changes sign.
This is the measure the docs have always claimed to care about — "a real effect degrades SMOOTHLY; an isolated spike is a fitting artefact" — and which nothing actually computed. Counting positive cells cannot see the difference between a coherent positive region and a checkerboard: a grid reading +0.8 / -0.2 / +0.9 / -0.7 scores 50% positive either way.
Neighbours are cells one step apart along a single axis. A rate near 0 means a smooth surface; near 0.5 means the sign carries no information.
Source code in nullres/robustness.py
count_gate_power ¶
How often a strategy exactly as good as the benchmark clears a count gate.
"Beat the benchmark in 60% of periods" sounds demanding until you count the periods. With five years it means three of five, and a strategy that is genuinely a coin flip against buy & hold clears that half the time. With four symbols it is three of four, which a coin flip clears 31% of the time.
A gate this noisy cannot carry a verdict by itself, so the number is printed beside every count so the reader knows what the count is worth.
Source code in nullres/robustness.py
excess_magnitude ¶
(mean, p) for "is the mean excess distinguishable from zero".
The count gates throw away magnitude, and that loses real information: on
the 4h config donchian and mean_reversion both beat hold in 40% of years
and score identically, while their mean excess Sharpes are -0.04 and -1.13.
One is indistinguishable from holding; the other is far worse. A test on the
magnitude separates them; counting signs cannot.
It is not a replacement for the count, because it is less decisive on noisy
series — mean_reversion's -1.13 carries p=0.31 across five volatile years.
Neither statistic dominates, so both are reported.
Source code in nullres/robustness.py
verdict ¶
verdict(neighbourhood: DataFrame, stability: DataFrame, transfer: DataFrame, benchmark_sharpe: float | None = None, flip_rate: float | None = None, flip_pairs: int | None = None) -> tuple[str, list[str]]
Turn the three tables into KILLED, SURVIVED or INCONCLUSIVE, with reasons.
Three states, not two, because two states forced a claim the evidence does not support. The count gates rest on four or five Bernoulli draws: at n=5 a strategy genuinely equal to buy & hold fails the stability gate half the time. Reporting that as KILLED dressed a coin flip as a finding, and the verdict then propagated into the ledger and warned future runs off the config.
So a gate now FAILS only on decisive evidence — the count went against it AND the magnitude of the shortfall is distinguishable from zero. A count that fails on its own returns WEAK, and the run comes out INCONCLUSIVE.
The decision rule stays deliberately aggressive: for research triage a false kill is cheap and a false survival is expensive, so a strategy has to earn SURVIVED by clearing every gate. That is a decision-theoretic stance, not a claim that the thresholds are statistically demanding. They are not, and each note now says so.
Source code in nullres/robustness.py
434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 542 543 544 545 546 547 548 549 550 551 552 553 554 555 556 557 558 559 560 561 562 563 564 565 566 567 568 569 570 571 572 573 574 575 | |
pivot_grid ¶
Render a two-parameter grid as a Sharpe matrix.
Source code in nullres/robustness.py
Cross-sectional panels¶
crosssec ¶
Cross-sectional (relative-value) research on a panel of symbols.
Everything else in this repo asks "will BTC go up?". That question has to clear a 24bps round trip out of a single asset's own move, and six lines of attack have now died against that wall (docs/05-graveyard.md).
This asks a different question: "which of these assets will outperform the others?" It is easier in three specific ways.
- The market move cancels. A dollar-neutral book is not betting on crypto going up, so it does not need to out-predict a 60% annualised drift.
- The label is balanced by construction — half the universe beats the median at every timestamp, in every regime.
- Errors are relative. Being wrong about BTC and wrong about ETH in the same direction costs nothing; only the ranking matters.
What it costs you is a new failure mode: survivorship bias. Choosing a universe by looking at what is liquid today is a test of "assets that survived", and it will produce a beautiful, meaningless equity curve. The universe here is fixed as of 2021-12 and includes LUNAUSDT, which went to zero in May 2022 and was delisted. If your cross-sectional backtest cannot lose money on LUNA, it is not measuring anything.
Long/short also requires PERPETUAL FUTURES — you cannot short spot — so this module uses USD-M perp bars and charges the 8-hourly funding rate on every position held.
Panel
dataclass
¶
Panel(features: DataFrame, y: Series, ret_next: DataFrame, funding: DataFrame, times: DatetimeIndex, horizon: int, symbols: list[str] = list(), delisted: dict[str, Timestamp] = dict())
A tidy panel: MultiIndex (ts, symbol) features, plus per-bar returns.
load_panel ¶
load_panel(cfg, symbols: list[str] | None = None, verbose: bool = True, top_n: int | None = None, screen_window: int = 180) -> Panel
Load symbols, screen for liquidity, build features, stack into a panel.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
top_n
|
int | None
|
keep only the top-N symbols by TRAILING dollar volume at each bar. Without this a wide universe ranks BTC against coins that traded $50k a day, which is a ranking you could not act on. Screening on full-sample volume instead would be lookahead — it selects the coins that went on to matter. |
None
|
Source code in nullres/crosssec.py
68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 | |
time_folds ¶
Yield (train_times, test_times). Purged by the label horizon.
Splitting on row position would be wrong here: rows are (ts, symbol) pairs, so a positional split would put BTC's Tuesday in training and ETH's Tuesday in test. The model would learn the day from ten correlated siblings.
Source code in nullres/crosssec.py
fit_predict_panel ¶
fit_predict_panel(panel: Panel, cfg, verbose: bool = True)
Walk-forward P(outperforms) for every (ts, symbol) in a test fold.
Source code in nullres/crosssec.py
panel_positions ¶
panel_positions(proba: Series, panel: Panel, top_k: int = 3, rebalance: int = 42, allow_short: bool = True) -> DataFrame
Long the top-k ranked symbols, short the bottom-k. Dollar neutral.
Weights are +1/k and -1/k so gross exposure is 2 and net is 0 — the book makes no bet on the market's direction, which is the entire point.
That gross of 2.0 is deliberate, and sizing.max_leverage does not
constrain it. That field clips a single-asset position to +/-1 and is
never read here; a dollar-neutral book is 100% long and 100% short by
definition, and capping gross at 1.0 would mean holding half of each side,
which is a different strategy rather than a safer version of this one. But
2x gross notional is 2x notional whatever the net is: it requires margin,
and it is why the tail census in panelaudit matters. summarize reports
gross_exposure so this is visible in the results table instead of being
inferable only from reading this function.
rebalance throttles turnover exactly as min_hold does in the
single-asset engine. Reshuffling an 11-symbol book every bar is the
cross-sectional version of the mistake that cost the original baseline 100%.
On the shape of this function. The obvious spelling is one loop over
every timestamp, and that is what this was: ~42 seconds on a 9,000 x 120
panel, because positions.loc[ts] = ... reallocates and realigns a row at
a time and _neutralise was called once per bar. The book only changes
at rebalances, though — 214 of those 9,000 bars — and everything in between
is a carry-forward, a liveness mask and a rescale, all of which are array
operations over the whole panel.
So the ranking still runs bar by bar, on rebalance bars only, and still
through Series.sort_values: the selection has to break ties exactly as
before, and reimplementing that with argsort risks silently reordering
near-equal probabilities. Only the parts with no ordering subtlety were
vectorised. That is 0.2 seconds, and tests/test_crosssec.py asserts the
output is bit-identical to the original loop rather than merely close —
this book's headline Sharpe has already been shown to move by 60% on a
0.0003% change in its inputs, so "close" is not a standard worth anything
here.
Source code in nullres/crosssec.py
backtest_panel ¶
backtest_panel(positions: DataFrame, panel: Panel, cost_cfg, charge_funding: bool = True)
Per-bar net log return of the book, with per-symbol costs and funding.
Source code in nullres/crosssec.py
benchmarks ¶
benchmarks(panel: Panel, cost_cfg, oos_times: DatetimeIndex | None = None, rebalance: int = 42, reference: str = 'BTCUSDT') -> dict
The books a cross-sectional model has to beat to be worth anything.
All are restricted to oos_times. That restriction is not cosmetic: the ML
book only trades inside its test folds, so an unmasked benchmark carries the
entire pre-OOS period. Here that meant equal-weight absorbing the 2022 bear
market the model never touched, reporting -83% and making a mediocre
strategy look excellent by comparison.
static_vs_alts is the important one. It is long the reference asset and
short everything else, rebalanced never — no model, three trades. BTC
outperformed alts massively over 2022-2025, so any model that has learned
"the lowest-volatility member outperforms" has learned this bet under
another name, and must be measured against it.
Note that static_vs_alts is itself hindsight-selected: BTC is the long leg
because we know how the period ended. It is a lower bound on what a model
must beat, not a strategy.
Source code in nullres/crosssec.py
Panel controls¶
panelaudit ¶
Controls that decide whether cross-sectional skill is real.
A panel AUC above 0.5 has more ways of being an artefact than a single-asset one, and the interesting failures are not leaks — they are the model learning something true but useless:
SHUFFLED LABEL Permute the target within each timestamp and refit. The permutation preserves the balanced structure, so anything left is a side channel, not signal.
SURVIVORS ONLY Refit with the delisted symbols removed. If the AUC collapses, the model was detecting death, not ranking assets — a real regularity you cannot trade, because by the time a coin is dying its borrow has vanished.
PER-SYMBOL Cross-sectional ranks never name a symbol, but if BTC is permanently the lowest-volatility member then "rank 1 by low vol" and "BTC" are the same column. A wide spread in per-symbol accuracy is that tell.
CONTRIBUTION Which symbols actually produced the P&L, and how much of it came from the ones that delisted.
TAIL CENSUS A short book that never got hit is not a book with no tail risk. Count how often the moves that would hurt occur, and how many short-name-bars were exposed to them; the product is the number of hits chance predicts.
These were run once by hand and quoted in RESEARCH.md, which meant the numbers underneath the project's strongest result were the only ones no command could regenerate. That is exactly backwards.
shuffled_label_auc ¶
Mean fold AUC after permuting the label within each timestamp.
Permuting within a timestamp rather than globally keeps the label balanced in every regime, so the control isolates the ranking signal instead of also destroying the panel's structure.
Each timestamp draws from a stream seeded by (seed, that timestamp), so a
given bar receives the same permutation no matter where it falls in the
iteration. Sharing one generator across groups would have made the result
depend on processing order — reproducible only so long as nothing upstream
changed how the panel is grouped, which is not a property worth relying on
in a control whose whole job is to be trustworthy.
Source code in nullres/panelaudit.py
survivors_only_auc ¶
Mean fold AUC with every delisted symbol dropped.
A model that only knows which coins are dying has found something real and untradable. None when the universe contains no corpses to remove.
Source code in nullres/panelaudit.py
per_symbol_accuracy ¶
Directional accuracy per symbol, with the count it rests on.
Skill concentrated in one or two names is skill that has learned those names, whatever the ranks pretend. But the spread only means that if every symbol has enough scored bars to have an accuracy worth reading.
The count is not decoration. On a screened wide universe, symbols drift
in and out of the tradable set and some are scored on a few dozen bars. The
first version of this returned bare accuracies, and the 136-symbol panel
duly produced a spread of 0.685 — seven times the narrow panel's, and
entirely an artefact of thin symbols. At n=30 a coin flip reaches 0.86
without trying. Callers filter on n before quoting a spread.
Accuracy alone is the wrong statistic here, and this is subtle. The label is "beats the cross-sectional median", so a coin that persistently underperformed has a lopsided base rate of its own — say 0.85 zeros. A model that learned nothing but "this one usually lags" scores 0.85 on it. High per-symbol accuracy can therefore be pure unconditional drift, and reading the raw spread as "skill is concentrated in these names" overstates it.
Lift over the symbol's own majority class removes that, but introduces its own bias in exactly this setting: a cross-sectional model MUST rank, so at every timestamp roughly half the universe is predicted low. It structurally cannot predict the majority class for a symbol that beats the median 79% of the time, and gets charged a large negative lift for a constraint rather than a mistake.
So the spread worth reading is in per-symbol AUC. It is threshold-free and base-rate invariant: 0.5 means the model cannot tell this symbol's good bars from its bad ones, whatever its unconditional tendency and whatever the ranking forced. Accuracy, base rate and lift are kept alongside because they are what a reader expects to see, but the AUC column is the one that answers "is the skill concentrated in particular names".
Source code in nullres/panelaudit.py
pnl_contribution ¶
Gross log P&L attributable to each symbol.
Source code in nullres/panelaudit.py
delisted_share ¶
Share of P&L ACTIVITY, in absolute terms, from symbols that later delisted.
This is a share of absolute P&L, not of net profit, and the distinction
changes what the number means. Each symbol contributes |its P&L|, so a
+30% winner and a -30% loser both count as 30 rather than cancelling to
zero. The question being asked is "how much of what this book did happened
in coins that were dying" — exposure, not profitability.
Netting would answer a different and weaker question. A book that made a fortune on one delisting and lost it on another would net to ~0% and look untouched by delisting, when in fact its entire outcome hinged on dying coins. For a survivorship control that is the wrong answer, so the metric deliberately does not net.
The consequence to keep in mind when reading it: this number cannot be compared against a return, and it can be large while the delisted names contributed nothing to the bottom line.
Source code in nullres/panelaudit.py
tail_census ¶
How many moves big enough to matter occurred, and how exposed the book was.
Observing no blow-up means nothing until you know how many blow-ups chance predicted. If the expected count is a fraction of one, zero hits is what chance produces and the tail is untested rather than absent.
threshold is one point on a curve, and a single point is a constant doing
analytical work — the thing this repo keeps having to remove. Prefer
tail_curve, which sweeps it and reports the capital each level would cost,
so no one number carries the argument.
Source code in nullres/panelaudit.py
concentration ¶
How often a dead leg left the book concentrated, and for how long.
_neutralise keeps the book dollar-neutral when a symbol delists by
rescaling the surviving side, so a k=2 book whose short partner dies holds
-1.0 in one name instead of -0.5 in two. Gross exposure does not change and
net stays at zero — what changes is that the move which ruins the book
halves, from +200% to +100%. Nine symbols delist in the wide universe, so
this is not hypothetical.
Keeping that behaviour is a choice: dollar neutrality is the book's defining constraint, and the alternatives (halve the long side, go flat) change the strategy rather than make it safer. What was missing is visibility. A maximum weight on its own cannot distinguish one bar from three thousand — the difference between a curiosity and the dominant risk in the book — so this reports the share of bars spent concentrated and the longest unbroken stretch of it.
Source code in nullres/panelaudit.py
tail_curve ¶
Expected tail hits across move sizes, with what each would cost.
Two questions the single-threshold census could not answer.
How often? A rate estimated at one move size is one point on a steeply
falling curve, and which point you pick decides whether the answer sounds
reassuring. Sweeping removes the choice — the same reason nullres sweep
prints a surface instead of its maximum.
How bad? The graveyard works this out by hand: "one UNFI-type event
(+274% in 4h) against a -0.5 weight is -137% of capital". That arithmetic is
mechanical and belongs in code. cost_of_one multiplies each move by the
largest short weight the book actually held, so the loss is the book's own
rather than an illustration, and ruinous marks the levels that would take
more than all of it.
This is as far as tail risk can honestly be taken here. It says how exposed the book was and what one hit would have cost — it does not model margin, liquidation price, or auto-deleveraging, none of which are in the archive.
Source code in nullres/panelaudit.py
format_report ¶
format_report(panel, cfg, proba, positions, mean_auc: float, min_obs: int = 200, nominal_weight: float | None = None) -> str
Run every control and render it. The order is cheapest-first.
Source code in nullres/panelaudit.py
337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 | |
The run ledger¶
runlog ¶
The evidence layer beneath the graveyard.
Two things are being kept, and keeping them apart is the whole design:
runs/*.json machine-written, append-only, never hand-edited.
What was measured, under exactly which config, at
which commit. Boring by design.
docs/05-graveyard.md hand-written. WHY it died and what it means. That
judgement is the actual work and stays human.
Neither half works alone. Prose without evidence rots — six months on, nobody can reproduce "mean AUC 0.5443". Evidence without prose is a spreadsheet that teaches nothing: no log entry will ever contain "one bear market wearing a trend-following costume". Graveyard entries cite run ids; run records point back at the graveyard.
The one capability that genuinely needs the machine layer is recognising that you are about to re-run something you already killed. Markdown cannot do that.
RunRecord
dataclass
¶
RunRecord(id: str, timestamp: str, command: str, config_name: str, config_hash: str, git_sha: str, git_dirty: bool, config: dict[str, Any], metrics: dict[str, Any] = dict(), verdict: str | None = None, notes: str = '', variants: int | None = None)
One execution of one command against one config.
flatten_config ¶
Flatten nested dataclasses to {"data.symbol": "BTCUSDT", ...}.
Source code in nullres/runlog.py
config_hash ¶
Stable fingerprint of everything that makes this a distinct experiment.
config_distance ¶
How many meaningful parameters differ, and which.
Used to answer "is this config a near-miss of something already killed". A distance of 0 means you are re-running an identical experiment; 1-3 means you are tuning around one that may already be dead.
Source code in nullres/runlog.py
count_trials ¶
count_trials(runs: list[RunRecord], prior: int = 0, pending: tuple[str, str, int] | None = None) -> int
Distinct parameter combinations explored, across every recorded run.
This is deliberately GLOBAL rather than scoped to one config. The question multiple-testing correction asks is "how many things did you look at before reporting this one", and a researcher who would have published whichever of six configs worked has tried all six — not one.
It counts DISTINCT experiments, not executions. Summing every record made
the correction a function of how often commands were run: repeating one
xsec five times took the count 230 -> 270 and quietly lowered every
reported deflated Sharpe without a single new hypothesis being tested, which
also means a number quoted in the docs could not be reproduced later. Each
(config fingerprint, command) pair therefore contributes once, at the
largest variant count seen for it — a re-run is the same look, and a wider
sweep of the same config is a bigger one.
prior declares trials that predate the ledger. Undercounting is the
failure this whole function exists to fix, so an honest estimate of past
work belongs here rather than a zero.
pending is the run about to happen, as (config_hash, command, variants).
It is folded into the same dedupe rather than added on top, because adding
it on top reintroduces the bug in miniature: re-running an experiment
already in the ledger would nudge its own trial count up by its own size
every time, so a reported deflated Sharpe drifted with each verification.
Folded in, a re-run adds nothing and only a wider sweep of the same config
raises the count.
Source code in nullres/runlog.py
unrecorded_variants ¶
unrecorded_variants(runs: list[RunRecord]) -> int
Records written before variants existed, each counted as a single look.
Reported rather than repaired: the ledger is append-only and back-filling a
guess would be worse than naming the gap. A non-zero count means the true
multiple-testing exposure is HIGHER than count_trials returns.
Source code in nullres/runlog.py
record_run ¶
record_run(cfg: Any, command: str, metrics: dict[str, Any] | None = None, verdict: str | None = None, notes: str = '', variants: int = 1, runs_dir: str = RUNS_DIR, repo: Path | None = None) -> RunRecord
Append one record. Never overwrites: the log is a ledger, not a cache.
Source code in nullres/runlog.py
load_runs ¶
load_runs(runs_dir: str = RUNS_DIR) -> list[RunRecord]
Every record on disk, newest last. Corrupt files are skipped, not fatal.
Source code in nullres/runlog.py
find_similar ¶
find_similar(cfg: Any, runs: list[RunRecord], max_distance: int = 3, verdict: str | None = 'KILLED') -> list[tuple[int, list[str], RunRecord]]
Past runs whose config is within max_distance parameters of this one.
This is the reason the machine layer exists. Nobody re-reads a 294-line markdown file before every experiment, so eighteen months from now the dead end gets re-run. A config comparison does not forget.
A change of data is a change of experiment. Counting every key equally
made data.symbol: BTCUSDT -> DOGEUSDT a distance of 1, identical to
min_hold: 84 -> 85 — so testing a dead rule on a completely different
asset was flagged as re-treading a dead end, which it is not. Whatever
killed a rule on BTC is not evidence about SOL. Runs whose data differs are
therefore not near-misses at all, no matter how close the rest looks. The
warning has to stay rare or it gets ignored, which costs more than it saves.
Source code in nullres/runlog.py
format_warning ¶
format_warning(hits: list[tuple[int, list[str], RunRecord]]) -> str
Render the near-miss warning. Empty string when there is nothing to say.