Commands¶
One function per command. Each computes and returns a result object; none of them print.
api ¶
The programmatic entry points. One function per command, returning data.
from nullres import load_config, run
result = run(load_config("configs/btc_4h.toml"))
result.metrics["donchian"]["sharpe"]
Every function here computes and returns a nullres.results object. None of
them print, and none of them format — nullres.report does that, and the CLI
is a thin layer over the two. That separation is what makes the commands
callable from a notebook, testable without capturing stdout, and documentable
without pasting a terminal transcript.
These functions append to the run ledger by default. That is not an
accident of implementation: the ledger is what deflated_sharpe reads to find
out how many variants were tried, and a run that goes unrecorded undercounts
the exposure and flatters every result that follows it. Pass record=False
when you are genuinely not testing a hypothesis — re-deriving a number for a
plot, say — and be honest about which case you are in.
run ¶
run(cfg: RunConfig, n_trials: int | None = None, ablate: str | None = None, record: bool = True, verbose: bool = True) -> RunResult
Backtest every configured strategy, plus buy & hold, out of sample.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
n_trials
|
int | None
|
override the multiple-testing count. Defaults to reading the
ledger, which is the honest answer — see |
None
|
ablate
|
str | None
|
drop a feature group after row alignment, so the rows, splits
and benchmark stay byte-identical and the only variable is the
feature set. Turning the data off in the config is NOT a controlled
comparison; see |
None
|
Source code in nullres/api.py
audit ¶
audit(cfg: RunConfig, verbose: bool = True) -> AuditResult
The five mechanical leak checks.
Run before believing any result. It takes a minute and it has a much better record than intuition.
Source code in nullres/api.py
budget ¶
budget(cfg: RunConfig) -> BudgetResult
What accuracy would this instrument and cost structure actually require?
Run this FIRST. It is arithmetic, it takes two seconds, and it will tell you whether the thing you are about to attempt is possible at all.
Source code in nullres/api.py
robust ¶
robust(cfg: RunConfig, strategy: str, symbols: list[str], transfer_start: str | None = None, record: bool = True, verbose: bool = False) -> RobustnessResult
Three independent attempts to kill a strategy that looked good once.
Neighbourhood, sub-period stability, and cross-symbol transfer. Passing all three does not make a strategy real — the only test that does is forward paper trading — but failing any one of them is cheap information.
Source code in nullres/api.py
sweep ¶
sweep(cfg: RunConfig, strategy: str, entries=None, holds=None, record: bool = True, verbose: bool = True) -> SweepResult
Threshold sensitivity — read the SHAPE, not the peak.
A real edge degrades smoothly as the entry threshold moves. A lone spike surrounded by losses is a fitting artefact, and picking it is how you turn a backtest into fiction.
Source code in nullres/api.py
ablate ¶
ablate(cfg: RunConfig, group: str = 'derivatives', record: bool = True, verbose: bool = False) -> AblationResult
Matched-sample A/B on AUC for one feature group.
Sharpe cannot answer this question. With ~80 trades an equity curve swings from -0.68 to +0.43 on feature sets whose AUC differs by one point.
Source code in nullres/api.py
xsec ¶
xsec(cfg: RunConfig, symbols: list[str] | None = None, universe_month: str | None = None, top_n: int | None = None, top_k: int | None = None, rebalance: int = 42, verify: bool = False, n_trials: int | None = None, hardcoded: bool | None = None, record: bool = True, verbose: bool = True) -> XsecResult
Cross-sectional long/short on a panel of symbols.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
symbols
|
list[str] | None
|
explicit universe. Overrides |
None
|
universe_month
|
str | None
|
enumerate the universe mechanically from the archive as of this month (YYYY-MM), including symbols that later died. This is the survivorship-honest option; a hardcoded list is not. |
None
|
top_k
|
int | None
|
symbols long and short per side. Defaults to a sweep, because the point of a wide universe is that the same signal can be expressed through diversification instead of concentration. |
None
|
verify
|
bool
|
run the controls in |
False
|
Source code in nullres/api.py
419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 | |
feature_importance ¶
feature_importance(cfg: RunConfig, verbose: bool = True) -> FeatureImportanceResult
Permutation importance on the last fold's test window.
In-sample importances tell you what the model memorised. This tells you what carried out of sample, which is a much shorter list.
Source code in nullres/api.py
fetch ¶
fetch(cfg: RunConfig) -> FetchResult
Download and cache bars, plus any configured auxiliary data.
Source code in nullres/api.py
ledger ¶
ledger(verdict: str | None = None, limit: int = 25) -> LedgerView
Read the run ledger. verdict filters; limit is a display hint.
Source code in nullres/api.py
resolve_universe ¶
resolve_universe(cfg: RunConfig, symbols: list[str] | None = None, universe_month: str | None = None) -> tuple[list[str], bool]
Decide which symbols a panel covers. Returns (symbols, hardcoded).
hardcoded is True when the universe came from a literal list rather than
being enumerated from the archive as of a date. It is not a detail:
audit.check_survivorship reports it, because a hardcoded list is exactly
how a universe ends up filtered by survival.
Split out from xsec so a caller can know the universe size before paying
for load_panel, which is the slowest thing in this repository.
Source code in nullres/api.py
verify_panel ¶
verify_panel(panel, cfg, proba, positions, mean_auc: float, min_obs: int = 200, nominal_weight: float | None = None) -> PanelVerification
Run every cross-sectional control and return the numbers.
These were run once by hand and quoted in RESEARCH.md, which meant the numbers underneath the project's strongest result were the only ones no command could regenerate. That is exactly backwards.
Source code in nullres/api.py
killed_warning ¶
Has something within a few parameters of this config already been killed?
Empty string when there is nothing to say. This is the reason the machine ledger exists: nobody re-reads a 500-line graveyard before every experiment, and eighteen months from now the dead end gets re-run.