Data and features¶
Configuration¶
config ¶
Typed configuration loaded from TOML.
TOML is read with the stdlib tomllib, so config costs no dependency. (It
landed in 3.11; the project's tested floor is 3.13 — see requires-python.)
Every experiment is fully described by one file in configs/ — that file is the
unit of reproducibility. If a number influenced a result, it belongs here and
not in a function default.
SizingConfig
dataclass
¶
SizingConfig(long_entry: float = 0.56, long_exit: float = 0.5, short_entry: float = 0.44, short_exit: float = 0.5, allow_short: bool = True, min_hold: int = 6, max_leverage: float = 1.0, vol_target: float = 0.0)
Signal -> position. This is where a real edge is kept or destroyed.
load_config ¶
Load and validate one experiment's config.
A missing or malformed file is a ConfigError, not a traceback. nullres run
-c typo.toml used to end in a raw FileNotFoundError from deep inside
pathlib, which tells the reader where Python gave up rather than what they
got wrong.
Source code in nullres/config.py
Loading¶
data ¶
Market data loading.
Everything here returns the same contract: a DataFrame indexed by UTC timestamp
with float columns [open, high, low, close, volume, trades], strictly increasing
index, no duplicates. load_bars is the only entry point callers should need.
fetch_month ¶
fetch_month(symbol: str, interval: str, month: str, cache_dir: str = 'data', retries: int = 3, market: str = 'spot') -> DataFrame | None
Return one month of klines, from cache when present. None if unavailable.
Source code in nullres/data/binance.py
load_binance ¶
load_binance(symbol: str, interval: str, start: str, end: str, cache_dir: str = 'data', verbose: bool = True, market: str = 'spot', required: bool = True) -> DataFrame | None
Load a contiguous range of months and validate the result.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
required
|
bool
|
when False, return None instead of raising if the symbol has no data at all. Delisted assets must be loadable-or-absent rather than fatal — excluding them because they stopped existing is the definition of survivorship bias. |
True
|
Source code in nullres/data/binance.py
load_funding ¶
load_funding(symbol: str, start: str, end: str, cache_dir: str = 'data', verbose: bool = True) -> DataFrame
8-hourly funding rates, indexed by settlement time (UTC).
The index is the moment the rate was SETTLED, which is the moment it became
known. Callers must join with that in mind — see features/derivatives.py.
Source code in nullres/data/futures.py
load_metrics ¶
load_metrics(symbol: str, start: str, end: str, cache_dir: str = 'data', workers: int = 8, verbose: bool = True) -> DataFrame
Open interest and long/short ratios at 5-minute granularity.
Binance publishes these as one archive PER DAY (the monthly path 404s), so
a six-year range is ~2,000 requests. They are fetched concurrently and
cached one parquet per month — caching per day would leave 2,000 files in
data/ for no benefit.
Source code in nullres/data/futures.py
synthetic_bars ¶
synthetic_bars(n: int = 40000, seed: int = 0, interval: str = '1h', sigma: float = 0.004, edge: float = 0.0, start: str = '2020-01-01') -> DataFrame
Geometric random walk in OHLCV form.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
edge
|
float
|
AR(1) coefficient on log returns. 0.0 means a pure martingale — unpredictable by construction. ~0.05 is a faint but genuinely learnable edge; real liquid markets sit near 0.0 to 0.02. |
0.0
|
Source code in nullres/data/synthetic.py
synthetic_funding ¶
Funding settlements every 8h, uncorrelated with future returns.
Exists so the null control exercises the SAME code path as a real run. Without it the random-walk check runs on 32 features while the live config runs on 47, and a broken funding join would never reach the one test whose whole job is to catch fabricated edge.
The values are noise by construction, so any strategy that profits from them on this data has found a bug in the join, not a signal.
Source code in nullres/data/synthetic.py
synthetic_metrics ¶
Open interest and positioning ratios, also pure noise.
Open interest is generated as a random walk with drift so it is
non-stationary like the real thing — that way the stationarity discipline
in features/derivatives.py is genuinely exercised.
Source code in nullres/data/synthetic.py
load_bars ¶
Dispatch on cfg.source and return bars matching the OHLCV contract.
Source code in nullres/data/__init__.py
load_auxiliary ¶
Return (funding, metrics), either of which may be None.
For synthetic data the auxiliary frames are generated as pure noise rather than skipped. That keeps the null control running the SAME feature pipeline as a live config — otherwise the random-walk check would exercise 32 features while the real run uses 46, and a broken funding join would never reach the one test designed to catch fabricated edge.
Source code in nullres/data/__init__.py
binance ¶
Binance public monthly kline archives (data.binance.vision).
No API key, no rate limit, no exchange account. Each month is cached to parquet so a re-run is offline and instant.
fetch_month ¶
fetch_month(symbol: str, interval: str, month: str, cache_dir: str = 'data', retries: int = 3, market: str = 'spot') -> DataFrame | None
Return one month of klines, from cache when present. None if unavailable.
Source code in nullres/data/binance.py
load_binance ¶
load_binance(symbol: str, interval: str, start: str, end: str, cache_dir: str = 'data', verbose: bool = True, market: str = 'spot', required: bool = True) -> DataFrame | None
Load a contiguous range of months and validate the result.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
required
|
bool
|
when False, return None instead of raising if the symbol has no data at all. Delisted assets must be loadable-or-absent rather than fatal — excluding them because they stopped existing is the definition of survivorship bias. |
True
|
Source code in nullres/data/binance.py
futures ¶
Binance USD-M futures data: funding rates and open-interest metrics.
This is the first information in the repo that is NOT a transform of OHLCV.
Everything in features/technical.py is thirty-two views of four numbers;
nullres features showed almost none of them carry out of sample. Funding and
open interest measure something the price series cannot express: how much
leverage is deployed, and on which side.
FUNDING RATE Perpetual futures have no expiry, so an 8-hourly payment
tethers them to spot. Positive funding means longs pay
shorts — the crowd is long and paying to stay there. It is a
direct read on positioning, and it is a PRICE, so it is set
by people with money at risk.
OPEN INTEREST Total outstanding contracts. Rising OI into a rally means
new money; falling OI means an unwind. Same price move,
opposite implication.
LONG/SHORT Binance publishes account- and position-weighted long/short
RATIOS ratios, including a top-trader subset.
Availability (probed, not assumed): funding monthly archives, BTCUSDT from 2020-01 metrics DAILY archives only, BTCUSDT from 2020-09; monthly 404s
Caveat worth stating plainly: these describe the PERPETUAL market, while the bars elsewhere in this repo are spot. That is a legitimate pairing — futures positioning predicting spot price is the whole idea — but they are different venues, and the perp can dislocate from spot precisely when it matters most.
load_funding ¶
load_funding(symbol: str, start: str, end: str, cache_dir: str = 'data', verbose: bool = True) -> DataFrame
8-hourly funding rates, indexed by settlement time (UTC).
The index is the moment the rate was SETTLED, which is the moment it became
known. Callers must join with that in mind — see features/derivatives.py.
Source code in nullres/data/futures.py
load_metrics ¶
load_metrics(symbol: str, start: str, end: str, cache_dir: str = 'data', workers: int = 8, verbose: bool = True) -> DataFrame
Open interest and long/short ratios at 5-minute granularity.
Binance publishes these as one archive PER DAY (the monthly path 404s), so
a six-year range is ~2,000 requests. They are fetched concurrently and
cached one parquet per month — caching per day would leave 2,000 files in
data/ for no benefit.
Source code in nullres/data/futures.py
universe ¶
Point-in-time universe construction.
Writing down a list of symbols from memory is hindsight dressed as data — you will recall the ones that survived. The universe here is built mechanically: enumerate everything the archive holds, then ask each symbol whether it had data in a given month. A coin that listed in 2023 fails that test; a coin that died in 2022 passes it, and belongs in the sample.
Liquidity screening is a separate problem and a subtler one. Ranking by
full-sample average volume is lookahead — it knows which coins would go on to
matter. liquidity_screen ranks on a TRAILING window only, so the universe at
each bar is the one you could actually have chosen at that bar.
list_symbols ¶
Every symbol the archive holds for market.
Source code in nullres/data/universe.py
universe_as_of ¶
universe_as_of(month: str, interval: str = '4h', market: str = 'um', workers: int = 24, cache_dir: str = 'data', verbose: bool = True) -> list[str]
Symbols that were trading in month — nothing about what came after.
Cached, because 787 HEAD requests is rude to repeat and the answer for a past month never changes.
Source code in nullres/data/universe.py
delisted_from_cache ¶
delisted_from_cache(symbols: list[str], interval: str, end: str, cache_dir: str = 'data', market: str = 'um', grace_months: int = 2) -> dict[str, str]
Symbols whose cached archive stops well before the sample ends.
Works entirely off local parquet files, so the survivorship check runs
offline and costs nothing. grace_months absorbs the normal lag between
the end of a range and the archive catching up — without it, every symbol
looks delisted in the current month.
Source code in nullres/data/universe.py
liquidity_screen ¶
liquidity_screen(volumes: DataFrame, top_n: int = 40, window: int = 180, min_history: int = 180) -> DataFrame
Boolean mask: is this symbol in the top-N by TRAILING dollar volume?
The trailing window is what makes this point-in-time. Screening on full-sample volume would quietly select the coins that went on to become important — a survivorship bias that hides inside what looks like ordinary data hygiene.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
volumes
|
DataFrame
|
ts x symbol quote volume per bar. |
required |
window
|
int
|
bars of history the ranking is computed over. |
180
|
min_history
|
int
|
a symbol needs at least this much history to be eligible, so newly listed coins are not ranked on three days of launch hype. |
180
|
Source code in nullres/data/universe.py
synthetic ¶
Synthetic bars with known ground truth.
Two uses, both essential:
synthetic_bars — a geometric random walk. By construction there is NO edge.
Any strategy that profits on this after costs has a bug.
This is the single most useful test in the repo.
synthetic_bars(edge=...) — a walk with a real, known autocorrelation. If your
pipeline CANNOT find this, it is too weak or mis-wired,
and a null result on real data tells you nothing.
Run both before trusting any result. A harness that fails either is not measuring what you think it is.
synthetic_funding ¶
Funding settlements every 8h, uncorrelated with future returns.
Exists so the null control exercises the SAME code path as a real run. Without it the random-walk check runs on 32 features while the live config runs on 47, and a broken funding join would never reach the one test whose whole job is to catch fabricated edge.
The values are noise by construction, so any strategy that profits from them on this data has found a bug in the join, not a signal.
Source code in nullres/data/synthetic.py
synthetic_metrics ¶
Open interest and positioning ratios, also pure noise.
Open interest is generated as a random walk with drift so it is
non-stationary like the real thing — that way the stationarity discipline
in features/derivatives.py is genuinely exercised.
Source code in nullres/data/synthetic.py
synthetic_bars ¶
synthetic_bars(n: int = 40000, seed: int = 0, interval: str = '1h', sigma: float = 0.004, edge: float = 0.0, start: str = '2020-01-01') -> DataFrame
Geometric random walk in OHLCV form.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
edge
|
float
|
AR(1) coefficient on log returns. 0.0 means a pure martingale — unpredictable by construction. ~0.05 is a faint but genuinely learnable edge; real liquid markets sit near 0.0 to 0.02. |
0.0
|
Source code in nullres/data/synthetic.py
cache ¶
Writing to the parquet cache without leaving corpses behind.
crosssec._guard_metrics_fetch refuses to start a download it estimates at
four hours, on the grounds that a silent multi-hour fetch is not something a
tool should do to you. The corollary went unhandled: a download that long WILL
be interrupted — Ctrl-C, a laptop lid, an OOM kill — and DataFrame.to_parquet
writes in place. An interrupt part-way through leaves a truncated file at
exactly the path every later run treats as authoritative.
That failure is worse than a missing file in three ways. It is permanent, since nothing ever rewrites a path that already exists. It is silent until the next read, which may be days later. And it surfaces as a parquet decode error deep inside pyarrow, naming neither the cache nor the download that produced it — so the obvious reading is "the library is broken", not "delete this one file".
Writing to a temporary file in the same directory and renaming it into place
fixes it. os.replace is atomic on POSIX and on Windows, so a reader sees
either the previous file or the complete new one, never a partial write.
write_parquet_atomic ¶
Write df to path so that readers never observe a partial file.
The temporary file is created beside the target rather than in the system
temp directory, because os.replace is only atomic within one filesystem
and a cache directory may well be on a different mount.
Source code in nullres/data/cache.py
read_parquet_or_discard ¶
Read a cached parquet, deleting and reporting it if it is unreadable.
Files written before write_parquet_atomic may already be truncated, and a
corrupt cache entry should cost one re-download rather than an afternoon of
reading tracebacks. Returning None means "treat this as a cache miss".
Source code in nullres/data/cache.py
Features¶
technical ¶
Feature engineering.
Two hard rules, both enforced by nullres.audit:
-
POINT-IN-TIME. Every value at bar t uses only bars <= t. In practice this means: rolling windows only, never
.shift(-k), never an expanding stat over the full sample, neverfillna(method="bfill"), never a global mean/std for scaling.audit.check_point_in_timerecomputes features on truncated data and asserts the last row is unchanged. -
STATIONARY. No raw price levels. BTC ran 4k -> 100k over this sample; a tree that learned "close > 60000" learned the calendar, not the market. Everything below is a ratio, a z-score, or a bounded oscillator.
rsi ¶
Wilder RSI, bounded 0..100.
The zero-loss case has to be handled explicitly. up / dn divides by zero
whenever the window contains no down moves, and mapping that to NaN — the
obvious defensive spelling — is wrong twice over. RSI is 100 there by
definition, not unknown; and because pipeline.prepare keeps only rows
where EVERY feature is present, one NaN here evicts the whole bar and the
other 45 features with it. A silent divide-by-zero became silent sample
loss, concentrated in exactly the strong uptrends a momentum feature is
supposed to describe.
Source code in nullres/features/technical.py
build_features ¶
Return the feature matrix, indexed identically to df.
Leading rows are NaN until the longest window fills; callers drop them.
funding and metrics are optional Binance futures frames. When supplied,
derivative features are appended — see features/derivatives.py, where the
point-in-time join is the part that matters.
Source code in nullres/features/technical.py
derivatives ¶
Features from funding rates and open interest.
THE JOIN IS THE WHOLE PROBLEM. Everything else here is arithmetic.
A bar indexed at time T covers [T, T+interval) and CLOSES at T+interval. So a funding settlement or OI reading may be used for bar T only if its timestamp is strictly before T+interval. Joining on the bar's OPEN time throws away a bar of information; joining on anything at or after the close is lookahead, and it is the kind that produces a beautiful equity curve.
We use merge_asof(direction="backward", allow_exact_matches=False) against
the bar's close instant. Exact matches are excluded because a settlement
stamped exactly at the close is simultaneous with it, and "simultaneous" is not
"available".
The auxiliary frames are also CLIPPED to the bar range before joining. That is
not cosmetic: audit.check_point_in_time truncates the bars and recomputes, and
without clipping the funding frame would still hold future rows, so a bad join
direction would silently produce identical output and the check would pass. With
clipping, the audit covers this surface too — tests/test_derivatives.py proves
it by injecting a forward join and asserting it gets caught.
build_derivative_features ¶
build_derivative_features(bars: DataFrame, funding: DataFrame | None = None, metrics: DataFrame | None = None) -> DataFrame
Stationary features from funding and open-interest data.
Levels are avoided throughout. Open interest grew ~10x over this sample; a model that learned "OI > 400k" learned the calendar, exactly as it would have from raw price.
Source code in nullres/features/derivatives.py
Labels¶
targets ¶
Label construction.
Every label returns a frame with a uniform contract:
y int 0/1 target, NaN where the bar is unlabelled (dropped later)
t_end int positional index of the bar at which the label RESOLVES
ret float the log return the label is derived from, for diagnostics
sigma float volatility estimate at decision time, known at bar t
t_end is the load-bearing column. A label spanning bars t..t+20 must not sit
in a training set whose test window begins at t+5 — the training label already
contains the answer to the test period. nullres.validation purges on this column.
A fixed purge constant is only correct when every label has the same horizon,
which stops being true the moment you use barriers.
On label choice: next_bar_sign is the honest version of the baseline's label,
and it is almost pure noise. A 1h BTC bar's empirical mean absolute move is
~0.40% against a ~0.24% round trip, so you are asking a model to call a coin
flip well enough to clear 60% of the move. (nullres budget quotes 45% instead,
because it uses the Gaussian E|move| of 0.54%; fat tails make the real move
smaller and the real bar higher — see docs/03.) triple_barrier instead asks a
question worth answering —
"does price travel 1.5 sigma up before it travels 1.5 sigma down" — which has
a real, if small, autocorrelation structure and a payoff that exceeds costs.
next_bar_sign ¶
1 if the next bar's close exceeds this one's. The baseline's label.
Kept for comparison, not recommended. Resolves one bar ahead.
Source code in nullres/labels/targets.py
fwd_return ¶
Sign of the vol-scaled return over horizon bars.
Bars whose move is smaller than deadband sigma are left unlabelled. That
matters: without it, roughly half the training set is noise the model tries
to fit, and the fit it finds is spurious.
Source code in nullres/labels/targets.py
triple_barrier ¶
López de Prado triple barrier, vectorised over the horizon.
From the close of bar t, place a profit barrier at +uppersigma and a stop
at -lowersigma, plus a vertical barrier horizon bars out. Label 1 if the
upper barrier is touched first, 0 if the lower is, and by the sign of the
realised return if the vertical barrier is reached first.
The barriers are volatility-scaled, so the label means the same thing in a calm 2023 and a violent March 2020 — a fixed 1% target is a different question in each regime, and mixing the two is why fixed-percent labels train models that only work in the regime that dominated the sample.
When both barriers fall inside one bar, OHLC cannot tell us which came first. We assume the STOP hit first. That is pessimistic by design: the alternative silently inflates every result you will ever produce here.
Source code in nullres/labels/targets.py
Models¶
classifier ¶
Model construction and out-of-sample prediction.
Every .fit() in this repository is a place that could accidentally train on
the future, so the set of them is kept small, deliberate, and pinned by
tests/test_packaging.py::test_no_unaudited_fit_sites. There are three:
classifier.fit_predict_walk_forward the single-asset walk-forward
classifier.feature_importance refits the last fold to permute it
crosssec.fit_predict_panel the panel walk-forward, split on TIME
The third is easy to miss and long went unmentioned — the docs claimed a single call site while the cross-sectional path, which produced the strongest result in the project, had its own. Each is purged independently, so nothing leaks; the risk was that a fourth could appear without anyone noticing. The test now fails if one does.
make_model ¶
Build an unfitted estimator from a ModelConfig.
Source code in nullres/models/classifier.py
fit_predict_walk_forward ¶
fit_predict_walk_forward(X: DataFrame, y: Series, t_end: ndarray, split_cfg, model_cfg, use_uniqueness: bool = True, verbose: bool = True) -> tuple[Series, list[dict]]
Out-of-sample P(class 1) for every bar in a test fold.
Bars outside every test window stay NaN — they are training-only and must
never appear in a backtest. Rows with a NaN label are predicted but not
trained on, which is how the deadband in fwd_return works.
Returns (proba, fold_reports).
Source code in nullres/models/classifier.py
feature_importance ¶
feature_importance(X: DataFrame, y: Series, t_end: ndarray, split_cfg, model_cfg, n_repeats: int = 3) -> Series
Permutation importance on the LAST fold's test window only.
In-sample importances tell you what the model memorised. This tells you what actually carried out of sample, which is a much shorter list.
Source code in nullres/models/classifier.py
Strategies¶
base ¶
Context
dataclass
¶
Context(bars: DataFrame, features: DataFrame, label: DataFrame, cfg: object, oos_mask: Series, diagnostics: dict = dict(), verbose: bool = True)
Everything a strategy is allowed to see.
Note what is absent: there is no handle on the future, and oos_mask marks
the bars a strategy is permitted to be judged on. Rule strategies could in
principle trade the whole sample, but they are masked to the same window as
the ML strategies so the comparison is fair — a rule evaluated over six
years against a model evaluated over five is not a comparison.
Strategy ¶
strategy_fingerprint ¶
A stable identity for a strategy instance, including its parameters.
repr will not do: these are plain classes, so the default repr embeds
id(obj) and changes every run. This reads the instance dictionary
instead, recursing into nested strategies — MLMeta holds a primary rule
whose own parameters decide what the model is trained on.
Source code in nullres/strategies/base.py
cached_proba ¶
cached_proba(ctx: Context, key: str, compute, extra: str = '')
Memoise walk-forward predictions across runs that share a context.
nullres sweep varies only sizing, which cannot change the model's output,
so refitting 25 times would be pure waste. The cache key includes the
label, split and model config, so any change that WOULD alter the
predictions misses the cache instead of silently returning stale ones.
extra is for anything else that feeds the feature matrix. It exists
because the fingerprint was incomplete: MLMeta appends primary_side to
the features, which depends on the parameters of its primary rule, and none
of those appeared in the key. Two MLMeta strategies with different
primaries, evaluated against one prepared context, would have taken each
other's predictions — silently, since a cache hit looks exactly like a fast
computation. Nothing in the shipped configs varies the primary, so this was
a trap rather than a live bug, which is the kind that survives longest.
Source code in nullres/strategies/base.py
crossover_state ¶
+1 while fast is above slow, -1 while below. Point-in-time safe.
rules ¶
Rule-based strategies — the benchmarks any model has to clear.
These are deliberately simple and deliberately not tuned. Their job is to set the bar. A tuned rule is not a benchmark, it is another overfit strategy with fewer parameters.
SMACross ¶
Long when the fast average is above the slow one, flat otherwise.
The oldest systematic strategy there is. It trades rarely, so costs barely register, which is exactly why it is hard to beat.
Source code in nullres/strategies/rules.py
DonchianBreakout ¶
VolTargetHold ¶
VolTargetHold(target: float = 0.5, vol_window: int = 30, band: float = 0.1, max_leverage: float = 1.0)
Always long, but sized so that RISK is constant rather than notional.
This strategy makes no directional claim at all. It exists because of a measured asymmetry in the data:
lag-1 autocorrelation of returns -0.029 (noise)
lag-1 autocorrelation of |returns| +0.227 (strong)
lag-1 autocorrelation of 30-bar vol +0.992 (near-deterministic)
Direction is unpredictable; volatility is extremely persistent. So rather than guessing which way the market goes, hold it continuously and vary the size by 1/sigma — cutting exposure when the market is violent and restoring it when it calms.
Note what this can and cannot do. It does not improve expected return; a lower-volatility path with the same drift compounds better, but the edge comes from risk management, not prediction. Judge it on Sharpe and drawdown, and expect total return at or slightly below buy & hold.
max_leverage=1.0 by default, so in calm regimes it is simply long and
never borrows. That makes it deliverable in a spot account with no margin.
Source code in nullres/strategies/rules.py
MeanReversionZ ¶
Fade stretched moves: long when the z-score is deeply negative, and vice versa.
Works in ranging regimes, gets destroyed in trending ones. Included partly as a benchmark and partly because its failure mode is instructive.
Source code in nullres/strategies/rules.py
ml ¶
Machine-learning strategies.
Two formulations, and the difference between them matters more than the model:
MLDirection Predict the direction. The model must answer "which way", which on liquid intraday crypto is close to unanswerable.
MLMeta Meta-labelling. A simple rule decides WHICH WAY to trade; the model only decides WHETHER TO TAKE the trade. This is a far easier question — the model is allowed to say "I don't know" by declining, and declining is free. It also turns an unbalanced 3-class problem into a clean binary one, and the model's output maps naturally onto position size.
If you only take one structural idea from this repo, take the second one.
MLMeta ¶
Meta-labelling on top of a moving-average trend filter.
The primary rule supplies the side. The label becomes "was the rule right?", which is trained only on bars where the rule actually had a position — the model never wastes capacity on bars it will not trade.
Source code in nullres/strategies/ml.py
Orchestration¶
pipeline ¶
End-to-end orchestration: bars -> features -> labels -> positions -> metrics.
One rule governs the ordering here. Features and labels are built on the FULL frame first, and only then are rows dropped and positions renumbered. Building them per-fold would be slower and no safer; building them after dropping rows would silently shorten every rolling window across the gaps.
prepare ¶
prepare(cfg, verbose: bool = True) -> Context
Load data, build features and labels, align them, and mark the OOS window.
Source code in nullres/pipeline.py
ablate ¶
Drop a feature group AFTER row alignment, for a matched-sample ablation.
Turning the data off in the config is not a controlled comparison: the
derivative features carry their own warmup (oi_z needs 168 bars), so
disabling them changes which rows survive the NaN mask, which changes the
fold boundaries and the out-of-sample window. The two runs then differ in
their samples as well as their features, and even buy & hold moves.
This drops the columns from an already-prepared context, so the rows, the splits and the benchmark are byte-identical and the only variable is the feature set.
Source code in nullres/pipeline.py
coherence_warnings ¶
Catch configurations that cannot work, before spending compute on them.
These are not style notes. Each one describes a setup where the backtest will produce a number that means nothing.
Source code in nullres/pipeline.py
trials_so_far ¶
Multiple-testing exposure: everything looked at before reporting this.
Reads the run ledger rather than counting strategies in the current run.
Counting only the current run is the mistake this replaces — it reported
n_trials=6 for a project that had explored well over a hundred parameter
combinations, which made every deflated Sharpe too generous.
extra is how many variants the run about to happen will evaluate. It is
folded into the ledger's own dedupe rather than added on top — see
runlog.count_trials — so verifying an existing result does not inflate
that result's own correction.
Source code in nullres/pipeline.py
trials_caveat ¶
Anything that makes the trial count a FLOOR rather than a measurement.
Source code in nullres/pipeline.py
run_pipeline ¶
run_pipeline(cfg, verbose: bool = True, ctx: Context | None = None, n_trials: int | None = None) -> dict[str, dict]
Run every configured strategy and return {name: metrics}.
A caller may pass a prepared ctx to avoid recomputing features when only
sizing or cost parameters change (see nullres sweep). The context's cfg is
repointed at cfg so those overrides actually take effect — strategies read
their parameters from ctx.cfg, not from the closure.