Raju Shanigarapu · October 2026. Part one showed that no strategy survived out-of-sample testing. This part is about the engineering that keeps the system honest while it runs: shadow trading, deterministic replay, a CI fidelity gate and per-call cost accounting.
On 30 September 2026 my trading system ran its first shadow cycle on a schedule. The job finished. Nothing raised an exception. The database commit succeeded.
The book was empty.
The run was supposed to seed a paper portfolio with $100,000 of cash and rebalance it into the target weights. What it actually wrote was a set of zero-quantity position rows, no trades, and a daily snapshot with a NAV of 0. Next to it, the "expected NAV" I compare against read 512,273, because that calculation had been compounding the whole backtest instead of starting at the seed.
A crash would have been easy. This is the failure I care about most as a tester: the system says "done" and the result is wrong. This post covers how I test for that, and what the testing does not prove yet.
autonomous-investor is a research and validation environment for equity strategies. It scores candidates through a signal pipeline and tests every idea out-of-sample against the passive alternative it would have to beat.
Some of that pipeline uses LLMs: a sentiment agent, a fundamentals thesis, signal fusion, a weekly report, and a seven-node bull/bear debate graph. LLM calls go through model chains that try free-tier models on Google AI Studio and OpenRouter first, then fall through to the next model on an error or an unusable reply. Paid Anthropic models are an opt-in fallback with a budget cap.
Two facts frame everything that follows:
INDEXED: no new signals, no new entries.What is left is the platform, including a shadow deployment that runs a frozen strategy forward every trading day. The shadow book trades two frozen rule-based sleeves (a CEF discount signal and a trend signal). It uses no LLM, deliberately: the safety paths (shadow, order state machine, ledger, halt manager, risk guard) import no LLM module.
flowchart LR
D[Data ingestion<br/>prices, news, macro, filings] --> A[Analysis agents<br/>technical, fundamental,<br/>sentiment, smart money, macro]
A --> F[Signal fusion<br/>+ debate graph]
L[LLM chains<br/>free-first fallback<br/>paid tier budget-gated] -.-> A
L -.-> F
F --> R[Risk guard<br/>hard limits, circuit breakers]
R --> E[Execution<br/>Alpaca paper]
S[Shadow cycle 15:30 ET<br/>frozen CEF + trend rules<br/>no LLM imports] --> B[(Shadow book<br/>trades, snapshots,<br/>recorded inputs)]
B --> P[Deterministic replay]
C[Cost guard<br/>usage log per call] -.-> L
The suite has about 1,700 tests. The last merged change reported 1,686 passing with 57.89% line coverage. The floor in pytest.ini is a ratchet: it is set to the measured total rounded down, and it only moves up. It went from 55 to 56 to 57 during September.
58% is not high, and I don't treat it as a quality signal. The ratchet stops backsliding; it says nothing about whether tests assert the right things. One bug proved that: the cost guard had an invalid Haiku model id, and the test for it never failed because the test kept its own copy of the constant instead of importing it.
The shadow runner has to do exactly what the frozen research harness did. If it doesn't, the shadow record measures my reimplementation rather than the strategy. A GitHub Actions workflow runs two jobs on every change to the app, tests, migrations or dependencies:
ruff check (since #39 a violation fails the build) and the full non-fidelity suite with the coverage floor enforced.The honest weakness is that the fidelity tests need live Yahoo data and skip themselves when it is unreachable. I chose that so CI does not fail on network flakiness. The cost is that a green run can mean "skipped". The workflow prints skip reasons (-rs), but a skip still does not block a merge.
The shadow cycle runs Monday to Friday at 15:30 ET, 30 minutes before the close, because post-close IEX quotes are unusable. Each run records its inputs in a shadow_cycle_inputs row: outcome, rebalance flag, seed capital, expected NAV, target weights, the book before the run, every quote it was served, and a SHA-256 hash plus window of the price history.
The full history is several MB a day and yfinance is not bit-stable across fetches, so I store only the hash. Replay doesn't need the history, because the two values derived from it (target weights, expected NAV) are stored as they were.
replay_shadow_day rebuilds a day's runs in order, in a scratch in-memory book, through the real cycle code. It then diffs the result against the recorded trades and snapshot, with a relative tolerance of 1e-9. It only reads the real database. The ops wrapper exits 0 on a match, 1 on a diff and 2 when nothing was recorded, so it can be scripted. The replay for 1 October returned match: one run, zero diffs.
Replay deliberately does not re-evaluate the rebalance rule. The book had been rebalancing on the first run of each month while its backtest rebalances on the last trading day, so the forward record drifted from its own benchmark. The fixed rule (#48) is the last full-session NYSE trading day, with one catch-up if missed and never two per month. Replay uses the is_rebalance_day flag recorded with each run, so days recorded under the old rule still replay unchanged; a test patches the rule to raise and proves it.
The monthly budget tier (normal, Haiku-only, paused) used to be display-only. The one method that acted on it had no callers. I made every paid Anthropic call site check the tier before it runs (#46). Then I found that the debate graph, the sentiment fallback, the bootstrap key check and the free chains wrote no usage rows at all, so the projection that sets the tier only saw part of the spend. #51 routes all of them through track_usage. Free chain responses are recorded at $0 with real token counts, so volume is visible without moving the tier; a chain overridden to a paid model is charged what OpenRouter reports.
The 30 September failure was in app/database.py. The async engine for the file database used StaticPool, so every session (scheduler jobs, web requests) shared one SQLite connection. SQLite transactions belong to the connection, not the session. When any other session closed, the pool rolled back the shared connection and undid the seed's flushed but uncommitted INSERT. The seed's commit then succeeded with nothing in it. The rebalance saw NAV 0. The partial set of six position rows for eight targets plus cash fits a second rollback landing during the row-by-row inserts.
Replay could not have caught this. It runs one session on a private in-memory database, so it cannot see cross-session effects. Worse, live captured positions_before after the seed, and replay read that empty book as "unseeded, seed it". So I wrote a test that goes through the real scheduler job wrapper, the real settings and the application's own engine, against a temporary file database. It opens and closes another session between the seed's insert and its commit. It fails with StaticPool and passes with the fix.
The fix (#49): StaticPool only for in-memory SQLite; the seed is read back after commit and the run fails loudly if it is missing; expected NAV starts at the seed; input capture records the book before the seed. The 30 September rows are kept, and their replay still reports a diff, which is the correct answer.
The trade-off: sessions now contend for the SQLite write lock (5-second busy timeout). A job that holds an uncommitted write across a slow await can now fail with "database is locked", where before it silently corrupted other sessions. I prefer the failure I can see.
While writing the rebalance rule (#48), I noticed that is_nyse_holiday treated 31 December as a holiday whenever 1 January fell on a Saturday. That is the usual weekend-observance rule, but NYSE Rule 7.2 makes an exception: when New Year's Day is a Saturday there is no Friday observance and the market is open. The old code marked 2021-12-31 and 2027-12-31 as closed. The scheduler would have skipped a real trading day, and last_trading_day(2027, 12) returned 30 December, moving the December 2027 rebalance a day early.
I kept it out of the rebalance PR and fixed it separately (#50). Every closure now lives in its own calendar year, and a test checks that no closure leaks into another year from 2020 to 2040. The bug would have stayed silent until December 2027 and then shown up as a one-day tracking difference that looks like noise.
StaticPool is correct. The bug only existed in the production configuration. A live-path test should have been the first test, not the post-mortem.The shadow stage has pre-committed exit criteria: at least 60 trading days, annualized tracking error under 400 bps, blended execution cost under 75 bps, zero unhandled failures and every rebalance fired. As of this draft the book has one clean recorded day. None of those criteria is met yet. A small real-money stage depends on meeting them.
Proven so far: the shadow runner reproduces the frozen research code at machine precision, when data is available; recorded days replay deterministically; the two calendar and concurrency bugs above are fixed with regression tests; and paid LLM spend is metered and capped at every call site I found.
Not proven: any edge or return (the research result is negative); behaviour with real money, real fills or real slippage; how the system holds up over a long horizon; and whether SQLite under the new locking behaviour holds up under heavier concurrent load.