An autonomous, paper-trading research system with a falsification harness at its core — built not to showcase a profitable backtest, but to answer the harder question honestly: does this strategy actually have an edge?
Part two: how I validate it before it touches money →
A solo retail participant trying to generate alpha is competing against firms with faster data, lower costs, and teams of quants who have already arbitraged the obvious signals. The intellectually honest question isn't "can I make the backtest green?" — it's "would any apparent edge survive multiple-testing, realistic costs, and out-of-sample reality?" The whole platform is built around answering that without fooling itself.
The deliverable isn't a profitable bot. It's a trustworthy falsification harness — and the discipline to index the capital when the harness says there's no edge.
A production-deployed research service: market-data ingestion, regime classification, multi-signal generation, risk gating, and broker paper-execution — with a scheduler driving the daily cycle and a web dashboard for observability. Live on a managed cloud host behind SSO.
Technical, fundamental, sentiment, and positioning signals fused per ticker, gated by a macro-regime classifier and a conviction threshold before any order is placed.
Atomic signal-claim, idempotent fills, stop/trailing/volatility exits, position sizing, drawdown and cost gates — paper-only, no real capital at risk.
Dashboards for prediction calibration (predicted vs realized hit-rate + ECE), per-signal P&L attribution, and equity vs benchmark — surfacing the honest numbers, not the flattering ones.
A pure, dependency-light statistics module that gates every candidate on multiple-testing, cost-floor, and autocorrelation before it could ever be trusted. Open-sourced as falsify →
The centerpiece. Every candidate strategy must clear all three gates — and the test budget is rationed so the discipline can't be gamed.
Large analyses are run as deterministic workflows that fan work out, verify it adversarially, then synthesize — so conclusions survive independent scrutiny instead of resting on a single pass.
Parallel auditors across subsystems surfaced findings; each was re-checked by an independent skeptic before any fix landed — separating real correctness bugs from plausible-but-wrong ones.
Before enabling a behavior change, a separate workflow quantified its effect on the real trade history — answering "does this help, or just change activity?" with numbers, not hope.
Candidate edges were screened by category, filtered through a pre-registration gate, and only the single survivor spent a real, pre-registered test slot.
Surprising "passes" were handed to skeptics told to break them. That step is what caught a real bug in the platform's own significance gate.
The honest scorecard — and why each result is a success of the method, not a failure of it.
31 candidates tested, 0 deployable. Across forecasting, calendar/flow effects, cross-asset signals, carry, and event-driven strategies — none survived hostile out-of-sample testing versus simply holding the index after costs. A clean kill is the harness doing its job.
The live system was confidently wrong. It reported high confidence on trades that realized a near-zero win rate — confidence was inversely related to outcomes. The dashboard was rebuilt to surface net expectancy and a predicted-vs-realized gap, so the metrics can no longer flatter.
The platform found a real defect in its own statistics core. An adversarial review caught a unit mismatch that was silently saturating the multiple-testing gate to a perfect score. It was fixed with a misuse-proof API (callers pass raw returns; the core computes statistics itself) and guards that reject degenerate input instead of masking it.
The verdict was acted on, not rationalized away. Rather than p-hack a result, the strategy engine was disarmed and the platform reframed as a passive falsification harness — the correct outcome when the evidence says there is no edge.
It demonstrates the competencies that matter for senior AI-engineering and quality-leadership work — the ability to build a system and the judgment to distrust it.
Anyone can show a green backtest. Far fewer can build the machinery to prove their own isn't real — and then trust it.