Case study · Quant research infrastructure

A trading research platform engineered to prove itself wrong.

An autonomous, paper-trading research system with a falsification harness at its core — built not to showcase a profitable backtest, but to answer the harder question honestly: does this strategy actually have an edge?

Most trading projects optimize for a good-looking backtest. Backtests are easy to fake and almost always overfit. This platform inverts the goal: it is a disciplined machine for falsifying its own ideas — and the engineering judgment to act on the answer it gives.

Part two: how I validate it before it touches money →

31
strategy candidates tested
0 deployable — every one falsified
self-caught
a real bug in its own stats gate
found by adversarial self-review
3 axes
of statistical rigor
multiple-testing · costs · autocorrelation
honest
verdict acted on
strategy disarmed → indexed

The premise

A solo retail participant trying to generate alpha is competing against firms with faster data, lower costs, and teams of quants who have already arbitraged the obvious signals. The intellectually honest question isn't "can I make the backtest green?" — it's "would any apparent edge survive multiple-testing, realistic costs, and out-of-sample reality?" The whole platform is built around answering that without fooling itself.

The deliverable isn't a profitable bot. It's a trustworthy falsification harness — and the discipline to index the capital when the harness says there's no edge.

What it is

A production-deployed research service: market-data ingestion, regime classification, multi-signal generation, risk gating, and broker paper-execution — with a scheduler driving the daily cycle and a web dashboard for observability. Live on a managed cloud host behind SSO.

Signal layer

Multi-source fusion

Technical, fundamental, sentiment, and positioning signals fused per ticker, gated by a macro-regime classifier and a conviction threshold before any order is placed.

Execution & risk

Guarded paper trading

Atomic signal-claim, idempotent fills, stop/trailing/volatility exits, position sizing, drawdown and cost gates — paper-only, no real capital at risk.

Observability

Calibration & attribution

Dashboards for prediction calibration (predicted vs realized hit-rate + ECE), per-signal P&L attribution, and equity vs benchmark — surfacing the honest numbers, not the flattering ones.

Research core

The falsification harness

A pure, dependency-light statistics module that gates every candidate on multiple-testing, cost-floor, and autocorrelation before it could ever be trusted. Open-sourced as falsify →

The falsification harness

The centerpiece. Every candidate strategy must clear all three gates — and the test budget is rationed so the discipline can't be gamed.

Extracted and open-sourced as a standalone, dependency-free library → github.com/RAJUSHANIGARAPU/falsify — MIT, pure-stdlib, tests passing across Python 3.9–3.13.

Multi-agent orchestration

Large analyses are run as deterministic workflows that fan work out, verify it adversarially, then synthesize — so conclusions survive independent scrutiny instead of resting on a single pass.

fan-out (parallel specialists)→ adversarial verify (default-to-refute)→ synthesize (ranked, deduped)
Pattern · review

Reliability audit

Parallel auditors across subsystems surfaced findings; each was re-checked by an independent skeptic before any fix landed — separating real correctness bugs from plausible-but-wrong ones.

Pattern · decision

Evidence before build

Before enabling a behavior change, a separate workflow quantified its effect on the real trade history — answering "does this help, or just change activity?" with numbers, not hope.

Pattern · discovery

Screen → gate → test

Candidate edges were screened by category, filtered through a pre-registration gate, and only the single survivor spent a real, pre-registered test slot.

Pattern · self-check

Refute, don't confirm

Surprising "passes" were handed to skeptics told to break them. That step is what caught a real bug in the platform's own significance gate.

What the harness found

The honest scorecard — and why each result is a success of the method, not a failure of it.

No edge

31 candidates tested, 0 deployable. Across forecasting, calendar/flow effects, cross-asset signals, carry, and event-driven strategies — none survived hostile out-of-sample testing versus simply holding the index after costs. A clean kill is the harness doing its job.

Calibration

The live system was confidently wrong. It reported high confidence on trades that realized a near-zero win rate — confidence was inversely related to outcomes. The dashboard was rebuilt to surface net expectancy and a predicted-vs-realized gap, so the metrics can no longer flatter.

Self-caught bug

The platform found a real defect in its own statistics core. An adversarial review caught a unit mismatch that was silently saturating the multiple-testing gate to a perfect score. It was fixed with a misuse-proof API (callers pass raw returns; the core computes statistics itself) and guards that reject degenerate input instead of masking it.

Discipline

The verdict was acted on, not rationalized away. Rather than p-hack a result, the strategy engine was disarmed and the platform reframed as a passive falsification harness — the correct outcome when the evidence says there is no edge.

Why this is the portfolio piece

It demonstrates the competencies that matter for senior AI-engineering and quality-leadership work — the ability to build a system and the judgment to distrust it.

Anyone can show a green backtest. Far fewer can build the machinery to prove their own isn't real — and then trust it.

Stack

Python · FastAPI async SQLAlchemy Alembic migrations scheduler-driven daily cycle broker paper-trading API pandas · numpy pure-stdlib statistics core multi-agent workflows pytest Fly.io · SQLite volume SSO-gated dashboard