NFL Sports Edge V2 — Execution Plan

Author: Jose "Joey" Moronta  Drafted: 2026-10-03  Status: Build starting. This is the bridge between the build spec and the actual repo, Incognito196/nfl-sports-edge.

Builder: Claude (this account, on your Max plan, no separate cost). Operator once stable: Hermes, on joey-dev, owning the weekly scheduled run and the alert if it breaks. Repo and Cloudflare Pages deploy stay exactly as they are — git push to main is still the deploy.

What's actually in the repo right now

Read every file before writing any of this. Verdict: better scaffolding than expected, and it confirms every structural worry in the build spec.

PieceStatusFinding
Site, PWA, salary ingestion, manifest, deployKEEPSolid. Installable, atomic commits, data hashing, freshness gate. No reason to touch it.
Starter gateKEEPHas real unit tests. A spot start can't override the market's QB hierarchy. Good instinct, keep it.
Lineup solverKEEP (tune)Uses scipy.milp to solve the exact legal DK lineup — not a greedy heuristic. Right tool. The problem is what feeds it, and how many times it runs.
Core projectionREPLACEMostly a pass-through: takes a public podcast site's projection when available, else 0.55×DK season avg + 0.25×recent + 0.20×xFP. This is the "replace manually chosen weights" problem, literally.
Expected Fantasy Points (xFP)REPLACEA Ridge regression trained AND predicted on the same 1-3 weeks of this season. That's leakage by definition — it's not forecasting, it's describing. No historical seasons used at all.
Simulation volumeREPLACE300 simulations, not 25,000+. Each one re-solves a full integer program with a 10-second cap — that's why it's capped at 300. Confirms the optimizer-cost warning from the start of this conversation.
Ownership / "field attention"REPLACEA hand-typed list of player names in the source file, honestly labeled as not real ownership. Rots the moment nobody edits it weekly.
Feature weights (role accel, mispricing, regression)REPLACEAll hand-assigned (0.45, 0.25 0.30, 0.20 0.30 etc). Exactly the "should this weight exist at all" question the spec raises.

The baseline gate — nothing ships without this

Before any V2 component replaces its V1 counterpart on the live site, it has to beat the thing it's replacing on data neither one was trained on. Specifically:

If V2 loses to the dumb pass-through on holdout, the dumb pass-through stays live and V2 goes back to the drawing board. No exceptions for "it should work in theory."

Build order

PhaseDeliverableWhy this order
0DuckDB/Parquet data layer on the T130's bulk drive, streaming — never load full seasons into pandas in memory8GB of RAM on the T130 is the hard constraint discussed earlier; this has to be right before anything else is built on top of it
1Historical player-game dataset, 2022–present, nflverse, every row stamped with what was knowable pre-kickoffEverything downstream is replaceable code sitting on this asset
2Role-change / usage-trend features with a leakage unit test (recompute pre-kickoff-only, assert equality)Cheapest phase to get wrong silently — rolling windows leak easiest
3Real out-of-sample xFP model, trained on prior seasons, evaluated on held-out weeks it never sawFixes the leakage found in the current repo
4Feature-survival pass — test each candidate feature's actual out-of-sample contribution per position, per percentile; drop what doesn't surviveThis is the "don't assign weights, let the data decide" instruction — the core philosophical change
5Distributional output (percentiles, P(20+)/P(30+)/P(multiplier) per player) replacing the single-number projectionTournament scoring needs the tail, not the mean
6Correlated game-environment simulation, replacing independent player samplingQB↔WR correlation is where most of the real edge lives per the spec
7Fast approximate lineup construction for bulk sims, calibrated against the exact MILP solver on a few thousand simsSolves the 300-sim ceiling found in the audit without pretending 100k exact solves is realistic on this hardware
8Ownership model — proxy from salary/projection/value/game-total now, replaced by real logged ownership from contests you actually enter going forwardNo free historical DK ownership exists; this was flagged as a hard data gap in the original spec, confirmed again in the audit
9Leverage, portfolio construction across your 3 entries, explainability panel per player (our number vs. industry vs. why)Final layer, only meaningful once 1–8 are trustworthy
10Audit trail + handoff runbook to Hermes for the weekly scheduled runBuilder (me) finishes, operator (Hermes) takes over reliability

Overfitting discipline

2022–present is roughly four seasons. Split by position and by percentile, that's a thin sample for a flexible model with dozens of candidate features. The SharpShooter NHL model found a "recalibration edge" that looked real on a small backtest and was pure noise. To not repeat that here: heavy regularization, a small surviving feature set per Phase 4, and nested validation — not one holdout season treated as final proof.

What happens next

Phase 0 and Phase 1 start now. Nothing touches the live site until a phase clears the baseline gate above. You'll get the full plain-English rundown by email once this is far enough along to be worth reading in one sitting, not before.