Options Buyer ML: Why One Model Fails (and the V2 Fix) A developer detailed the rebuild of an options-buyer prediction system, moving from a single XGBoost model that learned noise from raw premium data to a multi-head architecture that separates underlying mechanics from option contract evaluation. The V2 fix trains narrow heads for specific timeframes and strike multiples, uses shallow trees with heavy regularization, and gates model promotion on out-of-sample performance gaps. Lessons from a real rebuild of an options-buyer prediction system. No profit claims — just the architecture that fixes the chronic bugs of V1. V1 asked one XGBoost model one big fuzzy question: "CE ya PE?" — directly from raw CE/PE premium data. Premium is a transformed signal underlying move × delta × gamma × IV × theta × spread × strike distance × liquidity . The model learned noise as much as signal. Concrete evidence from the research logs: lr=0.02, depth=3 defaults used throughout; Optuna existed but was never run . iv change 1d shift inside single-row groups silently zeroed a whole feature for the entire history. php underlying mechanics -- side, range, ETA, invalidation option chain scanner -- is the buyer contract worth paying for? XGBoost many heads -- thin calibrated learner on clean mechanics Rule: underlying decides side; option contract decides execution eligibility. CE/PE premium is validated against, never learned as, direction. Instead of one CE/PE answer, V2 trains separate narrow heads: underlying up/down touch {15,30,60}m ce 1p3x / ce 1p5x / ce 2p0x and pe 1p3x / pe 1p5x / pe 2p0x SEPARATE CE and PE no trade quality This single change removes most of the CE/PE confusion V1 fought for months. learning rate = 0.015–0.035 n estimators = 800–2000 early stop max depth = 2–3 min child weight = 12–40 gamma = 0.1–2.0 subsample = 0.65–0.90 colsample bytree = 0.55–0.85 reg alpha = 0.5–3.0 reg lambda = 6.0–20.0 scale pos weight = min neg/pos, 8.0 V1's intraday head had only 8 of 1280 features with non-zero gain — most of the bloat was pure noise the regularizer had to prune. Shallow + hard-regularized is the answer. overfit gap = train metric − test metric . Flag if 0.15. A model is NOT promoted just because train metrics look good. Log the gap automatically on every head, every retrain. V2 is a cleaner architecture, but it is still research . The lesson that transfers: stop asking fuzzy questions, declare your nulls, keep trees shallow, and gate promotion on out-of-sample gap — not training score. Research only. Not investment advice.