Results & negatives
The default path verifies committed offline scores, then recomputes approve/review/decline bands, swap sets, queue overflow, expected loss, calibration, vintage drift, descriptive slices, and a policy audit record.
- 01
I expected the 240-tree XGBoost challenger to pull away from the calibrated logistic baseline on the 24,000 later-backtest rows. It didn't: Brier 0.159253 against 0.159280, log loss 0.494320 against 0.493403, ROC AUC 0.675262 against 0.674615. Better on two of the three, worse on one, every margin far too small to act on.
- 02
The workbench opens on a policy that cannot be published, and that is deliberate. At approve ≤12% / review ≤28% with a review capacity of 180, the fixture's final vintage sends 345 applications to manual review — 165 over capacity — so the capacity gate blocks publication. I did not slide the threshold until the queue fit.