COST-ACCURACY CASCADE · OFFLINE REPLAY
The expensive model
was the wrong default.
The expensive model was the wrong default: on held-out data Claude Sonnet 5 ties Haiku 4.5 at 2.8× the cost, and under a shared output budget it silently answers nothing on up to 2.5% of calls. A confidence cascade routes each complaint to the cheapest tier that can handle it, measured against 11 years of drift. The in-page demo replays those measured runs.
≈ 474 of every 1,000 complaints/month reach a human reviewer at this setting.
MISROUTE COST · CONFIDENCE THRESHOLD · PRECOMPUTED GRID
Move the sliders.
The strategy rewrites itself.
CONTROL
CURRENT READING
threshold0.747
Tier A — TF-IDF LogReg, 3 real complaints, evaluated against the current threshold:
- STOPS AT TF-IDFconf 0.777correct
They are attempting to collect a debt that doesn't belong to me. They are demanding immediate payment from me without providing any proof they have the right pe…
- STOPS AT TF-IDFconf 0.765correct
Fraud accounts on XXXX report
- ESCALATESconf 0.540predicted card, truth credit_reporting
On XX/XX/2022, I mailed my dispute letter to Citi cards CBNA ( see Exhibit A ) and received their Response dated XX/XX/2022 ( see Exhibit B ). Their response no…
escalate47.4%
Tier A — TF-IDF LogReg, 3 real complaints, evaluated against the current threshold:
- STOPS AT TF-IDFconf 0.777correct
They are attempting to collect a debt that doesn't belong to me. They are demanding immediate payment from me without providing any proof they have the right pe…
- STOPS AT TF-IDFconf 0.765correct
Fraud accounts on XXXX report
- ESCALATESconf 0.540predicted card, truth credit_reporting
On XX/XX/2022, I mailed my dispute letter to Citi cards CBNA ( see Exhibit A ) and received their Response dated XX/XX/2022 ( see Exhibit B ). Their response no…
macro-F10.919 [0.786, 0.976]
Tier A — TF-IDF LogReg, 3 real complaints, evaluated against the current threshold:
- STOPS AT TF-IDFconf 0.777correct
They are attempting to collect a debt that doesn't belong to me. They are demanding immediate payment from me without providing any proof they have the right pe…
- STOPS AT TF-IDFconf 0.765correct
Fraud accounts on XXXX report
- ESCALATESconf 0.540predicted card, truth credit_reporting
On XX/XX/2022, I mailed my dispute letter to Citi cards CBNA ( see Exhibit A ) and received their Response dated XX/XX/2022 ( see Exhibit B ). Their response no…
monthly cost (USD / 1k)USD 1,337.53
Tier A — TF-IDF LogReg, 3 real complaints, evaluated against the current threshold:
- STOPS AT TF-IDFconf 0.777correct
They are attempting to collect a debt that doesn't belong to me. They are demanding immediate payment from me without providing any proof they have the right pe…
- STOPS AT TF-IDFconf 0.765correct
Fraud accounts on XXXX report
- ESCALATESconf 0.540predicted card, truth credit_reporting
On XX/XX/2022, I mailed my dispute letter to Citi cards CBNA ( see Exhibit A ) and received their Response dated XX/XX/2022 ( see Exhibit B ). Their response no…
INTERPRETATION
Balance automated routing with the review queue at threshold 0.747: 47.4% escalate and sample macro-F1 is 0.919.
≈ 474 of every 1,000 complaints/month reach a human reviewer at this setting.
triage:balanced-queue
DETAIL
cascade threshold, cf. RouteLLM (Ong et al., 2024)
Disclosure: the escalation/cost grid is built from real offline replay of 86,972 calibration-slice complaints, not simulation (dated 2026-08-12). Misroute-cost sensitivity and normalized monthly-cost projections use the recorded USD basis; no currency conversion is applied. Monthly cost is normalized per 1,000 complaints and is not a production invoice. Per-threshold macro-F1 and its confidence interval are computed on a 200-sample curated paired set, not the full 104,443-row TEST-IID sweep. Macro-F1 near 100% escalation reflects human-credit accounting, not router quality.
model.int8.onnx67,575,183 Btokenizer.json711,494 Bort-wasm-simd-threaded.wasm13,479,978 B
RECORDED MISROUTES · A_TO_B
Eight real complaints,
the cascade sent the wrong way.
| Complaint | Route | Predicted | Truth | Confidence |
|---|---|---|---|---|
| 5773141 | A → escalated → B2 → answered | deposit_account | payday_personal_loan | 0.376 |
| 5819417 | A → escalated → B2 → answered | vehicle_loan | credit_reporting | 0.626 |
| 5819724 | A → answered | debt_collection | deposit_account | 0.667 |
| 5820251 | A → answered | card | credit_reporting | 0.652 |
| 5838416 | A → escalated → B2 → answered | card | deposit_account | 0.647 |
| 5872622 | A → escalated → B2 → answered | vehicle_loan | debt_collection | 0.633 |
| 5903957 | A → answered | deposit_account | money_service | 0.794 |
| 5905003 | A → escalated → B2 → answered | credit_reporting | vehicle_loan | 0.701 |
- tier_b2 routed this complaint to deposit_account at confidence 0.376; the recorded label is payday_personal_loan.
- tier_b2 routed this complaint to vehicle_loan at confidence 0.626; the recorded label is credit_reporting.
- tier_a_logreg routed this complaint to debt_collection at confidence 0.667; the recorded label is deposit_account.
- tier_a_logreg routed this complaint to card at confidence 0.652; the recorded label is credit_reporting.
- tier_b2 routed this complaint to card at confidence 0.647; the recorded label is deposit_account.
- tier_b2 routed this complaint to vehicle_loan at confidence 0.633; the recorded label is debt_collection.
- tier_a_logreg routed this complaint to deposit_account at confidence 0.794; the recorded label is money_service.
- tier_b2 routed this complaint to credit_reporting at confidence 0.701; the recorded label is vehicle_loan.
8 RECORDED TIERS · COST vs MACRO-F1
Accuracy and cost
share the same axis.
| Tier | macro-F1 | Cost / 1k | Run |
|---|---|---|---|
| Tier A — TF-IDF LogReg | 0.761 [0.756, 0.764] | $933.41 | 8e4d6345b8… |
| Tier A — TF-IDF ComplementNB | 0.727 [0.722, 0.731] | $1032.39 | c20cd14a7a… |
| Tier B1 — ModernBERT-base (seed a) | 0.788 [0.784, 0.792] | $862.60 | 8071d31db2… |
| Tier B1 — ModernBERT-base (seed b) | 0.788 [0.784, 0.791] | $865.58 | adb96307e2… |
| Tier B1 — ModernBERT-base (seed c) | 0.786 [0.782, 0.790] | $867.94 | a523049a7e… |
| Tier B2 — DistilBERT (deployment point) | 0.795 [0.791, 0.799] | $834.32 | 5517ebf1df… |
| Tier C — Claude Haiku 4.5 (TEST-IID) | 0.770 [0.750, 0.789] | $916.92 | 70a1b0c40e… |
| Tier C — Claude Sonnet 5 (TEST-IID subsample) | 0.742 [0.701, 0.773] | $955.66 | e1503146b7… |
DRIFT 2022-H2 – 2026-H1 · MACRO_F1
The model
aged out.
Every trained tier loses ground as the complaint mix moves past its training window; the two Claude tiers, scored zero-shot with no fitted vocabulary, do not. Tier B1's yearly series is not in the source yet and is shown as pending, not zero.
| Tier | 2022-H2 | 2023 | 2024 | 2025 | 2026-H1 |
|---|---|---|---|---|---|
| Tier A — TF-IDF LogReg | 0.761 [0.756, 0.764] | 0.758 [0.750, 0.766] | 0.748 [0.738, 0.757] | 0.729 [0.721, 0.739] | 0.666 [0.658, 0.673] |
| Tier B2 — DistilBERT | 0.795 [0.791, 0.799] | 0.790 [0.782, 0.798] | 0.774 [0.765, 0.782] | 0.760 [0.751, 0.769] | 0.726 [0.719, 0.733] |
| Tier C — Claude Haiku 4.5 | 0.770 [0.750, 0.789] | 0.736 [0.699, 0.769] | 0.762 [0.727, 0.792] | 0.764 [0.730, 0.790] | 0.728 [0.705, 0.751] |
| Tier C — Claude Sonnet 5 | 0.742 [0.701, 0.773] | 0.754 [0.719, 0.786] | 0.780 [0.746, 0.807] | 0.763 [0.727, 0.791] | 0.782 [0.759, 0.805] |
| Tier B1 — ModernBERT-base | PENDING — NOT MEASURED | PENDING — NOT MEASURED | PENDING — NOT MEASURED | PENDING — NOT MEASURED | PENDING — NOT MEASURED |
Source note: Tier B1 yearly drift is pending in the source; no values were measured.
SOURCE · RECEIPTS
Every number opens
the same file.
- What was verified
- Recorded routing policies expose the cost/error tradeoff, known failures and the curated Python-to-browser parity check.
- Evidence class
- Compact exports of recorded triage runs, plus a separately served quantized browser model.
- Boundary
- Pending Tier B1 drift rows remain pending. Curated parity is not general accuracy, and the model file is not committed to Git.
Files, hashes and methods
public/case-studies/triage-router/frontier.compact.jsonLucisZhang/portfolio-site ·acf05ae78859sha256:c7e6555d3b03d988da30403f6052abf5e0be778e17f26b6eaa9c6696fc5eefccpublic/case-studies/triage-router/policies.compact.jsonLucisZhang/portfolio-site ·acf05ae78859sha256:9c2992c53dd62c7412a396e5c175e68068235a37548f74898b5858f258dc5c12public/case-studies/triage-router/strategy-cards.jsonLucisZhang/portfolio-site ·acf05ae78859sha256:a26dd35e2c911aed93894bf7044be9c3a167ff9ea874cfbd4b432570ca6a868apublic/case-studies/triage-router/drift.compact.jsonLucisZhang/portfolio-site ·acf05ae78859sha256:ea70d32187602f644e895a151e3320991dd1dfc187441bf101c2e49587dc307fpublic/case-studies/triage-router/known-failures.jsonLucisZhang/portfolio-site ·acf05ae78859sha256:fdcadaafa15e885890984cbab4b5320b46eadf6a3f374f2dcb7d76bf70e35218public/case-studies/triage-router/samples.curated.jsonLucisZhang/portfolio-site ·acf05ae78859sha256:3c6771b65f44c567082857acd4db08a30004be2c4dc562974bef7a80cb386e4bpublic/case-studies/triage-router/python_int8_curated.jsonLucisZhang/triage-router ·b2734bbcbd75sha256:484b21fab582af7a8b00c98abd89ecb575a5b1330dc7af791bde333b59953486public/models/triage-tier-b2/model.int8.onnxServed model; no committed GitHub file.sha256:da931ec8310cf1280747e22fc6ebfd30fd5f92e312ede6544042e1190764bb4a
The compact payloads above are built by Triage Router's export_site_payloads.py from real recorded triage runs; the model file is copied on-site and never enters git. docs/evidence/r2-source-map.md pins every path and hash; npm run verify:r2-sources re-checks them on every run.
python scripts/export_site_payloads.py --out public/case-studies/triage-router
Architecture
I built the full ladder: calibrated TF-IDF baselines, ModernBERT and DistilBERT fine-tunes, Claude tiers over OpenRouter, the explicit cost model, the cascade router, and the append-only results log every number traces back to.
- Freeze
Snapshot the CFPB corpus by hash and split it by time: train, calibration, IID test, drift slices.
- Ladder
Train each tier on identical splits: TF-IDF, ModernBERT ×3 seeds, DistilBERT int8, Claude zero-shot.
- Cost
Price every path with an explicit model: inference dollars, misroute cost, human review.
- Route
Fit cascade thresholds on the calibration slice only — never on test.
- Drift
Replay 2023–2026 yearly slices and decompose what actually moved: class mix, not vocabulary.
Results & negatives
On held-out data the cascade beats the linear baseline by +0.0370 macro-F1 (95% CI +0.0337 to +0.0402) while cutting expected cost by $120.58 per thousand complaints. On the 2026-H1 drift slice the linear tier drops 9.2 points; Claude Sonnet gains ground instead.
I assumed the expensive tier was the safe default. On the same 1,500 held-out IID rows, Sonnet 5 minus Haiku 4.5 comes out at −0.0073 macro-F1, CI [−0.0427, +0.0263], McNemar p=1.00 — indistinguishable, at 2.8× the cost per thousand complaints ($3.66 against $1.32) and 2.3× the p50 latency. It earns its price only on the post-cutoff slice, where the paired gain is +0.0458 macro-F1.
Under the shared 64-token completion budget the stronger model sometimes answers with nothing at all: 12 of 1,500 calls on the IID slice (0.8%) and 37 of 1,500 on the post-cutoff slice (2.5%), every one with finish_reason "length" and no parseable JSON — its reasoning consumed the budget before an answer appeared. Haiku did that on 0 of 10,000 calls. The failures are worst on the drifted slice, which is exactly where you would escalate to the stronger model.
I thought Tier C was pinned to Anthropic. Reading the request builder showed nothing of the kind: the body never carried an OpenRouter provider preference — no order, no only, no allow_fallbacks — so the pin never was in effect. Across all 16,050 committed Tier C calls, 16,020 (99.81%) were served by Amazon Bedrock, and every Phase 3 latency figure had to be relabeled as a route through OpenRouter to Bedrock.
Limitations
A pre-registered hypothesis — that the fine-tuned transformer would behave like the LLM tiers under drift — was refuted: its 2026-H1 drop splits roughly evenly between class-mix shift and within-class decay.
The obvious explanation did not survive either: out-of-vocabulary rate barely moved, so lexical drift is ruled out; the cliff is class mix.
Results apply to this task, this corpus, and the price sheets at measurement time; the router's dollar savings follow from the disclosed cost model, not a universal claim.