COST-ACCURACY CASCADE · OFFLINE REPLAY

The expensive model
was the wrong default.

The expensive model was the wrong default: on held-out data Claude Sonnet 5 ties Haiku 4.5 at 2.8× the cost, and under a shared output budget it silently answers nothing on up to 2.5% of calls. A confidence cascade routes each complaint to the cheapest tier that can handle it, measured against 11 years of drift. The in-page demo replays those measured runs.

+0.037MACRO-F1
−$120.581K CALLS
2015–2026DRIFT
Tier frontier · 64 thresholds0.747 · macro-F1 0.919

≈ 474 of every 1,000 complaints/month reach a human reviewer at this setting.

MISROUTE COST · CONFIDENCE THRESHOLD · PRECOMPUTED GRID

Move the sliders.
The strategy rewrites itself.

CONTROL

Recorded sensitivity range: USD 0.75 — USD 24.00
Recorded threshold range: 0.1890.979

CURRENT READING

threshold0.747

Tier A — TF-IDF LogReg, 3 real complaints, evaluated against the current threshold:

  • STOPS AT TF-IDFconf 0.777correct

    They are attempting to collect a debt that doesn't belong to me. They are demanding immediate payment from me without providing any proof they have the right pe…

  • STOPS AT TF-IDFconf 0.765correct

    Fraud accounts on XXXX report

  • ESCALATESconf 0.540predicted card, truth credit_reporting

    On XX/XX/2022, I mailed my dispute letter to Citi cards CBNA ( see Exhibit A ) and received their Response dated XX/XX/2022 ( see Exhibit B ). Their response no…

escalate47.4%

Tier A — TF-IDF LogReg, 3 real complaints, evaluated against the current threshold:

  • STOPS AT TF-IDFconf 0.777correct

    They are attempting to collect a debt that doesn't belong to me. They are demanding immediate payment from me without providing any proof they have the right pe…

  • STOPS AT TF-IDFconf 0.765correct

    Fraud accounts on XXXX report

  • ESCALATESconf 0.540predicted card, truth credit_reporting

    On XX/XX/2022, I mailed my dispute letter to Citi cards CBNA ( see Exhibit A ) and received their Response dated XX/XX/2022 ( see Exhibit B ). Their response no…

macro-F10.919 [0.786, 0.976]

Tier A — TF-IDF LogReg, 3 real complaints, evaluated against the current threshold:

  • STOPS AT TF-IDFconf 0.777correct

    They are attempting to collect a debt that doesn't belong to me. They are demanding immediate payment from me without providing any proof they have the right pe…

  • STOPS AT TF-IDFconf 0.765correct

    Fraud accounts on XXXX report

  • ESCALATESconf 0.540predicted card, truth credit_reporting

    On XX/XX/2022, I mailed my dispute letter to Citi cards CBNA ( see Exhibit A ) and received their Response dated XX/XX/2022 ( see Exhibit B ). Their response no…

monthly cost (USD / 1k)USD 1,337.53

Tier A — TF-IDF LogReg, 3 real complaints, evaluated against the current threshold:

  • STOPS AT TF-IDFconf 0.777correct

    They are attempting to collect a debt that doesn't belong to me. They are demanding immediate payment from me without providing any proof they have the right pe…

  • STOPS AT TF-IDFconf 0.765correct

    Fraud accounts on XXXX report

  • ESCALATESconf 0.540predicted card, truth credit_reporting

    On XX/XX/2022, I mailed my dispute letter to Citi cards CBNA ( see Exhibit A ) and received their Response dated XX/XX/2022 ( see Exhibit B ). Their response no…

INTERPRETATION

Balance automated routing with the review queue at threshold 0.747: 47.4% escalate and sample macro-F1 is 0.919.

≈ 474 of every 1,000 complaints/month reach a human reviewer at this setting.

triage:balanced-queue

DETAIL

cascade threshold, cf. RouteLLM (Ong et al., 2024)

Disclosure: the escalation/cost grid is built from real offline replay of 86,972 calibration-slice complaints, not simulation (dated 2026-08-12). Misroute-cost sensitivity and normalized monthly-cost projections use the recorded USD basis; no currency conversion is applied. Monthly cost is normalized per 1,000 complaints and is not a production invoice. Per-threshold macro-F1 and its confidence interval are computed on a 200-sample curated paired set, not the full 104,443-row TEST-IID sweep. Macro-F1 near 100% escalation reflects human-credit accounting, not router quality.

  • model.int8.onnx67,575,183 B
  • tokenizer.json711,494 B
  • ort-wasm-simd-threaded.wasm13,479,978 B

RECORDED MISROUTES · A_TO_B

Eight real complaints,
the cascade sent the wrong way.

ComplaintRoutePredictedTruthConfidence
5773141A → escalated → B2 → answereddeposit_accountpayday_personal_loan0.376
5819417A → escalated → B2 → answeredvehicle_loancredit_reporting0.626
5819724A → answereddebt_collectiondeposit_account0.667
5820251A → answeredcardcredit_reporting0.652
5838416A → escalated → B2 → answeredcarddeposit_account0.647
5872622A → escalated → B2 → answeredvehicle_loandebt_collection0.633
5903957A → answereddeposit_accountmoney_service0.794
5905003A → escalated → B2 → answeredcredit_reportingvehicle_loan0.701
PredictedTruth

5773141 conf 0.376 · A → escalated → B2 → answered

deposit_accountpayday_personal_loan

5819417 conf 0.626 · A → escalated → B2 → answered

vehicle_loancredit_reporting

5819724 conf 0.667 · A → answered

debt_collectiondeposit_account

5820251 conf 0.652 · A → answered

cardcredit_reporting

5838416 conf 0.647 · A → escalated → B2 → answered

carddeposit_account

5872622 conf 0.633 · A → escalated → B2 → answered

vehicle_loandebt_collection

5903957 conf 0.794 · A → answered

deposit_accountmoney_service

5905003 conf 0.701 · A → escalated → B2 → answered

credit_reportingvehicle_loan
  • tier_b2 routed this complaint to deposit_account at confidence 0.376; the recorded label is payday_personal_loan.
  • tier_b2 routed this complaint to vehicle_loan at confidence 0.626; the recorded label is credit_reporting.
  • tier_a_logreg routed this complaint to debt_collection at confidence 0.667; the recorded label is deposit_account.
  • tier_a_logreg routed this complaint to card at confidence 0.652; the recorded label is credit_reporting.
  • tier_b2 routed this complaint to card at confidence 0.647; the recorded label is deposit_account.
  • tier_b2 routed this complaint to vehicle_loan at confidence 0.633; the recorded label is debt_collection.
  • tier_a_logreg routed this complaint to deposit_account at confidence 0.794; the recorded label is money_service.
  • tier_b2 routed this complaint to credit_reporting at confidence 0.701; the recorded label is vehicle_loan.

8 RECORDED TIERS · COST vs MACRO-F1

Accuracy and cost
share the same axis.

Tiermacro-F1Cost / 1kRun
Tier A — TF-IDF LogReg0.761 [0.756, 0.764]$933.418e4d6345b8…
Tier A — TF-IDF ComplementNB0.727 [0.722, 0.731]$1032.39c20cd14a7a…
Tier B1 — ModernBERT-base (seed a)0.788 [0.784, 0.792]$862.608071d31db2…
Tier B1 — ModernBERT-base (seed b)0.788 [0.784, 0.791]$865.58adb96307e2…
Tier B1 — ModernBERT-base (seed c)0.786 [0.782, 0.790]$867.94a523049a7e…
Tier B2 — DistilBERT (deployment point)0.795 [0.791, 0.799]$834.325517ebf1df…
Tier C — Claude Haiku 4.5 (TEST-IID)0.770 [0.750, 0.789]$916.9270a1b0c40e…
Tier C — Claude Sonnet 5 (TEST-IID subsample)0.742 [0.701, 0.773]$955.66e1503146b7…

DRIFT 2022-H2 – 2026-H1 · MACRO_F1

The model
aged out.

Every trained tier loses ground as the complaint mix moves past its training window; the two Claude tiers, scored zero-shot with no fitted vocabulary, do not. Tier B1's yearly series is not in the source yet and is shown as pending, not zero.

Tier B1 — ModernBERT-baseUNMEASURED
Tier2022-H22023202420252026-H1
Tier A — TF-IDF LogReg0.761 [0.756, 0.764]0.758 [0.750, 0.766]0.748 [0.738, 0.757]0.729 [0.721, 0.739]0.666 [0.658, 0.673]
Tier B2 — DistilBERT0.795 [0.791, 0.799]0.790 [0.782, 0.798]0.774 [0.765, 0.782]0.760 [0.751, 0.769]0.726 [0.719, 0.733]
Tier C — Claude Haiku 4.50.770 [0.750, 0.789]0.736 [0.699, 0.769]0.762 [0.727, 0.792]0.764 [0.730, 0.790]0.728 [0.705, 0.751]
Tier C — Claude Sonnet 50.742 [0.701, 0.773]0.754 [0.719, 0.786]0.780 [0.746, 0.807]0.763 [0.727, 0.791]0.782 [0.759, 0.805]
Tier B1 — ModernBERT-basePENDING — NOT MEASUREDPENDING — NOT MEASUREDPENDING — NOT MEASUREDPENDING — NOT MEASUREDPENDING — NOT MEASURED

Source note: Tier B1 yearly drift is pending in the source; no values were measured.

SOURCE · RECEIPTS

Every number opens
the same file.

What was verified
Recorded routing policies expose the cost/error tradeoff, known failures and the curated Python-to-browser parity check.
Evidence class
Compact exports of recorded triage runs, plus a separately served quantized browser model.
Boundary
Pending Tier B1 drift rows remain pending. Curated parity is not general accuracy, and the model file is not committed to Git.
Files, hashes and methods
public/case-studies/triage-router/frontier.compact.jsonLucisZhang/portfolio-site · acf05ae78859
sha256:c7e6555d3b03d988da30403f6052abf5e0be778e17f26b6eaa9c6696fc5eefcc
public/case-studies/triage-router/policies.compact.jsonLucisZhang/portfolio-site · acf05ae78859
sha256:9c2992c53dd62c7412a396e5c175e68068235a37548f74898b5858f258dc5c12
public/case-studies/triage-router/strategy-cards.jsonLucisZhang/portfolio-site · acf05ae78859
sha256:a26dd35e2c911aed93894bf7044be9c3a167ff9ea874cfbd4b432570ca6a868a
public/case-studies/triage-router/drift.compact.jsonLucisZhang/portfolio-site · acf05ae78859
sha256:ea70d32187602f644e895a151e3320991dd1dfc187441bf101c2e49587dc307f
public/case-studies/triage-router/known-failures.jsonLucisZhang/portfolio-site · acf05ae78859
sha256:fdcadaafa15e885890984cbab4b5320b46eadf6a3f374f2dcb7d76bf70e35218
public/case-studies/triage-router/samples.curated.jsonLucisZhang/portfolio-site · acf05ae78859
sha256:3c6771b65f44c567082857acd4db08a30004be2c4dc562974bef7a80cb386e4b
public/case-studies/triage-router/python_int8_curated.jsonLucisZhang/triage-router · b2734bbcbd75
sha256:484b21fab582af7a8b00c98abd89ecb575a5b1330dc7af791bde333b59953486
public/models/triage-tier-b2/model.int8.onnxServed model; no committed GitHub file.
sha256:da931ec8310cf1280747e22fc6ebfd30fd5f92e312ede6544042e1190764bb4a

The compact payloads above are built by Triage Router's export_site_payloads.py from real recorded triage runs; the model file is copied on-site and never enters git. docs/evidence/r2-source-map.md pins every path and hash; npm run verify:r2-sources re-checks them on every run.

python scripts/export_site_payloads.py --out public/case-studies/triage-router

Architecture

I built the full ladder: calibrated TF-IDF baselines, ModernBERT and DistilBERT fine-tunes, Claude tiers over OpenRouter, the explicit cost model, the cascade router, and the append-only results log every number traces back to.

  1. Freeze

    Snapshot the CFPB corpus by hash and split it by time: train, calibration, IID test, drift slices.

  2. Ladder

    Train each tier on identical splits: TF-IDF, ModernBERT ×3 seeds, DistilBERT int8, Claude zero-shot.

  3. Cost

    Price every path with an explicit model: inference dollars, misroute cost, human review.

  4. Route

    Fit cascade thresholds on the calibration slice only — never on test.

  5. Drift

    Replay 2023–2026 yearly slices and decompose what actually moved: class mix, not vocabulary.

Results & negatives

On held-out data the cascade beats the linear baseline by +0.0370 macro-F1 (95% CI +0.0337 to +0.0402) while cutting expected cost by $120.58 per thousand complaints. On the 2026-H1 drift slice the linear tier drops 9.2 points; Claude Sonnet gains ground instead.

  1. I assumed the expensive tier was the safe default. On the same 1,500 held-out IID rows, Sonnet 5 minus Haiku 4.5 comes out at −0.0073 macro-F1, CI [−0.0427, +0.0263], McNemar p=1.00 — indistinguishable, at 2.8× the cost per thousand complaints ($3.66 against $1.32) and 2.3× the p50 latency. It earns its price only on the post-cutoff slice, where the paired gain is +0.0458 macro-F1.

  2. Under the shared 64-token completion budget the stronger model sometimes answers with nothing at all: 12 of 1,500 calls on the IID slice (0.8%) and 37 of 1,500 on the post-cutoff slice (2.5%), every one with finish_reason "length" and no parseable JSON — its reasoning consumed the budget before an answer appeared. Haiku did that on 0 of 10,000 calls. The failures are worst on the drifted slice, which is exactly where you would escalate to the stronger model.

  3. I thought Tier C was pinned to Anthropic. Reading the request builder showed nothing of the kind: the body never carried an OpenRouter provider preference — no order, no only, no allow_fallbacks — so the pin never was in effect. Across all 16,050 committed Tier C calls, 16,020 (99.81%) were served by Amazon Bedrock, and every Phase 3 latency figure had to be relabeled as a route through OpenRouter to Bedrock.

Limitations

  1. A pre-registered hypothesis — that the fine-tuned transformer would behave like the LLM tiers under drift — was refuted: its 2026-H1 drop splits roughly evenly between class-mix shift and within-class decay.

  2. The obvious explanation did not survive either: out-of-vocabulary rate barely moved, so lexical drift is ruled out; the cliff is class mix.

  3. Results apply to this task, this corpus, and the price sheets at measurement time; the router's dollar savings follow from the disclosed cost model, not a universal claim.