• RECORDED EVALUATION
  • OFFLINE REPLAY

Know the frontier.
Then forge past it.

Scaling free rule labels from 1,450 to 20,000 raised frozen-eval task success from 66.35% to 99.05%: +32.70 percentage points, paired 95% CI [30.60, 34.50], in one training seed (seed 0), using 15.236 measured RTX 4090 GPU-hours ($4.571).

99.1%TASK SUCCESS
+32.7 ppGAIN VS R1
15.24RTX 4090 GPU-HOURS
$4.571MEASURED TRAINING COST

The release model is the rule-label scaling ablation, not GRPO. GRPO's paired 95% CI includes zero across both completed seeds; a third seed aborted on the zero-reward-variance guard.

  • FULL INSTRUMENT
  • RECORDED SERVING RUNS

The same instrument,
full size, zero clicks.

RECORDED EVALUATION · OFFLINE REPLAY

1.06s320.6 tok/s · 20 reqs · 95% success · RTX 4090phase4_spec_decode_r1b_bf16_native_mtp · sha256:7878b55f6f…

COMPLAINT TRANSCRIPT — NOT RECORDED. release.json ships aggregate serving metrics, not per-request text.

curl -s https://xiangguozhang.com/case-studies/frontier-forge/release.json |   jq '.serving.serving_at_4_qps[] | select(.run_id == "phase4_spec_decode_r1b_bf16_native_mtp")'

RELEASE.JSON · CLAIM REGISTRY · SHA-256

Every number on this page
opens the command that made it.

Blank command cells mean the receipt recorded no shell command; config paths appear only where present.

ClaimNumbern · CIGeneration commandSHA-256
PositiveFree-label SFT reached the selected task-success result99.05%n=2,000 · 95% CI 98.60%–99.45%make reproduce-headline673138d9888b
NegativeAPI distillation lost to the smaller free-label SFT run52.15% · −14.20 pp vs R195% CI 50.00%–54.40%make reproduce-headline673138d9888b
NegativeGRPO's paired interval includes zero+0.25 ppn=2,000 · 95% CI −0.10 pp–+0.65 pp673138d9888b
NegativeThe alternative training backend failed the agreement gate−3.85 ppn=2,000 · paired 95% CI −4.85 pp–−2.75 ppconfig: configs/r1_sft_rule.yaml06e3f0c235f6
PositiveGPTQ-int4 recorded the lowest 4 QPS p950.963 s p95n=20 · task-success Wilson 95% CI 76.39%–99.11%c99b42cf0e06
BoundaryNative MTP has a measured win/lose boundary0.25 QPS lose; 0.50-4.00 QPS win7878b55f6fe6
PositiveAt 3× overload the gateway rejected excess work with zero upstream 5xx215 HTTP 429 · 0.00% upstream 5xxn=707 scheduled per side9e547b3418ce
NegativeBare vLLM crashed in the matched 5× cell651 transport errors · gateway 687 HTTP 429n=1,177 scheduled per sidemake reproduce-headline9e547b3418ce
PositiveThe CPU gateway scaled out and returned to one replica1→3→1 replicasn=3,799 recorded HTTP 4299e547b3418ce
BoundaryGPU cold start is a batch boundary, not an interactive SLOp50 124.6 s · p95 127.1 sn=109e547b3418ce
  • 99.05%Positive

    Free-label SFT reached the selected task-success result

    n=2,000 · 95% CI 98.60%–99.45%

    make reproduce-headline

    SHA-256 673138d9888b

  • 52.15% · −14.20 pp vs R1Negative

    API distillation lost to the smaller free-label SFT run

    95% CI 50.00%–54.40%

    make reproduce-headline

    SHA-256 673138d9888b

  • +0.25 ppNegative

    GRPO's paired interval includes zero

    n=2,000 · 95% CI −0.10 pp–+0.65 pp

    SHA-256 673138d9888b

  • −3.85 ppNegative

    The alternative training backend failed the agreement gate

    n=2,000 · paired 95% CI −4.85 pp–−2.75 pp

    config: configs/r1_sft_rule.yaml

    SHA-256 06e3f0c235f6

  • 0.963 s p95Positive

    GPTQ-int4 recorded the lowest 4 QPS p95

    n=20 · task-success Wilson 95% CI 76.39%–99.11%

    SHA-256 c99b42cf0e06

  • 0.25 QPS lose; 0.50-4.00 QPS winBoundary

    Native MTP has a measured win/lose boundary

    SHA-256 7878b55f6fe6

  • 215 HTTP 429 · 0.00% upstream 5xxPositive

    At 3× overload the gateway rejected excess work with zero upstream 5xx

    n=707 scheduled per side

    SHA-256 9e547b3418ce

  • 651 transport errors · gateway 687 HTTP 429Negative

    Bare vLLM crashed in the matched 5× cell

    n=1,177 scheduled per side

    make reproduce-headline

    SHA-256 9e547b3418ce

  • 1→3→1 replicasPositive

    The CPU gateway scaled out and returned to one replica

    n=3,799 recorded HTTP 429

    SHA-256 9e547b3418ce

  • p50 124.6 s · p95 127.1 sBoundary

    GPU cold start is a batch boundary, not an interactive SLO

    n=10

    SHA-256 9e547b3418ce

1,450 → 20,000 RULE LABELS

Free labels moved
the frontier.

Scaling free rule labels from 1,450 to 20,000 raised frozen-eval task success from 66.35% to 99.05%: +32.70 percentage points, paired 95% CI [30.60, 34.50], in one training seed (seed 0), using 15.236 measured RTX 4090 GPU-hours ($4.571).

  1. R0 baseComplete0.00%

    95% CI 0.00%–0.00% · 0.98 GPU-h · $0.29

  2. R1 rule SFT (1,450)Complete66.35%

    95% CI 64.20%–68.40% · 3.48 GPU-h · $1.04

  3. R1b rule SFT (20,000)Release-selected99.05%

    95% CI 98.60%–99.45% · 15.24 GPU-h · $4.57

  4. R2 distilled SFTComplete — negative52.15%

    95% CI 50.00%–54.40% · 3.61 GPU-h · $1.08

  5. R3 DPOComplete55.95%

    95% CI 53.75%–58.25% · 1.93 GPU-h · $0.58

  6. R4 v2 GRPO seed 0Partial only56.20%

    95% CI 54.00%–58.50% · 1.86 GPU-h · $0.56

  7. R4 v2 GRPO seed 1Partial only56.20%

    95% CI 53.95%–58.50% · 1.66 GPU-h · $0.50

NEGATIVE RESULT

Distillation lost 14.2 pp to free rule labels: 52.15% versus the smaller R1 run's 66.35%.

NEGATIVE RESULT

GRPO's paired 95% confidence interval includes zero; one additional seed was stopped by the unchanged zero-reward-variance guard.

0.25–4.00 QPS · NATIVE MTP

Serving is a boundary,
not a badge.

At 4 QPS, three precisions were served and measured end to end. Native MTP's win/lose boundary: 0.25 QPS lose; 0.50-4.00 QPS win.

1.42sE2E P50
1.69sE2E P95
0.17sTTFT P50
307.0OUTPUT TOK/S
95%TASK SUCCESS
$0.0211COST / 1K TASKS
22,829 MiBVRAM PEAK
n=20REQUESTS
0.25 QPSLOSE+0.051s p95
0.5 QPSWIN-0.229s p95
1 QPSWIN-0.227s p95
2 QPSWIN-0.321s p95
4 QPSWIN-0.370s p95

SAME-BOX A10 · SUSTAINED OVERLOAD

Reject the work
before it becomes a crash.

Fixed-seed Poisson arrivals, same-box NVIDIA A10, every cell ran at least 120 seconds; each cell's queue was verified saturated.

MultiplierOffered QPSRequestsHTTP 429Gateway upstream 5xxBare vLLM upstream 5xxGate
2×4510190.0%0.0%PASS
3×67072150.0%0.0%PASS
5×101,1776870.0%3.1%PASS

At the highest recorded load (5×, 10 QPS): bare vLLM logged 651 transport errors while the gateway returned 687 bounded HTTP 429 rejects and zero upstream 5xx.

MODEL BOUNDARY

The model has a boundary.
It is drawn here.

4 capabilities hold up under measurement; 5 do not — the failing count is not hidden below the passing one.

  • Reaches 99.05% task success on complaint routing (n=2,000 paired, 95% CI 98.60%–99.45%)
  • Holds schema-valid tool calls at 100% under both xgrammar and outlines constraints
  • Gateway upstream 5xx stayed at 0% through 5× sustained overload (bare vLLM did not)
  • Native MTP wins on p95 latency across 0.50-4.00 QPS
  • Constrained decoding's simultaneous mode (no two-pass retry) has 0% task success under both xgrammar and outlines
  • API distillation lost 14.20 pp to the smaller free-label SFT run (52.15% vs 66.35%)
  • GRPO's paired 95% CI includes zero across both completed seeds; a third seed aborted on a zero-reward-variance guard
  • Below 0.5 QPS, native MTP is slower, not faster (0.25 QPS: p95 +0.051s, lose)
  • The alternative training backend (unsloth) failed the agreement gate against the default (trl): −3.85 pp, 95% CI −4.85 pp–−2.75 pp

KNOWN FAILURES

Per-sample misclassified complaints — NOT RECORDED. release.json ships aggregate task_success and paired confidence intervals only, not individual predictions.

The closest real, cited failure mode: neither constrained-decoding backend produced a usable first-pass answer without a second pass.

  • phase4_structured_r1b_bf16_xgrammar — xgrammar: simultaneous task success 0% (24 requests); two-pass recovers to 100% at +0.81s p50 latency.
  • phase4_structured_r1b_bf16_outlines — outlines: simultaneous task success 0% (24 requests); two-pass recovers to 100% at +0.88s p50 latency.

ff-qwen3.5-4b-r1b_sft_rule_20k_s0 · release.json sha256:9e547b3418ce1f63

HOW THIS WAS VERIFIED

Every claim opens
the same command.

What was verified
The release registry connects training outcomes, measured spend and the sustained overload test to recorded runs.
Evidence class
Recorded GPU experiments and a preserved A10 load-test receipt; the browser replays that receipt.
Boundary
The release uses rule-label scaling. GRPO's completed-seed confidence intervals include zero; the third seed aborted. Hardware and workload boundaries remain specific to each run.
Files, hashes and methods
release.jsonLucisZhang/frontier-forge · 06de6e5c1d02
sha256:9e547b3418ce1f633914e8dfe7b818fa6f7e064a891cee8e4c73837e6ce45f4d
Overload replay receiptLucisZhang/frontier-forge · 06de6e5c1d02
sha256:d132ebece698c98ab996b154067679ea1d74b52e9528834ed19a21a132137929
Reproduce the headline
make reproduce-headline
Generated
2026-08-22

The claim table is built from the copied Phase 7 release.json; the local manifest pins its exact SHA-256. The overload replay fetches the preserved Phase 7.1 sustained A10 receipt and does not call a model.

Architecture

I trained and evaluated the ladder, exported the selected checkpoint, served it with vLLM, wrote the C++20 gateway, and exercised the k3s runtime.

  1. Frozen input

    Hash-pinned CFPB splits feed rule labels and API distillation.

  2. Training ladder

    R0 → R1/R1b → R2 → R3 → R4, with paired confidence intervals.

  3. Model exports

    The selected R1b becomes BF16, GPTQ-int4, and MTP-preserved artifacts.

  4. vLLM serving

    Client/server timing records the precision and native-MTP boundaries.

  5. Bounded gateway

    A C++20 token-aware admission layer protects the OpenAI-compatible JSON/SSE path.

Results & negatives

Under 3× overload the gateway sheds load with 429s and zero upstream 5xx; bare vLLM crashed at 5×. Distillation lost 14.2 pp to free rule labels and GRPO's CI includes zero — both runs are kept on the page.

  1. Distillation lost 14.2 pp to free rule labels: 52.15% versus the smaller R1 run's 66.35%.

  2. GRPO's paired 95% confidence interval includes zero; one additional seed was stopped by the unchanged zero-reward-variance guard.

Limitations

  1. The lifted production block applies only to the measured single-node gateway overload contract; it does not establish cloud production or multi-GPU scaling.

  2. Only CPU gateway replicas scaled 1→3→1. The GPU deployment moved only between zero and one replica on one physical A10.

  3. The roughly 125-second GPU cold start is suited to batch and development workloads, not an interactive serving SLO.