Know the frontier.
Then forge past it.
Scaling free rule labels from 1,450 to 20,000 raised frozen-eval task success from 66.35% to 99.05%: +32.70 percentage points, paired 95% CI [30.60, 34.50], in one training seed (seed 0), using 15.236 measured RTX 4090 GPU-hours ($4.571).
The release model is the rule-label scaling ablation, not GRPO. GRPO's paired 95% CI includes zero across both completed seeds; a third seed aborted on the zero-reward-variance guard.
The same instrument,
full size, zero clicks.
RECORDED EVALUATION · OFFLINE REPLAY
phase4_spec_decode_r1b_bf16_native_mtp · sha256:7878b55f6f…COMPLAINT TRANSCRIPT — NOT RECORDED. release.json ships aggregate serving metrics, not per-request text.
curl -s https://xiangguozhang.com/case-studies/frontier-forge/release.json | jq '.serving.serving_at_4_qps[] | select(.run_id == "phase4_spec_decode_r1b_bf16_native_mtp")'RELEASE.JSON · CLAIM REGISTRY · SHA-256
Every number on this page
opens the command that made it.
Blank command cells mean the receipt recorded no shell command; config paths appear only where present.
| Claim | Number | n · CI | Generation command | SHA-256 |
|---|---|---|---|---|
| PositiveFree-label SFT reached the selected task-success result | 99.05% | n=2,000 · 95% CI 98.60%–99.45% | make reproduce-headline | 673138d9888b… |
| NegativeAPI distillation lost to the smaller free-label SFT run | 52.15% · −14.20 pp vs R1 | 95% CI 50.00%–54.40% | make reproduce-headline | 673138d9888b… |
| NegativeGRPO's paired interval includes zero | +0.25 pp | n=2,000 · 95% CI −0.10 pp–+0.65 pp | 673138d9888b… | |
| NegativeThe alternative training backend failed the agreement gate | −3.85 pp | n=2,000 · paired 95% CI −4.85 pp–−2.75 pp | config: configs/r1_sft_rule.yaml | 06e3f0c235f6… |
| PositiveGPTQ-int4 recorded the lowest 4 QPS p95 | 0.963 s p95 | n=20 · task-success Wilson 95% CI 76.39%–99.11% | c99b42cf0e06… | |
| BoundaryNative MTP has a measured win/lose boundary | 0.25 QPS lose; 0.50-4.00 QPS win | 7878b55f6fe6… | ||
| PositiveAt 3× overload the gateway rejected excess work with zero upstream 5xx | 215 HTTP 429 · 0.00% upstream 5xx | n=707 scheduled per side | 9e547b3418ce… | |
| NegativeBare vLLM crashed in the matched 5× cell | 651 transport errors · gateway 687 HTTP 429 | n=1,177 scheduled per side | make reproduce-headline | 9e547b3418ce… |
| PositiveThe CPU gateway scaled out and returned to one replica | 1→3→1 replicas | n=3,799 recorded HTTP 429 | 9e547b3418ce… | |
| BoundaryGPU cold start is a batch boundary, not an interactive SLO | p50 124.6 s · p95 127.1 s | n=10 | 9e547b3418ce… |
99.05%PositiveFree-label SFT reached the selected task-success result
52.15% · −14.20 pp vs R1NegativeAPI distillation lost to the smaller free-label SFT run
+0.25 ppNegativeGRPO's paired interval includes zero
−3.85 ppNegativeThe alternative training backend failed the agreement gate
0.963 s p95PositiveGPTQ-int4 recorded the lowest 4 QPS p95
0.25 QPS lose; 0.50-4.00 QPS winBoundaryNative MTP has a measured win/lose boundary
215 HTTP 429 · 0.00% upstream 5xxPositiveAt 3× overload the gateway rejected excess work with zero upstream 5xx
651 transport errors · gateway 687 HTTP 429NegativeBare vLLM crashed in the matched 5× cell
1→3→1 replicasPositiveThe CPU gateway scaled out and returned to one replica
p50 124.6 s · p95 127.1 sBoundaryGPU cold start is a batch boundary, not an interactive SLO
1,450 → 20,000 RULE LABELS
Free labels moved
the frontier.
Scaling free rule labels from 1,450 to 20,000 raised frozen-eval task success from 66.35% to 99.05%: +32.70 percentage points, paired 95% CI [30.60, 34.50], in one training seed (seed 0), using 15.236 measured RTX 4090 GPU-hours ($4.571).
- R0 baseComplete
0.00%95% CI 0.00%–0.00% · 0.98 GPU-h · $0.29
- R1 rule SFT (1,450)Complete
66.35%95% CI 64.20%–68.40% · 3.48 GPU-h · $1.04
- R1b rule SFT (20,000)Release-selected
99.05%95% CI 98.60%–99.45% · 15.24 GPU-h · $4.57
- R2 distilled SFTComplete — negative
52.15%95% CI 50.00%–54.40% · 3.61 GPU-h · $1.08
- R3 DPOComplete
55.95%95% CI 53.75%–58.25% · 1.93 GPU-h · $0.58
- R4 v2 GRPO seed 0Partial only
56.20%95% CI 54.00%–58.50% · 1.86 GPU-h · $0.56
- R4 v2 GRPO seed 1Partial only
56.20%95% CI 53.95%–58.50% · 1.66 GPU-h · $0.50
NEGATIVE RESULT
NEGATIVE RESULT
0.25–4.00 QPS · NATIVE MTP
Serving is a boundary,
not a badge.
At 4 QPS, three precisions were served and measured end to end. Native MTP's win/lose boundary: 0.25 QPS lose; 0.50-4.00 QPS win.
SAME-BOX A10 · SUSTAINED OVERLOAD
Reject the work
before it becomes a crash.
Fixed-seed Poisson arrivals, same-box NVIDIA A10, every cell ran at least 120 seconds; each cell's queue was verified saturated.
| Multiplier | Offered QPS | Requests | HTTP 429 | Gateway upstream 5xx | Bare vLLM upstream 5xx | Gate |
|---|---|---|---|---|---|---|
2× | 4 | 510 | 19 | 0.0% | 0.0% | PASS |
3× | 6 | 707 | 215 | 0.0% | 0.0% | PASS |
5× | 10 | 1,177 | 687 | 0.0% | 3.1% | PASS |
At the highest recorded load (5×, 10 QPS): bare vLLM logged 651 transport errors while the gateway returned 687 bounded HTTP 429 rejects and zero upstream 5xx.
MODEL BOUNDARY
The model has a boundary.
It is drawn here.
4 capabilities hold up under measurement; 5 do not — the failing count is not hidden below the passing one.
- Reaches 99.05% task success on complaint routing (n=2,000 paired, 95% CI 98.60%–99.45%)
- Holds schema-valid tool calls at 100% under both xgrammar and outlines constraints
- Gateway upstream 5xx stayed at 0% through 5× sustained overload (bare vLLM did not)
- Native MTP wins on p95 latency across 0.50-4.00 QPS
- Constrained decoding's simultaneous mode (no two-pass retry) has 0% task success under both xgrammar and outlines
- API distillation lost 14.20 pp to the smaller free-label SFT run (52.15% vs 66.35%)
- GRPO's paired 95% CI includes zero across both completed seeds; a third seed aborted on a zero-reward-variance guard
- Below 0.5 QPS, native MTP is slower, not faster (0.25 QPS: p95 +0.051s, lose)
- The alternative training backend (unsloth) failed the agreement gate against the default (trl): −3.85 pp, 95% CI −4.85 pp–−2.75 pp
KNOWN FAILURES
Per-sample misclassified complaints — NOT RECORDED. release.json ships aggregate task_success and paired confidence intervals only, not individual predictions.
The closest real, cited failure mode: neither constrained-decoding backend produced a usable first-pass answer without a second pass.
phase4_structured_r1b_bf16_xgrammar— xgrammar: simultaneous task success 0% (24 requests); two-pass recovers to 100% at +0.81s p50 latency.phase4_structured_r1b_bf16_outlines— outlines: simultaneous task success 0% (24 requests); two-pass recovers to 100% at +0.88s p50 latency.
ff-qwen3.5-4b-r1b_sft_rule_20k_s0 · release.json sha256:9e547b3418ce1f63…
HOW THIS WAS VERIFIED
Every claim opens
the same command.
- What was verified
- The release registry connects training outcomes, measured spend and the sustained overload test to recorded runs.
- Evidence class
- Recorded GPU experiments and a preserved A10 load-test receipt; the browser replays that receipt.
- Boundary
- The release uses rule-label scaling. GRPO's completed-seed confidence intervals include zero; the third seed aborted. Hardware and workload boundaries remain specific to each run.
Files, hashes and methods
- release.jsonLucisZhang/frontier-forge ·
06de6e5c1d02 sha256:9e547b3418ce1f633914e8dfe7b818fa6f7e064a891cee8e4c73837e6ce45f4d- Overload replay receiptLucisZhang/frontier-forge ·
06de6e5c1d02 sha256:d132ebece698c98ab996b154067679ea1d74b52e9528834ed19a21a132137929- Reproduce the headline
make reproduce-headline- Generated
- 2026-08-22
The claim table is built from the copied Phase 7 release.json; the local manifest pins its exact SHA-256. The overload replay fetches the preserved Phase 7.1 sustained A10 receipt and does not call a model.
Architecture
I trained and evaluated the ladder, exported the selected checkpoint, served it with vLLM, wrote the C++20 gateway, and exercised the k3s runtime.
- Frozen input
Hash-pinned CFPB splits feed rule labels and API distillation.
- Training ladder
R0 → R1/R1b → R2 → R3 → R4, with paired confidence intervals.
- Model exports
The selected R1b becomes BF16, GPTQ-int4, and MTP-preserved artifacts.
- vLLM serving
Client/server timing records the precision and native-MTP boundaries.
- Bounded gateway
A C++20 token-aware admission layer protects the OpenAI-compatible JSON/SSE path.
Results & negatives
Under 3× overload the gateway sheds load with 429s and zero upstream 5xx; bare vLLM crashed at 5×. Distillation lost 14.2 pp to free rule labels and GRPO's CI includes zero — both runs are kept on the page.
Distillation lost 14.2 pp to free rule labels: 52.15% versus the smaller R1 run's 66.35%.
GRPO's paired 95% confidence interval includes zero; one additional seed was stopped by the unchanged zero-reward-variance guard.
Limitations
The lifted production block applies only to the measured single-node gateway overload contract; it does not establish cloud production or multi-GPU scaling.
Only CPU gateway replicas scaled 1→3→1. The GPU deployment moved only between zero and one replica on one physical A10.
The roughly 125-second GPU cold start is suited to batch and development workloads, not an interactive serving SLO.