SQL WORKBENCH · CACHED STATE · BUILD 2026-08-22
Run the argument, query by query.
Architecture
SQL workbench
Six curated queries walk from raw scale to the study's null result and its one honest counterexample. Every result on this screen is cached from the full local build — the workbench replays it exactly, and says so.
This workbench renders pre-aggregated, cached views of the committed Iceberg snapshot; the full Spark batch results appear in the analysis exhibit below.
-- How large is the current contract-checked silver data plane? SELECT 'silver.interactions' AS dataset, count(*) AS row_count, count(DISTINCT user_id) AS entity_count FROM silver_interactions UNION ALL SELECT 'silver.items' AS dataset, count(*) AS row_count, count(DISTINCT parent_asin) AS entity_count FROM silver_items ORDER BY dataset LIMIT 20
179,901,744 rows scanned · 1922.04 ms · cached · build 2026-08-22
| dataset | row_count | entity_count |
|---|---|---|
| silver.interactions | 43,365,424 | 18,286,190 |
| silver.items | 1,610,012 | 1,610,012 |
ICEBERG
- snapshotId
- 884031112460958161
- committedAt
- 2026-08-05
- schema
- v2
- rowCount
- 43,365,424
- files
- 50
- bytes
- 1.36 GB
NDCG@10 BY HISTORY DEPTH · AMAZON ELECTRONICS + MOVIELENS-32M
One curve never crosses. The other crosses at n*=20.
Same metric, same paired-bootstrap discipline, two catalogs. The gap between a personalized model and recency-weighted popularity either closes or it doesn't — these two panels are that comparison, not a summary of it.
AMAZON ELECTRONICS · POPULARITY VS ALS
No crossover on Amazon
ALS stays below recency-weighted popularity at every observed history depth. The gap narrows from 1–4 interactions to 20+, but never reaches zero.
20260805T172047Z-035042bALS (rank 128, seed 20260805)20260806T082441Z-2f2f26d| series | 0 | 1-4 | 5-9 | 10-19 | 20+ |
|---|---|---|---|---|---|
| pop-t12m | 0.0075 | 0.0058 | 0.0052 | 0.0051 | 0.0037 |
| ALS (rank 128, seed 20260805) | 0.0000 | 0.0033 | 0.0029 | 0.0025 | 0.0023 |
MOVIELENS-32M · POPULARITY VS ITEM-KNN
The crossover appears at n*=20
With 6.40% catalog churn, trailing-window item-kNN overtakes popularity from 20 interactions onward on NDCG@10, after Benjamini–Hochberg correction.
20260820T221055Z-20d8ff9item-kNN t12m20260820T221701Z-20d8ff9| series | 0 | 1-4 | 5-9 | 10-19 | 20+ |
|---|---|---|---|---|---|
| pop-t12m | 0.4412 | 0.3610 | 0.2873 | 0.2457 | 0.0922 |
| item-kNN t12m | 0.0568 | 0.3471 | 0.2836 | 0.2017 | 0.1052 |
SOURCE · RECEIPTS
Every curve opens
the same six runs.
- What was verified
- Recorded evaluations retain the Amazon null result and the ML-32M crossover in ranking quality.
- Evidence class
- Single-machine batch evaluations with frozen splits and individual run receipts; the SQL workbench replays cached results.
- Boundary
- Recall did not confirm the ML-32M winning buckets. The dataset contrast is non-causal and provides no online or production evidence.
Files, hashes and methods
The Amazon curve resolves to runs 20260805T172047Z-035042b and 20260806T082441Z-2f2f26d; the 41.11% mechanism resolves to 20260817T095926Z-633d454.
The ML-32M curve resolves to runs 20260820T221055Z-20d8ff9 and 20260820T221701Z-20d8ff9. The page projection preserves config hashes, dataset hashes, snapshot IDs, seeds, hardware, and wall-clock time for all six source runs.
20260805T172047Z-035042bpopularity
- git SHA
035042bbc4ce5284a4b7afd46f5dca574b5ddee4- Recorded
- 2026-08-05T17:30:18.009394+00:00
- Config
configs/eval_pop_t12m_test.yamlLucisZhang/crossover-study ·035042bbc4ce- config SHA-256
sha256:1e60f2621f29dda053119282bfd8dbc3c824df0af7d95b8d9aeca2f86a1500ed- dataset SHA-256
sha256:4a553c348bfccb05918362192692ed5fe6965e3e2662fe751cf08d3b9f9da849- Frozen split
- 2026-08-05
- Runtime
- 570.663 s · arm64 · Darwin
20260806T082441Z-2f2f26dals
- git SHA
2f2f26d55129eabadfd84012146c8a202d61e1c2- Recorded
- 2026-08-06T08:46:12.872525+00:00
- Config
configs/eval_als_test_seed1.yamlLucisZhang/crossover-study ·2f2f26d55129- config SHA-256
sha256:c50679bf1b89ed1bb82d7914b089a887a6ff91332c228b5ff67724966606d31e- dataset SHA-256
sha256:4a553c348bfccb05918362192692ed5fe6965e3e2662fe751cf08d3b9f9da849- Frozen split
- 2026-08-05
- Runtime
- 1292.102 s · arm64 · Darwin
20260817T095926Z-633d454regime_map
- git SHA
633d4543605b68e03a6e770ade46764f598dd5e9- Recorded
- 2026-08-17T09:59:26.571790+00:00
- Config
configs/regime_map_test.yamlLucisZhang/crossover-study ·633d4543605b- config SHA-256
sha256:32c1c06a860eeb62ae42fe0c1c8e5c7d2b528aeb65d1a0fcf59adffbf0366806- dataset SHA-256
sha256:4a553c348bfccb05918362192692ed5fe6965e3e2662fe751cf08d3b9f9da849- Frozen split
- 2026-08-05
- Runtime
- 41.002 s · arm64 · Darwin
20260820T134403Z-e2263d2churn_contrast
- git SHA
e2263d24e03a0a61d2743314f60cdd21c289dadb- Recorded
- 2026-08-20T13:44:03.648361+00:00
- Config
configs/churn_contrast_ml32m.yamlLucisZhang/crossover-study ·e2263d24e03a- config SHA-256
sha256:682d2e89e2c86cd71d46c43b844b70fd54f5f6b3a92ec7cd1285a6599cf3505f- dataset SHA-256
sha256:1f830512534c39f3768e2f8b03a05d950c77fbb927e51b072a9eb541432e98a2- Frozen split
- 2026-08-19
- Runtime
- 26.124 s · x86_64 · Linux
20260820T221055Z-20d8ff9popularity
- git SHA
20d8ff995f4baceda31a9eedfc78effd632099ce- Recorded
- 2026-08-20T22:11:55.106578+00:00
- Config
configs/eval_pop_t12m_ml32m_test.yamlLucisZhang/crossover-study ·20d8ff995f4b- config SHA-256
sha256:0e8c7581986745417dfc534edfc0132a547303c80dbe8d90c9989002d31ee9b0- dataset SHA-256
sha256:1f830512534c39f3768e2f8b03a05d950c77fbb927e51b072a9eb541432e98a2- Frozen split
- 2026-08-19
- Runtime
- 58.601 s · x86_64 · Linux
20260820T221701Z-20d8ff9item_knn
- git SHA
20d8ff995f4baceda31a9eedfc78effd632099ce- Recorded
- 2026-08-20T22:18:30.876460+00:00
- Config
configs/eval_itemknn_t12m_ml32m_test.yamlLucisZhang/crossover-study ·20d8ff995f4b- config SHA-256
sha256:21212ee94a62fb72e684abc3a6d83355c96be6f257603996ca0e40c118d15d9a- dataset SHA-256
sha256:1f830512534c39f3768e2f8b03a05d950c77fbb927e51b072a9eb541432e98a2- Frozen split
- 2026-08-19
- Runtime
- 88.445 s · x86_64 · Linux
docs/evidence/digits-crossover.md pins every number on this page to one of these six runs or to the cached workbench's own committed JSON files.
public/case-studies/crossover-study/exhibits.jsonLucisZhang/portfolio-site · acf05ae78859
Results & negatives
Amazon produced no crossover at any observed depth. The mechanism measurement found 41.11% catalog churn. Under the same ladder on MovieLens-32M, where churn was 6.40%, item-kNN crossed popularity at n*=20 on NDCG@10.
On Amazon, ALS trails trailing-12-month popularity at every observed history depth. The deficit narrows from −0.00256 at 1–4 interactions to −0.00146 at 20+, but never reaches zero. The fitted policy is n*=∞: build no personalization-routing layer for this regime.
On ML-32M, item-kNN turns positive from 20 interactions onward on NDCG@10 after Benjamini–Hochberg correction. Recall@20 confirms none of the three winning buckets, so the result is ranking quality only; it is not a recall win.
The same ML-32M arm loses −0.3844 at zero history, a bucket containing 43.9% of test users. Popularity still wins globally. The crossover changes the deep-history tail; it does not reverse the deployment default.
Limitations
The Amazon null applies to this catalog, five-core population, temporal split, and observed history depths — not to recommenders in general.
Amazon and ML-32M differ in domain, density, catalog size, and feedback semantics. Their contrast measures a regime association; it does not isolate churn as a cause.
Single-machine batch evaluation only: no serving system, online metric, or A/B test. Raw reviews, per-user arrays, and MiniLM weights are not shipped with this page.