PorkiCoder Research · Tab titles

176k Flash-Lite CE

A quantity test after 6t. Four LRs, one pack, one Terminal packet. Ship stays 6t_argmax.

16 August 2026 · 18:20–18:29 PDT · Terminal-board only · Continues The Sniff Test, FLAN + centroid, mid GSG, four beams, session 4, shipping 6t_argmax

Still ship this

6t_argmax, sha 2ea02c13…. Same weights as production porkr1/resources/tab-namer/6t-argmax-centroid-v1. Beam-4 two-word + Hybrid A + centroid v2. This 176k run does not replace it.

Two boards

This page is Terminal-board only (300 terminal-dev). SO-board is not mixed in. Holdout 200 and sealed SO 1000 stay sealed. Product metric is blinded Gemini 3.5 Flash-Lite on Hybrid-A-glued titles, two decimals. CE has no ship weight.

Abstract

We stopped a 1M Flash-Lite labeling job at 219,800 raw rows (~$16) and trained on 175,642 vocab-legal titles. Four 35M continues (A 2e-5, B 5e-5, C pin+2e-5, D 1e-5) ran 8,000 steps on four Ada 6000s. CE fell 6.2 → ~1.5 (B to 0.91). Unique stayed ~290/300. On the first glue-only packet, A was 6.19 vs 6t 6.59 vs FLAN 7.05. After adding 6t’s beam-4 two-word decode, A is 6.05 vs 6t 6.16 vs FLAN 6.62 (Δ A−6t −0.112). Hottest CE (B) is still last. Quantity of teacher-CE did not beat the small judged set or FLAN.

What we trained on

GH Archive agent/PR pages, labeled by gemini-3.5-flash-lite (2–3 words, on-page overlap ≥ 0.5, 7k vocab-legal, ≤16 target tokens). 6t was judged-SFT on 2,040 reward≥6 titles. This pack is ~86× more rows and not pickier.

IDF 2-word stickers were tried first on the same 1M pages and produced gibberish. Flash-Lite labels replaced them. Next paid regen must say max 3 short common words — see required-specs.md.

Learning rates

Same pack, same 35M, 8k steps, batch 256. Notes written at 18:20 PDT before the 8k snap, confirmed by the packets below.

ArmInitLRFinal CEWhat we learned
A6t2e-51.55Workhorse. On-page held. Closest to 6t after beam-4.
B6t5e-50.91Prettiest curve. Raw on-page fell 0.32→0.18. Worst judge.
Cpin B-95002e-51.54A’s twin. LR dominates init on this teacher.
D6t1e-51.92Slow spare. Spikes were batch noise, not a reject.

Packet 1 — glue only (unfair to 6t)

Trainer beam + Hybrid A. 6t already had production beam-4. n=300, same judge, same centroid.

ArmMeanUseful ≥4Fail ≤2vs 6t
FLAN + centroid7.0593.3%1.7%
6t_argmax (ship)6.5992.3%2.3%
A 2e-56.1985.3%2.7%−0.399 (101–80–119)
D 1e-56.1386.0%3.0%−0.461
C pin 2e-56.1286.7%3.0%−0.467
B 5e-56.0684.7%3.3%−0.534

Do not compare these absolute means to packet 2. New judge draw. Paired Δ inside a packet is the fair number.

Packet 2 — full 6t stack (apples to apples)

A–D re-decoded with eval_snapshot_router.py: beam-4 two-word, Hybrid A bounded 32, centroid v2. Same 300, same 6t routed file, same FLAN Hybrid A, one blinded Flash-Lite packet.

ArmMeanUseful ≥4Fail ≤2vs 6t
FLAN + centroid6.6289.0%3.0%
6t_argmax (ship)6.1687.3%3.7%
A 2e-56.0585.0%3.0%−0.112 (73–149–78)
B 5e-56.0286.7%3.0%−0.135
C pin 2e-56.0085.3%3.0%−0.162
D 1e-55.9986.0%3.0%−0.164

Beam-4 closed most of the fake gap (−0.40 → −0.11). They still lose. B’s CE win is still a judge loss.

Learnings

Do not put back on the plate

Artifacts

LR notes: 2026-08-16_1820PDT_g300k_learning_rate_findings.md. Packets: output/g300k_eval/packet.summary.json, output/g300k_eval/beam4/packet.summary.json. Labels: data/derived/terminal/agent_1m_gemini/shards/ (Gitea). Weights: local output/g300k_train/*/final (not in git). Production 6t: porkr1/resources/tab-namer/6t-argmax-centroid-v1.