Session 4: we beat our own namer. FLAN is still ahead on terminal.
16 August 2026 · four-hour Terminal-board campaign · Flash-Lite judge on 300 held-out terminal tasks · Continues the SO-board papers FLAN + centroid, mid GSG, four beams (19/40 lives there and stays sealed)
Frozen names, use these in every later writeup. There are two product boards. They share a judge and a glue. They do not share a leaderboard.
- SO-board: 1,000 Stack Overflow clean-dev pages. Canonical FLAN bar is FLAN + centroid v2 = 6.11 / 88.5% useful. Our pin on that board is 5.96 / 85.0%. This is the near-miss. Cite only as SO-board.
- Terminal-board: 300 redacted terminal-dev tasks. Canonical FLAN bar is FLAN + centroid v2 = 7.11 / 93.3% useful (session 2.5). Our pin on that board is 5.94 / 83.3%. This is the open race. Cite only as Terminal-board.
Never write "we almost beat FLAN" without naming the board. Never average, rank, or promote across boards. A method that wins on SO-board is innocent until measured on Terminal-board (exact-three is the exhibit: +0.39 on SO-board, −0.36 on Terminal-board).
What this session did. Teaching the 35M namer to copy FLAN titles that can be spelled from the page nudged our terminal score up a little, and the gain repeated in a second blind packet. Copying our own best-of-8 or best-of-24 guesses did nothing. We are closer to our own pin, not to FLAN on this domain.
SO-board vs Terminal-board
Same judge, same 1 to 10 rubric, same Hybrid A + centroid glue. Different pages. A win on SO-board does not transfer to Terminal-board. That is why "we were almost there" and "we are 1.4 points back" are both true.
| Board (use this name) | Pages | FLAN + centroid v2 | Our pin (B-9500 + glue) | Gap |
|---|---|---|---|---|
| SO-board (near-miss, not the ship race) | 1,000 SO clean-dev | 6.11 / 88.5% useful | 5.96 / 85.0% useful | −0.15 |
| Terminal-board (ship race) | 300 terminal-dev | 7.11 / 93.3% useful | 5.94 / 83.3% useful | −1.17 |
SO-board source: 6-way Flash-Lite after centroid v2 (docs/gsg_mid_campaign_20260814.md). Terminal-board source: session 2.5 blinded packet. Session 4 trained and judged only on Terminal-board. SO-board appears later only as a regression check, not as a promotion surface.
Who is who
Three systems matter on this page. All three write a 1 to 3 word tab name, then fill leftover slots with centroid v2 (the current glue: trained B-9500 embeddings, SO IDF, residual, MMR).
- Pin (incumbent). The shipped 35M namer, B-9500, two-word beam, Hybrid A glue, centroid v2. This is what production already uses.
- r3 (this session). Same pin, then a short SFT pass on 971 FLAN titles that (a) can be spelled from on-page words and (b) scored at least 6 from the judge.
- FLAN + centroid v2 (the bar). Google's FLAN-T5-small writes a free title, then the same Hybrid A + centroid v2 glue we ship. Published terminal-dev numbers from session 2.5: mean 7.11, 93% useful. That is the benchmark this campaign exists to beat.
Today's result, still on the terminal board
Same 300 terminal-dev tasks, same judge, same 1 to 10 rubric. A title is useful if it scores 4 or higher. The FLAN row is copied from the session 2.5 packet above so you can see the bar; the pin and r3 rows are from this session's confirmation packet. Packet-to-packet noise on the same system is a few points. Read the FLAN column as the unfinished race. Read pin vs r3 as a small confirmed bump on our own namer.
| System | Mean | Useful (≥4) | Solid (≥6) | Strong (≥7) | Fails (≤2) |
|---|---|---|---|---|---|
| FLAN + centroid v2 (bar) | 7.11 | 93.3% | 79.0% | 67.3% | 2.7% |
| r3 (this session) | 5.67 | 79.0% | 55.3% | 35.3% | 8.0% |
| Pin (incumbent) | 5.56 | 78.0% | 52.0% | 33.0% | 8.0% |
r3 vs pin, same packet: +0.12 mean, 95% interval [+0.04, +0.20], 67 wins / 193 ties / 40 losses. An earlier independent packet said +0.07 [+0.01, +0.14]. Both exclude zero. That is progress against us. Against FLAN on this board we are still about 1.4 points back (7.11 vs 5.67). We did not fall off the SO board. We changed boards.
What the titles actually look like
Same 300 terminal-dev tasks. Left column is a shortened version of the user task. Middle is the new incumbent (r3 + centroid v2). Right is the FLAN + centroid v2 bar. These are real outputs, not cherrypicked winners.
| Task | Our incumbent (r3 + centroid v2) | FLAN + centroid v2 |
|---|---|---|
| npm start | npm start | Start Npm |
| Too many unused Supabase tables; check the unused ones and delete them | Check Tables Supabase | Tables Supabase Ones |
| Tell me where we are and what remains for this research paper | Right Me Tell | Research Tell Done |
| Ask GPT via MCP to review the work and find research gaps | mcp Gpt Far | Review Work Far |
| This skill may have lost brand guidelines and logos; check access | Check Access Still | Check Brand Still |
| Where is the [project] intake board? | Project Board [name] | Find Board [name] |
| [name] wants to test the intake form; pull the link | Test Link Form | Intake Form Pull |
| Cheaper alternatives to Google Search with grounding | Google Search Grounding | Cheaper Search Grounding |
| It named the tab "unknown command"; suggest a fix | Fix Instance Useful | Fix Tab Useful |
| Fetch org name and link for the company GitHub | Github Fetch Link | Fetch GitHub Link |
| There is no deadline or rush on this | No There Fantastic | No Deadline Fantastic |
| rm -rf gemm_ | rm rf Gemm | Gemm |
FLAN more often keeps the verb that names the job (Review, Find, Check Brand). Ours still leans on page nouns and leftover glue words (Far, Still, Ones, Fantastic). 37 of 300 displays are identical; the rest differ.
What we actually did this afternoon
The bet was: the pin already knows the page; it just learned the wrong dialect from Stack Overflow titles. So we tried giving it better labels, not a bigger pretrain.
- Round 1. Take the 4,338 existing judged rollouts (8 candidates each), keep the judge's favorite if it scored at least 6, train. Four recipe variants. Result: flat versus the pin.
- Round 2. Same idea with 24 candidates per task (fresh mining, 104k judged titles). Result: still flat. Picking the luckiest of our own guesses is not a teacher.
- Round 3. Decode FLAN on those 4,338 tasks, keep only titles the 35M can legally emit from on-page words (1,136), keep only those the judge scored at least 6 (971), train. Result: the confirmed bump above.
That last construction is the one that worked because the teacher is consistent. A noisy argmax over eight near-tied self-guesses is not.
A scoring bug that almost hid the result
The first evals said every arm lost to the pin by about 0.3 points. That was not training. The local decoder was pointed at a centroid file that does not cover the 300 terminal-dev ids, so every local title dropped its third glue word while the saved pin file still had it. After switching to the session 2.5 terminal centroid (the same centroid v2 table used above), the untrained pin matches the saved pin file (287 of 300 titles identical, mean delta −0.004). Every number on this page uses that corrected glue.
SO-board continuity: the old 19/40 usefulness audit
The four-beams paper's final gate failed on absolute usefulness: an identity-blind Codex review scored the pin useful on 19/40 rows (47.5%), short of the 24/40 bar, even though it beat FLAN 18/9/13 on that same packet. That 40-row sample sits on the consumed sealed SO 1,000. It was not regenerated. It will not be regenerated.
The reusable twin of that protocol is the later 200-row clean-dev Codex audit, where the same pin was useful on 80/200 (40.0%) under the human 0 to 4 rubric (useful = 3 or 4). Session 4 decoded r3 on those exact 200 SO-board ids (two-word + Hybrid A + centroid v2) and scored a fresh blinded Flash-Lite packet against the original pin titles and the original FLAN + centroid v2 titles.
| SO-board packet | Judge | Pin | r3 | FLAN + centroid v2 |
|---|---|---|---|---|
| Sealed 40 (consumed; not rerun) | Codex 0–4 | 19/40 useful (47.5%) | not run | 13/40 useful (32.5%) |
| Reusable 200 (human, historical) | Codex 0–4 | 80/200 useful (40.0%) | not run then | 74/200 useful (37.0%) |
| Reusable 200 (this session) | Flash-Lite 1–10 | 5.96 / 90.0% ≥4 | 5.78 / 88.5% ≥4 | 5.28 / 81.0% ≥4 |
| SO clean-dev 1,000 (this session) | Flash-Lite 1–10 | 5.94 / 86.7% ≥4 | 5.88 / 86.9% ≥4 | not in this packet |
On the reusable 200, r3 vs the original pin titles is −0.17 [−0.32, −0.02] (48/75/77). r3 still beats those original FLAN titles (+0.50). Pin re-decoded today matches the original pin titles on 173/200 rows (mean delta −0.01). r3's win is Terminal-board only. On SO-board usefulness, including the lineage of the 19/40 gate, it does not improve the pin. Do not read the Flash-Lite 90% ≥4 cut as a replacement for the human 40% (3-or-4) cut; different instruments, same rows.
Where the 60% / 80% goal actually sits
- On the cheap Flash-Lite "score at least 4" cut, the pin was already at 78% and r3 is at 79%. The 80% target is inside packet noise on this instrument. That cut is not the FLAN race. FLAN is already at 93% on the same cut.
- The 55.9% figure from session 2.5 was a stricter human review ("this clearly identifies the task"). The closest automated stand-in is score ≥6: pin 52%, r3 55%, FLAN 79%. That is the remaining gap.
- Short tasks are still the worst slice (session 2.5: 1 in 10 useful). Ranking-style training does not invent a name when the page barely has words.
What to do next
- Keep r3 as the promote candidate on the surface we measured: two-word beam, Hybrid A, centroid v2. Checkpoint:
output/s4_r3_train_a/teacher_argmax_ge6_lr2e5/final. - Scale the FLAN-teacher labels on a fresh terminal export. 971 rows already moved the pin; more of the same teacher is the bet with evidence. Do not spend another box on best-of-K self-distill.
- Leave terminal-holdout (200) and the sealed SO 1,000 sealed until a ship gate.
Cost
Two RTX 6000 Ada boxes, about 36 minutes each, then deleted ($1.88). Judge spend about $5.61 of a $15 cap. Holdout unused. Both campaign droplets are gone; the dashboard is stopped.
Contract: tab_namer_quality_roadmap/session4_contract.md. Findings: session4_findings.md. Packs: output/s4_r3/pack_teacher. Eval: output/s4_eval_fixed/.