Session 4: we beat our own namer. FLAN is still ahead on terminal.

16 August 2026 · four-hour Terminal-board campaign · Flash-Lite judge on 300 held-out terminal tasks · Continues the SO-board papers FLAN + centroid, mid GSG, four beams (19/40 lives there and stays sealed)

Frozen names, use these in every later writeup. There are two product boards. They share a judge and a glue. They do not share a leaderboard.

Never write "we almost beat FLAN" without naming the board. Never average, rank, or promote across boards. A method that wins on SO-board is innocent until measured on Terminal-board (exact-three is the exhibit: +0.39 on SO-board, −0.36 on Terminal-board).

What this session did. Teaching the 35M namer to copy FLAN titles that can be spelled from the page nudged our terminal score up a little, and the gain repeated in a second blind packet. Copying our own best-of-8 or best-of-24 guesses did nothing. We are closer to our own pin, not to FLAN on this domain.

SO-board vs Terminal-board

Same judge, same 1 to 10 rubric, same Hybrid A + centroid glue. Different pages. A win on SO-board does not transfer to Terminal-board. That is why "we were almost there" and "we are 1.4 points back" are both true.

Board (use this name) Pages FLAN + centroid v2 Our pin (B-9500 + glue) Gap
SO-board (near-miss, not the ship race) 1,000 SO clean-dev 6.11 / 88.5% useful 5.96 / 85.0% useful −0.15
Terminal-board (ship race) 300 terminal-dev 7.11 / 93.3% useful 5.94 / 83.3% useful −1.17

SO-board source: 6-way Flash-Lite after centroid v2 (docs/gsg_mid_campaign_20260814.md). Terminal-board source: session 2.5 blinded packet. Session 4 trained and judged only on Terminal-board. SO-board appears later only as a regression check, not as a promotion surface.

Who is who

Three systems matter on this page. All three write a 1 to 3 word tab name, then fill leftover slots with centroid v2 (the current glue: trained B-9500 embeddings, SO IDF, residual, MMR).

Today's result, still on the terminal board

Same 300 terminal-dev tasks, same judge, same 1 to 10 rubric. A title is useful if it scores 4 or higher. The FLAN row is copied from the session 2.5 packet above so you can see the bar; the pin and r3 rows are from this session's confirmation packet. Packet-to-packet noise on the same system is a few points. Read the FLAN column as the unfinished race. Read pin vs r3 as a small confirmed bump on our own namer.

System Mean Useful (≥4) Solid (≥6) Strong (≥7) Fails (≤2)
FLAN + centroid v2 (bar) 7.11 93.3% 79.0% 67.3% 2.7%
r3 (this session) 5.67 79.0% 55.3% 35.3% 8.0%
Pin (incumbent) 5.56 78.0% 52.0% 33.0% 8.0%

r3 vs pin, same packet: +0.12 mean, 95% interval [+0.04, +0.20], 67 wins / 193 ties / 40 losses. An earlier independent packet said +0.07 [+0.01, +0.14]. Both exclude zero. That is progress against us. Against FLAN on this board we are still about 1.4 points back (7.11 vs 5.67). We did not fall off the SO board. We changed boards.

What the titles actually look like

Same 300 terminal-dev tasks. Left column is a shortened version of the user task. Middle is the new incumbent (r3 + centroid v2). Right is the FLAN + centroid v2 bar. These are real outputs, not cherrypicked winners.

Task Our incumbent (r3 + centroid v2) FLAN + centroid v2
npm start npm start Start Npm
Too many unused Supabase tables; check the unused ones and delete them Check Tables Supabase Tables Supabase Ones
Tell me where we are and what remains for this research paper Right Me Tell Research Tell Done
Ask GPT via MCP to review the work and find research gaps mcp Gpt Far Review Work Far
This skill may have lost brand guidelines and logos; check access Check Access Still Check Brand Still
Where is the [project] intake board? Project Board [name] Find Board [name]
[name] wants to test the intake form; pull the link Test Link Form Intake Form Pull
Cheaper alternatives to Google Search with grounding Google Search Grounding Cheaper Search Grounding
It named the tab "unknown command"; suggest a fix Fix Instance Useful Fix Tab Useful
Fetch org name and link for the company GitHub Github Fetch Link Fetch GitHub Link
There is no deadline or rush on this No There Fantastic No Deadline Fantastic
rm -rf gemm_ rm rf Gemm Gemm

FLAN more often keeps the verb that names the job (Review, Find, Check Brand). Ours still leans on page nouns and leftover glue words (Far, Still, Ones, Fantastic). 37 of 300 displays are identical; the rest differ.

What we actually did this afternoon

The bet was: the pin already knows the page; it just learned the wrong dialect from Stack Overflow titles. So we tried giving it better labels, not a bigger pretrain.

That last construction is the one that worked because the teacher is consistent. A noisy argmax over eight near-tied self-guesses is not.

A scoring bug that almost hid the result

The first evals said every arm lost to the pin by about 0.3 points. That was not training. The local decoder was pointed at a centroid file that does not cover the 300 terminal-dev ids, so every local title dropped its third glue word while the saved pin file still had it. After switching to the session 2.5 terminal centroid (the same centroid v2 table used above), the untrained pin matches the saved pin file (287 of 300 titles identical, mean delta −0.004). Every number on this page uses that corrected glue.

SO-board continuity: the old 19/40 usefulness audit

The four-beams paper's final gate failed on absolute usefulness: an identity-blind Codex review scored the pin useful on 19/40 rows (47.5%), short of the 24/40 bar, even though it beat FLAN 18/9/13 on that same packet. That 40-row sample sits on the consumed sealed SO 1,000. It was not regenerated. It will not be regenerated.

The reusable twin of that protocol is the later 200-row clean-dev Codex audit, where the same pin was useful on 80/200 (40.0%) under the human 0 to 4 rubric (useful = 3 or 4). Session 4 decoded r3 on those exact 200 SO-board ids (two-word + Hybrid A + centroid v2) and scored a fresh blinded Flash-Lite packet against the original pin titles and the original FLAN + centroid v2 titles.

SO-board packet Judge Pin r3 FLAN + centroid v2
Sealed 40 (consumed; not rerun) Codex 0–4 19/40 useful (47.5%) not run 13/40 useful (32.5%)
Reusable 200 (human, historical) Codex 0–4 80/200 useful (40.0%) not run then 74/200 useful (37.0%)
Reusable 200 (this session) Flash-Lite 1–10 5.96 / 90.0% ≥4 5.78 / 88.5% ≥4 5.28 / 81.0% ≥4
SO clean-dev 1,000 (this session) Flash-Lite 1–10 5.94 / 86.7% ≥4 5.88 / 86.9% ≥4 not in this packet

On the reusable 200, r3 vs the original pin titles is −0.17 [−0.32, −0.02] (48/75/77). r3 still beats those original FLAN titles (+0.50). Pin re-decoded today matches the original pin titles on 173/200 rows (mean delta −0.01). r3's win is Terminal-board only. On SO-board usefulness, including the lineage of the 19/40 gate, it does not improve the pin. Do not read the Flash-Lite 90% ≥4 cut as a replacement for the human 40% (3-or-4) cut; different instruments, same rows.

Where the 60% / 80% goal actually sits

What to do next

Cost

Two RTX 6000 Ada boxes, about 36 minutes each, then deleted ($1.88). Judge spend about $5.61 of a $15 cap. Holdout unused. Both campaign droplets are gone; the dashboard is stopped.

Contract: tab_namer_quality_roadmap/session4_contract.md. Findings: session4_findings.md. Packs: output/s4_r3/pack_teacher. Eval: output/s4_eval_fixed/.