Session 10: every namer we have, scored against the clock

17 August 2026 · Terminal-board · reusable-dev 500, one blinded packet · Gemini 3.5 Flash-Lite · Continues session 9, techniques log

New ship candidate: Title-SFT FLAN raw (student-flan-t5-run5-capacity/best, no Hybrid A). Confirm-v1 335 unseen local tasks: 6.55 / 71.3% versus P1 6.41 / 64.5% and 6t 6.24 / 60.3%. Hybrid A hurts this SFT on that set (5.78). Human labels agree: SFT raw names the job; P1 and Hybrid A often scramble it. +193 MB and 29.3 ms are accepted. 6t stays in the app until the SFT worker is wired. Honorable mentions: P1, 6t, the two pickers.

6t and Dense P1 are not the same system. 6t_argmax is the shipped 35M encoder–decoder, constrained two-word beam, then Hybrid A. Dense P1 freezes that encoder, throws away the decoder, and adds an 8 MB pointer head. They share an encoder and disagree on 136–194 rows. Never average them. Never write “6t (including P1).”

Two boards. This page is Terminal-board only. Sealed SO 1000 stays closed. Terminal-holdout 200 stays closed. Sealed-v2 516 is diagnosis only; the configuration table below is reusable-dev 500.

The ship board — one packet, twelve arms

Seed session10-ship-board-v1. Knowledge is the training corpus and teacher, not the decode rule. Hybrid A / keep-namer / typed B add no weights. Pickers do not emit titles; they choose among titles another model already wrote.

Arm Family Knowledge / teacher Decode Glue Mean ≥6 ≤2 Unique Δ vs SFT+HA Δ vs 6t HA In the app
Title-SFT FLAN + Hybrid A Our 77M SFT google/flan-t5-small, then session-4 title SFT on terminal-pilot teacher titles. Distilled from our teachers, not from 6t. free 8 tok Hybrid A 7.41 89.2% 2 493 +0.77 No
Title-SFT FLAN + keep-namer Our 77M SFT Same SFT checkpoint. Glue only. free 8 tok keep-namer 7.35 86.8% 0 494 −0.06 +0.70 No
Title-SFT FLAN raw Our 77M SFT Same SFT checkpoint. No glue. free 8 tok none 7.28 87.8% 0 472 −0.13 +0.64 No
P1 ∪ 6t oracle Pick better display No weights. Flash-Lite already scored both titles; take the max. Reference only. both Hybrid A 7.04 84.0% 1 −0.37 +0.40 No — judge at tab-open
Neural ranker s101 over P1 vs 6t 6t encoder + 1.6 MB MLP Listwise CE on session-9 train_cands: teacher / 6t / centroid Flash-Lite lists. Not P1-HA vs 6t-HA. pick 2 titles Hybrid A 6.83 80.6% 1 489 −0.58 +0.18 ONNX MLP ready
CPU picker over P1 vs 6t 8-feature logistic Pairwise logistic on Flash-Lite scores of P1 top-3 HA + 6t HA, session-9 train 5,938. No namer weights. pick 2 titles Hybrid A 6.81 80.0% 1 489 −0.60 +0.16 378 B JSON ready
Dense P1 + Hybrid A 35M encoder + 2M head Frozen 6t encoder. Pointer head on 5,938 Flash-Lite schema lists (session 9 dense pack). Inherits 6t, does not see FLAN. schema pointer Hybrid A 6.79 78.2% 1 489 −0.63 +0.14 ONNX ready
6t_argmax + Hybrid A Shipped 35M From-scratch 35M. Technical pretrain + GSG, then session-5 judged 2-word CE on 4,338 terminal-pilot teacher titles. Not distilled from FLAN. constrained 2-word Hybrid A 6.65 75.6% 2 489 −0.77 Yes, ONNX
Dense P1 + keep-namer 35M encoder + 2M head Same P1 head. Glue only. schema pointer keep-namer 6.56 73.0% 0 483 −0.85 −0.09 No
agent1m + Hybrid A 6t + 1M CE 6t continued on Gemini Flash-Lite title CE (~300k–1M agent pages). Distilled from Gemini labels, not FLAN. constrained 2-word Hybrid A 6.50 69.6% 1 488 −0.91 −0.15 No
6t free + Hybrid A Shipped 35M Same 6t weights. Unconstrained decode only. unconstrained Hybrid A 6.39 68.6% 4 497 −1.02 −0.26 No
agent1m free + Hybrid A 6t + 1M CE Same agent1m weights. Unconstrained decode only. unconstrained Hybrid A 6.35 67.4% 5 493 −1.06 −0.30 No
Dense P1 typed B 35M encoder + 2M head Same P1 head. Slot rules, no new training. schema pointer typed slots 6.25 67.6% 2 447 −1.16 −0.40 No
agent1m free + keep-namer 6t + 1M CE Same agent1m weights. Glue only. unconstrained keep-namer 5.00 34.0% 27 495 −2.41 −1.64 No
Stock FLAN raw google/flan-t5-small Google FLAN instruction mix. Never saw PorkiCoder tabs. free 8 tok none 3.12 15.8% 217 373 −4.29 −3.53 No

Source: output/s10/ship_board/packet.raw.scored.jsonl. n=500. Judge gemini-3.5-flash-lite. $0.22. Do not mix these means with Session 9’s sealed 7.02 / 6.43 / 6.30 — different packet. Within this packet the ranking is the ship ranking.

ContrastΔ95% CIw / t / l
SFT raw − stock raw +4.17 [3.95, 4.38] 465 / 9 / 26
SFT + Hybrid A − SFT raw +0.13 [0.03, 0.22] 252 / 74 / 174
SFT keep-namer − SFT + Hybrid A −0.06 [−0.13, −0.00] 55 / 364 / 81
Dense P1 HA − 6t HA +0.14 [0.05, 0.23] 194 / 170 / 136
P1 ∪ 6t − P1 HA +0.26 [0.21, 0.31] 136 / 364 / 0
P1 ∪ 6t − SFT raw −0.24 [−0.37, −0.12]
SFT raw − P1 HA +0.50 [0.36, 0.64] 303 / 34 / 163
SFT raw − 6t HA +0.64 [0.50, 0.78] 316 / 29 / 155
P1 keep − P1 HA −0.23 [−0.30, −0.15] 90 / 224 / 186
P1 typed B − P1 HA −0.54 [−0.63, −0.45] 75 / 82 / 343
agent1m HA − 6t HA −0.15 [−0.23, −0.07] 112 / 217 / 171
6t free HA − 6t HA −0.26 [−0.38, −0.13] 164 / 109 / 227
agent1m free keep − agent1m free HA −1.34 [−1.49, −1.20] 29 / 158 / 313

Accuracy, usefulness, disk, inference

Four product numbers plus the teacher. Glue is a 0.02 ms string pass. Disk and clock are the decode graph plus any picker. The app already carries the 77 MB 6t INT8 bundle. Pick clocks are onnxruntime-node 1.24.3, CPU, 1 thread, n=35 after 5 warmup, M4 Max.

Arm Mean ≥6 Electron payload Added to today’s app Infer / title Knowledge / teacher Clock status
6t_argmax + Hybrid A 6.65 75.6% 77 MB INT8 (measured) 0 (already in the app) 10.1 ms From-scratch 35M. Technical pretrain + GSG + session-5 judged 2-word CE. Not FLAN. App worker, measured
Dense P1 + Hybrid A 6.79 78.2% 48 MB if it replaces 6t; 85 MB with 6t +8 MB ONNX head 3.88 ms Frozen 6t encoder + session-9 Flash-Lite schema pointer (5,938 lists). onnxruntime-node, measured
CPU picker over P1 vs 6t 6.81 80.0% 85 MB + 378 B +8 MB + 378 B 11.72 ms Logistic on Flash-Lite P1-top3 vs 6t (train 5,938). Chooses titles; writes none. Generate 11.7 + pick 0.02, measured
Neural ranker s101 over P1 vs 6t 6.83 80.6% 87 MB (77 + 8 + 1.6) +9.6 MB 17.5 ms Listwise on teacher/6t/centroid lists — not production P1-HA vs 6t-HA. Generate 11.7 + 2 encodes + MLP 5.82, measured
P1 ∪ 6t oracle 7.04 84.0% No model. Max of two already-judged displays. Reference only
6t free + Hybrid A 6.39 68.6% 77 MB (same graphs) 0 Same 6t weights. No unconstrained worker
agent1m + Hybrid A 6.50 69.6% ≈ 77 MB 0 if it replaces 6t 6t + Gemini Flash-Lite title CE (~300k–1M). Not exported to a worker
Title-SFT FLAN raw 7.28 87.8% ~193 MB INT8 (est.) +193 MB 29.3 ms FLAN-T5-small + session-4 terminal title SFT. onnxruntime-node, measured
Title-SFT FLAN + Hybrid A 7.41 89.2% ~210 MB INT8 (est.) +210 MB 29.3 ms Same SFT checkpoint + glue. Same SFT graph + 0.02 ms glue
Title-SFT FLAN + keep-namer 7.35 86.8% ~210 MB +210 MB 29.3 ms Same SFT checkpoint + glue. Same SFT graph + 0.02 ms glue
Stock FLAN raw 3.12 15.8% ~249 MB INT8 (est.) +249 MB 25.7 ms Google FLAN mix. Never saw our tabs. onnxruntime-node, measured

Keep-namer / typed B / free-keep share the parent decode payload. They are omitted here so the clock and the disk stay attached to a graph, not a string rule.

How each column was calculated

The 6.65 quality number and the 10.1 ms clock are not the same artifact. Mean / ≥6 for 6t in this packet are from the fp32 PyTorch 6t_argmax decode that Session 9 used. The 10.1 ms clock is the INT8 ONNX worker the app actually loads. We have not judged those INT8 titles against fp32 on this packet. Dynamic INT8 can drop usefulness. The next measurement is a same-packet three-way: PyTorch fp32 vs ONNX fp32 vs the shipped INT8 worker, for 6t and for any FLAN we consider shipping. Until that packet exists, “unquantize 6t” is an open quality lever, not a free lunch: fp32 graphs would grow the 58 MB INT8 encode/decode files back toward the 134 MB weight class and may slow the 10 ms clock.

How to read the tradeoff

The FLAN ladder Session 9 hid

Session 9 printed one FLAN number: title-SFT after Hybrid A. That made “FLAN” look like a generic 77M bar. It is not. Stock FLAN, asked to name a terminal tab, emits x_bot, cwd=/Us, and empty strings. Our SFT is the product.

Stock raw (3.12) Our SFT, raw (7.28) Our SFT + Hybrid A (7.41)
x_botCreate Beautiful MusicCreate Beautiful Music
x_botExplain DNSExplain DNS
x_bot: aCloudflare CutoverCloudflare Cutover
cwd=/UsPrint Training ResultsPrint Training Results

Human whiff — closest locals vs the FLAN variants

Same ship-board packet as the means above. Columns are the three FLAN variants plus the three local systems nearest SFT+HA (7.41): neural ranker 6.83, P1 6.79, 6t 6.65. CPU picker (6.81) is omitted — it almost always reprints P1. Grey number is that row’s Flash-Lite score. This is a taste check, not a new metric.

Task Stock FLAN 3.12 SFT raw 7.28 SFT + HA 7.41 Neural pick 6.83 P1 + HA 6.79 6t + HA 6.65
do not be shy, let's work together and make beautiful music x_bot1.00 Create Beautiful Music9.00 Beautiful Music Together9.00 let work Together5.00 make Together Let3.00 let work Together5.00
this is not an easy thing… lots of complicated DNS x_bot1.00 Explain DNS9.00 Explain DNS Easy9.00 Explain 50 Easy3.00 Explain 50 Easy3.00 Doing DNS Easy7.00
estimate wait time until we get next meaningful result? - a t1.10 Wait Time8.80 Wait Time Until9.10 next Wait Until8.10 next Wait Until8.10 Get Time Wait3.10
how is the training run on C doing, print the results cwd=/Us1.10 Print Training Results8.90 Print Training Results8.90 print results Doing6.40 print results Doing6.40 Print Results Doing6.40
ok what plan am i using i need to know what plan7.12 Plan Check7.89 Plan8.92 ok am Plan4.12 ok am Plan4.12 ok am Plan4.12
35 or 70 моете1.00 Fix Tab Namer4.10 (empty)1.00 357.50 357.50 357.50
get opus 5 to code review alongside your own review adolescent1.25 Code Review9.12 Code Review Yeah4.12 Review code Yeah8.78 Review code Yeah8.78 Check Changes Yeah5.12
dont kill gpu; write a compact handoff, fresh context gpu1.20 Write Plan4.20 Write Plan Handoff5.20 handoff Centroid Docu8.20 handoff Centroid Docu8.20 Docu handoff Centroid8.00
rename folder peggy; commit all work to gitea GitHub3.10 Add Git Name4.10 Git Name Gitea5.20 Change commit Gitea6.50 gitea commit Folder7.50 Change commit Gitea6.50
what is a cutover in this context, under 50 words cutover' is a7.50 Explain Cutover9.00 Explain Cutover Relevant8.00 Context words Explain3.00 cutover 50 Explain8.50 Context words Explain3.00
integrating with cloudflare without an account or api key Cloudflare integration9.50 Cloudflare Integration9.50 Cloudflare Simply Words7.50 Api Endpoints Simply7.50 account Simply Words3.00 Api Endpoints Simply7.50
professor jobs are not easy to come by idx - a1.20 Find Prof Jobs8.70 Find Prof Easy8.50 Find window Easy4.20 window Easy Come3.20 Find window Easy4.20

Read left to right on one row. Stock is usually junk. SFT names the job. Locals often keep a real noun (DNS, gitea, 35) but scramble word order or glue in “Yeah / Easy / Doing.” Hybrid A sometimes helps SFT and sometimes wrecks it (Code Review 9.12 → Code Review Yeah 4.12; empty title on “35 or 70”). If your eye prefers SFT raw over SFT+HA on those rows, that matches the sealed-v2 packet where glue hurt SFT.

6t versus Dense P1, again

On this packet P1 beats 6t by +0.14 [0.05, 0.23], the same direction as Session 9 sealed (+0.13). They display the same title on a minority of rows. The union of the two Hybrid A titles is 7.04. That is the argument for keeping both in a future ONNX graph (shared encoder, 6t decoder + P1 head, ~37M). It is not an argument that 6t “is” P1.

P1’s typed-B selector — keep off-page Review/Fix/Create, allow two-word titles, do not force centroid — lost by 0.54. Object-only titles in the earlier diagnostic were 4.65. The third word is still load-bearing for P1. Hybrid A stays the P1 glue until a learned post-finalizer ranker exists.

Sealed-v2 diagnosis (not the ship table)

A separate packet on the 516 sealed-v2 rows scored the FLAN ladder only. New judge draw. Do not compare those P1/6t means to Session 9.

ArmMean≥6
Title-SFT FLAN raw7.2983.1%
Title-SFT FLAN + Hybrid A6.9175.8%
Stock FLAN raw3.2217.2%

SFT raw − stock +4.07. Hybrid A − raw −0.38 on that draw. Glue on SFT is packet-sensitive; SFT versus stock is not.

Session 9’s original sealed-v2 packet (ten Hybrid A / keep-namer arms, no stock, no unglued SFT) remains: SFT+HA 7.02, P1 HA 6.43, 6t HA 6.30, agent1m 6.16. P1 ∪ 6t on those already-judged rows is 6.67, −0.35 versus that packet’s FLAN. Teacher schema oracle on an older reusable-dev packet is still 7.97 versus FLAN 6.45. The schema space can beat SFT FLAN; the current P1 top-1 cannot.

What we are not claiming

Kill permanently

agent1m and more generic title CE. Free decoding. Global keep-namer. Another learning-rate sweep. Forced three-word titles as policy. Opening holdout to inspect a 0.05 lead. Treating stock FLAN as the bar we have to beat — the bar we built is the SFT.

Confirm-v1 — unseen 335, then the ship call

New local tasks, not reusable-dev, not sealed-v2, not Terminal-holdout 200, not sealed SO. 311 grok / 15 claude / 9 codex. Flash-Lite. Title-SFT must load with tie_word_embeddings=False or it collapses to reheat/blackjack.

Arm Role Confirm mean Confirm ≥6 Ship-board mean Ship-board ≥6 Δ mean Δ ≥6 vs today’s app Infer
Title-SFT FLAN raw Ship candidate 6.55 71.3% 7.28 87.8% −0.73 −16.5 +193 MB 29.3 ms
Title-SFT FLAN + keep-namer Honorable mention 5.90 54.3% 7.35 86.8% −1.45 −32.5 +210 MB 29.3 ms
Title-SFT FLAN + Hybrid A Honorable mention 5.78 51.9% 7.41 89.2% −1.63 −37.3 +210 MB 29.3 ms
Dense P1 + Hybrid A Honorable mention 6.41 64.5% 6.79 78.2% −0.38 −13.7 −29 MB 3.88 ms
CPU picker over P1 vs 6t Honorable mention 6.41 64.2% 6.81 80.0% −0.40 −15.8 +8 MB 11.72 ms
Neural ranker s101 Honorable mention 6.40 64.2% 6.83 80.6% −0.43 −16.4 +10 MB 17.5 ms
6t + Hybrid A Honorable mention · still wired 6.24 60.3% 6.65 75.6% −0.41 −15.3 +0 10.1 ms

Confirm-v1 is harder and grok-heavy; every local arm dropped ~0.4. SFT raw dropped less than SFT+HA. Glue is not the ship rule for this checkpoint.

Human labels — SFT raw vs P1 (confirm-v1)

Twenty rows from the same packet. Grey number is that row’s Flash-Lite score.

Task Title-SFT FLAN raw · ship Dense P1 + Hybrid A
yes i meant basetenBaseten Check7.90baseten Yes Meant2.10
oh wow that ain't cheap, eh.Cheaper Sidebar7.50ain Wow Cheap2.06
tab naming is not working, 6t argmax should fireFix Tab Name8.50tab Naming Happening3.25
spin up the dashboard before sending workload to adaSpin Up Dashboard8.50ada Sending Spin3.40
lets do it now then. start the campaign.Start Campaign8.00it start Now2.90
how are the scores looking so farReview Scores8.30soounds Far3.80
how much will that cost and how long will it takeEstimate Cost9.00it Words Much4.50
everybody wins.Everybody Wins2.15Everybody Wins8.95
7250 steps still some time to go lad.Go LD1.50lad Steps Still7.90
early results from the 1200 that are doneStart Date2.10results 1200 Done8.20
ok no longer idle. thanks.Retire Tab Namer2.10ok Longer Idle8.15
would you recommend shipping Dense P1 + Hybrid ARecommend Package Shipping3.00Hybrid Recommend Schema8.90
fail count is lower. if we scale to 1M…Low Fetch Fetch1.501M Scale Might7.30
whats the eta buddyeta buddy1.50eta Whats Buddy7.00
finish @maxpork.md and ship itFinish Maxpork8.50md Maxpork Finish8.50
make a call to kimi k3 asking for an honest evaluationKimi K3 Assessment8.50Run k3 Ada8.50
the raison detre for this is fast inferenceFast Inference8.50inference detre Fast8.50
will you need ada boxes for itAda Boxes8.20ada Need Boxes8.20
Write a terminal tab title for each task below…Terminal Tab Title2.15Write array Terminal2.15
actually write a handoff document called maxpork.mdWrite Handoff Document8.50handoff Maxpork Write7.50

What to ship