Session 11: more epochs on the Aug-12 mix is the new ship
17 August 2026 · Terminal-board · two never-seen 1000-row holdouts · Gemini 3.5 Flash-Lite · Continues session 10, session 9, techniques log
New ship: Title-SFT FLAN raw, arm flat1e4 — continue the Session-10 ship on the leak-checked Aug-12 gold + Grok / Opus / Codex mix, 4 epochs, lr 1e-4, seed 1, flat weights. Same 77M untied FLAN graph. No Hybrid A. Holdout 1000: 7.59 / 85.7% versus the frozen Session-10 ship 7.38 / 83.5% (+0.21). Holdout 1000b: 7.55 / 84.5% versus 7.34 / 82.5% (+0.20). Package: porkr1/sft_flan_raw_ship_package/. The unconstrained FLAN worker is wired in porkr1 at resources/tab-namer/flat1e4-sft-v1/. Public 2.17.42 still serves 6t until the next notarized installer.
Two boards. Do not mix packets. This page is Terminal-board only. Select on holdout1000 and holdout1000b, not reusable-dev 500. Terminal-holdout 200 and sealed SO 1000 stay closed. Do not average with Session-10’s 7.28 / 7.95 ship numbers — those were different draws.
What we tried today
Three separate experiments. Only the third is the ship.
| Experiment | Init | Targets | Result |
|---|---|---|---|
| 35M CE distill | 6t_argmax |
Session-9 train decoded by Session-10 SFT (2-word schema dialect) | Learned English. Beat 6t-raw by +3.91 on reusable-dev. Still −1.82 vs teacher. Shelved. |
| Judged-schema continue-SFT | Session-10 ship | Session-9 Flash-Lite argmax, 2-word schema | Lost to ship on both holdouts (e.g. 6.70 vs 7.37). Wrong dialect. Do not repeat. |
| Aug-12 continue-SFT | Session-10 ship | Same free-form mix the ship was trained on (gold + Grok / Opus / Codex) | All four arms beat ship on both holdouts. flat1e4 is the new ship. |
Aug-12 continue-SFT vs frozen Session-10 ship
Raw greedy, max_new_tokens=12, no Hybrid A. Flash-Lite seeds session10-aug12-holdout1000-v1 and session10-aug12-holdout1000b-v1. 1000 / 1000 scored, zero failures.
Holdout 1000
| Arm | Recipe | Mean | ≥6 | ≥7 | Δ vs ship | win / tie / lose | Same title |
|---|---|---|---|---|---|---|---|
| flat1e4 (new ship) | flat weights, 4 ep, 1e-4, seed 1 | 7.59 | 85.7% | 80.1% | +0.21 | 296 / 551 / 153 | 482 |
| rewt5e5 | reweighted, 5 ep, 5e-5, seed 1 | 7.60 | 86.0% | 80.3% | +0.22 | 267 / 603 / 130 | 537 |
| flat5e5 | flat weights, 5 ep, 5e-5, seed 1 | 7.59 | 85.9% | 79.8% | +0.20 | 262 / 607 / 131 | 537 |
| flat5e5s2 | flat weights, 5 ep, 5e-5, seed 2 | 7.51 | 84.3% | 78.0% | +0.12 | 248 / 610 / 142 | 538 |
| Session-10 ship | frozen b798d78f… |
7.38 | 83.5% | 74.6% | — | — | — |
rewt5e5 is +0.01 mean on this draw. The owner picked flat1e4: it also wins on holdout 1000b, changes more titles (482 same vs 537), and is the simpler recipe (flat weights, 4 epochs).
Holdout 1000b
| Arm | Mean | ≥6 | ≥7 | Δ vs ship | win / tie / lose | Same title |
|---|---|---|---|---|---|---|
| flat1e4 (new ship) | 7.55 | 84.5% | 78.4% | +0.20 | 281 / 570 / 149 | 518 |
| flat5e5 | 7.52 | 84.1% | 78.1% | +0.17 | 246 / 632 / 122 | 581 |
| rewt5e5 | 7.51 | 83.6% | 77.7% | +0.17 | 247 / 627 / 126 | 575 |
| flat5e5s2 | 7.47 | 83.3% | 77.2% | +0.13 | 249 / 615 / 136 | 562 |
| Session-10 ship | 7.34 | 82.5% | 74.4% | — | — | — |
About half the titles are unchanged. The lift is the other half: mean words 2.74 vs ship 2.59 (more 3-word titles). Typical win: Fix Fix Bug → Fix SSE Disconnect, Apply Deadline → Apply Target File. Typical loss: Detect Code Blocks → Remove Banner Banner.
Accuracy, usefulness, disk, inference
Same four product numbers as session 10. Glue is a 0.02 ms string pass. Disk and clock are the decode graph. The app already carries the 77 MB 6t INT8 bundle. Clocks below are the Session-10 measurements: onnxruntime-node 1.24.3, Apple M4 Max, CPU, 1 thread, 40 rows, 5 warmup discarded. flat1e4 did not change the graph — 77M FLAN-T5-small, 307,867,048-byte fp32, same tokenizer — so payload and clock are the Session-10 SFT numbers, not a new bench.
| Arm | Mean | ≥6 | Electron payload | Added to today’s app | Infer / title | Knowledge / teacher | Clock status |
|---|---|---|---|---|---|---|---|
| Title-SFT FLAN raw · Session 11 flat1e4 | 7.59 | 85.7% | ~193 MB INT8 (est.) / 294 MB this fp32 | +193 MB | 29.3 ms | google/flan-t5-small → Session-10 Aug-12 SFT → 4 more epochs on the leak-checked Aug-12 mix, lr 1e-4, flat weights. Untied. No Hybrid A. | Same SFT graph as Session 10. Clock not re-run. |
| Title-SFT FLAN raw · Session 10 (superseded) | 7.38 | 83.5% | ~193 MB INT8 (est.) | +193 MB | 29.3 ms | Same 77M. First continue of FLAN on the 7,623-row Aug-12 mix. SHA b798d78f…. |
onnxruntime-node, measured Session 10 |
| 6t_argmax + Hybrid A | 6.65 | 75.6% | 77 MB INT8 (measured) | 0 (already in the app) | 10.1 ms | From-scratch 35M. Technical pretrain + GSG + session-5 judged 2-word CE. Not FLAN. | App worker, measured. Quality is Session-10 ship-board 500, not this holdout. |
| Dense P1 + Hybrid A | 6.79 | 78.2% | 48 MB if it replaces 6t; 85 MB with 6t | +8 MB ONNX head | 3.88 ms | Frozen 6t encoder + session-9 Flash-Lite schema pointer. | onnxruntime-node, measured. Quality is Session-10 ship-board 500. |
| CPU picker over P1 vs 6t | 6.81 | 80.0% | 85 MB + 378 B | +8 MB + 378 B | 11.72 ms | Logistic chooser. Writes no titles. | Generate 11.7 + pick 0.02, measured |
| Neural ranker s101 over P1 vs 6t | 6.83 | 80.6% | 87 MB (77 + 8 + 1.6) | +9.6 MB | 17.5 ms | Listwise MLP chooser. Writes no titles. | Generate 11.7 + 2 encodes + MLP 5.82 |
| Stock FLAN raw | 3.12 | 15.8% | ~249 MB INT8 (est.) | +249 MB | 25.7 ms | Google FLAN mix. Never saw our tabs. | onnxruntime-node, measured Session 10 |
flat1e4 / Session-10 ship means and ≥6 on the first two rows are holdout 1000 (this session). 6t / P1 / picker / stock means are the Session-10 ship-board 500 packet and must not be ranked against 7.59. Payload and infer/title are unchanged from Session 10 because the architecture did not change. The 6.65 quality number and the 10.1 ms clock are still not the same artifact (fp32 PyTorch vs INT8 ONNX worker).
How each column was calculated
- Mean / ≥6 for the two SFT rows. Blinded Gemini 3.5 Flash-Lite, seed
session10-aug12-holdout1000-v1, n=1000, rubric 1.00–10.00 two decimals. Identical strings on a row get the same score. - Mean / ≥6 for 6t, P1, pickers, stock. Copied from Session 10 seed
session10-ship-board-v1, n=500 reusable-dev. Context only. Not a promote board for this session. - Electron payload / added / infer. Copied from Session 10. 6t is the live 77.2 MB INT8 bundle at 10.1 ms. SFT FLAN is ~193 MB INT8 estimate at 29.3 ms. P1 +8 MB at 3.88 ms. No new ONNX export this session.
How the Aug-12 mix was made — copy this for Session 12
This is the recipe that produced the 7,623-row file the ship learned from. Session 11 only leak-checked it and trained more. The next training move is more data built the same way, not another lr sweep on these 6.4k rows, and not session-9 2-word schema argmax.
Title contract (every teacher, every gold row)
- 1 to 3 words. Max 32 characters.
[A-Za-z0-9 ]only. Title Case. - Prefer two words. Prefer content words that appear in the task. Name the user intent, not agent chrome.
- Do not spam
Code Review/Check Status/Fix Bugunless that is the task. - Input string is always synthesized:
title: agent=<a> [cwd=<c>] task=<z>. First paragraph, then 400 characters. Never train on the raw logged prompt.
Pipeline (exact scripts and knobs)
- Export chats.
./data/export_supabase.sh→data/raw/chat_interactions.latest.jsonl. For Session 12 also remine local Claude / Codex / Grok / OpenCode sessions the wayscripts/s10_build_holdout1000.pydoes, then drop anything already in train or a frozen eval. - Gold. Short production replies, mostly
kimi-fast, max 40 response characters, grammar-gated:.venv/bin/python data/clean_prod_titles.py \ --input data/raw/chat_interactions.latest.jsonl \ --prefer-model-key kimi-fast \ --max-response-chars 40 \ --task-mode paragraph
Historical keep: 681 gold titles (680 kimi-fast + 1 other). - Unlabeled tasks. Same bound, no title from
response:.venv/bin/python data/extract_tasks.py --task-mode paragraph
- Freeze shard inputs. Historical set is 12 shards, ~387 rows each,
data/derived/label_shards/shard_*_in.jsonl. Every teacher must see those exact task strings. Do not re-parse. For new data, cut new shards the same way and freeze them before any teacher runs. - Three teachers, one protocol. Instructions:
data/derived/label_shards/LABEL_INSTRUCTIONS.md(same text indata/run_cli_teacher_label.py). Teachers:- Grok Build · Grok 4.5 (high)
- Claude Code · Claude Opus 5 (high) —
data/run_cli_teacher_label.py --teacher claude - Codex · OpenAI high —
data/run_cli_teacher_label.py --teacher codex
- Merge each teacher. Grammar + token overlap τ = 0.5. Task text from the frozen shard inputs, never from a fresh extract:
scripts/run_multiteacher_after_labels.sh # which is: # merge_teacher_shards.py --min-overlap 0.5 --weight 0.5 # grok → 3184 kept / 1462 reject # opus → 3366 kept / 1280 reject # codex → 2974 kept / 1672 reject
- Union. Distinct
(task, title)pairs. Same task, different teacher titles → keep all, weight 0.5. Cap a title at 8 uses:.venv/bin/python data/build_multiteacher_pseudo.py \ --inputs data/derived/pseudo_labels_grok.jsonl \ data/derived/pseudo_labels_claude.jsonl \ data/derived/pseudo_labels_codex.jsonl \ --max-per-title 8 --weight 0.5Historical union: 6,116 rows. - Mix with gold. Dedupe
(task, title)preferring gold. Gold train ×3. Title cap 6 on the train pool:.venv/bin/python data/build_student_dataset.py \ --gold data/derived/prod_titles.jsonl \ --pseudo data/derived/pseudo_labels_multiteacher.jsonl \ --gold-dup 3 --max-per-title 6
Historical file: 7,623 train rows = 510 gold × 3 + 6,093 pseudo. Gold-only val 100 / test 62. Breakdown of the 7,623: gold 1,530 · Grok 2,854 · Opus 1,868 · Codex 1,371. - Reweight (optional recipe, used by rewt5e5, not by flat1e4).
.venv/bin/python data/reweight_by_title_freq.py \ --gold-weight 2.0 --pseudo-base 0.5 --floor 0.1
Gold weight 2.0. Pseudo0.5 / sqrt(title_freq)clipped to [0.1, 0.5]. - Leak-check before any new train (Session 11 addition, now mandatory). Drop task-hash overlap with terminal-dev 300, holdout 200, holdout 1000, holdout 1000b, reusable-dev 500, confirm-v1, sealed-v2. Ignore numeric-id collisions across supabase dumps. Session 11 keep: 6,419 (372 gold + 6,047 teacher) → 6,219 train / 200 val. Hash leaks dropped 1,148.
.venv/bin/python scripts/s10_build_aug12_flan.py \ --train data/derived/student_train_reweighted.jsonl \ --out-dir output/s10/flan_sft/aug12_flat --flat-weights
- Pack once, then train cache-only. Google FLAN tokenizer. No compact vocab. Missing
DONEis a hard fail..venv/bin/python scripts/pack_flan_title_jsonl.py \ --train output/s10/flan_sft/aug12_flat/train.jsonl \ --val output/s10/flan_sft/aug12_flat/val.jsonl \ --output-dir output/s10/flan_sft/pack_aug12_flat
- Continue the current ship, untied. Load
porkr1/sft_flan_raw_ship_package/namerwithtie_word_embeddings=false, share encoder/decoder embeds, refuse iflm_headis tied. Expected 76,961,152 params. Winning Session-11 recipe: batch 16, 4 epochs, lr 1e-4, seed 1, flat weights, 1,556 steps, first-batch CE ~0.13 (the ship already knows this mix).
Session 12 must add rows, not epochs. First-batch CE was already 0.13. Another lr/seed sweep on the same 6.2k will not move the product. Export newer chats, freeze new shards, run the same three teachers with the same instructions, union + gold-dup 3 + title-cap 6, leak-check the frozen evals, pack once, continue from flat1e4 (the package now in porkr1). Do not distill 35M. Do not imitate session-9 schema argmax.
Frozen names
- New ship:
porkr1/sft_flan_raw_ship_package/· armflat1e4· weight SHA0dbd3a3385ce252594fdfd0acd83c2917506a8004c97b74c31095f082c934598· 76,961,152 params · load untied or decode isreheat/blackjack. - Superseded Session-10 ship: SHA
b798d78f3ffcf83510b4e0af55388c75f4232fefbe1a3db17191cfe15b1f902d. Keep as provenance. Do not reload as the product init. - Live app: unconstrained FLAN worker at
porkr1/resources/tab-namer/flat1e4-sft-v1/(Gitea). Public 2.17.42 still serves 6t until the next notarized installer. - Holdout 1000: seed
session10-holdout1000-20260817, sha16c21871…. Holdout 1000b: seedsession10-holdout1000b-20260817, sha64935c8b…. - Holdout 200 and sealed SO 1000 stay closed. Reusable-dev 500 is not the selector.