SHLEX · CONTINUATION EXPERIMENT · 10 SEPTEMBER 2026 UTC

240/240 on new requests.
Older skills still regressed.

A full-replay continuation substantially improves six synthetic live-request families. It does not pass the predeclared release gate: the installed retention model remains the incumbent.

Decision: retain the candidate for diagnosis; do not promote.

New transfer outcomes rise from 49/240 to 240/240, but older-board outcomes fall from 145/160 to 139/160. Fifteen older cases lose a previously correct outcome or plan. Perfect synthetic transfer does not establish general user reliability.

New transfer outcomes49 → 240 / 240Same-run incumbent → candidate
Older-board outcomes145 → 139 / 160Six net correct outcomes lost
Training data19,408 rowsFull replay + 3,175 reviewed additions

Against which incumbent?

The baseline is the installed retention-repair checkpoint, not the older filename-only r1 model. Both models generated the same frozen 535 requests on the same RTX 6000 Ada, using bfloat16, greedy decoding, the production document context and the same inference code. Recovered predictions were replayed through the actual macOS search broker with gold-validated fixtures.

Paired evaluation: incumbent versus candidate
MeasureIncumbentCandidateInterpretation
New transfer outcomes49/240 (20.4%)240/240 (100%)+191 correct
New transfer plans40/240240/240+200 correct
Older-board outcomes145/160 (90.6%)139/160 (86.9%)12 regressions, 6 improvements
Older-board plans / controls129/160125/16015 regressions, 11 improvements
Historical filename validation, literal only125/135126/135Not an execution benchmark
All generated outputs, literal only273/535482/535Text identity; not file-search success

The older board contains fresh Boolean cases and historical retention controls. Its plan denominator here includes unsupported controls. The previous published 146/160 outcomes and 131/150 plans used a different inference/evaluation setup; do not subtract those figures from this paired run. The baseline’s same-run 145/160 is the relevant comparison.

Where the new coverage works

Transfer outcomes; 40 cases per family
FamilyIncumbentCandidate
Dotenv file location0/4040/40
Credential-file location0/4040/40
Creation/modification windows25/4040/40
Multiple document formats4/4040/40
Boolean composition8/4040/40
Unsupported-action / ambiguity controls12/4040/40

The candidate also gets 40/40 plans in each family. Supported cases require complete, exact file results. Unsupported controls are scored by the expected refusal or clarification response and are not executed. No generated filename-time correction is applied in this replay.

Where retention fails

Outcome regressions comprise five older OR cases, four AND cases, one scope/format case and two capability controls. Six Boolean cases improve, leaving a net loss of six. Three compound cases additionally lose exact-plan credit while retaining their outcome. The union is 15 regressed cases. This is evidence of lost older behavior despite preserving the full replay mixture; it does not isolate the causal contribution of learning rate, sampling or language distribution.

Data, preflight and recipe

The run preserves all 16,233 historical training rows and 1,491 validation rows. It adds 3,175 new training and 555 validation rows: totals of 19,408 and 2,046. Three unfinished language-review batches were excluded at the operator’s request. The frozen 240-case transfer denominator was retained in full. Cross-split phrase grouping quarantined 187 rows; exact deduplication removed 52. No conflicting query targets or tokenizer truncation remained.

Grok Build authored the synthetic language through the installed PorkiMCP provider. The same teacher performed semantic review: this is not independent human evaluation. Split-specific literals and normalized phrase-group filtering reduce obvious overlap, but they do not prove semantic independence. Templates, intent generators and synthetic fixtures are shared methodological constraints.

Before allocation, the real macOS gold gate passed 2,650 unique plans, 8,973 mutation checks and 2,648 empty-scope checks across the frozen intent specifications. Fixtures distinguish creation from modification, include outside-scope decoys, missing formats, Boolean alternatives and credential placeholders. These are generator/preflight checks, not additional held-out model scores.

One response-only full-model epoch used LR 1e-5, batch 4 × gradient accumulation 4, cosine schedule, 10 warmup steps, weight decay 0.1, seed 42, max length 1,792, bfloat16 and SDPA without cuDNN. All 1,213 expected optimizer steps completed in 1,466 seconds (24.4 minutes). There was one training arm; no post-hoc model selection among multiple runs.

Cost, speed and recovery

The Toronto RTX 6000 Ada was allocated at $1.57/hour. Baseline and candidate generation each took approximately 347 seconds for 535 requests; this provides no evidence of an inference speed improvement and is not an app latency benchmark. Total controller-estimated GPU cost, including setup, generation and recovery, was about $1.20; this is an estimate, not an invoice.

Both saved weight files were recovered to private Mini storage and checked at 3,087,467,144 bytes with identical SHA-256. Only then was the GPU deleted. A later account check confirmed zero GPU droplets. The candidate has not been installed in Shlex.

Reproducibility and next experiment

Candidate SHA-256: 1cdbff7ac98dfbc4abeb00e2f0a7b2eea124674d4eebc9fe5bee0357ea8b4502
Incumbent SHA-256: 984ff1aaed71652d9f53c91074e0787b3386bc3ae41900b4b830e6211805fc64
Frozen generation board SHA-256: c6605077bb26c30e3e16c5e964e73d5d9bc4771ea5f90d51f7640e52b17793f0

Next, diagnose the older OR/AND, scope/format, compound and capability failures without changing this frozen board. Prepare separate training examples and independent held-out language for those behaviors. A lower-exposure or replay-reweighted continuation is a hypothesis to test, not an established fix. Keep the no-regression gate and the 38/40-per-family threshold unchanged; do not ship on aggregate gains alone.

The full original filename execution benchmark, independent human requests and app launch/permission/latency checks remain outstanding. Signing and notarization are separate release requirements. Public artifacts contain only aggregate findings and hashes; training rows, private traces, fixtures and weights remain private.