# Bash SFT note 02 evidence packet

This is the sanitized public evidence packet for the cumulative Bash Session 2
research note published on 2026-08-21.

The packet was corrected on 2026-08-21 after we determined that the original
strict stock comparison conflated adherence to an explicitly stated raw-command
form with Bash semantics. The prompt gave no format demonstration and did not
disclose the judge's exact command-policy allowlist and restrictions. The
original 0/98 is retained as an audit trail. The fair comparisons are now the
basis for stock-versus-SFT claims.

## Files

- `summary.json` — aggregate pilot, data, arm, baseline, canary, and claim fields
- `arm-comparison.json` — paired component-bootstrap recipe comparisons
- `fair-eval-summary.json` — aggregate protocol and results for the symmetric
  adapter and identical contract-calibrated generation checks
- `fair-eval-comparison.json` — sanitized paired deltas and bootstrap intervals
- `data-manifest.json` — sanitized GFR promotion counts and known limits
- `canary-summary.json` — aggregate post-selection external canary result
- `report.md` — compact companion narrative
- `SHA256SUMS` — hashes for every packet file except itself

The packet intentionally excludes private traces, prompts, model completions,
fixtures, local or remote paths, machine addresses and identifiers,
credentials, and model weights.
The full experiment ran from a dirty local worktree; this packet is not a
hermetic reproduction bundle.

## Claim boundary

The selected 1.5B arm passed 88/98 within-generator execution tasks in the
original run but 0/1 on the only generator-independent historical canary.
Stock's original strict 0/98 is response-contract-confounded. A symmetric
whole-response adapter yields stock 9/98 versus Arm 2 at 88/98. With identical
contract-calibrated prompts, stock passes 8/98 and Arm 2 passes 89/98. The
raw-format confound is removed and the large in-domain gap remains. No SHELLper
benchmark was run.
