From f923b0abfe3a4bcca79e5b1a5eb04550d1697de9 Mon Sep 17 00:00:00 2001 From: norvalbv Date: Sun, 6 Sep 2026 01:08:25 +0100 Subject: [PATCH] docs(benchmarks): qualify typed-reply source evidence [sc-2866] The next source-family qualification attempt reproduces the archived kickoff-prompt resend, but withholds a clean repair label. The repaired extractor matches all 12 selected controls; a separate controlled run of the unmodified send function saves a hidden continuation before dispatch, at which point extraction already excludes the pending reply. Full application recovery remains unresolved, so these helper checks cannot certify a clean whole change. This adds a sanitized readout and receipt, a benchmark navigation link, and a native-CLI decision note for sc-2866. Exact source/claims and executable controls remain private. All 14 archive preimages match the selected source parent and the diff reconstructs byte-identically; 29 repair-source files are blob-verified. Historical unchanged context and the complete staged/capture roster are not authenticated. The incident retains its earlier scale-probe exposure. Validation: 12 extractor cases across three variants (original, explicit partial adaptation, source repair); controlled persistence-before-dispatch observation plus healthy continuation; current native census (3 chunks / 10 tasks, zero judge calls); 60 private hashes, source bindings, recorded outcomes, 28 local links and preserved decision-history bytes verified. Benchmark tracker and ship gates are reported below when complete. No full product suite is needed for this documentation-only delivery. Zero corpus admissions, qualified clean pairs, reviewer benchmark calls, score changes or production behavior changes. Parent sc-2832 remains open; the readout records the precise recovery/provenance questions required for reassessment. Shortcut: https://app.shortcut.com/benordlabs/story/2866 --- docs/benchmarks/README.md | 1 + ...ource-family-qualification-2026-09-06.json | 468 ++++++++++++++++++ .../source-family-qualification-2026-09-06.md | 151 ++++++ .../benchmarks-grow-from-telemetry.md | 1 + 4 files changed, 621 insertions(+) create mode 100644 docs/benchmarks/source-family-qualification-2026-09-06.json create mode 100644 docs/benchmarks/source-family-qualification-2026-09-06.md diff --git a/docs/benchmarks/README.md b/docs/benchmarks/README.md index c2e771cd..a4711614 100644 --- a/docs/benchmarks/README.md +++ b/docs/benchmarks/README.md @@ -58,6 +58,7 @@ Hand-written records that sit beside the ledger; each number in them names its p - [`reviewer-yield-vs-diff-size.md`](reviewer-yield-vs-diff-size.md) — the review gate returns ~1 issue per run above ~1k LOC and the corpus never measures that regime (2026-08-22). - [`corpus-growth.md`](corpus-growth.md) — how corpus rows are mined from telemetry and adapted. - [`source-family-census.md`](source-family-census.md) — source-qualified candidate families and exact native initial-evidence coverage (zero reviewer calls). +- [`source-family-qualification-2026-09-06.md`](source-family-qualification-2026-09-06.md) — source-controlled reply extraction and unresolved delivery-boundary qualification (sc-2866). - [`correctness-rebaseline-2026-09-04.md`](correctness-rebaseline-2026-09-04.md) — correctness re-baselined on the shipped gpt-5.6-sol pin: recall 0.92, clean-pass 0.60, and the precision loss sits entirely in the fix-pair decoys (sc-2494). - [`experiments/`](experiments/) — one folder per investigation with its inputs, arms, and reports (latest: [`2026-08-24-codex-judge-probe`](experiments/2026-08-24-codex-judge-probe/README.md) — OpenAI codex-CLI models as gate judges on the round-1 corpus: gpt-5.6-sol@chunk:400 13/23 and gpt-5.6-terra@xhigh WHOLE-diff 11/23 vs sonnet's best 8/23, ~perfect checklist compliance across 840 tasks (839 terminal); chunk size and reasoning effort proved ALTERNATIVE strategies, and terra's effort curve is a cliff at xhigh. Prior: [`2026-08-23-scale-probe`](experiments/2026-08-23-scale-probe/README.md) — whole-diff vs chunked correctness review on 8 real 1k–7.2k-LOC diffs: chunking pooled ≥ whole-diff under all five mining/scoring rulesets tried (shipped ruleset: chunk:1000 8/23 vs whole 5/23 on decontaminated labels), but the ratio moves 1.375×–2.0× with the ruleset so no registered bar is treated as cleared — decision moved to the confirmation round; ~half the round's original labels were test-retest contamination, since fixed; haiku judges failed their arms; full correction record inside. Prior: [`2026-08-22-ship-attempts-research`](experiments/2026-08-22-ship-attempts-research/README.md)). diff --git a/docs/benchmarks/source-family-qualification-2026-09-06.json b/docs/benchmarks/source-family-qualification-2026-09-06.json new file mode 100644 index 00000000..61b9003c --- /dev/null +++ b/docs/benchmarks/source-family-qualification-2026-09-06.json @@ -0,0 +1,468 @@ +{ + "version": 1, + "mode": "source-family-qualification", + "family": "family-005", + "exposure": "exposed-development", + "outcome": "withhold-clean-label", + "qualificationRetained": "unresolved", + "targetEvidence": "source-extractor-controlled", + "qualifiedCleanPairs": 0, + "admittedRows": 0, + "judgeCalls": 0, + "fiveRoundManifestCreated": false, + "sourceChecks": { + "scopedArchiveFiles": 14, + "matchedPreimages": 14, + "byteIdenticalReconstruction": true, + "originalCompleteStagedFiles": 22, + "originalHeadRecorded": false, + "unchangedContextAuthenticated": false, + "repairSourceFilesVerified": 29, + "changedLines": { + "additions": 436, + "deletions": 149, + "total": 585 + }, + "metricNote": "Changed source lines in scoped archive; old1054-line report is not reused as changed LOC." + }, + "controls": { + "extractorVariants": 3, + "casesPerVariant": 12, + "matchingCases": { + "bug": 5, + "partial": 8, + "repair": 12 + }, + "baseControl": "not-executed-helper-absent", + "partialVariant": "investigator-adapted-not-observed-repair", + "deliveryOrder": { + "wakeSaved": true, + "dispatchCallsWhilePaused": 0, + "pendingReplyBeforePersistence": true, + "pendingReplyAfterPersistence": false, + "healthyCompletionDispatchCalls": 1, + "actualSQLiteOrAppRestartExecuted": false + } + }, + "claims": { + "retainedOriginalLogStatements": 6, + "completeExactCapture": false, + "precisionEstimate": null + }, + "census": { + "manifestSha256": "ddd3616084b83c55595eba4987ef113a8391f48b23d3e4de2b4edc43c606f88f", + "implementationSha256": "991ed22dea56", + "runtime": "v24.19.0", + "condition": { + "cap": 400, + "groups": [ + [ + "state-transitions" + ], + [ + "concurrency-races" + ], + [ + "error-and-edge-classification" + ], + [ + "writer-reader-contracts" + ] + ], + "configSha256": "f12d7449ef4f1f0f495f44e4dc33b98df3926fe8685f8b79269b8a885fa54938", + "configSourceSha256": "31121e9e9fa0ce058190833df258ae0a88246de50bf96355eb37cf9dada55f05", + "model": "gpt-5.6-sol", + "judgeCalls": 0 + }, + "facts": { + "identityBytes": 31662, + "chunkCount": 3, + "taskCount": 10, + "cap": 400, + "groups": [ + [ + "state-transitions" + ], + [ + "concurrency-races" + ], + [ + "error-and-edge-classification" + ], + [ + "writer-reader-contracts" + ] + ], + "planHash": "2fb455dae335d02058bddc2ac1d6445422dd6a85adc2cc9cd3d2609f79965e9f", + "stagedTreeSha": "8672a3d7b3642cb17cf69433939269683c0cd94b", + "sourceChangedFiles": 14, + "selectedFiles": 14 + }, + "tasks": [ + { + "taskSha256": "cee5650cc2c53706896f9ba1cb09c03dcd646f027a8f52331db5fc8e1fe8a07a", + "lens": "state-transitions", + "targetLens": false, + "scope": "local-chunk", + "filesSha256": "a3bfffe63b3faba5d7250f93a0a0db679f206b458a74316a4217eccc63660ad7", + "fileCount": 6, + "diffSha256": "5be1ab3686cc6ce52426d50b71b3d58bc788b2afe985d8e2273cd4d5f96ef7d4", + "inputSha256": "1755444a71e2ff26cd9b5207987a058558d12ca56cbfc628debbff4445174b7e", + "inputBytes": 21447, + "required": [ + { + "spanSha256": "fdbf0585ba2660cb43f5005dda857f416f9a0a40f82c141a8003a20e5e4e7292", + "status": "out-of-scope", + "shownLines": 0, + "totalLines": 24, + "retrieval": "not-observed-zero-judge" + }, + { + "spanSha256": "ce5503f3d225c8ae67462d08c8cf15e43517e5ac193855e659742d6a3adbc5b3", + "status": "out-of-scope", + "shownLines": 0, + "totalLines": 19, + "retrieval": "not-observed-zero-judge" + }, + { + "spanSha256": "ad2d9066df534e0287eba7b4929620b8bbda5b1a11a709e51d8836d1b73bfef1", + "status": "out-of-scope", + "shownLines": 0, + "totalLines": 12, + "retrieval": "not-observed-zero-judge" + } + ] + }, + { + "taskSha256": "c5e0c86fb5b9f8fd9c32e9dc050845100913ea4143fba862f3f5f8e5b2e2d10d", + "lens": "concurrency-races", + "targetLens": false, + "scope": "local-chunk", + "filesSha256": "a3bfffe63b3faba5d7250f93a0a0db679f206b458a74316a4217eccc63660ad7", + "fileCount": 6, + "diffSha256": "5be1ab3686cc6ce52426d50b71b3d58bc788b2afe985d8e2273cd4d5f96ef7d4", + "inputSha256": "1755444a71e2ff26cd9b5207987a058558d12ca56cbfc628debbff4445174b7e", + "inputBytes": 21447, + "required": [ + { + "spanSha256": "fdbf0585ba2660cb43f5005dda857f416f9a0a40f82c141a8003a20e5e4e7292", + "status": "out-of-scope", + "shownLines": 0, + "totalLines": 24, + "retrieval": "not-observed-zero-judge" + }, + { + "spanSha256": "ce5503f3d225c8ae67462d08c8cf15e43517e5ac193855e659742d6a3adbc5b3", + "status": "out-of-scope", + "shownLines": 0, + "totalLines": 19, + "retrieval": "not-observed-zero-judge" + }, + { + "spanSha256": "ad2d9066df534e0287eba7b4929620b8bbda5b1a11a709e51d8836d1b73bfef1", + "status": "out-of-scope", + "shownLines": 0, + "totalLines": 12, + "retrieval": "not-observed-zero-judge" + } + ] + }, + { + "taskSha256": "af270d11457edd6bc77786f626296b34d4e923e99246287f80f4c21d39b4e0cb", + "lens": "error-and-edge-classification", + "targetLens": true, + "scope": "local-chunk", + "filesSha256": "a3bfffe63b3faba5d7250f93a0a0db679f206b458a74316a4217eccc63660ad7", + "fileCount": 6, + "diffSha256": "5be1ab3686cc6ce52426d50b71b3d58bc788b2afe985d8e2273cd4d5f96ef7d4", + "inputSha256": "1755444a71e2ff26cd9b5207987a058558d12ca56cbfc628debbff4445174b7e", + "inputBytes": 21447, + "required": [ + { + "spanSha256": "fdbf0585ba2660cb43f5005dda857f416f9a0a40f82c141a8003a20e5e4e7292", + "status": "out-of-scope", + "shownLines": 0, + "totalLines": 24, + "retrieval": "not-observed-zero-judge" + }, + { + "spanSha256": "ce5503f3d225c8ae67462d08c8cf15e43517e5ac193855e659742d6a3adbc5b3", + "status": "out-of-scope", + "shownLines": 0, + "totalLines": 19, + "retrieval": "not-observed-zero-judge" + }, + { + "spanSha256": "ad2d9066df534e0287eba7b4929620b8bbda5b1a11a709e51d8836d1b73bfef1", + "status": "out-of-scope", + "shownLines": 0, + "totalLines": 12, + "retrieval": "not-observed-zero-judge" + } + ] + }, + { + "taskSha256": "ad3aa0dc967c2ba18bcdbcd94e9e1c73b71add41fbfdc32b067cf8976e1ff8f2", + "lens": "state-transitions", + "targetLens": false, + "scope": "local-chunk", + "filesSha256": "dade4265ce287ad5b87caa58f30cbb8d8509efa9582ff88b651656cd394cc776", + "fileCount": 7, + "diffSha256": "3675fe20a156195312119b8bf0a48d2780003728b765dd48cb79c1f62bce8b6f", + "inputSha256": "da04d3f5f35ac1cfa7c8eac1061db6f45a2af900d16f3a9f77408a822fadf4c6", + "inputBytes": 19906, + "required": [ + { + "spanSha256": "fdbf0585ba2660cb43f5005dda857f416f9a0a40f82c141a8003a20e5e4e7292", + "status": "supplied", + "shownLines": 24, + "totalLines": 24, + "retrieval": "not-observed-zero-judge" + }, + { + "spanSha256": "ce5503f3d225c8ae67462d08c8cf15e43517e5ac193855e659742d6a3adbc5b3", + "status": "partial", + "shownLines": 14, + "totalLines": 19, + "retrieval": "not-observed-zero-judge" + }, + { + "spanSha256": "ad2d9066df534e0287eba7b4929620b8bbda5b1a11a709e51d8836d1b73bfef1", + "status": "out-of-scope", + "shownLines": 0, + "totalLines": 12, + "retrieval": "not-observed-zero-judge" + } + ] + }, + { + "taskSha256": "b96c3cde4e76dbf4bcd028a96a0c513d27088a0807340a1570e2107e7d753f5b", + "lens": "concurrency-races", + "targetLens": false, + "scope": "local-chunk", + "filesSha256": "dade4265ce287ad5b87caa58f30cbb8d8509efa9582ff88b651656cd394cc776", + "fileCount": 7, + "diffSha256": "3675fe20a156195312119b8bf0a48d2780003728b765dd48cb79c1f62bce8b6f", + "inputSha256": "da04d3f5f35ac1cfa7c8eac1061db6f45a2af900d16f3a9f77408a822fadf4c6", + "inputBytes": 19906, + "required": [ + { + "spanSha256": "fdbf0585ba2660cb43f5005dda857f416f9a0a40f82c141a8003a20e5e4e7292", + "status": "supplied", + "shownLines": 24, + "totalLines": 24, + "retrieval": "not-observed-zero-judge" + }, + { + "spanSha256": "ce5503f3d225c8ae67462d08c8cf15e43517e5ac193855e659742d6a3adbc5b3", + "status": "partial", + "shownLines": 14, + "totalLines": 19, + "retrieval": "not-observed-zero-judge" + }, + { + "spanSha256": "ad2d9066df534e0287eba7b4929620b8bbda5b1a11a709e51d8836d1b73bfef1", + "status": "out-of-scope", + "shownLines": 0, + "totalLines": 12, + "retrieval": "not-observed-zero-judge" + } + ] + }, + { + "taskSha256": "bc1c097c1feb50e327d98e9207dca1d44603e8f4a5f9101ab6bf4422b6823016", + "lens": "error-and-edge-classification", + "targetLens": true, + "scope": "local-chunk", + "filesSha256": "dade4265ce287ad5b87caa58f30cbb8d8509efa9582ff88b651656cd394cc776", + "fileCount": 7, + "diffSha256": "3675fe20a156195312119b8bf0a48d2780003728b765dd48cb79c1f62bce8b6f", + "inputSha256": "da04d3f5f35ac1cfa7c8eac1061db6f45a2af900d16f3a9f77408a822fadf4c6", + "inputBytes": 19906, + "required": [ + { + "spanSha256": "fdbf0585ba2660cb43f5005dda857f416f9a0a40f82c141a8003a20e5e4e7292", + "status": "supplied", + "shownLines": 24, + "totalLines": 24, + "retrieval": "not-observed-zero-judge" + }, + { + "spanSha256": "ce5503f3d225c8ae67462d08c8cf15e43517e5ac193855e659742d6a3adbc5b3", + "status": "partial", + "shownLines": 14, + "totalLines": 19, + "retrieval": "not-observed-zero-judge" + }, + { + "spanSha256": "ad2d9066df534e0287eba7b4929620b8bbda5b1a11a709e51d8836d1b73bfef1", + "status": "out-of-scope", + "shownLines": 0, + "totalLines": 12, + "retrieval": "not-observed-zero-judge" + } + ] + }, + { + "taskSha256": "7a8ad42c4dd540102e05167cbaad6470edcf209c0656036000c08fd67e798765", + "lens": "state-transitions", + "targetLens": false, + "scope": "local-chunk", + "filesSha256": "11cc9865e8da63913773ad49107a7ff358858a338aa3e7088bb1b7c1027e61c5", + "fileCount": 1, + "diffSha256": "45dc50a99b8f0b8c47a3e26bb3f5070a058d50f122f05853bff16a78c33a604e", + "inputSha256": "62b075e186250f9dc2bce9244572494ef468c3bf3904ecd54153b980a509ff0d", + "inputBytes": 1893, + "required": [ + { + "spanSha256": "fdbf0585ba2660cb43f5005dda857f416f9a0a40f82c141a8003a20e5e4e7292", + "status": "out-of-scope", + "shownLines": 0, + "totalLines": 24, + "retrieval": "not-observed-zero-judge" + }, + { + "spanSha256": "ce5503f3d225c8ae67462d08c8cf15e43517e5ac193855e659742d6a3adbc5b3", + "status": "out-of-scope", + "shownLines": 0, + "totalLines": 19, + "retrieval": "not-observed-zero-judge" + }, + { + "spanSha256": "ad2d9066df534e0287eba7b4929620b8bbda5b1a11a709e51d8836d1b73bfef1", + "status": "out-of-scope", + "shownLines": 0, + "totalLines": 12, + "retrieval": "not-observed-zero-judge" + } + ] + }, + { + "taskSha256": "bfeb0269976dacb702246a8ddaebcacbb715c4857deb6a38bf7da6627f4ac2b1", + "lens": "concurrency-races", + "targetLens": false, + "scope": "local-chunk", + "filesSha256": "11cc9865e8da63913773ad49107a7ff358858a338aa3e7088bb1b7c1027e61c5", + "fileCount": 1, + "diffSha256": "45dc50a99b8f0b8c47a3e26bb3f5070a058d50f122f05853bff16a78c33a604e", + "inputSha256": "62b075e186250f9dc2bce9244572494ef468c3bf3904ecd54153b980a509ff0d", + "inputBytes": 1893, + "required": [ + { + "spanSha256": "fdbf0585ba2660cb43f5005dda857f416f9a0a40f82c141a8003a20e5e4e7292", + "status": "out-of-scope", + "shownLines": 0, + "totalLines": 24, + "retrieval": "not-observed-zero-judge" + }, + { + "spanSha256": "ce5503f3d225c8ae67462d08c8cf15e43517e5ac193855e659742d6a3adbc5b3", + "status": "out-of-scope", + "shownLines": 0, + "totalLines": 19, + "retrieval": "not-observed-zero-judge" + }, + { + "spanSha256": "ad2d9066df534e0287eba7b4929620b8bbda5b1a11a709e51d8836d1b73bfef1", + "status": "out-of-scope", + "shownLines": 0, + "totalLines": 12, + "retrieval": "not-observed-zero-judge" + } + ] + }, + { + "taskSha256": "93991ea192cc83e0c8d7f25517d2475f9b9308153161e145faa9f2191a1f4285", + "lens": "error-and-edge-classification", + "targetLens": true, + "scope": "local-chunk", + "filesSha256": "11cc9865e8da63913773ad49107a7ff358858a338aa3e7088bb1b7c1027e61c5", + "fileCount": 1, + "diffSha256": "45dc50a99b8f0b8c47a3e26bb3f5070a058d50f122f05853bff16a78c33a604e", + "inputSha256": "62b075e186250f9dc2bce9244572494ef468c3bf3904ecd54153b980a509ff0d", + "inputBytes": 1893, + "required": [ + { + "spanSha256": "fdbf0585ba2660cb43f5005dda857f416f9a0a40f82c141a8003a20e5e4e7292", + "status": "out-of-scope", + "shownLines": 0, + "totalLines": 24, + "retrieval": "not-observed-zero-judge" + }, + { + "spanSha256": "ce5503f3d225c8ae67462d08c8cf15e43517e5ac193855e659742d6a3adbc5b3", + "status": "out-of-scope", + "shownLines": 0, + "totalLines": 19, + "retrieval": "not-observed-zero-judge" + }, + { + "spanSha256": "ad2d9066df534e0287eba7b4929620b8bbda5b1a11a709e51d8836d1b73bfef1", + "status": "out-of-scope", + "shownLines": 0, + "totalLines": 12, + "retrieval": "not-observed-zero-judge" + } + ] + }, + { + "taskSha256": "681f5b1d8ee30208e088da72fa1ce21839f3f72b6016dbcc30161cced804ad9c", + "lens": "writer-reader-contracts", + "targetLens": false, + "scope": "whole-diff-contracts", + "filesSha256": "4450f9bca42c50d37af733b5969d6e1d2d4eefe96614f7a75eb0f0389088c9b7", + "fileCount": 14, + "diffSha256": "032d886019d372296aa86749ef79a71b964379c7a6c1eb90fac1868b5de04e73", + "inputSha256": "aac33c2a982da0d92ecec0a6a7a15b1344575e5f138d3ddf8274d468b41cf62a", + "inputBytes": 43122, + "required": [ + { + "spanSha256": "fdbf0585ba2660cb43f5005dda857f416f9a0a40f82c141a8003a20e5e4e7292", + "status": "supplied", + "shownLines": 24, + "totalLines": 24, + "retrieval": "not-observed-zero-judge" + }, + { + "spanSha256": "ce5503f3d225c8ae67462d08c8cf15e43517e5ac193855e659742d6a3adbc5b3", + "status": "partial", + "shownLines": 14, + "totalLines": 19, + "retrieval": "not-observed-zero-judge" + }, + { + "spanSha256": "ad2d9066df534e0287eba7b4929620b8bbda5b1a11a709e51d8836d1b73bfef1", + "status": "out-of-scope", + "shownLines": 0, + "totalLines": 12, + "retrieval": "not-observed-zero-judge" + } + ] + } + ] + }, + "receipts": { + "original.private.diff": "032d886019d372296aa86749ef79a71b964379c7a6c1eb90fac1868b5de04e73", + "original-ship.private.log": "c8a2c2427cd0bf159329a299b80a3d4b6e087576640e98215ba9aa9815dbe043", + "original-events.private.json": "85408afcc1c96111a1c9bf448611fc1ecd685e9605a0fe8ceae019370598a41f", + "pr387.private.json": "42f352a852957c53a41e7a677e98d1a3eea9b48ebd1be558871bd27fe9b423af", + "preimage-check.private.json": "651ac4f9e7678d1bacc4f66ac5d2d2f2663365dd4272ef539810ae1f07fb4a2a", + "provenance.private.json": "d5e000eb7061d92a9744e91600db1d462198eda2e23bd4fd974480de818c2b19", + "base-caller.private.ts": "2d554e9fef87cd611a42871e56351dcfbacabfdd30a153567ddd1fa0d374ca7a", + "control-source-bindings.private.json": "1b8cc0a3e73ebbfb27f270a9bb0788d1419deef65771f96fa2b8fef0774f7898", + "repair-source-bindings.private.json": "5d03e7ba3b4e3acefaf3978c3abb56481208425afdb528db868c929beb0e4aa4", + "run-controls.mjs": "c3829386dd8617af6907133494456c627c681e8aaf9d475c8fff97e205c75a73", + "controls-ryw8hk/report.private.json": "6fcc98ac252eb8c9904a37526cefebba25576232c14ca2506bba0f07b52aa310", + "delivery-order-control.mjs": "99e2872ed7dee0af79015a6c51751a9385d6cf9f9a044d4bc147fa74120bc4ea", + "delivery-controls-QkOjma/report.private.json": "35e6c907b93b3c521fbfd756b25aec814d2a979c8d194347d33ef4875642e7e7", + "claim-ledger.private.json": "990dd65c051e3eeef092f462312700ed5926606f649fea3a217c924f0cf71a24", + "requirement.private.md": "5004bfc4c53cfb23583b087f93ae870a9efa26f343a6081fa43b95adeb00f573", + "assessment.private.json": "b28586b7f460567c06ee71a3dddf3269cf561f9a492a453426535a977be94b17", + "manifest.private.json": "fdc512f5c6e411bd12fa7d2dd448268996d84f4b246797c47d5a9267add43a44", + "census-output.json": "087ebca008acb3981a54b4d6df82211c6aa2448b62d3c62a28d952b233aef02a", + "prior-art.json": "970187d038c4bd50868ba436e9d49e839d27eac3dbb4785ff8a93b77dd3e7d70", + "plan-critique.json": "b6b13b5acbf0951534f7eb3962f4dc9a1b49f3868e61d69f6b15b6c0af37811e", + "verify-evidence.py": "c92055ecac565b1f46e5fcf2ef660df096e71dccd70a3985710736ae8d6485b1" + }, + "privateSourceInventorySha256": "18644f230db1a821d3df53904d02c37e463000660e54699dc26f9121121c9342" +} diff --git a/docs/benchmarks/source-family-qualification-2026-09-06.md b/docs/benchmarks/source-family-qualification-2026-09-06.md new file mode 100644 index 00000000..92e9a6ce --- /dev/null +++ b/docs/benchmarks/source-family-qualification-2026-09-06.md @@ -0,0 +1,151 @@ +# Typed-reply source qualification (sc-2866) + +**The original resend is reproduced; the clean label remains withheld.** This investigation adds +source evidence for another incident, `family-005`, using the existing qualification workflow. +It admits zero rows, creates no qualified five-round timeline, and runs no reviewer benchmark. +The [sanitized receipt](source-family-qualification-2026-09-06.json) preserves the controls and +native coverage census. [Source-control decisions](../decisions/benchmarks-grow-from-telemetry.md) +and [history-experiment decisions](../decisions/reviewer-history-experiments.md) govern the result. + +## What the example shows + +A task starts with its original instructions and crashes before its first assistant response. +On retry, the archived extractor treats those instructions as a new user reply. Its caller then +uses them instead of the continuation nudge. The source repair excludes the initial message while +retaining later replies. That narrow behavior is testable without running a model. + +| Transcript scenario | Archived extractor | Adapted partial repair | Source repair | +| --- | --- | --- | --- | +| Original instructions only | Resends instructions | Returns no new reply | Returns no new reply | +| Instructions, then a newer reply | Includes both | Keeps newer reply | Keeps newer reply | +| Replies before a delivered hidden continuation, then a newer reply | Includes old replies again | Includes old replies again | Keeps newer reply | +| A hidden continuation was saved but dispatch has not happened | Keeps preceding reply | Keeps preceding reply | Returns no new reply | + +The last row is a different condition from a **delivered** continuation. The source repair assumes +that saving a hidden continuation establishes a delivery boundary. Its actual send path persists +the message before awaiting machine identity and before notifying the execution listener. + +A separate control executes the unmodified source `sendMessage` with explicit dependency doubles. +It pauses the identity dependency after a file-backed persistence double has saved the hidden +message. At that point there are **zero execution-listener calls**, but the repaired source +extractor already returns no pending reply. Releasing the pause produces one execution call. +This demonstrates why persistence alone cannot certify dispatch. + +This is **not a complete application-loss reproduction**: actual SQLite, app restart, admission +recovery and the provider are not executed. The full recovery outcome remains unresolved. The +clean label is withheld rather than resolving that uncertainty in the repair's favor. This also +does not assert that today's production implementation still has the same behavior. + +## Controls and source identity + +The original correctness-scoped archive contains 14 files: 436 additions and 149 deletions, +**585 changed source lines**. The old scale readout's 1,054-line figure is not reused as a changed-LOC +count. No tests or documentation were added to inflate the selected input. + +All 14 archived preimages match the parent of the eventual source implementation commit. Applying +the archive there reproduces its complete diff bytes exactly, including Git's nine-character blob +abbreviations. Every captured repaired-source file was checked against that implementation commit: +29 files, with full Git blob IDs and SHA-256 hashes retained privately. + +Those checks establish a reproducible selected-source reconstruction. They do **not authenticate +all unchanged historical context**: the available original telemetry does not record HEAD. The +archive is the correctness scope, while the original ship formatted 22 staged files. The complete +original staged change and exact native checklist captures have not been recovered. Neither a +merely applicable parent nor the final PR description fills those gaps. + +The extraction control executes the unchanged source extractor, marker helper and synchronous +image readers. A fail-closed stub prevents access to the unused authentication dependency; there +are no real credentials, network requests or production database writes. Across 12 explicitly +chosen cases, the archived extractor matches five expected outputs, the adapted partial version +eight, and the source repair twelve. These are **control outcomes, not reviewer accuracy scores**. +Cases cover initial/newer/multiple replies, assistant and conditional delivered-wake boundaries, +both image shapes, image-only input, empty input and non-image files. + +The partial version changes only the initial-message boundary in the original extractor. It is an +investigator adaptation, not an observed intermediate repair or an independent incident. The final +source repair also changes surrounding modules; it is not transplanted into the original archive +and presented as an observed clean whole change. + +The helper is absent from the original source base. The actual base caller was inspected and sends +the continuation nudge directly, but was not executed by these controls. No executable base PASS +is claimed. Source inspection establishes the normal task-created, kickoff-first transcript path; +other imported or concurrent ordering remains unqualified. A type permitting system messages +alone does not prove a reachable system-prefix defect. + +## Exposure, other findings and native coverage + +This incident already appeared in the [August scale probe](experiments/2026-08-23-scale-probe/README.md) +as `032d8860`. It is separate from the four incidents investigated by [sc-2002](source-family-census.md), +but it is **exposed development evidence**, not a new unseen holdout. All future descendants must +retain that incident linkage. Repeated findings and adapted rounds cannot increase the independent +family count. + +The private ledger retains six statements visible in the original ship log: two correctness +statements and four completeness statements. The extraction finding has the controls above; the +mode-propagation claim, latch clearing and toast behavior remain unresolved. A separate capacity +classification is supported by inspected source but has no full admission execution control here. +The documentation complaint is outside the declared extractor scope. These six statements are +neither complete checklist captures nor a replacement for the older mined-label population; they +cannot support a precision estimate or establish that the remaining change is clean. + +A current, zero-judge native census of the reconstructed archive produces **three chunks and ten +tasks** under cap 400, with the Sol condition explicitly pinned. The relevant parsing task receives +the complete declared extractor span and part of the caller span; the unchanged marker helper is +outside its selected files. The contracts task has the same supplied/partial/out-of-scope pattern. +Other local tasks have those target spans out of scope. These are initial-input observations, +not measured retrieval, misses or ten historical reviewer outcomes. + +## Reproduction + +Use Node 24.19.0 and the private evidence directory +`~/.devkit/research/source-family-next-20260906/`. Exact source and claim text remain there. The +committed receipt hashes the private source inventory, source provenance, control scripts and +results, assessment, claim ledger, research and census. Missing or mismatched files mean unavailable +reproduction; never invent replacement historical evidence. + +1. Run `verify-evidence.py` from the private directory to check the committed receipt, private + hashes, Git source bindings, byte-identical archived diff and recorded outcomes. +2. Run the extraction and send-order controls. Each writes to a new private run directory; preserve + the original results. The recorded successful runs are `controls-ryw8hk` and + `delivery-controls-QkOjma`. Temporary paths and output hashes differ on rerun. + +```sh +node ~/.devkit/research/source-family-next-20260906/run-controls.mjs +node ~/.devkit/research/source-family-next-20260906/delivery-order-control.mjs +``` + +3. From a devkit checkout containing PR603, rerun the native census against the frozen source view: + +```sh +export GUARD_CORRECTNESS_MODEL=gpt-5.6-sol +node gate-engine/review/eval/reviewers/scale/corpus/census-cli.mts \ + ~/.devkit/research/source-family-next-20260906/manifest.private.json case-005 \ + ~/.devkit/research/source-family-next-20260906/source-bug \ + ~/.devkit/research/source-family-next-20260906 +``` + +The source view is sparse and privately owned. Its complete index records unchanged source objects; +only the materialized input and inspected dependencies are used. The census manifest intentionally +retains `unresolved` with no qualified-label evidence fields. No repaired-view census or complete +five-round execution is claimed. Qualification is a separate assertion from source visibility. + +Validation for this documentation delivery covers the source controls, native census, private and +public receipt reconciliation, source hashes, JSON, local links and benchmark tracking. It does +not require a full product test run because implementation, corpus and scores are unchanged. + +## Research and next action + +[SWE-Review v1, Appendix B.2](https://arxiv.org/html/2607.06065v1#A2.SS2) motivates checking faulty or +insufficient reproducers. [MalPR-Bench v1](https://arxiv.org/html/2608.25730v1) distinguishes target +diagnosis and paired complete-fix controls. [MCR-Bench v1](https://arxiv.org/html/2608.27442v1) motivates +lifecycle consistency checks. These papers motivate the procedure; they do not verify this source +case or prove that history helps Sol. Prior-art review recommended reusing the existing workflow. +Reference checkouts and the deep-research service were unavailable; broader prior art is unverified. + +The next reassessment must establish the delivery acknowledgement and recovery contract around +interrupted post-persistence/pre-dispatch sends, including concurrent typed intake, and validate +actual caller/base behavior and remaining source-context provenance. If that broad source slice +cannot be qualified, acquire another coherent source family. A smaller adapted extractor case +would still be the same exposed family and would need explicit admission; it must not silently +stand in for a clean large PR. [sc-2832](https://app.shortcut.com/benordlabs/story/2832) remains open: +qualified frozen cases and native execution integration are still required before measurement. diff --git a/docs/decisions/benchmarks-grow-from-telemetry.md b/docs/decisions/benchmarks-grow-from-telemetry.md index 93f6962f..ee5f2465 100644 --- a/docs/decisions/benchmarks-grow-from-telemetry.md +++ b/docs/decisions/benchmarks-grow-from-telemetry.md @@ -41,3 +41,4 @@ created: 2026-08-01 **Evidence-change:** Relative to the August 1 ruling, sc-2494 remeasurement localized 25/26 false blocks to repair twins and identified zero native chunk coverage. The owner-authorized sc-2500 explicitly re-aimed growth at large fix-pairs and forbade additional gold. The September 5 holdout-reset evidence independently retained the same pattern (27/30 repair false blocks). - 2026-09-05 — Retrospective capture-to-corpus boundary clarified by sc-2493 and sc-2500: ship evidence and same-lens FAIL→PASS/waiver mining create investigation candidates, not automatic truth labels. Exact archived source, original base, callers and producer constraints must establish the requirement; external executable controls distinguish base, buggy change and repair without leaking assertions to the reviewer. A passing targeted control establishes that invariant, not absence of all other introduced defects. Investigation/adaptation and admission are explicit workflows, not an automatic monthly promotion of telemetry. Existing [corpus-growth](../benchmarks/corpus-growth.md) and [claim adjudication](../benchmarks/claim-adjudication.md) provide the commands and evidence boundaries; this note does not expand the existing source-anchoring permission. - 2026-09-05 — Retrospective sc-2002 / sc-2843 outcomes (PR601–602, 2026-09-05): 577 candidate-labelled diffs among 780 archived scoped diffs were observations, not independent bugs. Four investigated families yielded one exposed target-controlled large diagnostic pair and zero qualified clean pairs; other families had explicit context/size/source-repair exclusions. PR602 then rejected that narrow repaired member as clean: its target dispatch control passed, but an actual source-CLI control reproduced loss of failed-refresh uncertainty on a later cached check. Both members remain target-controlled; zero rows were admitted. A targeted repair can therefore be valid for its defect while unsuitable as a clean benchmark control. [Source census](../benchmarks/source-family-census.md) and [qualification receipt](../benchmarks/source-family-qualification-2026-09-05.md) retain the investigation and original identities; do not backfill missing source evidence or call those cases unseen. +- 2026-09-06 — sc-2866 qualifies another exposed source incident after PR603 using the existing workflow: the original reply extractor reproduces the kickoff resend and the final source extractor matches all 12 selected control cases. A separate unmodified source send-function control, with explicit persistence and scheduling doubles, saves the hidden wake before dispatch; the extractor already excludes the pending reply. Full application recovery, historical unchanged context, executable base behavior and other claims remain unresolved, so no clean pair, corpus row or five-round manifest is admitted. Repaired helper success cannot substitute for delivery evidence or a clean whole change. Exact private provenance and limitations are in the [qualification readout](../benchmarks/source-family-qualification-2026-09-06.md); the old scale exposure and all prior receipts remain intact.