Skip to content

Prepare comparative correctness experiments and unscored mechanism probes - #605

Merged
norvalbv merged 2 commits into
mainfrom
codex/sc2872-correctness-reachable-witness
Sep 6, 2026
Merged

Prepare comparative correctness experiments and unscored mechanism probes#605
norvalbv merged 2 commits into
mainfrom
codex/sc2872-correctness-reachable-witness

Conversation

@norvalbv

@norvalbv norvalbv commented Sep 6, 2026

Copy link
Copy Markdown
Owner

Prepares alternatives to the original single prompt change: baseline B, independent reachability guidance P, conditional lens-rule bundle L, and source-context delivery C. Production agents, skills, corpus, runner and four-lens routing are unchanged.

Only eight existing source-backed guardrail rows may be scored: B/P/L × eight rows × two rounds = 48 proposed executions. Seven newly authored JS examples are separate, unscored probes: B/P/L/C × seven probes × two rounds = 56 proposed exploratory executions. They have no expected verdict or native scoring metadata and cannot produce benchmark metrics, rank candidates or reward an optimizer. C is unscored only. Passing authored controls cannot create gold provenance; a scored comparison of the motivating mechanisms still needs independently anchored source cases.

Includes two independent unapplied patches, exact unchanged-source packets, inert probe/control definitions, current arXiv/CodeRabbit/Macroscope and promptfoo research, and a native integration handoff. Records family-split/cache limitations, the evidence required for a fifth lens, and the provenance correction caught by ship review. Durable notes use the decisions CLI.

Readiness: the native scored/unscored adapter is NOT implemented or validated. The old CLI commands cannot execute this mixed plan. This preparation neither establishes improvement nor authorizes measurement or production promotion. No optimizer or dependency is installed.

Validation: JSON, 15 original source hashes, 8 original row hashes, 7 unscored probe hashes, 7 source packets, both independent patch applications/resulting hashes, family membership and 17 JS syntax checks passed. No control program, benchmark, probe, census, optimizer or evaluation was executed. Normal ship gates review the preparation artifacts; embedded fixture-control strings must not be executed.

The user's explicit execution pause remains in force after merge. Future execution needs permission and validated native separation/parity. Source-backed controls must qualify scored evidence; direct-mechanism/context scoring additionally needs eligible source cases. Probe observations never enter scored denominators.

Shortcut: https://app.shortcut.com/benordlabs/story/2872
Epic: https://app.shortcut.com/benordlabs/epic/2830

The preparation story closes only after owner merge. sc-2832's historical experiment remains separate and blocked on its own prerequisites.

…rk execution

Prepares one research-grounded correctness prompt experiment after the complete-claim replay exposed repeated unsupported caller/lifetime findings. The unapplied patch requires a proposed failing scenario to follow a reachable post-change path through actual guards, earlier writes and caller or exported-contract constraints. It preserves the existing adversarial PASS check and production reviewer behavior.

Includes primary arXiv and CodeRabbit/Macroscope research, an exact patch with before/after hashes, eight pinned exposed bug/repair rows, source-audit notes and a future matched-run protocol. The four pairs are regression guardrails, not evidence that the motivating lifetime/send-guard problem is fixed. Runtime settings, task-archive retention, missing outcomes and claim-vs-target scoring are explicit. Native decision notes retain the rationale and current benchmark pause.

Validation: JSON, all eight exact row hashes, 15 source hashes, exact prefix membership, patch applicability and resulting prompt hash checked; git diff --check passed. No benchmark, census, fixture control or evaluation was executed for preparation. Production agent/skill/corpus/runner files and shared baselines remain unchanged. Normal ship integrity checks validate artifacts without measuring reviewer performance. No implementation-source changes require a test-suite run.

The user's pause remains in force: merging does not activate the patch or authorize any benchmark. Direct source-controlled mechanism cases and separate confirmation are still required before production promotion. No improvement is claimed.

Shortcut: https://app.shortcut.com/benordlabs/story/2872
Epic: https://app.shortcut.com/benordlabs/epic/2830

Resolves preparation scope of sc-2872; sc-2832's history experiment remains separately blocked.
@coderabbitai

coderabbitai Bot commented Sep 6, 2026

Copy link
Copy Markdown
Contributor

Warning

Review limit reached

Next included review available in 46 minutes.

Check out review usage here.

View limit details

Limit details: You’ve used the included review currently available.

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

Learn how review limits work.

Review configuration:

⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Team

Run ID: a0ae19e3-d6bf-4047-ab40-171c8dce6864

📥 Commits

Reviewing files that changed from the base of the PR and between 9b08577 and 73d0516.

📒 Files selected for processing (14)
  • docs/benchmarks/experiments/2026-09-06-correctness-reachable-witness/README.md
  • docs/benchmarks/experiments/2026-09-06-correctness-reachable-witness/agent.patch
  • docs/benchmarks/experiments/2026-09-06-correctness-reachable-witness/comparison.md
  • docs/benchmarks/experiments/2026-09-06-correctness-reachable-witness/context-packets.json
  • docs/benchmarks/experiments/2026-09-06-correctness-reachable-witness/draft-cases.json
  • docs/benchmarks/experiments/2026-09-06-correctness-reachable-witness/lens-rules.patch
  • docs/benchmarks/experiments/2026-09-06-correctness-reachable-witness/protocol.json
  • docs/benchmarks/experiments/2026-09-06-correctness-reachable-witness/research.md
  • docs/benchmarks/experiments/2026-09-06-correctness-reachable-witness/source-audit.md
  • docs/benchmarks/experiments/2026-09-06-correctness-reachable-witness/tooling.md
  • docs/decisions/benchmarks-grow-from-telemetry.md
  • docs/decisions/correctness-lens-hole-instrument.md
  • docs/decisions/correctness-reviewer-precision.md
  • docs/decisions/reviewer-claim-measurement.md

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

Prepares alternatives to the original single prompt change: baseline B, independent reachability guidance P, conditional lens-rule bundle L, and source-context delivery C. Production agents, skills, corpus, runner and four-lens routing are unchanged.

Only eight existing source-backed guardrail rows may be scored: B/P/L × eight rows × two rounds = 48 proposed executions. Seven newly authored JS examples are separate, unscored probes: B/P/L/C × seven probes × two rounds = 56 proposed exploratory executions. They have no expected verdict or native scoring metadata and cannot produce benchmark metrics, rank candidates or reward an optimizer. C is unscored only. Passing authored controls cannot create gold provenance; a scored comparison of the motivating mechanisms still needs independently anchored source cases.

Includes two independent unapplied patches, exact unchanged-source packets, inert probe/control definitions, current arXiv/CodeRabbit/Macroscope and promptfoo research, and a native integration handoff. Records family-split/cache limitations, the evidence required for a fifth lens, and the provenance correction caught by ship review. Durable notes use the decisions CLI.

Readiness: the native scored/unscored adapter is NOT implemented or validated. The old CLI commands cannot execute this mixed plan. This preparation neither establishes improvement nor authorizes measurement or production promotion. No optimizer or dependency is installed.

Validation: JSON, 15 original source hashes, 8 original row hashes, 7 unscored probe hashes, 7 source packets, both independent patch applications/resulting hashes, family membership and 17 JS syntax checks passed. No control program, benchmark, probe, census, optimizer or evaluation was executed. Normal ship gates review the preparation artifacts; embedded fixture-control strings must not be executed.

The user's explicit execution pause remains in force after merge. Future execution needs permission and validated native separation/parity. Source-backed controls must qualify scored evidence; direct-mechanism/context scoring additionally needs eligible source cases. Probe observations never enter scored denominators.

Shortcut: https://app.shortcut.com/benordlabs/story/2872
Epic: https://app.shortcut.com/benordlabs/epic/2830

The preparation story closes only after owner merge. sc-2832's historical experiment remains separate and blocked on its own prerequisites.
@norvalbv norvalbv changed the title Prepare source-grounded correctness prompt experiment without benchmark execution Prepare comparative correctness experiments and unscored mechanism probes Sep 6, 2026
@norvalbv
norvalbv merged commit 05dcafa into main Sep 6, 2026
1 of 2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant