Prepare comparative correctness experiments and unscored mechanism probes - #605
Merged
Merged
Conversation
…rk execution Prepares one research-grounded correctness prompt experiment after the complete-claim replay exposed repeated unsupported caller/lifetime findings. The unapplied patch requires a proposed failing scenario to follow a reachable post-change path through actual guards, earlier writes and caller or exported-contract constraints. It preserves the existing adversarial PASS check and production reviewer behavior. Includes primary arXiv and CodeRabbit/Macroscope research, an exact patch with before/after hashes, eight pinned exposed bug/repair rows, source-audit notes and a future matched-run protocol. The four pairs are regression guardrails, not evidence that the motivating lifetime/send-guard problem is fixed. Runtime settings, task-archive retention, missing outcomes and claim-vs-target scoring are explicit. Native decision notes retain the rationale and current benchmark pause. Validation: JSON, all eight exact row hashes, 15 source hashes, exact prefix membership, patch applicability and resulting prompt hash checked; git diff --check passed. No benchmark, census, fixture control or evaluation was executed for preparation. Production agent/skill/corpus/runner files and shared baselines remain unchanged. Normal ship integrity checks validate artifacts without measuring reviewer performance. No implementation-source changes require a test-suite run. The user's pause remains in force: merging does not activate the patch or authorize any benchmark. Direct source-controlled mechanism cases and separate confirmation are still required before production promotion. No improvement is claimed. Shortcut: https://app.shortcut.com/benordlabs/story/2872 Epic: https://app.shortcut.com/benordlabs/epic/2830 Resolves preparation scope of sc-2872; sc-2832's history experiment remains separately blocked.
Contributor
|
Warning Review limit reachedNext included review available in 46 minutes. View limit detailsLimit details: You’ve used the included review currently available. You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository. Review configuration: ⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Team Run ID: 📒 Files selected for processing (14)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Prepares alternatives to the original single prompt change: baseline B, independent reachability guidance P, conditional lens-rule bundle L, and source-context delivery C. Production agents, skills, corpus, runner and four-lens routing are unchanged. Only eight existing source-backed guardrail rows may be scored: B/P/L × eight rows × two rounds = 48 proposed executions. Seven newly authored JS examples are separate, unscored probes: B/P/L/C × seven probes × two rounds = 56 proposed exploratory executions. They have no expected verdict or native scoring metadata and cannot produce benchmark metrics, rank candidates or reward an optimizer. C is unscored only. Passing authored controls cannot create gold provenance; a scored comparison of the motivating mechanisms still needs independently anchored source cases. Includes two independent unapplied patches, exact unchanged-source packets, inert probe/control definitions, current arXiv/CodeRabbit/Macroscope and promptfoo research, and a native integration handoff. Records family-split/cache limitations, the evidence required for a fifth lens, and the provenance correction caught by ship review. Durable notes use the decisions CLI. Readiness: the native scored/unscored adapter is NOT implemented or validated. The old CLI commands cannot execute this mixed plan. This preparation neither establishes improvement nor authorizes measurement or production promotion. No optimizer or dependency is installed. Validation: JSON, 15 original source hashes, 8 original row hashes, 7 unscored probe hashes, 7 source packets, both independent patch applications/resulting hashes, family membership and 17 JS syntax checks passed. No control program, benchmark, probe, census, optimizer or evaluation was executed. Normal ship gates review the preparation artifacts; embedded fixture-control strings must not be executed. The user's explicit execution pause remains in force after merge. Future execution needs permission and validated native separation/parity. Source-backed controls must qualify scored evidence; direct-mechanism/context scoring additionally needs eligible source cases. Probe observations never enter scored denominators. Shortcut: https://app.shortcut.com/benordlabs/story/2872 Epic: https://app.shortcut.com/benordlabs/epic/2830 The preparation story closes only after owner merge. sc-2832's historical experiment remains separate and blocked on its own prerequisites.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Prepares alternatives to the original single prompt change: baseline B, independent reachability guidance P, conditional lens-rule bundle L, and source-context delivery C. Production agents, skills, corpus, runner and four-lens routing are unchanged.
Only eight existing source-backed guardrail rows may be scored: B/P/L × eight rows × two rounds = 48 proposed executions. Seven newly authored JS examples are separate, unscored probes: B/P/L/C × seven probes × two rounds = 56 proposed exploratory executions. They have no expected verdict or native scoring metadata and cannot produce benchmark metrics, rank candidates or reward an optimizer. C is unscored only. Passing authored controls cannot create gold provenance; a scored comparison of the motivating mechanisms still needs independently anchored source cases.
Includes two independent unapplied patches, exact unchanged-source packets, inert probe/control definitions, current arXiv/CodeRabbit/Macroscope and promptfoo research, and a native integration handoff. Records family-split/cache limitations, the evidence required for a fifth lens, and the provenance correction caught by ship review. Durable notes use the decisions CLI.
Readiness: the native scored/unscored adapter is NOT implemented or validated. The old CLI commands cannot execute this mixed plan. This preparation neither establishes improvement nor authorizes measurement or production promotion. No optimizer or dependency is installed.
Validation: JSON, 15 original source hashes, 8 original row hashes, 7 unscored probe hashes, 7 source packets, both independent patch applications/resulting hashes, family membership and 17 JS syntax checks passed. No control program, benchmark, probe, census, optimizer or evaluation was executed. Normal ship gates review the preparation artifacts; embedded fixture-control strings must not be executed.
The user's explicit execution pause remains in force after merge. Future execution needs permission and validated native separation/parity. Source-backed controls must qualify scored evidence; direct-mechanism/context scoring additionally needs eligible source cases. Probe observations never enter scored denominators.
Shortcut: https://app.shortcut.com/benordlabs/story/2872
Epic: https://app.shortcut.com/benordlabs/epic/2830
The preparation story closes only after owner merge. sc-2832's historical experiment remains separate and blocked on its own prerequisites.