fix(dev): deduplicate unchanged watcher events by path signature (#122) - #329
Conversation
🦋 Changeset detectedLatest commit: d2c5ea1 The changes in this PR will be included in the next version bump. This PR includes changesets to release 1 package
Not sure what this means? Click here to learn what changesets are. Click here if you're a maintainer who wants to add another changeset to this PR |
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: d2c5ea11be
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| const defaultPathSignature = async (path: string): Promise<string | undefined> => { | ||
| try { | ||
| const source = await stat(path, { bigint: true }); | ||
| return `${source.dev}:${source.ino}:${source.size}:${source.mtimeNs}`; |
There was a problem hiding this comment.
Include permission changes in path signatures
After this watcher has reported a path once (for example, when a new non-executable claude.bin file is added), a subsequent POSIX chmod +x changes its mode/ctime but leaves every field in this signature unchanged, so the Chokidar event is discarded. This is a meaningful source change because config/normalize.ts reads the executable bits and the claude.bin.executable.required diagnostic explicitly directs users to repair the file with chmod +x; under agent-bundle dev, that repair therefore leaves the stale diagnostic/build in place until an unrelated change or manual rebuild. Include mode or ctime in the signature so permission-only repairs invalidate the build.
Useful? React with 👍 / 👎.
* fix(claude): validate numeric config bounds in artifacts (#287 r3911238592) * fix(claude): reject marketplace backslash traversal (#314 r3918535243) * fix(claude): validate authority-only archive hosts (#314 r3918535249) * fix(replay): preserve invocation provenance in renderers (#322 r3919341927) * fix(replay): derive workspace from cursor roots (#322 r3919341940) * fix(dev): invalidate watcher on chmod changes (#329 r3919499846) * chore: add framework review fixes changeset
…ples-real edits
Each source edit paired a file write with an immediate manual rebuild, so the
watcher's own rebuild of the same write raced it for a second epoch whose
timing depended on load; { retry: 2 } absorbed the fallout. Edits now go
through replaceWatchedSourceAndAwaitRebuild: one atomic replacement, then a
wait on the coordinator's published build attempt, so one edit is exactly one
build and the retries are gone.
Refs #122, #200, #329.
…ples-real edits
Each source edit paired a file write with an immediate manual rebuild, so the
watcher's own rebuild of the same write raced it for a second epoch whose
timing depended on load; { retry: 2 } absorbed the fallout. Edits now go
through replaceWatchedSourceAndAwaitRebuild: one atomic replacement, then a
wait on the coordinator's published build attempt, so one edit is exactly one
build and the retries are gone.
Refs #122, #200, #329.
…ples-real edits
Each source edit paired a file write with an immediate manual rebuild, so the
watcher's own rebuild of the same write raced it for a second epoch whose
timing depended on load; { retry: 2 } absorbed the fallout. Edits now go
through replaceWatchedSourceAndAwaitRebuild: one atomic replacement, then a
wait on the coordinator's published build attempt, so one edit is exactly one
build and the retries are gone.
Refs #122, #200, #329.
…der mounting and deterministic examples-real edits (#399) * feat(test): mount conventional providers in the agent-bundle/test harness The public harness (renderRoute, renderRouteEvents, invokeCli, in-memory MCP) now discovers and executes src/providers/* for every manifest-backed request scope exactly like the generated entries: same deterministic key order, same surface-specific invocation, same fail-closed messages, seeded processLifetime. context.providers still wins when passed. The execution contract the codegen emits and the harness runs now lives in one module and is pinned together. Refs #313, #366. * test(workbench): wait on the watcher rebuild instead of retrying examples-real edits Each source edit paired a file write with an immediate manual rebuild, so the watcher's own rebuild of the same write raced it for a second epoch whose timing depended on load; { retry: 2 } absorbed the fallout. Edits now go through replaceWatchedSourceAndAwaitRebuild: one atomic replacement, then a wait on the coordinator's published build attempt, so one edit is exactly one build and the retries are gone. Refs #122, #200, #329. * test(route-unit): reconcile the #371 provider seam pin with harness auto-mounting Explicit context.providers still mounts verbatim; a module rendered directly now observes the framework-owned processLifetime like a generated scope without providers, and manifest routes execute conventional providers. * test(workbench): replace the logs-real source edit atomically A truncating writeFile can split into two watcher invalidations under load, logging "Project source changed." twice and tripping the strict locator; the shared atomic replacement makes one edit one invalidation. * fix(test): derive executable surface for harness invocations and bump registry version renderRoute now hands providers and the request scope the surface the generated entries record — a routed CLI command's space-joined command path and a script's path-derived name — instead of the route id, so providers that branch on `command`/`name` behave identically in the harness and the artifact. The test registry version moves to 4 because the layout gained `providerLoaders`. The audiobook-curator degraded-catalog test opts out of the auto-mounted library provider explicitly. * fix(test): pin the tooling tool in the packed projection and ignore example artifacts The packed stdio projection lists the route-harness tools, which now include `tooling`. Every example's build writes `examples/<name>/artifact/`; the previous commit accidentally tracked three of them, so ignore the whole family and untrack the accidental copies. * fix(build): mount the compiled event route id as operationId in the Flight worker The generated worker resolved event routes by their hook identity and mounted that identity (`hook:event-route:tool-after`) as `invocation.operationId`, while the hook shell's request scope, the lifecycle replay, the test manifest, and `renderRoute` all use the compiled route id (`event:tool/after`). The worker record now carries the compiled id, so a route reading `invocation.operationId` sees one value on every surface; pinned by the warm-runtime integration test and the worker source digest. * fix(test): scope the harness process lifetime to each simulated executable A module-level lifetime made unrelated invokeCli calls, renders, and MCP sessions look like one warm process, so a provider branching on hits or instanceId could pass in the harness and fail in the artifact. The lifetime now lives on the logical executable exactly as the artifact scopes it: fresh per CLI invocation and per route-unit render, shared across the requests of one open in-memory MCP server, and fresh again for the open-call-close helper. * fix(test): snapshot the process hit count before awaiting provider loaders Concurrent requests on one in-memory MCP server could each increment the shared lifetime before the first request reached executeProviders, so every request observed the final count. Capture each request's hit right after the increment, as the generated worker does, and hand the snapshot to the shared execution helper. * fix(build,test): snapshot the process hit synchronously before state bindings Generated stateful scopes incremented processLifetime.hits, awaited state bindings, then snapshotted the value, so concurrent requests could observe the same count; the harness's in-memory server claimed its hit only after the bindings resolved. Both now claim and snapshot in one synchronous step before any await: the emitted scopes bind `processHit` at the increment, and the harness claims through claimProcessHit before requestBindings. The harness also looks up CLI command paths among authored commands only, since projected MCP commands carry their tool's route id and render through the tool branch. * fix(test): keep harness context optional under provider typegen and document module-evaluation scope Auto-mounting made #409's harness rule unreachable: once the generated augmentation declared provider keys, HarnessOptionsArguments and RenderRouteContextInit turned `options`/`context.providers` mandatory, so a typed project could never let the harness mount its real providers. The harness now keeps both optional (an explicit map must still carry every declared key; a direct runAgentRequest still requires providers) and framework-mode.md, entry-conventions.md, and the README describe the auto-mount contract instead of "the harness never executes provider modules". They also record what the harness does not simulate: provider modules are evaluated once per test worker, so module-level provider state is shared across simulated executables and is only proven cold by the proof levels that spawn the artifact. * test(typegen): pin that harness calls stay legal without context under provider typegen The #409 acceptance pinned `renderRoute(id)` as a compile error once the augmentation declares provider keys. With the harness mounting the project's providers itself that call is the artifact-faithful one, so it now typechecks clean alongside a call that passes only `input`, while a partial explicit fixture still fails on the missing key and a direct runAgentRequest still requires `providers`. * fix(test,changeset): scale the watcher e2e outer timeouts and rewrite the changeset as a release summary The three examples-real tests that wait on watcher rebuilds bounded their rebuild waits at 60s × timeScale but kept fixed 120s/150s outer timeouts, so in CI (timeScale 4) Rstest could end the test before its own readiness wait did. Their outer timeouts scale the same way now. The changeset is rewritten per AGENTS.md as an imperative user-facing summary naming the harness exports and the harness error, ending with the PR reference.
Summary
dev-watcher.test.tsflake tracked in flake: track remaining dev-watcher and MCP App gate failures #122: under full-pool load, chokidar delivers delayed duplicate events for an already-reported write; the debounce window has already flushed, so the duplicate mints a spurious extra invalidation (and a redundant dev rebuild). Confirmed with a deterministic regression test that failed pre-fix (2 invalidations for 1 write).ProjectWatchernow gates flushed paths on a per-path change signature (dev:ino:size:mtimeNsvia bigint stat; a deletion tombstone for missing paths). Unchanged paths are dropped at flush; an all-unchanged flush delivers no invalidation. Deletion tombstones keep duplicateunlinkevents inert while re-creations always report. Flushes are serialized so signature reads never interleave.{ retry: 2 }and its fixed 100ms settle sleep: the ignored-path proof is now event-driven (a sentinel write barriers the assertion), and the invalidation count is deterministic.examples-real.e2e.test.tsretries are kept: they guard a different race (first chokidar event vs. the immediate manual rebuild after each source write); comments now say exactly that.Evidence
dev-watcher.test.ts(2×10 serial loops) under load average 45-65, retries removed.check:local-ci --current-node-onlybecause its long leg name pushes AF_UNIX socket paths to 119 bytes > the 108-byte kernel cap on this many-worker machine; they pass 34/34 when run directly).overview.e2e(packageVersionexpectation),examples-contract(17→18 routes), andhooks— all reproduce identically on unmodifiedorigin/main(sibling-wave breakage, triaged in flake: track remaining dev-watcher and MCP App gate failures #122 wave notes), none touch this diff.Part of #122 (dev-watcher scope). The MCP App polling scope lands separately.