Skip to content

fix(daemon): stop the lifetime-lock probe from dropping a held lock - #2207

Open
AmirF194 wants to merge 1 commit into
DeusData:mainfrom
AmirF194:fix/lifetime-lock-probe-drops-held-lock
Open

AmirF194 wants to merge 1 commit into
DeusData:mainfrom
AmirF194:fix/lifetime-lock-probe-drops-held-lock

Conversation

@AmirF194

Copy link
Copy Markdown
Contributor

While chasing #2162 I found a real, separate bug in posix_lifetime_lock_probe (src/daemon/ipc.c): when the in-process registry says this process already holds the lifetime lock, it still opens a throwaway fd on the lock file just to close it again and return 1. fcntl(2) record locks are scoped to (process, inode), not (fd, inode), so that close silently drops the real lock too, even though the reservation's own fd stays open.

Two probes back to back are enough to lose it. I wrote a test that acquires the reservation, probes it twice, then forks a genuinely separate process that tries to take the same lock directly. On main it succeeds, the lock is gone. With the fix it fails as expected.

Fix: check the in-process claim before opening anything and return early. That branch only ever needed to trust the registry, not touch the file.

scripts/test.sh --suites daemon_ipc: the new test fails on main (child_result == 1), 50 passed / 1 skipped with the fix.

Not a fix for #2162 itself, I ran into this while reading that code, it's a separate bug.

posix_lifetime_lock_probe opened a throwaway fd on the lock file to
confirm an already-claimed lock, then closed it. fcntl(2) record locks
are scoped to (process, inode), so that close silently released the
real lock too, even though the reservation's own fd stayed open.

Check the in-process claim before opening anything and return early;
that branch only ever needed to trust the registry.

Signed-off-by: Amir Fathi <amirfathi.me@gmail.com>
@AmirF194
AmirF194 requested a review from DeusData as a code owner September 14, 2026 14:10
@github-actions

Copy link
Copy Markdown

Thanks for opening this — it has been seen, and it is queued.

This note is automated, but it is not a brush-off: it exists so you know where your PR stands instead of having to guess from silence.

Current review status: working through a backlog. 0.9.1-rc.1 is out, so the release freeze that held reviews is over — but it left a large queue of open pull requests behind it, and we are reading through them oldest-first. The background is in discussion #1144.

What that means for this PR, concretely:

  • It will not be closed for inactivity. No stale bot touches pull requests here.
  • It may still sit a while before a human reads it. That is on us, not on you.
  • Older PRs are read first, so a recent one is not being skipped — it is behind a queue.

Things that will genuinely speed it up whenever review does happen:

  • Keep it rebased on main — the tree is moving quickly right now, and a conflicting branch cannot be reviewed as the diff you intended.
  • Get CI green, or say which failures you believe are pre-existing.
  • Keep the change to one claim. Bundled features and refactors get split before they get merged, which costs you a round trip.
  • Every commit needs a sign-off (git commit -s) — CI enforces DCO.

If this fixes a bug, a reproduction we can run is worth more than a description of the symptom.

Thanks for contributing, and sorry in advance for the wait.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant