Skip to content

[ROUTE6-1:T3] Durable model-canary attribution across resumes and model switches - #6

Merged
nullStack65 merged 3 commits into
mainfrom
feat/durable-model-canary-attribution-20260923
Sep 28, 2026
Merged

nullStack65 merged 3 commits into
mainfrom
feat/durable-model-canary-attribution-20260923

Conversation

@nullStack65

@nullStack65 nullStack65 commented Sep 23, 2026 •

Copy link
Copy Markdown
Owner

What this is

Narrow follow-on to #4, restacked onto its current branch head
f863939bcd52b49f238982f593ea34c146f80e28. It does not
change #4; it adds the minimum durable identity/history and pre-execution
experiment metadata that #4 explicitly deferred, so the ROUTE6-1 coding-model
canary can carry and measure metadata across resumes, provider-session
replacement, model switches, child sessions, retries/fallbacks, multiple PRs,
and manager/agent assignments.

Draft; not for merge. No endpoint, UI, PostHog, collector, gateway, routing, or
transport change. Upstream pingdotgg/t3code is untouched. T3 is not the
canonical model router — agent-config remains policy authority; T3 only carries
and measures metadata.

Round-5 repair (R5-T3) — resolves ROUTE6-1:R4-REV-T3 CHANGES_REQUIRED

Independent review ROUTE6-1:R4-REV-T3 returned CHANGES_REQUIRED on the
reviewed head 4d5671597 (base af7049544). #4 has since advanced to
f863939bc, so this session restacked 4d5671597 → 3a1047a16 onto the current
#4 head first (git rebase --onto f863939 af7049544, clean, docs auto-merged),
then added one repair commit e0e6318a8. Repair delta: 11 files,
+1206/−133
.

Finding Disposition
P1-1 multi-table upsert not transactional ProviderSessionRuntime.upsert now wraps runtime cursor + provider_session_history + thread_route_events in one sql.withTransaction; all commit or none. onConflict: "ignore" stays a single atomic statement and still appends nothing.
P1-2 history identity ignores provider instance Unique identity is now (thread_id, provider_name, provider_instance_key, native_session_id), with a normalized non-null provider_instance_key (trimmed instance id, "" for unknown/null) so SQLite NULL-distinctness cannot collapse instances. Route session reports expose providerInstanceIds.
P1-3 054 backfill not faithful to nativeSessionIdOf Backfill accepts only JSON text values, trimmed non-empty, precedence resume → threadId → sessionId with fallthrough on absent/null/non-string/blank; numbers/booleans/objects never become ids.
P2-4 ambiguous multi-thread session double count Route sessions now carry base allocation + boundThreadIds; an ambiguous session is not summed into any candidate thread and is pooled once in the new route-view unallocated bucket. Per-thread totals + unallocated reconcile additively to each session's usage exactly once.
P2-5 automatic request first-write-wins Request record uses an explicit enrichment UPSERT: null→non-null enriches; equal→equal idempotent; non-null→null retains; conflicting non-null→different non-null retains the first value and sets selection_conflict.
P2-6 production wiring gaps Documentation/PR wording only: write seam + extraction seam exist; only the automatic request record has a production producer; declared kinds and parent_native_session_id remain future agent-config/harness wiring and are not claimed as produced.
P2-7 history input not de-duplicated extractAttributionHistory de-duplicates by canonical identity, merges monotonically, and surfaces a conflicting duplicate.
P2-8 escalation reason free text escalation_reason is bounded to a lowercase code/slug (max 64, allowlisted charset); anything else is rejected at the writer and extractor.
Adversarial gaps Added atomicity failure-injection, instance identity/collapse, request enrichment (4 cases), cross-thread declared id, reordered upsert, malformed cursor, migration rerun, malformed requested route, and ambiguous reconciliation tests.

Round-3 restack (R3-T3)

Restacked with git rebase --onto af7049544 daa55eb88; later replayed onto
f863939 in R5-T3. The only shared path, docs/internals/usage-attribution.md,
auto-merged with all Round-3/M3D cache-quality content preserved. The M3D
parser/cache files (usageTranscripts.ts, usageScanCache.ts, UsageService.ts
and their tests) remain byte-identical to the current #4 head, so nothing here
reverts zero-total retention, invalid/partial classification, predecessor-v4
cache freshness, or the warm-cache quality gate. Migrations 054/055 are still
free on the current #4 head (highest is 053).

Why #4 needed this

#4's own limitation: a native session maps to a thread only through the single current provider_session_runtime.resume_cursor_json. A resume, fork, or model switch overwrites it and earlier usage becomes unbound. #4 recorded that as the follow-on it did not adopt.

Schema decision (migration only where required)

A migration is used because the existing schema cannot truthfully preserve
history — the cursor is one row per thread. Two additive, forward-compatible
tables (existing databases are safe; existing thread/session behavior is
unchanged). Ids 054/055 were confirmed free on the current #4 base (its
highest existing migration is 053). No collision, no renumbering.

  • 054_provider_session_history — append-only identity,
    UNIQUE (thread_id, provider_name, provider_instance_key, native_session_id).
    ProviderSessionRuntime.upsert appends one row per distinct identity and only
    advances last_seen_at on a repeat. The backfill mirrors nativeSessionIdOf
    semantics exactly. A conflicting onConflict: "ignore" write appends nothing
    (that cursor was never applied). parent_native_session_id is structurally
    supported; no production caller populates it yet.
  • 055_thread_route_events — append-only, content-free pre-execution
    route/experiment metadata. Nullable route_event_kind keeps an automatic
    "requested" record distinct from a declared
    canary/fallback/review/escalation event. The automatic record is monotonically
    enrichable; selection_conflict surfaces a contradiction. escalation_reason
    is a bounded slug.

Changed paths

File Change
apps/server/src/persistence/Migrations/054_ProviderSessionHistory.ts new: history table (instance-aware unique key) + index + faithful current-cursor backfill
apps/server/src/persistence/Migrations/055_ThreadRouteEvents.ts new: route/experiment event table + index + selection_conflict
apps/server/src/persistence/Migrations.ts register 54, 55
apps/server/src/persistence/ProviderSessionRuntime.ts transactional upsert; history + enriching route events
apps/server/src/usage/routeMetadata.ts pure route/experiment types + parsers + instance-key/reason guards
apps/server/src/usage/usageAttributionSources.ts extract history (de-duped) + route events; snapshot merges them
apps/server/src/usage/usageAttribution.ts add sessionHistory binding origin
apps/server/src/usage/usageRouteAttribution.ts pure identity/route/experiment projection; carries allocation/ambiguity + provider-instance identity; additive unallocated reconciliation
apps/server/src/provider/Services/ProviderSessionDirectory.ts, Layers/ProviderSessionDirectory.ts carry parentNativeSessionId, requestedRoute, routeEvent on the binding
apps/server/src/provider/Layers/ProviderService.ts record the requested route at session start (from the model selection T3 was given)
tests ProviderSessionRuntime.history.test.ts (atomicity/instance/enrichment), 054_ProviderSessionHistory.test.ts (backfill/rerun), usageAttributionSources.test.ts, usageRouteAttribution.test.ts (incl. #4 quality regressions A–D + ambiguity)
docs/internals/usage-attribution.md durable-history and requested-vs-observed sections

Semantics encoded

  • Identity — T3 thread id, native provider session id, configured provider
    instance, child/parent lineage, and the existing repository-qualified PR link
    table (projection_thread_pull_requests, reused, not replaced). A readable
    manager/agent id (ROUTE6-1 / T3) is a label, never a substitute for the
    native session id and never a join key.
  • Requested vs observed — requested provider/model/effort are captured at session start from what T3 was actually asked to run; the observed model comes from measured usage only. Never copied into each other. No scanned source exposes effort, so actualEffort is null with quality unsupported.
  • Measurement/identity quality — each route session carries the [METRICS-M3] Prove prompt request session and PR usage attribution #4 axes
    (measurement, identity, prompt, request, recordIdentity, conflict)
    or null when no usage was measured, so a partial/invalid measurement or a
    legacy identity-erased row is never presented as exact.
  • Allocation — a session bound to more than one thread is ambiguous; its
    usage is not duplicated onto candidate threads and is pooled once in the route
    view's unallocated bucket.
  • Experiment — task stratum, experiment/cohort id, route event kind (normal | availability_fallback | canary | independent_review | quality_escalation), and bounded escalation/fallback reason. Unknown is allowed and is the default.
  • History — every contributing native session survives a cursor change, so a
    thread is never attributed wholly to the initial or most recent model.
  • Privacy — metadata only: no prompts, responses, code, tool-call bodies, secrets, or arbitrary user content.

Validation

# 17 files, 362 passed (baseline at restacked head: 342; +20 repair tests)
vp test run apps/server/src/usage/ \
  apps/server/src/persistence/ProviderSessionRuntime.history.test.ts \
  apps/server/src/persistence/Migrations/054_ProviderSessionHistory.test.ts \
  apps/server/src/persistence/RepositoryErrorCorrelation.test.ts \
  apps/server/src/provider/Layers/ProviderSessionDirectory.test.ts \
  apps/server/src/provider/Layers/ProviderSessionReaper.test.ts \
  apps/server/src/provider/Layers/ProviderService.test.ts \
  apps/server/src/project/AgentSessionImporter.test.ts
vp run --filter t3 typecheck        # exit 0 (effect suggestions only, no errors)
vp lint <changed files>             # exit 0
vp fmt --check <changed files>      # clean
git diff --check                    # clean

Exact per-test receipts and CI state are in the RESULT comment. CI is not
green: the self-hosted blacksmith-* jobs stay queued on the fork.

Model/harness: opencode-go / deepseek-v4.1-flash via opencode (see RESULT).

@github-actions github-actions Bot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:XXL labels Sep 23, 2026
@nullStack65

Copy link
Copy Markdown
Owner Author

RESULT ROUTE6-1:T3 — durable model-canary attribution

Status: complete (source + migrations + focused tests + docs; stacked draft PR). CI cannot be observed green on the fork (see Checks).

Relationship to t3code#4

Stacked on #4's branch head, not a rewrite. #4 remains the implementation line for the attribution proof; this PR is the narrow follow-on #4 itself named and deferred ("durable per-thread session history … specified but not adopted"). Committed diff is only the new commits on top of e4f36af5e; git diff e4f36af5e..97b4ceb5d is 16 files, +2231/−17. #4's branch was not force-pushed and its reviewed head is unchanged.

Schema / storage decision

A migration was used because existing schema cannot truthfully preserve the required history: provider_session_runtime is PRIMARY KEY (thread_id) and holds one current cursor. Two additive, forward-compatible tables; existing databases are safe and existing thread/session behavior is unchanged:

  • 054_ProviderSessionHistory — provider_session_history: UNIQUE (thread_id, provider_name, native_session_id), columns parent_native_session_id, origin, first_seen_at, last_seen_at. Backfills the current cursor on upgrade. Index on (thread_id, first_seen_at).
  • 055_ThreadRouteEvents — thread_route_events: append-only, content-free; nullable route_event_kind; task_stratum, experiment_id, manager_id, agent_id, requested_{provider,model,effort}, escalation_reason. Index on (thread_id, recorded_at).

Both registered in Migrations.ts (ids 54, 55); a migration test asserts the backfill and manifest registration.

Durable history semantics

  • ProviderSessionRuntime.upsert now appends one provider_session_history row per distinct native session on the update path, in the same logical write as the runtime row. A repeat of the same session only advances last_seen_at (CASE WHEN excluded > existing), never replacing a different session.
  • A conflicting onConflict: "ignore" write appends nothing (that cursor was never applied), so no invented history.
  • usageAttributionSources.extractAttributionHistory reads it into sessionHistory bindings, and extractAttributionSnapshot merges cursor + history + imported transcripts. A resume/model switch therefore leaves earlier sessions bound, so a thread is never attributed wholly to its initial or most recent model.
  • parent_native_session_id carries a true sub-agent parent when a caller knows it. OpenCode child ids are still not persisted by T3 (they exist only in the adapter's in-memory relatedSessionIds and native event log), so an OpenCode child is preserved as a distinct identity with usage: null — unknown, not flattened into the parent, not zero.

Experiment metadata semantics

  • thread_route_events is written through the existing session-binding seam (ProviderRuntimeBinding.routeEvent / requestedRoute → ProviderSessionRuntime.upsert options). No new layer, endpoint, or transport.
  • route_event_kind ∈ normal | availability_fallback | canary | independent_review | quality_escalation, nullable. kind === null marks an automatic "what was requested" record, kept distinct from a declared canary/fallback/review/escalation event.
  • Task stratum is a bounded set with unknown as a first-class, non-guessed value (investigation | docs | tests | simple_edit | implementation | review | ci_repair | architecture | security | unknown).
  • T3 carries the metadata only. It does not decide "DeepSeek is default / Luna is challenger / Sol is escalation" — that stays in agent-config policy. The declared event is authored by the route authority.

Requested vs actual model handling

  • Requested provider/model/effort are captured at session start from the model selection T3 was actually given (ProviderService.upsertSessionBinding → requestedRouteOf). Requested effort reads reasoningEffort then effort.
  • The observed model is read from measured usage records. The two are never copied into each other.
  • No scanned source exposes reasoning effort, so actualEffort is always null with quality unsupported; it is never derived from the requested effort.
  • A readable manager/agent id is a label, asserted distinct from the native session id and never used as a join key.

Request-identity rule preserved from #4

Unsupported request/prompt identity stays null/unsupported; aggregate turns are not split into imaginary requests. This PR adds no request inference.

Privacy boundary

Metadata only. No raw prompts, responses, source code, tool-call bodies, secrets, or arbitrary user content are stored or logged anywhere in this path.

Tests (17 files, 285 passed)

vp test apps/server/src/usage/ apps/server/src/persistence/ProviderSessionRuntime.history.test.ts apps/server/src/persistence/Migrations/054_ProviderSessionHistory.test.ts apps/server/src/persistence/RepositoryErrorCorrelation.test.ts apps/server/src/provider/Layers/ProviderSessionDirectory.test.ts apps/server/src/provider/Layers/ProviderSessionReaper.test.ts apps/server/src/provider/Layers/ProviderService.test.ts apps/server/src/project/AgentSessionImporter.test.ts
# Test Files 17 passed (17); Tests 285 passed (285)

Covers the twelve required semantics: (1) one thread DeepSeek→Luna keeps both; (2) resume into a new provider session keeps the earlier one; (3) child OpenCode identity distinct, usage: null, not flattened; (4) manager/agent id distinct from the session id; (5) one thread on multiple PRs sums each session once (asserts shared, no duplication); (6) requested ≠ observed; (7) actual effort null/unsupported; (8) availability fallback vs independent review distinct with separate reasons; (9) quality-escalation reason preserved; (10) historical contribution survives a cursor change; (11) stack-dismissed PR link creates no attribution; (12) missing usage stays null, not zero. Plus persistence-level tests for append-on-resume, last-seen advancement, no-append-on-ignore, idempotent route events, and the migration backfill.

Typecheck / lint / format

vp run --filter t3 typecheck        # no errors
vp lint <changed files>             # exit 0
vp fmt --check <changed files>      # clean

CI

Actual route used for THIS agent session

opencode-go / deepseek-v4.1-flash via opencode (the harness this session runs in). Gemini 3.8 Flash medium was the preferred independent-family canary but is not available through a current qualified route from here, and I did not mutate live routing to obtain it. Observed effort: unknown (not exposed).

Blockers / unknowns

  • Live child-session persistence: OpenCode child ids are not yet written to T3 storage by any adapter; this PR makes the projection ready (parent field + distinct identity) but wiring an adapter to persist childSessionId is a separate, larger change.
  • Observed effort is unsupported across all scanned sources; the requested effort is captured but actual effort remains unknown.
  • Experiment metadata has a real, tested write seam and durable storage, but no production caller populates the declared (kind != null) event yet; that authority is agent-config/Skill-owned. Only the automatic requested-route record is written live today.
  • CI is unpushable to green on the fork.

Exact next action

Request review of (a) the two append-only tables and the upsert-time write path, (b) the requested-vs-observed separation, and (c) the companion usageRouteAttribution projection. On approval, the follow-ons are: an adapter change to persist OpenCode child session ids, and the agent-config/Skill side that supplies declared route events. Do not merge.

@nullStack65

Copy link
Copy Markdown
Owner Author

ROUTE6-1:T3 (#6) — corrected #4 head to rebase onto

This follow-on is stacked on #4's old head e4f36af5e. #4 now has a corrected head after the METRICS-1:M3C round-3 repairs:

Compatibility notes for your rebase (your migrations 054/055 and durable-history work were not touched here):

  • UsageRecord / AttributionUsageRecord gained optional measurementCompleteness, invalidTokenFields, identityAvailable, and dedupeKeyScope. Existing fields are unchanged; the scan-cache v4 row format appended three columns after index 14, and measurement gained the invalid code (appended last, so older codes keep their values).
  • The projection now treats a legacy identityAvailable:false record as unavailable for request/prompt levels and as only-partial for measurement completeness.
  • readResumeCursor now prefers resume over threadId for a Claude cursor that carries both (imported-session shape); your usageAttributionSources extensions should assume a cursor may carry more than one field.
  • Warm scan-cache acceptance now requires identity === "declared", so an unchanged extant legacy file is cold re-parsed once.

Please rebase after review of #4. No merge.

@nullStack65

Copy link
Copy Markdown
Owner Author

START ROUTE6-1:R2-T3 - rebase durable attribution onto corrected #4 head

@nullStack65
nullStack65 force-pushed the feat/durable-model-canary-attribution-20260923 branch from 97b4ceb to 67b4c9c Compare September 24, 2026 00:35
@nullStack65

Copy link
Copy Markdown
Owner Author

ROUTE6-1:R2-T3 — RESULT

Status: complete (restacked source + migrations + focused tests + docs; draft stacked PR). CI cannot be observed green on the fork — see CI state. Fresh session, not a continuation of the Round-1 T3 session. No merge.

Heads

Rebase / restack method

  • Backup refs first: branch backup/r2t3-pre-rebase-97b4ceb5d and tag r2t3-pre-rebase-97b4ceb5d at the old head.
  • git rebase --onto daa55eb88966665faa8c1770c35fc3c53a68f6fa e4f36af5ef279246bcb0f8463adeee8a09b7bde1 feat/durable-model-canary-attribution-20260923.
  • One [ROUTE6-1:T3] Durable model-canary attribution across resumes and model switches #6 commit (97b4ceb5d → 633528bd8) replayed; then git commit --amend folded in the audit reconciliation (route-view quality axes + regression), giving final head 67b4c9c48.
  • Pushed with --force-with-lease=refs/heads/feat/durable-model-canary-attribution-20260923:97b4ceb5d… (expected old head confirmed by ls-remote immediately before). Single commit on top of corrected [METRICS-M3] Prove prompt request session and PR usage attribution #4.

Conflicts encountered

Exactly one conflicted path: apps/server/src/usage/usageAttributionSources.test.ts. Both #4 and #6 had appended a test to the same describe("extractAttributionSnapshot") block, and #6 changes the snapshot shape.

Resolution: kept both tests (the #4 M1C interchange fixture test and the #6 durable-history/route-event merge test). Because #6 adds history/routeEvents to ExtractedSnapshotExtraction, the #4 fixture's toEqual expectation was extended with the additive history: [], routeEvents: [], and diagnostics.history / diagnostics.routeEvents blocks — the fixture's inputs are unchanged. The three other overlapping paths auto-merged: usageAttribution.ts, usageAttributionSources.ts, docs/internals/usage-attribution.md. No conflict was resolved by taking stale #6 copies.

How each #4 correction was preserved

The #6 commit's deletion set is only its intended replacements (verified line-by-line):
usageTranscripts.ts, usageScanCache.ts, and UsageService.ts — the files that carry the Round-3 correctness work — are not touched by this PR at all, so the following are byte-for-byte preserved from corrected #4:

#4 correction Preserved how
measurementCompleteness, invalidTokenFields, identityAvailable, dedupeKeyScope on AttributionUsageRecord type remains in #4 head; #6 does not edit that block
measurement: "invalid" code in usageTranscripts.ts (#4); untouched
scan-cache v4 row format appended fields in usageScanCache.ts (#4); untouched
legacy identityAvailable:false → requestQuality:"unavailable" in usageAttribution.ts (#4 identityErased); #6 only adds the sessionHistory origin to the adjacent AttributionThreadBinding union
readResumeCursor prefers resume over threadId in usageAttributionSources.ts (#4 doc/order); #6 adds new functions around it, does not change it. The 054 backfill SQL COALESCE(resume, threadId, sessionId) matches the same order
warm legacy-cache enrichment (identity === "declared") in UsageService.ts (#4); untouched
content equality ≠ event identity; dedupeIdentity in usageAttribution.ts (#4); #6 does not touch the dedupe path
cost/provenance conflict behavior in observationContent (#4); untouched
snapshot replacement semantics in usageAttribution.ts (#4); untouched
orphan / reconciliation identity base projection (#4); #6 composes buildUsageAttribution and does not alter it

Round-2 reconciliation added: SessionRouteQuality (measurement, identity, prompt, request, recordIdentity, conflict) on SessionRouteReport, null when no usage was measured. This closes the one place the follow-on view could have dropped the #4 quality axes; the base is embedded in base.sessions and now mirrored per route session.

Migration id check

On corrected #4 (daa55eb88) the highest registered migration is 053 (PullRequestFilesViewed); there is no 054/055. No collision. #6 registers [54, "ProviderSessionHistory"] and [55, "ThreadRouteEvents"] unchanged. No renumbering was needed.

#6 intended additions preserved

  • provider_session_history — append-only native-session history (migration 054, current-cursor backfill at upgrade); upsert appends on the update path and a conflicting onConflict:"ignore" write appends nothing (no invented session). Earlier sessions survive resume/model switch.
  • thread_route_events — content-free pre-execution metadata (migration 055): normal | availability_fallback | canary | independent_review | quality_escalation, nullable kind.
  • requested-vs-observed split — requested provider/model/effort captured at session start; observed model only from measured usage; actualEffort null/unsupported.
  • manager/agent ids — labels only, never join keys (asserted distinct from the session id).
  • PR relation — reuses projection_thread_pull_requests; association ≠ allocation.
  • Privacy — metadata only; no prompts/responses/code/tool bodies.

Changed paths (all 16)

apps/server/src/persistence/{Migrations.ts, Migrations/054_ProviderSessionHistory.ts, Migrations/054_ProviderSessionHistory.test.ts, Migrations/055_ThreadRouteEvents.ts, ProviderSessionRuntime.ts, ProviderSessionRuntime.history.test.ts}, apps/server/src/provider/{Services/ProviderSessionDirectory.ts, Layers/ProviderSessionDirectory.ts, Layers/ProviderService.ts}, apps/server/src/usage/{routeMetadata.ts, usageAttribution.ts, usageAttributionSources.ts, usageAttributionSources.test.ts, usageRouteAttribution.ts, usageRouteAttribution.test.ts}, docs/internals/usage-attribution.md.

corrected #4..new #6 diff

git diff daa55eb88..67b4c9c48 → 16 files, +2412 / −17, 1 commit. Every removed line is an intended #6 replacement (the upsert ternary, the modelSelection?: unknown annotation, options, the two-line binding-origin union, the two limitations strings, and the snapshot return / docs paragraph). Nothing outside the follow-on scope changed.

Tests (17 files, 312 passed)

vp test run \
  apps/server/src/usage/ \
  apps/server/src/persistence/ProviderSessionRuntime.history.test.ts \
  apps/server/src/persistence/Migrations/054_ProviderSessionHistory.test.ts \
  apps/server/src/persistence/RepositoryErrorCorrelation.test.ts \
  apps/server/src/provider/Layers/ProviderSessionDirectory.test.ts \
  apps/server/src/provider/Layers/ProviderSessionReaper.test.ts \
  apps/server/src/provider/Layers/ProviderService.test.ts \
  apps/server/src/project/AgentSessionImporter.test.ts
# Test Files 17 passed (17); Tests 312 passed (312)

Added regression (usageRouteAttribution.test.ts, "durable history keeps #4 measurement and identity quality"): a durable session with measurementCompleteness:"partial" reports quality.measurement === "partial" (not measured); a legacy identityAvailable:false claude row reports quality.request === "unavailable" (not missing); measurement:"invalid" stays invalid; a cost-only duplicate reports quality.conflict === true; a history-only OpenCode identity has usage === null and quality === null (no fabricated measured quality). The three #6 test counts previously reporting 285 now report 312 because the corrected #4 base contributes its own Round-3 tests.

Typecheck / lint / format

vp run --filter t3 typecheck   # exit 0 (effect lint *suggestions* only, no errors)
vp lint <16 changed files>     # exit 0
vp fmt --check <16 changed files>  # "All matched files use the correct format." exit 0

CI state (truthful)

Not green. At head 67b4c9c48 the self-hosted jobs Test, Test Server 1/2/3, Rust, Release Smoke, Mobile Native Changes, Check, Native fingerprint diff remain pending (queued) — the fork has no blacksmith-* self-hosted runners. GitHub-hosted jobs passed (Label PR 6, Label PR size, Prepare PR size config, Collect PR targets); Deploy web preview / EAS Preview / Sync PR size label definitions skipped; CodeRabbit skipped (draft). No green CI is claimed; the queued state is the same runner-capacity gap #4 already reported.

Actual agent route

opencode-go / deepseek-v4.1-flash via opencode (the harness this session runs in). Observed effort: unknown — no scanned source exposes it. Gemini 3.8 Flash was not obtained; routing was not mutated. (Also: in this environment tests must run with the repo-local vp; the prebuilt global vp resolves the wrong vite-plus/test binding and fails collection.)

Remaining gaps

Exact next action

Request independent exact-head review of #6 @ 67b4c9c48e5def2f9bb97547ecb893d794d11d2a against base daa55eb88…: (a) the two append-only tables and the upsert-time append path, (b) requested-vs-observed separation, (c) the usageRouteAttribution view and the new SessionRouteQuality axes. Do not merge #4 or #6. Merge ordering is a hub decision after #4 is reviewed and Round-2 integration ordering is set.

@nullStack65

Copy link
Copy Markdown
Owner Author

METRICS-1:M3D dependency update for this history owner

T3 #4 (the usage attribution/measurement PR this draft depends on) moved from daa55eb88966665faa8c1770c35fc3c53a68f6fa to head af70495442678a06e007d13bba477f123170c0f5 by fast-forward. This PR's durable migrations (054/055) were not touched or rebased here.

Two changes are relevant when you eventually rebase onto the new head:

  1. usageScanCache.ts: CachedFile gained qualityMetadata: "declared" | "predecessor". Warm-cache acceptance and incremental resume now require identity === "declared" and qualityMetadata === "declared". Any scan-cache construction you carry over must set it (a freshly parsed entry sets "declared"); a decoded predecessor 15-field v4 row is "predecessor" and its completeness decodes as partial.
  2. usageTranscripts.ts: Codex and Grok now emit eligible zero-total events (complete zero, known-zero subset, all-invalid) instead of dropping them; only a container with no recognised token field is no-usage.

RESULT with before/after receipts: #4 (comment)

@nullStack65
nullStack65 force-pushed the feat/durable-model-canary-attribution-20260923 branch from 67b4c9c to 4d56715 Compare September 25, 2026 22:44
@nullStack65

Copy link
Copy Markdown
Owner Author

ROUTE6-1:R3-T3 — RESULT

Status: complete (restacked source + migrations + focused tests + docs; draft stacked PR). Fresh session; not a continuation of R2-T3. No merge. CI cannot be observed green on the fork — see CI state.

Heads

Ref At start Actually used / final
#4 head af70495442678a06e007d13bba477f123170c0f5 af70495442678a06e007d13bba477f123170c0f5 (unchanged; did not move beyond the manager-observed head)
#6 old stacked base daa55eb88966665faa8c1770c35fc3c53a68f6fa —
#6 old head 67b4c9c48e5def2f9bb97547ecb893d794d11d2a —
#6 new base — af70495442678a06e007d13bba477f123170c0f5
#6 new head — 4d56715971dde1c60ee67e8c3d322812faa575e9
t3code fork main bcc1a58b19a9d610a4f08fed191a364767bc65b3 bcc1a58b19a9d610a4f08fed191a364767bc65b3

PR #6 base branch feat/usage-attribution-proof-20260923 is at af7049544…, so the PR now shows base.sha = af7049544… and MERGEABLE.

Restack method

  • Fresh isolated worktree off the canonical clone: /Users/businessaccount/w-t3/r3-t3, uniquely named local branch r3-t3-restack-20260924, created from the fetched remote [ROUTE6-1:T3] Durable model-canary attribution across resumes and model switches #6 head 67b4c9c48.
  • Backup refs first: branch backup/r3t3-pre-restack-67b4c9c48 and tag r3t3-pre-restack-67b4c9c48; tag r3t3-old-base-daa55eb88 and branch backup/r3t3-pr4-old-base-daa55eb88.
  • git rebase --onto af70495442678a06e007d13bba477f123170c0f5 daa55eb88966665faa8c1770c35fc3c53a68f6fa → the single [ROUTE6-1:T3] Durable model-canary attribution across resumes and model switches #6 commit replayed cleanly with no conflicts, producing 344be10e7.
  • A second commit 4d5671597 (test(usage): prove the route view preserves M3D zero/invalid/cache quality) adds the requested M3D-preservation regressions. Final head 4d5671597, two commits on top of af7049544.
  • Pushed with an explicit lease: git push --force-with-lease=refs/heads/feat/durable-model-canary-attribution-20260923:67b4c9c48e5def2f9bb97547ecb893d794d11d2a origin r3-t3-restack-20260924:refs/heads/feat/durable-model-canary-attribution-20260923. The expected old head was re-confirmed by git ls-remote immediately before the push; no blind force push, no shared-branch reset, no duplicate PR.

Conflicts

None. The only path changed by both #4 (M3D) and #6 is docs/internals/usage-attribution.md; git auto-merged it in disjoint hunks, and the final docs contain both the M3D cache-quality/zero-total text and the #6 durable-history/requested-vs-observed sections. No conflict was resolved by taking stale #6 copies.

Migration ID check

On current #4 (af7049544) the highest registered migration is 053 (PullRequestFilesViewed); there is no 054/055. No collision. No renumbering. #6 registers [54, "ProviderSessionHistory"] and [55, "ThreadRouteEvents"]; files 054_ProviderSessionHistory.ts/.test.ts and 055_ThreadRouteEvents.ts are present and registered.

M3D semantics preservation table

The #6 delta touches only 16 files; the M3D-owned parser/cache/service files are byte-identical to af7049544 (verified: git diff af7049544..4d5671597 -- apps/server/src/usage/{usageTranscripts.ts,usageTranscripts.test.ts,usageScanCache.ts,usageScanCache.test.ts,UsageService.ts,UsageService.test.ts,usageAttribution.test.ts} is empty).

M3D behavior Preserved how Route-view regression
Codex/Grok no longer discard eligible zero-total observations usageTranscripts.ts untouched A (route view keeps a complete measured zero as measured, total 0, not missing)
Measured zero stays measured zero recordQuality in usageAttribution.ts untouched A
Invalid fields stay invalid; partial stays partial classification untouched B (invalid) and existing partial test
all-invalid eligible observation reaches malformedRecords UsageService.ts untouched B (all-invalid stays invalid, not missing, not measured)
predecessor-v4 missing completeness decodes as partial, never complete usageScanCache.ts untouched; qualityMetadata: "predecessor" C (decode predecessor + partial, route view stays partial)
Warm acceptance requires identity and quality metadata UsageService.ts untouched covered by existing UsageService.test.ts
legacy identityAvailable:false → request unavailable, never missing usageAttribution.ts identityErased untouched existing route test
dedupeKeyScope explicit; global vs source-local dedupe path unchanged existing base tests
cost-only conflict is a conflict, not a silent dedupe observationContent unchanged E — existing route test asserts quality.conflict === true
history-only identity with no usage is usage=null/quality=null route view step 1 D — new test (measured provider and unmeasured provider)
readResumeCursor prefers resume over threadId unchanged; 054 backfill COALESCE(resume, threadId, sessionId) matches the same order —

Regressions added / updated

apps/server/src/usage/usageRouteAttribution.test.ts (512 → 700 lines), four new cases:

  • A keeps an explicit measured zero as a measured zero, not missing
  • B keeps an all-invalid eligible observation invalid, not missing or zero
  • C does not promote predecessor-quality partial cache evidence to complete — decodes a 15-field predecessor v4 document via decodeScanCache, asserts qualityMetadata: "predecessor" and measurementCompleteness: "partial", then feeds the decoded records through buildUsageRouteAttribution and asserts quality.measurement === "partial" (never measured).
  • D keeps a history-only measured-provider identity at usage null / quality null

E and the legacy-identity-unavailable case were already present (surfaces a cost-only conflict instead of silently deduping, keeps a legacy identity-erased durable row unavailable, not missing).

Final #4 → #6 diff stats / changed paths

git diff af7049544..4d5671597 → 16 files, +2600/−17, 2 commits. The test file grew from 512 to 700 lines (+188). Every removed line (−17) is an intended #6 replacement (the upsert ternary, the modelSelection?: unknown annotation, options, the binding-origin union/doc, the two older limitation strings, the snapshot return, and the superseded docs paragraphs). No M3D-owned line is removed.

Changed paths (16):

  • apps/server/src/persistence/Migrations.ts
  • apps/server/src/persistence/Migrations/054_ProviderSessionHistory.ts
  • apps/server/src/persistence/Migrations/054_ProviderSessionHistory.test.ts
  • apps/server/src/persistence/Migrations/055_ThreadRouteEvents.ts
  • apps/server/src/persistence/ProviderSessionRuntime.ts
  • apps/server/src/persistence/ProviderSessionRuntime.history.test.ts
  • apps/server/src/provider/Layers/ProviderService.ts
  • apps/server/src/provider/Layers/ProviderSessionDirectory.ts
  • apps/server/src/provider/Services/ProviderSessionDirectory.ts
  • apps/server/src/usage/routeMetadata.ts
  • apps/server/src/usage/usageAttribution.ts
  • apps/server/src/usage/usageAttributionSources.ts
  • apps/server/src/usage/usageAttributionSources.test.ts
  • apps/server/src/usage/usageRouteAttribution.ts
  • apps/server/src/usage/usageRouteAttribution.test.ts
  • docs/internals/usage-attribution.md

Tests / validation

vp test run apps/server/src/usage/ \
  apps/server/src/persistence/ProviderSessionRuntime.history.test.ts \
  apps/server/src/persistence/Migrations/054_ProviderSessionHistory.test.ts \
  apps/server/src/persistence/RepositoryErrorCorrelation.test.ts \
  apps/server/src/provider/Layers/ProviderSessionDirectory.test.ts \
  apps/server/src/provider/Layers/ProviderSessionReaper.test.ts \
  apps/server/src/provider/Layers/ProviderService.test.ts \
  apps/server/src/project/AgentSessionImporter.test.ts
# Test Files 17 passed (17); Tests 338 passed (338)
#   baseline on the restacked #4 head before new tests: 334 passed; +4 new regressions
vp run --filter t3 typecheck        # EXIT=0 (effect lint *suggestions* only, no error diagnostics)
vp lint <16 changed files>          # 0 warnings, 0 errors
vp fmt --check <16 changed files>   # "All matched files use the correct format."
git diff --check                    # clean

Environment: macOS (darwin arm64), Node v25.6.0, project-local vite-plus 0.3.3 / vitest 4.1.11, invoked as ./node_modules/.bin/vp (the global vp v1.0.0-rc.0 was not used). Dependencies installed offline (pnpm install --frozen-lockfile --offline, store warm, zero downloads). TMPDIR=/private/var/folders/l9/z97_k2450xv4z04_6wwsw5zh0000gn/T is pinned for the two pre-existing macOS realPath UsageService tests.

CI state (truthful)

Not green. At head 4d5671597 the self-hosted jobs Test, Test Server 1/2/3, Rust, Release Smoke, Mobile Native Changes, Native fingerprint diff, Check are pending (queued) — the fork has no blacksmith-* self-hosted runners (run 36198123187). GitHub-hosted jobs passed (Prepare PR size config, Collect PR targets); Deploy web preview / EAS Preview / Sync PR size label definitions skipped; CodeRabbit pass (review skipped: draft). No green CI is claimed; runner/environment ownership stays with the environment lane (pingdotgg#237).

Scope guards honored

No endpoint, UI, PostHog, collector, gateway, routing, or transport change. No provider routing tables, OpenRouter pinning rules, billing policy, qualification receipts, or semantic model-selection logic were added — those remain agent-config policy/qualification responsibilities. No raw prompt/response/code/tool content: thread_route_events and routeMetadata.ts are content-free allowlisted metadata. Requested provider/model/effort are captured from the model selection T3 was given; observed model comes only from measured usage, and actualEffort is always null/unsupported (never fabricated from the request).

Actual model/provider/effort used for THIS agent

  • Provider opencode-go, model deepseek-v4.1-flash via opencode (the harness this session runs in).
  • Requested effort: unknown (not set by the agent); observed effort: unknown/unobservable — no scanned source exposes it, and routing was not mutated to obtain a different model.
  • Native session id is not exposed to the agent.

Remaining gaps

Exact next action

Independent exact-head review of #6 @ 4d56715971dde1c60ee67e8c3d322812faa575e9 against base af70495442678a06e007d13bba477f123170c0f5: (a) the two append-only tables and the upsert-time append path, (b) requested-vs-observed separation, (c) the usageRouteAttribution view with the new SessionRouteQuality axes and the A–D M3D-preservation regressions. Do not merge #4 or #6. Merge ordering remains a hub decision.

@nullStack65

Copy link
Copy Markdown
Owner Author

ROUTE6-1:R4-REV-T3 — REVIEW

Refresh: stack matches the manager-observed head exactly; neither PR moved.

Verdict: CHANGES_REQUIRED

Independence disclosure: I could not select GPT-6 Sol. This review ran on openrouter/deepseek/deepseek-v4.1-flash (opencode harness) — the same model family the PR body names as the authoring harness (opencode-go / deepseek-v4.1-flash). Independence is therefore limited / same-family, not a qualified independent model. Confidence is high on the mechanical evidence below (probes), lower on stylistic judgement.

Diff stats

16 files, +2600 / −17, 2 commits (344be10e7, 4d5671597). Matches the PR body.

Evidence gathered

  • Local source/tests: the exact claimed set runs green — 17 files / 338 tests (matches the claim). Server typecheck vp run --filter t3 typecheck exits 0 (only pre-existing suggestion diagnostics; no error lines in PR files).
  • Adversarial probes (run in a throwaway detached temp worktree at the PR head; worktree and probe files deleted afterward; the PR source was not modified):
    • A: same native session id under two provider instances on one thread.
    • C: request record write ordering.
    • D: migration 054 backfill with malformed/non-string/empty cursors.
    • E: migration rerun / idempotency.
    • F: route view per-thread rollup for a session bound to two threads.
  • GitHub-hosted checks on [ROUTE6-1:T3] Durable model-canary attribution across resumes and model switches #6: pass (Collect PR targets, Label PR size, Label PR 6); CodeRabbit skipped (draft). Self-hosted blacksmith-* jobs pending.
  • Fork full CI: unavailable. On [METRICS-M3] Prove prompt request session and PR usage attribution #4 the same jobs are recorded as fail / 24h0m0s (runner-queue timeout), not code failure. On [ROUTE6-1:T3] Durable model-canary attribution across resumes and model switches #6 they are still pending. No hosted Test / Test Server run has exercised this head.

P0

None.

P1

  1. The multi-table upsert is not transactional. ProviderSessionRuntime.upsert (ProviderSessionRuntime.ts:594-618) performs three independent autocommit writes: runtime upsert → provider_session_history append → thread_route_events insert. There is no transaction. The runtime cursor can commit while the history append (or a route-event insert) fails, leaving contradictory durable state and reporting failure to a caller that may retry against a different cursor — the exact durable-history loss this change exists to prevent. SqlClient.withTransaction is already used in this layer (PullRequestFilesViewed.ts:135, AuthSessions.ts:384, ProjectionTurns.ts:269). Failure ordering can go either way (runtime committed / history missing, or history committed / runtime stale).
  2. provider_session_history unique key is too coarse for provider-instance identity. UNIQUE (thread_id, provider_name, native_session_id) (054_ProviderSessionHistory.ts:33) omits provider_instance_id. Two configured instances of one driver that expose the same native session id on one thread collapse to one row, and the later instance overwrites provider_instance_id. Probe A (write opencode-go, then cliproxy-loopback, same ses_shared, same thread) yielded provider_session_history = [{ providerInstanceId: "cliproxy-loopback", n: 1 }] — opencode-go is erased. This directly weakens question 11 (native OpenCode Go vs CLIProxy loopback).
  3. Migration 054 backfill does not mirror the runtime writer. The backfill guards only native_session_id <> '' (054_ProviderSessionHistory.ts:91-92) after COALESCE(json_extract(resume|threadId|sessionId)). Probe D: {"resume":123} backfills native_session_id = "123" (a non-string the runtime nativeSessionIdOf rejects), and {"resume":"","threadId":"real-id"} backfills nothing although the runtime path would use "real-id". So the migration can invent an invalid identity or miss a valid one.

P2

  1. Route per-thread usage double counts a session bound to two threads, and the stated invariant is false. usageRouteAttribution.ts:118 claims the thread rollup "misses nothing, double counts nothing". Probe F: one session bound to th1 and th2 produced perThread = [th1:100, th2:100] while base.allocation = "ambiguous" and base pools it as unallocated. SessionRouteQuality (usageRouteAttribution.ts:70-83) also drops the base allocation / boundThreadIds axes, so the route view cannot flag the ambiguity.
  2. Automatic request record is first-write-wins and can permanently lose the requested model. The request event id is ${threadId}::request::${nativeSessionId} (ProviderSessionRuntime.ts:459) and does not encode the requested selection, so INSERT OR IGNORE (:387) keeps the first write. Probe C: write {codex, model: null} then {codex, model: "gpt-6-luna"} → stored requested_model stays null. This is reachable because modelSelection is passed only at startSession (ProviderService.ts:1552) while recovery/stop/rollback/stopAll upserts pass a model-less requestedRoute.
  3. Production wiring gaps (scope limits, not defects). No read path consumes either new table; buildUsageRouteAttribution and the history/route extractors have no non-test callers. parentNativeSessionId and routeEvent are never set by any production caller — grep of apps/server/src/provider shows only plumbing (Directory/ProviderService). So child lineage and every declared kind (normal/availability_fallback/canary/independent_review/quality_escalation) are structurally supported but always null/absent in practice today. The docs are honest about this; the PR body's "measured across child sessions, retries/fallbacks" is aspirational until a producer is wired.
  4. extractAttributionHistory does not de-duplicate. Duplicate (thread, provider, native) input rows (e.g. from a join) produce duplicate SessionRouteReports. The write-seam unique key prevents this in the DB, but the pure view trusts its input.
  5. escalation_reason is unconstrained free text. Every other column is content-free; privacy relies on the caller not putting content there. Acceptable for a metadata carrier, but worth a length/charset guard.

Primary-question answers

  1. Schema migration genuinely required? Yes. provider_session_runtime holds a single cursor row per thread; prior native sessions cannot be represented without a new append-only table. Justified.
  2. 054/055 — additive yes, safe yes, correctly registered yes (Migrations.ts:68-69,135-136; manifest asserted in test), deterministic yes, backward-compatible yes, collision-free no (P1-2), backfill truthful no (P1-3). Highest pre-existing id is 053 → no collision/renumbering.
  3. Preserves every distinct native session? For distinct ids, yes: a repeat advances last_seen_at, never overwrites a different session, and the onConflict: "ignore" path appends nothing (ProviderSessionRuntime.ts:596-601). Not for one driver + one native id across two instances (P1-2), and parent lineage is never populated in production (P2-6). last_seen_at ordering uses lexicographic excluded.last_seen_at > existing (:357) — correct for uniform ISO-8601 UTC, would misorder mixed offsets.
  4. Current-cursor backfill safe/correct? Mostly safe, not semantically exact (P1-3).
  5. Upsert append path transactionally sensible? Idempotent yes; transactional no (P1-1). Contradictory durable state is possible in both directions.
  6. Repeated observations idempotent? Yes — history last_seen advance, deterministic route-event ids, migration rerun no-op (Probe E). Caveat: the request record is idempotent but can be stale (P2-5).
  7. thread_route_events content-free/low-cardinality? Yes — only ids, label selection, stratum, cohort, reason; no bodies.
  8. Does nullable route_event_kind distinguish automatic request vs declared normal/fallback/canary/review/escalation? Yes at schema/parser/extractor level (routeMetadata.ts:19-26,108-118; extractor splits declared vs requests). No production producer writes declared events yet (P2-6).
  9. Requested vs observed strictly separated? Yes. requested comes only from route events; actualProvider/actualModel only from measured usage/usageProvider; never copied (usageRouteAttribution.ts:254-276,303-317). requestedQuality is never observed.
  10. Actual effort unsupported/null? Yes — always null, quality unsupported (:275-276, :426).
  11. Provider-instance identity sufficient? No (P1-2). SessionRouteReport also drops providerInstanceId even though base sessions retain providerInstanceIds (usageAttribution.ts:968).
  12. [METRICS-M3] Prove prompt request session and PR usage attribution #4 / M3D semantics preserved? Yes. usageTranscripts.ts, usageScanCache.ts, UsageService.ts are byte-identical across the delta (absent from name-status). The route view mirrors measurement/identity/prompt/request/recordIdentity/conflict (usageRouteAttribution.ts:70-83,256-283). Regressions A–D are present and pass locally.
  13. Route/session quality propagates without false upgrade? Yes for the quality axes; the allocation axis is dropped (P2-4). A history-only identity correctly gets quality = null, usage = null.
  14. Multiple-PR association non-additive? Yes — base shared/attributed logic is unchanged; the thread rollup sums each session once in the non-ambiguous case (Probe F shows the ambiguous exception).
  15. History-only identity = usage null / quality null (not measured zero)? Yes (:304-311; tests "reports no measurement quality…" and case D).
  16. OpenCode child sessions truthful? Yes — parent_native_session_id is a dedicated field and child identities are never flattened; docs explicitly state the adapter does not yet persist every child id. No false claim found.
  17. Privacy? Yes — no prompt/response/code/tool-body/secret columns; reason strings are metadata by convention (P2-8 caveat).
  18. T3 metadata carrier, not policy owner? Yes — routeMetadata.ts:1-16 and the 055 header state agent-config decides; no routing decision code added.

Missing adversarial cases

Coverage requested vs actual:

  • concurrent / reordered upserts — not covered.
  • transaction failure — not covered (and cannot be, without a transaction: P1-1).
  • duplicate route-event IDs — covered only same-thread/same-session; cross-thread primary-key collision not tested.
  • malformed requested route — normalizeRouteSelection only lightly exercised.
  • provider-instance change under same provider/model — not covered (fails, Probe A).
  • same native id under different provider instances — not covered (P1-2).
  • backfill from malformed cursor — not covered (P1-3).
  • migration rerun / idempotency — 054 test runs to 53 then 55 once; no explicit second-run assertion.
  • current-cursor / history disagreement — covered only for the onConflict: "ignore" path, not for partial-write failure.

Test counts are not treated as correctness proof; the probes above were needed to surface P1-2/P1-3/P2-4/P2-5.

CI limitations

Ready for integration after #4?

No. CHANGES_REQUIRED on P1-1..P1-3; P2-4 and P2-5 should be resolved (or explicitly accepted) before canary measurements are trusted.

Exact repair requirements (not implemented here)

  • R1 Wrap the runtime + history + route-event writes for one upsert in a single transaction.
  • R2 Include provider_instance_id (with defined collapse semantics) in the provider_session_history identity, and expose it on the route session report.
  • R3 Make the 054 backfill require text-typed, non-empty id fields and mirror nativeSessionIdOf precedence.
  • R4 Carry allocation/boundThreadIds into the route view; stop double counting ambiguous sessions; fix or qualify the "double counts nothing" claim.
  • R5 Make the automatic request record updatable (selection-aware id or DO UPDATE on a richer selection).
  • R6 Optional: de-duplicate history input; constrain escalation_reason.
  • R7 Add the missing adversarial tests (concurrency/reorder, cross-thread duplicate eventId, instance change, backfill malformed, migration rerun, transaction failure).

No source edited, no push, no merge. #4 and #6 unchanged.

nullStack65 and others added 2 commits September 26, 2026 19:01
T3's single resume cursor lost earlier native sessions on a resume, fork, or model switch, so a thread could not say which models contributed. Add an append-only provider_session_history (migration 054, with a current-cursor backfill) written alongside the runtime upsert, and read it through usageAttributionSources as sessionHistory bindings.

Add a content-free thread_route_events store (migration 055) for pre-execution route/experiment metadata: requested provider/model/effort captured at session start, task stratum, experiment/cohort, route event kind, manager/agent ids, and escalation reason. Actual values are derived from measured usage, never copied from the request. usageRouteAttribution composes the base projection with this view and keeps child sessions distinct.

T3 carries the metadata only; agent-config remains the routing authority. No prompts, responses, code, or tool bodies are stored.
…ality

Add regressions that the durable session-history route view cannot regress the
#4 (METRICS-1:M3D) measurement/identity quality axes:

- A: an explicit complete measured zero stays a measured zero, not missing;
- B: an all-invalid eligible observation stays invalid, not missing and not a
  measured zero, and mirrors the base projection;
- C: a predecessor v4 cache entry decodes as `predecessor` with `partial`
  completeness and the route view never promotes it to complete;
- D: a history-only identity for a measurable provider stays usage=null and
  quality=null, never a fabricated measured zero.

No production behavior changes; the M3D parser/cache files remain byte-identical
to the restacked #4 head.
@nullStack65
nullStack65 force-pushed the feat/durable-model-canary-attribution-20260923 branch from 4d56715 to 1a1993b Compare September 26, 2026 23:33
@nullStack65

Copy link
Copy Markdown
Owner Author

ROUTE6-1:R5-T3 — RESULT

Status: complete (repair of #6 in place). Fresh session; not a continuation of R3-T3. No merge. CI cannot be observed green on the fork — see CI state.

Heads

Ref At start Final
#4 head (observed by manager) af70495442678a06e007d13bba477f123170c0f5 —
#4 head (actual, after refresh) f863939bcd52b49f238982f593ea34c146f80e28 — #4 had moved unchanged (f863939…)
#6 old head (reviewed) 4d56715971dde1c60ee67e8c3d322812faa575e9 —
#6 restacked head — 3a1047a16 (= 4d5671597 replayed onto f863939)
#6 new head — 1a1993b516a9e56802ddf3f1f8dbf8795c2dc551

PR #6 now shows base.sha = f863939…, head = 1a1993b51…. Backups created before rewriting: branch/tag backup/route6-r5-pre-restack-4d5671597 and backup/route6-r5-pr4-old-base-af7049544.

Restack method (required — #4 moved)

R1–R7 disposition

# Finding Disposition
R1 multi-table upsert not transactional Fixed. See transaction design.
R2 history identity too coarse for provider instance Fixed. See provider-instance identity.
R3 054 backfill not faithful to nativeSessionIdOf Fixed. See backfill semantics.
R4 ambiguous multi-thread double count + false invariant Fixed. See ambiguous-thread accounting.
R5 automatic request first-write-wins loses model Fixed. See automatic-request behavior.
R6a extractAttributionHistory not de-duplicated Fixed. See history de-dup.
R6b escalation_reason unconstrained free text Fixed. See reason guard.
R7 missing adversarial tests Added (atomicity, instance change, cross-thread event id, reorder, malformed cursor, rerun, malformed requested route, ambiguity).

Findings were repaired against the R4 review text, not against a repair implementation from an earlier T3 session. PR body updated to remove the aspirational "measured across child sessions/retries" wording (P2-6).

Transaction design

ProviderSessionRuntime.upsert now wraps one logical upsert in the repository's existing sql.withTransaction (same mechanism used by PullRequestFilesViewed.ts, AuthSessions.ts). For onConflict: "update" the transaction covers: runtime cursor row, the provider_session_history append, and every implied thread_route_events row. Any failure rolls back all of them. For onConflict: "ignore" the write stays a single statement (ON CONFLICT(thread_id) DO NOTHING) and still appends nothing — a cursor that was never applied produces no history/events. Invariant: for a single accepted upsert, either all durable effects commit or none.

Atomicity failure tests (actual SQLite, not mocks)

Run against SqlitePersistenceMemory with real triggers that RAISE(ABORT) at the persistence layer:

  • A history BEFORE INSERT trigger fails → upsert fails, runtime row is absent, history is empty.
  • B route-event BEFORE INSERT trigger fails → runtime and history are absent, and no route event exists.
  • C successful path commits runtime + history + request event together.
  • D after dropping the trigger, the retry commits exactly one history row and the expected cursor (idempotent).

Provider-instance identity design

  • provider_session_history unique identity is now (thread_id, provider_name, provider_instance_key, native_session_id).
  • provider_instance_key is a normalized non-null string: a trimmed instance id, or "" for unknown/null. NULL is not used as an identity axis because SQLite UNIQUE treats each NULL as distinct. "" is a deterministic bucket, so two unknown-instance observations of one (thread, provider, native session) collapse to one row while a real instance stays distinct.
  • Same thread/provider/native session + same instance → one idempotent row; different instance → two rows; a later instance cannot overwrite an earlier instance's row (verified).
  • Route/session report now exposes providerInstanceIds (union of base binding evidence and history), so instance identity is never dropped by the view; SessionRouteReport also carries allocation and boundThreadIds.
  • Two instances of one unmeasured session (native OpenCode Go vs CLIProxyAPI loopback) stay distinct session reports with usage: null.

Migration-054 backfill semantics

Mirrors runtime nativeSessionIdOf exactly: only JSON text values, trimmed non-empty, precedence resume → threadId → sessionId, with fallthrough to the next candidate on absent/null/non-string/blank. Numbers, booleans, and objects never become ids. Explicit cases tested: valid resume; empty resume + valid threadId; non-string resume + valid threadId; blank resume/threadId + valid sessionId; whitespace-only ignored; all-malformed → no row; null provider instance → "" key; whitespace-padded instance → trimmed key; and migration rerun is idempotent.

Ambiguous-thread accounting behavior

SessionRouteReport carries base allocation and boundThreadIds. A session bound to more than one thread is ambiguous; it is assigned threadId: null and its usage is not summed into any candidate thread. Per-thread totals plus the new route-view unallocated bucket reconcile additively to each session's usage exactly once. Regression: one measured session bound to th1 and th2 → both thread totals 0, unallocated holds the session once, sum(threads) + unallocated = session usage. The old "misses nothing, double counts nothing" comment is replaced with the precise statement.

Automatic-request enrichment / conflict behavior

INSERT OR IGNORE replaced with an explicit UPSERT policy on the automatic request record:

  • null → non-null: enrich;
  • equal → equal: idempotent;
  • non-null → null: retain (a model-less recovery/stop/rollback write never erases);
  • conflicting non-null → different non-null: retain the first value and set persisted selection_conflict = 1, surfaced through extraction (RouteEventMetadata.selectionConflict) and diagnostics (conflictingSelections).

All four cases tested at the real persistence layer, plus the model-less recovery case.

History de-dup behavior

extractAttributionHistory de-duplicates input by the canonical identity (thread, provider, normalized instance, native session), merges monotonically (min firstSeen, max lastSeen, fill null parent), and counts duplicatesDropped; a duplicate whose adapterKey disagrees increments conflictingDuplicates rather than being silently merged.

Reason guard

escalation_reason is bounded to a low-cardinality lowercase code/slug: max 64 chars, allowlisted charset [a-z0-9_.:-] starting alphanumeric, no whitespace/multiline. Non-conforming values are rejected to null at the writer and the extractor. Existing prose reason in the #6 test was updated to canary_measured.

Declared event id scoping (adversarial #8)

Declared event ids are persisted scoped by thread (<threadId>::declared::<id>), so a globally reused id no longer collides on the global primary key and can no longer suppress another thread's event. Tested with the same declared id on two threads → both rows survive.

Exact diff stats

Changed paths (repair delta, 11)

  • apps/server/src/persistence/Migrations/054_ProviderSessionHistory.ts
  • apps/server/src/persistence/Migrations/054_ProviderSessionHistory.test.ts
  • apps/server/src/persistence/Migrations/055_ThreadRouteEvents.ts
  • apps/server/src/persistence/ProviderSessionRuntime.ts
  • apps/server/src/persistence/ProviderSessionRuntime.history.test.ts
  • apps/server/src/usage/routeMetadata.ts
  • apps/server/src/usage/usageAttributionSources.ts
  • apps/server/src/usage/usageAttributionSources.test.ts
  • apps/server/src/usage/usageRouteAttribution.ts
  • apps/server/src/usage/usageRouteAttribution.test.ts
  • docs/internals/usage-attribution.md

Tests / validation

vp test run apps/server/src/usage/ \
  apps/server/src/persistence/ProviderSessionRuntime.history.test.ts \
  apps/server/src/persistence/Migrations/054_ProviderSessionHistory.test.ts \
  apps/server/src/persistence/RepositoryErrorCorrelation.test.ts \
  apps/server/src/provider/Layers/ProviderSessionDirectory.test.ts \
  apps/server/src/provider/Layers/ProviderSessionReaper.test.ts \
  apps/server/src/provider/Layers/ProviderService.test.ts \
  apps/server/src/project/AgentSessionImporter.test.ts
# Test Files 17 passed (17); Tests 362 passed (362)
#   baseline at restacked head before repair: 342 passed; +20 new repair tests
vp run --filter t3 typecheck        # EXIT=0; 0 `error TS` lines (effect suggestions only)
vp lint <10 changed source files>   # EXIT=0
vp fmt --check <11 changed files>   # "All matched files use the correct format."
git diff --check                    # clean

Environment: macOS (darwin arm64), Node v25.6.0, project-local vite-plus 0.3.3 / vitest 4.1.11, invoked as ./node_modules/.bin/vp. Dependencies installed offline (pnpm install --frozen-lockfile --offline, zero downloads). TMPDIR=/private/var/folders/l9/... pinned for the macOS realPath UsageService tests.

M3D / #4 preservation

usageTranscripts.ts, usageScanCache.ts, UsageService.ts and their tests are byte-identical across the full stack delta (absent from #4 f863939 → #6 1a1993b51 name-only except usageScanCache/UsageService which appear only because #4 itself advanced; verified the route-view A–D regressions and the existing base tests still pass). Explicit zero retained, invalid/partial retained, predecessor-cache quality partial, dedupeKeyScope, cost-only conflict, identity-erased unavailable, warm-cache quality gate all pass. No requested/observed mixing, actualEffort still null/unsupported.

CI state (truthful)

  • GitHub-hosted: Collect PR targets pass, Label PR 6 pass, Label PR size pass, Prepare PR size config pass; Deploy web preview / EAS Preview / Sync PR size label definitions skipped; CodeRabbit pass (review skipped: draft).
  • Self-hosted blacksmith-*: Test, Test Server 1/2/3, Rust, Release Smoke, Mobile Native Changes, Check, Native fingerprint diff remain queued/pending. Fork full CI unavailable — a runner/environment gap, not a code failure. Runner architecture not modified.

Actual agent route

  • Provider opencode-go, model deepseek-v4.1-flash via opencode (this harness).
  • Requested effort: unknown (not set by the agent); observed effort: unknown/unobservable. Routing was not mutated to obtain a different model. Native session id not exposed to the agent.

Remaining scope limitations (not defects; P2-6)

  • New tables have a tested write seam + extraction seam but no production read consumer yet.
  • Only the automatic request record has a production producer; declared kinds (normal/availability_fallback/canary/independent_review/quality_escalation) remain future agent-config/harness wiring.
  • parent_native_session_id is structurally supported but no production adapter populates it; OpenCode child ids are not persisted by any adapter.
  • Provider-instance identity is carried through the view but is still not recoverable from a transcript scan.
  • Fork CI cannot go green without self-hosted runners (environment lane).

Exact next action

Independent exact-head review of #6 @ 1a1993b516a9e56802ddf3f1f8dbf8795c2dc551 against base f863939bcd52b49f238982f593ea34c146f80e28, focused on: (a) the single-transaction upsert and the A–D failure-injection tests, (b) the instance-aware provider_session_history identity and null-instance collapse rule, (c) the 054 backfill precedence, (d) ambiguous-thread reconciliation, (e) automatic-request enrichment/conflict and cross-thread declared-id scoping. Do not merge #4 or #6. Merge ordering remains a hub decision.

…lict-safe

Repair the ROUTE6-1:R4-REV-T3 review findings on #6:
- P1-1 wrap runtime cursor + provider_session_history + thread_route_events
  in one sql.withTransaction so one accepted upsert is all-or-nothing.
- P1-2 include a normalized non-null provider_instance_key in the history
  unique identity and expose provider-instance ids on the route session.
- P1-3 mirror nativeSessionIdOf precedence in the 054 backfill (text-only,
  trimmed non-empty, resume -> threadId -> sessionId with fallthrough).
- P2-4 stop double counting a session bound to two threads; pool it once in
  a route-view unallocated bucket and carry allocation/boundThreadIds.
- P2-5 make the automatic request record monotonically enrichable with a
  surfaced selection_conflict instead of INSERT OR IGNORE.
- P2-7 de-duplicate history input by canonical identity.
- P2-8 bound escalation_reason to a low-cardinality slug.
- scope declared event ids per thread so a shared id cannot suppress
  another thread's event.
- add atomicity/enrichment/instance/backfill/reorder adversarial tests and
  refresh docs.
@nullStack65
nullStack65 force-pushed the feat/durable-model-canary-attribution-20260923 branch from 1a1993b to e0e6318 Compare September 27, 2026 01:02
@nullStack65

Copy link
Copy Markdown
Owner Author

ROUTE6-1:R5-T3 — CHECKPOINT (head supersedes the RESULT reference)

The R5-T3 RESULT above referenced head 1a1993b516a9e56802ddf3f1f8dbf8795c2dc551. One small refinement was made after that post and pushed with an explicit lease (--force-with-lease=…:1a1993b51):

  • Final head: e0e6318a84c6efb6afb46c14cf0e48f4906a1a4e (base f863939bcd52b49f238982f593ea34c146f80e28).
  • Refinement: the route view's thread list now includes every thread named by the binding evidence, so an ambiguous session's candidate threads remain visible with zero usage instead of disappearing. Regression tightened to assert both candidate threads exist with usage.totalTokens = 0 and empty sessionLabels; sum(threads) + unallocated = session usage unchanged.
  • Repair delta is now 11 files, +1206 / −133, 1 commit; full stack delta 16 files, +3665 / −16, 3 commits.
  • Re-validated at e0e6318a8: 17 files / 362 passed; typecheck exit 0 (0 error TS); lint exit 0; vp fmt --check clean; git diff --check clean.
  • All other R1–R7 dispositions, the transaction design, instance identity, backfill semantics, ambiguity accounting, enrichment policy, de-dup, and reason guard are unchanged from the RESULT above.

Exact next action: review #6 @ e0e6318a84c6efb6afb46c14cf0e48f4906a1a4e against base f863939bcd52b49f238982f593ea34c146f80e28. No merge.

@nullStack65

Copy link
Copy Markdown
Owner Author

ROUTE6-1:R6-REV-T3 — REVIEW

Verdict: PASS (with non-blocking residuals; #6 is ready for integration after #4)

Exact anchors / stack proof

#4 accepted-cache / lq preservation — PASS

R1 — atomic multi-table upsert — PASS

ProviderSessionRuntime.upsert (ProviderSessionRuntime.ts:700-747):

  • The update path wraps upsertRuntimeRow + recordSessionHistoryRow + the route-event writes in one sql.withTransaction (:713). All inner statements execute on the same sql client; effect-smol's withTransaction injects the transaction connection into the fiber context and getConnection selects it, so every write is on the transaction's connection.
  • Verified the failure-injection tests pass: history-abort rolls back runtime (runtime absent, history 0); route-event-abort rolls back runtime + history + request row; retry after rollback succeeds and is idempotent (upsert is ON CONFLICT DO UPDATE, history merges last_seen, request enriches, declared is INSERT OR IGNORE).
  • onConflict: "ignore" stays a single INSERT … ON CONFLICT(thread_id) DO NOTHING statement and appends no history/events for a conflicting (not-applied) cursor.
  • No path lets runtime commit without history, or runtime+history commit without route events.

R2 — provider-instance-aware identity — PASS

  • Unique key is (thread_id, provider_name, provider_instance_key, native_session_id) in 054 and in the writer (:355); key is normalized non-null via normalizeProviderInstanceKey (trim, "" for null/unknown).
  • A: repeat → one row, last_seen advances. B: same native id under two instances → two rows (test). C: null/unknown → deterministic "" bucket, collapses (test). D: a later instance cannot overwrite/erase an earlier instance's row (test). E: route view exposes providerInstanceIds per session plus base providerInstanceIds. F: unmeasured (OpenCode) two instances of one native id stay two distinct reports; measured sessions keep one report with both instance ids listed (information preserved, not erased).
  • SQLite NULL-distinctness is correctly neutralized. I found no accepted path where instance identity is silently lost.

R3 — migration 054 backfill — PASS

  • Precedence resume → threadId → sessionId, JSON text-only, trimmed non-empty, fallthrough otherwise — verified against runtime nativeSessionIdOf and with adversarial rows: {resume:123,threadId:"real"}→real, {resume:"",threadId:"real"}→real, {resume:" ",threadId:"",sessionId:"final"}→final, bool/object/array candidates→no row, valid resume wins, null/absent cursor→no row.
  • provider_instance_key normalization matches JS (trim / ""); migration rerun is idempotent (CREATE TABLE/INDEX IF NOT EXISTS + INSERT OR IGNORE over the unique key), asserted by test. Migrations registered 54/55; highest pre-existing was 53 (no collision).

R4 — ambiguous thread accounting — PASS (Probe F reproduced)

One measured session (usage=100 tokens) bound to th1+th2:

  • session allocation=ambiguous, threadId=null, boundThreadIds=[th1,th2];
  • thread[th1].usage=0, thread[th2].usage=0, thread[].sessionLabels=[] (the ambiguous session is not listed under each candidate thread);
  • unallocated=100; sum(per-thread additive)+unallocated = 100, not 200.
    Test route-view allocation and provider-instance identity asserts exactly this plus the unallocated limitation text. Pooling is exactly once.

R5 — automatic-request monotonic enrichment — PASS

All four cases verified with real SQLite persistence (test monotonically enriches an automatic request record): null→non-null enriches; equal→equal idempotent; non-null→null retains (model-less recovery/stop does not erase gpt-6-luna); conflicting non-null→different non-null retains first and sets selection_conflict=1. extractAttributionRouteEvents projects selectionConflict; the route view surfaces routeSelectionConflict and identity.conflictingSelections. The upsert CASE logic implements first-non-null-wins correctly.

R6 — de-dup + reason guard — PASS

  • extractAttributionHistory de-dups by canonical identity (thread/provider/normalized-instance/native), merges firstSeen min and lastSeen max, fills a null parent monotonically, and counts an adapter-identity disagreement as conflictingDuplicates while keeping the first adapter key (no silent switch). Null-instance collapse and distinct-instance separation both tested.
  • escalation_reason: normalizeEscalationReason bounds to ≤64 chars, ^[a-z0-9][a-z0-9_.:-]*$, rejects whitespace/multiline/prose. Applied at the writer (ProviderSessionRuntime.ts:597) and the extractor (usageAttributionSources.ts:574); extractor test rejects prose/oversized and keeps a bounded code.

R7 — adversarial coverage — PASS

Present and passing: transaction rollback (both history and route-event), retry-after-rollback idempotency, provider-instance change / same-native-different-instance, malformed cursor rows, migration rerun, malformed requested route, reordered repeated upserts, cross-thread declared-id, ambiguous reconciliation, enrichment 4 cases, #4 quality regressions A–D. I did not treat count as proof; I mapped each to the behavior and ran the files.

Declared event id scoping — PASS with a flagged residual

Cross-thread: persisted key "<threadId>::declared::<id>" prevents one thread's reused caller id from suppressing another's (test asserts two rows). Flagged residual (non-blocking): same thread + same declared id but different native session/kind/content is INSERT OR IGNORE and silently dropped with no conflict diagnostic (unlike the request record's selection_conflict). This cannot occur in production today because no production caller supplies routeEvent yet; recommend either a distinct-content check or a conflict marker when declared producers are wired.

Requested-vs-observed — PASS

requested comes only from thread_route_events/session-start metadata; actualModel only from measured usage; actualEffort is always null with quality unsupported and is never copied from requested. No requested→observed copying found; requestedQuality/observed axes are explicit.

Privacy — PASS

provider_session_history and thread_route_events store only ids, timestamps, enum/route metadata, and the bounded reason. No prompts, responses, code, tool bodies, secrets, or arbitrary content. The extractor drops payload fields (snapshot test asserts a secret marker and path never serialize).

Production-wiring limitation assessment — TRUTHFUL

Verified in source: extractAttributionSnapshot/buildUsageRouteAttribution have no production read consumer (tests only). The only production producer is the automatic request record (ProviderService.ts:1098 → requestedRoute). No production caller passes routeEvent or parentNativeSessionId (tests only). Docs state all of this explicitly ("no production read consumer yet", declared kinds + parent future wiring); PR body states the write/extraction seam and the single production producer. These are truthful scope limits, not integration blockers.

Test / CI evidence

  • Local (clean tree at e0e6318a8, w-t3/r5-t3): PR6-specific files 67 passed (ProviderSessionRuntime.history, 054_…, usageAttributionSources, usageRouteAttribution); provider-layer + importer + correlation 123 passed. All coverage above reproduces.
  • vp fmt --check on changed files: clean. vp lint on changed files: clean. Typecheck was not re-run to completion locally (server typecheck exceeded my 15-minute budget; no type errors surfaced in test compile). Author's typecheck-clean claim is unrefuted.
  • Pre-existing environmental failures: UsageService.test.ts has 2 failures ("customer path include", "in-flight scan timeout"). I reproduced the same 2 failures at accepted [METRICS-M3] Prove prompt request session and PR usage attribution #4 head f863939 (byte-identical file), so they are environmental/pre-existing, not [ROUTE6-1:T3] Durable model-canary attribution across resumes and model switches #6 regressions.
  • GitHub-hosted checks: non-self-hosted labeling jobs pass; CodeRabbit skipped (draft). Self-hosted jobs (Test, Test Server 1-3, Rust, Release Smoke, Native fingerprint diff, Mobile Native Changes) are pending/queued on the fork. Runner absence is not a code failure.

Ready for integration after #4?

Yes. #6 is restacked exactly on the accepted #4 head, preserves #4 cache/lq behavior, and repairs R1–R7. Suggested (non-blocking) follow-ups: (1) declared same-thread id conflict surfacing, (2) SQLite TRIM() vs JS trim() divergence (SQLite trims only spaces; a tab/newline-only cursor/instance would diverge — see note), (3) append history on the onConflict:"ignore" insert path or document it.

Exact repairs required

None blocking.

Independence

Fresh session; no R4-REV-T3/R5-T3 context resumed; no source edits, no commits/pushes, #4/#6 not modified; tests run read-only in a clean detached worktree already at the exact head (and a control run at #4 head). Smallest-trim divergence: reviewed strictly the current stack after refresh.

Reviewer model/provider: openrouter/deepseek/deepseek-v4.1-flash via the OpenCode harness. The R5-author reported running deepseek-v4.1-flash (opencode-go) — same model family, so this rereview is not model-independent; it is independently executed on a fresh context with its own reproductions and a #4-head control. No routing/spend was mutated to obtain a specific reviewer.

Minor note (SQLite whitespace): migration 054 uses TRIM() (space-only) and json_type(...)= 'text', while nativeSessionIdOf uses JS String.trim() (all whitespace). A cursor/instance value consisting solely of a tab/newline/NBSP would be accepted as an id by the backfill but rejected by the runtime. Non-material (no real session id is whitespace-only) but worth a TRIM(x, char(9)||char(10)||char(13)||char(32)) alignment if exact parity is desired.

@nullStack65
nullStack65 changed the base branch from feat/usage-attribution-proof-20260923 to main September 28, 2026 05:24
@nullStack65
nullStack65 marked this pull request as ready for review September 28, 2026 05:25
@nullStack65
nullStack65 merged commit f5d3fc6 into main Sep 28, 2026
9 of 22 checks passed
@coderabbitai

coderabbitai Bot commented Sep 28, 2026

Copy link
Copy Markdown

Important

  • 🔍 Trigger review

This repository does not receive automatic reviews because it has fewer than 10 stars.

⚙️ Run configuration

Configuration used: Repository: nullStack65/t3code/.coderabbit.yaml

Review profile: CHILL

Plan: Advanced

Run ID: 4beccd23-f0c4-43b6-8458-5c8323e4de60


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@nullStack65

Copy link
Copy Markdown
Owner Author

ROUTE6-1:R7-LAND-T3 — RESULT

Status: complete — both #4 and #6 landed on T3 main as ancestry-preserving merge commits. Fresh session; no R5-T3 / R6-REV-T3 context resumed. No source rewrite, no rebase, no force push.

Exact anchors

Ref SHA Verified
starting main bcc1a58b19a9d610a4f08fed191a364767bc65b3 unchanged pre-merge
#4 reviewed/accepted head f863939bcd52b49f238982f593ea34c146f80e28 matched, not moved
#6 reviewed head e0e6318a84c6efb6afb46c14cf0e48f4906a1a4e matched, not moved
git merge-base(#4,#6) f863939bcd52b49f238982f593ea34c146f80e28 == #4 head ✅

Exact-head gate held and neither PR moved during this session, so the #6 delta was exactly #4→#6.

Runner / check state

  • The fork lacks the upstream blacksmith-* self-hosted runners. R6 explicitly treated this as an environment limitation, not a code defect.
  • Self-hosted jobs (Test, Test Server 1/2/3, Rust, Release Smoke, Mobile Native Changes, Native fingerprint diff, Check) are cancelled/queued on the fork and never ran.
  • GitHub-hosted checks passed on both heads: Collect PR targets, Prepare PR size config, Label PR 4 / Label PR 6, Label PR size; CodeRabbit pass (review skipped: draft).
  • Branch protection: main is NOT protected — GitHub API returns "Branch not protected" (HTTP 404). There are no required checks, so the unavailable self-hosted jobs were never a merge blocker.
  • No runner created, no workflow policy changed, no paid CI, no protected check bypassed.

Step 1 — #4 landed (merge commit)

Step 2 — #6 retargeted to main (no source rewrite)

Step 3 — #6 landed (merge commit)

Final state

Post-merge validation

Branch-protection blocker

None. main is unprotected; an ordinary merge proceeded for both PRs. No required check was bypassed, and none was skipped by policy change.

Exact next dependency

closura-agent-config#231 must refresh its #4/#6 fixture pins onto the landed T3 main (f5d3fc66016d54a16fd8872321d7722d4457526e) before pingdotgg#248 integration. The bridge extractAttributionSnapshot fixture pin (historically daa55eb88 / af7049544) should be regenerated/confirmed against landed main.

Not touched in this session

agent-config pingdotgg#231 (referenced only), pingdotgg#248, pingdotgg#246/pingdotgg#249/pingdotgg#250, upstream pingdotgg/t3code. No force push; #4 and #6 feature branches left intact.

Model/harness: openrouter/deepseek/deepseek-v4.1-flash via the OpenCode harness. Agent label ROUTE6-1:T3.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XXL vouch:trusted PR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant