Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
22 commits
Select commit Hold shift + click to select a range
e6a85a8
feat(dashboard): report host context evidence without implying live s…
pacphi Sep 9, 2026
2c1b651
feat(dashboard): preserve repository and desktop origin dimensions
pacphi Sep 9, 2026
7f775c6
refactor(dashboard): clarify maintenance project card identity and la…
pacphi Sep 9, 2026
9302882
test(status): normalize inspection timestamps and record context report
pacphi Sep 9, 2026
5ae65f3
feat(maintenance): report repository groups and independent session o…
pacphi Sep 9, 2026
d61f172
feat(dashboard): group projects and compact host context evidence
pacphi Sep 9, 2026
167d1b4
feat(intelligence): group and alphabetize learning project picker
pacphi Sep 9, 2026
5a3e517
test(dashboard): verify grouped Intelligence picker in desktop and mo…
pacphi Sep 9, 2026
d35e8bb
feat(dashboard): surface cached model capacities across supported hosts
pacphi Sep 9, 2026
06176e4
test(status): record cached context inventory metadata in golden fixture
pacphi Sep 9, 2026
9231b13
refactor(dashboard): show concise context controls and evidenced mode…
pacphi Sep 9, 2026
dae932d
feat(usage): group project costs by verified repository identity
pacphi Sep 9, 2026
aad6387
test(dashboard): exercise Usage project groups and worktrees in browser
pacphi Sep 9, 2026
5720009
fix(dashboard): route model inventory links and simplify picker labels
pacphi Sep 9, 2026
fe39ffc
docs(dashboard): record grouping design and browser evidence
pacphi Sep 9, 2026
0e691a2
test(dashboard): canonicalize Windows short paths in identity assertions
pacphi Sep 9, 2026
163778c
fix(usage): rank only verified Git projects with real parent roots
pacphi Sep 9, 2026
22d766e
fix(usage): restore a top-ten Git-project spend ranking
pacphi Sep 9, 2026
7feb010
feat(intelligence): subgroup machine-wide learning with bounded scrol…
pacphi Sep 9, 2026
031a55f
fix(context): distinguish missing windows from absent sessions
pacphi Sep 9, 2026
b506e25
docs(dashboard): document grouped learning and context coverage
pacphi Sep 9, 2026
0acebf2
refactor(projects): remove redundant directory group headings
pacphi Sep 9, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
31 changes: 21 additions & 10 deletions docs/DASHBOARD.md
Original file line number Diff line number Diff line change
Expand Up @@ -134,7 +134,8 @@ Overview keeps status and routing in one health-first area:
Direct Ruflo agents must explicitly select OpenRouter or Ollama together with a provider-native
model, and the Ruflo/MCP process must inherit the required credential environment. Served-provider
and served-model claims come from **Usage → Scorecard** evidence instead.
- **Runtime** presents operational services, processes, and MCP readiness.
- **Runtime** presents operational services, processes, MCP readiness, and cached
context configuration with host-specific native controls.
- **Intelligence** presents memory, learning, and quality-improvement signals machine-wide: an
always-visible rollup folded across every project on this machine where memory or intelligence has
been activated — a `.claude-flow`, `.agentic-qe` or `.swarm` directory, whichever host created it
Expand All @@ -149,13 +150,15 @@ Overview keeps status and routing in one health-first area:
[Project intelligence](ddd/project-intelligence.md) and
[ADR-0024](adr/0024-project-intelligence-telemetry.md) for the full model and the two learning
metrics' load-bearing distinction, and [ADR-0027](adr/0027-shared-project-census.md) for project
discovery.
discovery. The machine-wide table and picker share alphabetized Git repository,
worktree, user-level, and other/unclassified subgroups. Each table subgroup shows
five rows before scrolling; all rows remain available inside the bounded panel.

### Why project counts differ between tabs

Every area derives its project list from one census, so a project means the same thing everywhere —
a session run in `myrepo/backend` belongs to `myrepo`, not to a project called `backend`, and an
agent worktree is not a peer of the repository it was cut from.
Project counts reflect different populations. Intelligence uses the learning census, while Usage
Score ranks verified Git identities from the selected session window. Evidence-backed worktree
associations keep a worktree with its repository; display names alone do not establish identity.

The totals still differ, because the tabs ask different questions:

Expand Down Expand Up @@ -204,6 +207,12 @@ evidence at all is structurally zero rather than cheap, and folding it in would
toward zero for a reason that is not about spend. A real figure under a cent renders `<$0.01`, never
`$0.00`.

**Projects** ranks the top 10 discovered Git projects by spend in the selected
timeframe. Verified worktree usage rolls into its parent project; standalone
worktrees, user-level locations, and unclassified directories are excluded from
this panel. Overall Usage totals retain all activity, including usage outside
these ten rows.

**Rhythm & responsiveness** puts two histograms side by side, session length and response latency,
each with its percentile markers laid over the bars. A percentile that lands in the open-ended
top bucket renders with a `≥` prefix — the bucket has no upper edge, so the honest claim is a floor
Expand Down Expand Up @@ -276,13 +285,15 @@ sources: [Usage scorecard metrics](USAGE-SCORECARD-METRICS.md) §2a, §2b, §20
Context answers how much of a runtime-observed window the retained sessions used. The policy strip
shows the canonical startup, dynamic and reserve bands. The summary reports exactly how many
sessions have a paired input/window pressure observation and how many lack a denominator. Claude,
Codex and OpenCode each keep their own card with evidence state, p90 peak pressure, paired-sample
count, p90 peak input and median observed window.
Codex and OpenCode each keep their own card with coverage state, p90 peak pressure, number of
sessions with pressure measurements, p90 peak input and median observed window.

A percentage is rendered only when input and window were observed together for that session.
Claude and OpenCode commonly carry input-only evidence, so they may read **Partial evidence** while
Codex is observed. Missing evidence renders `unknown`; an unknown ARIA meter omits `aria-valuenow`
rather than announcing zero. The attention projection is capped to the top 20 sessions before
The cards distinguish **Input only**, **Partial coverage**, **Not recorded**, and **No sessions**.
Missing token/window values render as an em dash. A pressure meter appears only for a measured
value, and each card explains its coverage gap. Claude transcript input records do not include a
paired window; older Codex records can contain only cumulative totals. OpenCode may have no
sessions in the chosen window even when its installation and model catalog are available. The attention projection is capped to the top 20 sessions before
presentation. The browser renders one disclosure row per bounded project and keeps each sanitized
conversation label inside the expanded session table, then exposes explicit column headers, an opaque session reference, host, policy-derived
recommendation, pressure/input/window, and start date. The session reference links to that retained
Expand Down
10 changes: 5 additions & 5 deletions docs/TRANSCRIPTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -179,8 +179,8 @@ The same parsers serve two very different callers, switched by `withTurns`:

| Path | Entry point | `withTurns` | Message bodies | Cached? |
|---|---|---|---|---|
| **Scan** — the aggregate index behind the Scorecard/Findings/Sessions views | `buildIndex` (`usage-index.mjs:566`) → `parseFile` (`usage-index.mjs:389`) | `false` | no turn list is built, and no body is retained — holding them would balloon memory across 3,000+ files (`parseClaude`'s own doc comment, `usage-parsers.mjs:679-684`). Since v14 the scan path does *read* one narrow slice: opencode's USER text parts, so a prompt can be fingerprinted (`loadTextParts`, `usage-opencode.mjs:276`). Only the fingerprint is kept; the text is discarded with the row. Measured at 45 µs/session materializing 0.6 MB on a 300-session store, against 125 µs and 61 MB for the reader path's unfiltered join | yes: per-file derived records in `~/.config/agentic-kit/usage-index.json`, keyed `(path, mtime, size)`, invalidated wholesale by `SCHEMA_VERSION` (`usage-index.mjs:142`) |
| **Reader** — one transcript for the Transcript view | `readSession` (`usage-index.mjs:960`) | `true` | full turn list built | **never** — every call re-reads and re-parses the one file |
| **Scan** — the aggregate index behind the Scorecard/Findings/Sessions views | `buildIndex` (`usage-index.mjs:578`) → `parseFile` (`usage-index.mjs:401`) | `false` | no turn list is built, and no body is retained — holding them would balloon memory across 3,000+ files (`parseClaude`'s own doc comment, `usage-parsers.mjs:679-684`). Since v14 the scan path does *read* one narrow slice: opencode's USER text parts, so a prompt can be fingerprinted (`loadTextParts`, `usage-opencode.mjs:276`). Only the fingerprint is kept; the text is discarded with the row. Measured at 45 µs/session materializing 0.6 MB on a 300-session store, against 125 µs and 61 MB for the reader path's unfiltered join | yes: per-file derived records in `~/.config/agentic-kit/usage-index.json`, keyed `(path, mtime, size)`, invalidated wholesale by `SCHEMA_VERSION` (`usage-index.mjs:154`) |
| **Reader** — one transcript for the Transcript view | `readSession` (`usage-index.mjs:972`) | `true` | full turn list built | **never** — every call re-reads and re-parses the one file |

![Figure: one parser, two read paths — the scan path (withTurns false) caches per-file records keyed by path, mtime and size; the reader path (withTurns true) builds full turns and is never cached](assets/transcript-read-paths.svg)

Expand Down Expand Up @@ -317,15 +317,15 @@ string.

## 4. The `readSession` pipeline — how one session becomes a payload

`readSession(id, opts)` (`usage-index.mjs:960-987`) is the only way
`readSession(id, opts)` (`usage-index.mjs:972-987`) is the only way
transcript content leaves the module, and every step is a gate:

### 4.1 Locate, contain, bound

1. **Id grammar before any filesystem access** — an id must match one of
exactly two shapes, or it is rejected with `ERR_INVALID_SESSION_ID` at
this call: `invalidId(id)` (`usage-index.mjs:961`) before any read happens:
* `VALID_ID` (`/^[A-Za-z0-9._-]{1,128}$/`, `usage-index.mjs:153`) — a plain
* `VALID_ID` (`/^[A-Za-z0-9._-]{1,128}$/`, `usage-index.mjs:165`) — a plain
session id;
* `VALID_SUBAGENT_ID` (`usage-index.mjs:169`) — a namespaced nested
subagent id, EXACTLY `<parentId>/<stem>` with one slash, where the parent
Expand All @@ -345,7 +345,7 @@ transcript content leaves the module, and every step is a gate:
symlinks; a symlink planted inside a root pointing at `/etc/anything`
passes a lexical `startsWith` but fails this. Roots are realpath'd too so
a symlinked dotfiles setup still works.
4. **Size cap** — `MAX_SESSION_BYTES` (64 MB, `usage-index.mjs:152`): a
4. **Size cap** — `MAX_SESSION_BYTES` (64 MB, `usage-index.mjs:164`): a
transcript is read whole and JSON-expands ~5×, so an unbounded read is a
memory-amplification primitive. Oversized reads as unavailable, not risky.

Expand Down
50 changes: 25 additions & 25 deletions docs/USAGE-SCORECARD-METRICS.md
Original file line number Diff line number Diff line change
Expand Up @@ -73,7 +73,7 @@ Every metric section below follows the same shape:
Two transcript stores, read-only, parsed at most once per file — the derived
record is cached "keyed by (path, mtime, size)" (`src/lib/usage-index.mjs:10`),
and the whole cache is invalidated on a `SCHEMA_VERSION` change
(`usage-index.mjs:142`):
(`usage-index.mjs:154`):

| Transcript host | Store | Format |
|---|---|---|
Expand Down Expand Up @@ -784,7 +784,7 @@ session that runs from 23:58 local to 00:05 local is billed to the day its
*first* row landed on (test:
`tests/kit/usage-index.test.mjs:738`, "a session that opens before midnight
is counted on its first billed day"). Accumulation, at this call: `dayBucket(byDay,
row.day)` then `d.cost = round(d.cost + rowCost)` (`usage-aggregate.mjs:747-752`). Bar height:
row.day)` then `d.cost = round(d.cost + rowCost)` (`usage-aggregate.mjs:760-767`). Bar height:
`h = maxDay ? max(2, cost/maxDay*100) : 2` (`dashboard/client.mjs`) —
every non-empty day gets a visually nonzero bar (floor of 2%), so a very
cheap day is never rendered as invisible.
Expand Down Expand Up @@ -976,25 +976,25 @@ ranking entirely rather than merely re-labelled in place.

## 11. Projects

**Displayed as:** a ranked bar list, top 8 shown, note reading `"top 8 of
N"` when more exist; each row shows `cost`, `N sess · minutes`.
**Displayed as:** a simple ranked bar list of the top 10 discovered Git projects
in the selected timeframe. Each row shows cost, session count, and minutes.
No standalone worktree, user-level, or unclassified directory appears here.

**Formula:** identical shape to §10 (`byProject[project]`), plus a `project
= 'unknown'` fallback and a repo/worktree collapsing rule:
`projectLabel(cwd)` collapses `<repo>/<marker>/worktrees/<rest>` (marker ∈
`.autopilot`, `.claude`, `.git`) to `<repo>`, keeping `rest` as a separate
`worktree` field on the session rather than either discarding it or letting
it masquerade as a sibling project. (The mislabelling this rule corrected is
recorded in [Appendix A](#appendix-a--fix-history).)
**Formula:** the additive `gitProjects` projection uses the same filtered session
population as the rest of Usage. It sums session cost, count, tokens, and minutes
by evidenced repository identity. A worktree contributes to its parent only when
Git common-directory/backlink evidence establishes an existing project root.
Exact user/host-state roots and unverifiable repository associations are excluded
from this ranking. Names and remotes alone do not establish eligibility.

**Source:** ranking and truncation, `dashboard/client.mjs`
(`shown = projects.slice(0,8)`); accumulation via the same `addTo()`/
`entries()` machinery as §10, keyed by `s.project` instead of `s.models`.
**Source:** `usage-project-evidence.mjs`, `usage-project-groups.mjs`, and
`dashboard/client/usage.mjs` (`renderScoreProjects`). The existing `byProject`
aggregate remains available to other consumers, but does not establish Git identity.

**What this does not model:** a session whose working directory could not be
determined (e.g. missing `cwd` in the transcript) lands in a literal
`"unknown"` bucket rather than being dropped — visible in the project list
rather than silently absent from the total.
**Population:** the overall Usage totals still include all recorded usage. They
can exceed the sum of these ten Git-project rows. Missing identity remains in
those totals; it is not guessed into this ranking. Older payloads without Git
identity request a usage refresh instead of displaying arbitrary directories.

---

Expand Down Expand Up @@ -1753,7 +1753,7 @@ cacheSavedUsd = Σ rows (costOf(1M as input) - costOf(1M as cacheRead))
**Source:** the derived block is `finishTotals` (`usage-aggregate.mjs:1033-1073`),
which the previous-window projection calls too so a baseline is never derived a
second, drifting way. `median` and `percentile` are exact over the values
(`usage-aggregate.mjs:998-1009`), unlike §15's bucketed percentiles.
(`usage-aggregate.mjs:1020-1032`), unlike §15's bucketed percentiles.
Active days come from `byDay`'s key count and the streak from `activeStreak` in
`src/lib/dashboard/client/usage.mjs`; the tiles are `cadenceCells` there, and
`printScoreCadence` (`src/commands/usage.mjs:219-242`) in the CLI.
Expand Down Expand Up @@ -1802,9 +1802,9 @@ positive figure that rounds away at two decimals prints `<$0.01`, never
"nothing" are different claims.

**What the cache saved, asked as a difference.** `cacheSavingPerMillion`
(`usage-aggregate.mjs:718-727`) prices one million tokens twice through the
(`usage-aggregate.mjs:731-742`) prices one million tokens twice through the
*injected* pricer — once as fresh input, once as cache reads — and takes the
gap; `cacheSavedFor` (`usage-aggregate.mjs:717-720`) scales that to the tokens
gap; `cacheSavedFor` (`usage-aggregate.mjs:731-749`) scales that to the tokens
a row actually read from cache. Nothing in that path knows what the cache
multiplier is, so the saving cannot drift out of step with §3's table the way a
hard-coded "0.9 × input" would the day the multiplier changed. Both probes
Expand Down Expand Up @@ -2155,13 +2155,13 @@ unbounded observation list. Codex reads the gross `last_token_usage.input_tokens
not add its cached-input subset again. Claude and OpenCode sum their split fresh/cache fields but
usually have no runtime window, so their coverage is commonly partial.

The current cache schema is v18 because controlled prompt intent/topic facets also require parser
output. A v17 cache is reparsed; the context evidence contract itself is unchanged.
Cache schema v20 also retains Git-project eligibility evidence. Older cache records are reparsed;
the context evidence contract itself is unchanged.

The Context view contains:

- the canonical startup/dynamic/reserve policy bands;
- counted coverage, including paired samples and sessions missing a window;
- counted coverage, including sessions with paired measurements and sessions missing a window;
- one host card each for Claude, Codex and OpenCode;
- p90 peak pressure/input and median observed window where supported; and
- at most 20 attention rows carrying a deterministic opaque session reference plus bounded project
Expand Down Expand Up @@ -2334,7 +2334,7 @@ promise.
`byModel` on the first run after the change, purely because the cache
predated it; every unit test still passed, since tests only exercise a
fresh parse. `SCHEMA_VERSION` went to `4` specifically to force the one-time
re-parse; the constant now reads `17` (`usage-index.mjs:142`), each bump since
re-parse; the constant now reads `17` (`usage-index.mjs:154`), each bump since
having forced its own re-parse the same way.
Re-querying the same live server after the bump returned
`totals.exceptions: 20` with `<synthetic>` absent from `byModel` —
Expand Down
Loading
Loading