The dominant cost driver in an agent session is the number of assistant turns, not the size of the work done. cached_read is charged once per assistant turn, and each charge re-reads the entire prefix — system prompt, tool defs, preloaded skill text, and the transcript accumulated so far — not just what changed. Cost is therefore Σ(prefix size per turn), and because the prefix grows monotonically as the transcript accumulates, cost scales with the number of turns and the transcript’s growth, not with the size of the work: a turn late in a long session is more expensive than an identical turn early in a short one, regardless of what either turn accomplishes.
A turn carrying several parallel tool calls pays that prefix-read cost once, not once per call. So tool calls are an upper bound on turns, and the ratio between them (calls ÷ turns) is the batching factor — the degree to which a procedure is using parallel calls to avoid paying for extra turns. This is why “fewer tool calls” and “fewer turns” are not the same lever: shrinking what’s preloaded on each turn cuts a flat per-turn constant, but eliminating whole turns removes that constant’s re-read entirely, which is the larger win whenever the transcript is already nontrivial. #98 measured this directly on wiki-ingest: trimming preloaded skill text from ~10K to ~4K tokens was projected at roughly 8.7% of a 53-turn run’s cache-read bill, versus 50%+ available from removing turns outright.
The metric a procedure should be designed against is turns, but turns are what gets instrumented as tool calls, not turns exactly. #99 spiked the PostToolUse hook payload to check whether turns are recoverable and found they are not: subagent tool calls do reach the hook, but under the parent’s session_id, not their own; agent_id/agent_type distinguish a subagent’s calls from the parent’s, but no field identifies the originating assistant message (prompt_id is scoped to the whole user-prompt turn, not a single assistant message within it), and the payload carries no timestamp at all — so not even approximate wall-clock clustering is available as a fallback. Separately, subagent turns are not observable locally by any means: across every session JSONL under ~/.claude/projects/, isSidechain appears on 278 lines and is false on all of them, so a transcript parser counting turns has nothing to parse for delegated work. This is why #100’s instrument is a PostToolUse hook counting tool calls, not a transcript parser counting turns: it’s the only signal available that also captures subagent activity, at the cost of being a proxy rather than an exact count.
Every downstream design choice in #98 follows from optimizing turns while measuring calls:
discover.py). A script call that returns everything a following decision needs (title, relative link, tags, volatility) collapses what would otherwise be a script call plus N exploratory Read calls into one turn.ingest.py reports a tool-call count that is nothing to do with ingestion (#100). The count is surfaced in the ingest manifest purely to make turn-count regressions visible on every run, not because ingestion itself has any use for the number.The revisit trigger for this ADR is the same fact that makes tool calls a proxy in the first place: if a future CLI version exposes a payload field that groups tool calls by originating assistant message (per-message identifier, or even a timestamp precise enough for clustering), turns become exactly recoverable and the instrumentation in #100 should switch from counting calls to counting turns directly, no design change needed elsewhere.