Field report · 2026-08-08 → 2026-09-07
One machine, 30 days,
1,660 sessions.
This is not a benchmark run. It is the complete usage record of one development machine over 30.4 days, every number recomputed from local transcripts — nothing estimated, no samples cherry-picked. Fleet does not sell you a faster model; it spreads one person's attention across more parallel work. This page measures how far.
81.6×
Each human intervention buys 81.6 model round trips. 382,874 API calls answer to just 4,693 human inputs.
Sessions live at once (time-weighted mean)
6.98
Peak 26. A bare terminal caps at 1 — one terminal, one session.
Share of busy time with 2+ sessions live
99.4%
Across 30 days there was almost never only one thing running.
Relay handoffs across context windows
524
A full context window is not the end of the task: a successor session picks it up with the briefing.
List-price-equivalent usage
$108,981
Recomputed per usage record at published prices — not a bill; the account is on a subscription.
Concurrency is the only throughput you can buy.
How many engineering threads one person can push depends on how many they can watch at once. These two charts are the actual distribution and the day-by-day curve over 30.4 days — not a peak screenshot.
Sessions live simultaneously during busy time
Share of 729.8 active wall-clock hours; only sessions a human actually typed into and that ran longer than 60s
Table view
| Sessions in parallel | Share of active wall clock | Cumulative hours |
|---|---|---|
| 2+ | 99.4% | 725.4 h |
| 3+ | 85.3% | 622.5 h |
| 5+ | 71.8% | 523.9 h |
| 10+ | 25.0% | 182.4 h |
| peak | 26 | — |
Model round trips per day
Fleet-spawned sessions, each call binned by its own timestamp; 380,284 across 31 days
Table view
| Date | API calls | Sessions live |
|---|---|---|
| 2026-08-08 | 2,042 | 10 |
| 2026-08-09 | 560 | 5 |
| 2026-08-10 | 39 | 2 |
| 2026-08-11 | 35 | 2 |
| 2026-08-12 | 32 | 2 |
| 2026-08-13 | 1,796 | 10 |
| 2026-08-14 | 3,155 | 17 |
| 2026-08-15 | 8,205 | 32 |
| 2026-08-16 | 4,337 | 23 |
| 2026-08-17 | 6,093 | 22 |
| 2026-08-18 | 14,710 | 42 |
| 2026-08-19 | 10,560 | 43 |
| 2026-08-20 | 16,563 | 51 |
| 2026-08-21 | 20,434 | 69 |
| 2026-08-22 | 15,451 | 62 |
| 2026-08-23 | 23,074 | 80 |
| 2026-08-24 | 21,135 | 77 |
| 2026-08-25 | 21,287 | 87 |
| 2026-08-26 | 16,311 | 79 |
| 2026-08-27 | 16,079 | 64 |
| 2026-08-28 | 12,847 | 37 |
| 2026-08-29 | 11,852 | 31 |
| 2026-08-30 | 16,013 | 34 |
| 2026-08-31 | 17,363 | 33 |
| 2026-09-01 | 16,870 | 31 |
| 2026-09-02 | 26,901 | 58 |
| 2026-09-03 | 12,716 | 32 |
| 2026-09-04 | 15,612 | 32 |
| 2026-09-05 | 13,305 | 36 |
| 2026-09-06 | 20,560 | 87 |
| 2026-09-07 | 14,347 | 77 |
Machine work bought by one human intervention
Every square is one call; the orange one is the human input
Table view
| Count | Total over 30.4 days | Relative to human input |
|---|---|---|
| 382,874 API calls | 382,874 | 81.6× |
| 4,693 human inputs | 4,693 | 1× |
| 4,215 decision cards | 4,215 | 0.90× |
What that throughput buys.
Every entry carries a measured number from this machine and what a bare terminal does in the same spot. Capabilities without a measurement are not on this list.
6.98
Parallel engineering threads
Time-weighted sessions live at once, peaking at 26. Each has its own worktree, its own plan, its own verification.
Bare terminal: 1, and it stops when you switch away.
524
The context window stops being the end
As a window fills, the session registers a handoff and Fleet spawns a successor with the briefing and the plan. 524 times in 30 days.
Bare terminal: dig through the compaction, or restart by hand.
45
Work that ran unattended
Sessions started by schedules and loops with nobody at the keyboard. A cheap probe runs first; no LLM starts unless it passes.
Bare terminal: 0 — close the terminal and it is gone.
4,215
Decisions collapsed into cards
Every judgement call becomes one card with options, answered in a median of 147 seconds — from the phone too.
Bare terminal: mixed into the output, for you to spot.
382,874
Every call is on the books
All calls land with their usage record, so cost can be recomputed per session, model or period — that is where this page comes from.
Bare terminal: transcripts too, but no aggregate view.
97.9%
Long context held up by caching
97.9% of input tokens hit the prompt cache — the reason a 190k-token call is economically viable at all.
Capability source: Claude Code, not introduced by Fleet.
Method and provenance.
This section states how each number was collected and computed, and where the data applies.
How it was measured
- Usage and cost: every usage record under
~/.claude/projects/*/*.jsonl, recomputed at current published prices, with cache writes priced separately for the 5-minute and 1-hour TTLs. No pre-existing ledger is trusted. - Concurrency: each session's first and last timestamps form an interval; a sweep line weights the result by wall clock, excluding Claude Code's own seconds-long helper runs.
- Daily curve: each API call is binned by its own timestamp in local time, not by when its session started.
- Wait time: the delta between a decision card's
tool_useand its matchingtool_result. - Session provenance: the
entrypointfield on each transcript's first user record; Fleet-spawned sessions carry their own marker.
Scope of these figures
Study type. An observational usage record, not a same-task controlled experiment. The bare-terminal values for concurrency, relay and unattended work are taken from its structural ceiling (one terminal one session, stop at a full window, gone when closed) rather than from measured samples.
Sample size. One operator, one machine, 30.4 days; n = 1. Suitable for describing real usage on this configuration, not for direct extrapolation across environments; a multi-sample controlled study is separate work.
Cost basis. $108,981 is list-price-equivalent usage recomputed from local records, not a billed amount; the account is on a subscription and paid a different amount.
Data window: 30.4 days is this machine's transcript retention period, not the total history of use. Overhead: Fleet's injected discipline and its bookkeeping round trips add roughly 23% to input tokens, 9% to output and 4–6% to machine wall clock, listed here alongside the throughput above.
Questions
Does Fleet burn more tokens?
Yes. It injects discipline and plan context into every call and adds round trips for bookkeeping, handoffs and cards: about +23% input, +9% output, +7% calls. What you get back is the concurrency and unattended work measured on this page; token volume is not what this design optimises for.
Can I recompute these numbers on my own machine?
Yes. They all come from local transcripts under ~/.claude/projects; the collection and accounting scripts ship with this page and read your own usage records.
Do the hooks slow the machine down?
Not measurably: the per-turn plan injection takes about 10 ms, plan bookkeeping a median 0.0s, handoff and worktree operations a median 1.4s. The real waiting is a decision card waiting on a person.
Why is there no same-task A/B comparison?
This machine has only 2 bare-terminal sessions, too few to form a control group, and a control assembled after the fact is not included here. A full same-task A/B needs a fixed task set, both modes run over it, and control for model version and the operator's own learning; that is separate work. What this page does support is the differential measurement: Fleet's injected text subtracted from the same call, so the error comes from tokenization rather than task variance. The concurrency, relay and unattended figures are compared against a bare terminal's structural ceiling, marked in the scope note as not measured samples.