Claw Fleet

Field report · 2026-08-08 → 2026-09-07

One machine, 30 days,
1,660 sessions.

This is not a benchmark run. It is the complete usage record of one development machine over 30.4 days, every number recomputed from local transcripts — nothing estimated, no samples cherry-picked. Fleet does not sell you a faster model; it spreads one person's attention across more parallel work. This page measures how far.

81.6×

Each human intervention buys 81.6 model round trips. 382,874 API calls answer to just 4,693 human inputs.

Sessions live at once (time-weighted mean)

6.98

Peak 26. A bare terminal caps at 1 — one terminal, one session.

Share of busy time with 2+ sessions live

99.4%

Across 30 days there was almost never only one thing running.

Relay handoffs across context windows

524

A full context window is not the end of the task: a successor session picks it up with the briefing.

List-price-equivalent usage

$108,981

Recomputed per usage record at published prices — not a bill; the account is on a subscription.


Concurrency is the only throughput you can buy.

How many engineering threads one person can push depends on how many they can watch at once. These two charts are the actual distribution and the day-by-day curve over 30.4 days — not a peak screenshot.

Sessions live simultaneously during busy time

Share of 729.8 active wall-clock hours; only sessions a human actually typed into and that ran longer than 60s

2+ in parallel

99.4%

3+ in parallel

85.3%

5+ in parallel

71.8%

10+ in parallel

25.0%

025%50%75%100%
The bare terminal has no bar here because its ceiling is structural, not distributional: one terminal, one session, and it stops when you look away — so it is 1 session 100% of the time. Method: each session's first and last timestamps form an interval; a sweep line weights them by wall clock. Claude Code's own seconds-long helper runs are excluded; including them puts the peak reading at 199.
Table view
Sessions in parallelShare of active wall clockCumulative hours
2+99.4%725.4 h
3+85.3%622.5 h
5+71.8%523.9 h
10+25.0%182.4 h
peak26—

Model round trips per day

Fleet-spawned sessions, each call binned by its own timestamp; 380,284 across 31 days

0k10k20k30k 26,901 08-0808-1508-2208-2909-052026-08-08 · 2,042 calls · 10 sessions live2026-08-09 · 560 calls · 5 sessions live2026-08-10 · 39 calls · 2 sessions live2026-08-11 · 35 calls · 2 sessions live2026-08-12 · 32 calls · 2 sessions live2026-08-13 · 1,796 calls · 10 sessions live2026-08-14 · 3,155 calls · 17 sessions live2026-08-15 · 8,205 calls · 32 sessions live2026-08-16 · 4,337 calls · 23 sessions live2026-08-17 · 6,093 calls · 22 sessions live2026-08-18 · 14,710 calls · 42 sessions live2026-08-19 · 10,560 calls · 43 sessions live2026-08-20 · 16,563 calls · 51 sessions live2026-08-21 · 20,434 calls · 69 sessions live2026-08-22 · 15,451 calls · 62 sessions live2026-08-23 · 23,074 calls · 80 sessions live2026-08-24 · 21,135 calls · 77 sessions live2026-08-25 · 21,287 calls · 87 sessions live2026-08-26 · 16,311 calls · 79 sessions live2026-08-27 · 16,079 calls · 64 sessions live2026-08-28 · 12,847 calls · 37 sessions live2026-08-29 · 11,852 calls · 31 sessions live2026-08-30 · 16,013 calls · 34 sessions live2026-08-31 · 17,363 calls · 33 sessions live2026-09-01 · 16,870 calls · 31 sessions live2026-09-02 · 26,901 calls · 58 sessions live2026-09-03 · 12,716 calls · 32 sessions live2026-09-04 · 15,612 calls · 32 sessions live2026-09-05 · 13,305 calls · 36 sessions live2026-09-06 · 20,560 calls · 87 sessions live2026-09-07 · 14,347 calls · 77 sessions live
The mid-August trough is real: for those days only two long sessions were running. What follows is not a one-day sprint but three weeks of steady load — 12k calls a day on average, peaking at 27k.
Table view
DateAPI callsSessions live
2026-08-082,04210
2026-08-095605
2026-08-10392
2026-08-11352
2026-08-12322
2026-08-131,79610
2026-08-143,15517
2026-08-158,20532
2026-08-164,33723
2026-08-176,09322
2026-08-1814,71042
2026-08-1910,56043
2026-08-2016,56351
2026-08-2120,43469
2026-08-2215,45162
2026-08-2323,07480
2026-08-2421,13577
2026-08-2521,28787
2026-08-2616,31179
2026-08-2716,07964
2026-08-2812,84737
2026-08-2911,85231
2026-08-3016,01334
2026-08-3117,36333
2026-09-0116,87031
2026-09-0226,90158
2026-09-0312,71632
2026-09-0415,61232
2026-09-0513,30536
2026-09-0620,56087
2026-09-0714,34777

Machine work bought by one human intervention

Every square is one call; the orange one is the human input

1 human input81.6 model round trips
382,874 API calls4,693 human inputs4,215 decision cards
Decision cards and human inputs come out nearly equal — Fleet collapses every "needs a person" moment into one card with options instead of leaving it buried in scrolling output. Median time to an answer: 147 seconds.
Table view
CountTotal over 30.4 daysRelative to human input
382,874 API calls382,87481.6×
4,693 human inputs4,6931×
4,215 decision cards4,2150.90×

What that throughput buys.

Every entry carries a measured number from this machine and what a bare terminal does in the same spot. Capabilities without a measurement are not on this list.

6.98

Parallel engineering threads

Time-weighted sessions live at once, peaking at 26. Each has its own worktree, its own plan, its own verification.

Bare terminal: 1, and it stops when you switch away.

524

The context window stops being the end

As a window fills, the session registers a handoff and Fleet spawns a successor with the briefing and the plan. 524 times in 30 days.

Bare terminal: dig through the compaction, or restart by hand.

45

Work that ran unattended

Sessions started by schedules and loops with nobody at the keyboard. A cheap probe runs first; no LLM starts unless it passes.

Bare terminal: 0 — close the terminal and it is gone.

4,215

Decisions collapsed into cards

Every judgement call becomes one card with options, answered in a median of 147 seconds — from the phone too.

Bare terminal: mixed into the output, for you to spot.

382,874

Every call is on the books

All calls land with their usage record, so cost can be recomputed per session, model or period — that is where this page comes from.

Bare terminal: transcripts too, but no aggregate view.

97.9%

Long context held up by caching

97.9% of input tokens hit the prompt cache — the reason a 190k-token call is economically viable at all.

Capability source: Claude Code, not introduced by Fleet.


Method and provenance.

This section states how each number was collected and computed, and where the data applies.

How it was measured

  1. Usage and cost: every usage record under ~/.claude/projects/*/*.jsonl, recomputed at current published prices, with cache writes priced separately for the 5-minute and 1-hour TTLs. No pre-existing ledger is trusted.
  2. Concurrency: each session's first and last timestamps form an interval; a sweep line weights the result by wall clock, excluding Claude Code's own seconds-long helper runs.
  3. Daily curve: each API call is binned by its own timestamp in local time, not by when its session started.
  4. Wait time: the delta between a decision card's tool_use and its matching tool_result.
  5. Session provenance: the entrypoint field on each transcript's first user record; Fleet-spawned sessions carry their own marker.

Scope of these figures

Study type. An observational usage record, not a same-task controlled experiment. The bare-terminal values for concurrency, relay and unattended work are taken from its structural ceiling (one terminal one session, stop at a full window, gone when closed) rather than from measured samples.

Sample size. One operator, one machine, 30.4 days; n = 1. Suitable for describing real usage on this configuration, not for direct extrapolation across environments; a multi-sample controlled study is separate work.

Cost basis. $108,981 is list-price-equivalent usage recomputed from local records, not a billed amount; the account is on a subscription and paid a different amount.

Data window: 30.4 days is this machine's transcript retention period, not the total history of use. Overhead: Fleet's injected discipline and its bookkeeping round trips add roughly 23% to input tokens, 9% to output and 4–6% to machine wall clock, listed here alongside the throughput above.


Questions

Does Fleet burn more tokens?

Yes. It injects discipline and plan context into every call and adds round trips for bookkeeping, handoffs and cards: about +23% input, +9% output, +7% calls. What you get back is the concurrency and unattended work measured on this page; token volume is not what this design optimises for.

Can I recompute these numbers on my own machine?

Yes. They all come from local transcripts under ~/.claude/projects; the collection and accounting scripts ship with this page and read your own usage records.

Do the hooks slow the machine down?

Not measurably: the per-turn plan injection takes about 10 ms, plan bookkeeping a median 0.0s, handoff and worktree operations a median 1.4s. The real waiting is a decision card waiting on a person.

Why is there no same-task A/B comparison?

This machine has only 2 bare-terminal sessions, too few to form a control group, and a control assembled after the fact is not included here. A full same-task A/B needs a fixed task set, both modes run over it, and control for model version and the operator's own learning; that is separate work. What this page does support is the differential measurement: Fleet's injected text subtracted from the same call, so the error comes from tokenization rather than task variance. The concurrency, relay and unattended figures are compared against a bare terminal's structural ceiling, marked in the scope note as not measured samples.