keeping the receipts: what the git log doesn't say

there's a directory on my laptop called ~/dev. it holds every personal project i've made in the last ten years.. seventy-odd git repos, most of them dormant, a few of them finished, a handful still moving. this summer i've been reviving some of them with agentic coding and starting new ones the same way, and a few things fell out of that which i want to write down.

the short version: a turn late in a long agent session costs about five times what a turn early in one does, and once i could see that, tokens per commit dropped by about 60%. the trick wasn't a better model or a better prompt. it was keeping the receipts. the rest of this is how i found that out, and what the receipts say about what all of this actually costs.

the git log is the source of truth

this was true long before agents. when i'm chasing a bug the first thing i do is go back in time: git log, git blame, find the commit that changed the line, read the message, follow it to the ticket, read the comments. a good commit message says what changed and why, and a project with a good log is readable as a history on its own. it's the one record that survives editors, machines, and memory.

the conversation is ephemeral

with an agent there's a whole conversation going on that produces the commits. what i asked for, what it tried, what got rejected and why. that's the detail behind every commit message, and the agent doesn't keep it. claude code writes one jsonl file per session and prunes it after thirty days. so the log says what, the conversation says how, and the how was evaporating.

while building a life-casting app i was sharing with friends, right off the bat i asked for a way to view the transcripts and commits in the app itself, so anyone could read how the thing was being built. that grew into a habit, and then into a hook: when a session ends its raw transcript is copied into the repo and committed, and each commit carries a trailer naming the session it came out of. now i have both records side by side.. the discussion and the code change and the commit message.. and one points at the other.

an aside, because it was fucking cool, drumpy, a drum practice app i've been building for about five months, had months of sessions behind it before any of this existed, and by the time i went to archive them agent sessions, they had been pruned all but three. i asked about recovering them from backups. the agents mounted the mac time machine snapshots, found 58 of the 65 sessions, wrote a small tool to pull the fullest copy of each one into the repo, and suggested raising the prune window so it wouldn't happen again. drumpy's history came back mostly whole.

what it cost

on top of the transcript, every turn carries its token usage, so once the sessions were in the repo the total was a sum away. with the life-casting app, after a week of development, was close to $900 at token list price. 620 million tokens by the time it settled. i had not expected a number like that for a side project's first week.

these agentic coding tools with subscriptions are hiding the true cost. a flat monthly rate, and the meter doesn't show it; a long session costs the same as a short one, and the marginal session reads as free. the transcripts are the only place the meter still runs. list price isn't what the subscription bills either; it's what the meter would read if there were one.

drumpy is the bigger number: 2.2 billion tokens over five months, about $1,600 at list, for 55,000 lines of code across 280 commits.

step chart of cumulative tokens across all drumpy claude code sessions from april through early july, rising to 2.1 billion; two steep stretches
cumulative tokens across drumpy's claude code sessions from its first tracked session through early summer, one point per day with a session; the last month is left off. drawn from the per-turn usage in the transcripts.

98% of those tokens are cache reads. long sessions are expensive because every turn re-reads the whole context, not because the model writes a lot. the lever is context × turns.

across all of ~/dev

so i applied the same tooling to every live project in ~/dev. twenty repos, about 140,000 lines of code, 1,400 commits going back to the first commit in the oldest of them. 3.0 billion tokens and roughly $2,900 at list, and that's a floor.. seven of the smaller repos had already been pruned to nothing when the tooling landed, so their tokens are simply gone.

project commits loc tokens list
drumpy 280 55,000 2.2 B $1,600
the life-casting app 230 12,000 620 M $870
this blog 310 6,300 48 M $83
everything else (17) 570 69,000 140 M $300

two projects are 94% of the tokens. the rest is the long tail of things i poked at with an agent for an afternoon.

what i didn't expect is how useful the drill-down is. pick a line of code, blame it to a commit, follow the commit to its session, and read the conversation that created it. half the time the "why" is right there in my own prompt. it's also a mirror: reading back what i asked for, and what i had to ask twice, has done more for my prompting than anything i've read about prompting.

why long sessions cost more

here are the first four days of the life-casting app, from the usage on each turn, with what got committed each day:

day turns turns ≥ 500k commits
1 870 73 60
2 800 300 77
3 270 0 11
4 200 9 19

day 2 is the expensive one: 300 turns past half a million of context, and every one of them re-reads the whole conversation. a turn at 500k costs about five times one at 100k. by day 4 the cost per commit had roughly halved. day 3 is the catch.. fewer turns, but so few commits that per commit it was the worst of the four. shorter sessions are the lever, but only if the work gets done in them.

it changed the way i code

so now the projects i'm actively working on have three files: TODO.md for intake, PLAN.md for the work in progress, DONE.md for what shipped, dated. one item is one unit of work. i finish it, commit it, and end the session.. the transcript gets archived on the way out. the next unit starts cold, and cold starts are cheap because the plan and the log already say where things stand. big reads go to subagents whose context dies with them.

did it work? same app, before and after the change, counted per commit and per line changed so that the app slowing down as it matured doesn't get credited to the discipline:

life-casting app sessions turns / commit avg ctx tokens / commit tokens / 100 lines
before 6 13 280k 3.7 M 2.9 M
after 11 10 140k 1.4 M 1.6 M

per commit that's a 2.7x drop, about 60%. it comes from two places: turns per commit fell a little, 13 to 10, and the context under each turn fell a lot, 280k to 140k. the two aren't independent.. context is what piles up when a session runs long, so ending the session is the one lever that pulls both.

the closest thing i have to a controlled comparison is a change i made to nine repos in one go recently.. the same linting setup in each, one short session per repo, no plan file, in and out. those nine sessions came to 360 thousand tokens per commit. the two biggest all-day sessions on drumpy in early summer came to 8.2 and 7.2 million per commit. different work, so not a clean test, but a twenty-fold gap is more than the work explains. that's way cool.

a couple of weeks later

i kept going, so i reran the numbers. the life-casting app held right where the "after" row left it, and context stayed down everywhere else too.

what crept back was long sessions. a handful of them on infrastructure work ran past a hundred turns, one left going overnight, and cost about three times as much per commit as the rest. a small context wasn't enough.. ending the session is still the lever, and when i didn't, the receipts said so.

at what cost

i'm proud of these things. drumpy and the life-casting app are real, they work, and i made them short order, where either would have taken me years on nights and weekends, if it ever got finished at all.

but the receipts say 3.0 billion tokens for one person's side projects. the dollar figure is a list-price estimate from a hand-maintained pricing table, and what the subscription actually bills is a different number. the energy behind it isn't an estimate i know how to make at all. i can see exactly how much it cost me, and only vaguely what it cost. am i influencing climate change, one drum practice app at a time? i don't know. keeping the receipts at least means i'll be able to ask.