A kaizen review of your AI conversations
It's probably the main demo in YouTube tutorials at the moment, touted as one of the great practical use cases for AI, especially since Claude Cowork picked up scheduled tasks recently: set up a scheduled task in Cowork that emails you every morning with your calendar for the day and your outstanding to-dos. On the face of it, that's a fair description. It runs while you're asleep, the email's waiting when you wake up, and you didn't have to lift a finger.
But have a think about what's actually happening. Every morning, a large language model wakes up, reads your calendar, reads your task list, and writes them into a tidy email. That is a data pull and a bit of formatting. A script could do it for nothing. Instead it's spending model capacity, and on the newer metered plans real budget, on a job that has no judgement in it at all.
Software developers have a rule of thumb for this. Do something once and you just do it. By the third time you've done the same thing by hand, that's the signal to stop and automate it. The chore itself is rarely the real cost; what adds up is the same chore repeating, quietly, with nobody deciding that it should.
The same logic applies to how we work with AI assistants, and almost nobody is applying it. We use AI to run tasks. The more useful move, most of the time, is to use AI to find the rule behind the tasks we keep running, and then build the boring bit of infrastructure that means we never have to run it by hand again. Find the rule, don't run the rule.
Kaizen, briefly
Kaizen is a Japanese word meaning "improvement", absorbed into manufacturing practice through Toyota's production system and into the IT world through Lean and ITIL. The principle is small, regular improvements rather than occasional big rewrites. You don't tear the line down and rebuild it. You walk it regularly and fix the small things you notice. The accumulated effect, over months, is more reliable than any single big change would have been.
In practice it's more a habit than a moment of insight. The discipline is to keep looking, and to keep making small fixes. The cost of any single kaizen pass is low. The cost of never running one is that small inefficiencies compound until they're large ones.
That's what this piece is about, applied to how you use an AI assistant. Rather than a one-off overhaul, the move is a periodic look at what your sessions are actually costing you, with a small fix or two each time.
The cost behind it
I've written before about the AI cost squeeze: the price rise most people are waiting for is already here, just routed through quieter channels than a headline number on a pricing page. One of those channels is agentic usage. The tools that do the most work for you, the ones running multi-step tasks on your behalf, are also the ones quietly consuming the most tokens per task.
That stopped being abstract recently. From 15 June 2026, Claude subscriptions get a separate metered credit for Agent SDK usage. Headless and programmatic usage now draws from a credit that runs out, rather than sitting inside a flat subscription. When the credit's gone you stop, or you pay overage. That changes the maths. Every token you spend re-explaining something you've already explained, or running a task by hand that a script could do, is now coming out of a finite pot.
There's a sharper question to ask than "how do I use AI more", and it is this: where am I spending tokens on things that shouldn't cost tokens at all? The answer is sitting in your history.
The technique
In theory I try to apply this principle in real time. As I'm working, if I notice I'm doing the same kind of task a second or third time, I stop and turn it into something reusable. In practice, I miss things. The patterns that quietly drain tokens are harder to see in the moment than they are in aggregate, and the small annoyances rarely feel worth pausing the actual work for. So I rely on a second mechanism: a periodic pass over the record to surface what I missed while I was busy.
Claude Code keeps a full record of your conversations on your own machine, in the ~/.claude directory. Every project, every session, the lot. Most people never look at it; it just sits there, accumulating.
That record is the raw material. You can point Claude Code at its own history and ask it to read back through what you've actually been doing, looking for the patterns that are costing you. It reads back what you actually do with the tool, session after session, in your own words, which is rarely quite what you think you do.
It's a bit like going back through a project's change log after the fact. Any single change looked reasonable on the day. It's only when you read the whole run together that you notice you've made the same sort of fix eleven times, and that the eleven fixes were really one missing thing.
Be warned that the output can be a little confronting. You will probably recognise yourself in it, asking for the same thing a dozen different ways, pasting in the same context, sending the assistant off to rediscover something it had already worked out a week ago. That recognition is the point. The review itself is cheap: one focused session. What it can save you is the same cost paid over and over for months.
The prompt
Here's a prompt you can use more or less as-is. Adjust the specifics to your own setup, but the shape is the point.
Context. Over time I've built up a lot of conversation history with you, and
I suspect a fair amount of it is repetitive: the same kinds of requests, the
same context pasted in, the same setup steps, done again and again. Each
repeat costs tokens and time. I'd like to find those patterns and replace
them with something reusable.
What I'd like you to do. Please have a look through my Claude Code
conversation history in ~/.claude (the projects and history folders, plus any
CLAUDE.md files). Look for:
- Requests I make repeatedly, or tasks that follow the same shape each time
- Long context or explanations I paste in regularly that could live in a file
- Multi-step sequences that could be a single script, slash command, or subagent
- Information you have to rediscover each session because it isn't written down
What I'd like back. A ranked list of the patterns you found. For each one: a
short description, roughly how often it comes up, a rough token or time cost,
and a specific proposed fix (a script, a CLAUDE.md entry, a slash command, an
MCP, a reference doc, a subagent). Put the highest-impact, lowest-effort
changes at the top. If you spot anything else worth automating that I haven't
framed here, please add it.
The goal is to set up the enabling infrastructure once, so next time I can get
things running with as few tokens as possible, rather than paying the same
cost on every conversation.
A few things about why it's written this way.
It leads with context before the ask, because the assistant does better work when it knows why you want something, not just what. Telling it the goal is to cut repeated cost shapes everything it then looks for.
The output is ranked deliberately. You want the highest-impact, lowest-effort changes at the top, because those are the ones you'll actually do this week, rather than a flat wall of observations you skim once and forget.
Asking for a rough token or time cost against each pattern matters more than it looks. The estimates won't be precise, and that's fine. You're after a sense of scale, enough to tell a five-minute annoyance apart from a real drain.
The last line leaves room for the assistant to add things you didn't ask about. The whole point of the exercise is that you can't see your own patterns clearly, so some of what comes back should be things you'd never have framed yourself.
DadOps, a worked example
The name is, admittedly, a bit daft. DadOps started as a parenting thing. I kept asking the assistant for the same kinds of help over and over: a play date finder, age-appropriate activity suggestions for a rainy Saturday, a way to keep track of what had worked the last time. It was the same shape of task again and again, and I wanted a single place to house the small tools that solved it rather than re-explaining what I needed each time.
Over time the same place became a hub for everything else. Sites I run, services I have set up, the VPSs I host them on, what lives on each one, how they relate. The original "dad" framing stuck, but the toolset grew into something closer to a personal operations layer.
When I ran this review on my own history, the biggest pattern was embarrassingly simple. I kept re-explaining my own setup. Which sites I run, where they're hosted, how they relate to each other, which analytics property maps to which domain. Every time a session needed that context, I either typed it out again or the assistant went and worked it out. Same information, paid for repeatedly.
So DadOps now holds a piece of live documentation of all of it. Kept current, structured for the assistant to read rather than for me to read, because I don't need a pretty document; I already know my own estate. A session that needs to understand how things fit together reads the document instead of interrogating me or rebuilding it from scratch.
The same logic extended to external services. I was regularly asking the assistant to pull something from Google Analytics, or Search Console, or Ezoic, and each time it had to work out how to get in, what to ask for, and how to make sense of what came back. So DadOps now has structured proxies for those services: a defined, repeatable way in, so "get me the GA numbers for this site" is a known route rather than a small research project every time.
A more recent step was a small Python CLI. There are jobs where the assistant needs information that only my machine has, or only my network can reach, and the scriptable bits of getting that information out are tedious and repeat across many sessions. The CLI runs locally, does the boring part, and hands the answer back in a no-fluff, machine-readable shape. Claude reads exactly the thing it would have rebuilt from scratch otherwise, rather than a polished report it doesn't need.
There's also a Claude Code skill that wraps the lot. Rather than dropping all of this into a top-level CLAUDE.md and bloating the context window every session, the skill knows when DadOps is relevant and pulls in just the right slice of detail when it is. It's a small example of the same principle one layer up: the rule for when DadOps matters is itself encoded once, in the skill, rather than reloaded into every conversation.
None of this is sophisticated. It's a folder of documentation, a small CLI, and a Claude Code skill that wires them in. But it came directly from reading my own history and noticing what I kept paying for. I didn't design DadOps up front. The review told me what it needed to be.
Categories of fix
When you run the review, the fixes it suggests tend to fall into a few shapes. It's worth knowing them, because once you've seen the categories you start spotting candidates yourself.
The first is a reference document. Something you keep explaining becomes a file the assistant reads instead. DadOps' network documentation is one of these. When I ran the same review on my Cowork history, the equivalent finding was that nearly every content session opened by rediscovering the same things: my brand structure, which audience sits where, my writing workflow. All of that should be, and now is, written down once and pointed at. In Claude Code the natural home for it is a CLAUDE.md file the tool reads automatically.
Next, the script. A sequence of mechanical steps you keep asking for by hand gets written once as code and runs for nothing after that. The morning-schedule email from the top of this piece is exactly this. So is anything that's really a data pull plus formatting.
Third comes the slash command or the subagent. This is for the repeated task that does need some judgement, so you can't reduce it to a plain script, but it follows the same shape every time. My Cowork history was full of one of these: take an external thing, a report, a competitor's post, a piece of legislation, and run it through my usual set of lenses. Same shape, over and over. That's a packaged command, not a thing to re-describe from cold each time.
The fourth is a structured access point, a proxy or an MCP, for an outside service. DadOps' analytics proxies are the example. Rather than the assistant working out how to reach Google Analytics every time, there's a defined route in.
And there's a fifth that's less about saving tokens than saving mistakes. The review can surface a rule you've already written down but that keeps getting ignored. In my case, my writing workflow says to agree an outline before drafting, and my history showed it being skipped anyway, more than once. That's a useful signal in itself. A rule that doesn't get followed usually needs to become infrastructure, a checklist or a step that's genuinely hard to skip, rather than a line in a document everyone means to honour.
One more thing worth saying about all of these. You don't have to build them by hand. The same assistant that found the rule can write the script, the skill, the CLAUDE.md entry, the proxy. Finding the rule and building the rule are both, increasingly, Claude's job. That's a whole subject in its own right and worth its own piece, so I'll come back to it. For this one, the move is just to do the review.
Beyond Claude Code
This works best where you can get at the raw material. Claude Code keeps full transcripts on your machine, so the review can read exactly what happened. Cowork is similar: it keeps your session history, and you can ask it to read back through it the same way. I did exactly that while writing this piece. Two of the examples above came straight out of it: the rediscovery of my brand structure in every content session, and the recurring "take an external source and run it through my usual lenses" shape. Same technique, different surface, same kinds of fix on the other side.
The web version of Claude is a softer case. Claude.ai now has a memory feature, free for everyone since March 2026, which summarises your conversations and carries a synthesis forward. On the paid plans there's also search across past chats. Both are useful. But they give you a tidied, summarised layer rather than the raw logs, so you can still ask Claude to look for things you repeat, you just can't point it at the full unedited record and have it count the cost precisely.
The principle holds across all of them, though, and well beyond Claude specifically. Any assistant you use regularly is accumulating a record of how you use it. That record is the most honest description of your own workflow you'll ever get, because you didn't write it to look good. It's just what you did. Reading it back, with the specific question "what here should have been infrastructure", is the move.
This is really an old discipline wearing new clothes. Anyone who's worked in a managed IT environment will recognise it as a problem review, or a lessons-learned exercise, or a kaizen pass: you stop, you look at the record of what actually happened, and you fix the system rather than the symptom. The novelty is only that the record is now a folder of AI conversations, and the cost of not looking is metered by the token.
When I ran the review on my own history, what stuck with me most was the price of it. One focused session, set against months of paying the same costs over again. It paid for itself the first time I didn't have to explain my own network from scratch.
Reading your history back is something you can do yourself, and I'd encourage it. If you'd rather have help turning what it surfaces into actual infrastructure, that's the kind of work I do through Solutions Delivered, and it's the backbone of my AI Foundations service. I also write these techniques up most weeks in Practically AI, my weekly newsletter.