Your AI Needs an Org Chart
Running Claude Code like a real engineering team: an architect, seniors, mids, juniors, and a five-minute tool that staffs each project for you.
Every software team has levels. The architect decides what gets built and how. Seniors build the hard parts. Mids take well-defined tickets. Juniors do the mechanical work. Nobody asks the architect to rename a variable, and nobody asks a junior to design the payment flow.
Claude's models map onto those levels. Each one costs a different amount and is good at a different size of problem. Run everything on one model and you either overpay for chores or under-power design. This post describes the team, how work flows down and problems flow up, and a small tool, claude-stuffer, that sets the whole thing up by asking you a few questions in the project folder.
The team
| Level | Model | Runs as | Does |
|---|---|---|---|
| Architect and engineering lead | Fable 5.1 | the main session | Talks to you. Works out requirements, designs, writes specs. Decides who gets each task and, per project, who is on the team at all. Reads the reviewer's verdicts. Last stop for problems nobody below could solve. |
| Senior | Opus | senior subagent |
Builds features from a spec. New modules, changes that cross several files, tricky logic, anything with judgement still in it. Writes and runs tests. |
| Mid | Sonnet | mid subagent |
Well-defined, contained changes. A bug with a clear repro, a single-function change, a refactor already covered by tests, docs with substance. Runs tests. |
| Junior | Haiku | junior subagent |
Mechanical chores. Renames, formatting, moving files, README updates, git housekeeping, running a command and reporting the output. |
The rule for picking a level is the same as on a real team: hand the task to the most junior level that can do it reliably. Not the cheapest level that might manage. A task that bounces back costs more than one done right the first time.
Some signs a task is too big for the level you had in mind:
- It needs a decision the spec doesn't make. That's architect work, or at least senior.
- It touches more than two or three files for reasons other than a rename. Senior.
- "Fix the failing test" with no idea why it fails. Mid at least, and read the escalation section.
- The spec says "clean up" or "improve" without saying what done looks like. Not ready for anyone yet.
Roles beside the ladder
A team is not only a ladder. Some people review, some test, some go and find things out. Four roles do that here, each placed at a level by how much judgement it needs, not by how important it sounds.
| Role | Level | Model | Does |
|---|---|---|---|
| Reviewer | Senior | Opus, no code edits | Reads a finished change against its spec, runs the tests itself, returns PASS or FAIL with reasons tied to files and lines. Writes only its review. |
| QA engineer | Mid | Sonnet | Writes the tests from the spec before anyone builds, runs them, confirms they fail for the right reason. Never writes the implementation. |
| Scout | Junior | Haiku, no code edits | Answers the architect's questions about the codebase: where is this used, what calls that, which tests cover it. Facts with file paths and line numbers. Writes only its notes. |
| Project manager | Mid | Sonnet, no code | Keeps the backlog and the ticket log, grooms items until they are ticket-ready, writes status. Never sets priorities. |
Why those levels: a reviewer weaker than the author misses exactly the bugs you care about, so the reviewer is Opus. QA sits one level below the builder because turning a good spec into tests is well-defined work; the hard decision, what "done" means, was already made in the spec. The scout does mechanical reading.
Two things make the roles work. They cannot touch the code: Claude Code lets an agent's tools be restricted, and the reviewer and the scout write only their own report file, so they physically cannot "fix it while I'm here". And they are not on the ladder: roles don't escalate through each other. If QA's tests are themselves broken, that is a mid-level failure and goes to the senior like any other.
Skip all three for junior chores. A rename does not need tests written in advance or a senior review, and paying for them is how a cheap setup turns expensive.
Project roles, and who decides them
The seven agents above are the bench every project gets. A real company does not staff a game and a system utility with the same people. Deciding who is on the team is the engineering lead's job, and here the architect wears that hat too.
Two axes, not one. A level says how much judgement a task needs and picks the model. A role says what domain knowledge the agent carries, and that lives in its prompt: the engine's conventions, the style guide, the toolkit, the packaging rules. "Graphics designer" is a role. Its level depends on the task:
| Task | Level | Model |
|---|---|---|
| Decide the visual direction, palette, how UI and game art fit together | Senior | Opus |
| Produce an SVG icon set or a CSS theme from an approved style guide | Mid | Sonnet |
| Resize sprites, rename asset files, regenerate a sprite sheet | Junior | Haiku |
So a project does not get "a designer". It gets a designer prompt, and the architect still picks the level per task.
The rule that keeps staffing honest: a role exists only if the backlog has tasks for it. Six agents nobody calls are six ways to misroute work. Start with the bench, add a role when the third task of that kind shows up.
One honest limit: a design agent produces what can be written as code or text. SVG, CSS, shaders, procedural or pixel art as data, style guides, asset briefs. It cannot paint. For raster art the role writes the brief and judges the result.
Setting it up: claude-stuffer
Writing seven bench agents, a routing rule and a handful of project roles by hand is twenty minutes of copy and paste, and the second project is where you start cutting corners. So the setup is a single script. It has no dependencies beyond Python 3 and it lives at github.com/kpatr1981/claude-stuffer.
curl -fsSL https://raw.githubusercontent.com/kpatr1981/claude-stuffer/main/claude-stuffer -o /tmp/claude-stuffer
sudo install -m 0755 /tmp/claude-stuffer /usr/local/bin/claude-stuffer
claude-stuffer init
init installs the bench and the routing rule into ~/.claude. It runs once per machine. Then, in each project:
cd ~/code/my-project
claude-stuffer
It looks at the folder first (languages, test command, docs, whether it smells like production) and asks you to confirm or correct. Then it proposes roles for the kind of project, lets you keep or drop each one, offers add-ons such as a read-only security reviewer when the project touches servers or credentials, and takes any custom roles you want. It ends by listing exactly what it will write and asking once.
What it writes is small and readable: one Markdown file per role under the project's .claude/agents/, and a short "Team" block in the project's CLAUDE.md that tells the architect who is available and what the routing is. Each role file carries the project's description, the documents to read first, the stack, the test command, the role's rules, and the same two-attempt failure protocol as the bench. Nothing is hidden. Edit the files freely; re-running the tool keeps your edits unless you ask it to overwrite. Claude Code's /agents wizard has been removed; the official way to add or change subagents is now to edit these files, which is what this tool does.
For each piece of work, claude-stuffer ticket <slug> creates the handoff folder described below, with the spec template and the two shared record files.
Running the same command in the same folder later is how you revisit the team when the backlog changes shape. It never deletes anything; it tells you which generated files are no longer part of the team so you can remove them by hand.
A worked example: logsieve
Say I have a small Go tool called logsieve. It tails logs from several hosts over SSH, filters them with rules, and raises alerts. It has tests, and it talks to production boxes. Here is the staffing session, exactly as the tool prints it:
$ cd ~/code/logsieve && claude-stuffer
claude-stuffer 0.2.0: staffing a project. Enter accepts the default in brackets.
Project name [logsieve]:
What is it, in one line [Tails and filters logs from several hosts over SSH.]:
Kinds:
1. cli cli-developer
2. service backend-developer, db-engineer
3. web frontend-developer, backend-developer, ui-designer
4. mobile app-developer, ui-designer
5. game game-designer, gameplay-programmer, artist
6. data data-engineer, analyst
7. infra site-ops
8. library library-developer, docs-writer
9. content tech-writer, editor
10. other (custom roles only)
Kind [infra]: cli
Languages/frameworks, comma separated [Go, Shell]:
Test command ("none" if there is no test suite) [go test ./...]:
Touches production systems (servers, cron, live databases)? [Y/n] y
Documents a role must read first, comma separated [README.md, docs/design.md]:
Keep cli-developer (senior): Builds the command-line tool: flags, help text, exit codes and the non-interactive mode? [Y/n] y
Add site-ops (mid): Runs and changes servers and deployments (use senior for production changes)? [y/N] n
Add security-reviewer (senior): Reviews changes for leaked secrets, trust boundaries, dependency fetching and credential permissions; no code edits? [y/N] y
Custom roles (blank name to finish).
Role name (lowercase, hyphens): alert-rules-author
What it does (one line): writes alert rules and their fixtures from real log samples
What it must know (comma-separated paths or notes): docs/rules.md, samples/
Level (senior, mid, junior) [mid]:
Role name (lowercase, hyphens):
Files for logsieve (cli-developer, security-reviewer, alert-rules-author):
would write .claude/agents/cli-developer.md
would write .claude/agents/security-reviewer.md
would write .claude/agents/alert-rules-author.md
would update CLAUDE.md
Write these files? [Y/n]
Staffed logsieve:
wrote .claude/agents/cli-developer.md
wrote .claude/agents/security-reviewer.md
wrote .claude/agents/alert-rules-author.md
updated CLAUDE.md
Check with claude-stuffer show, or start a new session; agents are loaded at startup.
Notice the kind: the folder has an ops/ directory, so the tool guessed infra, and I corrected it to cli. Detection gives defaults, it does not decide. The block it added to the project's CLAUDE.md is what the architect reads at the start of every session:
## Team (staffed 2026-10-10 by claude-stuffer)
Global bench applies (`senior`, `mid`, `junior`, `reviewer`, `qa`, `scout`, `pm`). Project roles, level chosen per task by the architect:
- `cli-developer` (Opus; senior): Builds the command-line tool: flags, help text, exit codes and the non-interactive mode.
- `security-reviewer` (Opus; senior): Reviews changes for leaked secrets, trust boundaries, dependency fetching and credential permissions; no code edits.
- `alert-rules-author` (Sonnet; mid): Writes alert rules and their fixtures from real log samples.
- Production: changes are specified, backed up and verified one at a time; the user applies anything the spec marks as manual.
The resulting team, bench plus project roles:
The sequence of work for one feature
I ask for a --since 2h flag that only shows lines newer than a relative time. Here is who does what, in order:
Six handoffs. Fable wrote one spec and read three short reports. Opus did the building and the two reviews. Sonnet wrote the tests. Haiku answered two questions. Nothing ran on a model bigger than it needed.
The same project, a chore that goes wrong
Now a rename: package filter becomes sieve. That is junior work. Here is what happens when it breaks:
Had the junior's report said "looks like a logic problem", the architect would have skipped the mid and gone straight to the senior. That is the one skip the ladder allows, and it is based on the report, not a hunch.
How work flows down
The example shows the general shape:
- You talk to the architect. That's the normal Claude Code session.
- The architect asks the scout what it needs to know, then writes a spec. A spec has four parts, and if any is missing the task is not ready: the files to touch, the constraints, what "done" means, and which tests must pass.
- QA writes the tests from the spec. They fail, which is the point.
- The architect hands the spec and the tests to the right level or role. A subagent starts with no memory of your conversation, so the spec has to be complete on its own.
- The builder does the work, runs the tests, reports.
- The reviewer, and any read-only add-on such as a security reviewer, check the diff and return a verdict.
- The architect reads the verdicts and tells you.
For a junior chore it collapses to two steps: the junior does it, the architect glances at the report.
Parallel work
The ladder is sequential, the team is not. On a real team several people work at once, and the architect here does the same: it can launch several subagents in one turn and gets each report as it lands. Run in parallel anything that is independent: two mids on unrelated bugs, QA writing the next ticket's tests while the senior builds this one, the reviewer and a security reviewer on the same diff. Both of those last two are read-only, so they cannot collide.
Two things never run in parallel. Ladder steps: QA before the builder, the builder before the reviewer, each escalation after the previous report, because the next step's input is the previous step's output. And two writers on overlapping files: subagents share the working tree, so either serialize them or give each its own git worktree and merge afterwards.
Parallel saves time, not tokens. A subagent costs the same alone or beside another, because the cost is per agent, not per minute, and the agents do not pay for talking to each other: a handoff is one prompt in and one line back, and the reports go through files. What each agent pays for is its own reading. Every tool call is a new request that carries the whole transcript so far, so a scout that answers eleven questions over 147 tool calls costs about 300k tokens whether or not a junior is working next to it. The main session is a transcript too, and the longest one of the day paid for its full length on every turn.
The same arithmetic says when not to use the ladder. Each handoff is a cold start that re-reads the spec and the files, so a ticket that passes through the scout, QA, the builder and the reviewer pays for four readings of what one session would read once. On a feature that is cheap insurance. On a one-file fix it is most of the bill, and slower too. If the spec would take longer to write than the change, give it to one junior without QA or review, or make the change in place.
The architect does not write code. Its value is in the design, the spec and the review. If it is typing implementation, something below it should have been asked instead.
Isn't this waterfall?
It looks like it: spec, then build, then test, in that order. The difference is scale and feedback. Waterfall designs the whole system up front and never loops back. Here a spec is one ticket, usually an hour of work or less, and the loop back is built in: tests fail, the report comes up the ladder, the architect learns something and the next spec is better. It is closer to a kanban board with a strict definition of ready than to a phase-gated project.
The one place it can turn into waterfall is the "not ready" gate. If you find yourself writing long specs for work you don't understand yet, stop. Exploration, spikes and "what would it take" questions have no spec and no tests, and they belong in the main session, not on the ladder.
How problems flow up: the escalation ladder
The person who wrote the code gets the first look. The author has the context and is the cheapest person to put on it. Each level gets two attempts. If the tests are still failing after that, it stops and writes up what it found. No thrashing.
The write-up goes to the architect, which hands it one level up, together with the original spec:
junior -> mid -> senior -> architect
Each step up gets the spec, the failure report from the level below, and the same two attempts. If the senior can't solve it either, the architect takes it on itself. That is the one case where the architect works on code directly, and it arrives with three failure reports' worth of context.
Rules that keep the ladder honest:
- Don't skip levels. Most failures are small and the mid solves them. Jumping to the senior every time turns a cheap ladder into an expensive one.
- The one allowed skip is evidence-based. The junior's report says whether the failure looks mechanical or a logic problem. If it says logic, the mid is skipped.
- Never send it back to the same level. A second mid with the same spec will most likely fail the same way.
- No tests, no ladder. In a project without a test suite, code changes go to the senior regardless of size, and the spec says how to verify instead.
A failure report contains: what failed, with the actual output; what was tried, each attempt; the author's hypothesis and whether it looks mechanical or logic; and exactly which files are modified, because the changes stay in place for the next level.
How handoffs carry context
The obvious objection to all this: a subagent starts cold, so doesn't every handoff lose what the previous one knew? It would, if the handoff were a chat message. It isn't. Subagents share the working tree, so what one agent writes to a file the next one reads.
One folder per ticket, docs/tickets/<slug>/, one file per step:
- spec.md, written by the architect, read by everyone.
- scout.md, what the scout found.
- tests.md, QA's tests and their failing output.
- build.md, the builder's report. On escalation the next level reads it as it stands, with the real test output, not a paraphrase.
- review.md, the reviewer's verdict and reasons.
The architect's handoff shrinks to one line: the folder, the role, which files to read. Each agent's chat reply is one line too: the verdict and the path. The file is the record. Keep it that way, or the architect ends up reading everything twice.
claude-stuffer ticket <slug> creates the folder with a spec template that has the four parts, plus What and Notes, already as headings, plus two files the team keeps across tickets: LOG.md, one line per finished ticket, and FACTS.md, where the scout appends durable facts about the codebase so the next scout doesn't re-derive them.
Keeping the team efficient
A few more rules, each one cheap and each one there because of a cost that showed up in practice:
- Scope guard. A builder that needs a file the spec does not list stops and reports instead of widening the change. Scope creep inside a subagent is invisible until the review, and by then it has cost a full run.
- Batch chores. Three renames in one project are one junior ticket, not three. Handoff overhead is per ticket, so size the ticket to the overhead.
- A project manager on the bench.
pm, Sonnet, keeps the backlog and the ticket log, grooms items until they are ticket-ready (what, where, done), writes a one-screen status on request, and chases what is blocked on whom. No code, no priority changes, closes nothing without a reviewer PASS or your word. It takes the bookkeeping off the architect, which is the cheapest work on the architect's plate and the easiest to forget. - Facts outlive tickets. The scout reads
FACTS.mdbefore searching and appends what it learns. The tenth ticket in a project should cost less research than the first. - Narrow scout questions. One scout per ticket with three or four questions, not one sweep that answers eleven. A broad sweep reads most of the codebase and then re-sends all of it on every later step. The next ticket's scout starts from
FACTS.mdanyway. - Compact between tickets. The main session re-reads its whole conversation every turn. Once a ticket is closed,
/compact. The next task starts from the ticket folder, which is the point of keeping handoffs in files. - The log makes patterns visible. When every junior ticket in
LOG.mdshows an escalation, the tickets aren't junior-sized. When the senior keeps escalating to the architect, the specs are leaving too much open. You don't have to remember this; the file shows it.
Pros and cons
Pros
- Cost tracks difficulty. Most of a day's work is renames, small fixes and docs. Those run on Haiku or Sonnet, and the expensive models only see work that needs them.
- The architect stays sharp. The main session never fills its context with diffs and test output. It keeps the whole picture of the project and the conversation.
- Specs get written. Because a subagent starts cold, the architect has to produce a complete spec before anything happens. That catches vague requests early.
- Failures land on the cheapest capable level first, and attempts are bounded, so a stuck agent cannot burn an hour of expensive tokens going in circles.
- Each report is a review point. Mistakes are caught between levels instead of at the end.
- Staffing is a five-minute conversation, not a copy-and-paste session, so you actually do it per project.
Cons
- Latency. A task that one model would do in a single pass now takes a spec, tests, a handoff, a report and a review. For a one-line change in a file you have open, that is slower than just asking.
- Context has to be written down. A subagent knows only what is in the ticket folder. The files fix most of the loss, but anything the architect knew and did not write is still gone, and a weak spec still produces weak work.
- Escalation is cheap on average, expensive in the worst case. The evidence-based skip helps only when the junior classifies the failure correctly.
- The architect is still the router. Subagents cannot launch each other, so every escalation passes through the main session. With file handoffs it carries a path and a verdict rather than the content, which keeps its context small, but it is still one serial decision point.
- Level boundaries are fuzzy. Misrouting is the most common failure.
- It needs tests to work well. Without them the ladder is replaced by "everything to the senior", which is safe but loses the savings.
- Staffing can be over-done. The tool makes adding a role cheap, which is exactly when you should be stingy with it.
When to use it
Use it when you spend real money on Claude Code, your projects have tests, and most of your requests are concrete changes. Even then, size the process to the task: the ladder is for features, and a chore smaller than its spec is cheaper done by one junior or in place. Skip it, or run only the architect plus the senior, when you mostly explore or work in codebases without tests. Trying it for a week and removing it costs nothing: the tool never deletes anything, and undoing it is deleting the files it told you it wrote.
Checking it works
- Start a new session. Agents and rules are loaded at startup.
- Run
claude-stuffer showin the project folder and check that the seven bench agents and your project roles are listed. - Ask for a small feature. The main session should write a spec, start
qa, then a builder, thenreviewer. It should not write the code itself. - Open a project with no
.claude/agents/folder and describe it. The main session should tell you to runclaude-stuffer, or propose a team and wait for your OK. - Ask for a rename across a package. It should go to
junior. If the rename breaks a test, watch where the fix goes.
Tips
- Pass the full spec, every time. On escalation, pass the failure report too, verbatim.
- Two attempts is a default, not a law. For a flaky integration suite allow three. For a one-line rename, one. Put it in the spec.
- Keep the architect out of the code. If the main session is editing files, ask which level should have had the task. Usually the answer is a missing or vague spec.
- Read the verdicts. Check that "tests pass" means the tests in the spec, not a subset.
- Don't let the reviewer become a second builder. If a reviewer report reads like a patch, its read-only tool list has been loosened.
- Watch the ladder for patterns. If every junior task escalates, the tasks aren't junior-sized. If the senior keeps escalating to the architect, the specs are leaving too much open.
$ ls ./see-also/
$ ./buy-me-a-coffee ☕ // found this useful? I write these in my spare time