An agentic design team in Claude Cowork
Five AI agents, each named after someone I admire, each with a distinct charter — strategy, voice, design systems, team operations, and research. The harder problem turned out not to be the agents. It was the structure underneath them: what gets remembered, who owns it, what has to be reviewed before it counts as true, and how the whole thing stays cheap to run. This is the ongoing log of what's working, what's breaking, and what I'm learning about delegating judgment, not just tasks.

Hypothesis
Delegation, not automation
I lead a distributed design team across five countries. Most of what fills my week isn't design — it's context-switching between strategy, content, systems thinking, team operations, and research, often inside the same hour. AI tools promised to speed that up. What I actually needed was narrower: a way to delegate the way I would to a trusted senior peer, not a way to generate more output faster.
That distinction came out of a tension I kept running into. The pressure to move faster with AI was real, but speed on its own wasn't the goal. I mapped what "impact" actually meant across five different registers: user impact, strategic impact, revenue impact, differentiation, and team capability. My goal was to find out how AI could help me improve my systems thinking — relating strategy, decisions, team management, and different product visions at the same time.
So I built a team of five agents using Claude Code — each with a named identity, a distinct charter, and explicit instructions for how they should work with me, not just what they should produce.
Setup
Five charters, one discipline
Each agent is named after someone I admire, which isn't decoration — it's a constraint. Naming an agent after Julie Zhuo or Teresa Torres forces a specific point of view into the brief instead of a generic "helpful assistant" register.
Maite (strategic partner, named for my mother) exists to stress-test my reasoning, not support it. Her only brief is to help me think clearly, to help me raise the bar and be better. Chimamanda (voice and content) writes and shapes public-facing work in a voice that still sounds like me, not like AI smoothing my rough edges into nothing. Ruth (design systems) holds the token architecture and component decisions for our design system, and asks whether a decision scales before I commit to it. Julie (team operations) is focused entirely on whether my team has what it needs to do its best work — rituals, onboarding, collaboration frameworks. Teresa (research) thinks in opportunity trees and evidence, and asks the question everyone else skips: are we researching the right thing at all.
Maite stays in the main thread with me, always. The other four run as separate subagents I invoke on demand. That split isn't cosmetic. Maite's job needs the full history of a conversation to see the whole board — an isolated call would hand her a blank page and strip her of exactly that. Ruth, Julie, and Teresa need the opposite: a clean, focused context that isn't cluttered with whatever I was discussing with someone else five minutes earlier. Deciding which parts of a system get memory and which get a fresh start every time turned out to be its own design decision, not a technical detail.
The standing rule across all five: the team's job is not to agree with me. It's to think harder than I would alone. Honesty over comfort, every time. And when something needs real alignment — a framework, an article, a system decision — we discuss before we build. Not build, show, and revise endlessly. Align once, then move.
Structuring the second brain
Five points of view immediately created a harder problem: where does everything they know actually live, and how do I stop that pile of information from turning into its own mess?
Each agent has one file that is its territory, not the whole system. Design systems knowledge doesn't leak into team-operations knowledge, and neither agent has to read the other's file to do its job. Every file opens with a small frontmatter block — who owns it, when it last changed, where the update came from. That sounds minor. In practice it's the single most useful discipline in the whole setup, because I can look at any fact my second brain holds and answer "is this still true, and where did it come from" in five seconds instead of scrolling back through months of conversation.
Files also outgrow themselves, and when they do, I split them. One file covering everything I track across my product portfolio kept growing until it stopped being fast to navigate. I split the two areas getting the most attention into their own files and left the rest in the shared one. The structure follows how the work is actually distributed, not a taxonomy I designed on day one and never revisited.
None of the agents load everything by default either. Each one is told which file to open, and — more importantly — when not to bother: a simple question gets a direct answer, no file loading, no performing thoroughness it doesn't need. That single rule is what keeps the token cost of the system down. An agent that reads five files before saying good morning isn't a thinking partner. It's a slow, expensive one.
Nothing enters memory unreviewed
What I underestimated when I started is that a second brain fed automatically will eventually contradict itself. Meeting notes, research reports, and tickets now flow into this system through a few different tools, and every one of those sources is wrong sometimes, incomplete sometimes, or captures a conversation that never actually resolved into a decision.
So nothing goes straight from a source into permanent memory. Everything lands as a draft first. I read it, keep what's accurate, correct or drop what isn't, and only then does it get merged into the file it belongs to. Not every source earns the same trust once it's in, either: a meeting summary stays marked unconfirmed until I've actually confirmed it, because a transcript isn't the same thing as a decision. A finished piece of research is stronger evidence — but it's still one signal about a specific group of people at a specific time, not a verdict that overrides everything else I know for other reasons.
The system can also be asked to check itself for staleness — comparing what a file says against how long it's actually been since the thing it describes last changed. Written down is not the same as still true.
What I'm learning
This is a pattern, not a personal quirk
None of this is really about having five agents with names. It's about deciding, on purpose, how information should be structured before you let AI act on it. The transferable version is smaller than it looks: one owned file per domain of the work, a lightweight way to know who wrote a fact and when, a review step before anything becomes shared truth, and a way to notice when something has gone stale.
That's exactly as useful at product level as it is for a team of one. A product manager and a product designer already know things about a product that live in different heads — strategy in one, user evidence in the other, decisions nobody wrote down anywhere. Giving that shared context the same treatment — one place it lives, one owner per piece, one review step before it counts as fact — isn't a research project. It's a decision about how a team wants to work, worth making on purpose rather than letting whichever tool arrives first decide it by default.
From prompting, to context, to harness
When I started this project, the conversation around AI was mostly about prompting: how you word the ask. Then it moved to context engineering: what you feed the model, how you structure the information it has access to. Both still matter here — every file and every frontmatter block above is context engineering.
But the layer that changed the most while I was building this was neither of those. It was the harness: the system around the model that decides which parts run in isolation and which stay in the same thread with full history, which processes get encoded as a repeatable skill instead of re-explained every time, where a human has to approve something before it becomes permanent, and how the whole thing checks its own freshness instead of assuming it's still correct. That layer isn't a prompt, and it isn't a context file. It's closer to engineering a small piece of infrastructure — and it's what decides whether a system like this gets more useful over a year, or quietly rots.
I expect this to keep moving. The tools now feeding this system didn't exist when I started it. The next layer will probably make parts of what I built here look primitive too. That's fine. The point was never to finish the system. It's to keep it honest about what it actually knows, at whatever layer that requires.
Speed on gathering information for human decisions
The goal of this team is to serve one person, and that means it keeps changing shape as I do. When I started, none of the MCPs I rely on now existed. I now have a meeting-transcription MCP, a ticketing MCP for what each product actually has in flight, and a research-reporting MCP — all three feeding the review pipeline above: drafted, checked by me, then merged. What used to take an hour of reading meeting notes and half-remembering a decision from three weeks ago now takes minutes, and I trust it more, not less, because nothing reaches the parts I actually rely on without me looking at it first.
Still figuring out how much of this scales past a team of one. But the parts that do, don't need five names attached to them. They just need a team willing to structure what it knows on purpose.

Next experiment