Insights

How to Build an AI Employee (And Why Yours Keeps Handing You Slop)

I ran an audit on my own AI setup last week, expecting to tidy up a couple of things.

What I found was sixty skills available to it at any given moment, twenty-one of which were duplicates. One plugin had been saved twice, ninety seconds apart, back in April, and had been sitting there quietly ever since. So every time I asked for something, six near-identical descriptions were competing to answer, and one of them won. Sometimes the one that knew my brand. Sometimes the generic one that didn't.

That explained rather a lot about output I'd been blaming on the model.

If you've been working with AI long enough to be past the novelty of it, long enough to have built a few things, and you keep receiving work that is technically correct and entirely useless, you may be somewhere similar. This is not a knowledge problem. You know how to prompt. It's an execution and architecture problem, and those have architecture solutions.

Where this starts

For about a year I had been treating my AI as a very capable stranger I could brief from scratch each morning.

That approach has a ceiling, and I'd hit it without noticing. Every conversation began from nothing. Every piece of context I cared about — who my clients are, how I like work done, what we never compromise on — lived in my head and got retyped, differently each time, depending on how much attention I had that day.

What I needed was to treat it as someone who works here.

That turns out to require almost exactly what a new hire requires, and the list is unglamorous. A place to work. Context about the business. A brief before a task. A task with a finish line. Some way of checking its own work. A standard to be measured against. A few recurring responsibilities. And clarity about what it can decide alone.

None of that is optional with a person. It's what makes someone's first month useful rather than expensive. I don't know why I assumed it would be optional here.

The audit found sixty skills and twenty-one duplicates. When AI hands you something generic, the model usually isn't the problem — you've built a room where six people claim the same job and nobody decided who owns it.

The project has to explain itself

The first change I made was to write the business down.

Not in a prompt I retype. In three markdown files sitting in the folder Claude Code opens, which it reads before anything else happens.

CLAUDE.md says how I like work done: small pieces, plan before producing anything client-facing, stay inside the task you were given, and when you finish, tell me what changed, what you checked, and what still needs my judgment.

roadmap.md says what matters this week, along with a list of what explicitly doesn't.

review.md says what my standards are — what stops a piece from going out, what has to be fixed first, and what is fine.

Then folders for the things that make the work specific rather than plausible: clients/ for who we serve and the words they actually use, context/ for what the business currently knows, routines/ for the work that repeats.

The out-of-scope list was the part I nearly left out, and it's the part that changed most. I'd assumed it was housekeeping. Naming what you don't want is what forces depth on what you do. Without it, ambition leaks — you ask for one thing and receive four, and the three extras are always the ones nobody needed.

There's a useful question underneath all of this. Where in your business does the context live only in your head? Whatever comes to mind first is the thing to write down.

A brief, then a task

Once the workspace could explain itself, the output improved noticeably. It still went sideways on anything substantial, and it took me a while to see that I was the cause.

I had been handing over assignments the way you'd throw something over a wall. Improve the landing page. Make this stronger. Tighten the messaging. Then feeling let down by what came back, as though the disappointment belonged to the AI.

You wouldn't do this to a person. If you gave a new team member “make the landing page better” and walked away, you'd expect exactly what I was getting, which is a confident guess.

Two things fixed it, and the order matters.

The first is a brief. Claude Code has a plan mode for exactly this, and it is the single most underused thing in the product. Before anything substantial gets built, ask for the approach rather than the output. Tell it to read CLAUDE.md, roadmap.md and review.md first, which is not a formality, and then to come back with what it would change, the smallest clean way to do it, the risks, how you'll know it worked, and what it's leaving out of the first version deliberately.

What comes back won't be right. It doesn't need to be. It needs to be something you can correct — keep it to the hero section, leave the pricing page alone, simplify the form — and correcting a proposal takes two minutes. Correcting finished work takes an afternoon. Measure twice, cut once is an old idea that survives contact with AI perfectly well.

The second is the task itself, with a visible finish line. Not “improve the landing page” but “add a waitlist form collecting name, email and company, with a success state, consistent with the existing brand.” One task, one finish line, one thing to review.

When there's no finish line, the AI has to guess what matters, and once it's guessing you've stopped managing the work and started cleaning it up.

Most of us have one task we've been handing over vaguely, because writing it out properly felt like more effort than doing it ourselves. That's usually the one to start with.

A standard, and something that runs on a schedule

When the first two pieces are working, output arrives quickly, which moves the problem somewhere new.

The constraint stops being production and becomes judgment. Is this right? Does it match what I asked for? Did something change that I didn't ask about? If everything needs reading with full attention, you've swapped a writing problem for a reviewing one and gained very little.

So the standard has to live in the workspace too, rather than in your instincts on the day. That is what review.md is for.

Mine grades everything three ways.

  • Critical, and it doesn't go out: wrong voice, an unsubstantiated claim, contradicts the client's brand foundation, another client's context has found its way in.
  • Major, and it gets fixed before anything else: my language about the client instead of their own words, no clear next step, a weak opening.
  • Minor, and it gets fixed if there's time.

Writing that down did something I hadn't anticipated. It meant I could say review this against review.md and receive a triage I could act on in ninety seconds. It also stopped the standard drifting, because it was no longer being rebuilt from memory each time.

One habit belongs alongside it. Read the diff — the before and after view of every file that changed — not only the finished output. Look at what actually moved. Surprising changes are where the risk sits, and if you asked for a headline rewrite and six other things shifted, that's the thing worth noticing — which you'll only ever do by looking.

Then there's the part that changes the shape of a week.

Every business has work that keeps things moving and never feels urgent. Somebody should read the client notes. Somebody should notice which problems keep recurring across accounts. Somebody should look at the open list and say what today ought to be.

That work suits a schedule, and it's the right place to begin. Claude calls these scheduled tasks, and the first one you set up should be useful and entirely safe rather than something that ships client work while you sleep. Mine reads clients/, roadmap.md and the open task list every weekday at seven and writes four things into context/morning-brief.md: the most useful move available today, the most important thing a client said this week that I haven't acted on, one thing quietly going wrong, and one question worth asking a client. It publishes nothing and changes nothing. It means I start the day with context rather than a blank screen.

A second one runs on Fridays and looks across every client at once, grouping problems that appear in more than one place. That one earns its keep by itself, because a problem showing up across three clients isn't a client problem. It's a gap in the system, and you can't see it from inside a single conversation.

What nobody tells you about working this way

I'd rather be straight about the parts that are harder than they sound.

Setting this up takes real time, and the first attempt won't be right. The workspace I built still has placeholders I haven't filled. Writing down who your buyer actually is, in one sentence, is difficult, which is precisely why it's worth doing. If you can't write it, your AI can't infer it, and neither can your audience.

Maintenance is not optional either, because the system decays without it. A roadmap that's three weeks stale is worse than none at all, because everything built from it inherits the wrong priorities while looking authoritative. The same is true of placeholders. A document full of brackets reads as confident and says nothing. Ten minutes on a Monday keeps it honest.

Then there's the question of what stays yours, which no system decides for you. I keep a clear line.

  • Reading, drafting, planning, reviewing, producing briefs — all fine, go ahead.
  • Publishing anything, sending to any audience, scheduling a post, touching credentials, changing an approved brand document — ask me first.
  • Billing, pricing, contracts, and anything going out under my name that I haven't read stay with me, always.

That last category matters more than the efficiency does. The point of building this well is that it frees you to be more present in the parts that need you, not less.

And none of it makes the strategy for you. A beautifully organised system pointed at the wrong buyer produces beautifully consistent content that doesn't convert. The architecture amplifies whatever you give it. If the foundation is vague, you've built a faster route to vague.

What this is actually about

It would be easy to read all of this as a productivity story. Fewer hours, more output, tidier files.

What I built was a business that holds its own context. Everything that used to live only in my head — who we serve, what good looks like, what matters this quarter, what we don't compromise on — now sits somewhere outside me, written down, readable, and open to improvement.

The consequences of that have very little to do with speed.

I can bring someone in and have them useful in a week rather than a quarter, because the answers exist rather than being trapped in whichever conversation I last had. My standards apply consistently whether I'm rested or stretched, present or away. The business doesn't quietly stop when I take a holiday. And what I've built has value that isn't only me, which for anyone whose business currently runs on their own judgment and availability is not a small thing.

There's a particular kind of tiredness that comes from being the only place the knowledge lives. It feels like being busy. It's closer to being irreplaceable in the wrong way.

Writing it down is how you stop being the constraint in your own business. The AI is simply what makes writing it down finally worth the afternoon.

Your invitation

If you take one thing from this, take the smallest one.

Open a file. Write three sentences: what your business does, who it's for, and what you're working on this week. Save it where your AI can read it. Five minutes, and it will change the quality of the next thing you ask for more than any prompt technique will.

Then add the standard. Then the schedule.

You don't have to build the whole system this week. You have to stop briefing a stranger every morning.

If this landed, I'd like to hear which part you recognised. Most people know exactly which piece is missing the moment it's named. And if you'd like to build this properly, with the architecture designed around your business rather than borrowed from someone else's, that's the work I do.

You're not here to play small. Neither is the thing you're building.

Build this around your business, not someone else's

An honest, confidential conversation about what you're building and whether I'm the right partner for it. No cost, no commitment.