AI writes fast. Someone still has to check it.

What I learned helping my engineering team start using AI: the writing got quicker, the checking did not, and that is where the real work is.

When I ran a three-day training to help the engineers on my team start using AI tools, I told them something before the first session that surprised a few people: the first few weeks will probably feel slower, not faster.

I said it because I'd already felt it myself. AI can write code in seconds. But somebody still has to read that code, make sure it does what it should, and make sure it won't break something else. That checking takes time, and at first it takes more time than anyone expects.

Google's DORA research team, who study how software teams work, have a name for this. They call it the "verification tax". I like the name because it's honest. The AI gives you something for almost nothing, and then you pay for it in checking.

What the research says

For their 2025 report, DORA surveyed almost 5,000 people who work in technology. Nine out of ten said they use AI at work, and most felt it made them more productive. But about three in ten said they had little or no trust in the code it writes.

Both of those are true at once. The same report found that teams using AI ship more, and that their releases also get less stable.

It's a bit like hiring a very fast new team member who never gets tired but also never says "I'm not sure." Their work is often good. You still read it before it goes out the door.

In our first weeks, two things kept happening.

First, the AI did far more than we asked. Someone would ask it to fix one thing, and because we hadn't told it where to stop, it went through the codebase and "fixed" every file and every bit of logic it could find. A small change turned into a big one that nobody had planned to review.

Second, it got stuck. A lot. Engineers would keep one conversation open all day and try to do everything in it, one task after another. The AI's memory filled up with old, unrelated work, and it started losing track of what it was actually meant to be doing.

Neither was really the AI's fault. We hadn't given it a boundary, and we hadn't given it a clean start. Both turned out to be habits we had to learn.

The training mattered less than what we built after

Three days of training shows people the tools work. It doesn't change how they work on a normal Tuesday.

What did change things was writing down how we do things, in a form the AI could read. We built seven small instruction packs, which Anthropic calls "skills". One explains how we write automated browser tests. One explains how we turn a design into a web page. One explains how we write a ticket.

Think of a skill as the onboarding notes you'd give a new hire, except the AI rereads them every time. So an engineer who has never thought hard about how to phrase a request still gets work that follows our team's rules. And when our rules change, we update one file instead of retraining the whole team.

A nice detail: the AI only opens a skill when it needs it. Until then it only knows the skill's name and a one-line description. So you can keep lots of them around without slowing anything down.

Connecting AI to other systems, carefully

There's another way to extend these tools, called MCP. Where a skill is a set of instructions, MCP is more like a plug that lets the AI reach into other systems, such as your ticket tracker or a database.

The simple way I think about it: if the answer changes every time you ask, like "which tickets are open right now?", that's a job for a plug. If it's "how do we do things here?", that's a skill.

Plugs have two costs people don't see at first.

The first is memory. Every plug you connect takes up part of the AI's working memory before you've even asked it anything. Anthropic measured one setup where five connected systems used up a big chunk of that memory before the first question, and the AI got noticeably better at choosing the right tool once they trimmed it down.

The second is security. A plug is a door into your systems, and doors can be abused. Researchers have shown that a harmful plug can hide instructions in its own description, and the AI will follow them. So we treat adding a new one the same way we treat any change to production: someone reviews it first.

Tests: where checking matters most

Tests are small programs that check the main program does what it should. They're tedious to write, which makes them the first thing everyone hands to AI.

And AI writes them well. A 2026 study looked at over two thousand real changes across ten projects and found that AI-written tests covered about as much ground as human-written ones.

But there's a trap. If you ask the AI to write tests by reading the code, and the code has a mistake in it, the AI writes a test that expects the mistake. Everything shows green. The bug now looks like the plan.

It's like asking a student to write the answer key by copying their own homework.

So our test skill tells the AI to work from what the feature is supposed to do, the requirement in the ticket, and not from the code. Then comes the step I care about most: we break the feature on purpose and run the tests again. If a test still passes when the feature is broken, it wasn't checking anything, and we throw it away.

Keeping an eye on it

I've spent years setting up monitoring for the applications I work on: dashboards that show how fast pages load, how often things fail, and when to wake someone up. For a while, the AI tools on my team sat outside all of that. Nobody was watching them the way we watch everything else.

That's easier to fix now. OpenTelemetry, the common standard most monitoring tools use, added a way to record what AI tools are doing, and products like Datadog can now show it next to everything else.

The number I'd watch first isn't speed or cost. It's how often a person had to step in and fix what the AI did. That's the verification tax, measured. If it goes down over a few months, the rollout is working. If it doesn't, no amount of speed makes up for it.

If your team is starting now

  • Tell people the first weeks will feel slower. It's normal, and it passes.
  • Write down how your team works, somewhere the AI can read it. Knowledge that lives in people's heads doesn't scale.
  • Be careful what you connect the AI to, and review new connections like any other change.
  • Don't let the AI mark its own homework. Tests should come from what the feature is meant to do.
  • Watch how often people have to step in. It tells you more than any speed number.

AI made the writing part of my job faster. It didn't make the checking go away. Most of my work now is deciding what's safe to trust, and building the habits that let my team trust it too.

Sources

aiteamstesting

All writing