AI-Assisted Building

Shouldn’t a website be cheaper now that AI builds it?

Partly, yes. Building takes less time than it used to, and that should show up in what you pay. But deciding what to build, holding scope and catching the confident mistake didn’t get cheaper, and that’s where your money goes now.

Shouldn’t a website be cheaper now that AI builds it?

In short: partly, yes. Production work has genuinely compressed, and that should show up in what you pay. But the newest AI coding model arrived last week with a manual telling developers to stop asking it to check its own work, which says something about where the cost actually sits now. Deciding what to build, holding the scope, and catching the confident mistake didn’t get cheaper. Those were always the expensive part.

Anthropic released Claude Opus 5 on July 24. On the numbers it’s the strongest coding model available, close to the frontier at half the price of the model above it. That part isn’t in dispute.

What’s more telling is what shipped alongside it. In its own migration notes, Anthropic tells developers to delete instructions like “include a final verification step” from prompts written for older models, because this one already verifies its own work and doubles up. The same notes say it delegates to other agents more readily and writes longer than its predecessor.

Read that again. The company that built the best coding model available is documenting that it does more than you asked.

What that looks like in practice

CodeRabbit, a company that reviews code for a living, ran it against roughly a hundred real error patterns from open source projects. It was more precise than their existing setup, and it caught fewer of the known issues. It produced four times as many trivial comments, 92 against 23, and used 50 to 65% more tokens per review. Their verdict was that it works as a specialist alongside another model, not as a sole reviewer.

The developer forums ran on the same theme all week. Reddit threads with hundreds of comments describing a model that codes superbly and takes real effort to keep on task: escalating small problems as emergencies, relitigating settled decisions, building things nobody asked for. The workaround people converged on isn’t a better prompt. It’s a second model running in front of the first, planning the work and holding the boundaries while the strong one implements inside them.

Read that as a staffing decision, because that’s what it is. Faced with a worker who’s exceptional and hard to direct, they hired it a manager. Which is roughly what I’ve argued this technology has been all along: a great junior and a dangerous senior.

A brief feeding a model that holds scope, which in turn directs a larger model that implements
The fix wasn’t a better prompt. It was an org chart.

My own few days with it line up with that, from a different angle. It isn’t as capable as Fable 5, the more expensive model above it, and it’s noticeably reliable. On the work my team and I actually ship, site builds, CMS structures, content passes, it does what I ask and holds its footing. That’s a few days of one person’s work rather than a benchmark. But reliable is worth more than brilliant on paid production, and that trade is the whole point of what follows.

Where the cost moved

For most of my fourteen years, production was the bottleneck. Building the thing took the time, so that’s what everyone priced and scheduled. Estimates were really estimates of typing.

That bottleneck is largely gone. The machine can write the code. What it can’t do is decide what should exist, hold a scope when the goal gets fuzzy, or notice that its own output is confidently wrong. Those three were always the hard part, hidden behind the production work the way a foundation is hidden behind the time it takes to build the walls.

And when building is cheap, the wrong decision gets built faster. Speed doesn’t forgive a bad brief, it just delivers it sooner.

So what should you actually pay less for?

Build time. Things that took a week take a day, and my team and I pass that on in scope and speed rather than pretending it hasn’t happened.

What hasn’t moved: working out what the site needs to do, structuring content so people find what they came for, and arguing you out of the feature that will quietly cost you conversions. Review has gotten harder, not easier, because the volume of plausible-looking output went up and the confidence attached to it went up with it. Reviewing this work properly takes real time, and it’s the first step that disappears when someone quotes you a number that seems too good.

The cheap quote isn’t cheap because they found a faster way to build. It’s cheap because someone stopped reviewing.

So when you’re comparing proposals, ask three things. How they use AI, because everyone credible does now and evasiveness is the signal rather than the usage. Who reviews what it produces, and how, because “we check it” is not a process. And who’s accountable in month eight, which is the question that hasn’t changed in fourteen years and still decides everything.

Compare the thinking, not the build estimate. The build is no longer the expensive part, which means the brief, the scope and the review are where your money either works or evaporates. That gap between producing something and knowing whether it’s any good is the whole job now.

If you want a straight read on a project you’re scoping, or on a quote that came in surprisingly low, that’s a conversation my team and I have most weeks. I’m glad to have it with you.

Quick questions, quick answers

Does AI make websites cheaper to build?

Partly. Building genuinely takes less time than it used to, and that should show up in your quote. Discovery, information architecture, content decisions and review haven’t moved, and on some projects review takes longer than it used to. Be suspicious of a price that assumes everything got faster, and read what a website actually costs before you compare numbers.

Should I be worried if my web partner uses AI?

No. Worry if they won’t tell you how, or can’t name who reviews the output. Used with supervision it raises quality and speed. Used without it, it produces work that looks finished and isn’t.

What still needs a person?

Deciding what should be built, holding scope when requirements get vague, catching the confident mistake, and being accountable after launch. The strongest coding model available is an excellent worker and still a poor pilot.

Related reading

Keep reading

Have a project in mind?

Let's build something that lasts

Get in touch

Based in Manila, working with teams across time zones.