Book a call

AI-Assisted Engineering, Not Vibe Coding

We use AI on every project. We don't vibe code on any of them. This article is about the difference, and why it matters once your product has real users.

Isometric illustration: a glass AI block sends a row of cubes through a black inspection frame before they join a tidy grey building, one rejected cube aside, on a pistachio background

Why can't our product move at weekend speed?

It started with a good question from one of our clients. They had built a small app over a weekend, just by prompting, and it worked. So they asked what any hands-on founder would ask: if this takes two days, why does our main product move so much slower? It's a fair question. Their product is live, complex and growing, and the team delivers on plan. Just not at weekend speed.

If you haven't been there yet, let me tell you: the first time you watch an app appear from a single prompt, it feels like magic. It turns your world upside down and leaves you in shock. And then the question comes: why keep paying engineers significantly more than the couple of hundred dollars a month that a Claude Code subscription costs, and then wait months to see results?

A fair warning before I answer: I run a company that sells engineering work, so of course I'd say you need engineers. Keep that in mind and check my reasoning against it.

What is vibe coding, anyway?

Andrej Karpathy coined the term in February 2025. He described accepting every change the model suggests, not reading the diffs, and forgetting that the code exists.

Simon Willison later gave the definition I find most useful: you're vibe coding when you build software with an LLM and don't review what it writes. His article is worth reading in full.

I like this definition because it draws the line in the right place: between code a person has read and code nobody has read. Who typed it doesn't matter.

Could it work?

Oh yes, and it does. The question is where:

  • A prototype you plan to throw away
  • A weekend experiment
  • A demo for a pitch
  • A website built through MCP servers for WordPress, Framer or Webflow. This works pretty neatly.
  • A whole app, if it's a simple tool that doesn't store user data
  • With a certain degree of care, even an MVP or a proof of concept

A founder can now sit down in the evening and have a working app by midnight. That's real, and it's fun. I wouldn't talk anyone out of it.

But it also sets the wrong expectation. The evening produced a demo. A product that carries customers, payments and other people's data is a different kind of work, and code nobody has read shouldn't go into it.

The tricky part is the moment a demo quietly becomes a product. Here's my rule: once it has real users, takes payments or stores someone else's data, it's a product. From that point on, somebody has to read the code.

What happens when nobody reads the code?

Models get better every few months, so I only looked at recent data. The short version: more code gets written, and somebody still has to check it.

Two caveats first. Nobody has studied a mature product that switched to vibe coding, so none of this is direct proof. And two of the sources below, Veracode and GitClear, sell tools for the problems they describe. I still use their numbers, because they measure real code at scale.

Veracode's 2026 report, published in July, tested current models on security tasks. They pass 56% of them, against 55% in the previous edition. The best model still fails nearly one task in three. So models have pretty much learned to write code that runs. Writing code that's safe is another story.

GitClear's 2026 research covers 623 million code changes from 2023 to 2026. Refactoring is down to 3.8% of changed lines, from 21% in 2022. Duplicated blocks are up 81% since 2023. So more code gets written, and less of it gets cleaned up.

Google's DORA research reached a similar conclusion in late 2025: AI improves throughput, often at the cost of stability when the engineering foundation is weak.

The most honest note comes from Willison himself. In May 2026 he wrote that as agents became more reliable, he stopped reviewing every line they produce, and that this troubles him. I think he's right to be troubled. An agent can't be held accountable for what it ships, so somebody still has to be.

Screenshot of a coding agent asking for an INTERNAL_DISTINCT_IDS variable in Vercel, Production and Preview, so the team's two iPhones count as internal; the device IDs are blurred
One of the newest coding agents suggests hiding the team's test devices by putting their IDs into an environment variable. How many workarounds like this vibe coding leaves behind is a separate question.

Why is a mature product the worst candidate?

A mature system is its code plus the record of decisions behind it. Why is this edge case handled this way? Which customer depends on which behaviour? What was tried and rolled back? On a well-run project that record is written down, in decision records, runbooks and tests, so that it doesn't live in anyone's head and create a dependency. (I'm building my whole business around the idea of not creating dependencies.)

Vibe coding breaks that record in two ways.

First, nobody checks the change against it. A model works on the task in front of it, and in a large system it can miss a constraint documented three services away. It writes a new function where one already exists. It fixes a symptom in one place and leaves three copies untouched. A reviewer would catch this. With no reviewer, it goes straight to production.

Hand-drawn example: an agent fixes a price rounding bug in the Orders service with a new formatPrice() that rounds with float math; a formatPrice() already exists in a shared library, the same rounding code stays untouched in Cart, Invoices and Reports, and an ADR in Payments says amounts are integer cents, never floats

Second, the record stops growing. Every unread change carries decisions that no one made on purpose and no one wrote down. After a few months the documentation describes a system that no longer exists.

That's a dependency too, and a worse one than depending on a vendor. You depend on a tool to keep changing a product nobody can explain. A new team can't take it over, because there's nothing accurate to hand them.

If you work under HIPAA, SOC 2 or a similar regime, there's one more problem. An auditor will ask who reviewed a change. Good luck answering "nobody".

What do we do instead?

Our engineers use AI assistants and coding agents every day. We keep the rules short.

The engineer who commits a change is responsible for every line, whoever wrote it. Code the author can't explain doesn't get merged. AI-assisted code goes through the same review, tests and security scans as everything else.

Agents work on a branch in a development environment. They have no production credentials and they don't merge. A person approves anything that reaches beyond the working copy.

The tools run on company accounts. Training models on our clients' code is not allowed and is switched off. And each client has the right not to use AI on their project at all.

That doesn't mean every change gets the same amount of attention. There's a lot we hand to agents much more freely: tests, internal tools, isolated modules, prototypes of new features. The closer a change gets to money, user data or access control, the more human eyes it gets.

So if you want more speed, don't ask your team whether they can stop reading the code. Ask them where they can afford a lighter review. That one actually has good answers.

Are we getting faster with AI-assisted engineering?

I'm leaving vibe coding out of this comparison, because it's a no-go for a serious commercial product. That leaves two ways of working: classic hand-written code and AI-assisted engineering.

My personal judgment: yes, people work more efficiently with AI. I wouldn't say they're dramatically faster, though. Engineers who use AI can take on more. On some tasks they're up to 80% faster. On others the gain almost disappears, because they spend that time reviewing what the AI wrote and fixing it.

The numbers back this up. A Stanford study of more than 100,000 engineers at 600+ companies measured where the gain actually shows up. On simple tasks in a brand-new project, AI made people 30–35% faster. On complex tasks in an existing, mature codebase, the gain was 5–10%.

Why not more? Because typing code was never the whole job. Understanding the task, reading what's already there, reviewing, testing and deciding what to build take as long as they always did. AI speeds up the writing part. The thinking part is still on us.

Overall, my estimate (an estimate, not a measurement) is that our engineers work 15–20% faster, which is right in the range the Stanford study found for existing codebases.

Hand-drawn bar chart of AI productivity gains from the Stanford study: 30–35% on simple tasks in a new project, 15–20% on simple tasks in an existing codebase, 5–10% on complex tasks in an existing codebase; a yellow band marks MetaProject's estimate of 15–20%

And yes, clients see it in the bill. That 15–20% shows up in delivery: we either move faster or do the same work with a smaller team.

So it's not the magic pill it seems to be when you vibe-code a mini app over a weekend. But the gain is there. Efficiency grows and models keep evolving. I expect we'll get faster still, and that's not decades away. We're just not there yet.

So where's the line for your product?

Ask yourself five questions:

  1. If this breaks in production, who can explain what the code does?
  2. Does the system handle money, personal data or access control?
  3. Who reviews changes, and how large are the changes they're asked to review?
  4. What can the agent reach: production, secrets, the deploy pipeline?
  5. Could a different team take over this codebase a year from now?

If the answers are comfortable, you're probably not vibe coding. You're doing engineering with better tools.

If they're not, slow down. For the sake of stability and control, keep unread code out of your product.

If you're weighing this decision for your own product, we're happy to compare notes. You can find us at metaproject.pro.

Read next

planning a new project
We’ll help you choose the right discovery depth and map a realistic starting plan.
Isometric blue glass blueprint plate with glass blocks rising from it and one yellow block

Recent blog posts

Read More Articles
Isometric illustration: a glass AI block sends a row of cubes through a black inspection frame before they join a tidy grey building, one rejected cube aside, on a pistachio background
AI-Assisted Engineering, Not Vibe Coding
View insights

We use AI on every project. We don't vibe code on any of them. This article is about the difference, and why it matters once your product has real users.