← Journal
Engineering practice

AI Coding Doesn't Replace the Engineer: Who Owns Architecture, Tests, and Production

AI writes code faster than I do — and that changes nothing about who owns the outcome. An honest breakdown: where assistants carry research, routine, and documentation, and where the territory begins on which there is nothing left to delegate.

Yuri Eliseev
19
AI Coding Doesn't Replace the Engineer: Who Owns Architecture, Tests, and Production
Article contents10
Three in the morning. Phone rings. It's the client.

Payments are failing. Forty minutes now. I open the laptop. Logs are clean. Tests are green — I ran them myself that evening. Forty-seven tests, 92% coverage. The billing module was written by an agent, I reviewed it, everything lined up.

And payments are still failing.

It took twenty minutes to find the cause. The payment provider's webhook arrived twice. The first time, we processed it. The second time, we tried to process it again — and there was no idempotency in the code. The client's money was charged, the subscription never activated.

Not one of those forty-seven tests checked for this. Because neither the agent nor I had thought that the same webhook could arrive twice.

That night is when I started seriously working out where AI coding genuinely helps, and where it creates the illusion that the work is done when it isn't.

Where the assistant truly delivers

Let me start with the good part. I won't lie — it speeds things up.

Research. It used to take me half a day to get comfortable with an unfamiliar library: read the docs, then the GitHub issues, then write throwaway scripts. Now an agent pulls what I need in ten minutes. Sometimes it gets things wrong, sometimes it mixes up versions, but the starting line has moved dramatically closer.

Routine. Migrations. Refactoring repetitive code. API plumbing. Configs. What used to eat a day now takes an hour. I'm not exaggerating one bit.

Documentation. This is where I stopped suffering entirely. It used to get written at the end, on whatever energy was left, and it showed. Now it gets written as we go. And more importantly, it gets written in a boring, detailed way — not a pretty, useless one.

Tests. Not all of them, but coverage on the obvious scenarios, absolutely. Edge cases I would have forgotten, the agent often remembers on its own.

That's a decent list. For those four things, I'm willing to forgive AI coding a lot.

But past that point, we get into what nobody talks about in a demo.

The trap of speed

Before I get to the limitations, let me flag a danger I see in everyone who's recently picked up agentic coding.

Speed is intoxicating.

One evening and you've built what used to take a week. Two days and you've built what used to take a month. And at some point you start to feel like anything is possible now. That there are no limits. That the old world — architecture, tests, code review — is a relic for people who don't know how to use the tools.

That feeling is dangerous. It comes to everyone. And it's the most expensive thing to pay for later.

Because speed is not the same as quality. Writing code fast and building a product fast are two different things. The first, AI can do. The second, it can't.

Where acceleration ends

Listen carefully. This is the main point.

AI doesn't own the decisions. It proposes. Sometimes brilliantly, sometimes way off. The decision is yours. And if it turns out to be wrong, you answer for it — not the model.

An example from practice. The agent proposed an architecture: a monolith, background jobs in the same process. Reasonable? Sure. Fast? Absolutely. We'd have shipped an MVP in a week.

I said no.

Because I knew: in three months there'd be a second client with data isolation, then a third, then we'd want to scale background jobs separately. And the monolith would have to be carved up — which isn't refactoring anymore, it's a rewrite. With data migration, with downtime, with explanations to clients.

We lost a week up front. We saved two months later.

This isn't a story about "AI bad, human good." It's the difference between optimizing for today and architecting for years. The agent doesn't see the horizon. It sees the task you put in front of it. If you framed it wrong, that's not on the agent. That's on the person who framed it.

Tests that are always green

A separate conversation about tests. There's a treacherous quality to AI coding here.

It writes tests that pass. Sounds like a compliment. It's actually a diagnosis.

A test that's always green often verifies the wrong thing. It verifies that the code does what the code does. Not what it's supposed to do by design.

A story from that night. Forty-seven tests. 92% coverage. All green. Beautiful.

Then the client couldn't pay for their subscription. Because the tests walked the happy path, and not one of them checked what happens if the webhook arrives twice. The agent didn't know that could happen. I didn't think about it either — and that part is on me, not on it.

Since then, the rule is simple. AI writes tests — I write the scenarios for tests. Every edge-case test isn't about coverage. It's about the question: what happens when things go wrong? The agent won't ask that question on your behalf. It doesn't have the experience that question comes from.

Experience is the one thing you can't delegate.

Production: where delegation ends

There's nothing to argue about here.

Production is the territory where AI coding goes quiet. Not because it can't. Because the consequences belong to a human.

When a service goes down at three in the morning, they call you, not the model. When a client loses data, they sue you, not the agent. When the infrastructure bill suddenly jumps tenfold, you're the one explaining it, not the prompt.

One thing matters here. AI doesn't remove responsibility. It concentrates it.

You used to be able to hide behind the team. Developer missed it. Architect chose poorly. QA let it slip. Now, when one person with agents does what a team used to do, there's nowhere to hide.

Sometimes it works out that way: you become the architect, the developer, the tester, and the on-call engineer, all in one night. That's not a bad thing. It's just the new reality. And with it comes the full weight of responsibility.

So who actually owns it

Back to the question I was asked a year ago: why am I needed if AI writes code faster?

Simple. AI writes code. The engineer owns whether that code solves the right problem, whether it's been tested against what could break, and whether it holds up when things don't go to plan.

These aren't grand words about human value. It's a division of labor. The agent is the executor. The engineer is the one who sets the task, chooses between options, verifies, answers, and, if needed, redoes it.

The difference isn't visible in a demo. It's visible six months in — when the system is alive, growing, and not falling apart.

The resolution

AI coding isn't going anywhere. It's going to get better, faster, smarter. And that's a good thing.

But the more powerful the tool, the more expensive the mistake of the person wielding it.

Engineers used to be expensive because they could write code. Now they're expensive because they can own the outcome. Subtle difference, but it's the one that separates someone who "uses AI" from someone who builds products.

If you're looking for someone to press a button and get code, AI will manage without you. If you're looking for someone who owns architecture, tests, and production — it would be my pleasure. That's the job.

Glossary of terms

  • AI coding — the practice of development with heavy use of AI code-generation tools.
  • Vibe coding — an approach where the human sets direction and the model writes much of the code.
  • AI agent — a model-based system capable of performing tasks autonomously: writing code, calling tools, making decisions within defined boundaries.
  • Engineer — a specialist responsible for architecture, quality, tests, and operation of a system, regardless of who wrote the code.
  • Architecture — the structure of a system: components, their connections, distribution of responsibility, and boundaries.
  • Production (prod) — the working environment where the system is used by real users.
  • Testing — the process of verifying a system against requirements and its resilience to errors.
  • Test coverage — the share of code checked by automated tests. Doesn't guarantee quality by itself.
  • Unit test — a test of a single module or function in isolation.
  • Integration test — a test of interaction between several system components.
  • Regression test — a test verifying that new changes haven't broken previously working behavior.
  • Idempotency — the property of an operation producing the same result when repeated.
  • Webhook — a notification an external system sends to your service when an event occurs.
  • Edge case — a rare but possible situation easily missed in design and testing.
  • Code review — a check of code by another specialist before it enters the main branch.
  • Debugging — the process of finding and fixing the causes of errors in a system.
  • Incident — an event that caused a failure or degradation of the system in production.
  • Production responsibility — an engineer's duty to be accountable for the system's behavior with real users.
  • Automation — replacing manual actions with programmatic processes: build, tests, deploy, monitoring.
  • Technical debt — accumulated compromises slowing development and requiring rework.
  • LLM (Large Language Model) — a large language model, the foundation of modern generative AI systems.

Respectfully,

Yuri Eliseev

AI Systems Architect · Full-Stack Product Engineer

Collaboration

Need a project of any complexity?

Let’s discuss an idea, product, AI system or technical challenge and define a realistic first step.

Start a conversation