Why AI-generated code is hard to maintain, and what fixes it
I’ve been skeptical of AI-built work because it tends to be fragile. Something gets built quickly, then rebuilt when it breaks, and the rework eats the time you saved.
I hear the same thing from friends on in-house teams who got stuck maintaining apps someone built with AI and then left. The person who prompted it is gone. Nobody knows why it works the way it does, and the first real change request turns into a rewrite.
Where the fragility comes from
A model with no context builds the way the average of its training data builds. For a one-off script that’s fine. For a data pipeline or a finance app that other people have to run and change, it’s a problem, because the average answer skips the parts that make something maintainable.
In data work that usually looks like this:
- Writes that aren’t transactional, so a failure halfway through leaves a table partly updated.
- Checks that log a warning and keep going when they should stop the run.
- Duplicate records handled with a quiet dedupe, so the pipeline picks a winner nobody chose.
- Naming and structure that change from one build to the next.
None of these show up in a demo. The output looks right on day one. They show up the first time the source data is wrong, or when someone else has to change the code.
The reviewer is the other half of it. If you don’t already know how a publish step should behave when a check fails, the model’s version looks fine. That’s how fragile work gets shipped by people who are trying to do it right.
What fixes it
Better prompting helps some. The bigger change is giving the model your standards before it starts, every time, without relying on the person at the keyboard to remember them.
I’ve spent most of this year at OVG building Meridian. This is the problem it solves for us. Our design standards are part of what the model knows before it starts, so the first version comes out the way we would have built it ourselves. When someone asks for a step that publishes a table, the model already knows how we want a failed check handled and how the run gets logged.
Two things change. We redo less, because the first version already follows the pattern a reviewer would have pushed it toward. And new folks are working to the same standard on day one, instead of picking it up over a year of projects.
What counts as a standard
The standards that matter most are the boring ones. How a table gets published. What happens when a validation check fails. How a run gets logged so someone can tell later what happened. Where credentials live. A senior engineer makes these calls without thinking about it, and a model skips them when nobody tells it.
Writing them down is most of the work. A standard has to be specific enough that a model can follow it and a reviewer can check it. “Handle errors properly” does nothing. “If the source has two conflicting records for the same key, fail the run and name the key” is something a model can follow and a reviewer can test.
They also need upkeep. When a pattern turns out to be wrong, it has to come out, or the model repeats the mistake on every build after that.
If you’re building with AI in-house
You don’t need a platform to start on this. A few things that help on their own:
- Write down the patterns your team keeps correcting in code review, and put them in front of the model on every build.
- Make failures loud. A pipeline that stops and names the problem is easier to maintain than one that finishes green with bad data in it.
- Have someone who knows the pattern review the first few builds against it, line by line.
- Assume someone else will maintain it.
What I can’t tell you yet
I don’t have a number for how much rework this saves, and I’m not going to guess at one here.
If you’re a finance or data leader working out where context fits alongside your CPM data, reach out. Happy to compare notes.