Skip to content
Tillbaka till bloggen
The Shipping Gap: Why AI Writes Code Nobody Uses
Byggarloggen

The Shipping Gap: Why AI Writes Code Nobody Uses

F
Fredrik BrunnbergVD & Skribent
19 september 20265 min läsning

Right now, somewhere in San Francisco, a team just merged a pull request nobody reviewed, generated by an agent nobody fully understands, into a production system that touches real money. And right now, somewhere in Jönköping, an engineer is still arguing in a Slack thread about whether a function name is clear enough. Guess which company still exists in three years.

That's not a joke. That's the actual state of software right now, and the data backs it up.

The gap nobody wants to talk about

Wharton just published research asking a question that should be embarrassing for the entire AI industry: AI is producing more software, so why isn't it being used? CEPR ran parallel numbers comparing writing output to shipping output across generations of AI coding tools. Same conclusion, different lens: the code volume curve and the deployment curve have split apart. Hard.

Everyone in this industry has been obsessed with one metric for two years: how fast can AI write code. Lines per minute. Tickets closed per sprint. Tokens generated per dollar. Nobody was asking the only question that actually matters to a business: how much of that code is trusted enough to run in production and stay there.

Turns out the answer is: not much. Companies are sitting on mountains of AI-generated code that engineers won't touch, can't fully explain, and don't trust under load. DevOps.com flagged the quality risk side of this just this week. Meanwhile Trend Micro caught threat actors weaponizing Claude's shared-chat feature to distribute malware, which tells you the attack surface around AI-assisted development is expanding faster than any team's ability to govern it.

So here's the real bottleneck, and it was never code generation: it's judgment. Knowing what to ship, when, to whom, with what guardrails. That's a human skill. It doesn't scale with GPU count.

The Swedish accident

I've spent my whole career watching American tech move fast and Swedish tech move carefully, and for a decade I thought we were just slow. Consensus culture. Everyone needs to weigh in. Nobody wants to be the one who broke prod. It felt like a handicap against Silicon Valley's "move fast, ask forgiveness" approach.

Now it looks like the opposite. Our instinct to slow down, review, argue about names in Slack threads, get three people to sign off before something touches customers, that instinct is exactly the missing layer the AI coding boom skipped. We built the review culture before we had the tool that desperately needs one.

Nobody here is writing think pieces about this. Breakit isn't running headlines about the "shipping gap." DI isn't covering it. That silence is itself the story. Sweden isn't loudly solving a problem the US is loudly creating, we're just quietly not creating it in the first place, because our engineering culture never let volume substitute for trust.

That's not virtue. It's a structural accident. But accidents you can build a strategy on are still strategies.

What this looks like in practice at HEIMLANDR

When we do AI agent development for a client, the agent writing code is the easy 20%. The other 80% is deciding what the agent is allowed to touch, what gets human review before merge, what gets rolled back automatically, and what never ships without a person signing their name to it. That's not bureaucracy. That's the actual engineering.

Same with Rapid MVP work. Clients come to us wanting speed, and we give it to them, but speed without judgment is just how you build a liability faster. We ship fast because we know exactly what we're shipping and why. Not because we skipped the part where someone checks.

Where this goes: the next 2-5 years

This gap doesn't close by itself. It gets worse before it gets better, and here's the trajectory I'm watching.

Short term, the volume of AI-generated code keeps climbing, agent frameworks get more capable, and the number of companies with a code review bottleneck grows faster than the number of companies that solve it. Most won't solve it. They'll paper over it with more AI, asking a second model to review the first model's output, which is a house of cards dressed up as governance.

Medium term, this is where regulation matters and where the EU and Sweden are dangerously behind. The EU AI Act covers model risk categories, but it says almost nothing concrete about code provenance, review requirements, or liability when AI-generated software fails in production. If an AI agent writes a smart contract that gets exploited, or writes an update that takes down a hospital system, who is liable? The vendor? The company that deployed it? The model provider? Right now in Swedish and EU law, that's a foggy answer, and foggy answers are exactly what attackers and bad actors exploit. The Trend Micro finding about Claude's shared-chat feature being weaponized isn't an edge case, it's a preview of a much bigger governance failure coming.

Longer term, as we get closer to systems with real agentic autonomy, closer to whatever AGI actually ends up meaning in practice, this judgment gap becomes the single most important competitive axis in software. Not who has the best model. Everyone will have access to roughly the same models within 18 months, that race is commoditizing fast. The winners will be the organizations that built the trust infrastructure: review pipelines, rollback discipline, human accountability chains, audit trails that actually mean something. Sweden's engineering culture, boring as it looks from a pitch deck, is a head start on exactly that infrastructure.

I'd bet money that within three years, "AI-generated but unreviewed" becomes a liability category insurers price into contracts, the same way uninsured software risk got priced after major breaches in the 2010s. The companies that can prove a human judgment layer sits on top of their AI output will have a real, quantifiable advantage. Not a marketing advantage. An underwriting advantage.

What to actually look at

If you're a CTO trying to close this gap instead of widening it, here's where I'd point you this week:

  • ponytail: a repo that's gaining real traction because it does the opposite of what most agent tools do. It makes your AI agent think like the laziest senior dev in the room, meaning it optimizes for the code you never have to write or review at all. That's judgment encoded into tooling. Worth studying even if you don't adopt it directly.
  • opencode: open source coding agent, transparent enough that you can actually audit what it's doing instead of trusting a black box. Auditability is the whole game right now.
  • n8n: not flashy, but if you're automating workflows with AI in the loop, having a visual, inspectable pipeline beats a chain of invisible agent calls every time. You can actually see what fired and why.
  • CEPR's writing-vs-shipping research: read it before your next planning cycle. If your team's velocity metrics only track output, you're measuring the wrong half of the problem.

What to actually do about it

Stop measuring AI coding tools by lines generated. Start measuring them by lines that survived code review, shipped, and stayed in production 90 days without an incident. That number is smaller and it's the only one that matters.

Build the review layer before you scale the generation layer. If you're doing SaaS development or fullstack development with AI assistance, and you don't have a clear answer for who signs off on agent-generated code before it touches customers, you don't have an AI strategy. You have an incident waiting for a date.

And if you're a founder in the Nordics reading US tech press and feeling behind because you're not shipping as fast as the Bay Area teams, stop. Look at what they're actually shipping. A lot of it is slop with good marketing. Our slowness is starting to look like discipline, and discipline is the only thing that scales when the model race stops mattering and the trust race begins.

Fredrik Brunnberg is the CEO of HEIMLANDR.IO, building AI and software solutions from Jönköping, Sweden. This is the daily HEIMLANDR briefing. If you found this valuable, share it with someone who builds things.

#AI coding#software development Sweden#AI agent development#tech company Jönköping#engineering culture#AI governance
F
Fredrik Brunnberg

VD & Skribent

VD för HEIMLANDR.IO. Punk rock-teknik från Jönköping. Bygger AI-system och blockkedjeinfrastruktur och skriver om vart branschen faktiskt är på väg. Ingen ekokammare, ingen hype.

// vad vi bygger

Vill du ha något liknande byggt?

Vi bygger AI-agenter och privat AI på servrar som vi driver inom EU. Från Jönköping.