Systems programming with AI coding agents

By Pekka Enberg — October 2026

I was chatting with Marco Bambini, the founder of SQLite AI, and he asked me how I use AI for programming. How much code do I write myself, and how much do I let AI write for me? I told him I write 100% of the code myself. I just don't type it out character by character anymore; AI agents do that for me. As a systems programmer, I need strong guarantees that the code I ship works in production, so I still need to understand it deeply. That's why I think little of the essence has changed. I just have more degrees of freedom with AI coding agents.

Working incrementally

The best way to build reliable software is to work incrementally. I picked up this skill a long time ago from Linux kernel hackers, who make small, incremental changes that keep the code correct and running while moving toward a larger goal. For example, the Linux kernel started out as a single-processor architecture, and multi-core support was added and refined incrementally over a decade. All the while, new releases shipped and ran in production. That taught me that building reliable systems means keeping them running at all times while you improve them step by step.

With AI coding agents, it's important to resist the temptation to let the agent do too much. Frontier models are pretty amazing at code generation, but they don't generate perfect code. More importantly, the more work you let an agent do in one batch, the harder it is for you as a human to keep up with the system. Worse, you lose that understanding gradually, without noticing. One day, you find yourself debugging a bug you have absolutely no idea about.

So building systems with AI coding agents requires the human discipline of working incrementally. You need to resist the urge to delegate your understanding to the agent and instead steer it to work in increments that make sense to you. But how do you do that?

Prototyping

One amazing capability of frontier models is that you can ask them to do pretty wild things, and they'll do them. As a systems programmer, I take full advantage of that: I often tell the agent what I want and let it build the whole thing. Something that would have taken me a week now takes maybe an hour. Instead of spending a lot of time planning up front and doing the mundane work of refactoring, I get to a result very quickly. This kind of high-level experimentation lets me find the best technical direction.

However, I treat the prototype as a prototype and never use the code as is. At minimum, I ask the agent to split the work into small, reviewable commits, just like I've always done. More often than not, though, as I review those commits, both with agents and with my own eyes, I find tons of edge cases. That's my signal to stop and start over with a simpler scope and tighter steering. The point is to really build incrementally, not to pretend by artificially chopping the work into smaller commits.

Verification

Building a verification harness is critical when working with AI coding agents. Whenever I build something, I start with a conformance test suite that gives the agent a binary answer on whether something works as expected. I let the agent write the conformance tests, but I steer the process pretty heavily to make sure we're building the right thing. The best conformance test suites test the system end to end, including the human or agent interface.

I'm a big proponent of deterministic simulation testing, because when you get it right, it's a great unlock for working with agents. Deterministic simulation testing lets you explore your system's state space and reproduce any fault it finds. That makes it a perfect interface for AI coding agents: they can easily debug an issue, fix it, and verify the fix.

To fill the gaps the simulation leaves, I usually implement a non-deterministic testing tool as well. Agents can work with these, but non-determinism sometimes confuses them and leads them to wrong decisions. Tools like Antithesis are great here because they make non-deterministic testing deterministic. The testing cycle is often longer, but agents can help debug those bugs too.

I'm increasingly turning to formal methods for systems programming, too. Whereas other verification strategies explore the state space to find faults, formal methods take a different approach and prove properties of the system. At work, we use Aretta, a low-friction tool that maps the informal specifications we write to formal ones. I've also explored using TLA+ and Lean to verify some components and protocols. Formal methods are still an area I'm experimenting with as I figure out how to fit them into my systems programming workflow.

Performance

AI coding agents are really, really good at performance optimization. As long as you provide a well-designed benchmark, you can basically ask an agent to keep improving performance indefinitely. Of course, in systems programming, understanding is still critical. Agents will sometimes fool you with optimizations that don't make sense or that make the code more complex for no reason. So performance work needs to be incremental too, with every optimization carefully quantified. Human judgment and taste matter here, especially for low-level optimizations. For example, it's often better to work with the language and ecosystem and fix high-level performance problems first, before resorting to low-level bit-fiddling.

Summary

To sum up, I use AI coding agents heavily for systems programming. Still, I don't buy the idea of not reading the code. Most hard systems problems require careful consideration, and while AI coding agents are helpful assistants, they don't always make the right decisions. I also find that they work in batches that are way too large to produce systems software I can ship.

I think the goal with AI coding agents isn't to produce more code, but better code: better structured (because refactoring is free), more robust (because writing tests is cheaper), and more performant (because agents automate the grunt work). But again, at a fundamental level, nothing essential has really changed, at least not for me.