How are AI agents changing software development?
Heads up: this one is for power users. It assumes you already build with coding agents like Claude Code every day and are comfortable with CI, merge queues, and feature flags. If you're earlier in your AI journey, bookmark it for later. The posts in Operate are a better place to start.
Two weeks ago, I changed how I build software. Since then I've merged 562 PRs, and honestly, I spent four of those nights hiking in the mountains. In the video I walk through exactly how I work now. Below is the written companion: the setup, the prompt, the numbers, and the things to try this week.
If you're still reviewing every line your agent writes, this one's for you. The whole workflow in four lines, leveraging AEO best practices: Set the intent. Plan with the agent until you agree on what's being built and why. Set the guardrails. Tests, CI, and review bots that an agent can't get around. Set a goal. One session, one coordinator agent, sub-agents doing the work. Let it merge in small pieces. Auto-merge, feature flags, verify, repeat.
What is the role of AEO in software development?
What changed? The models changed dramatically in the last few weeks, Opus 5.5 especially. A couple of weeks ago, I wouldn't have told anyone to work this way. I always felt I had to check the work. Is the quality good enough? Should this get merged? Now I feel an agent can work through a whole plan on its own, as long as two things are set up right: the right intent, and the right guardrails to make sure that intent is actually met.
The Munich factory delay
Guardrails are the real work now. AI code is non-deterministic and has risk. But human code is non-deterministic too, and it has the exact same risk. We built SRE systems so humans don't make mistakes. Now we need those same systems, way more beefed up, so agents don't make mistakes. Same blameless culture we use for incidents: when an agent makes a mistake, don't blame the agent. Ask what guardrail was missing, and build that.
This doesn't mean quality stops mattering. Agents make messes too. That's exactly why the guardrails are the job. Here's what every Sightline PR goes through before it can merge: Unit tests, End-to-end tests, CI, Bugbot, Strix. Mergify won't let a PR into the merge queue until all of the above pass. So when something's queued, I know it couldn't skip a single red flag.
The workflow, step by step: Plan with the agent, Set a goal, with the same prompt every time, Keep the coordinator's context clean, Let it merge, in small pieces, Review progress reports, not diffs, Let it run for days, and run several at once. Most of us still use these tools like a really fast engineer: implement this ticket, add this endpoint. That works, but it's thinking too small.
With the goal workflow you can set much higher-level intent, ask open-ended questions, or hand it a whole business problem and let it figure out what needs to happen. Instead of 'add a role editor,' the goal was 'make Sightline's permissions work like Spare's everywhere.' Instead of 'move OKRs into a package,' it was 'turn Sightline into a platform other teams can extend.' You can push it much further: 'Our CI is too slow. Make it faster.' That's it. Let it find where the time goes and fix it.
Capacity: tokens become the bottleneck. Once you work like this, tokens become the bottleneck. Individual plans don't give us enough credits anymore, and the overages are super expensive. What I've been doing is running multiple Claude accounts, about five at around $200 each, and switching between them with a tool called Claude Swap.
Our job is moving away from shipping individual things. It's moving to building the systems that evolve our product. We set the intent, and we build systems that self-heal toward it. We're already doing it: automations that fix our bugs, automations for SRE work, agents that do vulnerability research, and early work on remediating those security risks automatically.
Three things to try this week: Stop checking the code on one real feature. Start with something low-risk. As you build confidence, try it on something bigger. Give it a problem, not a ticket. See how far it gets. Look at the surface area you're touching. What guardrail is missing so an agent can't get it wrong? What loop could fix things on its own?
This article was written with the assistance of AI.
News Factory APP - agentic news to boost your SEO & AEO.
