Don't Let AI Guess Where Your Monolith Breaks

sleroy · Jul 31, 2026 · 11 min read

I have seen a lot of “AI modernizes your monolith” demos, and they all trip on the same rock. You ask the model where a capability starts and stops. It answers with total confidence. And it is wrong in the one spot you didn’t check: a class loaded by reflection, a SQL query reaching into another module, one more caller of a method you were sure was private. The cut looks clean, the build goes green, and three weeks later something falls over in production.

So I stopped asking the AI that question.

Here is the thing most people miss. A boundary in a codebase is a fact, not an opinion. A tool that reads the compiled call graph will find it and give you the exact same answer every single time. Generating the new service, writing the docs, grinding through the boring refactors — that is where AI is genuinely good. So let a compiler-grade tool draw the line, and let AI build on top of it. In that order. Never the other way around.

For the record, the tools I reached for were CodeQL for the boundary, Kiro for the code generation, and AWS Transform Custom for the long documentation and refactoring passes. I will name them as they come up.

This is Part 1 of my series Monolith Modernization. I extracted one capability — let’s call it Events — out of a Java 8 monolith into a Java 21 microservice. You pay the setup cost once for the whole codebase, then you run the extraction loop once per capability, and here is the punchline: every capability after the first one is cheap.

This is Part 1 of the Monolith Modernization series.

  1. Don’t Let AI Guess Where Your Monolith Breaks (this post)
  2. Excavating the Monolith: What Is a Feature?

Know what you are fighting first

The monolith fights back for four reasons, and if you have touched one of these beasts you already know them:

  • One shared thread pool, so a slow query in one corner starves everything else.
  • One shared database with no table ownership, so there is no clean line to cut along.
  • A frozen public API that other systems depend on, so you can’t reshape it just to make your life easier.
  • A runtime so old the libraries have all moved on without it.

That is the textbook case for the strangler fig. Keep the public seam, replace what sits behind it one piece at a time, and stay able to back out at any step. Nothing new here — Martin Fowler named this pattern years ago. What is new is who draws the boundaries.


Setup: once per codebase

Coupling graph

Step 1 — get a baseline. Run the whole thing against seeded data and write down exactly what it spits out, right down to the final count of exceptions. Skip this and you will never be able to prove an extraction kept the behavior instead of quietly changing it. This captured output is what every later step has to reproduce, character for character. “The build is green” and “the behavior is identical” are two very different claims, and only one of them matters.

Step 2 — find the boundary with static analysis. This is the step everything else stands on. I used CodeQL for this. A static-analysis engine reads the compiled code and tells you what calls what, which package leans on which, and where the SQL lives. Why not just ask the model? Because the model guesses badly in exactly the places that bite: reflection, dynamic dispatch, config-driven wiring, cross-module SQL. And you won’t find out until the cut breaks. The tool gives the same answer on the same input, so you can hand a reviewer both the answer and the query behind it. Writing those queries costs you an afternoon, once. Every future extraction reuses them.

Here is the whole idea in one small query. This one lists every call that leaves the capability you want to extract — exactly the set an LLM tends to under-count:

import java

from MethodCall call, Callable caller, Method callee
where
  caller = call.getEnclosingCallable() and
  callee = call.getMethod() and
  caller.getDeclaringType().getPackage().getName() = "com.example.app.events" and
  callee.getDeclaringType().getPackage().getName() != "com.example.app.events"
select caller.getDeclaringType().getName(), caller.getName(),
       callee.getDeclaringType().getPackage().getName(),
       callee.getDeclaringType().getName(), callee.getName()

Swap the package name and you get the outbound edges for any capability. A second query counts those calls per target package, so you see at a glance what the capability leans on:

from string toPackage, int calls
where
  calls = strictcount(MethodCall call |
    call.getEnclosingCallable().getDeclaringType().getPackage().getName() = "com.example.app.events" and
    call.getMethod().getDeclaringType().getPackage().getName() = toPackage) and
  toPackage != "com.example.app.events"
select toPackage, calls order by calls desc

A third one, not shown, greps the capability’s code for string literals that touch the shared database schema — how you catch the SQL that couples two capabilities through the data layer. None of this is guesswork. It runs against the compiled call graph and returns the same rows every time.

Step 3 — let a long AI pass write the docs. Now switch tools on purpose. This is where AWS Transform Custom earns its place: a cloud-side agent chews through the whole codebase and produces the architecture notes, the behavior summaries, the business rules. Great job for a patient agent. Just be clear on what it is: context, not truth. The boundary came from Step 2. The prose sits on top of it. Don’t let a nicely written paragraph overrule a fact you verified.


Extraction: once per capability

Step 1 — pick the target. That is your job, not the AI’s. A human decides what to pull out next. A Kiro agent then runs the CodeQL boundary query and reports back in plain English: Events owns two classes, writes one table, has one caller, makes two calls into another capability. Low coupling, good first pick. Notice the order — facts first, then the story. That is the reverse of how AI normally works, and honestly that reversal is the whole trick.

Step 2 — write the slice. Let the same Kiro agent turn the CodeQL output into a short written spec, always the same shape: what moves, what data it owns and who else reads it, the seam that isolates it, its dependencies, and the shared utilities that can’t come along for the ride. Same shape every time means anyone can review it, and it means capability number two is basically fill-in-the-blanks. Every claim in that spec cites a row from the analysis output, so nothing is hand-waved. A trimmed slice for the Events capability looks like this:

# Events Capability — Extraction Slice
_Source of truth: static-analysis output. Every claim cites a row._

## Classes to move
- EventProcessingWorker (owns the worker loop)
- EventService (owns the business logic + one table)

## Data owned
- Table `meter_event` — written only by EventService.
  Still read by the outage integration → needs a read contract after the cut.

## Cross-boundary calls to resolve (4 of 34 edges)
- EventService → common.DatabaseConnectionFactory → replace with a service-local datasource
- EventService → exceptions.ExceptionManagementService.raise → replace with an async event

## Verdict
Clean, low-coupling seam. Only real decision: the in-process call to
`raise(.)` becomes an event on the bus.

The number that matters is on the coupling line: 34 outbound call edges, and only 4 of them cross a capability boundary. That is the whole extraction — four edges to think about instead of the entire package. You do not get that count from reading the code. You get it from CodeQL.

Step 3 — set up your workspace once. This is where casual AI users and serious ones part ways. I did this in Kiro, whose spec mode and per-save hooks are built exactly for this. The casual crowd retypes their guardrails into every single prompt. Don’t be that person. Bake the rules into the workspace up front: agents each scoped to one job, hooks that compile and test on every save, a library of reusable prompts. Pin your target versions here too — and that brings me to the one war story worth telling.

Left on its own, the code agent tried to write against an older framework version than the spec asked for, because its training data still thought the old version was current. It would have compiled. It would have passed the tests. And it would have blown up months later on a feature that only exists in the new version. The fix wasn’t to patch the code. It was to patch the prompt: check the published version before you swap one in, and print the version you used as proof. Fix the prompt, not the output, and the fix travels to the next capability instead of dying with this one.

Step 4 — scaffold an empty service. Point Kiro at the target spec and generate a skeleton that builds green and does nothing useful: project file, main class, a health check, real config with no fallback, one placeholder test. The agent prints the framework version it picked, so your guardrail proves itself before a single line of real code exists.

Step 5 — build the safety net. Before you touch any logic, generate tests that pin what the system does today. This is a good job for AWS Transform Custom, which reads the existing behavior and writes the characterization tests for you: this event type maps to that severity, that one produces nothing. They pass against the untouched monolith, and then you freeze them forever. That is your ratchet. Anything that breaks them later is wrong by definition, and it is exactly what lets you delete code later without sweating.

Step 6 — move the logic and rewire the monolith. Build the real capability in the new service with Kiro: its data access, a pure classifier you can unit test, the event publisher, an orchestrator that dedupes on a stable id. Then change the monolith without touching a single public signature or test. The call-site surgery on the old code is where OpenRewrite pays off, because its type-aware edits rewrite every caller correctly instead of by pattern-matching. Add the client that calls the service, add the listener that consumes its events, and gut the old class so its API is identical but its body just delegates. No fallback anywhere. If the service is down, the call throws, that capability fails loudly, and its siblings keep running. The tests are supposed to fail when the service is unreachable. That is the rule, not a bug. I mean it — resist the urge to add a graceful fallback here.

Step 7 — prove it. Have a Kiro agent reset the data, start the service, run the monolith, and check the output. The exception count at the end matches the baseline exactly, even though the work now takes a completely different road: monolith to service over REST, service to an event bus, bus to a queue, queue back into the monolith’s original exception path. Same result, different route entirely. When that last line matches the baseline word for word, the extraction is invisible to everyone who calls it. That is the moment worth screenshotting.

Step 8 — modernize what’s left. With the capability out, point AWS Transform Custom at the mechanical cleanups: old logging to a modern facade, hand-rolled data access to a framework, an old test library to the current one. Codify once, replay everywhere. For the purely mechanical bits a deterministic tool like OpenRewrite beats the agent hands down, so save the agent for the parts that actually need a judgment call on naming and placement.

And here is where it pays off. The prompts and the CodeQL boundary query never mention Events by name. So the next capability is the same workflow pointed at a new target. No rewrites. You get a new service with its own event flow, the old class shrunk to a delegator, and every test still green. The first extraction is the investment. The second takes hours instead of weeks. So does the third. You step in twice — to pick the target and to judge the result. Everything in between just runs.


What I want you to remember

  • Let a compiler-grade tool draw the boundary. AI writes lovely docs, but it will not guarantee a sound call graph, and the boundary is the one thing you cannot afford to get wrong. Facts first.
  • Treat your prompts like code. When the output is wrong, fix the prompt so the fix survives to the next job.
  • Refuse graceful degradation in the middle of a migration. A loud failure is a gift. A silent fallback is a bug you will meet again later, at a much worse time.
  • Judge the payoff across the whole migration, not the first cut. That first extraction looks more expensive than just letting the AI wing it. By the second one, the method is the plan. That is when this stops being a demo and becomes how you work.

Further reading

  • Strangler Fig Application — Martin Fowler’s original write-up of the pattern.
  • CodeQL — the static-analysis engine used to draw the boundary.
  • Kiro — spec-driven AI code generation with per-save hooks.
  • AWS Transform — cloud-side agentic passes for documentation and large-scale modernization.
  • OpenRewrite — deterministic, type-aware refactoring recipes for the purely mechanical rows.

This post is based on an article I originally published on the AWS Builder Center. Any opinions in this article are my own.

comments powered by Disqus