Since June this year, OptiSol has been spending roughly $400 a day on Claude usage supporting presales and proof-of-concept work across our iBEAM modernisation leads. What started as a manageable trickle has become a fast-growing line item in our cost of service. At the same time, clients — now genuinely knowledgeable about what these agents can do, in a way they weren’t a year ago — have started pushing back harder on token spend.
I wrote earlier in this series about our first Sovereign LLM evaluation, where we benchmarked open-weight models against Claude Sonnet as a baseline. That evaluation began, honestly, as a response to client requirements in regulated industries wanting on-premises control over data and access. It has since become something else entirely: a genuine cost optimisation initiative, driven as much by our own economics as by any single client’s compliance requirement.
This piece picks up where that evaluation left off — because we hit a wall, and the way we got past it taught me something I did not expect about how our own agents had actually been built.
The wall we hit
Our team narrowed in on the Kimi model family — K2.7 and K3 — as the strongest open-weight candidates for our BA and QA agents. On our evaluation scale, they were consistently scoring around 16 out of 20, against Claude Sonnet’s baseline of roughly 20 out of 20. Respectable. Usable for some cases. But short of where we needed to be.
The obvious next move — throw a bigger model at the problem — was not available to us. Kimi is already a trillion-parameter model. There was no meaningfully larger open-weight option to reach for. We had hit a ceiling that model scale alone was not going to lift.
So we stopped looking at the model and started looking at the architecture around it.
Why the ceiling was there at all
The pattern our team found was specific: on-premises models performed comparably to Claude right up until a task required what we started calling a reasoning “leap” — a large, complex inferential jump the model had to make in a single pass. Frontier models like Claude Sonnet absorbed these leaps far more gracefully than the open-weight alternatives did.
The obvious question was why these leaps existed in our agents at all. Our BA, QA, and Dev agents were built from scratch on frontier models. As our use cases evolved, as the legacy applications we encountered grew more complex, we kept adding skills and instructions to help the agents cope — and Claude, being a frontier model, simply absorbed the growing complexity without complaint. The leaps were never planned. They accumulated.
A twin separated at birth
I want to use an analogy here, and I’ll admit upfront that I’m anthropomorphising — something that’s become common shorthand in how people describe AI behaviour, myself included on occasion. But this one earned its place, because it clarified the problem for my team faster than any technical description did.
There’s a familiar trope in Indian cinema from the eighties and nineties: two siblings, often twins, separated at birth. One is raised in wealth, with every resource available. The other grows up poor, in a hard-scrabble life, building resilience and resourcefulness out of necessity. Decades later, when the rich twin’s circumstances suddenly change and resources disappear, they struggle to cope. The poor twin, who never had the luxury of abundance, handles the same adversity with equanimity — because they built the skills for it from day one.
Our BA and QA agents, as they existed, were the rich twin. Born and raised entirely on frontier-model reasoning, they never had to develop the discipline of breaking a hard problem into smaller, well-defined steps — because Claude never forced that discipline on them. Pair that same agent, unmodified, with an on-premises model, and it struggles exactly the way the rich twin struggles: not because it lacks the underlying skill, but because it was never built to need it.
So my team made a decision: raise the poor twin. Build the agent from the ground up on open-weight, on-prem-compatible models only, from day one. Not adapt the frontier-raised agent to a smaller model — start over, and let the constraints of the smaller model force better design from the outset.
From leaps to hops
Once we committed to that, the reasoning leaps stopped being invisible. Building on a model without frontier-scale reasoning headroom forced us to confront every place where the agent had been relying on a single large inferential jump — and to redesign those moments into a sequence of smaller, well-defined steps instead.
For our BA agent specifically, that meant restructuring what had been a single monolithic document-generation pass into a phased, checkpointed workflow. The rebuilt version starts by setting only the input file and output location — no upfront assumptions — then extracts structured data from the legacy source only if it isn’t already available, and pulls simple counts of fields, program units, and blocks before touching the model at all, deliberately without loading the full extracted JSON into context. Only then does it read the relevant skill and template, without any ground-truth reference to lean on. Generation itself happens section by section: the largest, most error-prone sections are written in small batches against an explicit zero-miss verification gate, rather than in one continuous sweep, with every phase writing its output to file and a verification check running immediately afterward, comparing counts against the extracted source of truth before the agent is allowed to proceed. A separate, final validation phase reconciles the completed output against the original source and produces a traceability report, and the workflow only completes once both the output document and that traceability report exist.
The next step, now underway, is enforcing this structure at the runtime level with an orchestrator that issues a fresh model call for each phase, with a hard checkpoint between hops — rather than trusting a single long-running agent session to hold the discipline on its own. That comes with a real trade-off: each hop has less visibility into what came before it, since it isn’t carrying forward a long conversational memory. Cross-section consistency — identifiers, cross-references, tone — has to be maintained through the generated artefacts and structured data slices passed explicitly between hops, rather than relying on the model to simply remember.

What the numbers said
The results validated the redesign faster than we expected. Across a small set of representative legacy forms, Kimi K2.7 scored in the 16 to 17 out of 20 range under the new phased architecture. Kimi K3, on the same restructured workflow, reached 17 to 18 out of 20 — what our evaluation framework classifies as the sovereign-ready band.
The pattern held across every test case: near-perfect structural completeness — full field coverage, correct section shape, zero-miss inventories — with the remaining gap concentrated in domain depth rather than structure. Business rule completeness, message catalogue granularity, and non-English string retention in one of our test applications were the areas still trailing the frontier baseline, not the underlying architecture.
We have since pushed further. The current version of the BA agent is now consistently hitting 18 out of 20, with roughly a 10 percent improvement in overall requirements-coverage compared to where we started. Critically, none of that gain came from a bigger model. It came entirely from redesigning how the existing model was asked to do the work.
Why I believe this generalises
My team is now testing this same phased architecture against legacy codebases from other engagement leads, to see whether the accuracy gains hold outside the specific applications we first tested on. I’m quietly confident they will — because nothing about the phased-workflow redesign was specific to one legacy system. It was a response to a general property of smaller models: less headroom for large, ungoverned reasoning leaps, and a correspondingly higher payoff from explicit, checkpointed structure.
We’ve since applied the same unpack-and-rebuild approach to our Dev and QA agents, and early results are showing a similar pattern of improvement. I don’t think this is a coincidence specific to the BA agent. I think it’s a real property of how you have to build agents when you don’t have unlimited reasoning headroom to lean on.
Why this can't come fast enough
I’ll end with the plain business reality driving all of this. Our first iBeam agent-assisted legacy modernisation project only began in February this year — these agents are still, in the scheme of things, young. But the shift in how closely both OptiSol and our clients watch token cost has been sharp and recent: it has really only sharpened over the past quarter, once the capability question had largely settled in everyone’s mind. Nobody doubts what these agents can do anymore. Which means the conversation has moved, almost overnight, from “can this work” to “what does this cost, and does it need to cost this much.”
For a CTO watching a similar cost curve, the lesson I’d pass on is this: if you hit a capability ceiling moving from a frontier model to a smaller or on-premises alternative, resist the instinct to conclude you simply need a bigger model. Ask first whether the capability gap is actually an architecture gap — whether your agent was ever designed to work without frontier-scale reasoning to lean on, or whether, like ours, it just grew up rich, and nobody noticed until the money got tight.


