How to Swap LLMs Without Breaking Your Product
Frontier models ship every few weeks and old ones get retired. How Israeli product teams build AI features that survive a model swap without a quality drop.
The email arrives on a Tuesday. The model your main feature runs on is being retired in ninety days, and the suggested replacement is “broadly equivalent.” It isn’t. Your carefully tuned extraction prompt starts returning a slightly different JSON shape, your summaries get 30% longer, and the eval suite you never built can’t tell you any of this.
We’ve watched this play out enough times to treat it as a design constraint rather than an incident. The model behind your AI feature is the most volatile dependency in your stack. It changes on someone else’s schedule, and it changes behaviour without changing its interface.
Why Model Swaps Break Products Quietly
Your prompts are tuned to one model’s habits
Every prompt that survived contact with production has been shaped against a specific model. You added “respond only with JSON” because one model liked to preface things. You shortened an instruction because another over-complied. None of that tuning is documented anywhere except in the prompt itself, and none of it transfers.
Swap the model and you keep all the workarounds for the old one’s quirks while inheriting a new set you haven’t met yet.
Nothing throws an exception
This is what makes model changes different from a library upgrade. A breaking API change fails loudly at build time. A model change compiles, deploys, returns a 200, and produces output that is 8% worse in a way no test asserts on. Support tickets show up three weeks later and nobody connects them to the deploy.
If your only signal is “did the call succeed,” you’re flying blind through the exact change that matters most.
Build the Seam Before You Need It
One interface, thin adapters
Every AI call in your product should go through a single internal function — something like complete(task, input) — not through the provider SDK directly. Behind it, keep thin adapters per provider that handle the differences in message format, tool-calling syntax, and token accounting.
When we built Agents Army, a platform where users compose teams of agents from different providers and talk to all of them at once, this wasn’t an optimisation. It was the architecture. Multi-provider was the product. But the same seam pays off in a single-provider app the first time you need to move.
Pin versions and keep the model in config
Use dated model identifiers in production, never floating aliases. An alias means your provider can change your product’s behaviour without you shipping anything, and you’ll hear about it from a customer.
Then put that identifier in configuration, not code. Changing a model should be an environment variable and a restart, so a swap and a rollback take the same thirty seconds. This is ordinary deployment and infrastructure hygiene applied to a dependency most teams forget is a dependency.
Separate the task from the model
Different jobs in your product deserve different models — classification on something small and cheap, the user-facing generation on something stronger. If your code routes by task name rather than calling one model everywhere, you can move a single task to a new model without touching the rest. You also get cost routing for free, which usually pays for the work on its own.
Make Evals the Gate, Not the Vibe Check
Score the candidate on real inputs
You need a set of real production inputs with known-good outputs — a few dozen is enough to start, a few hundred is comfortable. Run the candidate model against it and compare on what your users actually feel: correctness, format compliance, refusal rate, latency, and cost per action.
If you don’t have that set yet, building LLM evals is the prerequisite for everything else here. Without it, “the new model seems better” is a feeling, and feelings don’t survive a board question about why retention dipped.
Shadow traffic tells you what a staging environment can’t
Once the offline numbers look fine, run the candidate on a copy of live traffic without showing users the result. Log both outputs, diff them, and look at the disagreements. That’s where you’ll find the 3% of inputs your eval set never imagined — the malformed PDF, the mixed Hebrew and English message, the customer who writes in all caps.
Roll It Out Like an Infrastructure Change
Canary, and watch cost alongside quality
Move 5% of traffic, then 25%, then the rest, with a day at each step. Watch your quality metrics and your spend together — a model that scores better and costs 3x more per action is a margin decision, not just an engineering one.
This only works if you already have per-request tracing and cost logging in place. Retrofitting observability during a migration is how deadlines get missed.
Keep the old path warm
Don’t delete the previous adapter the day the new model ships. Leave it behind a config flag for a release cycle or two. The failure mode you’re protecting against isn’t a crash — it’s discovering in week three that a low-volume but high-value workflow got quietly worse.
Model churn isn’t going to slow down. The teams that handle it well aren’t the ones who pick the perfect model — they’re the ones who made picking a model a reversible decision. If you’re building AI features you expect to run for years, that’s the part we design for first.
Yaniv Amrami is founder of quickdev. He has led model migrations for Israeli startups running AI features in production across fintech, retail, and B2B SaaS.
Work with us
Ready to build something?
quickdev is a full-service software studio based in Tel Aviv. We build MVPs, SaaS platforms, mobile apps, and AI-powered products — fast and without compromise.
Let's Talk