Every week now there is a new “model router” promising to send each prompt to the perfect model. The pitch is seductive: cheap models for easy turns, expensive models for hard ones, and you stop overpaying. I run a routing setup in my own tooling, nothing fancy, a primary model with a fallback, and I want to separate what an LLM gateway genuinely does from what the marketing implies.

Start with the plumbing, because that is the part that actually works. An LLM gateway sits in front of a coding agent like Claude Code, Codex or Cursor and gives every tool one API, one key, one billing dashboard, and automatic failover when a provider has a bad day. That is real and useful. My own config has a primary provider and a named fallback, and the single time it earned its place was not a clever routing decision, it was the primary returning errors and the fallback quietly taking over while I kept working. Failover is the feature. The intelligence is the garnish.
The part that is oversold
The headline promise is dynamic routing: the gateway reads each prompt, judges its difficulty, and picks the model that fits. RouteLLM, the open-source version, makes this concrete. You configure a strong_model and a weak_model, and it scores each prompt’s complexity to decide which one answers.
The idea is sound. The problem is that “complexity” is exactly the thing that is hard to measure before you have the answer. A one-line prompt can require deep reasoning, and a long prompt can be trivial boilerplate. A router guessing wrong in the cheap direction gives you a bad answer to a hard question, which is the most expensive mistake there is, because you do not notice until later. Guessing wrong in the expensive direction just wastes money on an easy turn, which is the failure the router was supposed to prevent. So the router has to be very good to come out ahead, and “very good at predicting difficulty from the prompt alone” is close to the whole unsolved problem.
What the tiers actually cost
The reason anyone bothers is the price spread, and it is genuinely large. Rough 2026 input pricing per million tokens: budget models like DeepSeek V3 or Qwen3 sit around $0.07 to $0.40, mid-range like Claude Sonnet 4.6 or GPT-4o around $3 to $6, and premium like Claude Opus around $15 and up.
That is a 40x to 200x gap between the floor and the ceiling, so the incentive to route is real. But look at what the spread implies. If your work is mostly easy, pin the budget model and stop paying for a router to tell you what you already know. If your work is mostly hard, pin the strong model, because a router that occasionally downgrades a hard turn is actively hurting you. Dynamic routing only pays off in the messy middle, where your workload is a genuine mix and you cannot predict the ratio in advance. Most individual developers are not in that middle. Teams with many contributors and mixed workloads are the ones who actually have the distribution routing is built for.
How the tools split
There is a useful distinction hiding under the word “routing.” Some tools route between models for capability, and some sit in front for operations, and they are not the same product.
Cursor routes between Claude, GPT, Gemini and its own Composer model, and its argument is a hedge: if model leadership changes next quarter, you are not stranded on the loser. Codex is GPT-only, so it makes the opposite bet. That is a capability choice about which brains you can reach.
A gateway like OpenRouter is the operations version: one API for 200-plus models, one key, one bill, failover across providers. Aider, Cline and Roo Code accept a custom base URL, so you can point them at any OpenAI-compatible gateway and get that layer for free. That is not about picking the smartest model per turn, it is about not rebuilding auth and billing five times.
If you are choosing, know which problem you have. “I want to switch models without switching tools” is the gateway. “I want the machine to pick the model for me” is dynamic routing, and it is the harder, less proven half.
What I would actually do
Put a gateway in front of your agent for the operational wins, because those are real and boring and they work: one key, one bill, failover. That alone is worth the setup.
Then be honest about dynamic routing. Turn it on only if your workload is a real mix and you are willing to measure whether it saves money without quietly degrading answers, which means comparing outputs, not just invoices. If you cannot tell whether a cheap turn was wrong, you cannot tell whether routing is working, and a router you cannot audit is just a source of quiet regressions.
For most people the honest answer is smaller than the pitch. Pin one good model, keep a fallback for the day your provider breaks, and skip the complexity scoring until you have the workload that justifies it. The gateway is the part that pays for itself. The intelligence is a bet you should make on purpose, not by default.
Sources: Entelligence, 9 best LLM routers and model routing tools; Maxim, best LLM gateways for coding agents; Augment, AI model routing guide (RouteLLM strong/weak config); cowork.ink, model routing for AI agents (tier pricing); Requesty, agentic coding tools compared. Setup facts are from the author’s own machine on 27 September 2026: a primary-plus-fallback model config and litellm installed locally.







