A leading product engineering company, creating adaptive software solutions to improve operations, providing businesses with expert development services from across domain.
A leading product engineering company, creating adaptive software solutions to improve operations, providing businesses with expert development services from across domain.
LLM cost optimization from real gateway data: open-weight models ran 56% of tokens but 14% of spend. How to tier models by task and cut API costs.

The most useful data on LLM cost optimization this month does not come from a pricing page. It comes from the traffic of thousands of production applications. Vercel's AI Gateway routes tens of trillions of tokens a month between apps and model providers, and its September report shows that open-weight models now handle 56% of those tokens — while taking only 14 cents of every dollar spent. Production teams have already worked out the pattern: send volume to cheap models, pay frontier prices only where the task earns it.
Vercel's Production Index, covering activity through August 2026, reports:
As The New Stack notes, token volume and money are different measures. Vercel counts input, output, reasoning, cached-input and cache-creation tokens. A cheap model can process most of the tokens while an expensive model takes most of the budget, and both can be the right choice.
The 56%-versus-14% gap is the whole lesson. It tells you that teams are not replacing frontier models with open-weight ones; they are adding open-weight models for the high-volume work — classification, extraction, summarisation, routing, first drafts — and keeping frontier models for reasoning-heavy steps, where errors are expensive.
It also warns against optimising the wrong number. A 23% drop in average price per token is mostly a mix effect: more of the traffic moved to cheaper models. The median team saw 7.6%. If you are budgeting from headline price cuts, you will over-estimate your savings.
From the gateway data and our own AI engineering work, the pattern that works has five parts.
A support-ticket triage, a document extraction, a code review: each is a task with a quality bar and a price. A cheaper model that needs two retries, or a longer prompt, can cost more per completed task than an expensive model that gets it right first time. Log tokens, model and outcome per task, and compare on that. Our guide to agent observability and cost controls covers the instrumentation.
Sort your workloads into those that need frontier reasoning and those that need throughput. Run an evaluation set for each task against two or three candidate models, including at least one open-weight model. Many teams find that a large share of their calls pass their quality bar on a model costing a fraction of their default.
The Fable-to-Opus shift is instructive: when Anthropic offered a cheaper model with a similar profile, teams moved within the lab because behaviour stayed consistent. Vercel's authors put it plainly: "Lab loyalty doesn't follow brand, it follows model profile, and consistency wins." Trying the next tier down from your current model is usually the cheapest experiment you can run, because prompts and output formats tend to carry over.
The same report shows the risk of standing still. Google's Gemini 3 Flash lost most of its gateway token share after a successor changed its profile, and more than three-quarters of the lost volume moved to other labs. A routing layer that can send a task to a second provider — tested, not theoretical — protects you from both price changes and model changes. We describe one approach in AI gateway model routing.
Prompt caching and shorter prompts cut cost on any model. Cached-input tokens are typically billed at a steep discount, so stable system prompts and shared context should be structured to hit the cache. That work pays off whichever model you end up on.
When errors are costly and hard to detect: multi-step reasoning, code changes across a codebase, decisions a customer sees directly, or tasks where a wrong answer looks plausible. The gateway's spend data suggests production teams agree — they still pay frontier prices for a minority of tokens because those tokens carry most of the risk. The mistake is using that model for everything by default.
Per token through an API, usually yes. Self-hosting them is a different calculation: GPUs, operations and utilisation decide the real price, and at low volume a hosted API is almost always cheaper, as we found when we checked claims about self-hosted LLM cost. Open-weight also brings choices about where the model is served and under which jurisdiction — relevant for regulated data, and for regional pricing quirks like DeepSeek's peak-hour pricing.
If your AI bill is growing faster than usage, our AI engineering services team can run that evaluation and build the routing layer, so the saving shows up per task rather than only on the price list.
Reduce LLM costs by measuring cost per completed task, testing cheaper models against a small evaluation set for each task, routing high-volume work to cheaper or open-weight models, keeping frontier models for reasoning-heavy steps, and structuring prompts to benefit from prompt caching.
Open-weight models are usually cheaper per token through hosted APIs. On Vercel's AI Gateway in August 2026 they ran 56% of tokens but took only 14% of spend. Self-hosting them is a separate calculation driven by GPU cost and utilisation.
A frontier model is worth the price when errors are costly or hard to detect, such as multi-step reasoning, codebase-wide changes or customer-facing decisions. For classification, extraction and summarisation, cheaper models often meet the quality bar at a fraction of the cost.
LLM model routing sends each request to the model best suited to the task, based on quality, cost and latency, rather than using one default model for everything. A routing layer can also fall back to a second provider when a model changes or fails.
The average price per token on Vercel's AI Gateway fell 23.2% in August 2026, largely because more traffic moved to cheaper open-weight models. Among teams with steady volume, the median saw a smaller 7.6% drop, so headline averages overstate typical savings.
Ready to take the first step towards unlocking opportunities, realizing goals, and embracing innovation? We're here and eager to connect.
11th Floor, O-Hub, Chandaka Industrial Estate, Infocity, Bhubaneswar, Odisha 751024
Level 4, 11 York Street Sydney Startup Hub Sydney, NSW – 2000
30 N. Đinh Nghệ, Phước Mỹ Sơn Trà, Đà Nẵng / Da Nang City – 550000
Level 25, AIDP Business Tower, Dubai Marina, United Arab Emirates
50 Beauchamp Street, Wellington, WGN 5028, New Zealand