A leading product engineering company, creating adaptive software solutions to improve operations, providing businesses with expert development services from across domain.

A leading product engineering company, creating adaptive software solutions to improve operations, providing businesses with expert development services from across domain.

Mobile App Development

DeepSeek API pricing: the peak-hour trap for Indian teams

DeepSeek API pricing added peak hours on 16 August. Four of the seven peak hours fall inside the Indian and UAE working day. What that does to costs.

DeepSeek API pricing: the peak-hour trap for Indian teams

DeepSeek API pricing stopped being a single number on 16 August 2026. It is now time-of-day dependent, and if your team works from India or the Gulf, the expensive window is your working day.

That second part is the bit nobody writing about this has bothered to convert into a local timezone, and it is the only part that changes what you should do.

What changed

The old structure was a flat rate. V4 Flash cost $0.14 per million cache-miss input tokens and $0.28 per million output tokens, all day, every day. From 16 August 2026 at 16:00 UTC that became a two-tier structure.

V4 Flash, per 1M tokensOld flatOff-peakPeak
Cache hit input$0.0028$0.007$0.014
Cache miss input$0.14$0.22$0.44
Output$0.28$0.66$1.32

V4 Pro moved the same way, from $0.435 and $0.87 flat to $0.66 and $1.98 off-peak, and $1.32 and $3.96 at peak. The largest single jump is V4 Pro cache-hit reads, which go up 12.1 times at peak.

Note what happened to the off-peak column. This was not a peak surcharge bolted onto existing prices. Off-peak is itself higher than the old flat rate — output went from $0.28 to $0.66 before any peak multiplier applies. If you budgeted on the old numbers, your cheapest hour is now more than twice what you planned for.

The seven hours that matter

Peak runs 01:00 to 04:00 and 06:00 to 10:00 UTC — seven hours a day, set to a Chinese working day.

Converted:

Peak window (UTC)India (IST)UAE (GST)
01:00 – 04:0006:30 – 09:3005:00 – 08:00
06:00 – 10:0011:30 – 15:3010:00 – 14:00

The second window is the problem. For a team in India, 11:30 to 15:30 is the core of the working day — the hours when developers are running the most requests, when your product's Indian users are most active, and when any interactive feature gets hammered. For a team in the UAE, 10:00 to 14:00 is the same story.

Four of the seven peak hours land inside the working day in both markets. A US-based team, by contrast, sleeps through all seven. The same price change costs a Bengaluru team roughly double what it costs a San Francisco one for identical usage patterns, purely on clock alignment.

Check which price table you are reading

While researching this we found a widely-linked pricing page still presenting the old flat rates and stating that the surcharge, though announced on 30 June, remained inactive. Multiple other sources report it live since 16 August.

We could not resolve this from the official API documentation, which did not render a price table to us at all. So treat this as the practical warning rather than a footnote: the number in your cost model probably came from a third-party page, and third-party pages disagree with each other right now. Pull your actual rate from the platform console and your own invoices before you commit to anything.

What to do about it

Three responses, in descending order of how much they are worth. None of them require changing model, which is where most teams start and where the least money is.

Move what you can off the clock. Batch work — embeddings, summarisation backfills, evaluation runs, content generation — does not care what time it happens. Anything that can be queued should run outside 01:00–10:00 UTC, which for an Indian team means overnight or late afternoon. This is the single largest lever and it costs one scheduler change.

Fix your caching before your model choice. Cache-hit input is still the cheapest thing on the table by two orders of magnitude, even at peak. A prompt whose stable prefix actually hits cache is dramatically cheaper than the same prompt with a volatile prefix, and most teams have never checked their hit rate. Do that before switching provider over price.

Then reconsider the provider. Because these APIs are OpenAI-compatible, moving is largely a base URL and model name change — the same property that makes a provider-agnostic integration layer worth the week it takes to build — which cuts both ways. It makes switching easy, and it means your provider knows switching is easy, so price moves like this one will keep happening. Build the abstraction that lets you move; do not assume today's cheapest stays cheapest.

The wider point about cheap models

The reason this matters beyond one vendor is that a lot of 2026 architecture decisions were justified on a price that no longer exists. "We can afford to call the model on every request" was true at $0.28 per million output tokens. At $1.32 during your working day it is a different conversation, and the system was designed around the first number.

When we build AI features for clients, the cost model is part of the architecture review rather than something finance discovers later. The specific question worth asking is not what the tokens cost. It is what happens to the unit economics if the price triples — because on this evidence, it can, with about six weeks' notice.

If your product depends on a per-request model call and you have not run that number, the arithmetic is worth doing before the next repricing rather than after.

Frequently asked questions

16 August 2026 at 16:00 UTC, which is 17 August at 00:00 Beijing time. The flat rate was replaced with peak and off-peak tiers across all billing items.

01:00 to 04:00 and 06:00 to 10:00 UTC, seven hours a day, aligned to a Chinese working day. In India that is 06:30–09:30 and 11:30–15:30 IST; in the UAE, 05:00–08:00 and 10:00–14:00 GST.

V4 Flash output went from $0.28 flat to $0.66 off-peak and $1.32 at peak per million tokens. The largest single increase is V4 Pro cache-hit input, up 12.1 times at peak.

No. Off-peak is higher than the old flat rate. V4 Flash output was $0.28 flat and is now $0.66 off-peak, so even your cheapest hour costs more than twice what it did before the change.

Move batch work outside 01:00–10:00 UTC, then audit your prompt cache hit rate — cache-hit input remains the cheapest tier by a wide margin. Consider switching provider only after both.

Written by

Akash Mohapatra

Akash Mohapatra

Co Founder & Director

10 Sep 2026

·

5 min read

Share

LET'S CONNECT

Connect with Creuto!

Ready to take the first step towards unlocking opportunities, realizing goals, and embracing innovation? We're here and eager to connect.

Contact Us

We don't just aim to fit in – we strive to stand out. Experience the perfect blend of innovation, excellence, and trust that makes us truly unforgettable. Discover the difference with Creuto.

© 2026 Creuto All Rights Reserved