A leading product engineering company, creating adaptive software solutions to improve operations, providing businesses with expert development services from across domain.
A leading product engineering company, creating adaptive software solutions to improve operations, providing businesses with expert development services from across domain.
DeepSeek API pricing added peak hours on 16 August. Four of the seven peak hours fall inside the Indian and UAE working day. What that does to costs.

DeepSeek API pricing stopped being a single number on 16 August 2026. It is now time-of-day dependent, and if your team works from India or the Gulf, the expensive window is your working day.
That second part is the bit nobody writing about this has bothered to convert into a local timezone, and it is the only part that changes what you should do.
The old structure was a flat rate. V4 Flash cost $0.14 per million cache-miss input tokens and $0.28 per million output tokens, all day, every day. From 16 August 2026 at 16:00 UTC that became a two-tier structure.
| V4 Flash, per 1M tokens | Old flat | Off-peak | Peak |
|---|---|---|---|
| Cache hit input | $0.0028 | $0.007 | $0.014 |
| Cache miss input | $0.14 | $0.22 | $0.44 |
| Output | $0.28 | $0.66 | $1.32 |
V4 Pro moved the same way, from $0.435 and $0.87 flat to $0.66 and $1.98 off-peak, and $1.32 and $3.96 at peak. The largest single jump is V4 Pro cache-hit reads, which go up 12.1 times at peak.
Note what happened to the off-peak column. This was not a peak surcharge bolted onto existing prices. Off-peak is itself higher than the old flat rate — output went from $0.28 to $0.66 before any peak multiplier applies. If you budgeted on the old numbers, your cheapest hour is now more than twice what you planned for.
Peak runs 01:00 to 04:00 and 06:00 to 10:00 UTC — seven hours a day, set to a Chinese working day.
Converted:
| Peak window (UTC) | India (IST) | UAE (GST) |
|---|---|---|
| 01:00 – 04:00 | 06:30 – 09:30 | 05:00 – 08:00 |
| 06:00 – 10:00 | 11:30 – 15:30 | 10:00 – 14:00 |
The second window is the problem. For a team in India, 11:30 to 15:30 is the core of the working day — the hours when developers are running the most requests, when your product's Indian users are most active, and when any interactive feature gets hammered. For a team in the UAE, 10:00 to 14:00 is the same story.
Four of the seven peak hours land inside the working day in both markets. A US-based team, by contrast, sleeps through all seven. The same price change costs a Bengaluru team roughly double what it costs a San Francisco one for identical usage patterns, purely on clock alignment.
While researching this we found a widely-linked pricing page still presenting the old flat rates and stating that the surcharge, though announced on 30 June, remained inactive. Multiple other sources report it live since 16 August.
We could not resolve this from the official API documentation, which did not render a price table to us at all. So treat this as the practical warning rather than a footnote: the number in your cost model probably came from a third-party page, and third-party pages disagree with each other right now. Pull your actual rate from the platform console and your own invoices before you commit to anything.
Three responses, in descending order of how much they are worth. None of them require changing model, which is where most teams start and where the least money is.
Move what you can off the clock. Batch work — embeddings, summarisation backfills, evaluation runs, content generation — does not care what time it happens. Anything that can be queued should run outside 01:00–10:00 UTC, which for an Indian team means overnight or late afternoon. This is the single largest lever and it costs one scheduler change.
Fix your caching before your model choice. Cache-hit input is still the cheapest thing on the table by two orders of magnitude, even at peak. A prompt whose stable prefix actually hits cache is dramatically cheaper than the same prompt with a volatile prefix, and most teams have never checked their hit rate. Do that before switching provider over price.
Then reconsider the provider. Because these APIs are OpenAI-compatible, moving is largely a base URL and model name change — the same property that makes a provider-agnostic integration layer worth the week it takes to build — which cuts both ways. It makes switching easy, and it means your provider knows switching is easy, so price moves like this one will keep happening. Build the abstraction that lets you move; do not assume today's cheapest stays cheapest.
The reason this matters beyond one vendor is that a lot of 2026 architecture decisions were justified on a price that no longer exists. "We can afford to call the model on every request" was true at $0.28 per million output tokens. At $1.32 during your working day it is a different conversation, and the system was designed around the first number.
When we build AI features for clients, the cost model is part of the architecture review rather than something finance discovers later. The specific question worth asking is not what the tokens cost. It is what happens to the unit economics if the price triples — because on this evidence, it can, with about six weeks' notice.
If your product depends on a per-request model call and you have not run that number, the arithmetic is worth doing before the next repricing rather than after.
16 August 2026 at 16:00 UTC, which is 17 August at 00:00 Beijing time. The flat rate was replaced with peak and off-peak tiers across all billing items.
01:00 to 04:00 and 06:00 to 10:00 UTC, seven hours a day, aligned to a Chinese working day. In India that is 06:30–09:30 and 11:30–15:30 IST; in the UAE, 05:00–08:00 and 10:00–14:00 GST.
V4 Flash output went from $0.28 flat to $0.66 off-peak and $1.32 at peak per million tokens. The largest single increase is V4 Pro cache-hit input, up 12.1 times at peak.
No. Off-peak is higher than the old flat rate. V4 Flash output was $0.28 flat and is now $0.66 off-peak, so even your cheapest hour costs more than twice what it did before the change.
Move batch work outside 01:00–10:00 UTC, then audit your prompt cache hit rate — cache-hit input remains the cheapest tier by a wide margin. Consider switching provider only after both.
Ready to take the first step towards unlocking opportunities, realizing goals, and embracing innovation? We're here and eager to connect.
11th Floor, O-Hub, Chandaka Industrial Estate, Infocity, Bhubaneswar, Odisha 751024
Level 4, 11 York Street Sydney Startup Hub Sydney, NSW – 2000
30 N. Đinh Nghệ, Phước Mỹ Sơn Trà, Đà Nẵng / Da Nang City – 550000
Level 25, AIDP Business Tower, Dubai Marina, United Arab Emirates
50 Beauchamp Street, Wellington, WGN 5028, New Zealand