Creuto is now an OpenAI Select Partner Read More

Software Architecture & Technical

OpenAI Ultrafast: when is 6x the price worth the speed?

OpenAI Ultrafast costs 6x standard: $300 per million output tokens against $50. The arithmetic, the limits, and the workloads where it pays.

OpenAI Ultrafast: when is 6x the price worth the speed?

OpenAI Ultrafast is a straight trade of money for latency: GPT-6 Astra runs at six times the standard token price on every line of the bill, and the only thing it makes faster is token generation. It is worth buying where a human is waiting on the output, and a waste everywhere else.

The figures below come from OpenAI's own API pricing page, read on 3 October 2026, for short-context requests of 272K input tokens or fewer. Per million tokens:

GPT-6 Astra service tierInputCached inputOutput
Batch / Flex$5.00$0.50$25.00
Standard$10.00$1.00$50.00
Fast$20.00$2.00$100.00
Ultrafast$60.00$6.00$300.00

Long-context requests above 272K input tokens cost more again on every tier: Ultrafast is $120 input and $450 output per million against standard's $20 and $75. The multiple holds at six either way. GPT-6 Astra is the only model with a published Ultrafast price at all.

The strongest case for paying six times standard

Start with the version of the argument that actually wins, because it does win in places.

Token cost is not the expensive part of an interactive product. A developer sitting in Codex waiting eleven seconds for a diff, forty times a day, is burning salary that dwarfs the API line. OpenAI states that GPT-6 Astra Ultrafast generates tokens up to 8x faster than GPT-6 Astra in Standard mode in Codex, and up to 6x in the API. If that removes five seconds from a wait a person has forty times a day, the arithmetic on an engineer's hourly rate is not close.

The same logic holds wherever latency is the product rather than a property of it. A live support reply that lands in two seconds instead of nine is a different experience, not a cheaper one. A voice or sales-floor assistant that stalls mid-sentence gets abandoned. In those cases the right comparison is not Ultrafast against standard tokens — it is Ultrafast against the abandonment rate, and token price barely enters the calculation.

There is a quieter version too. Agentic loops multiply: twenty sequential tool calls at six times the generation speed compound into an end-to-end difference a user can feel, where a single call would not have been noticed.

What the premium costs on a realistic month

Work it with your own numbers; these are illustrative, built from published rates rather than anyone's actual bill. Take an interactive assistant doing 50,000 requests a month, each with 1,500 input tokens and 500 output tokens — 75M input and 25M output.

TierInput costOutput costMonthly total
Standard$750$1,250$2,000
Fast$1,500$2,500$4,000
Ultrafast$4,500$7,500$12,000

The premium over standard is $10,000 a month, or $120,000 a year, for one workload. That is a hire. It is defensible if the workload is the one your customers are staring at. It is indefensible if half those 50,000 requests are a nightly classification job nobody is watching.

Run the same volume as an overnight batch and the output side costs $625 instead of $7,500 — Ultrafast output is twelve times the Batch and Flex rate for the same model. Paying a 12x multiple to finish a job at 3am faster than 3am is the single most expensive mistake available on this pricing page.

So the first job is not choosing a tier. It is splitting the traffic: the path a human waits on, and everything else. Most teams have never drawn that line, which is also why model overuse shows up on the bill before it shows up in a dashboard.

What Ultrafast does not make faster

OpenAI is unusually direct about the limits of its own claim. The Codex documentation attaches this to the 8x figure: the comparison "measures token generation speed, not billing rates or overall task completion time".

Generation is one segment of a user-visible wait. The others are untouched:

  • Time to first token. OpenAI does not publish a time-to-first-token figure for Ultrafast. The tier is described as the fastest service tier and documented in terms of generation rate; if your p95 is dominated by the gap before the first character appears, nothing on that pricing page promises to move it.
  • Your retrieval step. If a RAG lookup takes 900ms before the model is called, it still takes 900ms. A team whose latency budget is mostly vector search will pay six times for the smallest slice of it.
  • Network round trips. OpenAI's own Ultrafast guide warns that "without a persistent connection, network overhead can reduce the latency gains", and strongly recommends WebSockets for agentic applications making many tool calls.
  • Framework overhead. Orchestration layers, serialisation, guardrail passes and your own middleware sit outside the inference call entirely.

The order of operations follows from that. Measure where the p95 actually goes before you buy a tier that only addresses one segment of it. If generation is 40% of your wait, a 6x generation speedup caps out at a 33% improvement end to end — for six times the token cost.

Does cached input pricing apply to Ultrafast?

Yes, and OpenAI publishes the number: cached input on GPT-6 Astra Ultrafast is $6.00 per million against $1.00 on standard, with cache writes at $75.00 against $12.50. Every column scales by the same six.

That matters more than it first looks for prompt-heavy work. The multiple does not improve, but the absolute premium shrinks. Take the same 50,000-request month with 80% of input tokens served from cache: standard lands at roughly $1,460 and Ultrafast at roughly $8,760, so the gap narrows from $10,000 to about $7,300. Prompt caching is worth doing on its own merits, and it is worth doing before you price Ultrafast rather than after, because it changes the number you are deciding against.

Who can actually buy OpenAI Ultrafast today

Availability is narrower than the launch coverage suggests, and it has a hard regional limit.

  • API. Available to all API users for gpt-6-astra with service_tier: "ultrafast", at default limits of 500,000 tokens per minute on usage tiers 1–3, 1M on tier 4 and 5M on tier 5.
  • ChatGPT Work and Codex. Pro $500 and eligible Enterprise and Edu plans only. OpenAI states that other self-serve plans "don't have access to Ultrafast at launch, even with purchased credits" — that excludes Plus, Pro $100, Pro $200 and Business. On Enterprise it is off by default until a workspace owner enables it.
  • Region. Ultrafast supports US data residency and global processing only, and is not available to workspaces that require inference residency outside the United States. For a regulated EU or UAE workload pinned to a regional endpoint, this decision is already made for you.

Two further details are worth having straight. Inside ChatGPT the multiplier is not uniformly six: included subscription usage is charged at 8x the standard rate while purchased credits and Enterprise pay-as-you-go are charged at 6x. And OpenAI's sources do not agree on which model is next. The API guide describes preview access for GPT-5.6 Sol; the ChatGPT documentation lists GPT-6.1 Sol as supporting Standard and Fast only; OpenAI's developer account said Ultrafast for GPT-6.1 Sol is "coming soon". Coming soon is not available, and no model other than GPT-6 Astra has a published Ultrafast price. VentureBeat's figure of roughly 300 tokens per second in Codex, and its derived $12/$60 estimate for a future Sol Ultrafast, are the outlet's, not OpenAI's.

The cheaper things to try first

OpenAI's latency optimisation guide contains the most awkward fact for the 6x tier. Its heuristic: "cutting 50% of your output tokens may cut ~50% of your latency", while "cutting 50% of your prompt may only result in a 1–5% latency improvement".

Shortening the output costs nothing and buys roughly half the generation time. The 6x tier, at a theoretical 6x generation speed, removes at most five-sixths of it. One of those is free.

Three more are close to free. A persistent WebSocket connection is OpenAI's own recommendation for tool-heavy loops, and its documentation reports up to roughly 40% faster end-to-end execution on rollouts with 20 or more tool calls. Streaming cuts perceived wait to the first token rather than the last. And sending the easy traffic to a smaller model is usually the largest single win available — GPT-6.1 Sol is $2 input and $10 output against Astra's $10 and $50, which is why routing by request type beats buying speed for all of it.

There is also a middle rung most of the coverage skipped. Fast mode on GPT-6 Astra is 2x standard — $20 input, $100 output — for up to 2.5x faster processing. Against Ultrafast's 6x it is a third of the price, and for many interactive paths it is the honest answer. We wrote about what a saved second costs on Fast mode when that tier launched, and the same test applies here with a steeper multiple.

What we would do

Split the traffic, then price each path separately. In the work we do on API development and integrations, the useful question is never which tier is fastest — it is which requests have a person on the other end of them, and that is usually a minority of the volume.

  1. Instrument before you switch. Record p50 and p95 split into retrieval, time to first token, generation and post-processing. If generation is not the largest segment, Ultrafast is the wrong purchase.
  2. Take the free wins. Shorter outputs, streaming, prompt caching, a persistent connection for tool-heavy loops.
  3. Route the interactive path only. One tier flag on the requests a human waits on, standard or Fast on the rest.
  4. Leave the batch path alone. Classification jobs, index builds and async summarisation queues belong on Batch or Flex at half standard, not at twelve times it.
  5. Re-measure, then decide. Compare the new p95 against the new invoice. If the wait did not move, the tier flag comes back off.

The decision Ultrafast forces is a useful one even if you never buy it: you cannot route by latency until you know which of your requests anyone is actually waiting for. Most teams find that out by reading the bill, which is the expensive order to do it in.

Frequently asked questions

OpenAI Ultrafast for GPT-6 Astra costs $60 per million input tokens, $6 per million cached input tokens and $300 per million output tokens on short-context requests, against $10, $1 and $50 on the standard tier. Long-context requests above 272K input tokens cost $120 input and $450 output.

The Ultrafast tier is worth paying for only where a person is waiting on the output, such as an interactive coding assistant or a live support reply. For nightly batch classification, RAG index builds or async summarisation queues, a six-times premium buys speed nobody is present to notice.

OpenAI does not publish a time-to-first-token figure for Ultrafast. The tier is documented in terms of token generation speed, and OpenAI's Codex documentation states the comparison measures generation speed rather than overall task completion time, so retrieval, network and framework overhead are unaffected.

OpenAI Ultrafast is available for GPT-6 Astra to all API users, and in ChatGPT Work and Codex on Pro $500 and eligible Enterprise and Edu plans. OpenAI states other self-serve plans have no access at launch even with purchased credits, and it is off by default on Enterprise.

Yes, cached input pricing applies to Ultrafast. OpenAI publishes a cached input rate of $6 per million tokens for GPT-6 Astra Ultrafast, against $1 on standard, with cache writes at $75 per million. Caching does not change the six-times multiple, but it shrinks the absolute premium on prompt-heavy workloads.

Ultrafast supports US data residency and global processing only. OpenAI states it does not support EU or other non-US regional processing endpoints, and is unavailable to workspaces that require inference residency outside the United States, which rules it out for some regulated deployments.

Written by

Akash Mohapatra

Akash Mohapatra

Co Founder & Director

2 Oct 2026

·

9 min read

Share

LET'S CONNECT

Connect with Creuto!

Ready to take the first step towards unlocking opportunities, realizing goals, and embracing innovation? We're here and eager to connect.

We don't just aim to fit in – we strive to stand out. Experience the perfect blend of innovation, excellence, and trust that makes us truly unforgettable. Discover the difference with Creuto.

© 2026 Creuto All Rights Reserved