OpenAI Ultrafast pricing: when six times the cost pays
Astra Ultrafast costs six times Standard API rates. Worked examples show the premium, while official docs explain Pro 500 access and subscription usage.
OpenAI Ultrafast charges $300 per million output tokens for GPT-6 Astra, six times the ordinary $50 rate. OpenAI advertises up to 300 generated tokens per second in Codex.
Less waiting sounds good. Before paying for it, check whether token generation is what keeps you waiting in the first place.
Ultrafast is a processing tier. Its API prices and subscription-usage rules are separate, and mixing them can leave you overestimating how much work a Pro 500 plan buys.
OpenAI Ultrafast pricing by token category
These are US dollar prices per million tokens for GPT-6 Astra. The table uses requests with at most 272,000 input tokens.
| Token category | Standard | Fast | Ultrafast |
|---|---|---|---|
| Ordinary input | $10 | $20 | $60 |
| Cached input reads | $1 | $2 | $6 |
| Cache writes | $12.50 | $25 | $75 |
| Output | $50 | $100 | $300 |
The official API pricing table lists all four categories. Ultrafast costs six times Standard and three times Fast at these rates. For cache-write tokens, you pay the write rate in place of ordinary input charges, as the prompt-caching guide explains.
Above 272,000 input tokens, Astra Ultrafast charges $120 input, $12 cached input, $150 cache writes and $450 output per million tokens. Those prices apply to the full request. Budget for the whole conversation you're sending, including its history.
Output billing includes reasoning tokens, as OpenAI's reasoning guide explains. A concise answer can still leave you with a large output bill.
Eight times faster generation doesn't mean eight times faster work
The DevDay recap advertises up to eight times faster token generation in Codex, reaching 300 tokens per second, and up to six times faster in the API.
Those claims cover generation. Your coding agent can still wait for a shell command, a test run, a web page or a slow external service. Faster text won't shorten those waits, so the figures don't promise an eightfold or sixfold reduction in total task time.
Suppose a task takes 120 seconds: 60 seconds generating tokens and 60 seconds in tools or other overhead. With sixfold faster generation, that first portion falls to 10 seconds. The whole task takes 70 seconds, roughly 1.7 times faster overall. This is a hypothetical example, not a benchmark.
If almost all the job's time goes into generating tokens, the improvement can be much larger. Measure that split before paying for the tier. Record both the first useful output and the completed result; they answer different questions.
What you pay for a fresh request, a cache hit and a batch
The examples below use invented token budgets and published rates. Each stays below the long-context threshold and excludes separate tool charges, regional premiums and retries. Billable output includes reasoning tokens.
A fresh interactive request
With caching disabled, 20,000 input tokens and 5,000 output tokens cost $0.45 on Astra Standard: $0.20 input plus $0.25 output.
The same counts cost $2.70 on Ultrafast: $1.20 input plus $1.50 output. The premium is $2.25 for the request.
If the premium saves one minute of useful waiting time, you're paying an implied $135 per hour saved. That's this scenario's arithmetic; productivity depends on what you do with the minute. Reviewing another task may be worth more to you than simply receiving an answer sooner.
A request with a reusable prefix
Assume 50,000 cached input tokens, 5,000 ordinary input tokens and 2,000 output tokens, with no cache writes on this request.
Standard costs $0.20: $0.05 for cache reads, $0.05 for fresh input and $0.10 for output. Ultrafast costs $1.20: $0.30, $0.30 and $0.60 respectively.
The premium falls to $1, but you still need to budget for the earlier write that populated the cache. Include that setup cost and check how often later requests actually hit.
A background batch of repeated jobs
For 1,000 copies of the first example's token budget, Astra Standard totals $450 and Ultrafast totals $2,700. The difference is $2,250 before any other charges.
I'd struggle to justify that premium for a background job with a loose deadline. A firm deadline and a measurable cost of delay give you something to weigh against it. Compare a suitable cheaper model too, before paying to run Astra faster.
What the $500 Pro 500 plan actually buys
The ChatGPT Work and Codex pricing page lists Pro 500 at $500 per month. The DevDay recap describes its allowance as 25 times the Plus allowance.
Pro 500 is the self-serve plan with Ultrafast at launch. The speed documentation also lists eligible Enterprise and Edu plans; Enterprise access is off by default and requires workspace permission. Buying credits won't give other self-serve plans access.
OpenAI distinguishes two usage multipliers:
| Billing route | Astra Ultrafast relative to Standard |
|---|---|
| Included subscription usage | 8 times |
| Purchased credits and eligible Enterprise usage | 6 times |
These multipliers tell you how usage is billed. They happen to resemble the speed claims, but they don't guarantee that speed. Work and Codex draw from the same allowance, so activity in both counts against it.
Pro 500's 25 times Plus usage doesn't buy you 25 times as many Ultrafast jobs. Model, context, tools and reasoning affect consumption. Check your usage dashboard with the workflow you intend to run.
Request Astra with service_tier: "ultrafast"
The API guide specifies model: "gpt-6-astra" and service_tier: "ultrafast" in a Responses API request. You select the tier while keeping Astra's model ID.
The guide says Astra access is broadly available to API users at initial rate limits. It recommends WebSockets for agents that make frequent tool calls, because a persistent connection can reduce overhead between requests. HTTP is also supported.
Keep connection behavior consistent when comparing tiers, or a networking difference may look like a model speedup. Use the same prompts and tools, and inspect the returned service tier alongside usage.
Ultrafast supports global processing and US data residency. EU and other non-US regional processing endpoints aren't supported, so the tier can't currently meet a requirement for inference residency outside the US.
Cerebras is confirmed for the earlier preview, not Astra
OpenAI's August Ultrafast preview explicitly says GPT-5.6 Sol Ultrafast runs on Cerebras. That preview advertised up to 750 output tokens per second and up to 14 times Standard speed for that earlier model.
The GPT-6 Astra launch materials and API guide we checked don't specify its hardware. Cerebras is confirmed for the older preview, but those primary sources don't establish the supplier for Astra or every later Ultrafast model.
Keep the speed numbers separate too: GPT-5.6 Sol's preview and GPT-6 Astra's DevDay release have different models and availability.
Pay for urgency when the saved time is worth the premium
In a Reddit discussion about Ultrafast, users worry about burning through their allowances. BannedGoNext points to an expensive software outage as a reason to pay for speed. That's a user's argument, without a measured return on investment.
Urgency changes the calculation. Interactive debugging, a live research session or a time-sensitive investigation can justify paying to reduce waits. Routine overnight work usually has a different budget.
Test one representative workflow and compare total task time, acceptance rate and dollars spent. Keep the tier if the faster usable result is worth the premium. A faster token stream is pleasant; a result you can use sooner is what you're paying for.
FAQ
How much does OpenAI Ultrafast cost?
Astra Ultrafast costs $60 input, $6 cached input, $75 cache writes and $300 output per million tokens for short-context requests. Longer prompts cost more.
Do I need Pro 500 for the Ultrafast API?
No. API access and billing are separate. Pro 500 is the eligible self-serve subscription route for Ultrafast in Work and Codex; eligible Enterprise and Edu plans also have access.
Is GPT-6.1 Sol Ultrafast available?
OpenAI says it is coming soon. On September 30, the published Ultrafast pricing table lists GPT-6 Astra, not GPT-6.1 Sol.
Is every Ultrafast task eight times faster?
No. OpenAI's claim concerns token generation in Codex. Tools, input processing and other overhead affect total completion time.
Sources
- OpenAI API pricing
- OpenAI prompt-caching guide
- OpenAI: DevDay 2026 recap
- OpenAI API Ultrafast guide
- ChatGPT Work and Codex speed documentation
- ChatGPT Work and Codex pricing
- OpenAI reasoning documentation
- OpenAI: August GPT-5.6 Sol Ultrafast preview
- Reddit: Ultrafast allowance and pricing discussion