All posts
guide

GPT-6.1 Sol cached input: what $0.10 saves you

The cached-input rate is 95% below ordinary input, but that is not a whole-bill discount. A monthly agent example includes writes, reads, fresh input and output.

EiliyaOctober 1, 20267 min read

GPT-6.1 Sol cached input costs $0.10 per million tokens, half GPT-6 Sol's $0.20 rate. If your agent keeps sending the same instructions and reference material, this is the price change most likely to matter.

Fresh input remains $2 and output remains $10 per million tokens under Standard short-context pricing. The 95% discount applies to cached reads against ordinary input, rather than your whole bill. You still pay for writes, changing input and output, and repeated prompts save money only when the API reuses their prefix.

GPT-6.1 Sol cached input: reads and writes have different prices

The OpenAI API pricing table lists these Standard rates in US dollars per million tokens:

Token category GPT-6 Sol GPT-6.1 Sol
Ordinary input $2 $2
Cache reads $0.20 $0.10
Cache writes $2.50 $2.50
Output $10 $10

These rates apply at up to 272,000 input tokens. Above that threshold, Sol 6.1 charges $4 input, $0.20 cache reads, $5 cache writes and $15 output, with higher rates applied to the full request.

The model specification puts writes at 1.25 times ordinary input and reads at 5% of ordinary input. Each token uses one input billing category. For a cache write, you pay $2.50 instead of the $2 ordinary rate.

If you write context and never reuse it, you've paid more than ordinary input would have cost. The cheap read rate helps only when you actually read it again.

The cache reuses your prompt prefix, then generates a new answer

The prompt-caching guide describes reusing computed model state for an unchanged prefix at the start of the rendered prompt. The model processes your new material and generates a new answer.

The complete rendered prefix has to match, including tool definitions and relevant request settings. Repeating a paragraph near the end of two different prompts won't establish that prefix.

For GPT-5.6 and later, the minimum cacheable prefix is 1,024 visible input tokens. Hidden OpenAI instructions do not count toward that minimum. Sol 6.1 belongs to this newer caching system, with both implicit and explicit breakpoints.

Check older caching advice against the current model. Before copying a retention setting or cache-key strategy, read the documentation that applies to Sol 6.1.

Put stable instructions first and timestamps later

For a coding agent, put stable developer instructions and shared reference material first. Keep timestamps, request IDs and newly retrieved snippets later, where changes won't disturb that prefix. Preserve earlier conversation messages and append new ones instead of rebuilding the history with altered wording.

Suppose your system prompt starts with the current time, followed by a long project guide. Every timestamp changes the prefix before the guide begins. Move it after the stable material to give the guide a chance to be reused.

Keep tool definitions, schemas and ordering consistent too. To make a tool unavailable for one turn, OpenAI recommends restricting which tools can run while keeping their definitions in the supplied list. Changes to tools, input, model and service tier can all cause an expected prefix to miss, according to the cache diagnostics guide.

For explicit caching, put the breakpoint where stable context ends, so you don't pay to write a changing tail you won't reuse. Top-level instructions can't contain an explicit breakpoint. Put reusable developer instructions in an input_text block inside an input developer message.

This request layout is illustrative. Replace the placeholder with stable material that meets the minimum length:

{
  "model": "gpt-6.1-sol",
  "prompt_cache_options": { "mode": "explicit" },
  "input": [
    {
      "role": "developer",
      "content": [{
        "type": "input_text",
        "text": "Stable project instructions and reference material...",
        "prompt_cache_breakpoint": { "mode": "explicit" }
      }]
    },
    {
      "role": "user",
      "content": "The changing task goes here."
    }
  ]
}

We haven't run this example. In explicit mode, omitting the breakpoint means the request neither creates cache writes nor uses prompt caching.

A monthly agent bill: $164.80 with assumed reuse

Take an invented workload of 10,000 requests in a month. Each has a 20,000-token stable prefix, 2,000 ordinary input tokens after it and 1,000 billable output tokens, including reasoning. Every request stays below the long-context pricing threshold.

Assume 100 independent groups of 100 requests, each writing its prefix once and fully reusing it on the other 99 requests while the entry is available. The changing tail is processed without cache writes. This is an assumed reuse pattern, not a measured hit rate.

The monthly prefix consists of 2 million write tokens and 198 million read tokens. The changing tail totals 20 million input tokens, and output totals 10 million tokens.

Monthly component No caching on Sol 6.1 GPT-6 Sol with assumed reuse GPT-6.1 Sol with assumed reuse
Stable prefix, ordinary input $400 $0 $0
Prefix cache writes $0 $5 $5
Prefix cache reads $0 $39.60 $19.80
Changing ordinary input $40 $40 $40
Billable output $100 $100 $100
Total $540 $184.60 $164.80

These calculations use published token rates, exclude tool fees, regional premiums and retries, and hold usage constant across models. Your real workload may use different token counts on each model.

Caching cuts this hypothetical Sol 6.1 bill by $375.20, about 69.5%, against disabling caching. Moving from the older Sol's cache rate saves a further $19.80, about 10.7% of that cached workload's total. That's what the read-rate discount does to a bill that also includes writes, fresh input and output.

If output dominates your bill, halving cache-read prices may save relatively little. If the prefix dominates and you reuse it reliably, the saving grows. Your application's mix of billing categories decides the result.

How soon does a cache write pay back?

For a 20,000-token prefix, an ordinary input pass costs $0.04, a cache write costs $0.05 and a Sol 6.1 cache read costs $0.002.

One write followed by one full read costs $0.052, compared with $0.08 for two ordinary passes. Under these assumptions, a single successful reuse already recovers the write premium. A write with no later reuse costs $0.01 more than ordinary input.

Use that calculation to choose which documents deserve a breakpoint. Stable material reused across turns is a better candidate than a tool result you'll use once. Check which tokens the API actually reports as writes or reads when measuring the bill.

A 30-minute lifetime makes reuse possible, not certain

For GPT-5.6 and later, OpenAI documents prompt_cache_options.ttl: "30m". A prefix remains eligible for reuse for at least 30 minutes after its most recent write or reuse, and reuse refreshes its lifetime without another write charge.

An eligible prefix can still miss. Cached state lives on individual machines, so your request has to reach one with a matching entry. Cache entries aren't shared between organizations or regional processing boundaries.

OpenAI handles routing automatically for this model generation. The optional prompt_cache_key separates cache accounting between customers or users; you don't need it as an optimization trick. Changing it can affect reported reuse, so keep it stable within the group you're measuring.

Budget for fresh writes when sessions resume after gaps. One write at the start of the month followed by perfect reuse would depend on a retention assumption your estimate needs to spell out.

Check cached tokens and dollars before changing your budget

Log usage.input_tokens_details.cached_tokens, cache-write tokens and total billable output. Compare reused tokens with all input tokens, including the first write and failed attempts. Looking only at the prefix you hoped would hit gives you an incomplete picture.

When reuse falls short, set prompt_cache_options.comparison_response_id to a recent baseline response in the diagnostics API. It compares requests and can explain a mismatch. The setting requests diagnostics without loading the earlier conversation or guaranteeing a hit.

Keep managing context even when it reduces reuse. Compaction can lower total input cost while replacing an existing prefix. Compare dollars and accepted results; chasing a higher cache-hit percentage while the conversation grows can miss those savings.

Cost labels matter in the launch discussion too. In a Reddit comparison of Sol and Opus, Background-Web-6312 reports that lower API task cost didn't translate into better subscription-quota value for their setup. That's an individual report with different effort settings, rather than an API pricing measurement.

For this guide, check your API bill. The cheaper cache rate gives you a chance to save; stable prompts and recorded usage show whether your app does.

FAQ

How much is GPT-6.1 Sol cached input?

Standard cache reads cost $0.10 per million tokens at up to 272,000 input tokens. Above that threshold, the rate is $0.20.

Do I pay extra to write the cache?

Cache-write tokens cost $2.50 per million under short-context Standard pricing. That replaces the ordinary input rate for those tokens; it is not an additive fee.

Does prompt caching return an old answer?

No. It reuses computed state for the matching prompt prefix. The model processes new input and generates a new response.

Why do repeated prompts miss the cache?

The rendered prefix or a relevant setting may differ, the entry may be unavailable, or routing may prevent reuse. Check actual cached-token usage and diagnostics.

Sources