PRICING MECHANICS · 2 OCTOBER 2026

How image and audio models are actually billed

A walk through what providers charge, where the usage numbers come from, and why most cost tools get multimodal wrong. Written from every model price we track, read on 2 October 2026.

Image and audio models are billed in tokens, not per image or per minute. An image becomes output tokens; a second of speech becomes audio tokens. Audio tokens are priced separately from text tokens on the same model, and typically cost more — on OpenAI's realtime model, audio input costs eight times text input.

Cost tools that apply one input rate and one output rate to every request will understate a voice workload badly, and will quietly misprice an image workload too.

If you run a text feature, cost is easy to reason about. Tokens in, tokens out, two rates. Most people carry that model over to images and voice, and it breaks in three separate places at once.

This post covers all three: how providers actually calculate multimodal charges, where the numbers live in the API response, and what goes wrong in the tools people use to track them. Everything here is checked against the live model catalogue rather than written from memory, and the figures are dated so you can tell when they go stale.

Images are not billed per image

The intuition is that an image costs a fixed amount. It doesn't. The provider converts the image into tokens and bills those at the model's output rate, which means the same request can cost different amounts depending on size and quality settings.

Here are two real requests to the same model, at the same 1024×1024 size:

REQUESTINPUT TOKENSOUTPUT TOKENSRATIO
Image A244,160173×
Image B591,05618×

Same model, same declared dimensions, and a four-fold difference in output tokens. Any estimate built on "about N cents per image" is wrong for one of these two.

The rates make the shape clearer. Image-capable models carry an ordinary input rate and a very high output rate, because the output is where the picture lives:

MODELINPUT / 1MOUTPUT / 1MOUT ÷ IN
gemini-3-pro-image$2.00$120.0060×
gemini-3.1-flash-image$0.50$60.00120×
gemini-3.1-flash-lite-image$0.25$30.00120×
gpt-image-2$5.00$30.006×
gemini-2.5-flash-image$0.30$30.00100×

A text model usually prices output at four to six times input. An image model can price it at a hundred. If you are tracking an image feature with a blended rate, or with the "assume 75% input, 25% output" heuristic that several tools fall back on when they lack a breakdown, you are not approximating — you are producing a number with no relationship to the bill.

Audio tokens are a separate price on the same model

This is the one that costs people real money, and it is the least visible.

A model that handles speech has two input rates and sometimes two output rates: one pair for text, one pair for audio. The model's headline price is the text price. The audio price sits in a different field, and it is usually higher.

MODELTEXT INAUDIO IN×TEXT OUTAUDIO OUT×
gpt-realtime-2.1$4.00$32.008×$24.00$64.002.7×
phi-4-multimodal$0.08$4.0050×$0.32——
gemini-3.1-flash-live-preview$0.75$3.004×$4.50$12.002.7×
gemini-3.1-flash-lite$0.25$0.502×$1.50——
gemini-3.5-live-translate-preview$3.50$3.501×$21.00$21.001×

Three things worth drawing out of that table.

There is no standard multiplier. It ranges from 1× to 50×. Any tool that estimates audio cost by multiplying the text rate by a fixed factor is guessing, and the guess is wrong by an order of magnitude at both ends of that range.

Plenty of models have an audio input rate and no audio output rate. That usually means the model listens but answers in text. Transcription and classification workloads look like this.

Audio pricing is not rare, but it is unevenly published. Across every model we price, 135 carry a separate audio input rate and only 18 carry a separate audio output rate. So for a lot of speech-capable models, there simply is no published audio rate to apply — which is a fact a cost tool should tell you rather than paper over.

The practical consequence

Take a voice agent on gpt-realtime-2.1 handling 500 minutes of conversation a month. Priced at the text rate, the audio input looks like a rounding error. Priced correctly at $32 per million, it is eight times that. The bill does not arrive until the end of the month, and when it does, nothing in it explains which feature caused it.

Where the numbers come from, and the double-counting trap

Every provider returns a usage object. The trap is that providers disagree about whether the detailed counts are inside the headline counts or beside them.

OpenAI-style (inclusive). prompt_tokens contains everything. The audio, cached and reasoning counts are breakdowns of it:

{
  "prompt_tokens": 2000,
  "prompt_tokens_details": { "audio_tokens": 1500, "cached_tokens": 200 },
  "completion_tokens": 800,
  "completion_tokens_details": { "audio_tokens": 600, "reasoning_tokens": 0 }
}

Here the text input is 2000 − 1500 − 200 = 300 tokens, not 2000. Add the audio term on top of the full prompt_tokens and you charge those 1,500 audio tokens twice, once at the text rate and once at the audio rate.

Anthropic-style (exclusive). input_tokens excludes the cache counts, which arrive in their own fields:

{
  "input_tokens": 2000,
  "cache_read_input_tokens": 8000,
  "cache_creation_input_tokens": 1000,
  "output_tokens": 500
}

Here the subtraction you needed a moment ago is exactly wrong. Those 8,000 cached tokens are not in the 2,000, and subtracting them produces a negative count.

So the same formula is wrong in opposite directions depending on the provider. Applying the inclusive rule to Anthropic undercharges. Applying the exclusive rule to OpenAI overcharges. There is no single formula that works for both, which is why this has to be resolved per request rather than assumed once.

And the resolution cannot be done on provider name alone. An Anthropic model served through a cloud reseller keeps Anthropic's response shape while arriving under a different provider name — the catalogue contains entries like amazon-bedrock / au.anthropic.claude-sonnet-4-6. Resolve by provider and you will get those wrong every time.

Two more charges that hide in the same request

Cached tokens

If a model supports prompt caching, repeated prefixes are billed at a reduced rate — but the discount is not a constant. On claude-sonnet-5, a cache read costs $0.20 against a $2.00 input rate, a 90% discount. On gpt-4o-mini it costs $0.075 against $0.15, only 50%. Cache writes cost more than ordinary input, typically 1.25× for a short time-to-live and 2× for a long one.

So caching can make a bill worse. A feature that writes to cache constantly without reading back pays a premium for nothing.

Context tiers

Some models charge more once your prompt crosses a size threshold. The rate is not a single number; it is a staircase:

CONTEXT SIZEINPUT / 1MOUTPUT / 1M
base$2.50$7.50
over 32,000 tokens$5.00$15.00
over 128,000 tokens$6.25$18.50

That is a real model's published schedule. A long-document feature that creeps past 32,000 tokens doubles its unit cost without anything in the code changing. In the catalogue we read, 615 models publish tiered pricing of this kind and another 520 publish a separate rate for contexts over 200,000 tokens.

This one is worth checking even if you never touch images or audio. It is a text-model problem that most tools ignore.

How most tools calculate this, and where each breaks

APPROACHWHAT IT DOESWHERE IT BREAKS
Blended rateOne price per 1M tokens, input and output averagedImage models price output up to 120× input. The average is meaningless.
Two rates, no detailinput × rate + output × rateAudio tokens billed at text rates. Cached tokens billed at full price. Tiers ignored.
75/25 splitAssumes a ratio when only a total is reportedFails on any multimodal workload, where the real split is nothing like 75/25.
Gateway passthroughReports whatever the gateway chargedShows you the gateway's price, not the provider's. You cannot see the spread.
Hand-rolled loggingYour own table of ratesWorks until prices change. They change often, and silently.

The last one deserves a word, because it is what most teams actually do and it is a reasonable choice at first. Logging usage to your own database for one model and one feature is genuinely fine. The part that rots is the price table. Prices move, models are deprecated, new ones launch, and cache and audio rates get published weeks after the model does. Nobody notices a stale rate, because a stale rate still produces a confident number.

How we calculate it

Four decisions, and they are the reason our numbers differ from a spreadsheet.

1. Every component is priced separately

Text input, audio input, cached reads, cache writes, reasoning, text output and audio output each get their own rate from the model's own published schedule. No blending, no assumed ratios. A request to a realtime model with 1,500 audio input tokens and 300 text input tokens is priced as two different things, because the provider bills it as two different things.

2. The inclusive-or-exclusive question is resolved per request

Before any arithmetic, we determine which convention the response follows, checking the model identity before the provider name so that resold and gateway-served models resolve correctly. When we genuinely cannot tell — an unfamiliar vendor, an unusual model string — we do not guess. We price what we are sure of, leave out what we are not, and mark the request so you can see that we did.

3. Prices are re-synced daily, and cost is computed server-side

The price table is rebuilt from the upstream catalogue every day, so a price cut or a new model appears without anyone hand-editing anything. Cost is computed on our side at the moment the event arrives, from the token counts you send. We never accept a cost figure from the client, because a client-supplied cost is only as current as the client's own rate table.

4. An approximation is labelled as one

This is the decision that matters most, and it is the one that makes our numbers occasionally less tidy than a competitor's.

Some rates are not published. Of the active models we price, a large share have no published cache-read rate, most have no cache-write rate, and the overwhelming majority have no separate reasoning rate. Where a rate is missing we fall back to the standard industry multiplier — and then we flag that row as having used a derived rate rather than a published one.

Likewise, when a request carries audio tokens for a model with no published audio rate, we do not invent a multiplier. The observed range runs from 1× to 50×; a made-up number in that range is not an estimate, it is a fiction. We price those tokens at the text rate, which we know understates them, and we mark the request so the understatement is visible rather than hidden.

A number you can see is uncertain is worth more than a number that is quietly wrong.

The same principle governs what we do with a model we have never seen. It does not get priced at zero and filed as a success. It gets a status that says the model is unknown, so the gap shows up as a gap.

What we cannot price, and will not pretend to

Three honest gaps. If a tool tells you it has none, it has not looked hard enough.

Worked example: a voice agent

A support voice agent on gpt-realtime-2.1. One conversation turn reports:

prompt_tokens        2,000   (1,500 audio, 200 cached, 300 text)
completion_tokens      800   (600 audio, 200 text)

Rates: text in $4.00, audio in $32.00, cache read $0.40, text out $24.00, audio out $64.00, all per million tokens.

COMPONENTTOKENSRATE / 1MCOST
Text input300$4.00$0.00120
Audio input1,500$32.00$0.04800
Cached input200$0.40$0.00008
Text output200$24.00$0.00480
Audio output600$64.00$0.03840
Total$0.09248

The same request priced with two rates and no breakdown — 2,000 input at $4.00 and 800 output at $24.00 — comes to $0.0272. That is less than a third of the real cost, and the error grows with every second of speech.

Questions people ask

Are images billed per image or per token?

Per token. The provider converts the generated image into output tokens and bills them at the model's output rate. The same declared image size can produce very different token counts depending on quality settings, which is why a flat per-image estimate drifts.

Why is audio so much more expensive than text?

A second of speech carries far more tokens than a second of reading, and the model does more work per token. Providers price that separately. On OpenAI's realtime model audio input is 8× text input; on one Azure-hosted multimodal model it is 50×.

How do I convert minutes of audio into tokens?

You do not have to. The provider reports the audio token count in the usage object on every response. Conversion factors are only needed for forecasting before you have real traffic, and they should be treated as estimates rather than rates.

Does prompt caching work on multimodal models?

Often yes, and the cached tokens are reported alongside the audio and text counts. The discount varies by model — 90% on one flagship, 50% on another — so applying a single assumed discount misprices it.

Why does my cost tracker show $0.00 for an image request?

Usually the model is not in its price table. A missing price should produce a visible gap, not a zero. A zero recorded as a success is indistinguishable from a request that genuinely cost nothing, which means the error can sit in your data for months.

Can I just log usage to my own database?

Yes, and for one model and one feature you probably should. The part that becomes work is the price table: rates change, models are deprecated, and audio and cache rates are often published after the model launches. A stale rate still produces a confident number, which is why nobody notices.

Do tiered context prices affect me?

They do if your prompts grow. Several hundred models charge a higher rate above a context threshold — commonly 32,000, 128,000 or 200,000 tokens — so a feature that slowly accumulates context can double its unit cost with no code change.

What to do with this

If you run any multimodal feature, three checks are worth an afternoon.

  1. Find out whether your tracking separates audio tokens from text tokens. If it reports a single input number, your voice costs are understated and you do not know by how much.
  2. Check one image request by hand. Take the output token count from the usage object, multiply by the model's output rate, and compare it against whatever your dashboard says. If they disagree, the dashboard is using an assumption.
  3. Look at what your prompts actually cost at their current length. If a feature is near a tier boundary, the next prompt-engineering change could step it over.

None of that requires a product. It requires the breakdown the provider already gives you, applied against rates that are current today rather than on the day you wrote the code.

That is the whole job LLMtrack does: one call after each request, every component priced separately against rates re-synced daily, and anything we had to approximate marked as approximate. There is a live workspace with real multimodal traffic in it if you would rather see the numbers than read about them, and the free plan tracks the same way the paid one does.

Last updated 2 October 2026. Rates quoted were read from the upstream model catalogue on that date and change frequently; the figures in your own dashboard are synced daily and will be more current than this page.