Provider guide

LLM inference pricing for providers

Understand token prices, cached input, per-request charges and the difference between traffic estimates and settlement.

By LLM Gateway · Updated September 11, 2026

How does token pricing work?

LLM inference pricing commonly charges separately for input and output tokens. To estimate one request, multiply each token count by its corresponding rate per million, divide by one million and add any flat request charge. Cached input uses its own rate when that rate applies to the request.

Which price belongs in each field?

FieldUnitApplies to
InputUSD per million tokensUncached input tokens
OutputUSD per million tokensGenerated output tokens
Cached inputUSD per million tokensInput tokens served from an eligible cache
Per requestUSD per requestEach billable request, in addition to token charges

Use the token cost calculator to model your workload. Total input includes cached tokens: subtract the cached portion before charging the uncached input rate.

Compare like-for-like deployments

Start with the same canonical model, then compare context limits, quantization, supported features and service behavior. A low headline token price can represent a different deployment. Check the live catalogue for current mappings and prices.

Tiered, regional, time-dependent and media pricing need their own calculations. Airside’s catalogue-price shortcut only offers flat tariffs it can represent. It never turns a multi-tier offer into a single rate.

How do discount and gateway margin affect routing?

Airside lets a provider file a routing discount and the gateway margin it accepts. Approved settings affect routing economics alongside availability and performance. Changing a price or margin is not a promise of traffic: competing routes, supported capabilities and request requirements also matter.

Plan volume without mistaking capacity for demand

Multiply estimated request cost by the number of requests you expect to serve. Use measured input and output distributions from your own service, including longer requests, instead of assuming every request is average. Compare demand on the public rankings page and check the ceiling implied by your rate limits.

Is estimated traffic value a provider payout?

No. A token estimate describes usage at the rates you enter. It excludes infrastructure costs, discounts, gateway margin, credits and settlement adjustments. Airside’s traffic view is an operating report; payment and settlement follow the provider’s written agreement.

When do updated prices go live?

A price change is filed for review and applies after approval. Review your units, optional charges and regional prices before filing. Follow the listing guide to prepare a new deployment.

Put your API in front of gateway users.

Verify your provider, submit models and pricing, and track the traffic routed to your API.

Register your provider