Inference Cost
Also called: token cost
What it costs to run a model in production, billed per token in and per token out.
Input and output tokens are usually priced differently, with output the expensive side. Costs scale with conversation length because history is resent each turn, which is why naive chat apps get expensive fast. Caching, shorter context, and routing to smaller models are the standard levers.
In practice: Resending a 50-page document with every follow-up question, and paying for it every time.
Where this comes up
- DeepSeek V4.1 Flash Replaces V4 Pro: Pricing, Benchmarks, and What Changed Since V4 Flash
- GPT-6 Sol Benchmarks: How to Read the Scores and Run Your Own Evaluation
- GPT-6 Sol vs GPT-6 Astra: A Workload-Based Comparison
- GPT-Live-1 in the API: Pricing, Full-Duplex Voice, and the Configuration Behind Every Benchmark
- Will AI Replace DevOps Engineers? Reading the Divergence in the Data