Inference Cost
Also called: token cost
What it costs to run a model in production, billed per token in and per token out.
Eingabe- und Ausgabe-Token werden meist unterschiedlich bepreist, wobei die Ausgabe die teure Seite ist. Die Kosten wachsen mit der Gesprächslänge, weil die Historie bei jedem Zug erneut mitgeschickt wird, und deshalb werden naive Chat-Apps schnell teuer. Caching, kürzerer Kontext und Routing auf kleinere Modelle sind die üblichen Hebel.
In der Praxis: Ein 50-seitiges Dokument bei jeder Rückfrage erneut mitschicken — und jedes Mal dafür bezahlen.
Where this comes up
- DeepSeek V4.1 Flash Replaces V4 Pro: Pricing, Benchmarks, and What Changed Since V4 Flash
- GPT-6 Sol Benchmarks: How to Read the Scores and Run Your Own Evaluation
- GPT-6 Sol vs GPT-6 Astra: A Workload-Based Comparison
- GPT-Live-1 in the API: Pricing, Full-Duplex Voice, and the Configuration Behind Every Benchmark
- Will AI Replace DevOps Engineers? Reading the Divergence in the Data