Quantization
Storing model weights at lower numeric precision to cut memory and cost, with a small quality trade-off.
La quantizzazione riduce ogni peso da, per esempio, 16 bit a 8 o 4. Il modello diventa molto più piccolo e veloce con un degrado di solito modesto, ed è questo che permette a modelli capaci di girare su un laptop o su un telefono. Quanta qualità perdi dipende dal metodo e da quanto sei aggressivo.
In pratica: Un modello da 70B compresso per girare su una singola GPU consumer.
Where this comes up
- AI Technology Trends 2026: The Complete Guide to What's Next
- An AI Just Designed an AI Chip in Two Weeks: What Redwood Really Proves
- DeepSeek V4.1 Flash Replaces V4 Pro: Pricing, Benchmarks, and What Changed Since V4 Flash
- How to Set Up ChatGPT Locally for Work in 2026
- Top AI Image Generators in 2026: Tools, Pricing & Prompts