Quantization
Storing model weights at lower numeric precision to cut memory and cost, with a small quality trade-off.
Quantization shrinks each weight from, say, 16 bits to 8 or 4. The model gets dramatically smaller and faster with usually modest degradation, which is what allows capable models to run on a laptop or a phone. How much quality you lose depends on the method and how aggressive you get.
In practice: A 70B model compressed to run on a single consumer GPU.
Where this comes up
- AI Technology Trends 2026: The Complete Guide to What's Next
- An AI Just Designed an AI Chip in Two Weeks: What Redwood Really Proves
- DeepSeek V4.1 Flash Replaces V4 Pro: Pricing, Benchmarks, and What Changed Since V4 Flash
- How to Set Up ChatGPT Locally for Work in 2026
- Top AI Image Generators in 2026: Tools, Pricing & Prompts