Technical articles, tutorials, and insights
AirLLM streams transformer layers one at a time to the GPU, making trillion-parameter models runnable on consumer hardware — no quantization, distillation, or pruning required. 4 GB VRAM for Llama 3 70B, 8 GB for 405B, 12 GB for DeepSeek-V3 671B, and just 3.72 GB for Kimi K3 2.8T. 33.5k Stars, Apache 2.0.