AirLLM is a Python library that lets 70B parameter language models run on a single 4GB GPU, without
Arynnraj · LinkedIn · source
AirLLM is a Python library that lets 70B parameter language models run on a single 4GB GPU, without quantization, distillation, or pruning.
The problem it’s solving is access. Large open-source models keep getting released, but running them normally requires enough GPU memory to hold the entire model at once, which puts most of them out of reach for anyone without high-end or multiple GPUs.
AirLLM’s approach: during inference, the original model is first decomposed and saved layer-wise. That’s how it avoids needing the full model resident in memory at once. For MoE models specifically, it goes further and streams one expert at a time rather than a whole layer, since a token typically doesn’t need every expert to run.
It works with almost any popular model, you just pass a Hugging Face repo ID, and it works the same way regardless of model size or family, Llama, Qwen, DeepSeek, Mistral, Phi, Gemma, ChatGLM, Baichuan, InternLM, and others are all supported. There’s also an optional compression feature, block-wise quantization that delivers up to 3x faster inference with almost ignorable accuracy loss.
It’s Apache 2.0 licensed, has 25.5k stars, and has been actively maintained, with new model support added on an ongoing basis.
Here’s the GitHub Repo: https://lnkd.in/dV7FDsWG