← Back to the directory

Individual

JustVugg


Coverage1

Open source · Sep 26, 2026

Open-source framework Colibrì runs 744B-parameter GLM-5.2 on a 25GB laptop without a GPU

Developer JustVugg has released Colibrì, a pure-C, dependency-free inference framework for MoE models that keeps a model's dense layers in RAM and lazy-loads expert weights from NVMe SSD. A 12-core machine with 25GB RAM can run 744B-parameter GLM-5.2; 128GB pure CPU reaches about 1.8 tokens/s, and six RTX 5090s reach 5.8–6.8 tokens/s.

Colibrì supports GLM-5.2/5.3, DeepSeek V4 Flash, Qwen, OLMoE, 975B-parameter Inkling and 2.8T-parameter Kimi K3, with multi-tier VRAM/RAM/SSD caching and next-layer expert prediction at 71.6%. The GitHub repository has about 32k stars; a license is not stated.

Original sources (Chinese)

笔记本跑7000亿参数GLM!无GPU也行? SSD当显存用火爆GitHubqbitai