An open-source inference engine called Colibrì has climbed past 37,000 stars and 4,000 forks on GitHub. What it does is straightforward: it lets MoE models with hundreds of billions to over a trillion parameters run on an ordinary person's computer. Using Zhipu AI's GLM-5.2 as an example, the project lists a minimum requirement of 16GB of RAM, with 24GB recommended, and a GPU is optional.
The project was created by Vincenzo Fornaro on July 1, is written entirely in C, has zero external runtime dependencies, and is licensed under Apache 2.0. Version 1.12.0, released on September 20, merged 81 pull requests; version 1.12.1 on September 24 merged another 96, 80 of them from outside contributors.
Weights live on disk, read layer by layer as needed
MoE models have a lot of total parameters but use only a small fraction of them per token. GLM-5.2 has 744 billion total parameters, of which only about 40 billion are activated per call. Colibrì's approach is to avoid loading the whole model into memory: the full weights, quantized to int4, sit on an NVMe SSD — about 372GB for GLM-5.2 and roughly 419GB for GLM-5.3.
The project treats VRAM, RAM, and SSD as one unified storage tier, which the author calls "just-in-time compilation of weights." The techniques include a per-layer least-recently-used cache, pinning frequently used experts permanently in memory, prefetching the next layer's likely experts one layer ahead, merging the experts needed by multiple tokens into a single disk read, overlapping disk reads with computation, and striping reads across two SSDs in parallel.
The README is upfront about the tradeoff on speed: when fast storage isn't available, things just get slower — model accuracy never drops — and speed itself comes with "no guarantees."
How fast does it actually run
The project provides several measured benchmarks:
- 6x RTX 5090 with all weights resident in memory: 5.8 to 6.8 tokens per second
- A CPU-only desktop with 128GB of RAM, after cache warm-up: about 1.8 tokens per second
- A single RTX 5070 Ti: 1.07 tokens per second
- The author's own 25GB-RAM development machine, cold start: 0.05 to 0.1 tokens per second
Version 1.11.0, released in mid-September, added support for DeepSeek V4.1 Flash — 552 billion parameters, 510GB of disk usage — reading the official weights directly on a CPU-only machine with no format conversion. The author recorded cold-start latency for a single turn dropping from 78.7 seconds to 25.1 seconds, with five-turn conversations stabilizing at 1.14 to 1.58 tokens per second.
The main new feature in 1.12.0 is called brio mode: the caller supplies a fixed set of answer options, and the engine reads out the probability of each option directly instead of sampling generated text — output token count is zero, and the answer can never fall outside the given options. For most local users, that's far more practical than waiting for a reply to be typed out one token at a time.
The supported-model list reads like a roster of Chinese open models
Colibrì currently supports nine model families: GLM-5.2/5.3 and the vision-capable GLM-5.3-Flash, DeepSeek V4 Flash and V4.1 Flash, Kimi K3, Qwen3.6 and Qwen3.8-Flash-Next, plus Inkling and OLMoE. Apart from the last two, all of them come from Chinese companies — Zhipu AI, DeepSeek, Moonshot AI, and Alibaba.
For developers in China, this has two practical uses. One is keeping data on-premises: for law firms, hospitals, and manufacturers that can't send documents to the cloud, a 128GB-RAM workstation with a 1TB SSD is enough to run a 744-billion-parameter GLM model locally. The other is lowering the barrier to entry: a model like Kimi K3, at 2.8 trillion parameters with int4 weights around 1.6TB and a 32GB RAM minimum, was previously all but impossible for an individual to run on their own machine.
The limitations are just as clear. At one to two tokens per second, a thousand-word reply takes over ten minutes to generate; the project's documentation doesn't estimate the wear from long, high-intensity random reads on an SSD. The author also cautions that the speedups from speculative decoding are empirical results — they may not reproduce on a different machine.
Sources: Colibrì project README, Colibrì release notes for versions 1.11 through 1.12.1, QbitAI, CocoLoop; speed figures reflect the benchmark configurations published by the project, and star counts are taken from the GitHub page.