← All articles

A 744B-parameter model now runs on a laptop: Colibrì streams its experts from an SSD

A glowing glass core next to a huge translucent structure of many glowing cubes on a bright surface

The open-source Colibrì engine, written in pure C with no external dependencies, runs models from 744 billion to 2.8 trillion parameters on an ordinary laptop — with no GPU required. It does not achieve that by crushing the model beyond recognition: the engine keeps the "experts" — the parts a token does not always need — on disk and loads only the ones the current token actually uses. The project has nearly 38,000 stars on GitHub and already supports nine model families.

What the project is

Colibrì is an inference engine for large models, written by a single developer under the handle JustVugg. It began as an experiment on a 12-core laptop with 25 GB of RAM and grew into a project with 37,951 stars and 4,124 forks on GitHub. The licence is Apache 2.0 and all the code is open.

The engine itself is one C file plus small headers. No BLAS, no Python at runtime, no mandatory GPU. It builds with plain gcc or clang and OpenMP, and prebuilt binaries exist for Linux, macOS and Windows.

Why this is possible at all

The secret is in how modern models are built. GLM-5.2 has 744 billion parameters, but only about 40 billion are active for any given token: a router picks which "experts" will compute that token. And only a small slice changes from token to token — roughly 11 GB of weights.

Hence the project's central idea: the model does not have to fit in fast memory — it has to be placed. The dense part (attention, shared experts, embeddings — about 17 billion parameters) stays resident in RAM and takes about 9.9 GB at int4. The 19,456 routed experts (75 MoE layers of 256 each plus the MTP head, roughly 19 MB each at int4) live on disk — about 370 GB — and are streamed on demand.

A glowing glass core next to a huge translucent structure of many glowing cubes on a bright surface

A JIT for weights

The author describes the engine's core as a JIT compiler, but for weights. A normal JIT does not compile the whole program — it watches what actually runs and compiles the hot paths. Colibrì makes the same bet: parameters are not state to be held, they are data to be staged just in time.

Measured routing heat decides which expert earns which tier. Each layer has an LRU cache, a learned pinned hot-store, and one-layer-ahead prefetch: the router runs before its results are needed, so the staging latency hides behind compute. The longer you run it, the better the engine predicts your hot experts.

The three memory tiers — VRAM, RAM and NVMe — are placement tiers for the same weights. Limited fast memory changes speed, not the model: weight precision and router choices stay the same, and an expert served from VRAM and one served from disk produce the same result.

The numbers

All measurements use the same GLM-5.2 model in an int4 container; only the hardware changes, meaning where the experts live.

  • 6× RTX 5090, full residency: 5.8–6.8 tokens per second, time to first token about 13 seconds.
  • 128 GB CPU-only desktop: about 1.8 tokens per second on a warm cache.
  • Laptop with a single RTX 5070 Ti: 1.07 tokens per second.
  • The machine it all started on — 12 cores and 25 GB of RAM: 0.05–0.1 tokens per second on a cold cache. That is the project's honest floor.

Two independent NVMe drives were measured separately: +37.5% decode speed. A third, slower drive added nothing after weighted striping.

What it supports

The engine now covers nine model families, each with its own C file, while shared components (I/O, cache, tokenizer) are reused:

  • GLM-5.2 and GLM-5.3 — 744B parameters, about 372 and 419 GB on disk, from 16 GB of RAM;
  • GLM-5.3-Flash — 321B, with vision;
  • Inkling — 975B;
  • Kimi K3 — 2.8T parameters, about 1.6 TB on disk and from 32 GB of RAM, no GPU needed;
  • DeepSeek V4 Flash — 284B and DeepSeek V4.1 Flash — 552B, both with vision;
  • Qwen3.8-Flash-Next — 125B plus 51B in n-gram memory;
  • Qwen3.6 — 35B and OLMoE — 7B.

There is a web dashboard with live metrics and a separate Brain page that visualises all 19,456 experts of GLM-5.2: you can see which expert the router just called and which memory tier it currently sits in. There is also a local cluster mode: a coordinator handles tokens and routing while expert workers on other machines take on the compute.

What the authors say honestly

The project deliberately promises no stable speed. The README states it plainly: there is no SLA on speed, but there is a hard guarantee on semantics — experiments must earn their place through reproducible measurements, and the default policy never silently changes model precision or router behaviour.

Some ideas remain hypotheses. Learned pinning of hot experts wins on repeatable workloads but can overfit a specific prompt. Speculative decoding on some configurations measured a 32% loss rather than a gain at around 85% expert hit rate. Direct disk access depends heavily on the drive model. The author invites the community to test these hypotheses and publish negative results too.

What it means

The main takeaway is simple: working with a large model does not require renting a server rack. If the model is sparse, it can be placed across memory tiers and run on what is already under your desk. Speed will be modest — from a fraction of a token per second on a weak machine to a few tokens on a fast one — but it is a real local run, not a demo.

If you would rather just talk to modern models without juggling disks and builds, there is NeuralSpace Chat; image generation lives in Images, and video in Video.

FAQ

What is Colibrì?

An open-source inference engine for large models, written in pure C. It does not hold the whole model in fast memory; it places it across tiers — some in RAM, some on disk, and some in VRAM when available.

Do I need a GPU?

No. A GPU only makes it faster, and none of the supported models requires one. Speed is set mainly by the drive the experts are streamed from.

How fast is it?

A fraction of a token per second on a weak machine, a few tokens per second on a fast one. It is not a speed replacement for a cloud API, but it is a full local run with no rented servers.

Where do I get it?

The code and prebuilt releases are in the open JustVugg/colibri repository on GitHub. Source: 量子位 (QbitAI).