The vLLM team made Kimi K3 run 57% faster on Google TPUs than on NVIDIA GPUs
Inferact, a startup founded by the creators and core maintainers of the open-source vLLM engine, has posted a striking result: on 16 Google TPU v7 tensor accelerators, the Kimi K3 model serves 709 tokens per second, while 16 NVIDIA GB200 GPUs manage only 452. That is a 57% edge for the TPU — even though NVIDIA's hardware looks stronger on paper. The trick is not in the silicon but in the software layer: the team wrote a "megakernel" that fuses hundreds of small inference operations into one large program. The code is open source.
What exactly was compared
Inferact put TPUs and GPUs on identical footing: 16 accelerators per side, the same Kimi K3 model, the same vLLM engine. Only the chips and the low-level kernel implementation changed. With speculative decoding enabled, the TPU delivered 709 tokens per second against 452 for the GB200.
Turn speculative decoding off and look at raw speed, and the gap does not disappear. At batch size 1 the TPU does 249 tokens per second versus 127 for the GB200 — nearly double. At batch size 8 it is 865 versus 636. On a different model, Qwen 3.8 27B, the gap is even wider: 4 TPU chips reach 1,515 tokens per second while 4 GPUs reach 695.

Why the "weaker" chip came out ahead
On paper the TPU v7 Ironwood and the GB200 are close: peak BF16 throughput of 2.31 versus 2.5 PFLOPS, FP8 of 4.61 versus 5 PFLOPS. On memory bandwidth the TPU actually loses — 7,380 GB/s against 8,000 GB/s. So NVIDIA's memory is faster while its inference is slower. The difference is not in the hardware but in how that hardware is used.
Language model inference is not bound by compute but by moving data. To generate a single token, every weight in the model has to pass through the chip: pulled from HBM into on-chip cache, used for math, then pulled again for the next token. Normally this is split into hundreds of separate small programs — kernels. While one kernel finishes and the next has not started yet, memory sits idle. Those pauses add up to a real loss.
A megakernel: one program instead of a hundred
Inferact's idea is to remove the boundaries between kernels. The entire Kimi K3 model — 92 mixture-of-experts layers — fits into a single program written in Pallas, the low-level TPU language that plays the role CUDA plays for NVIDIA. With no boundaries, the next layer's weights can start loading early while the current layer is still computing. While layer N is busy with its experts, layer N+1's attention weights are already moving from memory onto the chip. Memory is almost never idle.
The trick fits TPU architecture well. Each TPU v7 tensor core has 64 MiB of fast on-chip memory, and the programmer decides what lives there and when to free it. GPUs are organized differently: roughly 38 MiB across the whole card, split among 152 blocks and managed by hardware — there is nowhere to stage weights ahead of time.
A side benefit is compile speed. A conventional TPU compiler takes more than 30 minutes to build a large model, while Inferact's megakernel compiles in under 90 seconds. For an engineer, that is the difference between checking a change tomorrow and checking it a minute from now.
Speculative decoding and accuracy
The second layer of speedup is speculative decoding via the DSpark scheme proposed by the DeepSeek team. A small draft model quickly proposes several candidates for the next token, and the large model verifies them as a batch: correct guesses are accepted immediately, wrong ones are discarded. On the TPU the average accepted chain reached 6 tokens, and a single decode step takes about 8.5 milliseconds.
Importantly, none of that speed came at the cost of quality. With greedy decoding and maximum reasoning effort, the megakernel scored 94.4% on GPQA-Diamond and 97.2% on GSM8K — exactly the same numbers as on GPU.
Who did it
Inferact is a company of people from the vLLM team, the most popular open-source engine for serving neural networks. CEO Simon Mo is one of the vLLM maintainers, co-founder Woosuk Kwon created the project and the PagedAttention algorithm, and chief scientist You Kaichao came from Tsinghua University. This year the company raised $150 million in seed funding at an $800 million valuation. The megakernel work is a joint project with Google Cloud, and all the code is open.
What it means in practice
The main takeaway is simple: in the race for inference speed, hardware is no longer the only thing that decides. Two chips of the same class, one model, one engine — and a 1.5x difference that lives entirely in the software layer. For anyone paying for compute, that means code optimization can deliver more than buying new accelerators.
The caveats are honest too. The megakernel is currently tailored to the specific Kimi K3 architecture and a specific 16-chip configuration — a different model or a different number of accelerators means rewriting it. This is a first step, not a finished universal solution.
FAQ
What is a megakernel?
One large program that runs the whole model pass at once instead of hundreds of separate small programs. That removes the pauses between them, and memory gets to prefetch data ahead of time.
Are TPUs now better than GPUs for AI?
Not so clear-cut. This compared a narrow scenario — serving one model at a small batch. Other workloads, especially training, may look different. But the result shows that software optimization can outweigh a hardware gap.
Can I try Kimi K3?
Yes, the model weights are open. You can talk to it in NeuralSpace Chat and try agent-style workflows in Code.