A model's throughput cannot be higher than the speed of loading the weights used for compute. That speed depends on the hardware used. Here are the numbers for a few leading GPUs and TPUs:
| Hardware | Memory bandwidth | Speed of light, Gemma 4 26B-A4B |
|---|---|---|
| NVIDIA H100 | 3.35 TB/s | 427 tok/s |
| NVIDIA H200 | 4.8 TB/s | 612 tok/s |
| NVIDIA B200 | 8.0 TB/s | 1,020 tok/s |
| Google TPU v5e, 8 chips | 6.55 TB/s | 835 tok/s |
| Google TPU v6e, 8 chips | 13.1 TB/s | 1,671 tok/s |
Speed of light is memory bandwidth divided by the bytes read per token: 7.8 GB for one user with a 1K prompt and BF16 weights.
This is the theoretical limit for decode speed. Inference engineers call it the speed of light and chase it by hiding the compute cost behind the weight loads. It doesn't always work, as it depends on the model's architecture and the hardware it's served on.
Popular serving stacks like vLLM or SGLang are built to be easily extensible, which makes it harder for them to reach this speed. A megakernel closes this gap with a single kernel running the whole decode step, keeping the memory busy the entire time. We worked on this for two Gemma models, Gemma 4 31B (dense) and Gemma 4 26B-A4B (MoE, 128 experts).
All numbers are greedy decoding with no speculative decoding, for one user and a 1K prompt. The megakernel runs on a TPU v5e-8 with BF16 weights. The baseline is vLLM 0.30 on one H100, with FP8 weights for the MoE model and BF16 for the dense model.
When a megakernel pays off
A megakernel implemented correctly will match performance of vLLM/SGLang, and it is clearly faster only in specific scenarios. The time spent building the kernel and the maintenance after is not worth it, if the difference is a few ms. So, after trying this out for different architectures, we decided a megakernel is worth building in cases like if the model has many small kernels (think MoE / small models) or vLLM/SGLang out of box is below half the theoretical max or if we are not building a high batch use case, as decode becomes compute bound on a single GPU in such cases.
The measurements behind each of these:
| Verdict | Situation | Measured against a general stack |
|---|---|---|
| Works | MoE model, one user | 3.7× faster on Gemma 4 26B-A4B, 2.0× on Kimi K31 |
| Works | Small model, one user | almost 2.5× faster on Llama 1B2 |
| Works | Long prompts | 4.4× faster at 8K on Gemma 4 26B-A4B |
| Helps less | MoE model, eight users | 1.6× faster on Gemma 4 26B-A4B, 1.4× on Kimi K31 |
| Helps less | Large dense model | 2.5× faster than vLLM on an H100 on Gemma 4 31B, but only 1.2× faster than vLLM on the same TPU |
We tested this by building a megakernel for Gemma 4 dense and MoE to compare the savings.
Megakernels for Gemma 4
We started with the open-source TPU megakernel1 and iterated on it for both Gemma models. On the MoE model, when we assigned whole experts to a single chip, we started seeing straggler chips at the all-reduce after MoE routing: one chip at 2.4 experts per layer against an average of 1.0. We resolved this by splitting each expert across all eight chips instead of assigning it to one, and the shared MLP is split the same way. The other optimization is overlapping the MoE all-reduce with the next weight load. Together with a smaller payload in the all-reduce, these cut the per-step time from 3.46 to 1.56 ms.
vLLM on an H100 runs 1,494 GPU kernels per token on this model and reaches 21% of the speed of light. The megakernel runs one and reaches 77%. Here is where the time goes in each:
| Component | vLLM, H100, FP8 | Megakernel, TPU v5e-8 |
|---|---|---|
| Experts | 1.76 ms | 0.51 ms |
| Sliding-window attention, 25 layers | 1.40 ms | 0.08 ms |
| Norms, residuals, quantization | 0.86 ms | 0.02 ms |
| Shared MLP | 0.52 ms | in the stream |
| Output head | 0.48 ms | 0.003 ms |
| Full attention, 5 layers | 0.45 ms | 0.03 ms |
| Router | 0.16 ms | 0.02 ms |
| Whole step | 5.77 ms | 1.56 ms |
The vLLM column is GPU kernel time from the profiler. The megakernel column is how much the step shrinks when that component is switched off, so it does not add up to the whole step.
For the dense 31B model, vLLM already reaches 74% of the speed of light on the same TPU and the megakernel reaches 90%. Here is a graph of tps vs token length for both the models:
What's next
A megakernel is worth building when out-of-the-box vLLM or SGLang's decode speeds are far below the speed of light. For the Gemma 4 MoE model, 640 tok/s per user on eight TPU v5e chips is a considerable upgrade over 173 on an H100. A few other ideas we are exploring from here:
Speculative decoding: Fusing speculative decoding into the megakernel should raise the speed even further. Inferact reports over 700 tok/s on Kimi K3 this way1.
More users per GPU: At eight users the gain drops from 3.7× to 1.6×, because attention and vector arithmetic grow with every row. This work needs to be split across chips for higher concurrency.
Quantization and better GPU: With INT8 experts the same kernel already reaches 760 tok/s. TPU v6e has twice the memory bandwidth per chip, which puts the speed of light for this model near 1,670 tok/s.
- The megakernel design for TPUs was published by Inferact in A Case for TPU Megakernels. Our kernels are derived from their Apache-2.0 release. Thanks to them for publishing it. The Kimi K3 figures are theirs. The ones in the table are without speculative decoding.
- Hazy Research, Look Ma, No Bubbles! The Llama 1B figures are theirs.