Close-up of vintage circuit boards with rows of colour-banded resistors

towards the speed of light with megakernels

A model's throughput cannot be higher than the speed of loading the weights used for compute. That speed depends on the hardware used. Here are the numbers for a few leading GPUs and TPUs:

Hardware Memory bandwidth Speed of light, Gemma 4 26B-A4B
NVIDIA H1003.35 TB/s427 tok/s
NVIDIA H2004.8 TB/s612 tok/s
NVIDIA B2008.0 TB/s1,020 tok/s
Google TPU v5e, 8 chips6.55 TB/s835 tok/s
Google TPU v6e, 8 chips13.1 TB/s1,671 tok/s

Speed of light is memory bandwidth divided by the bytes read per token: 7.8 GB for one user with a 1K prompt and BF16 weights.

This is the theoretical limit for decode speed. Inference engineers call it the speed of light and chase it by hiding the compute cost behind the weight loads. It doesn't always work, as it depends on the model's architecture and the hardware it's served on.

Popular serving stacks like vLLM or SGLang are built to be easily extensible, which makes it harder for them to reach this speed. A megakernel closes this gap with a single kernel running the whole decode step, keeping the memory busy the entire time. We worked on this for two Gemma models, Gemma 4 31B (dense) and Gemma 4 26B-A4B (MoE, 128 experts).

All numbers are greedy decoding with no speculative decoding, for one user and a 1K prompt. The megakernel runs on a TPU v5e-8 with BF16 weights. The baseline is vLLM 0.30 on one H100, with FP8 weights for the MoE model and BF16 for the dense model.

When a megakernel pays off

A megakernel implemented correctly will match performance of vLLM/SGLang, and it is clearly faster only in specific scenarios. The time spent building the kernel and the maintenance after is not worth it, if the difference is a few ms. So, after trying this out for different architectures, we decided a megakernel is worth building in cases like if the model has many small kernels (think MoE / small models) or vLLM/SGLang out of box is below half the theoretical max or if we are not building a high batch use case, as decode becomes compute bound on a single GPU in such cases.

The measurements behind each of these:

Verdict Situation Measured against a general stack
WorksMoE model, one user3.7× faster on Gemma 4 26B-A4B, 2.0× on Kimi K31
WorksSmall model, one useralmost 2.5× faster on Llama 1B2
WorksLong prompts4.4× faster at 8K on Gemma 4 26B-A4B
Helps lessMoE model, eight users1.6× faster on Gemma 4 26B-A4B, 1.4× on Kimi K31
Helps lessLarge dense model2.5× faster than vLLM on an H100 on Gemma 4 31B, but only 1.2× faster than vLLM on the same TPU

We tested this by building a megakernel for Gemma 4 dense and MoE to compare the savings.

Megakernels for Gemma 4

We started with the open-source TPU megakernel1 and iterated on it for both Gemma models. On the MoE model, when we assigned whole experts to a single chip, we started seeing straggler chips at the all-reduce after MoE routing: one chip at 2.4 experts per layer against an average of 1.0. We resolved this by splitting each expert across all eight chips instead of assigning it to one, and the shared MLP is split the same way. The other optimization is overlapping the MoE all-reduce with the next weight load. Together with a smaller payload in the all-reduce, these cut the per-step time from 3.46 to 1.56 ms.

vLLM on an H100 runs 1,494 GPU kernels per token on this model and reaches 21% of the speed of light. The megakernel runs one and reaches 77%. Here is where the time goes in each:

Component vLLM, H100, FP8 Megakernel, TPU v5e-8
Experts1.76 ms0.51 ms
Sliding-window attention, 25 layers1.40 ms0.08 ms
Norms, residuals, quantization0.86 ms0.02 ms
Shared MLP0.52 msin the stream
Output head0.48 ms0.003 ms
Full attention, 5 layers0.45 ms0.03 ms
Router0.16 ms0.02 ms
Whole step5.77 ms1.56 ms

The vLLM column is GPU kernel time from the profiler. The megakernel column is how much the step shrinks when that component is switched off, so it does not add up to the whole step.

For the dense 31B model, vLLM already reaches 74% of the speed of light on the same TPU and the megakernel reaches 90%. Here is a graph of tps vs token length for both the models:

Gemma 4 26B-A4B, MoEmegakernel, TPU v5e-8vLLM, H100020040060080002K4K6K8Kprompt tokens587132 Gemma 4 31B, densemegakernel, TPU v5e-8vLLM, H100030609012002K4K6K8Kprompt tokens9334
Tokens per second for one user against prompt length. Higher is better. The dense H100 run stops at 7K.
Gemma 4 26B-A4B, MoEmegakernel, TPU v5e-8vLLM, H100020040060080016401732430163334514343131505270132624313072231278207132batch size Gemma 4 31B, densemegakernel, TPU v5e-8vLLM, H10003060901201953829338391374883758637684377813688136batch size
Tokens per second per user against batch size, 1K prompt. Higher is better.

What's next

A megakernel is worth building when out-of-the-box vLLM or SGLang's decode speeds are far below the speed of light. For the Gemma 4 MoE model, 640 tok/s per user on eight TPU v5e chips is a considerable upgrade over 173 on an H100. A few other ideas we are exploring from here:

Speculative decoding: Fusing speculative decoding into the megakernel should raise the speed even further. Inferact reports over 700 tok/s on Kimi K3 this way1.

More users per GPU: At eight users the gain drops from 3.7× to 1.6×, because attention and vector arithmetic grow with every row. This work needs to be split across chips for higher concurrency.

Quantization and better GPU: With INT8 experts the same kernel already reaches 760 tok/s. TPU v6e has twice the memory bandwidth per chip, which puts the speed of light for this model near 1,670 tok/s.

  1. The megakernel design for TPUs was published by Inferact in A Case for TPU Megakernels. Our kernels are derived from their Apache-2.0 release. Thanks to them for publishing it. The Kimi K3 figures are theirs. The ones in the table are without speculative decoding.
  2. Hazy Research, Look Ma, No Bubbles! The Llama 1B figures are theirs.