Back to articles
Inference & Serving

Helion in vLLM: A Unified Route to Tunable Linear Kernels

3 min read

Introduction

Quantized linear layers are among the most performance-sensitive components in large language model serving. vLLM already connects formats such as FP8, INT8, INT4, and NVFP4 to specialized implementations from CUTLASS, DeepGEMM, FlashInfer, and other libraries. The difficulty is that the best matrix-multiplication strategy can change with the model, batch shape, quantization format, and GPU. A kernel that is broadly effective may not be optimal for a particular deployment.

A PyTorch team describes an alternative in its latest work: integrating Helion into vLLM’s linear backend. Helion is a PyTorch-native, hardware-agnostic kernel DSL based on tile programming. Instead of maintaining many separately specialized kernels, developers expose important choices as tunable parameters and allow an ahead-of-time autotuner to search for an appropriate implementation.

What the integration changes

  • One GEMM description, several strategies. The implementation can cover ordinary GEMM as well as Split-K and Swap-AB. Split-K divides the K dimension across more thread blocks, which can help when M or N is too small to provide enough parallelism. Swap-AB rewrites the multiplication through transposes and can improve utilization for small-M shapes.
  • Algorithm choice becomes part of tuning. Traditional backends often benchmark variants and encode shape-based dispatch heuristics by hand. In the Helion version, algorithmic choices are exposed alongside tile sizes, memory layouts, and scheduling parameters. The autotuner can therefore select both a variant and its detailed configuration for each shape.
  • A focused Hopper evaluation. The current work uses Helion’s Triton backend and concentrates on FP8 dynamic quantization, W8A8 INT8, and Block-FP8. The team also reports initial competitive GEMM results from a CuteDSL backend on NVIDIA Blackwell, while noting that broader support depends on that backend’s maturity.
  • Hybrid dispatch matters. On Hopper, the Helion linear backend combines per-shape tuning with hybrid dispatch. Across the evaluated models it outperformed vLLM’s default CUTLASS and DeepGEMM backends, with more than a 10% end-to-end throughput improvement for some workloads. These results are bounded by the tested hardware, models, and shapes rather than being a universal guarantee.

The costs behind the gains

Autotuning does not remove optimization work; it organizes it. Fine-grained ahead-of-time searches can still take hours. During vLLM startup, CUDA Graph capture can trigger Helion JIT compilation and increase cold-start latency. Caching compiled artifacts can substantially reduce that cost on warm starts. Outside the CUDA Graph capture region, additional CPU-side dispatch and launch work may also offset some of the GPU advantage, making graph-covered execution particularly important.

There is a maintenance issue as well. Shipping many pre-tuned configurations for popular models can make upstream files large and difficult to validate through ordinary unit tests and CI. This reflects a broader performance triangle: higher workload-specific performance usually requires more tuning effort or more configuration maintenance, while usability and maintainability push in the opposite direction.

Why it matters

The integration points toward a different division of labor in inference infrastructure. Kernel authors can describe a broader family of algorithms once, while the tuning system handles part of the hardware- and workload-specific search. Deployment teams may then optimize for their own models without implementing GPU kernels from scratch. The approach is most attractive when shapes are stable, CUDA Graphs are practical, and a deployment can amortize an initial tuning cost. For highly dynamic traffic or latency-sensitive cold starts, established pre-optimized backends may remain the simpler choice. Helion’s practical value will depend on whether its measured GPU gains outweigh compilation, dispatch, and maintenance overhead.

PyTorch Blog

Comments

Checking sign-in status...

Loading comments...

Related articles

CCTest · Blog
Wavefront Decoding Brings Parallelism to Looped Language Models
Inference & Serving
cctest.ai

Wavefront Decoding Brings Parallelism to Looped Language Models

Looped language models gain effective depth by repeatedly applying a shared block, but this also creates a sequential decoding bottleneck. Wavefront Decoding uses intermediate states as drafts and schedules states from different positions and recurrence depths in a diagonal, batched wavefront without additional training.

Read more