Skip to content

Latest commit

 

History

History
1614 lines (1555 loc) · 193 KB

File metadata and controls

1614 lines (1555 loc) · 193 KB

Query: By Architecture

Auto-generated. Do not edit manually. Exact sections contain only pages carrying that exact value; family-only pages and validated source-PR unknowns are intentionally separate. Non-PR sources with an empty architecture list have no validated unknown disposition and are outside the unknown lane.

Turing family-only

Pages with explicit generic Turing evidence but no supported exact SM in that family.

Page Path
perf: refactor fa2 prefill template sources/prs/flashinfer/PR-776.md
Add POD-Attention to FlashInfer sources/prs/flashinfer/PR-858.md
[Quantization] add humming quantization kernel sources/prs/vllm/PR-34556.md
[MoE] Move various experts classes to fused_moe/experts/ sources/prs/vllm/PR-41979.md

Ampere family-only

Pages with explicit generic Ampere evidence but no supported exact SM in that family.

Page Path
[TRTLLM-11092][feat] add support for visual gen FA4 attention backend sources/prs/TensorRT-LLM/PR-11697.md
[https://nvbugs/6108841][fix] add hidden_dim=6144 router GEMM instantiation for GLM-5 sources/prs/TensorRT-LLM/PR-13740.md
support fp16 accmulator for sm89 fp8 mma sources/prs/cutlass/PR-2378.md
bugfix: fix fused-temperature softmax IMA issue sources/prs/flashinfer/PR-1596.md
feat: add GDN Attention sources/prs/flashinfer/PR-2276.md
refactor: reduce hopper's gdn prefill compilation time and fix docstring. sources/prs/flashinfer/PR-2422.md
perf: refactor fa2 prefill template sources/prs/flashinfer/PR-776.md
perf: FlashAttention-3 style MLA PageAttention sources/prs/flashinfer/PR-887.md
Move fa4 from sgl-kernel to jit kernel sources/prs/sglang/PR-17353.md
[Attention Backend] TurboQuant: 2-bit KV cache compression with 4x capacity sources/prs/vllm/PR-38479.md
[Kernel] (1/N) Machete - Hopper Optimized Mixed Precision Linear Kernel sources/prs/vllm/PR-7174.md

Ada family-only

Pages with explicit generic Ada evidence but no supported exact SM in that family.

Page Path
[Quantization] add humming quantization kernel sources/prs/vllm/PR-34556.md

Hopper family-only

Pages with explicit generic Hopper evidence but no supported exact SM in that family.

Page Path
Use swizzling instead of padding sources/prs/DeepGEMM/PR-86.md
[https://nvbugs/5799917][fix] Recover from CUTLASS MoE doActivation perf regression for MXFP4/NVFP4 dtype sources/prs/TensorRT-LLM/PR-11165.md
Improvements for: Groupwise scaling along M for FP8 gemm sources/prs/cutlass/PR-2095.md
Flash MLA support sources/prs/cutlass/PR-2130.md
Flash MLA Support - Step 2 sources/prs/cutlass/PR-2134.md
Add local attention in Hopper FAv3 sources/prs/flash-attention/PR-1233.md
Fix FA3 Varlen Performance regression sources/prs/flash-attention/PR-1361.md
Support hdimQK != hdimV backward sources/prs/flash-attention/PR-1604.md
Add sorting and head swizzle to varlen scheduler sources/prs/flash-attention/PR-1823.md
Improve causal backward determinism perf with SPT schedule sources/prs/flash-attention/PR-1893.md
[Cute] Block sparse support Sm100 sources/prs/flash-attention/PR-1985.md
Add score-mod bwd support sources/prs/flash-attention/PR-2070.md
misc: fix instrument code for mla profiler sources/prs/flashinfer/PR-1014.md
Add blockwise-scaled FP8 GEMM via TRTLLM-Gen. sources/prs/flashinfer/PR-1320.md
bugfix: fix fused-temperature softmax IMA issue sources/prs/flashinfer/PR-1596.md
refactor: pass hopper deepgemm include directory through python sources/prs/flashinfer/PR-2090.md
Selective State Update kernel (mamba) sources/prs/flashinfer/PR-2301.md
Ameyn/gdn bf16 tolerance parallel reduction sources/prs/flashinfer/PR-2610.md
checkpointing_ssu kernel: fused replay + conditional state-write for Mamba2 sources/prs/flashinfer/PR-3324.md
perf: tweak the pipeline design of mla kernel sources/prs/flashinfer/PR-901.md
feat: experimenta support of PDL sources/prs/flashinfer/PR-930.md
bugfix: fix potential issues of FA3 template loading nans for PageAttention sources/prs/flashinfer/PR-945.md
perf: Use 2WG pipeline design for MLA implementation on Hopper sources/prs/flashinfer/PR-952.md
perf: prefetch page indices for mla kernel sources/prs/flashinfer/PR-991.md
3rdparty: upgrade cutlass to 3.9 sources/prs/flashinfer/PR-997.md
[SGLang-Diffusion] Fix custom op fake impl missing eps default for torch.compile sources/prs/sglang/PR-19725.md
Add compile-time 256-bit vector guard for pre-Blackwell sources/prs/sglang/PR-19794.md
[nvidia] Gemma4 nvfp4 fix sources/prs/sglang/PR-22079.md
[KDA] Optimize prefill kernels with diagonal and recompute fuse sources/prs/sglang/PR-24271.md
[codex] Split GEMM implementations by backend sources/prs/tilelang/PR-2153.md
[TIR][IR] Update to use tirx sources/prs/tilelang/PR-2216.md
[Kernel] Integrate CUTLASS MoE kernel with PPLX sources/prs/vllm/PR-18762.md
[Feature] Support Decode Context Parallel (DCP) for MLA sources/prs/vllm/PR-23734.md
[Kernel]Support W4A8 Grouped GEMM on Hopper sources/prs/vllm/PR-29691.md
[Kernel] Apply 256bit LDG/STG To Activation Kernels sources/prs/vllm/PR-33022.md
[Bugfix] Fix quant RMS norm fusion for quantization with TMA-aligned scales sources/prs/vllm/PR-33255.md
[Attention Backend] TurboQuant: 2-bit KV cache compression with 4x capacity sources/prs/vllm/PR-38479.md

Blackwell family-only

Pages with explicit generic Blackwell evidence but no supported exact SM in that family.

Page Path
NVIDIA CUDA Toolkit 13.3 overview sources/docs/nvidia-cuda-13.md
[None][perf] Add more optimization options for MOE CuteDSL finalized kernel sources/prs/TensorRT-LLM/PR-10042.md
[TRTLLM-9992][perf] Enable PDL for CuteDSL kernels and overlap MoeOutputMemset sources/prs/TensorRT-LLM/PR-10043.md
[None][feat] CuteDSL MOE FC1 Enhancement sources/prs/TensorRT-LLM/PR-10088.md
[None][fix] impl fused triton kernel for e8m0 resmooth to reduce memory footprint sources/prs/TensorRT-LLM/PR-10327.md
[None] [feat] Add test script and raster M for gather fc1 kernel sources/prs/TensorRT-LLM/PR-10429.md
[TRTLLM-9831][perf] Use TMA.RED to improve effective memory bandwidth sources/prs/TensorRT-LLM/PR-10987.md
[TRTLLM-10407][perf] Enable CuteDSL indexer_top_k in model sources/prs/TensorRT-LLM/PR-12236.md
[None][feat] Add PDL support to CuTE DSL top-k kernels sources/prs/TensorRT-LLM/PR-12506.md
fix gqa issue for blackwell fmha.py sources/prs/cutlass/PR-2599.md
Add tutorial fp16_gemm_1 sources/prs/cutlass/PR-2750.md
Update blackwell tutorial to be compatible with 4.5-dev version sources/prs/cutlass/PR-3130.md
feat: ragged tensor padding kernel for blackwell kernel alignment sources/prs/flashinfer/PR-1025.md
[nvidia] Add Blackwell FMHA decode kernel from TRT-LLM sources/prs/flashinfer/PR-1051.md
bugfix: temporally disable split-kv in blackwell mla sources/prs/flashinfer/PR-1055.md
feat: masked layout fp4 gemm using cute-dsl sources/prs/flashinfer/PR-1331.md
bugfix: fix persistent attention kernel correctness on blackwell sources/prs/flashinfer/PR-1559.md
bugfix: fix fused-temperature softmax IMA issue sources/prs/flashinfer/PR-1596.md
bugfix: fix the register overflow issue for topk renorm kernels on blackwell sources/prs/flashinfer/PR-1597.md
perf: Port the separate reduce kernel mode from trtllm. sources/prs/flashinfer/PR-1685.md
feat: Add TRTLLM-Gen Skip-Softmax kernels for prefill and decode sources/prs/flashinfer/PR-2477.md
Ameyn/gdn bf16 tolerance parallel reduction sources/prs/flashinfer/PR-2610.md
[feat] trtllm-gen mxfp8 gemm sources/prs/flashinfer/PR-2653.md
feat: enable glm5 router gemm sources/prs/flashinfer/PR-3185.md
Add dynamic tokens-per-page TRTLLM-GEN GQA kernels sources/prs/flashinfer/PR-3259.md
SM-constraint-GEMM by triton persistent kernel sources/prs/flashinfer/PR-982.md
3rdparty: upgrade cutlass to 3.9 sources/prs/flashinfer/PR-997.md
Optimize nvfp4 block scaled gemm kernel when M is small. sources/prs/sglang/PR-10101.md
[SGLang-Diffusion] Fix custom op fake impl missing eps default for torch.compile sources/prs/sglang/PR-19725.md
Add compile-time 256-bit vector guard for pre-Blackwell sources/prs/sglang/PR-19794.md
Support fp8 gemm for blackwell sources/prs/sglang/PR-4558.md
Blackwell Cutlass MLA kernel sources/prs/sglang/PR-5142.md
[Fix][Ready]Fix register spilling in cutlass nvfp4 gemm kernel on Blackwell sources/prs/sglang/PR-8127.md
Update CUTLASS 4.2 & Enable K-Major Scale Factor for SM90 FP8 Blockwise Group GEMM sources/prs/sglang/PR-9559.md
[TIR][IR] Update to use tirx sources/prs/tilelang/PR-2216.md
[NVIDIA] Support Cutlass MLA for Blackwell GPUs sources/prs/vllm/PR-16032.md
[Hardware/NVIDIA/Kernel] [Functional Enablement] [1/N] Enable nvidia/DeepSeek-R1-FP4 Model sources/prs/vllm/PR-16362.md
[Bug] Fix DeepGemm for EP low latency case sources/prs/vllm/PR-20833.md
[Perf] Use Triton instead of Torch for DeepGEMM Per Token Group Quant sources/prs/vllm/PR-20841.md
[NVFP4][Perf] Tune NVFP4 input quant kernel for small batch size sources/prs/vllm/PR-30897.md
[Kernel] Apply 256bit LDG/STG To Activation Kernels sources/prs/vllm/PR-33022.md
[Attention Backend] TurboQuant: 2-bit KV cache compression with 4x capacity sources/prs/vllm/PR-38479.md
Use CU_MEMCPY_SRC_ACCESS_ORDER_ANY for batch KV cache swaps sources/prs/vllm/PR-39306.md
Faster per-token fp8 group quant packed kernel for blackwell sources/prs/vllm/PR-41326.md
[Kernel] (1/N) Machete - Hopper Optimized Mixed Precision Linear Kernel sources/prs/vllm/PR-7174.md

Architecture unknown

Source-PR pages whose validated evidence review found no family-level or exact architecture signal.

Page Path
Solving bank conflict via padding and TMA 3D store sources/prs/DeepGEMM/PR-78.md
Use 1D TMA store instead of 3D sources/prs/DeepGEMM/PR-83.md
Support TMA multicast on B with m_grouped_gemm_contiguous. sources/prs/DeepGEMM/PR-88.md
[https://nvbugs/5669671][fix] Support GuidedDecoder with sharded logits (pick #10698) sources/prs/TensorRT-LLM/PR-10742.md
[None][feat] fuse shared to sparse experts in TRT-LLM Gen MoE sources/prs/TensorRT-LLM/PR-11143.md
[None][feat] TRT-LLM Gen MoE finalize kernel optimization sources/prs/TensorRT-LLM/PR-11501.md
[https://nvbugs/5799917][fix] Recover from CUTLASS MoE doActivation perf regression for MXFP4/NVFP4 dtype sources/prs/TensorRT-LLM/PR-11733.md
[TRTLLM-10421][perf] Add fused cat+fp8_quantize CUDA kernel for DSA indexer sources/prs/TensorRT-LLM/PR-11899.md
[TRTLLM-11540][feat] Add EAGLE3 dynamic tree speculative decoding support sources/prs/TensorRT-LLM/PR-12062.md
[None][feat] Add Mamba2 MTP SSM cache CUDA kernel for tree-based speculative decoding sources/prs/TensorRT-LLM/PR-12537.md
[https://nvbugs/5983390][perf] Multiple host perf optimizations for DSA part sources/prs/TensorRT-LLM/PR-12581.md
[None][feat] Update rms_norm + fp4_qaunt kernel supporting more dim sources/prs/TensorRT-LLM/PR-13033.md
[None][perf] Extend customMoeRouting kernel to support Qwen3.5 sources/prs/TensorRT-LLM/PR-13433.md
[None][perf] Optimize DeepSeek-V4 compressor BF16 input sources/prs/TensorRT-LLM/PR-13761.md
[None][perf] Add CUDA q_b norm for DeepSeek V4 sources/prs/TensorRT-LLM/PR-13975.md
[None][fix] support topk autotuner input for expert slot per group larger than 32 sources/prs/TensorRT-LLM/PR-9087.md
[None][feat] TRT-LLM Gen MoE optimize DeepSeek Fp8 activation kernel sources/prs/TensorRT-LLM/PR-9175.md
[None][feat] Fused kernels (qknormrope + moe routing) and two-model MTP support for glm4moe sources/prs/TensorRT-LLM/PR-9852.md
[None][feat] Port fp4 quantization kernel optimization from FlashInfer sources/prs/TensorRT-LLM/PR-9854.md
Fix the vectorized loading of BlockLoad sources/prs/cccl/PR-3517.md
Split Optimize Warp Reduce PR - CUB part sources/prs/cccl/PR-4716.md
CUB - Add internal integer utils and tests (Split WarpReduce PR) sources/prs/cccl/PR-5314.md
Add dynamic CUB dispatch for segmented_sort sources/prs/cccl/PR-6069.md
[CUB] Use BlockLoadToShared in DeviceMerge sources/prs/cccl/PR-6077.md
Fix debug section around line 390 of dispatch_topk sources/prs/cccl/PR-6152.md
Split fixed-size segmented reduce dispatch header sources/prs/cccl/PR-6597.md
Implement new tuning API arch dispatching sources/prs/cccl/PR-7093.md
Two-phase reduction for fixed size segmented reduction for very large segment sizes sources/prs/cccl/PR-7114.md
Optimize non fixed size segmented reduce for small segments using max_segment_size sources/prs/cccl/PR-7718.md
Add env SegmentedReduce (non fixed-size overloads) sources/prs/cccl/PR-7795.md
Forward policy hub from dispatch_streaming_arg_reduce_t to reduce::dispatch sources/prs/cccl/PR-7805.md
Implement the new tuning API for detail::reduce::dispatch_streaming_arg_reduce_t sources/prs/cccl/PR-7807.md
Use the new tuning API internally for detail::transform::dispatch sources/prs/cccl/PR-7810.md
[Backport branch/3.3.x] Forward policy hub from dispatch_streaming_arg_reduce_t to reduce::dispatch sources/prs/cccl/PR-7814.md
Implement the new tuning API for DispatchSegmentedRadixSort sources/prs/cccl/PR-7844.md
[cuda.compute]: Fix faulty pointer arithmetic calculation in CUB dispatch sources/prs/cccl/PR-7940.md
Use the new tuning API for detail::radix_sort::dispatch sources/prs/cccl/PR-7949.md
Adds support for non-fundamental types via decomposer to DeviceTopK sources/prs/cccl/PR-8040.md
Optimized Device-to-Device Tensor Copy (cudax) - Transpose Case sources/prs/cccl/PR-8125.md
[STF] Move unstable_unique from STF to generic cudax utility sources/prs/cccl/PR-8190.md
[cub]: implement utilities for policy selection sources/prs/cccl/PR-8355.md
[[CUB] Replace `Shuffle(Up Down
Replace detail::scan::dispatch by CUB's public API sources/prs/cccl/PR-8495.md
Replace detail::segmented_reduce::dispatch by the public API sources/prs/cccl/PR-8695.md
Use the new tuning API internally for detail::topk::dispatch sources/prs/cccl/PR-8742.md
Use the new tuning API internally for detail::reduce_by_key::dispatch sources/prs/cccl/PR-8756.md
Use the new tuning API internally for detail::reduce[_nd]::dispatch[_nd] sources/prs/cccl/PR-8826.md
Fix Warpspeed scan shifted output store sources/prs/cccl/PR-8839.md
[cub] Simplify arch dispatch sources/prs/cccl/PR-8861.md
[STF] Add per-handle exec_place stream resources sources/prs/cccl/PR-8905.md
[Use the new tuning API internally for `detail::select three_way_partition::dispatchandDevicePartition`](../sources/prs/cccl/PR-8925.md)
Use the new tuning API internally for detail::segmented_radix_sort::dispatch sources/prs/cccl/PR-8927.md
[libcu++] Always suppress C++ extensions warnings in prologue sources/prs/cccl/PR-9019.md
Vectorize contiguous iterators in cub::BlockLoad/Store sources/prs/cccl/PR-9056.md
Fix epilogue::thread::Convert cannot be used with DefaultEpilogue sources/prs/cutlass/PR-2333.md
[cute] Add constexpr specifier to make_tiled_copy sources/prs/cutlass/PR-2875.md
[Cute-DSL] Add option for issue_clc_query without multicast sources/prs/cutlass/PR-3021.md
feat: update decode attention APIs sources/prs/flashinfer/PR-1007.md
feat: Softmax free sampling sources/prs/flashinfer/PR-1035.md
fix: top_k_mask_logits hangs on -inf inputs sources/prs/flashinfer/PR-1050.md
Fix KV chunking for POD. sources/prs/flashinfer/PR-1054.md
Parameterize prefix mask call (needed by POD-Attention) sources/prs/flashinfer/PR-1059.md
bugfix: Fix test and output shape of fp4 quantize sources/prs/flashinfer/PR-1114.md
bugfix: softmax NaN results caused by large -inf masks sources/prs/flashinfer/PR-1178.md
update trtllm-gen decode attention kernel launcher sources/prs/flashinfer/PR-1189.md
[fix] fix BatchAttention CTA_TILE_KV mask issue sources/prs/flashinfer/PR-1206.md
Fix the issue with auxillary kernel launch and grid dim calculation sources/prs/flashinfer/PR-1208.md
Enable cudnn decode and add tests for the cudnn decode kernel sources/prs/flashinfer/PR-1221.md
feat: add trtllm-gen mla cubin sources/prs/flashinfer/PR-1222.md
bugfix: support uint8_t for vec_t class template sources/prs/flashinfer/PR-1234.md
Patch fp8 cubin availability sources/prs/flashinfer/PR-1240.md
Add trtllm-gen attention mha kernel with FP8 Q/K/V and FP8 output sources/prs/flashinfer/PR-1242.md
Bug fix: fix duplicate launch in POD sources/prs/flashinfer/PR-1267.md
Bug fix: guard fp8 e8m0 and e2m1 compile sources/prs/flashinfer/PR-1287.md
[fix] fix integer overflow in FA2 customized_mask & add buffer overflow warning. sources/prs/flashinfer/PR-1290.md
refactor: Improved metainfo for trtllm-gen fmha sources/prs/flashinfer/PR-1292.md
[Feature] SM level profiler sources/prs/flashinfer/PR-1305.md
Fix the bug of the kernel-selection heuristic in trtllm-gen sources/prs/flashinfer/PR-1307.md
feat: Add k_scale and v_scale to persistent attention sources/prs/flashinfer/PR-1322.md
feat: Support logits_soft_cap for Persistent attn; fix kv split limit sources/prs/flashinfer/PR-1324.md
feat: Fused rope fp8 quantize kernel for MLA sources/prs/flashinfer/PR-1339.md
feature: add fp4 mm using trtllm backend sources/prs/flashinfer/PR-1355.md
support trtllm-gen prefill fp4 output sources/prs/flashinfer/PR-1360.md
bugfix: fixed cutlass fused moe usage of FP4QuantizationSFLayout::SWIZZLED sources/prs/flashinfer/PR-1371.md
bugfix: Add guard for fp4/fp8 related include headers sources/prs/flashinfer/PR-1376.md
Add alignment in MxFP8Quantization sources/prs/flashinfer/PR-1445.md
Add python API for masked grouped gemm sources/prs/flashinfer/PR-1481.md
perf: add fast path to TopPRenormProbKernel for top_p >= 1.0, significantly boosting SGLang workloads sources/prs/flashinfer/PR-1483.md
fix: update cutedsl masked moe gemm sources/prs/flashinfer/PR-1488.md
feat: Support fp8 qkv, fp16/bf16 out MHA for trtllm-gen. sources/prs/flashinfer/PR-1490.md
feat: scaling at fp4 gemm epilogue sources/prs/flashinfer/PR-1498.md
fix: Replace cub Max/Min with cuda::maximum/minimum for cuda 13 compatibility sources/prs/flashinfer/PR-1500.md
bugfix: Fix stream handling in cutedsl gemm sources/prs/flashinfer/PR-1509.md
backend: Refactor trtllm-gen fmha metainfo loading sources/prs/flashinfer/PR-1518.md
Add GeGLU support to trtllm-gen NVFP4 Fused MoE Kernel sources/prs/flashinfer/PR-1525.md
bugfix: Fix compile error for undefined swizzle enum. sources/prs/flashinfer/PR-1530.md
bugfix: Fix Persistent kernel precision for masked output sources/prs/flashinfer/PR-1533.md
perf: replace cudaGetDeviceProperties with cudaDeviceGetAttribute sources/prs/flashinfer/PR-1547.md
Backend: downgrade trtllm-gen kernel to cuda-12 sources/prs/flashinfer/PR-1567.md
bugfix: fix cuda version guard macros sources/prs/flashinfer/PR-1571.md
update trtllm-gen fp4 autotuner and routing sources/prs/flashinfer/PR-1573.md
bugfix: Fix arg passing to TORCH_CHECK and TORCH_WARN macros sources/prs/flashinfer/PR-1582.md
bugfix: fix fp4 quantization with 8x4 scale factor layout sources/prs/flashinfer/PR-1611.md
bugfix: fix merge_attention_state in BatchAttention w/ gqa-group-size in Qwen family sources/prs/flashinfer/PR-1614.md
feat: Add variant.OutputTransform() to decode kernels sources/prs/flashinfer/PR-1670.md
feat: Batch-size invariant FA2 Prefill & Decode sources/prs/flashinfer/PR-1675.md
Support output signals for overlapping for cutedsl gemm sources/prs/flashinfer/PR-1677.md
Update TGV GEMM default kernel and TGV code cleanup. sources/prs/flashinfer/PR-1682.md
Support Kimi-K2 for TRT: templatize number of experts sources/prs/flashinfer/PR-1696.md
Fix DeepSeek quality for TRTLLM fused MoE routing sources/prs/flashinfer/PR-1723.md
Bugfix: Fix data hazard in persistent reduce sources/prs/flashinfer/PR-1826.md
[Quantization] Add per-expert global scaling factor for fp4 batched quantize sources/prs/flashinfer/PR-1835.md
Bugfix: fix o_strides in persistent kernel sources/prs/flashinfer/PR-1865.md
Add layernorm op for inputs of mixed dtype sources/prs/flashinfer/PR-1926.md
Update trtllm-gen fused moe routing kernel and add more kernels sources/prs/flashinfer/PR-1955.md
fix: correct PDL parameter handling in RopeQuantize kernel sources/prs/flashinfer/PR-1982.md
[feat] Refactor trtllmgen MOE and add Bf16 trtllmgen moe sources/prs/flashinfer/PR-2014.md
perf: Speed up fp4 quantization for small batch with swizzling for cutlass MoE sources/prs/flashinfer/PR-2025.md
perf: improve sampling/mask/softmax performance (part 1/2) sources/prs/flashinfer/PR-2044.md
[BUG] Fix trtllm-gen fp4 moe renormalize routing sources/prs/flashinfer/PR-2049.md
Add support for topkPacked input in block-level renormalize sources/prs/flashinfer/PR-2051.md
Fix: several bugs/issues with trtllm-gen attention kernels. sources/prs/flashinfer/PR-2062.md
perf: TRT-LLM MoE Block-FP8 activation optimization sources/prs/flashinfer/PR-2063.md
perf: TRT-LLM Gen finalize kernel optimization sources/prs/flashinfer/PR-2092.md
feat: support more head dim in RoPE kernel sources/prs/flashinfer/PR-2109.md
feat: MxInt4 x Bf16 TRT-LLM Gen MoE support sources/prs/flashinfer/PR-2159.md
[feat] Integrate SGLang concat_mla_k kernel into flashinfer sources/prs/flashinfer/PR-2237.md
fix: support int64 IdType for RoPE part argument in rope_quantize_fp8_append_paged_kv_cache sources/prs/flashinfer/PR-2255.md
[WIP] Refactor: simplify torch -> cute-dsl boilerplate and enable tvm-ffi for cute-dsl kernels sources/prs/flashinfer/PR-2279.md
feat: Support Fused MoE non gated Relu2 NVFP4 & FP8 and support Nemotron sources/prs/flashinfer/PR-2304.md
Fix: FilteredTopKUnifiedKernel read value out of length sources/prs/flashinfer/PR-2308.md
bugfix: fix multi-cta top-k implementation when k value is different for different row sources/prs/flashinfer/PR-2325.md
fix: guard batchWarpReduceSum with ENABLE_FP8 to fix compilation without FP8 sources/prs/flashinfer/PR-2328.md
bugfix: hotfix of PR 2366 (mamba kernel) sources/prs/flashinfer/PR-2378.md
fix: ensure each CTA processes full numHeadsQPerKv for trtllm decode kernel sources/prs/flashinfer/PR-2380.md
fix: In-place Residual Update for add_rmsnorm_fp4quant sources/prs/flashinfer/PR-2385.md
feat: Add output_both_sf_layouts option to add_rmsnorm_fp4quant API sources/prs/flashinfer/PR-2395.md
fix: Sampling: CUDA Graph fix sources/prs/flashinfer/PR-2432.md
fix: fix illegal memory access for NaN input in sampling kernels sources/prs/flashinfer/PR-2456.md
feat: Support Fused MoE non gated Relu2 NVFP4 & FP8 and support Nemotron, fixed sources/prs/flashinfer/PR-2462.md
Ameyn/gdn decode cutedsl kernel sources/prs/flashinfer/PR-2498.md
perf: cache cudaGetDeviceProperties in gdn_prefill to avoid per-call overhead sources/prs/flashinfer/PR-2509.md
Feat/gdn decode pooled sources/prs/flashinfer/PR-2521.md
[Bug] Fix spark unit test failures for test_add_rmsnorm_fp4_quant_cute_dsl sources/prs/flashinfer/PR-2573.md
perf(gdn): optimize MTP kernel with ILP rows and SMEM v caching sources/prs/flashinfer/PR-2618.md
fix: cute dsl nvfp4 moe routing index error sources/prs/flashinfer/PR-2629.md
[fp8_blockwise]Fix int32 overflow in TRTLLM fused MoE activation kernel sources/prs/flashinfer/PR-2642.md
Add varlen and speculative decoding support to selective state update sources/prs/flashinfer/PR-2700.md
feat: Support padding tokens with seqlen=0 for rope+quant+kv cache update fusion kernel sources/prs/flashinfer/PR-2792.md
CuteDSL MoE fix redundant output buffer zeroing sources/prs/flashinfer/PR-2811.md
feat: add pdl support for cute dsl mla decode kernel support sources/prs/flashinfer/PR-2901.md
[Fmha] support nvfp4 output keepsMmaAb generation kernels sources/prs/flashinfer/PR-2988.md
fix: use sym_int64 for strides in rmsnorm CuTe DSL kernels to prevent int32 overflow sources/prs/flashinfer/PR-3007.md
Prevent MoE autotuner buffer overflow on large token buckets sources/prs/flashinfer/PR-3025.md
perf: Add no-bias path for tinygemm_bf16 sources/prs/flashinfer/PR-3151.md
cute-dsl fmha prefill (cubin integration): remove front-padding, add attention_sink, and pdl support sources/prs/flashinfer/PR-3181.md
fix(cute_dsl/moe): unbias autotuner profiling for tile_size enumeration sources/prs/flashinfer/PR-3252.md
feat(cute_dsl/moe): add moe_output_memset_inplace dense memset wrapper sources/prs/flashinfer/PR-3328.md
perf: fix the iteration bound of SWA in FA2 prefill template sources/prs/flashinfer/PR-714.md
bugfix: FusedAddRMSNorm kernels might require more than 48KB shared memory when d is large. sources/prs/flashinfer/PR-718.md
Align KV chunk size binary search with actual KV chunk splitting. sources/prs/flashinfer/PR-728.md
Change apply_rope_with_cos_sin_cache to accept cos_sin_cache sources/prs/flashinfer/PR-754.md
bugfix: Ensure Loop Termination by Enforcing IEEE-754 Compliance in Sampling Kernels sources/prs/flashinfer/PR-774.md
bugfix: drop CTA_TILE_Q=32 sources/prs/flashinfer/PR-785.md
bugfix: MLA decode should multiply sm_scale by math::log2e sources/prs/flashinfer/PR-787.md
fix rope logic in mla decoding sources/prs/flashinfer/PR-793.md
feat: support f32 attention output in FA2 template sources/prs/flashinfer/PR-799.md
bugfix: mla page-attention kernel for different page sizes sources/prs/flashinfer/PR-810.md
bugfix: fix the behavior of MLA kernel when kv-length is 0 sources/prs/flashinfer/PR-868.md
feat - support mla kvcache store sources/prs/flashinfer/PR-888.md
perf: dual pivot top-p/top-k renorm sources/prs/flashinfer/PR-974.md
Triton rms_norm kernels sources/prs/flashinfer/PR-983.md
[CUDA] Only use vec128 if CUDA version is newer than 12.8 sources/prs/pytorch/PR-150705.md
Fix correction bias undefined behavior for nvfp4 models sources/prs/sglang/PR-10426.md
[sgl-kernel] Optimize concat_mla_k kernel sources/prs/sglang/PR-10543.md
Use trtllm_mla decode kernel for draft extend in speculative decoding sources/prs/sglang/PR-11664.md
[Fix] concat_mla_absorb_q_kernel fails for long inputs sources/prs/sglang/PR-12453.md
[NVIDIA] Fix CUDA arch requirement in nvfp4 cast sources/prs/sglang/PR-12581.md
Support moe topk sigmoid kernel sources/prs/sglang/PR-13049.md
[kernel][moe] add moe topk fast sources/prs/sglang/PR-13969.md
[LoRA][III] Add LoRA support for MoE layers and enable TP sources/prs/sglang/PR-14105.md
Add new moe wna16 marlin gemm sources/prs/sglang/PR-14122.md
Opt moe align block size kernel sources/prs/sglang/PR-14133.md
[diffusion] kernel fusion: gated residual layernorm scale shift and layernorm scale shift kernel fusion for Qwen-Image, WAN and HunyuanVideo sources/prs/sglang/PR-14717.md
[sgl-kernel][1/2] Fused qk_norm_rope for GLM4.6 sources/prs/sglang/PR-15141.md
Fix warp illegal instruction in kimi k2 thinking PCG sources/prs/sglang/PR-15306.md
Optimize FP8 MLA KV cache writes with Triton kernel sources/prs/sglang/PR-15522.md
MoE: Skip SiLU/GELU activation for masked experts sources/prs/sglang/PR-15539.md
[JIT kernel] Apply jit per_tensor_quant_fp8 kernel sources/prs/sglang/PR-15836.md
[diffusion] model: support TurboWan2.1-T2V-1.3B/14B SLA sources/prs/sglang/PR-15888.md
optimize get_topk_ragged by fusing get k and k_scale triton kernel sources/prs/sglang/PR-16043.md
[Feature] add aligned_vector type for JIT kernel sources/prs/sglang/PR-16162.md
Kernel: optimize decoding metadata in NSA multi-spec backend with fused kernels sources/prs/sglang/PR-17554.md
Feature/support longcat flash lite sources/prs/sglang/PR-17838.md
[Move sgl-kernel Kernel to JIT] Add JIT concat MLA kernels sources/prs/sglang/PR-17889.md
[Hicache & JIT_kernel] Support page first layout & mla jit kernel sources/prs/sglang/PR-18311.md
Tilelang sparse decode fwd for dsv32 mi355 sources/prs/sglang/PR-18488.md
[jit_kernel] Add fused_qknorm_rope JIT kernel sources/prs/sglang/PR-19059.md
[DeepSeek-V3.2][JIT-kernel] Support nsa fuse store indexer k cache sources/prs/sglang/PR-19148.md
[diffusion][llm] macOS support sources/prs/sglang/PR-19549.md
[AMD] Tilelang sparse fwd for dsv32 mi355/mi300 sources/prs/sglang/PR-19945.md
[diffusion] fix bug of copy_if sources/prs/sglang/PR-20094.md
[Kernel] Fuse temperature + softmax in sampling for decode speedup sources/prs/sglang/PR-20501.md
[Diffusion] Add qknorm rope fuse kernel sources/prs/sglang/PR-21440.md
[AMD] Enable FP8 KV cache and FP8 attention kernel for NSA on MI300/MI355 with TileLang backend sources/prs/sglang/PR-21511.md
[XPU] Enable qwen3.5 on XPU sources/prs/sglang/PR-21668.md
[VLM] Optimize Gemma4 VLM with PCG and fuse RMSNorm + residual add + scalar sources/prs/sglang/PR-24048.md
Amd/deepseek v4 rebase main 0509 sources/prs/sglang/PR-24933.md
amd/deepseek_v4 27/N [fix] Reduce Triton autotune configs for faster first-time server launch sources/prs/sglang/PR-25554.md
fix (jit kernel): elementwise activation C++ error sources/prs/sglang/PR-25695.md
linear support deepgemm sources/prs/sglang/PR-4199.md
Accelerate FP8 CUDA Kernel by 20-28% sources/prs/sglang/PR-4215.md
fix per_token_group_quant_fp8 illegal memory when num_groups % 16 != 0 sources/prs/sglang/PR-4231.md
Add deepseek style fused moe group gate selection kernel sources/prs/sglang/PR-4530.md
Feat/support encoder model (like bert) sources/prs/sglang/PR-4887.md
fix: solve cu118 issue for cutlass mla sources/prs/sglang/PR-5331.md
[perf] experimental enhance fp8 per-tensor quant sources/prs/sglang/PR-5370.md
Fuse MLA set kv cache kernel sources/prs/sglang/PR-5748.md
Cutlass MLA: Disable split kv due to https://github.com/NVIDIA/cutlass/issues/2274 sources/prs/sglang/PR-6101.md
Upgrade CUTLASS 4.0 sources/prs/sglang/PR-6336.md
reduce torch.zeros overhead in moe align block size kernel sources/prs/sglang/PR-6369.md
Fix bug of deepseek-v3 under DP+EP mode with large batchsize/seqlen sources/prs/sglang/PR-6449.md
Refine pre_reorder_triton_kernel slightly to improve performance sources/prs/sglang/PR-6627.md
[EP] Add cuda kernel for moe_ep_pre_reorder sources/prs/sglang/PR-6699.md
Support token-level quantization for EP MoE sources/prs/sglang/PR-6782.md
[EP] Add cuda kernel for moe_ep_post_reorder sources/prs/sglang/PR-6837.md
Fix AWQ Dequant and Weight Loading of deepseek v2 sources/prs/sglang/PR-6842.md
[sgl-kernel] Add cuda kernel for moe_ep_silu_and_mul sources/prs/sglang/PR-6919.md
Fuse routed scaling factor in deepseek sources/prs/sglang/PR-6970.md
Tiny fix cutlass_mla_get_workspace_size stub incorrect signature sources/prs/sglang/PR-7057.md
Support new DeepGEMM sources/prs/sglang/PR-7172.md
Fuse sorted_token_ids padding to moe_align_block_size kernel sources/prs/sglang/PR-7437.md
fix: fix apply_shuffle_mul_sum sources/prs/sglang/PR-7444.md
[kernel] opt moe align block kernel by block/warp scan algorithm sources/prs/sglang/PR-7884.md
[sgl-kernel] Opt per_token_quant_fp8 with warp reduce sources/prs/sglang/PR-8130.md
Optimize moe_sum_reduce_kernel sources/prs/sglang/PR-9477.md
Add swizzle layout detection and automatic merging for layout conflicts sources/prs/tilelang/PR-1736.md
[Feature] Support TMA store in T.tma_copy() sources/prs/tilelang/PR-1981.md
[CUDA] Support int4 T.gemm sources/prs/tilelang/PR-2063.md
[CUDA] Improve int4 GEMM lowering and packed codegen support sources/prs/tilelang/PR-2073.md
[Python] Drop Python 3.9 support sources/prs/tilelang/PR-2218.md
[ROCm] Faster Custom Paged Attention kernels sources/prs/vllm/PR-12348.md
[Attention] MLA decode optimizations sources/prs/vllm/PR-12528.md
[Kernel] port sgl moe_align_block_size kernels sources/prs/vllm/PR-12574.md
Expert Parallelism (EP) Support for DeepSeek Models sources/prs/vllm/PR-12583.md
[Kernel][Quantization] Integrate block-quantized CUTLASS kernels for DeepSeekV3 sources/prs/vllm/PR-12587.md
[Attention] MLA with chunked prefill sources/prs/vllm/PR-12639.md
[Kernel] Make rotary_embedding ops more flexible with input shape sources/prs/vllm/PR-12777.md
Optimize moe_align_block_size for deepseek_v3 sources/prs/vllm/PR-12850.md
[core] Perf improvement for DSv3 on AMD GPUs sources/prs/vllm/PR-13718.md
[Kernel] optimize performance of gptq marlin kernel when n is small sources/prs/vllm/PR-14138.md
dynamic distpatch of fp8 kernels sources/prs/vllm/PR-14245.md
[Kernel] allow non-contiguous input for marlin kernel sources/prs/vllm/PR-14658.md
[Kernel] Fix conflicting macro names for gguf kernels sources/prs/vllm/PR-15456.md
Use Cache Hinting for fused_moe kernel sources/prs/vllm/PR-15511.md
[ROCM][KERNEL] Paged attention for V1 sources/prs/vllm/PR-15720.md
[Bugfix] fix use_atomic_add support of marlin kernel when using v1 engine sources/prs/vllm/PR-15946.md
Modularize fused experts and integrate PPLX kernels sources/prs/vllm/PR-15956.md
[ROCM] Add gfx950 to the custom attention archs sources/prs/vllm/PR-16034.md
[Kernel] support merge_attn_states CUDA kernel, 3x speedup sources/prs/vllm/PR-16173.md
[Kernel] Support W8A8 channel-wise weights and per-token activations in triton fused_moe_kernel sources/prs/vllm/PR-16366.md
Allocate kv_cache with stride order sources/prs/vllm/PR-16605.md
[Kernel] GGUF MoeVec kernel sources/prs/vllm/PR-16780.md
[BugFix] Accuracy fix for llama4 int4 - improperly casted scales sources/prs/vllm/PR-16801.md
[Kernel] Add expert_map support to Cutlass FP8 MOE sources/prs/vllm/PR-16861.md
[ROCm][Kernel][V1] Enable AMD Radeon GPU Custom Paged Attention on v1 sources/prs/vllm/PR-17004.md
Fix numel() downcast in vllm/csrc/moe/moe_align_sum_kernels.cu +2 sources/prs/vllm/PR-17082.md
[ROCm][FP8][Kernel] FP8 quantization fused into Custom Paged Attention sources/prs/vllm/PR-17139.md
[Kernel] fp4 marlin kernel sources/prs/vllm/PR-17687.md
[Kernel] Have rotary embeddings support tensors sources/prs/vllm/PR-18046.md
[Hardware][AMD] integrate aiter chunked prefill into vllm sources/prs/vllm/PR-18596.md
[Kernel] Enable fp8 support for pplx and BatchedTritonExperts. sources/prs/vllm/PR-18864.md
[Bugfix] Fix some narrowing conversion warnings sources/prs/vllm/PR-20141.md
[Bugfix] Fix topk_ids indices_type for CUTLASS w8a8 FP8 MoE sources/prs/vllm/PR-20166.md
[Kernel][Bugfix] Fixup some warnings in nvfp4_blockwise_moe when CUDA < 12.8 sources/prs/vllm/PR-20324.md
[Performance] Performance improvements in non-blockwise fp8 CUTLASS MoE sources/prs/vllm/PR-20762.md
[Kernel] DeepGemm MoE : Integrate triton permute / unpermute kernels sources/prs/vllm/PR-20903.md
[perf] Add fused MLA QKV + strided layernorm sources/prs/vllm/PR-21116.md
[Feature][Kernel]FusedMoE LoRA sources/prs/vllm/PR-21229.md
[v1] - Mamba1 Attention Metadata sources/prs/vllm/PR-21249.md
[Bug] Fix Compressed Tensor NVFP4 cutlass_fp4_group_mm illegal memory access sources/prs/vllm/PR-21465.md
[Kernel] Improve machete memory bound perf sources/prs/vllm/PR-21556.md
Fp8 paged attention update sources/prs/vllm/PR-22222.md
[BugFix] Fix triton compile error in kernel_unified_attention_2/3d caused by attention sinks sources/prs/vllm/PR-22368.md
[Fix] enable swap_ab for pplx problem size computation sources/prs/vllm/PR-22991.md
[Kernel] CUTLASS MoE FP8: Integrate cuda moe permute/unpermute sources/prs/vllm/PR-23045.md
Optimize input preparation for FlashInfer [2/N] sources/prs/vllm/PR-23174.md
[Compile] Fix Compile Warning for w4a8_mm_entry.cu sources/prs/vllm/PR-23660.md
[Bugfix][Misc] Fix silu_and_mul_nvfp4_quant issue and extract common utils for nvfp4 kernel source files sources/prs/vllm/PR-23727.md
[Kernel] cuda kernels for upcoming decode context parallel feature sources/prs/vllm/PR-23791.md
[Kernel] Faster pre-processing time for W4A8 sources/prs/vllm/PR-23972.md
[Model] Add LongCat-Flash sources/prs/vllm/PR-23991.md
[Bugfix] Fix accuracy issue for silu_mul + nvfp4 quant fusion kernel sources/prs/vllm/PR-24833.md
[Compile] Fix Compile Warning for Ignoring MIN_BLOCK_PER_SM sources/prs/vllm/PR-25193.md
Fuse RoPE and MLA KV-cache write sources/prs/vllm/PR-25774.md
[Attention] Use sparse prefill kernel for fp8 kv-cache in DeepSeek-v3.2 sources/prs/vllm/PR-27532.md
[Model] Add support for openPangu moe model sources/prs/vllm/PR-28775.md
bugfix: correct attn output with base 2 or e sources/prs/vllm/PR-28840.md
Lora MoE Align Improvements sources/prs/vllm/PR-29257.md
Add unpermute-aware fused MoE path and small-batch fallback sources/prs/vllm/PR-29354.md
[Kernel][MoE] optimize moe_align_block_size sources/prs/vllm/PR-29642.md
Add llmcompressor fp8 kv-cache quant (per-tensor and per-attn_head) sources/prs/vllm/PR-30141.md
gptq marlin quantization support for fused moe with lora sources/prs/vllm/PR-30254.md
[LoRA] Support Quantized Adapters sources/prs/vllm/PR-30286.md
OffloadingConnector: Support kernel_block_size != block_size sources/prs/vllm/PR-30692.md
[Bugfix] [Kernel] Triton attention kernels: mask out V blocks that fall outside sliding window sources/prs/vllm/PR-30887.md
[Bugfix] Fix incorrect tiles creation for mm prefix triton attention sources/prs/vllm/PR-30974.md
[Bugfix][ROCm]Fix Qwen3-Next-80B-A3B-Thinking inference and optimize non-standard block size (544) support under rocm_atten sources/prs/vllm/PR-31380.md
[Perf] Fuse stride preparation for NVFP4 cutlass_moe sources/prs/vllm/PR-31837.md
[PERF] Change GDN Attention State Layout from [N, HV, K, V] to [N, HV, V, K] sources/prs/vllm/PR-33291.md
[Bugfix]fix output Nan/Inf in marlin if dtype=float16 sources/prs/vllm/PR-33972.md
[Custom Ops] Add functional + out variant for scaled_fp4_quant sources/prs/vllm/PR-34389.md
[Bugfix] Fix expert_ids padding values in moe_align_block_size kernel sources/prs/vllm/PR-35161.md
[Attention][Perf] Optimize cp_gather_and_upconvert_fp8_kv_cache - DeepSeek-v3.2 sources/prs/vllm/PR-35290.md
Add 320 dimension size support to MLA sources/prs/vllm/PR-36161.md
[MTP][Sparse MLA] Take advantage of native MTP support in indexer when possible sources/prs/vllm/PR-36982.md
[Bugfix] Fix broken explicit unquantized kv cache dtype support sources/prs/vllm/PR-38922.md
[Perf] Fuse Zero Initializer for FP8 DeepGemm Block Quant Kernel sources/prs/vllm/PR-39547.md
[Bugfix] moe lora align kernel grid sources/prs/vllm/PR-40131.md
[Performance][DSR1]: Fused RoPE+KVCache+q_concat for MLA sources/prs/vllm/PR-40392.md
[DSV4] Add silu clamp limit to shared expert sources/prs/vllm/PR-40950.md
[Bugfix] Add swiglu limits to deepgemm fp8 methods sources/prs/vllm/PR-41986.md
[feat] Add FP8 per-tensor Q scale support to Triton attention backend sources/prs/vllm/PR-42080.md
[Perf] Use 2D-grid to eliminate divmod in W8W8 group quant sources/prs/vllm/PR-42153.md
[Kernel] Pack topk id/weights triton kernel sources/prs/vllm/PR-42527.md
[6/n] Migrate activation kernels, gptq, gguf, non cutlass w8a8 to libtorch stable ABI (continued) sources/prs/vllm/PR-42663.md
[Perf] Padded nvfp4 quant kernel to remove additional copy, 2.4%~5.7% e2e performance improvement sources/prs/vllm/PR-42774.md
add cutedsl dsv4 indexer fp8 kernel sources/prs/vllm/PR-42899.md

Exact architectures

sm100

Page Path
Twelve Attempts at an FP4 Kernel sources/blogs/amandeep-nvfp4-attempts.md
Colfax Article Source Kernels sources/blogs/colfax-article-source-kernels.md
Colfax CUTLASS Tutorial: GEMM Kernels Using Tensor Memory for Blackwell sources/blogs/colfax-cutlass-blackwell.md
Colfax CUTLASS Kernels sources/blogs/colfax-cutlass-kernels.md
DeepGEMM tensor-core kernel library sources/blogs/deepgemm.md
FlashAttention-4 Blog sources/blogs/flash-attention-4.md
FlashMLA upstream README sources/blogs/flashmla.md
Anatomy of a Reward Hack sources/blogs/gpu-mode-reward-hack.md
Writing High-Performance Matrix Multiplication Kernels for Blackwell with JAX Pallas sources/blogs/jax-pallas-blackwell-matmul.md
Modular: Matrix Multiplication on Blackwell, Part 3 sources/blogs/modular-blackwell-matmul.md
NVIDIA Developer Code Samples sources/blogs/nvidia-code-samples.md
NVIDIA Qwen3-Next Architecture Announcement sources/blogs/qwen3-next-architecture.md
NVFP4 GEMV sources/blogs/simon-nvfp4-gemv.md
simveit effective_transpose sources/blogs/simveit-effective-transpose.md
simveit load_and_store sources/blogs/simveit-load-and-store.md
tcgen05 for dummies sources/blogs/tcgen05-tutorial.md
TFLOPS Gap: Why FP4 MoE Kernel Engineering Matters on Blackwell sources/blogs/tflops-gap-fp4-moe.md
Tilus: A Tile-Level GPGPU Programming Language for Low-Precision Computation sources/blogs/tilus-nvidia.md
DeepSeek-V3.2-Exp in vLLM: Fine-Grained Sparse Attention in Action sources/blogs/vllm-deepseek-v3-sparse-attention.md
Blackwell NVFP4 Kernel Hackathon Journey sources/blogs/yue-nvfp4-hackathon.md
FlashInfer MLSys 2026 - Track A: Fused MoE FP8 sources/contests/flashinfer-mlsys26/track-a-fused-moe.md
FlashInfer MLSys 2026 - Track B: Sparse Attention sources/contests/flashinfer-mlsys26/track-b-sparse-attention.md
FlashInfer MLSys 2026 - Track C: Gated Delta Net sources/contests/flashinfer-mlsys26/track-c-gated-delta-net.md
GPU Mode NVFP4 Hackathon - Problem 1: Batched GEMV sources/contests/gpu-mode-nvfp4/problem-1-gemv.md
GPU Mode NVFP4 Hackathon - Problem 2: NVFP4 GEMM sources/contests/gpu-mode-nvfp4/problem-2-gemm.md
GPU Mode NVFP4 Hackathon - Problem 3: Gated Dual GEMM sources/contests/gpu-mode-nvfp4/problem-3-gated-dual-gemm.md
GPU Mode NVFP4 Hackathon - Problem 4: Grouped GEMM sources/contests/gpu-mode-nvfp4/problem-4-grouped-gemm.md
NVIDIA Blackwell Compatibility Guide sources/docs/blackwell-compatibility-guide.md
Microbenchmarking NVIDIA's Blackwell Architecture sources/docs/blackwell-microbenchmarking.md
CUDA Programming Guide: Programmatic Dependent Launch sources/docs/cuda-programming-guide-pdl.md
cuTile Python Documentation sources/docs/cutile-python-dsl.md
CUTLASS Changelog: SM100/Blackwell Entries sources/docs/cutlass-changelog-sm100.md
CUTLASS Blackwell Cluster Launch Control sources/docs/cutlass-clc-documentation.md
CUTLASS CuTe DSL Documentation sources/docs/cutlass-cute-dsl.md
FlashAttention-4: Algorithm and Kernel Pipelining Co-Design for Asymmetric Hardware Scaling sources/docs/flash-attention-4.md
K-Search: LLM Kernel Generation via Co-Evolving Intrinsic World Model sources/docs/k-search-kernel-generation.md
NVIDIA Blackwell Tuning Guide sources/docs/nvidia-blackwell-tuning-guide.md
NVIDIA CUTLASS Blackwell support map sources/docs/nvidia-cutlass-blackwell.md
PTX ISA Fifth-Generation Tensor Core and CLC Reference sources/docs/nvidia-ptx-isa-sm100.md
Triton 3.6 Release Notes — Blackwell Backend Work sources/docs/triton-3.6-blackwell.md
Fix multicast bug and optimize masked GEMM sources/prs/DeepGEMM/PR-193.md
[Public release 26/04] Introducing Mega MoE, FP4 Indexer and other features/fixes sources/prs/DeepGEMM/PR-304.md
Sync nv_dev with upstream #316 (Mega MoE optimizations & benchmarks) sources/prs/DeepGEMM/PR-328.md
[None][feat] sm100 weight-only kernel sources/prs/TensorRT-LLM/PR-10190.md
[TRTLLM-9831][perf] Enable 2CTA with autotune for CuteDSL MoE and Grouped GEMM optimizations sources/prs/TensorRT-LLM/PR-10201.md
[TRTLLM-10276][feat] Integrate cutedsl argmax kernel sources/prs/TensorRT-LLM/PR-10476.md
[None] [feat] Add densegemm backend for MoE sources/prs/TensorRT-LLM/PR-10479.md
[https://nvbugs/5799917][fix] Recover from CUTLASS MoE doActivation perf regression for MXFP4/NVFP4 dtype sources/prs/TensorRT-LLM/PR-11165.md
[None][feat] Optimize super-v3 nvfp4 for better perf sources/prs/TensorRT-LLM/PR-11273.md
[None][feat] Optimize by fuse nvfp4_quant to layernorm_gated for mamba2_mixer sources/prs/TensorRT-LLM/PR-11473.md
[TRTLLM-11092][feat] add support for visual gen FA4 attention backend sources/prs/TensorRT-LLM/PR-11697.md
[TRTLLM-11119][feat] Blackwell SageAttention, Integrate into AttentionOp API sources/prs/TensorRT-LLM/PR-11718.md
[None][feat] Add fused DiT QK Norm + RoPE CUDA kernel for FLUX sources/prs/TensorRT-LLM/PR-11869.md
[TRTLLM-10990][feat] Fuse SwiGLU and quant into shared expert sources/prs/TensorRT-LLM/PR-11897.md
[TRTLLM-10407][feat] Integrate CuTE DSL top-k kernel for Blackwell sources/prs/TensorRT-LLM/PR-11900.md
[TRTLLM-11289][feat] Integrate CuteDSL's bf16 dense GEMMs sources/prs/TensorRT-LLM/PR-12074.md
[None][feat] CuteDSL MOE: Add raster along M/N support for blockscaled contiguous backbone kernel sources/prs/TensorRT-LLM/PR-12079.md
[None][feat] Add DWDP (Distributed Weight Data Parallelism) support for MoE inference sources/prs/TensorRT-LLM/PR-12136.md
[None][feat] Support update weight for nvfp4 sources/prs/TensorRT-LLM/PR-12320.md
[https://nvbugs/5983390][perf] Kernel fusions in _gather_k_cache_for_chunk of Indexer in DSA sources/prs/TensorRT-LLM/PR-12322.md
[TRTLLM-10407][perf] Add cute dsl single pass multi cta cluster topk sources/prs/TensorRT-LLM/PR-12354.md
[None][feat] Temporally-Correlated Heuristic-guided Indexer TopK for Sparse Attention sources/prs/TensorRT-LLM/PR-12385.md
[None][feat] Trtllm-gen FMHA JIT support sources/prs/TensorRT-LLM/PR-12612.md
[None][feat] Add triton paged attention for AutoDeploy sources/prs/TensorRT-LLM/PR-12642.md
[TRTLLM-11585][feat] Add CUTEDSL moe backend for nemotron-h sources/prs/TensorRT-LLM/PR-12884.md
[TRTLLM-11485][feat] Feature rework: Add SageAttention refreshed kernels (attentionOp only) sources/prs/TensorRT-LLM/PR-12937.md
[#12716][feat] Fused cross-head QK Norm + RoPE kernel for WAN sources/prs/TensorRT-LLM/PR-13052.md
[None][feat] Add FP4 residual quantization kernel without channel reo… sources/prs/TensorRT-LLM/PR-13117.md
[TRTLLM-34871][feat] Add cute dsl FP8 paged MQA logits decode kernel sources/prs/TensorRT-LLM/PR-13219.md
[None][feat] Integrate FP4 indexer for DSA on Blackwell sources/prs/TensorRT-LLM/PR-13340.md
[None][perf] Scheme X L2-aware dispatcher and PDL launchers for sparse-attention GVR Top-K sources/prs/TensorRT-LLM/PR-13477.md
[None][feat] Fuse FP8 1x128 quantize + UE8M0 scale pack on SM100 sources/prs/TensorRT-LLM/PR-13628.md
[https://nvbugs/6108841][fix] add hidden_dim=6144 router GEMM instantiation for GLM-5 sources/prs/TensorRT-LLM/PR-13740.md
[None][fix] Plumb swiglu_limit through DeepGEMM and TRTLLMGen FP8 fused MoE sources/prs/TensorRT-LLM/PR-13767.md
[None][fix] Fix fused MHC for DeepSeek-V4-Pro hidden size sources/prs/TensorRT-LLM/PR-13771.md
[None][feat] Indexer topk opt sources/prs/TensorRT-LLM/PR-13811.md
[None][perf] FC2 DenseGEMM autotune: split-K, swap_ab, fine-grained tuning buckets sources/prs/TensorRT-LLM/PR-13833.md
[None][perf] mHC fused_hc kernel optimizations + DS-V4 entry-boundary RMSNorm fold-in sources/prs/TensorRT-LLM/PR-13892.md
[TRTLLM-35237][feat] Add cute dsl FP4 paged MQA logits decode kernel sources/prs/TensorRT-LLM/PR-13929.md
[None][feat] Keep DSv4 o_a_proj as FP8, and port vLLM's fused_inv_rope_fp8_quant sources/prs/TensorRT-LLM/PR-13938.md
[None][feat] Add chunked prefill support for Gemma4 (text + vision multimodal) sources/prs/TensorRT-LLM/PR-14134.md
[None][feat] Update the logic of FMHA JIT path sources/prs/TensorRT-LLM/PR-14291.md
feat: Add w4a8_mxfp4_fp8 quantization recipe. sources/prs/TensorRT-LLM/PR-4867.md
[None][chore] Fix kernel launch param and add TRTLLM MoE backend test sources/prs/TensorRT-LLM/PR-7524.md
[None][fix] Fix and add test for TRTLLM MoE backend sources/prs/TensorRT-LLM/PR-7755.md
[TRTLLM-8535][feat] Support DeepSeek V3.2 with FP8 + BF16 KV cache/NVFP4 + BF16 KV cache sources/prs/TensorRT-LLM/PR-8405.md
[TRTLLM-9685] [feat] Add gather fc1 kernel by cuteDSL sources/prs/TensorRT-LLM/PR-9618.md
[None][feat] Adding torch ext API for FusedAddRMSNormQuant kernel sources/prs/TensorRT-LLM/PR-9905.md
Add b200 tunings for scan.exclusive.sum sources/prs/cccl/PR-3559.md
Fix SM100 histogram tunings sources/prs/cccl/PR-3691.md
Add nondeterministic reduce that uses atomics sources/prs/cccl/PR-4961.md
Combine block_reduce_warp_reduction_nondeterministic.cuh specialization with original deterministic one sources/prs/cccl/PR-5408.md
Integrate decoupled lookahead warpspeed scan sources/prs/cccl/PR-6811.md
Radix-selection based BlockTopK specialization sources/prs/cccl/PR-7384.md
Implement the new tuning API for DeviceRleDispatch sources/prs/cccl/PR-7669.md
Optimized Device-to-Device Tensor Copy (cudax) sources/prs/cccl/PR-7823.md
Avoid passing uninitialized values to scan_op sources/prs/cccl/PR-8184.md
[Port `thrust::min max_element` to CUB](../sources/prs/cccl/PR-8291.md)
Implement the new tuning API for DispatchSelectIf sources/prs/cccl/PR-8311.md
simplify dispatch segmented reduce to use latest dispatch and new tunings API sources/prs/cccl/PR-8332.md
Apply some random warpspeed tunings sources/prs/cccl/PR-8352.md
Vectorize mbarrier initialization in warpspeed scan sources/prs/cccl/PR-8423.md
Blockwise and Groupwise GEMM for Blackwell and Improvements for Hopper sources/prs/cutlass/PR-2139.md
Blockwise Improvement and Programmatic Dependent Launch sources/prs/cutlass/PR-2161.md
[ex77] fix mla split; add fwd lse; add bwd varlen sources/prs/cutlass/PR-2366.md
Example 77 add blackwell fmha bwd for MLA shape sources/prs/cutlass/PR-2466.md
Add Blackwell MLA forward (shape: d=192, dv=128) implementation sources/prs/cutlass/PR-2472.md
fix: examples/cute/tutorial/blackwell/04_mma_tma_2sm_sm100.cu GridDim miscalculated sources/prs/cutlass/PR-2492.md
DistGEMM bug fixes sources/prs/cutlass/PR-2713.md
Support for GEMM-K=0 for Blackwell Grouped GEMMs sources/prs/cutlass/PR-2746.md
Blockscaled Ragged Contiguous Grouped Gemm for MoEs sources/prs/cutlass/PR-2790.md
[Bug Fix]Bypass launch grids for SM120 Kernel with SM90 Mainloop & SM100 TileScheduler sources/prs/cutlass/PR-2865.md
Fix incorrect tensor layout strides in Blackwell MMA tutorial comments sources/prs/cutlass/PR-2921.md
[CuTeDSL] Fix: SM100 block-scale gemm overlapping accumulator sources/prs/cutlass/PR-2995.md
[Cute,Fwd,Sm100] Implement SplitKV sources/prs/flash-attention/PR-1940.md
Blackwell FlashAttention-BWD (v1.0) sources/prs/flash-attention/PR-1945.md
[Cute] Block sparse support Sm100 sources/prs/flash-attention/PR-1985.md
[Cute,Fwd,Sm100] Support paged attention sources/prs/flash-attention/PR-1999.md
[Cute,Sm100,Fwd] use correction warps for epi when not using TMA sources/prs/flash-attention/PR-2014.md
[Cute,Bwd,Sm100] enable deterministic mode for sm100 bwd and fix race conditions sources/prs/flash-attention/PR-2033.md
[Cute,Fwd] Extend score_mod to variable sequence length sources/prs/flash-attention/PR-2043.md
Add score-mod bwd support sources/prs/flash-attention/PR-2070.md
Add blocksparse support for bwd on blackwell sources/prs/flash-attention/PR-2085.md
[Cute,Fwd,Sm100] distributed offset calculation for paged KV sources/prs/flash-attention/PR-2104.md
[Cute,Fwd,Sm100] fp8 e4m3 and e5m2 support sources/prs/flash-attention/PR-2109.md
[Cute,Fwd,Sm100] support irregular qhead / kvhead ratios sources/prs/flash-attention/PR-2186.md
[Ai-assisted] CLC work stealing sources/prs/flash-attention/PR-2218.md
[Fwd,Sm90] Add paged KV attention support (tma and cp.async) sources/prs/flash-attention/PR-2360.md
Feat([FA4][CUTE DSL]) Add head_dim=256 support (forward + backward) sources/prs/flash-attention/PR-2412.md
[Cute,Sm100,Fwd] add MLA 64/512 with topk sparsity for MQA 128 heads sources/prs/flash-attention/PR-2441.md
[hd256] Improve forward kernel with exp2 FMA emulation (3% to 9% performance gain) sources/prs/flash-attention/PR-2488.md
[hd256] Add TMA paged KV support to SM100 2CTA forward kernel sources/prs/flash-attention/PR-2489.md
[FA4][hd256] Backward TMA bulk-store epilogue + LSE/dpsum coalesce sources/prs/flash-attention/PR-2497.md
[nvidia] initial support for blackwell kernels sources/prs/flashinfer/PR-1039.md
bugfix: adding lse output to blackwell fmha kernels sources/prs/flashinfer/PR-1071.md
bugfix: follow user-specified sm_scale for blackwell cutlass fmha sources/prs/flashinfer/PR-1072.md
perf: accelerate blackwell grouped gemm sources/prs/flashinfer/PR-1086.md
bugfix: host-precomuted plan function for blackwell fmha sources/prs/flashinfer/PR-1106.md
Add CUTLASS fused moe kernels from TensorRT-LLM. sources/prs/flashinfer/PR-1113.md
hotfix: fix the blackwell fmha stream sources/prs/flashinfer/PR-1116.md
Add more logging to TRTLLM-GEN debug trace (NFC) sources/prs/flashinfer/PR-1158.md
bugfix: fix blackwell fmha hanging issue for empty kv_len sources/prs/flashinfer/PR-1198.md
feat: trtllm-gen fp8 moe kernels sources/prs/flashinfer/PR-1212.md
Feature/sm100 low latency nvfp4 kernels sources/prs/flashinfer/PR-1214.md
Fix missing hash in the cudnn cubin path sources/prs/flashinfer/PR-1227.md
feat: Add non-causal cudnn prefill kernels sources/prs/flashinfer/PR-1230.md
feat: Support MXFP8 x MXFP4 CUTLASS grouped GEMM sources/prs/flashinfer/PR-1241.md
Reduce the JIT compilation time of gen_gemm_sm100_module sources/prs/flashinfer/PR-1251.md
refactor: refactor trtllm-gen attention kernel integration code sources/prs/flashinfer/PR-1289.md
Update cutlass fp4 moe kernels sources/prs/flashinfer/PR-1294.md
add cutlass backend for mm_fp4 sources/prs/flashinfer/PR-1296.md
Add blockwise-scaled FP8 GEMM via TRTLLM-Gen. sources/prs/flashinfer/PR-1320.md
gpt-oss: Add MXFP8 x MXFP4 CUTLASS MOE for SM100 and BF16 x MXFP4 CUTLASS for SM90 + SwigluBias Activation sources/prs/flashinfer/PR-1396.md
feature: add cutlass as bmm_fp8 backend. sources/prs/flashinfer/PR-1397.md
fix shared memory alignment conflict in sampling.cuh sources/prs/flashinfer/PR-1402.md
Remove getEnvEnablePDL in favor of enable_pdl parameter sources/prs/flashinfer/PR-1446.md
tuner: Trtllm-gen Fp4 MoE Autotunner sources/prs/flashinfer/PR-1475.md
refactor fp4 masked gemm cute-dsl implementation and add manual cache sources/prs/flashinfer/PR-1521.md
feat: Integrate TRTLLM varlen kernel for deepseek R1 prefill sources/prs/flashinfer/PR-1537.md
fix: separate out fp4 lib into sm90 and sm100 versions, add oob checking in fused moe sources/prs/flashinfer/PR-1565.md
feat: initial support for SM103, SM110, SM120, SM121 sources/prs/flashinfer/PR-1608.md
feat: cutlass fp4 gemm bringup for SM120 & SM121 sources/prs/flashinfer/PR-1609.md
bugfix: trtllm-gen fmha sm101 and sm100 compatibility sources/prs/flashinfer/PR-1631.md
TGV GEMM as a BF16 backend alternative to cuBLAS sources/prs/flashinfer/PR-1668.md
bugfix: partially fix tests/test_trtllm_gen_fused_moe.py unit test failure sources/prs/flashinfer/PR-1724.md
TVM: support TVM binding for GroupedGemm sources/prs/flashinfer/PR-1725.md
feat: trtrllm-gen global scaled FP8 GEMMs sources/prs/flashinfer/PR-1829.md
Add head_dim=64 for blackwell cutlass fmha implementation sources/prs/flashinfer/PR-1850.md
Tune kernel compilation parameters for https://github.com/flashinfer-ai/flashinfer/pull/1850 sources/prs/flashinfer/PR-1878.md
feat: Add FP4 TRTLLM-Gen throughput MOE batched gemms sources/prs/flashinfer/PR-1882.md
Feature: Support Relu2 activation in fused MoE sources/prs/flashinfer/PR-1954.md
[DSV3] Optimized Router Gemm sources/prs/flashinfer/PR-2019.md
update trtllm cutlass moe sources/prs/flashinfer/PR-2020.md
[NVIDIA] Thor & Spark Support sources/prs/flashinfer/PR-2028.md
feat: Add flashinfer.rope.rope_quantize_fp8_append_paged_kv_cache (fused RoPE + Q + KV cache, supports MLA/GQA/MHA) sources/prs/flashinfer/PR-2037.md
Rebase FP8 SM100 Cutlass FMHA Attention to main (original PR#1238) sources/prs/flashinfer/PR-2047.md
perf: Optimize helper max/minmax function in sampling.cuh sources/prs/flashinfer/PR-2058.md
feat: BF16 GEMM using CUTLASS backend for SM100 sources/prs/flashinfer/PR-2070.md
feat: add trtllm-gen per-tensor sparseMla kernels. sources/prs/flashinfer/PR-2138.md
feat: further optimize top-k and add fused top-k page construction kernels for DSA sources/prs/flashinfer/PR-2215.md
feat: Fused RMSNorm + FP4 Quantization Kernels in CuTe-DSL sources/prs/flashinfer/PR-2233.md
Remove cudaStreamSynchronize from gemm_groupwise_sm120.cuh for CUDA graph compatibility sources/prs/flashinfer/PR-2244.md
fix: Add global scale support and optional output allocation for RMSNorm+FP4Quant fusion kernels sources/prs/flashinfer/PR-2260.md
[TRTLLM-Gen Fmha] add optimized trtllm-gen decode kernels for high throughput + speculative decoding sources/prs/flashinfer/PR-2265.md
[Perf][Feature] Add SM103-specific schedulers for NVFP4 CUTLASS kernels sources/prs/flashinfer/PR-2303.md
[ML3] Optimized Router Gemm sources/prs/flashinfer/PR-2323.md
[perf] Improve gemm_fp8_nt_groupwise (cutlass backend) by 10-40% for batch sizes <= 32 sources/prs/flashinfer/PR-2327.md
Optimize quantization function in large problem size sources/prs/flashinfer/PR-2343.md
Enable fp16/bf16/f32 support for selective_state_update (mamba) sources/prs/flashinfer/PR-2366.md
feat: [Qwen3-Next] Add Cute DSL GDN decode kernel and tests sources/prs/flashinfer/PR-2370.md
A Blackwell-optimized version of selective_state_update (mamba) sources/prs/flashinfer/PR-2387.md
feat: cuteDSL fp4 moe for better DSR1 performance. sources/prs/flashinfer/PR-2398.md
perf: improve gdn decode cute-dsl kernels sources/prs/flashinfer/PR-2405.md
refactor: simplify fp4 rmsnorm sources/prs/flashinfer/PR-2421.md
refactor: refactoring cuda code to cute-dsl (part 1) sources/prs/flashinfer/PR-2428.md
fix: Fix NaN output in mxfp8_quantize for very small input values sources/prs/flashinfer/PR-2441.md
Add cute-dsl backends to mxfp[8,4]_quantization for future refactor sources/prs/flashinfer/PR-2443.md
MTP for mamba sources/prs/flashinfer/PR-2444.md
feat: Add MXFP8 GEMM mm_mxfp8 (cutlass) sources/prs/flashinfer/PR-2464.md
refactor: Port upstream CUTLASS fixes and refactor grouped_gemm_nt_masked GEMM module location sources/prs/flashinfer/PR-2503.md
feat: cute dsl mmfp4 for blackwell sources/prs/flashinfer/PR-2540.md
Implement cutlass_fused_moe mxfp8 sources/prs/flashinfer/PR-2581.md
feat: trtllm tinygemm2 in flashinfer as bf16 routergemm sources/prs/flashinfer/PR-2587.md
Mamba SSU: better automatic kernel selection + algorithm selection optionally exposed to the user. sources/prs/flashinfer/PR-2591.md
int16 Block-Scaled State and Stochastic Rounding for SSU (mamba) sources/prs/flashinfer/PR-2645.md
feat: support mxfp4 & mxfp8 entrypoint for blackwell cutedsl dense gemm sources/prs/flashinfer/PR-2660.md
Add NVFP4 KV cache quantization support for SM100 sources/prs/flashinfer/PR-2702.md
Mamba2 SSD Combined Forward Pass (Blackwell CuTe DSL Kernel) sources/prs/flashinfer/PR-2709.md
feat: Add DiT-oriented kernels where Qk (Bmm1) type can be reinterpreted into Int8 or BFloat16 sources/prs/flashinfer/PR-2711.md
[gdn] support non-contiguous state for decoding sources/prs/flashinfer/PR-2727.md
Add cute dsl mla decode op sources/prs/flashinfer/PR-2743.md
feat: Add FP4 KV cache quant/dequant kernels sources/prs/flashinfer/PR-2757.md
perf: Performance tune cute dsl RMSNorm variants sources/prs/flashinfer/PR-2777.md
feat: FP8 output support for CUTLASS MLA paged attention sources/prs/flashinfer/PR-2779.md
[CuTe DSL] Add modular FMHA prefill and MLA decode attention kernels sources/prs/flashinfer/PR-2805.md
[Fmha] Sparse MLA decode kernel selection heuristics sources/prs/flashinfer/PR-2836.md
feat: Add CuTe-DSL backend for NVFP4 quantization sources/prs/flashinfer/PR-2838.md
Mamba SSU: horizontal MTP kernel (+ DSTATE=96 support) sources/prs/flashinfer/PR-2865.md
perf: Optimize CuTe-DSL fp4 and fp8 quantization kernels sources/prs/flashinfer/PR-2904.md
[NVIDIA] fix(jit): enable GDC for CUTLASS fused MoE PDL — prevent random crashes on SM12x sources/prs/flashinfer/PR-2913.md
feat: Add cuBLASLt backend for mm_bf16 and enable multi-tactic autotuning for FP8/MXFP8 runners sources/prs/flashinfer/PR-2914.md
feat: Add CuTe DSL grouped-gemm + combine fusion support sources/prs/flashinfer/PR-2944.md
fix: use float instead of double in sampling binary search to avoid FP64 bottleneck on SM103 sources/prs/flashinfer/PR-2945.md
Improved simple mamba SSU kernel sources/prs/flashinfer/PR-2962.md
Add flashinfer.fused_rmsnorm_silu() with native kernel backend sources/prs/flashinfer/PR-2965.md
[feat] Add blackwell GDN prefill kernel sources/prs/flashinfer/PR-3001.md
[feat] Add routing_replay_out support to MoE kernels and Python API sources/prs/flashinfer/PR-3024.md
perf: Port TRT-LLM SM120/SM121 FP4 CUTLASS GEMM optimizations. Add PDL sources/prs/flashinfer/PR-3026.md
[feat] Trtllm-gen Per-token Nvfp4 MoE sources/prs/flashinfer/PR-3027.md
feat: Add b12x CuTe DSL fused MoE for SM120 sources/prs/flashinfer/PR-3066.md
feat: DiT layer norm fusions for WAN: flashinfer.diffusion_ops sources/prs/flashinfer/PR-3157.md
fix(cute_dsl/moe): make autotuner bucket configuration adapt to runtime input sources/prs/flashinfer/PR-3216.md
Support Kimi K2.5 H64 CuTe DSL MLA decode sources/prs/flashinfer/PR-3235.md
Ameyn/gdn bf16 dispatcher and 4d pool sources/prs/flashinfer/PR-3268.md
feat(cute_dsl/moe): deterministic balanced autotune profile inputs sources/prs/flashinfer/PR-3286.md
checkpointing_ssu kernel: fused replay + conditional state-write for Mamba2 sources/prs/flashinfer/PR-3324.md
[CUDA][avgpool2d] Fix backward launch bounds again for sm100, sm120 sources/prs/pytorch/PR-150640.md
[CUDA][avgpool2d] Fix backward launch bounds again for sm100, sm120 sources/prs/pytorch/PR-150676.md
feat: Add FP4 (E2M1) KV Cache Support with Quantization Utilities for MLA sources/prs/sglang/PR-10078.md
[NVIDIA] Add new SMs support for Spark & Thor sources/prs/sglang/PR-11287.md
[sgl-kernel][1/N]Support Expert Specialization Grouped GEMM sources/prs/sglang/PR-11432.md
[DeepseekV32] Enable flashmla_prefill kernel with fp8 kvcache sources/prs/sglang/PR-11655.md
support cutlass fp4 kernel in sm120 sources/prs/sglang/PR-11737.md
[sgl-kernel][Feat][B200][1/N]Support MXFP8 Grouped GEMM in Blackwell sources/prs/sglang/PR-13731.md
[sgl-kernel][Feat][B200][2/N] Support MXFP8 Grouped GEMM in Blackwell sources/prs/sglang/PR-14640.md
[NVIDIA] upstream FA4 sources/prs/sglang/PR-15182.md
[jit-kernel] Add CuTe DSL GDN Decode Kernel sources/prs/sglang/PR-15631.md
Move fa4 from sgl-kernel to jit kernel sources/prs/sglang/PR-17353.md
Add mxfp8 support for online quantization, Triton dense linear, and CUTLASS MoE sources/prs/sglang/PR-17449.md
[Diffsuion & JIT_kernel] QKNorm cross heads kernel sources/prs/sglang/PR-18073.md
feat: add FA4 SM90 paged KV decode support & update attention docs sources/prs/sglang/PR-18442.md
[FIX] Correct JIT kernel compilation on newer GPUs with outdated driver metadata. sources/prs/sglang/PR-18496.md
[Kernel Slimming] Migrate NVFP4 kernels to JIT sources/prs/sglang/PR-19437.md
[Feature] NVFP4 Marlin fallback for non-Blackwell GPUs (SM75+) sources/prs/sglang/PR-19652.md
[JIT Kernel] Reland NVFP4 kernels to JIT sources/prs/sglang/PR-20012.md
Fix(jit): support rmsnorm for hidden_size in {64, 128, 256} sources/prs/sglang/PR-20661.md
CUTLASS NVFP4 GEMM improvement of SM120 sources/prs/sglang/PR-21314.md
Fused_qknorm_rope kernel optimization: up to 2.4× faster sources/prs/sglang/PR-21654.md
[Diffusion] Fix weight scale swizzle and add large-M kernel config for FLUX.2-dev-NVFP4 sources/prs/sglang/PR-22064.md
[nvidia] Gemma4 nvfp4 fix sources/prs/sglang/PR-22079.md
[Refactor] Rename NSA → DSA: user-facing aliases, file/class/import rename sources/prs/sglang/PR-25821.md
Support FP4 gemm (1/2) sources/prs/sglang/PR-3899.md
Support Blackwell Block Scale FP8 Gemm sources/prs/sglang/PR-4278.md
support cmake for sgl-kernel sources/prs/sglang/PR-4706.md
[Build] Fix cuda12.8 build error in nvfp4_scaled_mm_kernels.cu sources/prs/sglang/PR-4953.md
[1/2] Add FP8 Blockscale MoE CUTLASS kernel for Blackwell sources/prs/sglang/PR-5281.md
[2/2] Add python wrapper for CUTLASS FP8 Blockscale MoE Kernel. sources/prs/sglang/PR-5694.md
[1/2] Add Kernel support for Cutlass based Fused FP4 MoE sources/prs/sglang/PR-6093.md
Add a CUDA kernel for fusing mapping and weighted sum for MoE. sources/prs/sglang/PR-6916.md
[perf][sgl-kernel] extend cutlass_mla_decode to support num_head < 128 sources/prs/sglang/PR-6929.md
chore: upgrade flashinfer v0.2.6.post1 jit sources/prs/sglang/PR-6958.md
Add CUTLASS FP8 Blockscale MoE kernel for Hopper architecture sources/prs/sglang/PR-7278.md
[Perf] Tunings for SM100 FP8 CUTLASS kernel sources/prs/sglang/PR-8818.md
optimize: reduce shulffle and quantization overhead in cutlass_moe sm90 sources/prs/sglang/PR-8962.md
[NVIDIA] [2/N] Optimize silu_and_mul_scaled_fp4_grouped_quant perf sources/prs/sglang/PR-9556.md
Make sm100 fp8 kernels available on sm103 sources/prs/sglang/PR-9789.md
Make fp4_quantize kernels work on sm103 sources/prs/sglang/PR-9807.md
[Model] Support Meituan LongCat-Flash && LongCat-Flash-MTP sources/prs/sglang/PR-9824.md
[WIP] support more dtypes for tcgen05 sources/prs/tilelang/PR-1229.md
[Enhancement] add more dtype and fix mma.ws for fp16 for tcgen05 sources/prs/tilelang/PR-1327.md
[CUDA] Support tcgen5mma gemm ts sources/prs/tilelang/PR-1866.md
[Feature] Support cluster launch, query, synchronization and barrier operations sources/prs/tilelang/PR-1874.md
[Feature] 2-SM support for TMA, TMEM and TCGEN5MMA on Blackwell sources/prs/tilelang/PR-1882.md
[Feature] Block-scaled GEMM support for MXFP8 on Blackwell sources/prs/tilelang/PR-1945.md
[Transform] Add InjectTcgen05Fence pass sources/prs/tilelang/PR-2003.md
[codex] Split GEMM implementations by backend sources/prs/tilelang/PR-2153.md
[NVIDIA] Support nvfp4 cutlass gemm sources/prs/vllm/PR-13571.md
add cutlass support for blackwell fp8 gemm sources/prs/vllm/PR-13798.md
Add cutlass support for blackwell fp8 blockwise gemm sources/prs/vllm/PR-14383.md
[NVIDIA] Support Cutlass w8a8 FP8 for Blackwell Geforce GPUs (sm120) sources/prs/vllm/PR-17280.md
Sm100 blockwise fp8 swap ab sources/prs/vllm/PR-18564.md
[Perf] Tunings for SM100 FP8 CUTLASS kernel sources/prs/vllm/PR-18778.md
[Hardware][NVIDIA] FP4 MoE kernel optimization sources/prs/vllm/PR-19110.md
[Hardware][NVIDIA][kernel] Fp4 MOE quant kernel optimization sources/prs/vllm/PR-19500.md
[Perf] Further tunings for SM100 FP8 CUTLASS kernel sources/prs/vllm/PR-19566.md
[feat]: CUTLASS block scaled group gemm for SM100 sources/prs/vllm/PR-19757.md
[Feature] Integrate SM100 DeepGEMM support sources/prs/vllm/PR-20087.md
[Kernel] SM90 CUTLASS FP8 GEMM: add support for swap AB + kernel tuning sources/prs/vllm/PR-20396.md
[feat]: add SM100 support for cutlass FP8 groupGEMM sources/prs/vllm/PR-20447.md
SM100 Cutlass MLA decode with unrestricted num_heads (< 128) for DeepSeek TP sources/prs/vllm/PR-20769.md
[fix]: disable cutlass block scaled group gemm for EP sources/prs/vllm/PR-20781.md
[Perf] Cuda Kernel for Per Token Group Quant sources/prs/vllm/PR-21083.md
[Bug] Fix B200 DeepGEMM E8M0 Accuracy Issue sources/prs/vllm/PR-22399.md
[Perf] Use upstream CUTLASS for SM90 Block FP8 kernel sources/prs/vllm/PR-23280.md
[Compile] Fix Compile Warning SM100 Cutlass MLA sources/prs/vllm/PR-23287.md
[NVIDIA] Support SiluMul + NVFP4 quant fusion sources/prs/vllm/PR-23671.md
[Kernel] Support decode context parallelism on Blackwell with CUTLASS MLA sources/prs/vllm/PR-24385.md
[Transform] [Quantization] Add QuTLASS support to vLLM sources/prs/vllm/PR-24440.md
[NVIDIA] Blackwell Family sources/prs/vllm/PR-24673.md
[Kernel][Quantization] add w4a8 support for marlin kernel sources/prs/vllm/PR-24722.md
[Perf] SM100 - add swap AB optimization to CUTLASS FP8 GEMM sources/prs/vllm/PR-27284.md
[Perf][DeepSeek] Add sigmoid+bias fusion to fused_grouped_topk from TRTLLM sources/prs/vllm/PR-28124.md
[Performance][B200] silu_mul_quant: pack scales in int32 sources/prs/vllm/PR-28358.md
[Kernel] Add NVFP4 MoE CUTLASS support for SM120 sources/prs/vllm/PR-29242.md
[Perf][Kernel] Optimize FP4 quantization kernels (SM100F) sources/prs/vllm/PR-32520.md
[Performance] Tune Mamba selective scan kernel for B200 sources/prs/vllm/PR-32873.md
[Spec Decode] Unified Parallel Drafting sources/prs/vllm/PR-32887.md
[Kernel] Integrate SM100 MXFP8 blockscaled grouped MM and quant kernels sources/prs/vllm/PR-34448.md
[Quantization] add humming quantization kernel sources/prs/vllm/PR-34556.md
[Bugfix] Gate 256-bit instructions to CUDA 12.9+ sources/prs/vllm/PR-34791.md
[BUGFIX][Mamba][Qwen3.5] Zero freed SSM cache blocks on GPU sources/prs/vllm/PR-35219.md
[Mamba] Add stochastic rounding support sources/prs/vllm/PR-35753.md
[Kernel] Fuse FP8 output quantization into merge_attn_states sources/prs/vllm/PR-36518.md
[Feat][Spec Decode] DFlash sources/prs/vllm/PR-36847.md
[Kernel] Add non-gated support for NVFP4 CUTLASS MoE sources/prs/vllm/PR-37320.md
Add nvfp4 support to reshape_and_cache_flash sources/prs/vllm/PR-37332.md
[Perf][Kernel] Persistent TopK scheduler: unified CUDAGraph-safe kernel with dynamic per-row dispatch - DeepSeek-V3.2 DSA decode sources/prs/vllm/PR-37421.md
[Kernel] Add MXFP4 W4A4 CUTLASS MoE kernel for SM100 sources/prs/vllm/PR-37463.md
[4/n] Migrate FP4/W4A8 CUTLASS kernels to torch stable ABI sources/prs/vllm/PR-37503.md
[Perf] triton bilinear_pos_embed kernel for ViT sources/prs/vllm/PR-37948.md
[Perf] FP8 FlashInfer Attn for ViT sources/prs/vllm/PR-38065.md
[NVIDIA] Bugfix NVFP4 DGX Spark and RTX50 sources/prs/vllm/PR-38423.md
[Refactor] Improve indexer decode path metadata preparation sources/prs/vllm/PR-38865.md
[Perf][GDN] Align TMA usage with upstream FLA sources/prs/vllm/PR-38981.md
[Bugfix] Guard mxfp4_experts_quant bindings on ENABLE_NVFP4_SM100 sources/prs/vllm/PR-40191.md
[Perf] Batch invariance with Cutlass fp8 support, 28.9% E2E latency improvement sources/prs/vllm/PR-40408.md
[DSv4] Improved fused Indexer Q quant kernel sources/prs/vllm/PR-41428.md
[DSv4] Improved dequant gather K cache kernel sources/prs/vllm/PR-42236.md
Two-CTA Cooperative MMA wiki/hardware/2sm-cooperative.md
Cluster Launch Control (CLC) wiki/hardware/clc.md
mbarrier (Memory Barrier Primitives) wiki/hardware/mbarrier.md
NVFP4 and block-scaled narrow precision wiki/hardware/nvfp4.md
Programmatic Dependent Launch / Grid Dependency Control wiki/hardware/pdl-gdc.md
tcgen05.mma — Fifth-Generation Tensor Core MMA wiki/hardware/tcgen05-mma.md
Tensor Memory Accelerator (TMA) wiki/hardware/tma.md
Tensor Memory (TMEM) wiki/hardware/tmem.md
DeepGEMM — runtime-JIT tensor-core kernels wiki/kernels/deepgemm.md
FlashAttention-4 wiki/kernels/flash-attention-4.md
FlashAttention SM100 MLA TopK Sparse Forward wiki/kernels/flash-attention-sm100-mla-topk.md
FlashMLA attention kernels wiki/kernels/flashmla.md
FP8 block-scale GEMM wiki/kernels/fp8-block-scale-gemm.md
Fused MoE — Expert GEMM and Adjacent Operations wiki/kernels/fused-moe.md
Gated Delta Network kernels wiki/kernels/gated-delta-net.md
Gated Dual GEMM (Gate-Up + Activation) wiki/kernels/gated-dual-gemm.md
Grouped GEMM for MoE wiki/kernels/grouped-gemm.md
NVFP4 GEMM wiki/kernels/nvfp4-gemm.md
NVFP4 batched GEMV wiki/kernels/nvfp4-gemv.md
Sparse MLA wiki/kernels/sparse-mla.md
TensorRT-LLM Blackwell FP4 DSA Indexer wiki/kernels/tensorrt-llm-blackwell-indexer.md
CUDA C++ for Blackwell Kernels wiki/languages/cuda-cpp.md
CuTe DSL for Blackwell wiki/languages/cute-dsl.md
PTX for SM100 wiki/languages/ptx-sm100.md
Triton on Blackwell wiki/languages/triton-blackwell.md
PTX Cache Policy Differentiation wiki/techniques/cache-policy.md
CCCL CUB SM100 Scan Tuning wiki/techniques/cccl-memory-primitives.md
Double/Multi-Buffering Patterns wiki/techniques/double-buffering.md
Epilogue fusion wiki/techniques/epilogue-fusion.md
External Source-Map Research For Kernel Edits wiki/techniques/external-source-map-research.md
Fine-grained FP8/FP4 scaling wiki/techniques/fine-grained-quantization.md
Kernel fusion wiki/techniques/kernel-fusion.md
Persistent Kernels with Cluster Launch Control wiki/techniques/persistent-kernels.md
Ping-Pong Scheduling wiki/techniques/ping-pong-scheduling.md
Software Pipelining and Multi-Stage Buffering wiki/techniques/pipeline-stages.md
Register budgeting wiki/techniques/register-budgeting.md
Software-Emulated Exponential wiki/techniques/software-exp.md
Shared Memory Swizzling wiki/techniques/swizzling.md
Tile Scheduling Strategies wiki/techniques/tile-scheduling.md
Wide Vectorized Loads and Cache Policies wiki/techniques/vectorized-loads.md
Warp Specialization on Hopper and Blackwell wiki/techniques/warp-specialization.md

sm100a

Page Path
TFLOPS Gap: Why FP4 MoE Kernel Engineering Matters on Blackwell sources/blogs/tflops-gap-fp4-moe.md
NVIDIA Blackwell Compatibility Guide sources/docs/blackwell-compatibility-guide.md
CUTLASS Changelog: SM100/Blackwell Entries sources/docs/cutlass-changelog-sm100.md
CUTLASS CuTe DSL Documentation sources/docs/cutlass-cute-dsl.md
NVIDIA CUTLASS Blackwell support map sources/docs/nvidia-cutlass-blackwell.md
PTX ISA Fifth-Generation Tensor Core and CLC Reference sources/docs/nvidia-ptx-isa-sm100.md
Triton 3.6 Release Notes — Blackwell Backend Work sources/docs/triton-3.6-blackwell.md
[TRTLLM-11119][feat] Blackwell SageAttention, Integrate into AttentionOp API sources/prs/TensorRT-LLM/PR-11718.md
[None][feat] Support sparse mqa/gqa attention sources/prs/TensorRT-LLM/PR-12470.md
[None][feat] Trtllm-gen FMHA JIT support sources/prs/TensorRT-LLM/PR-12612.md
[TRTLLM-11485][feat] Feature rework: Add SageAttention refreshed kernels (attentionOp only) sources/prs/TensorRT-LLM/PR-12937.md
[None][feat] Add DeepSeekV4 attention kernels sources/prs/TensorRT-LLM/PR-13652.md
[None][feat] Update the logic of FMHA JIT path sources/prs/TensorRT-LLM/PR-14291.md
[OMNIML-2336][feat] Add NVFP4 x FP8 sources/prs/TensorRT-LLM/PR-6809.md
new example with TMA prefetch feature targeting for DRAM latency boun… sources/prs/cutlass/PR-2881.md
[CLI] add cutedsl fp16 gemm tutorial from 2 to 6 sources/prs/cutlass/PR-3106.md
[nvidia] initial support for blackwell kernels sources/prs/flashinfer/PR-1039.md
bugfix: host-precomuted plan function for blackwell fmha sources/prs/flashinfer/PR-1106.md
Add CUTLASS fused moe kernels from TensorRT-LLM. sources/prs/flashinfer/PR-1113.md
bugfix: fix blackwell fmha hanging issue for empty kv_len sources/prs/flashinfer/PR-1198.md
Feature/sm100 low latency nvfp4 kernels sources/prs/flashinfer/PR-1214.md
feat: Support MXFP8 x MXFP4 CUTLASS grouped GEMM sources/prs/flashinfer/PR-1241.md
GPT-OSS Support: Add Blackwell MoE mxfp4 implementation from TRTLLM and Attention Sink sources/prs/flashinfer/PR-1389.md
gpt-oss: Add MXFP8 x MXFP4 CUTLASS MOE for SM100 and BF16 x MXFP4 CUTLASS for SM90 + SwigluBias Activation sources/prs/flashinfer/PR-1396.md
fix: separate out fp4 lib into sm90 and sm100 versions, add oob checking in fused moe sources/prs/flashinfer/PR-1565.md
feat: initial support for SM103, SM110, SM120, SM121 sources/prs/flashinfer/PR-1608.md
feat: cutlass fp8 gemm bringup for SM120 & SM121 sources/prs/flashinfer/PR-1610.md
Masked batch nvfp4 quantization sources/prs/flashinfer/PR-1774.md
silu_and_mul nvfp4 quanization fusion rework sources/prs/flashinfer/PR-1927.md
Rebase FP8 SM100 Cutlass FMHA Attention to main (original PR#1238) sources/prs/flashinfer/PR-2047.md
Add cute-dsl backends to mxfp[8,4]_quantization for future refactor sources/prs/flashinfer/PR-2443.md
Mamba2 SSD Combined Forward Pass (Blackwell CuTe DSL Kernel) sources/prs/flashinfer/PR-2709.md
Support for MXFP4 and NVFP4 group GEMMs on GeForce and Spark sources/prs/flashinfer/PR-2738.md
Add cute dsl mla decode op sources/prs/flashinfer/PR-2743.md
feat: FP8 output support for CUTLASS MLA paged attention sources/prs/flashinfer/PR-2779.md
[CuTe DSL] Add modular FMHA prefill and MLA decode attention kernels sources/prs/flashinfer/PR-2805.md
Mamba SSU: horizontal MTP kernel (+ DSTATE=96 support) sources/prs/flashinfer/PR-2865.md
perf: Optimize CuTe-DSL fp4 and fp8 quantization kernels sources/prs/flashinfer/PR-2904.md
fix: use float instead of double in sampling binary search to avoid FP64 bottleneck on SM103 sources/prs/flashinfer/PR-2945.md
[feat] Add blackwell GDN prefill kernel sources/prs/flashinfer/PR-3001.md
checkpointing_ssu kernel: fused replay + conditional state-write for Mamba2 sources/prs/flashinfer/PR-3324.md
[NVIDIA] Add new SMs support for Spark & Thor sources/prs/sglang/PR-11287.md
support cutlass fp4 kernel in sm120 sources/prs/sglang/PR-11737.md
[nvidia] Gemma4 nvfp4 fix sources/prs/sglang/PR-22079.md
Support Blackwell Block Scale FP8 Gemm sources/prs/sglang/PR-4278.md
support cmake for sgl-kernel sources/prs/sglang/PR-4706.md
[1/2] Add FP8 Blockscale MoE CUTLASS kernel for Blackwell sources/prs/sglang/PR-5281.md
[NVIDA] [1/N] Nvfp4 Masked Gemm: Add quant op for the flashinfer grouped gemm sources/prs/sglang/PR-9200.md
Make fp4_quantize kernels work on sm103 sources/prs/sglang/PR-9807.md
[NVIDIA] Support nvfp4 quantization sources/prs/vllm/PR-12784.md
Two-CTA Cooperative MMA wiki/hardware/2sm-cooperative.md
mbarrier (Memory Barrier Primitives) wiki/hardware/mbarrier.md
NVFP4 and block-scaled narrow precision wiki/hardware/nvfp4.md
tcgen05.mma — Fifth-Generation Tensor Core MMA wiki/hardware/tcgen05-mma.md
Tensor Memory Accelerator (TMA) wiki/hardware/tma.md
Tensor Memory (TMEM) wiki/hardware/tmem.md
Fused MoE — Expert GEMM and Adjacent Operations wiki/kernels/fused-moe.md
NVFP4 GEMM wiki/kernels/nvfp4-gemm.md
CUDA C++ for Blackwell Kernels wiki/languages/cuda-cpp.md
CuTe DSL for Blackwell wiki/languages/cute-dsl.md
PTX for SM100 wiki/languages/ptx-sm100.md

sm100f

Page Path
[TRTLLM-9457][feat] Add cute dsl fp8 gemm for Blackwell sources/prs/TensorRT-LLM/PR-10130.md
[None][feat] Optimize super-v3 nvfp4 for better perf sources/prs/TensorRT-LLM/PR-11273.md
[TRTLLM-11289][feat] Integrate CuteDSL's bf16 dense GEMMs sources/prs/TensorRT-LLM/PR-12074.md
[None][feat] Support sparse mqa/gqa attention sources/prs/TensorRT-LLM/PR-12470.md
[None][feat] Trtllm-gen FMHA JIT support sources/prs/TensorRT-LLM/PR-12612.md
[None][feat] Optimize mamba SSD prefill and extend flashinfer dispatch sources/prs/TensorRT-LLM/PR-12731.md
[TRTLLM-11485][feat] Feature rework: Add SageAttention refreshed kernels (attentionOp only) sources/prs/TensorRT-LLM/PR-12937.md
[#12784][feat] AutoDeploy: Optimize DeepSeek-R1 model performance sources/prs/TensorRT-LLM/PR-12946.md
[TRTLLM-34871][feat] Add cute dsl FP8 paged MQA logits decode kernel sources/prs/TensorRT-LLM/PR-13219.md
[None][feat] Add DeepSeekV4 attention kernels sources/prs/TensorRT-LLM/PR-13652.md
[TRTLLM-35237][feat] Add cute dsl FP4 paged MQA logits decode kernel sources/prs/TensorRT-LLM/PR-13929.md
[None][feat] Keep DSv4 o_a_proj as FP8, and port vLLM's fused_inv_rope_fp8_quant sources/prs/TensorRT-LLM/PR-13938.md
[None][feat] Update the logic of FMHA JIT path sources/prs/TensorRT-LLM/PR-14291.md
[TRTLLM-8535][feat] Support DeepSeek V3.2 with FP8 + BF16 KV cache/NVFP4 + BF16 KV cache sources/prs/TensorRT-LLM/PR-8405.md
feat: BF16 GEMM using CUTLASS backend for SM100 sources/prs/flashinfer/PR-2070.md
[TRTLLM-Gen Fmha] add optimized trtllm-gen decode kernels for high throughput + speculative decoding sources/prs/flashinfer/PR-2265.md
[Perf][Kernel] Optimize FP4 quantization kernels (SM100F) sources/prs/vllm/PR-32520.md
[Kernel] Add MXFP4 W4A4 CUTLASS MoE kernel for SM100 sources/prs/vllm/PR-37463.md

sm103

Page Path
PTX ISA Fifth-Generation Tensor Core and CLC Reference sources/docs/nvidia-ptx-isa-sm100.md
Sync nv_dev with upstream #316 (Mega MoE optimizations & benchmarks) sources/prs/DeepGEMM/PR-328.md
[TRTLLM-9831][perf] Enable 2CTA with autotune for CuteDSL MoE and Grouped GEMM optimizations sources/prs/TensorRT-LLM/PR-10201.md
[None] [feat] Add densegemm backend for MoE sources/prs/TensorRT-LLM/PR-10479.md
[None][feat] Optimize causal_conv1d prefill and decode kernels sources/prs/TensorRT-LLM/PR-13103.md
[None][fix] Plumb swiglu_limit through DeepGEMM and TRTLLMGen FP8 fused MoE sources/prs/TensorRT-LLM/PR-13767.md
[None][perf] mHC fused_hc kernel optimizations + DS-V4 entry-boundary RMSNorm fold-in sources/prs/TensorRT-LLM/PR-13892.md
[None][feat] DSv4: enable GVR Heuristic Top-K for compress_ratio=4 sources/prs/TensorRT-LLM/PR-14219.md
[TRTLLM-9685] [feat] Add gather fc1 kernel by cuteDSL sources/prs/TensorRT-LLM/PR-9618.md
[Cute,Fwd,Sm100] fp8 e4m3 and e5m2 support sources/prs/flash-attention/PR-2109.md
feat: initial support for SM103, SM110, SM120, SM121 sources/prs/flashinfer/PR-1608.md
bugfix: partially fix tests/test_trtllm_gen_fused_moe.py unit test failure sources/prs/flashinfer/PR-1724.md
Add head_dim=64 for blackwell cutlass fmha implementation sources/prs/flashinfer/PR-1850.md
MLA RoPE + quantization fused kernel: shape generalization for MHA / GQA sources/prs/flashinfer/PR-1924.md
[NVIDIA] Thor & Spark Support sources/prs/flashinfer/PR-2028.md
feat: add trtllm-gen per-tensor sparseMla kernels. sources/prs/flashinfer/PR-2138.md
enable sm103 moe dsl backend sources/prs/flashinfer/PR-2149.md
[TRTLLM-Gen Fmha] add optimized trtllm-gen decode kernels for high throughput + speculative decoding sources/prs/flashinfer/PR-2265.md
[Perf][Feature] Add SM103-specific schedulers for NVFP4 CUTLASS kernels sources/prs/flashinfer/PR-2303.md
feat: cute dsl mmfp4 for blackwell sources/prs/flashinfer/PR-2540.md
Implement cutlass_fused_moe mxfp8 sources/prs/flashinfer/PR-2581.md
feat: Add DiT-oriented kernels where Qk (Bmm1) type can be reinterpreted into Int8 or BFloat16 sources/prs/flashinfer/PR-2711.md
[Fmha] Sparse MLA decode kernel selection heuristics sources/prs/flashinfer/PR-2836.md
[NVIDIA] fix(jit): enable GDC for CUTLASS fused MoE PDL — prevent random crashes on SM12x sources/prs/flashinfer/PR-2913.md
feat: Add cuBLASLt backend for mm_bf16 and enable multi-tactic autotuning for FP8/MXFP8 runners sources/prs/flashinfer/PR-2914.md
fix: use float instead of double in sampling binary search to avoid FP64 bottleneck on SM103 sources/prs/flashinfer/PR-2945.md
[feat] Add routing_replay_out support to MoE kernels and Python API sources/prs/flashinfer/PR-3024.md
[feat] Trtllm-gen Per-token Nvfp4 MoE sources/prs/flashinfer/PR-3027.md
feat: Enable FP8 (E4M3/E5M2) in concat_mla_k for optimize long-context prefill performance and refactor type dispatch for BF16/FP16 sources/prs/flashinfer/PR-3129.md
feat: DiT layer norm fusions for WAN: flashinfer.diffusion_ops sources/prs/flashinfer/PR-3157.md
Disable kernel cutlass_mla_decode on SM103 sources/prs/sglang/PR-10058.md
support cutlass fp4 kernel in sm120 sources/prs/sglang/PR-11737.md
[Refactor] Rename NSA → DSA: user-facing aliases, file/class/import rename sources/prs/sglang/PR-25821.md
Make sm100 fp8 kernels available on sm103 sources/prs/sglang/PR-9789.md
Make fp4_quantize kernels work on sm103 sources/prs/sglang/PR-9807.md
[Attention][Perf][Kernel] Replace torch.cat with vectorized CUDA kernel MLA query concat - DeepSeek-V3.2 sources/prs/vllm/PR-34917.md
[Perf] FP8 FlashInfer Attn for ViT sources/prs/vllm/PR-38065.md
[MLA] Optimize mla indexer prepare uniform decode for MTP > 1 sources/prs/vllm/PR-39458.md
[DSV4] Fuse norm and router for low latency scenario sources/prs/vllm/PR-41263.md
Cluster Launch Control (CLC) wiki/hardware/clc.md

sm103a

Page Path
[TRTLLM-11119][feat] Blackwell SageAttention, Integrate into AttentionOp API sources/prs/TensorRT-LLM/PR-11718.md
[None][feat] Support sparse mqa/gqa attention sources/prs/TensorRT-LLM/PR-12470.md
[None][feat] Trtllm-gen FMHA JIT support sources/prs/TensorRT-LLM/PR-12612.md
[TRTLLM-11485][feat] Feature rework: Add SageAttention refreshed kernels (attentionOp only) sources/prs/TensorRT-LLM/PR-12937.md
[None][feat] Add DeepSeekV4 attention kernels sources/prs/TensorRT-LLM/PR-13652.md
[None][feat] Update the logic of FMHA JIT path sources/prs/TensorRT-LLM/PR-14291.md
feat: initial support for SM103, SM110, SM120, SM121 sources/prs/flashinfer/PR-1608.md
fix: use float instead of double in sampling binary search to avoid FP64 bottleneck on SM103 sources/prs/flashinfer/PR-2945.md
[NVIDIA] Add new SMs support for Spark & Thor sources/prs/sglang/PR-11287.md
support cutlass fp4 kernel in sm120 sources/prs/sglang/PR-11737.md
Make fp4_quantize kernels work on sm103 sources/prs/sglang/PR-9807.md

sm110

Page Path
cuTile Python Documentation sources/docs/cutile-python-dsl.md
PTX ISA Fifth-Generation Tensor Core and CLC Reference sources/docs/nvidia-ptx-isa-sm100.md
feat: initial support for SM103, SM110, SM120, SM121 sources/prs/flashinfer/PR-1608.md
[NVIDIA] Thor & Spark Support sources/prs/flashinfer/PR-2028.md
[NVIDIA] fix(jit): enable GDC for CUTLASS fused MoE PDL — prevent random crashes on SM12x sources/prs/flashinfer/PR-2913.md
feat: DiT layer norm fusions for WAN: flashinfer.diffusion_ops sources/prs/flashinfer/PR-3157.md
Cluster Launch Control (CLC) wiki/hardware/clc.md

sm110a

Page Path
feat: initial support for SM103, SM110, SM120, SM121 sources/prs/flashinfer/PR-1608.md
Rebase FP8 SM100 Cutlass FMHA Attention to main (original PR#1238) sources/prs/flashinfer/PR-2047.md
Add cute-dsl backends to mxfp[8,4]_quantization for future refactor sources/prs/flashinfer/PR-2443.md
Add cute dsl mla decode op sources/prs/flashinfer/PR-2743.md
feat: FP8 output support for CUTLASS MLA paged attention sources/prs/flashinfer/PR-2779.md
perf: Optimize CuTe-DSL fp4 and fp8 quantization kernels sources/prs/flashinfer/PR-2904.md

sm120

Page Path
CUDA Programming Guide: Programmatic Dependent Launch sources/docs/cuda-programming-guide-pdl.md
cuTile Python Documentation sources/docs/cutile-python-dsl.md
CUTLASS CuTe DSL Documentation sources/docs/cutlass-cute-dsl.md
NVIDIA Blackwell Tuning Guide sources/docs/nvidia-blackwell-tuning-guide.md
PTX ISA Fifth-Generation Tensor Core and CLC Reference sources/docs/nvidia-ptx-isa-sm100.md
[https://nvbugs/5854860][fix] Fix cutedsl argmax on sm120 sources/prs/TensorRT-LLM/PR-11181.md
[None][feat] Fuse FP8 1x128 quantize + UE8M0 scale pack on SM100 sources/prs/TensorRT-LLM/PR-13628.md
feat: Add w4a8_mxfp4_fp8 quantization recipe. sources/prs/TensorRT-LLM/PR-4867.md
[None][feat] GPT-OSS Sm120/Sm121 Support sources/prs/TensorRT-LLM/PR-7937.md
[None][feat] Enable nvfp4 cuda core for sm120 sources/prs/TensorRT-LLM/PR-8620.md
Use integer promotion for warp_reduce sources/prs/cccl/PR-6819.md
Replace detail::merge::dispatch by CUB's public API sources/prs/cccl/PR-8381.md
Replace detail::merge_sort::dispatch by CUB's public API sources/prs/cccl/PR-8473.md
Implement the new tuning API for detail::batched_topk::dispatch_batched_topk sources/prs/cccl/PR-8538.md
Replace detail::for_each::dispatch by CUB's public API sources/prs/cccl/PR-8565.md
Use the new tuning API internally for detail::select::dispatch and DeviceSelect sources/prs/cccl/PR-8880.md
[Bug Fix]Bypass launch grids for SM120 Kernel with SM90 Mainloop & SM100 TileScheduler sources/prs/cutlass/PR-2865.md
Replace std::min with cute::min in sm120 blockwise scaling device functions sources/prs/cutlass/PR-3055.md
Support for Group GEMM in CUTLASS Profiler for GeForce and Spark sources/prs/cutlass/PR-3092.md
Small Tile N BlockScaled GEMM + Grouped GEMM on SM12x sources/prs/cutlass/PR-3176.md
Add SM120 varlen attention support sources/prs/flash-attention/PR-2333.md
Add CUTLASS fused moe kernels from TensorRT-LLM. sources/prs/flashinfer/PR-1113.md
Update cutlass fp4 moe kernels sources/prs/flashinfer/PR-1294.md
add cutlass backend for mm_fp4 sources/prs/flashinfer/PR-1296.md
gpt-oss: Add MXFP8 x MXFP4 CUTLASS MOE for SM100 and BF16 x MXFP4 CUTLASS for SM90 + SwigluBias Activation sources/prs/flashinfer/PR-1396.md
feat: integrate xqa attention backend sources/prs/flashinfer/PR-1503.md
feat: initial support for SM103, SM110, SM120, SM121 sources/prs/flashinfer/PR-1608.md
feat: cutlass fp4 gemm bringup for SM120 & SM121 sources/prs/flashinfer/PR-1609.md
feat: cutlass fp8 gemm bringup for SM120 & SM121 sources/prs/flashinfer/PR-1610.md
feat: add xqa fp8 mha and fp8 kv cache sources/prs/flashinfer/PR-1769.md
Feature: Add support for L40 FusedMoE in cutlass path sources/prs/flashinfer/PR-1973.md
minor fix for xqa sources/prs/flashinfer/PR-1994.md
feat: add xqa backend and completes NHD/HND coverage for trtllm-gen/xqa backend sources/prs/flashinfer/PR-2001.md
update trtllm cutlass moe sources/prs/flashinfer/PR-2020.md
use scalar for kv_scale in xqa sources/prs/flashinfer/PR-2033.md
feat: add xqa mla backend sources/prs/flashinfer/PR-2053.md
add tensor scale input for xqa sources/prs/flashinfer/PR-2110.md
fix flaky xqa test sources/prs/flashinfer/PR-2126.md
feat: Fused RMSNorm + FP4 Quantization Kernels in CuTe-DSL sources/prs/flashinfer/PR-2233.md
Remove cudaStreamSynchronize from gemm_groupwise_sm120.cuh for CUDA graph compatibility sources/prs/flashinfer/PR-2244.md
feat: Add TRTLLM fmha_v2 library for SM90 attention with Skip-Softmax sources/prs/flashinfer/PR-2446.md
perf: add fp4 GEMM tile configs and streamK scheduler for SM120 sources/prs/flashinfer/PR-2460.md
Support NVFP4 KV cache decode on SM120 sources/prs/flashinfer/PR-2520.md
Implement cutlass_fused_moe mxfp8 sources/prs/flashinfer/PR-2581.md
fix: add SM121 support to SM120 version guards sources/prs/flashinfer/PR-2631.md
fix: reduce smem allocation for tinygemm2 kernel in SM120 sources/prs/flashinfer/PR-2670.md
Support for MXFP4 and NVFP4 group GEMMs on GeForce and Spark sources/prs/flashinfer/PR-2738.md
feat: Add FP4 KV cache quant/dequant kernels sources/prs/flashinfer/PR-2757.md
feat: add MXFP8 GEMM support for SM120 sources/prs/flashinfer/PR-2902.md
[NVIDIA] fix(jit): enable GDC for CUTLASS fused MoE PDL — prevent random crashes on SM12x sources/prs/flashinfer/PR-2913.md
perf: Optimize CUTLASS MoE helper kernels for small-batch decode workloads sources/prs/flashinfer/PR-3014.md
perf: Port TRT-LLM SM120/SM121 FP4 CUTLASS GEMM optimizations. Add PDL sources/prs/flashinfer/PR-3026.md
feat: Add backend="b12x" for mm_fp4 on SM120 sources/prs/flashinfer/PR-3051.md
feat: Add b12x CuTe DSL fused MoE for SM120 sources/prs/flashinfer/PR-3066.md
Integrate CUTLASS Small Tile N Blockscaled GEMMs/Grouped GEMMs for SM120 and SM121 sources/prs/flashinfer/PR-3152.md
fix(sm12x): fix micro-kernel workspace sizing when routed_rows > num_local_experts sources/prs/flashinfer/PR-3191.md
feat(moe): add SM120 W4A16 b12x kernels sources/prs/flashinfer/PR-3271.md
[CUDA][avgpool2d] Fix backward launch bounds again for sm100, sm120 sources/prs/pytorch/PR-150640.md
[CUDA][avgpool2d] Fix backward launch bounds again for sm100, sm120 sources/prs/pytorch/PR-150676.md
support cutlass fp4 kernel in sm120 sources/prs/sglang/PR-11737.md
[Fix] add block size logic for sm120 smem size sources/prs/sglang/PR-14311.md
[Kernel Slimming] Migrate NVFP4 kernels to JIT sources/prs/sglang/PR-19437.md
Support Triton MLA FP8 KV cache sources/prs/sglang/PR-20479.md
CUTLASS FP8 Blockwise GEMM improvement of SM120 sources/prs/sglang/PR-20887.md
CUTLASS NVFP4 GEMM improvement of SM120 sources/prs/sglang/PR-21314.md
[Intel GPU] Enable DeepSeek V4 Inference on XPU sources/prs/sglang/PR-25336.md
[sgl-kernel] feat: Support sm120 cutlass fp8 gemm kernel sources/prs/sglang/PR-9403.md
CUTLASS fp8 blockwise gemm support of sm120 sources/prs/sglang/PR-9969.md
[NVIDIA] Support Cutlass w8a8 FP8 for Blackwell Geforce GPUs (sm120) sources/prs/vllm/PR-17280.md
Support CUTLASS NVFP4 (w4a4) for Blackwell Geforce GPUs (SM120) sources/prs/vllm/PR-21309.md
[Kernel] Add support for block FP8 on SM120 (NVIDIA 5090 and RTX PRO 6000) sources/prs/vllm/PR-22131.md
[Perf] Use upstream CUTLASS for SM90 Block FP8 kernel sources/prs/vllm/PR-23280.md
[Kernel][Quantization] add w4a8 support for marlin kernel sources/prs/vllm/PR-24722.md
Update launch_bounds_utils.h for correct compile on Multiple Cuda Arch - PTXAS out of range Warning sources/prs/vllm/PR-25843.md
[Kernel] Add NVFP4 MoE CUTLASS support for SM120 sources/prs/vllm/PR-29242.md
[Kernel] Add enable_sm120_or_later for SM121 (DGX Spark) CUTLASS support sources/prs/vllm/PR-33517.md
Triton MLA perf fixes sources/prs/vllm/PR-33529.md
[Quantization] add humming quantization kernel sources/prs/vllm/PR-34556.md
[Kernel] Add FP8 KV cache support to Triton MLA decode attention sources/prs/vllm/PR-34597.md
[Bugfix] Gate 256-bit instructions to CUDA 12.9+ sources/prs/vllm/PR-34791.md
[BugFix] Fix fp4 quant kernel on CUDA 12.8 sources/prs/vllm/PR-35210.md
Add nvfp4 support to reshape_and_cache_flash sources/prs/vllm/PR-37332.md
[Kernel] Add MXFP4 W4A4 CUTLASS MoE kernel for SM100 sources/prs/vllm/PR-37463.md
[4/n] Migrate FP4/W4A8 CUTLASS kernels to torch stable ABI sources/prs/vllm/PR-37503.md
[Kernel] Optimize SM120 CUTLASS blockwise FP8 GEMM sources/prs/vllm/PR-37970.md
[Kernel] Add swapAB support for SM120 CUTLASS blockwise FP8 GEMM sources/prs/vllm/PR-38325.md
[NVIDIA] Bugfix NVFP4 DGX Spark and RTX50 sources/prs/vllm/PR-38423.md
[Bugfix] Guard mxfp4_experts_quant bindings on ENABLE_NVFP4_SM100 sources/prs/vllm/PR-40191.md
[Perf] Batch invariance with Cutlass fp8 support, 28.9% E2E latency improvement sources/prs/vllm/PR-40408.md
Cluster Launch Control (CLC) wiki/hardware/clc.md
Programmatic Dependent Launch / Grid Dependency Control wiki/hardware/pdl-gdc.md

sm120a

Page Path
feat: cutlass fp4 gemm bringup for SM120 & SM121 sources/prs/flashinfer/PR-1609.md
feat: cutlass fp8 gemm bringup for SM120 & SM121 sources/prs/flashinfer/PR-1610.md
feat: Add TRTLLM fmha_v2 library for SM90 attention with Skip-Softmax sources/prs/flashinfer/PR-2446.md
support cutlass fp4 kernel in sm120 sources/prs/sglang/PR-11737.md
CUTLASS fp8 blockwise gemm support of sm120 sources/prs/sglang/PR-9969.md

sm120f

Page Path
[NVIDIA] Bugfix NVFP4 DGX Spark and RTX50 sources/prs/vllm/PR-38423.md

sm121

Page Path
PTX ISA Fifth-Generation Tensor Core and CLC Reference sources/docs/nvidia-ptx-isa-sm100.md
[None][feat] GPT-OSS Sm120/Sm121 Support sources/prs/TensorRT-LLM/PR-7937.md
Small Tile N BlockScaled GEMM + Grouped GEMM on SM12x sources/prs/cutlass/PR-3176.md
feat: initial support for SM103, SM110, SM120, SM121 sources/prs/flashinfer/PR-1608.md
feat: cutlass fp4 gemm bringup for SM120 & SM121 sources/prs/flashinfer/PR-1609.md
feat: cutlass fp8 gemm bringup for SM120 & SM121 sources/prs/flashinfer/PR-1610.md
Implement cutlass_fused_moe mxfp8 sources/prs/flashinfer/PR-2581.md
fix: add SM121 support to SM120 version guards sources/prs/flashinfer/PR-2631.md
feat: add MXFP8 GEMM support for SM120 sources/prs/flashinfer/PR-2902.md
[NVIDIA] fix(jit): enable GDC for CUTLASS fused MoE PDL — prevent random crashes on SM12x sources/prs/flashinfer/PR-2913.md
perf: Optimize CUTLASS MoE helper kernels for small-batch decode workloads sources/prs/flashinfer/PR-3014.md
perf: Port TRT-LLM SM120/SM121 FP4 CUTLASS GEMM optimizations. Add PDL sources/prs/flashinfer/PR-3026.md
feat: Add backend="b12x" for mm_fp4 on SM120 sources/prs/flashinfer/PR-3051.md
feat: Add b12x CuTe DSL fused MoE for SM120 sources/prs/flashinfer/PR-3066.md
Integrate CUTLASS Small Tile N Blockscaled GEMMs/Grouped GEMMs for SM120 and SM121 sources/prs/flashinfer/PR-3152.md
feat: DiT layer norm fusions for WAN: flashinfer.diffusion_ops sources/prs/flashinfer/PR-3157.md
fix(sm12x): fix micro-kernel workspace sizing when routed_rows > num_local_experts sources/prs/flashinfer/PR-3191.md
support cutlass fp4 kernel in sm120 sources/prs/sglang/PR-11737.md
CUTLASS NVFP4 GEMM improvement of SM120 sources/prs/sglang/PR-21314.md
[Kernel] Add enable_sm120_or_later for SM121 (DGX Spark) CUTLASS support sources/prs/vllm/PR-33517.md
[Bugfix] Fix DSV3 kernels breaking _C and _moe_C on unsupported arches sources/prs/vllm/PR-35123.md
[NVIDIA] Bugfix NVFP4 DGX Spark and RTX50 sources/prs/vllm/PR-38423.md
Cluster Launch Control (CLC) wiki/hardware/clc.md

sm121a

Page Path
Add SM120 varlen attention support sources/prs/flash-attention/PR-2333.md
feat: cutlass fp4 gemm bringup for SM120 & SM121 sources/prs/flashinfer/PR-1609.md
feat: cutlass fp8 gemm bringup for SM120 & SM121 sources/prs/flashinfer/PR-1610.md

sm75

Page Path
Implement the new tuning API for deterministic (rfa) reduce dispatch sources/prs/cccl/PR-7346.md
simplify dispatch segmented reduce to use latest dispatch and new tunings API sources/prs/cccl/PR-8332.md
fix: put sampling kernel launch into macro sources/prs/flashinfer/PR-1727.md
[Feature] NVFP4 Marlin fallback for non-Blackwell GPUs (SM75+) sources/prs/sglang/PR-19652.md
CUTLASS NVFP4 GEMM improvement of SM120 sources/prs/sglang/PR-21314.md
Support cutlass Int8 gemm sources/prs/sglang/PR-2752.md
support cmake for sgl-kernel sources/prs/sglang/PR-4706.md
[CUDA] Add native SM75 MMA GEMM support for FP16, INT8 and INT4 sources/prs/tilelang/PR-2198.md
[TIR][IR] Update to use tirx sources/prs/tilelang/PR-2216.md
[Kernel][Quantization][MoE] add marlin kernel support for turing (sm75) sources/prs/vllm/PR-29901.md

sm80

Page Path
cuTile Python Documentation sources/docs/cutile-python-dsl.md
CUTLASS CuTe DSL Documentation sources/docs/cutlass-cute-dsl.md
Native Sparse Attention (NSA) sources/docs/nsa.md
[#12716][feat] Fused cross-head QK Norm + RoPE kernel for WAN sources/prs/TensorRT-LLM/PR-13052.md
[None][feat] GPT-OSS Sm120/Sm121 Support sources/prs/TensorRT-LLM/PR-7937.md
Implement the new tuning API for DeviceRleDispatch sources/prs/cccl/PR-7669.md
Optimized Device-to-Device Tensor Copy (cudax) sources/prs/cccl/PR-7823.md
Implement the new tuning API for DispatchSegmentedSort sources/prs/cccl/PR-7874.md
Implement the new tuning API for DispatchSelectIf sources/prs/cccl/PR-8311.md
Flash MLA support sources/prs/cutlass/PR-2130.md
Support hdimQK != hdimV backward sources/prs/flash-attention/PR-1604.md
feat: Adding varlen support to cute-dsl sm80 bwd sources/prs/flash-attention/PR-1934.md
Add SM120 varlen attention support sources/prs/flash-attention/PR-2333.md
Add CUTLASS fused moe kernels from TensorRT-LLM. sources/prs/flashinfer/PR-1113.md
[Feature] Support PDL for batch Prefill and Decode sources/prs/flashinfer/PR-1117.md
feat: trtllm-gen fp8 moe kernels sources/prs/flashinfer/PR-1212.md
Update cutlass fp4 moe kernels sources/prs/flashinfer/PR-1294.md
add cutlass backend for mm_fp4 sources/prs/flashinfer/PR-1296.md
feat: integrate xqa attention backend sources/prs/flashinfer/PR-1503.md
perf&bugfix: skip kv-tile computation out of sliding window in FA2; fix __syncthreads in mergestate sources/prs/flashinfer/PR-1661.md
update trtllm cutlass moe sources/prs/flashinfer/PR-2020.md
[Feature] Support batch prefill for POD Attention sources/prs/flashinfer/PR-2079.md
[bugfix] Fix FilteredTopK overflow correctness sources/prs/flashinfer/PR-2605.md
feat: implement deterministic topk sources/prs/flashinfer/PR-2661.md
feat: Add FP4 KV cache quant/dequant kernels sources/prs/flashinfer/PR-2757.md
Support NVFP4 KV for prefill and batch attention kernels sources/prs/flashinfer/PR-3097.md
feat: DiT layer norm fusions for WAN: flashinfer.diffusion_ops sources/prs/flashinfer/PR-3157.md
checkpointing_ssu kernel: fused replay + conditional state-write for Mamba2 sources/prs/flashinfer/PR-3324.md
perf: memory efficient deepseek mla fused page-attention kernel sources/prs/flashinfer/PR-804.md
feat: unlocking MLA for A100 sources/prs/flashinfer/PR-812.md
feat: unlock MLA attention for sm89 (L40/L40s/4090) sources/prs/flashinfer/PR-814.md
perf: MLA decode kernel implemented by CuTe targeted to SM80 sources/prs/flashinfer/PR-844.md
Add POD-Attention to FlashInfer sources/prs/flashinfer/PR-858.md
[ATen][CUDA] Optimize 128 bit vectorization sources/prs/pytorch/PR-152967.md
[Feature] NVFP4 Marlin fallback for non-Blackwell GPUs (SM75+) sources/prs/sglang/PR-19652.md
Support cutlass Int8 gemm sources/prs/sglang/PR-2752.md
Feature DeepSeek V3/R1 INT8 Quantization (block-wise) sources/prs/sglang/PR-3730.md
[Feature] DeepSeek V3/R1 INT8 Quantization (channel-wise) sources/prs/sglang/PR-3888.md
support cmake for sgl-kernel sources/prs/sglang/PR-4706.md
[Backend] Refactor gemm_sp sources/prs/tilelang/PR-2048.md
feat: auto-vectorize bf16/fp16 reduce with packed add2 intrinsics sources/prs/tilelang/PR-2112.md
[TIR][IR] Update to use tirx sources/prs/tilelang/PR-2216.md
[Kernel] add triton fused moe kernel for gptq/awq sources/prs/vllm/PR-12185.md
[Misc][Kernel]: Add GPTQAllSpark Quantization sources/prs/vllm/PR-12931.md
[Kernel] moe wna16 cuda kernel sources/prs/vllm/PR-13321.md
[Kernel] moe wna16 marlin kernel sources/prs/vllm/PR-14447.md
[Kernel] some optimizations for dense marlin and moe marlin sources/prs/vllm/PR-16850.md
[BugFix] FA2 MLA Accuracy Issue sources/prs/vllm/PR-18807.md
[Kernel][Quantization] add w4a8 support for marlin kernel sources/prs/vllm/PR-24722.md
[Quantization] add humming quantization kernel sources/prs/vllm/PR-34556.md
[Bugfix] Gate 256-bit instructions to CUDA 12.9+ sources/prs/vllm/PR-34791.md
mbarrier (Memory Barrier Primitives) wiki/hardware/mbarrier.md
Native Sparse Attention (NSA) wiki/kernels/nsa.md

sm86

Page Path
Implement the new tuning API for deterministic (rfa) reduce dispatch sources/prs/cccl/PR-7346.md
Implement the new tuning API for DeviceRleDispatch sources/prs/cccl/PR-7669.md
Implement the new tuning API for DispatchSegmentedSort sources/prs/cccl/PR-7874.md
Implement the new tuning API for DispatchSelectIf sources/prs/cccl/PR-8311.md
simplify dispatch segmented reduce to use latest dispatch and new tunings API sources/prs/cccl/PR-8332.md
feat: integrate xqa attention backend sources/prs/flashinfer/PR-1503.md
[Feature] NVFP4 Marlin fallback for non-Blackwell GPUs (SM75+) sources/prs/sglang/PR-19652.md

sm87

Page Path
feat: integrate xqa attention backend sources/prs/flashinfer/PR-1503.md

sm89

Page Path
[None][feat] GPT-OSS Sm120/Sm121 Support sources/prs/TensorRT-LLM/PR-7937.md
simplify dispatch segmented reduce to use latest dispatch and new tunings API sources/prs/cccl/PR-8332.md
support fp16 accmulator for sm89 fp8 mma sources/prs/cutlass/PR-2378.md
feat: integrate xqa attention backend sources/prs/flashinfer/PR-1503.md
feat:enable fp8 blockscale moe for fused cultass for sm90 sources/prs/flashinfer/PR-1819.md
Feature: Add support for L40 FusedMoE in cutlass path sources/prs/flashinfer/PR-1973.md
refactor: refactoring cuda code to cute-dsl (part 1) sources/prs/flashinfer/PR-2428.md
feat: DiT layer norm fusions for WAN: flashinfer.diffusion_ops sources/prs/flashinfer/PR-3157.md
checkpointing_ssu kernel: fused replay + conditional state-write for Mamba2 sources/prs/flashinfer/PR-3324.md
feat: unlock MLA attention for sm89 (L40/L40s/4090) sources/prs/flashinfer/PR-814.md
bugfix: bugfix on sm89 MLA sources/prs/flashinfer/PR-821.md
Optimize cutlass int8 gemm kernel for large M on SM89 Ada GPU sources/prs/sglang/PR-10714.md
[FIX] Correct JIT kernel compilation on newer GPUs with outdated driver metadata. sources/prs/sglang/PR-18496.md
Support cutlass Int8 gemm sources/prs/sglang/PR-2752.md
support w8a8 fp8 kernel with CUTLASS sources/prs/sglang/PR-3047.md
Apply sgl w8a8 fp8 kernel sources/prs/sglang/PR-3148.md
support cmake for sgl-kernel sources/prs/sglang/PR-4706.md
[sgl-kernel] feat: Support sm120 cutlass fp8 gemm kernel sources/prs/sglang/PR-9403.md
[Backend] Refactor gemm_sp sources/prs/tilelang/PR-2048.md
permute/unpermute kernel for moe optimization sources/prs/vllm/PR-14568.md
[Kernel][Quantization] add w4a8 support for marlin kernel sources/prs/vllm/PR-24722.md
[Bugfix] Gate 256-bit instructions to CUDA 12.9+ sources/prs/vllm/PR-34791.md
[NVIDIA] Bugfix NVFP4 DGX Spark and RTX50 sources/prs/vllm/PR-38423.md
[Perf] Batch invariance with Cutlass fp8 support, 28.9% E2E latency improvement sources/prs/vllm/PR-40408.md

sm90

Page Path
Colfax Article Source Kernels sources/blogs/colfax-article-source-kernels.md
Colfax CUTLASS Kernels sources/blogs/colfax-cutlass-kernels.md
DeepGEMM tensor-core kernel library sources/blogs/deepgemm.md
FlashMLA upstream README sources/blogs/flashmla.md
NVIDIA Developer Code Samples sources/blogs/nvidia-code-samples.md
NVIDIA Qwen3-Next Architecture Announcement sources/blogs/qwen3-next-architecture.md
simveit effective_transpose sources/blogs/simveit-effective-transpose.md
simveit load_and_store sources/blogs/simveit-load-and-store.md
DeepSeek-V3.2-Exp in vLLM: Fine-Grained Sparse Attention in Action sources/blogs/vllm-deepseek-v3-sparse-attention.md
CUDA Programming Guide: Programmatic Dependent Launch sources/docs/cuda-programming-guide-pdl.md
cuTile Python Documentation sources/docs/cutile-python-dsl.md
CUTLASS CuTe DSL Documentation sources/docs/cutlass-cute-dsl.md
K-Search: LLM Kernel Generation via Co-Evolving Intrinsic World Model sources/docs/k-search-kernel-generation.md
Tiled Flash Linear Attention (TFLA) sources/docs/tfla.md
Fix performance issue of m-grouped contiguous GEMMs. sources/prs/DeepGEMM/PR-168.md
Fix multicast bug and optimize masked GEMM sources/prs/DeepGEMM/PR-193.md
[Public release 26/04] Introducing Mega MoE, FP4 Indexer and other features/fixes sources/prs/DeepGEMM/PR-304.md
Sync nv_dev with upstream #316 (Mega MoE optimizations & benchmarks) sources/prs/DeepGEMM/PR-328.md
[TRTLLM-10022][feat] Add hopper xqa decode support for skip softmax attention sources/prs/TensorRT-LLM/PR-10264.md
[None][feat] Add support for expert_number<=2048 and K<=32 sources/prs/TensorRT-LLM/PR-11510.md
[TRTLLM-11092][feat] add support for visual gen FA4 attention backend sources/prs/TensorRT-LLM/PR-11697.md
[None][perf] add Dynamic SMEM block routing in MOE sources/prs/TensorRT-LLM/PR-12456.md
[None][feat] Add triton paged attention for AutoDeploy sources/prs/TensorRT-LLM/PR-12642.md
[#13580][fix] AutoDeploy: Support Gemma3n/4 E2B variants sources/prs/TensorRT-LLM/PR-13630.md
[None][feat] Indexer topk opt sources/prs/TensorRT-LLM/PR-13811.md
[None][feat] Update the logic of FMHA JIT path sources/prs/TensorRT-LLM/PR-14291.md
[TRTLLM-8637][feat] Optimize the routing kernel for DeepseekV3 (MoE CUTLASS backend); Add support for 384 experts (MoE TRTLLM backend) sources/prs/TensorRT-LLM/PR-7761.md
[None][feat] GPT-OSS Sm120/Sm121 Support sources/prs/TensorRT-LLM/PR-7937.md
[TRTLLM-8535][feat] Support DeepSeek V3.2 with FP8 + BF16 KV cache/NVFP4 + BF16 KV cache sources/prs/TensorRT-LLM/PR-8405.md
[None][fix] Fix the performance issue of FP8 blockwise grouped GEMM when using attention DP sources/prs/TensorRT-LLM/PR-8501.md
[https://nvbugs/5726962][feat] Apply fusion for W4AFP8_AWQ MoE sources/prs/TensorRT-LLM/PR-9838.md
[None][feat] Adding torch ext API for FusedAddRMSNormQuant kernel sources/prs/TensorRT-LLM/PR-9905.md
[TRTLLM-9493][feat] Add helixPostProcessNative kernel for cp_dim=2 sources/prs/TensorRT-LLM/PR-9924.md
fix thread-reduce performance regression sources/prs/cccl/PR-2944.md
Fix scan / sm90 perf regression sources/prs/cccl/PR-3236.md
Add nondeterministic reduce that uses atomics sources/prs/cccl/PR-4961.md
Combine block_reduce_warp_reduction_nondeterministic.cuh specialization with original deterministic one sources/prs/cccl/PR-5408.md
Use integer promotion for warp_reduce sources/prs/cccl/PR-6819.md
Implement the new tuning API for deterministic (rfa) reduce dispatch sources/prs/cccl/PR-7346.md
Implement the new tuning API for DeviceRleDispatch sources/prs/cccl/PR-7669.md
Optimized Device-to-Device Tensor Copy (cudax) sources/prs/cccl/PR-7823.md
Implement the new tuning API for DispatchTopK sources/prs/cccl/PR-7928.md
Implement the new tuning API for DispatchSelectIf sources/prs/cccl/PR-8311.md
simplify dispatch segmented reduce to use latest dispatch and new tunings API sources/prs/cccl/PR-8332.md
Improve sm90 mixed dtype kernel sources/prs/cutlass/PR-1883.md
[EVT] Add support for Row/Col broadcast PtrArray sources/prs/cutlass/PR-2033.md
Groupwise scaling along M for FP8 gemm sources/prs/cutlass/PR-2037.md
Hopper Grouped GEMM support for FP8 Accum sources/prs/cutlass/PR-2123.md
Blockwise and Groupwise GEMM for Blackwell and Improvements for Hopper sources/prs/cutlass/PR-2139.md
Set EpiTile correctly when TileN is not divisible by 32 sources/prs/cutlass/PR-2220.md
hopper-blockwise-generalization-optimization sources/prs/cutlass/PR-2270.md
DistGEMM bug fixes sources/prs/cutlass/PR-2713.md
Support PDL for SM90 Array TMA GEMM sources/prs/cutlass/PR-2719.md
[Bug Fix]Bypass launch grids for SM120 Kernel with SM90 Mainloop & SM100 TileScheduler sources/prs/cutlass/PR-2865.md
[Bug Fix]Set NumSplitsM to 1 when TileShapeM < 128 in sm90 fp8 blockwise scaling CollectiveMma sources/prs/cutlass/PR-2965.md
Small Tile N BlockScaled GEMM + Grouped GEMM on SM12x sources/prs/cutlass/PR-3176.md
Add var-seq-len to FA3 fp16 / bf16 fwd sources/prs/flash-attention/PR-1072.md
Fp8 kernel with "in-kernel" transpose of V in producer sources/prs/flash-attention/PR-1100.md
FA3 FP8 qkv descales + restore max offset for h128 causal + added sync for producer WG sources/prs/flash-attention/PR-1173.md
FA3 kvcache + split kv + gqa parallelization sources/prs/flash-attention/PR-1236.md
Paged Attention support for FA3 sources/prs/flash-attention/PR-1268.md
FA3 paged attention: Readiness for Cutlass 3.6 / default value for block_table sources/prs/flash-attention/PR-1331.md
feat: Adding varlen support to cute-dsl sm80 bwd sources/prs/flash-attention/PR-1934.md
[Cute,Fwd,Sm100] Implement SplitKV sources/prs/flash-attention/PR-1940.md
[Cute,Fwd] Extend score_mod to variable sequence length sources/prs/flash-attention/PR-2043.md
[CUTE][SM90]Enable pack-gqa with broadcasted maskmods sources/prs/flash-attention/PR-2145.md
[Cute,Fwd,Sm100] support irregular qhead / kvhead ratios sources/prs/flash-attention/PR-2186.md
[Fwd,Sm90] Add paged KV attention support (tma and cp.async) sources/prs/flash-attention/PR-2360.md
[Cute,Sm100,Fwd] add MLA 64/512 with topk sparsity for MQA 128 heads sources/prs/flash-attention/PR-2441.md
add multi-item scoring sources/prs/flashinfer/PR-1015.md
fix: add zero init for KV tiled copy sources/prs/flashinfer/PR-1029.md
feat: add functional per-head FP8 quantization for FA3 sources/prs/flashinfer/PR-1033.md
bugfix: fix fp8 attention kernels aot compilation issue sources/prs/flashinfer/PR-1087.md
Add CUTLASS fused moe kernels from TensorRT-LLM. sources/prs/flashinfer/PR-1113.md
[Feature] Support PDL for batch Prefill and Decode sources/prs/flashinfer/PR-1117.md
[feat] add unified batch attention w/ correctness tests. sources/prs/flashinfer/PR-1137.md
Fix FA2 and FA3 multi-item scoring and cuda illegal memory access error sources/prs/flashinfer/PR-1140.md
feat: Fused temperature online softmax kernel sources/prs/flashinfer/PR-1153.md
[feat] optimize persistent batch attention perf. sources/prs/flashinfer/PR-1200.md
feat: trtllm-gen fp8 moe kernels sources/prs/flashinfer/PR-1212.md
Reduce the JIT compilation time of gen_gemm_sm100_module sources/prs/flashinfer/PR-1251.md
add cutlass backend for mm_fp4 sources/prs/flashinfer/PR-1296.md
gpt-oss: Add MXFP8 x MXFP4 CUTLASS MOE for SM100 and BF16 x MXFP4 CUTLASS for SM90 + SwigluBias Activation sources/prs/flashinfer/PR-1396.md
feat: integrate xqa attention backend sources/prs/flashinfer/PR-1503.md
feat: Integrate TRTLLM varlen kernel for deepseek R1 prefill sources/prs/flashinfer/PR-1537.md
fix: separate out fp4 lib into sm90 and sm100 versions, add oob checking in fused moe sources/prs/flashinfer/PR-1565.md
bugfix: collect all modules to aot sources/prs/flashinfer/PR-1622.md
perf&bugfix: skip kv-tile computation out of sliding window in FA2; fix __syncthreads in mergestate sources/prs/flashinfer/PR-1661.md
feat: add xqa fp8 mha and fp8 kv cache sources/prs/flashinfer/PR-1769.md
feat:enable fp8 blockscale moe for fused cultass for sm90 sources/prs/flashinfer/PR-1819.md
Update the routing for TRTLLMGEN to support kimi k2 and qwen sources/prs/flashinfer/PR-1831.md
Feature: Support Relu2 activation in fused MoE sources/prs/flashinfer/PR-1954.md
feat: enable deepgemm jit for fp8 block-scale on SM90 sources/prs/flashinfer/PR-1969.md
minor fix for xqa sources/prs/flashinfer/PR-1994.md
feat: add xqa backend and completes NHD/HND coverage for trtllm-gen/xqa backend sources/prs/flashinfer/PR-2001.md
update trtllm cutlass moe sources/prs/flashinfer/PR-2020.md
[NVIDIA] Thor & Spark Support sources/prs/flashinfer/PR-2028.md
use scalar for kv_scale in xqa sources/prs/flashinfer/PR-2033.md
feat: Add flashinfer.rope.rope_quantize_fp8_append_paged_kv_cache (fused RoPE + Q + KV cache, supports MLA/GQA/MHA) sources/prs/flashinfer/PR-2037.md
[Feature] Support batch prefill for POD Attention sources/prs/flashinfer/PR-2079.md
enable xqa fp8 output sources/prs/flashinfer/PR-2081.md
enable xqa speculative decoding sources/prs/flashinfer/PR-2105.md
add tensor scale input for xqa sources/prs/flashinfer/PR-2110.md
refactor: update fa3 codebase and fix hopper unittest [part 1] sources/prs/flashinfer/PR-2111.md
feature: make the LSE returned by MLA support base 2 or e #2113 sources/prs/flashinfer/PR-2114.md
perf: bunch of features and optimizations for top-k (sampling + sparse attention) sources/prs/flashinfer/PR-2119.md
fix flaky xqa test sources/prs/flashinfer/PR-2126.md
make DeepGEMM swapAB available for linear gemm SM90 sources/prs/flashinfer/PR-2131.md
feat: TRTLLM FMHAv2 backend for ctx attention sources/prs/flashinfer/PR-2142.md
feat: RMSNorm/Fused RMSNorm + FP8 Quantization kernels sources/prs/flashinfer/PR-2243.md
feat: add GDN Attention sources/prs/flashinfer/PR-2276.md
feat: [Qwen3-Next] Add Cute DSL GDN decode kernel and tests sources/prs/flashinfer/PR-2370.md
A Blackwell-optimized version of selective_state_update (mamba) sources/prs/flashinfer/PR-2387.md
Remove cudaMalloc/Free in GDN prefill kernel sources/prs/flashinfer/PR-2415.md
refactor: reduce hopper's gdn prefill compilation time and fix docstring. sources/prs/flashinfer/PR-2422.md
refactor: refactoring cuda code to cute-dsl (part 1) sources/prs/flashinfer/PR-2428.md
feat: Add TRTLLM fmha_v2 library for SM90 attention with Skip-Softmax sources/prs/flashinfer/PR-2446.md
fix: W4A8 autotune crash in cutlass_fused_moe profiler workspace sources/prs/flashinfer/PR-2564.md
Implement cutlass_fused_moe mxfp8 sources/prs/flashinfer/PR-2581.md
[feat] trtllm-gen mxfp8 gemm sources/prs/flashinfer/PR-2653.md
fix: reduce smem allocation for tinygemm2 kernel in SM120 sources/prs/flashinfer/PR-2670.md
[gdn] support non-contiguous state for decoding sources/prs/flashinfer/PR-2727.md
[feat] Add 2048 experts and 32 Top K sources/prs/flashinfer/PR-2744.md
[feat] Add air top-p algorithm sources/prs/flashinfer/PR-2752.md
feat: Add FP4 KV cache quant/dequant kernels sources/prs/flashinfer/PR-2757.md
perf: Performance tune cute dsl RMSNorm variants sources/prs/flashinfer/PR-2777.md
feat(gdn): state checkpointing in chunk_gated_delta_rule sources/prs/flashinfer/PR-2908.md
[NVIDIA] fix(jit): enable GDC for CUTLASS fused MoE PDL — prevent random crashes on SM12x sources/prs/flashinfer/PR-2913.md
fix: tinygemm2 hang issue due to barrier sync sources/prs/flashinfer/PR-2996.md
feat: add PDL support to rmsnorm_fp4quant and add_rmsnorm_fp4quant CuTe DSL kernels sources/prs/flashinfer/PR-3008.md
feat: DiT layer norm fusions for WAN: flashinfer.diffusion_ops sources/prs/flashinfer/PR-3157.md
perf: optimize per-token nvfp4 quantization kernel. sources/prs/flashinfer/PR-3237.md
Ameyn/gdn bf16 dispatcher and 4d pool sources/prs/flashinfer/PR-3268.md
fix(fmha_v2): fix FP8 V-scratch pipeline and varlen scheduler on SM90 sources/prs/flashinfer/PR-3276.md
feat: support deepseek prefill attention shape sources/prs/flashinfer/PR-765.md
feat: apply sm_scale at logits instead of q in FA2 template sources/prs/flashinfer/PR-801.md
perf: memory efficient deepseek mla fused page-attention kernel sources/prs/flashinfer/PR-804.md
feat: unlock MLA attention for sm89 (L40/L40s/4090) sources/prs/flashinfer/PR-814.md
Add POD-Attention to FlashInfer sources/prs/flashinfer/PR-858.md
perf: dynamic split-k for MLA sources/prs/flashinfer/PR-863.md
Naive Support for Hopper FP8 Prefill Kernel with Per-Head Quantization sources/prs/flashinfer/PR-869.md
perf: FlashAttention-3 style MLA PageAttention sources/prs/flashinfer/PR-887.md
perf: fix MLA split-k performance bug sources/prs/flashinfer/PR-898.md
feat: flashinfer intra-kernel profiler sources/prs/flashinfer/PR-913.md
perf: Fix python API overhead when CUDAGraph is not enabled sources/prs/flashinfer/PR-969.md
Update CUTLASS. Refine KernelSchedule for fp8 (grouped) gemm. sources/prs/sglang/PR-10491.md
[sgl-kernel][1/N]Support Expert Specialization Grouped GEMM sources/prs/sglang/PR-11432.md
[DeepseekV32] Enable flashmla_prefill kernel with fp8 kvcache sources/prs/sglang/PR-11655.md
[sgl-kernel][4/N]Support Expert Specialization Grouped GEMM sources/prs/sglang/PR-12080.md
[DeepSeek v3.2] opt Context Parallelism: support fused moe, multi batch and fp8 kvcache sources/prs/sglang/PR-13959.md
[bug fix] fix ima with get_mla_kv_buffer_kernel overflow sources/prs/sglang/PR-14224.md
[NVIDIA] upstream FA4 sources/prs/sglang/PR-15182.md
[sgl-kernel][6/7]Support Expert Specialization Grouped GEMM sources/prs/sglang/PR-15471.md
[jit-kernel] Add CuTe DSL GDN Decode Kernel sources/prs/sglang/PR-15631.md
Add SwapAB Optimization for triton fused_moe_kernel on SM90. sources/prs/sglang/PR-15712.md
[Feature] JIT Fused QK norm + qk norm clean up sources/prs/sglang/PR-15835.md
[Rework] Add SwapAB Optimization for triton fused_moe_kernel on SM90. sources/prs/sglang/PR-16723.md
Move fa4 from sgl-kernel to jit kernel sources/prs/sglang/PR-17353.md
Add mxfp8 support for online quantization, Triton dense linear, and CUTLASS MoE sources/prs/sglang/PR-17449.md
[Diffsuion & JIT_kernel] QKNorm cross heads kernel sources/prs/sglang/PR-18073.md
feat: add FA4 SM90 paged KV decode support & update attention docs sources/prs/sglang/PR-18442.md
[FIX] Correct JIT kernel compilation on newer GPUs with outdated driver metadata. sources/prs/sglang/PR-18496.md
[diffusion] Diffusion norm fusion for z-image sources/prs/sglang/PR-18762.md
[Feature] NVFP4 Marlin fallback for non-Blackwell GPUs (SM75+) sources/prs/sglang/PR-19652.md
[Feature][JIT Kernel] Fused TP QK norm For Minimax sources/prs/sglang/PR-20673.md
[Qwen3.5] Fuse split/reshape/cat ops in GDN projection with Triton kernel sources/prs/sglang/PR-21019.md
[KDA] Support CuTeDSL KDA decode kernel sources/prs/sglang/PR-21203.md
CUTLASS NVFP4 GEMM improvement of SM120 sources/prs/sglang/PR-21314.md
[GDN] Fuse GDN kkt + solve_tril into one kernel sources/prs/sglang/PR-21411.md
[Feature] JIT rmsnorm update (with claude) sources/prs/sglang/PR-21834.md
diffusion: add HunyuanVideo GroupNorm+SiLU fast path sources/prs/sglang/PR-22814.md
[Fix/Kernel] Add JIT rmsnorm_hf kernel to fix transformers backend MMLU accuracy regression sources/prs/sglang/PR-22931.md
Deepseek_v4 support w4(mxfp4)a16 on hopper sources/prs/sglang/PR-23686.md
Optimize large GroupNorm SiLU apply sources/prs/sglang/PR-23938.md
[feat] Init true on policy with qwen_dense sources/prs/sglang/PR-23961.md
Enable PDL for various kernels in DSV32/GLM5 sources/prs/sglang/PR-23965.md
[diffusion] Fuse LTX2 split rotary embedding sources/prs/sglang/PR-24411.md
Port MXFP4 Marlin MoE support to JIT kernel path sources/prs/sglang/PR-24490.md
[Gemma4] Optimize Gemm4 with fused Q/K/V RMSNorm + per-expert FP8 ckpt loader sources/prs/sglang/PR-24696.md
[codex] Optimize hidden-size 512 RMSNorm dispatch sources/prs/sglang/PR-24710.md
[rebase]Deepseek_v4 support w4(mxfp4)a16 on hopper sources/prs/sglang/PR-24986.md
[Intel GPU] Enable DeepSeek V4 Inference on XPU sources/prs/sglang/PR-25336.md
[fp8] SM90 swap-AB scaled_mm dispatch (~1.16x kernel geomean, +5.8-18.5% end-to-end) sources/prs/sglang/PR-25532.md
[Refactor] Rename NSA → DSA: user-facing aliases, file/class/import rename sources/prs/sglang/PR-25821.md
Support cutlass Int8 gemm sources/prs/sglang/PR-2752.md
Support sm90 Int8 gemm sources/prs/sglang/PR-3035.md
support w8a8 fp8 kernel with CUTLASS sources/prs/sglang/PR-3047.md
Apply sgl w8a8 fp8 kernel sources/prs/sglang/PR-3148.md
add tensorrt_llm common and cutlass_extensions as 3rdparty sources/prs/sglang/PR-3216.md
support blockwise fp8 matmul kernel sources/prs/sglang/PR-3267.md
integrate blockwise fp8 kernel sources/prs/sglang/PR-3529.md
support cmake for sgl-kernel sources/prs/sglang/PR-4706.md
reduce moe_align_block_size_kernel small batch mode overhead sources/prs/sglang/PR-5086.md
Support MHA with chunked prefix cache for DeepSeek chunked prefill sources/prs/sglang/PR-5113.md
[perf] introduce deep gemm group_gemm_masked as bmm sources/prs/sglang/PR-5432.md
[2/2] Add python wrapper for CUTLASS FP8 Blockscale MoE Kernel. sources/prs/sglang/PR-5694.md
cutlass 3.9 supported to improve fp8_blockwise_gemm sources/prs/sglang/PR-5820.md
chore: upgrade cutlass 3.9.2 sources/prs/sglang/PR-6004.md
Set num_fused_shared_experts as num_shared_experts when shared_experts fusion is not disabled sources/prs/sglang/PR-6736.md
feat: integrate deepgemm into EPMoE sources/prs/sglang/PR-6821.md
fix ep_moe_reorder kernel bugs sources/prs/sglang/PR-6858.md
chore: upgrade flashinfer v0.2.6.post1 jit sources/prs/sglang/PR-6958.md
Add CUTLASS FP8 Blockscale MoE kernel for Hopper architecture sources/prs/sglang/PR-7278.md
Add dsv3 router gemm kernel sources/prs/sglang/PR-7627.md
Add dsv3 fused a gemm to sgl-kernel sources/prs/sglang/PR-7630.md
feat: support DeepSeek-R1-W4AFP8 model with ep-moe mode sources/prs/sglang/PR-7762.md
[1/n]: add cutlass W4A8 moe kernel for hopper architecture sources/prs/sglang/PR-7772.md
[feat] Support tp mode for DeepSeek-R1-W4AFP8 sources/prs/sglang/PR-8118.md
optimize: reduce shulffle and quantization overhead in cutlass_moe sm90 sources/prs/sglang/PR-8962.md
[fix]: fix cutlass moe ut and and Opt H20 cutlass groupGemm performance sources/prs/sglang/PR-9272.md
Update CUTLASS 4.2 & Enable K-Major Scale Factor for SM90 FP8 Blockwise Group GEMM sources/prs/sglang/PR-9559.md
[Feature] Support cp.reduce.async.bulk.tensor sources/prs/tilelang/PR-1667.md
[Feature] Support cluster launch, query, synchronization and barrier operations sources/prs/tilelang/PR-1874.md
[Feature] Add Producer-Consumer Warp Specialization and T.tma_copy() API sources/prs/tilelang/PR-1909.md
[Backend] Refactor gemm_sp sources/prs/tilelang/PR-2048.md
[Kernel]: Cutlass 2:4 Sparsity + FP8/Int8 Quant Support sources/prs/vllm/PR-10995.md
[Kernel] Update cutlass_scaled_mm to support 2d group (blockwise) scaling sources/prs/vllm/PR-11868.md
[Perf] Mem align KV caches for CUDA devices (MLA perf improvement) sources/prs/vllm/PR-12676.md
[Kernel]Add streamK for block-quantized CUTLASS kernels sources/prs/vllm/PR-12978.md
[Kernel] moe wna16 cuda kernel sources/prs/vllm/PR-13321.md
[Kernel] CUTLASS grouped gemm fp8 MoE kernel sources/prs/vllm/PR-13972.md
Add cutlass support for blackwell fp8 blockwise gemm sources/prs/vllm/PR-14383.md
[Kernel] GGUF MoE kernel sources/prs/vllm/PR-14613.md
[Kernel] Unified Triton kernel that doesn't distinguish between prefill + decode sources/prs/vllm/PR-16828.md
Enable V1 for Hybrid SSM/Attention Models sources/prs/vllm/PR-20016.md
[Feature] Integrate SM100 DeepGEMM support sources/prs/vllm/PR-20087.md
[Kernel] Optimize Prefill Attention in Unified Triton Attention Kernel sources/prs/vllm/PR-20308.md
[Kernel] SM90 CUTLASS FP8 GEMM: add support for swap AB + kernel tuning sources/prs/vllm/PR-20396.md
[feat]: add SM100 support for cutlass FP8 groupGEMM sources/prs/vllm/PR-20447.md
[Perf] Add swap_ab to SM90 FP8 non-block CUTLASS moe grouped gemm sources/prs/vllm/PR-20911.md
[Perf] Cuda Kernel for Per Token Group Quant sources/prs/vllm/PR-21083.md
[Kernel] Enable Hybrid Model Support in Triton Unified Attention Kernel sources/prs/vllm/PR-21197.md
[kernel] Support W4A8 on Hopper sources/prs/vllm/PR-23198.md
[Perf] Small optimizations for silu_mul_fp8_quant_deep_gemm sources/prs/vllm/PR-23265.md
[Kernel] Add fused grouped_topk kernel for MoE sources/prs/vllm/PR-23274.md
[Perf] Use upstream CUTLASS for SM90 Block FP8 kernel sources/prs/vllm/PR-23280.md
[Bugfix] Fixing division by zero in triton_attn if query_heads/kv_heads > 16 sources/prs/vllm/PR-23424.md
[Kernel][Quantization] add w4a8 support for marlin kernel sources/prs/vllm/PR-24722.md
[Perf] SM100 - add swap AB optimization to CUTLASS FP8 GEMM sources/prs/vllm/PR-27284.md
[Performance] Fused blockwise quant RMS norm sources/prs/vllm/PR-27883.md
[Kernel] Optimize rms_norm kernel sources/prs/vllm/PR-27931.md
[SpecDecode] Simplified alternative padded-speculation acceptance rate fix sources/prs/vllm/PR-29845.md
[Kernel] Add topk_sigmoid kernel sources/prs/vllm/PR-31246.md
Add TMA support to fused_moe_lora kernel sources/prs/vllm/PR-32195.md
[Performance] Tune Mamba selective scan kernel for B200 sources/prs/vllm/PR-32873.md
[Kernel] Optimize grouped topk kernel sources/prs/vllm/PR-34206.md
[ModelBash][DSV3] Add TRTLLM DSV3 Router GEMM kernel (6% B1 Speedup) sources/prs/vllm/PR-34302.md
[Quantization] add humming quantization kernel sources/prs/vllm/PR-34556.md
[Model Bash] DeepSeek R1 BF16 Min Latency QKV A GEMM (0.5% E2E Speedup) sources/prs/vllm/PR-34758.md
[Bugfix] Gate 256-bit instructions to CUDA 12.9+ sources/prs/vllm/PR-34791.md
[Kernel] Add fused_sigmoid_gating_delta_rule_update kernel for Qwen3 Next sources/prs/vllm/PR-35777.md
[Kernel] Fuse FP8 output quantization into merge_attn_states sources/prs/vllm/PR-36518.md
[Kernel] Add gpt-oss Router GEMM kernel sources/prs/vllm/PR-37205.md
[4/n] Migrate FP4/W4A8 CUTLASS kernels to torch stable ABI sources/prs/vllm/PR-37503.md
[NVIDIA] Bugfix NVFP4 DGX Spark and RTX50 sources/prs/vllm/PR-38423.md
[Perf] Batch KV cache swap copies via cuMemcpyBatchAsync sources/prs/vllm/PR-38460.md
[Perf][GDN] Align TMA usage with upstream FLA sources/prs/vllm/PR-38981.md
fix: clamp NaN/Inf in topk_softmax to prevent duplicate expert IDs sources/prs/vllm/PR-39391.md
[Perf] Batch invariance with Cutlass fp8 support, 28.9% E2E latency improvement sources/prs/vllm/PR-40408.md
[Perf] Add do_not_specialize in fused FP8 RoPE kernel sources/prs/vllm/PR-42849.md
[Kernel] (2/N) Machete - Integrate into CompressedTensorsWNA16 and GPTQMarlin sources/prs/vllm/PR-7701.md
mbarrier (Memory Barrier Primitives) wiki/hardware/mbarrier.md
Programmatic Dependent Launch / Grid Dependency Control wiki/hardware/pdl-gdc.md
Tensor Memory Accelerator (TMA) wiki/hardware/tma.md
DeepGEMM — runtime-JIT tensor-core kernels wiki/kernels/deepgemm.md
FlashMLA attention kernels wiki/kernels/flashmla.md
FP8 block-scale GEMM wiki/kernels/fp8-block-scale-gemm.md
Fused MoE — Expert GEMM and Adjacent Operations wiki/kernels/fused-moe.md
Gated Delta Network kernels wiki/kernels/gated-delta-net.md
Gated Dual GEMM (Gate-Up + Activation) wiki/kernels/gated-dual-gemm.md
Grouped GEMM for MoE wiki/kernels/grouped-gemm.md
Sparse MLA wiki/kernels/sparse-mla.md
Triton on Blackwell wiki/languages/triton-blackwell.md
PTX Cache Policy Differentiation wiki/techniques/cache-policy.md
Chunk-Based Parallelism for Linear Recurrent Models wiki/techniques/chunk-parallelism.md
Double/Multi-Buffering Patterns wiki/techniques/double-buffering.md
Epilogue fusion wiki/techniques/epilogue-fusion.md
External Source-Map Research For Kernel Edits wiki/techniques/external-source-map-research.md
Fine-grained FP8/FP4 scaling wiki/techniques/fine-grained-quantization.md
Kernel fusion wiki/techniques/kernel-fusion.md
Software Pipelining and Multi-Stage Buffering wiki/techniques/pipeline-stages.md
Register budgeting wiki/techniques/register-budgeting.md
Shared Memory Swizzling wiki/techniques/swizzling.md
Tile Scheduling Strategies wiki/techniques/tile-scheduling.md
Wide Vectorized Loads and Cache Policies wiki/techniques/vectorized-loads.md
Warp Specialization on Hopper and Blackwell wiki/techniques/warp-specialization.md

sm90a

Page Path
CUTLASS CuTe DSL Documentation sources/docs/cutlass-cute-dsl.md
[Hopper CuTeDSL] Add grouped GEMM kernel example sources/prs/cutlass/PR-3091.md
GPT-OSS Support: Add Blackwell MoE mxfp4 implementation from TRTLLM and Attention Sink sources/prs/flashinfer/PR-1389.md
gpt-oss: Add MXFP8 x MXFP4 CUTLASS MOE for SM100 and BF16 x MXFP4 CUTLASS for SM90 + SwigluBias Activation sources/prs/flashinfer/PR-1396.md
feat: initial support for SM103, SM110, SM120, SM121 sources/prs/flashinfer/PR-1608.md
feat:enable fp8 blockscale moe for fused cultass for sm90 sources/prs/flashinfer/PR-1819.md
refactor: update fa3 codebase and fix hopper unittest [part 1] sources/prs/flashinfer/PR-2111.md
perf: bunch of features and optimizations for top-k (sampling + sparse attention) sources/prs/flashinfer/PR-2119.md
make DeepGEMM swapAB available for linear gemm SM90 sources/prs/flashinfer/PR-2131.md
feat: add GDN Attention sources/prs/flashinfer/PR-2276.md
feat: Add TRTLLM fmha_v2 library for SM90 attention with Skip-Softmax sources/prs/flashinfer/PR-2446.md
[feat] Add blackwell GDN prefill kernel sources/prs/flashinfer/PR-3001.md
perf: FlashAttention-3 style MLA PageAttention sources/prs/flashinfer/PR-887.md
Support FP4 gemm (1/2) sources/prs/sglang/PR-3899.md
support cmake for sgl-kernel sources/prs/sglang/PR-4706.md
[Kernel]: Cutlass 2:4 Sparsity + FP8/Int8 Quant Support sources/prs/vllm/PR-10995.md
[Kernel] (1/N) Machete - Hopper Optimized Mixed Precision Linear Kernel sources/prs/vllm/PR-7174.md
mbarrier (Memory Barrier Primitives) wiki/hardware/mbarrier.md
Tensor Memory Accelerator (TMA) wiki/hardware/tma.md