Skip to content

Performance regression in causal_conv1d #596

Description

@ChristosMatzoros

PR #592 replaced the merged causal_conv1d kernel (#572) without benchmarking against it, causing degradation.

Timeline

Date PR Name Description
2026-06-15 #555 "add gdn custom conv1d" (@zhaozx-cn) opened — still open, never merged
2026-06-25 #572 Variant PTO (#572) "causal_conv1d: PTO-ISA rewrite" (@ChristosMatzoros) opened
2026-07-05 #572 Variant PTO (#572) Variant PTO merged into main by @ping1jing2
2026-07-09 #592 Variant ASC (#592) "Custom causal conv1d" (@zhaozx-cn) opened in a new PR (same code as #555)
2026-07-10 #592 Variant ASC (#592) Variant ASC merged into main by @iforgetmyname, replacing Variant PTO

Variant PTO (the PTO-ISA causal_conv1d kernel) was merged into main on 2026-07-05. Four days later, Variant ASC was merged (2026-07-10) and replaced it with the implementation from the still-open #555, reusing that PR's description and benchmark tables unchanged, and removing the merged kernel entirely (main today contains none of it). The creators of Variant ASC never benchmarked against Variant PTO, the kernel actually on main at the time. Its performance tables compare against an older, superseded baseline, not the incumbent it replaced.

Note that Variant PTO had already been benchmarked head-to-head against #555 as part of its own review, where that older kernel measured slower. So Variant ASC re-submitted an implementation that had already been compared against Variant PTO and measured slower, without re-running that comparison. Re-measuring today, the kernel now on main is slower than Variant PTO, so this merge was a performance regression.

Head-to-head: Variant ASC vs VariantPTO

We benchmarked the Variant ASC (kernel now on main) against VariantPTO for identical shapes and hardware, same input x=[B, seq, dim], weight, bias, output, and conv_states write-back, so both kernels do identical total work (only the internal tiling differs).

Method: Ascend 910B2, fp16. e2e latency via torch.npu.Event, min over 5 windows, 100 iters/window, 5 warmups.

Result for 69 shapes: VariantPTO is faster on all 69, median 1.49× (range 1.03×–1.90×). Both kernels are correct on every shape, so this is a pure performance regression.

Full per-shape sweep (69 shapes)
B dim seq B·seq·dim Variant PTO #572 (µs) Variant ASC #592 (µs) Variant ASC / Variant PTO
64 1024 512 33554432 358.6 408.6 1.14×
64 1024 1024 67108864 707.0 785.3 1.11×
64 1024 2048 134217728 1400.7 1541.6 1.10×
64 2048 512 67108864 602.4 628.1 1.04×
64 2048 1024 134217728 1191.7 1232.0 1.03×
64 2048 2048 268435456 2370.3 2444.8 1.03×
64 4096 512 134217728 898.9 1084.7 1.21×
64 4096 1024 268435456 1782.6 2146.8 1.20×
64 4096 2048 536870912 3550.8 4269.3 1.20×
64 6144 512 201326592 1261.6 2144.6 1.70×
64 6144 1024 402653184 2508.8 4265.8 1.70×
64 6144 2048 805306368 5001.8 8513.3 1.70×
128 1024 512 67108864 535.3 787.6 1.47×
128 1024 1024 134217728 1055.7 1540.7 1.46×
128 1024 2048 268435456 2096.7 3051.8 1.46×
128 2048 512 134217728 899.0 1234.5 1.37×
128 2048 1024 268435456 1783.4 2444.6 1.37×
128 2048 2048 536870912 3551.0 4867.7 1.37×
128 4096 512 268435456 1788.7 2147.7 1.20×
128 4096 1024 536870912 3558.0 4271.9 1.20×
128 4096 2048 1073741824 7095.0 8517.7 1.20×
128 6144 512 402653184 2514.9 4268.9 1.70×
128 6144 1024 805306368 5009.3 8515.6 1.70×
128 6144 2048 1610612736 9997.5 17006.4 1.70×
256 1024 512 134217728 1061.4 1545.0 1.46×
256 1024 1024 268435456 2102.2 3056.3 1.45×
256 1024 2048 536870912 4183.5 6079.2 1.45×
256 2048 512 268435456 1789.0 2447.6 1.37×
256 2048 1024 536870912 3557.0 4871.7 1.37×
256 2048 2048 1073741824 7093.8 9717.6 1.37×
256 4096 512 536870912 3272.4 4274.0 1.31×
256 4096 1024 1073741824 6515.2 8520.7 1.31×
256 4096 2048 2147483648 12998.8 17013.1 1.31×
256 6144 512 805306368 4605.6 8524.3 1.85×
256 6144 1024 1610612736 9177.6 17015.0 1.85×
256 6144 2048 3221225472 18321.6 34002.0 1.86×
512 1024 512 268435456 1936.8 3062.3 1.58×
512 1024 1024 536870912 3846.9 6083.9 1.58×
512 1024 2048 1073741824 7660.7 12128.2 1.58×
512 2048 512 536870912 3271.5 4879.3 1.49×
512 2048 1024 1073741824 6516.3 9724.1 1.49×
512 2048 2048 2147483648 12997.9 19416.1 1.49×
512 4096 512 1073741824 6538.4 8529.9 1.30×
512 4096 1024 2147483648 13021.3 17023.3 1.31×
512 4096 2048 4294967296 25991.8 34008.8 1.31×
512 6144 512 1610612736 9201.1 17030.7 1.85×
512 6144 1024 3221225472 18344.8 34017.8 1.85×
512 6144 2048 6442450944 36628.2 67990.1 1.86×
1024 1024 512 536870912 3860.1 6098.9 1.58×
1024 1024 1024 1073741824 7676.3 12145.4 1.58×
1024 1024 2048 2147483648 15314.1 24232.3 1.58×
1024 2048 512 1073741824 6535.0 9739.5 1.49×
1024 2048 1024 2147483648 13019.8 19431.2 1.49×
1024 2048 2048 4294967296 25985.3 38813.8 1.49×
1024 4096 512 2147483648 12772.7 17040.8 1.33×
1024 4096 1024 4294967296 25438.1 34029.8 1.34×
1024 6144 512 3221225472 17976.0 34049.2 1.89×
1024 6144 1024 6442450944 35847.5 68024.1 1.90×
2048 1024 512 1073741824 7536.5 12170.5 1.61×
2048 1024 1024 2147483648 14992.4 24260.6 1.62×
2048 1024 2048 4294967296 29905.8 48437.2 1.62×
2048 2048 512 2147483648 12761.5 19463.1 1.53×
2048 2048 1024 4294967296 25437.4 38846.1 1.53×
2048 4096 512 4294967296 25520.0 34068.8 1.33×
2048 6144 512 6442450944 35936.0 68083.4 1.89×
4096 1024 512 2147483648 15053.3 24315.3 1.62×
4096 1024 1024 4294967296 29971.4 48496.5 1.62×
4096 2048 512 4294967296 25509.2 38909.7 1.53×
8192 1024 512 4294967296 29921.6 48602.3 1.62×

Request

  1. Please confirm our benchmark findings.
  2. After confirming our results, we request to revert main back to the Variant PTO implementation.

Cc

@ping1jing2 @iforgetmyname @zhaozx-cn @raphael-s-steiner

Links: #572 (merged) · #555 (open) · #592 (merged)

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions