PR #592 replaced the merged causal_conv1d kernel (#572) without benchmarking against it, causing degradation.
Timeline
Variant PTO (the PTO-ISA causal_conv1d kernel) was merged into main on 2026-07-05. Four days later, Variant ASC was merged (2026-07-10) and replaced it with the implementation from the still-open #555, reusing that PR's description and benchmark tables unchanged, and removing the merged kernel entirely (main today contains none of it). The creators of Variant ASC never benchmarked against Variant PTO, the kernel actually on main at the time. Its performance tables compare against an older, superseded baseline, not the incumbent it replaced.
Note that Variant PTO had already been benchmarked head-to-head against #555 as part of its own review, where that older kernel measured slower. So Variant ASC re-submitted an implementation that had already been compared against Variant PTO and measured slower, without re-running that comparison. Re-measuring today, the kernel now on main is slower than Variant PTO, so this merge was a performance regression.
Head-to-head: Variant ASC vs VariantPTO
We benchmarked the Variant ASC (kernel now on main) against VariantPTO for identical shapes and hardware, same input x=[B, seq, dim], weight, bias, output, and conv_states write-back, so both kernels do identical total work (only the internal tiling differs).
Method: Ascend 910B2, fp16. e2e latency via torch.npu.Event, min over 5 windows, 100 iters/window, 5 warmups.
Result for 69 shapes: VariantPTO is faster on all 69, median 1.49× (range 1.03×–1.90×). Both kernels are correct on every shape, so this is a pure performance regression.
Full per-shape sweep (69 shapes)
| B |
dim |
seq |
B·seq·dim |
Variant PTO #572 (µs) |
Variant ASC #592 (µs) |
Variant ASC / Variant PTO |
| 64 |
1024 |
512 |
33554432 |
358.6 |
408.6 |
1.14× |
| 64 |
1024 |
1024 |
67108864 |
707.0 |
785.3 |
1.11× |
| 64 |
1024 |
2048 |
134217728 |
1400.7 |
1541.6 |
1.10× |
| 64 |
2048 |
512 |
67108864 |
602.4 |
628.1 |
1.04× |
| 64 |
2048 |
1024 |
134217728 |
1191.7 |
1232.0 |
1.03× |
| 64 |
2048 |
2048 |
268435456 |
2370.3 |
2444.8 |
1.03× |
| 64 |
4096 |
512 |
134217728 |
898.9 |
1084.7 |
1.21× |
| 64 |
4096 |
1024 |
268435456 |
1782.6 |
2146.8 |
1.20× |
| 64 |
4096 |
2048 |
536870912 |
3550.8 |
4269.3 |
1.20× |
| 64 |
6144 |
512 |
201326592 |
1261.6 |
2144.6 |
1.70× |
| 64 |
6144 |
1024 |
402653184 |
2508.8 |
4265.8 |
1.70× |
| 64 |
6144 |
2048 |
805306368 |
5001.8 |
8513.3 |
1.70× |
| 128 |
1024 |
512 |
67108864 |
535.3 |
787.6 |
1.47× |
| 128 |
1024 |
1024 |
134217728 |
1055.7 |
1540.7 |
1.46× |
| 128 |
1024 |
2048 |
268435456 |
2096.7 |
3051.8 |
1.46× |
| 128 |
2048 |
512 |
134217728 |
899.0 |
1234.5 |
1.37× |
| 128 |
2048 |
1024 |
268435456 |
1783.4 |
2444.6 |
1.37× |
| 128 |
2048 |
2048 |
536870912 |
3551.0 |
4867.7 |
1.37× |
| 128 |
4096 |
512 |
268435456 |
1788.7 |
2147.7 |
1.20× |
| 128 |
4096 |
1024 |
536870912 |
3558.0 |
4271.9 |
1.20× |
| 128 |
4096 |
2048 |
1073741824 |
7095.0 |
8517.7 |
1.20× |
| 128 |
6144 |
512 |
402653184 |
2514.9 |
4268.9 |
1.70× |
| 128 |
6144 |
1024 |
805306368 |
5009.3 |
8515.6 |
1.70× |
| 128 |
6144 |
2048 |
1610612736 |
9997.5 |
17006.4 |
1.70× |
| 256 |
1024 |
512 |
134217728 |
1061.4 |
1545.0 |
1.46× |
| 256 |
1024 |
1024 |
268435456 |
2102.2 |
3056.3 |
1.45× |
| 256 |
1024 |
2048 |
536870912 |
4183.5 |
6079.2 |
1.45× |
| 256 |
2048 |
512 |
268435456 |
1789.0 |
2447.6 |
1.37× |
| 256 |
2048 |
1024 |
536870912 |
3557.0 |
4871.7 |
1.37× |
| 256 |
2048 |
2048 |
1073741824 |
7093.8 |
9717.6 |
1.37× |
| 256 |
4096 |
512 |
536870912 |
3272.4 |
4274.0 |
1.31× |
| 256 |
4096 |
1024 |
1073741824 |
6515.2 |
8520.7 |
1.31× |
| 256 |
4096 |
2048 |
2147483648 |
12998.8 |
17013.1 |
1.31× |
| 256 |
6144 |
512 |
805306368 |
4605.6 |
8524.3 |
1.85× |
| 256 |
6144 |
1024 |
1610612736 |
9177.6 |
17015.0 |
1.85× |
| 256 |
6144 |
2048 |
3221225472 |
18321.6 |
34002.0 |
1.86× |
| 512 |
1024 |
512 |
268435456 |
1936.8 |
3062.3 |
1.58× |
| 512 |
1024 |
1024 |
536870912 |
3846.9 |
6083.9 |
1.58× |
| 512 |
1024 |
2048 |
1073741824 |
7660.7 |
12128.2 |
1.58× |
| 512 |
2048 |
512 |
536870912 |
3271.5 |
4879.3 |
1.49× |
| 512 |
2048 |
1024 |
1073741824 |
6516.3 |
9724.1 |
1.49× |
| 512 |
2048 |
2048 |
2147483648 |
12997.9 |
19416.1 |
1.49× |
| 512 |
4096 |
512 |
1073741824 |
6538.4 |
8529.9 |
1.30× |
| 512 |
4096 |
1024 |
2147483648 |
13021.3 |
17023.3 |
1.31× |
| 512 |
4096 |
2048 |
4294967296 |
25991.8 |
34008.8 |
1.31× |
| 512 |
6144 |
512 |
1610612736 |
9201.1 |
17030.7 |
1.85× |
| 512 |
6144 |
1024 |
3221225472 |
18344.8 |
34017.8 |
1.85× |
| 512 |
6144 |
2048 |
6442450944 |
36628.2 |
67990.1 |
1.86× |
| 1024 |
1024 |
512 |
536870912 |
3860.1 |
6098.9 |
1.58× |
| 1024 |
1024 |
1024 |
1073741824 |
7676.3 |
12145.4 |
1.58× |
| 1024 |
1024 |
2048 |
2147483648 |
15314.1 |
24232.3 |
1.58× |
| 1024 |
2048 |
512 |
1073741824 |
6535.0 |
9739.5 |
1.49× |
| 1024 |
2048 |
1024 |
2147483648 |
13019.8 |
19431.2 |
1.49× |
| 1024 |
2048 |
2048 |
4294967296 |
25985.3 |
38813.8 |
1.49× |
| 1024 |
4096 |
512 |
2147483648 |
12772.7 |
17040.8 |
1.33× |
| 1024 |
4096 |
1024 |
4294967296 |
25438.1 |
34029.8 |
1.34× |
| 1024 |
6144 |
512 |
3221225472 |
17976.0 |
34049.2 |
1.89× |
| 1024 |
6144 |
1024 |
6442450944 |
35847.5 |
68024.1 |
1.90× |
| 2048 |
1024 |
512 |
1073741824 |
7536.5 |
12170.5 |
1.61× |
| 2048 |
1024 |
1024 |
2147483648 |
14992.4 |
24260.6 |
1.62× |
| 2048 |
1024 |
2048 |
4294967296 |
29905.8 |
48437.2 |
1.62× |
| 2048 |
2048 |
512 |
2147483648 |
12761.5 |
19463.1 |
1.53× |
| 2048 |
2048 |
1024 |
4294967296 |
25437.4 |
38846.1 |
1.53× |
| 2048 |
4096 |
512 |
4294967296 |
25520.0 |
34068.8 |
1.33× |
| 2048 |
6144 |
512 |
6442450944 |
35936.0 |
68083.4 |
1.89× |
| 4096 |
1024 |
512 |
2147483648 |
15053.3 |
24315.3 |
1.62× |
| 4096 |
1024 |
1024 |
4294967296 |
29971.4 |
48496.5 |
1.62× |
| 4096 |
2048 |
512 |
4294967296 |
25509.2 |
38909.7 |
1.53× |
| 8192 |
1024 |
512 |
4294967296 |
29921.6 |
48602.3 |
1.62× |
Request
- Please confirm our benchmark findings.
- After confirming our results, we request to revert
main back to the Variant PTO implementation.
Cc
@ping1jing2 @iforgetmyname @zhaozx-cn @raphael-s-steiner
Links: #572 (merged) · #555 (open) · #592 (merged)
PR #592 replaced the merged
causal_conv1dkernel (#572) without benchmarking against it, causing degradation.Timeline
mainby @ping1jing2mainby @iforgetmyname, replacing Variant PTOVariant PTO (the PTO-ISA
causal_conv1dkernel) was merged intomainon 2026-07-05. Four days later, Variant ASC was merged (2026-07-10) and replaced it with the implementation from the still-open #555, reusing that PR's description and benchmark tables unchanged, and removing the merged kernel entirely (maintoday contains none of it). The creators of Variant ASC never benchmarked against Variant PTO, the kernel actually onmainat the time. Its performance tables compare against an older, superseded baseline, not the incumbent it replaced.Note that Variant PTO had already been benchmarked head-to-head against #555 as part of its own review, where that older kernel measured slower. So Variant ASC re-submitted an implementation that had already been compared against Variant PTO and measured slower, without re-running that comparison. Re-measuring today, the kernel now on
mainis slower than Variant PTO, so this merge was a performance regression.Head-to-head: Variant ASC vs VariantPTO
We benchmarked the Variant ASC (kernel now on
main) against VariantPTO for identical shapes and hardware, same inputx=[B, seq, dim], weight, bias, output, andconv_stateswrite-back, so both kernels do identical total work (only the internal tiling differs).Method: Ascend 910B2, fp16. e2e latency via
torch.npu.Event, min over 5 windows, 100 iters/window, 5 warmups.Result for 69 shapes: VariantPTO is faster on all 69, median 1.49× (range 1.03×–1.90×). Both kernels are correct on every shape, so this is a pure performance regression.
Full per-shape sweep (69 shapes)
Request
mainback to the Variant PTO implementation.Cc
@ping1jing2 @iforgetmyname @zhaozx-cn @raphael-s-steiner
Links: #572 (merged) · #555 (open) · #592 (merged)