Skip to content

[ExecuTorch][WebGPU] cat op test suite (cases.py op-test framework)#20566

Merged
JulianCloudNTH merged 5 commits into
gh/JulianCloudNTH/47/origfrom
gh/JulianCloudNTH/48/orig
Jun 27, 2026
Merged

[ExecuTorch][WebGPU] cat op test suite (cases.py op-test framework)#20566
JulianCloudNTH merged 5 commits into
gh/JulianCloudNTH/47/origfrom
gh/JulianCloudNTH/48/orig

Conversation

@pytorchbot

Copy link
Copy Markdown
Collaborator

This PR was created by the merge bot to help merge the original PR into the main branch.
ghstack PR number: #20399 by @JulianCloudNTH
^ Please use this as the source of truth for the PR details, comments, and reviews
ghstack PR base: https://github.com/pytorch/executorch/tree/gh/JulianCloudNTH/48/base
ghstack PR head: https://github.com/pytorch/executorch/tree/gh/JulianCloudNTH/48/head
Merge bot PR base: https://github.com/pytorch/executorch/tree/gh/JulianCloudNTH/47/orig
Merge bot PR head: https://github.com/pytorch/executorch/tree/gh/JulianCloudNTH/48/orig

@diff-train-skip-merge

Pull Request resolved: #20399

Registers `aten.cat.default` in the `cases.py` op-test framework: a `_cat_suite` of 4 configs (3 inputs along dim 0, 2 inputs along dim 1, 3 inputs along dim 2, uneven split along dim 1) that `generate_op_tests` exports via `VulkanPartitioner` and compares to a torch golden on Dawn. Also adds `test/ops/cat/test_cat.py` (`CatModule` + `CONFIGS` + `_op_delegated` smoke test, with distinct per-input value ranges to catch cross-slab contamination) and the `aten.cat.default` partitioner-allowlist entry in `tester.py`.
ghstack-source-id: 397534688
@exported-using-ghexport

@diff-train-skip-merge

Differential Revision: [D108793163](https://our.internmc.facebook.com/intern/diff/D108793163/)
@pytorch-bot

pytorch-bot Bot commented Jun 27, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/20566

Note: Links to docs will display an error until the docs builds have been completed.

❌ 136 Cancelled Jobs, 6 Pending, 2 Unrelated Failures

As of commit 178b4d9 with merge base 200f64a (image):

CANCELLED JOBS - The following jobs were cancelled. Please retry:

FLAKY - The following jobs failed but were likely due to flakiness present on trunk:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-cla meta-cla Bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jun 27, 2026
…>.py

Pull Request resolved: #20435

Pure test-file relocation: moves the already-landed ops' tests from nested `test/ops/<op>/test_<op>.py` to flat `test/ops/test_<op>.py`, matching the ExecuTorch convention (XNNPACK uses flat `test/ops/test_<op>.py`; Vulkan uses flat `test/test_*.py`) and completing the flatten applied to the new ops in the stack below. Drops the per-op `__init__.py`; the parent `test/ops/__init__.py` is kept. Ops: `add`, `rms_norm`, `sdpa` (`test_sdpa` + `test_update_cache`), `dispatch_order`, `quantized_linear`, `embedding_q4gsw`, `rope`, `prepack`.

No behavior change — the test modules and their export/golden functions are unchanged; only their path moves. Every reference to the old paths is updated: the `cases.py` op-test imports (`add`, `rms_norm`), `test/TARGETS` (`test_add` srcs), `test/ops/test_dispatch_order.py`'s internal `rms_norm` import, and the build/CI scripts that import the per-op export functions (`test/test_build_webgpu.sh`, `scripts/test_webgpu_native_ci.sh`).

Nothing required the per-op subdirectory: the codegen framework imports only `cases.py`, the one buck target uses a literal path, and the native-golden scripts import the modules by path — each resolves identically at the flat path.
ghstack-source-id: 397534695
@exported-using-ghexport

@diff-train-skip-merge

Differential Revision: [D109349894](https://our.internmc.facebook.com/intern/diff/D109349894/)
Pull Request resolved: #20463

`aten.clone.default` is a pure flat copy on the buffer-only WebGPU backend, identical to `view_copy`: `clone_impl` reuses the existing `add_flat_copy` helper (`output[i] = input[i]`) and registers a handler under `aten.clone.default`. No new shader, generated WGSL header, or CMake source — it shares the `view_copy` flat-copy compute pipeline. Required for end-to-end Llama 3.2 1B (4-bit, KV cache): the exported model serializes 2 `aten.clone.default` ops into its runtime operator chain (the RoPE-frequency clones reused across all 16 transformer layers), so without a handler the partition graph-breaks at those nodes. Mirrors the Vulkan delegate, which registers the same op and routes a buffer clone to a flat view-copy.
ghstack-source-id: 397534700
@exported-using-ghexport

@diff-train-skip-merge

Differential Revision: [D109477717](https://our.internmc.facebook.com/intern/diff/D109477717/)
Pull Request resolved: #20464

Adds the WebGPU delegate handler for aten.index.Tensor, the 1D-self advanced-index
gather out[i] = self[index[i]] (output shape == index shape). This is the form the
VulkanPartitioner delegates -- it requires a 1D self and exactly one non-None index
(op_registry.py); 2D mask/freqs gathers stay on CPU. It mirrors the Vulkan delegate's
index_tensor op (IndexTensor.cpp + index_tensor_buffer.glsl) as a single compute
dispatch over the output elements, each reading the int32 index and gathering the
corresponding fp32 self element.

The op is composed as:
- index.wgsl: one workgroup-strided pass, out[i] = self[u32(index[i])], guarded by a
  numel bound; buffer-only, fp32 self/out, int32 index, 1D dispatch via the shared
  WebGPUUtils helpers (clamp workgroup size + 1D count).
- Index.cpp: validates the args (self/out tensors; indices ValueList with exactly one
  index tensor; fp32 self/out; int32 index; out numel == index numel), failing loud on
  any violation, then records the dispatch. row_width is dropped (always 1 for 1D self).
ghstack-source-id: 397756251
@exported-using-ghexport

@diff-train-skip-merge

Differential Revision: [D109478967](https://our.internmc.facebook.com/intern/diff/D109478967/)
…lden)

Pull Request resolved: #20465

Adds the test suite for the aten.index.Tensor op (stacked on the op diff):
- test/ops/index/test_index.py: exports a module computing x[idx] through
  VulkanPartitioner for four configs (reorder/repeat indices over distinct self
  values, so a wrong-gather is visible), asserts a VulkanBackend delegate with
  index.Tensor absorbed (not a CPU fallback), and writes per-config .pte +
  .self/.idx/.golden.bin.
- test/native/test_index.cpp: a standalone Dawn binary that loads each .pte, feeds
  self (fp32) + index (int64 at the program boundary, narrowed to the int32 buffer)
  and compares the gather against the torch golden at 1e-3, with a single-output
  shape guard.
- Wired into CMake (webgpu_index_test), test/TARGETS (python_unittest test_index),
  and the Dawn native CI script.
ghstack-source-id: 397763261
@exported-using-ghexport

@diff-train-skip-merge

Differential Revision: [D109479000](https://our.internmc.facebook.com/intern/diff/D109479000/)
@JulianCloudNTH JulianCloudNTH merged commit 4a301d0 into gh/JulianCloudNTH/47/orig Jun 27, 2026
8 checks passed
@JulianCloudNTH JulianCloudNTH deleted the gh/JulianCloudNTH/48/orig branch June 27, 2026 21:33
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants