Skip to content

tests: enumerate CUDA devices for GPU-only system checks - #2916

Merged
isVoid merged 1 commit into
NVIDIA:mainfrom
isVoid:codex/filter-non-cuda-system-tests
Sep 18, 2026
Merged

isVoid merged 1 commit into
NVIDIA:mainfrom
isVoid:codex/filter-non-cuda-system-tests

Conversation

@isVoid

@isVoid isVoid commented Sep 18, 2026

Copy link
Copy Markdown
Contributor

Description

cuda 1.2.0 contains some tests that queries using CUDA API, which would fail on N1X systems due to the non-CUDA devices. Previously we had fixes for some of the test to enumerate only on CUDADevice.get_all_devices(), but they don't cover all tests impacted. This PR fixes that.

@github-actions github-actions Bot added the cuda.core Everything related to the cuda.core module label Sep 18, 2026

@mdboom mdboom left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@isVoid isVoid self-assigned this Sep 18, 2026
@isVoid isVoid added test Improvements or additions to tests P1 Medium priority - Should do labels Sep 18, 2026
@isVoid isVoid added this to the cuda.core 1.3.0 milestone Sep 18, 2026
@isVoid

isVoid commented Sep 18, 2026

Copy link
Copy Markdown
Contributor Author

@github-actions

This comment has been minimized.

@isVoid
isVoid merged commit f59229e into NVIDIA:main Sep 18, 2026
119 of 122 checks passed
@github-actions

This comment has been minimized.

@github-actions

Copy link
Copy Markdown
Contributor
Doc Preview CI
Preview removed because the pull request was closed or merged.

Andy-Jost added a commit that referenced this pull request Sep 23, 2026
* cuda.core: host-only memory resources need no CUDA context (#2773)

Since 1.2.0, Buffer.from_handle binds a default-stream deallocation
token to the current context for every owning memory resource. Memory
the device cannot access has no stream ordering to preserve, so skip the
binding when mr.is_device_accessible is False. Such buffers record no
deallocation stream and never call the driver.

Fixes #2769

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
(cherry picked from commit 3fef50b)

* Add Windows-on-ARM support

(cherry picked from commit 9d762dd)

* Fix Windows on ARM build

(cherry picked from commit 7cfdad7)

* Exclude Python 3.10 on Windows-on-ARM

(cherry picked from commit a6f6adc)

* make platform tag explicit (#2832)

(cherry picked from commit c280d47)

* tests: enumerate CUDA devices for GPU-only system checks (#2916)

(cherry picked from commit f59229e)

* Prepare cuda.core v1.2.1 with Windows on ARM support

Release branch off the cuda-core-v1.2.0 tag. Windows on ARM wheels need
the CUDA 13.4 toolkit, so the build moves to CUDA 13.4.2 for every
platform (ci/versions.yml and the pixi cuda-version pins). The WoA build
jobs themselves are cherry-picked from main in the preceding commits.

From the CUDA 13.4 integration on main (#2437, #2788, merged via #2789):
- Map the ClocksEventReasons BOARD_LIMIT and RELIABILITY values that
  CUDA 13.4 added. Device.current_clock_event_reasons and
  Device.supported_clock_event_reasons raised ValueError for an unknown
  reason bit on devices that report them.
- Map the NVIDIA DLA, vGameDev and NPU brand types instead of "Unknown".
- Test updates: accept the "DLA-" UUID prefix, exclude the new 13.4
  memory-location enum value from the coverage check, and cover
  Device.get_all_devices.

The 1.2.1 release notes and the install page describe the new wheels:
CUDA 13 only, Python 3.11 and newer, cuda-bindings 13.4.1 or newer.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* chore: refresh pixi.lock for cuda_core (#2830)

Scheduled pixi update --no-install for this workspace only.

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
(cherry picked from commit 04d4753)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* chore: refresh pixi.lock for cuda_bindings (#2828)

Scheduled pixi update --no-install for this workspace only.

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
(cherry picked from commit 7aa0648)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* ci: drop the Python 3.15 test jobs on the 1.2.1 release branch

This branch builds with CUDA 13.4.2 but tests on 13.3.0, so the test
jobs install published cuda-bindings 13.3.x to cover the real-world
combination. No 13.3.x wheel exists for Python 3.15, so the three 3.15
and 3.15t test jobs fail at pip install before any test runs. 1.2.1
ships no 3.15 wheels, so nothing is lost; the 3.15 build jobs remain.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

---------

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Co-authored-by: Michael Droettboom <mdroettboom@nvidia.com>
Co-authored-by: Leo Fang <leof@nvidia.com>
Co-authored-by: Michael Wang <13521008+isVoid@users.noreply.github.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

cuda.core Everything related to the cuda.core module P1 Medium priority - Should do test Improvements or additions to tests

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants