ci: clean up recurring nightly optional-deps failures - #2968
Merged
Merged
Conversation
…mlir Two targeted fixes for the recurring failures tracked in NVIDIA#2748: - nightly-numba-cuda: pin numpy<2.5. numba-cuda 0.30.4 references np.row_stack, removed in NumPy 2.5, so every nightly numba-cuda job crashes at collection. Tracked upstream in NVIDIA/numba-cuda#907; drop the cap once a fixed wheel is on PyPI. - nightly-numba-cuda-mlir: add cccl to cuda-toolkit extras. NVRTC currently fails with "catastrophic error: cannot open source file 'nv/target'" when compiling cooperative_groups/details/info.h and curand_kernel.h; the cccl extra pulls nvidia-cuda-cccl, which drops the missing header at nvidia/cu13/include/nv/target.
Contributor
Revert before merging. Same trick as NVIDIA@8d51cf7
test_frozen_driver_table_covers_all_curesult_members was removed on main by NVIDIA#2792 (2026-09-09) but neither cuda-core-v1.2.0 nor v1.2.1 carry the revert. main cuda-bindings 13.4.1 adds three new CUresult members (MULTICAST_RESOURCE_FULL, INSUFFICIENT_LOADER_VERSION, FABRIC_NOT_READY) not present in the frozen table those releases ship, so the nightly-cuda-core job trips it on every run. Mirror the existing v1.0.1 NvlinkVersion pattern with a version-gated --deselect that drops automatically once the next cuda-core release ships the NVIDIA#2792 revert.
leofang
commented
Sep 30, 2026
| DESELECTS+=( | ||
| --deselect 'tests/test_utils_enum_explanations_helpers.py::test_frozen_driver_table_covers_all_curesult_members' | ||
| ) | ||
| fi |
Member
Author
There was a problem hiding this comment.
Unfortunately #2792 did not catch any cuda.core v1.2.x release and only lives on the main branch, so we have to fix the nightly CI this way.
Member
Author
|
/ok to test feb41f7 |
Contributor
|
nightly-numba-cuda: switch from `python -m numba.runtests` to `pytest` (numba-cuda's own CI uses pytest; PR NVIDIA#1987 landed the runtests form as a bring-up fix) so we can skip individual tests. Two workarounds: - pin pytest<9 (matches upstream's numba-cuda cap; NVIDIA/numba-cuda#637 covers the subTest breakage on 9). - --ignore-glob "**/test_ipc.py" — numba-cuda v0.30.4 accesses CUipcMemHandle.reserved, which cuda-bindings intentionally removed; numba-cuda is EOL upstream so no fix is coming (NVIDIA#2748). nightly-numba-cuda-mlir: add --deselect for TestCudaDeviceRecordWithRecord.test_device_record_copy — the WithRecord fixture uses np.recarray (uninitialized memory), so when the float32 field encodes a NaN, `np.testing.assert_equal` trips on NaN != NaN even though the copy round-trip is correct. Filed upstream as NVIDIA/numba-cuda-mlir#341.
Keeps pip constraints in the install step alongside numpy<2.5, instead of a separate `pip install` in the workflow.
Member
Author
|
/ok to test f0682bf |
cuda_core's test-cuXX group hard-pins pytest==9.1.0, so pip can't satisfy that and pytest<9 in one solve. Split the numba-cuda pytest pin into a second `pip install "pytest<9"` after the main install.
Member
Author
|
/ok to test 31f7af7 |
Default `prepend` import mode walks up until it finds a directory without `__init__.py`. numba-cuda's on-disk layout is `site-packages/numba_cuda/numba/cuda/tests/...`, and `numba/` here is a namespace subdir with no `__init__.py`, so pytest treats it as the rootpath and imports each test file as `cuda.tests.<...>`. That collides with cuda-bindings' top-level `cuda` package (which has no `tests` subpackage), giving `ModuleNotFoundError: No module named 'cuda.tests'` at collection. `--import-mode=importlib` sidesteps the sys.path walk and imports each test module by its correct dotted path.
Same shape as the numba-cuda-mlir step. run-tests exposes NUMBA_CUDA_VER; the workflow checks out NVIDIA/numba-cuda at the matching tag and runs pytest from numba-cuda-released/testing/, so upstream's testing/pytest.ini kicks in (consider_namespace_packages=true + --pyargs numba.cuda.tests) and we stop caring about the site-packages namespace-package layout. Also mirror the wheel-only gate: in wheels mode, set NUMBA_CUDA_TEST_WHEEL_ONLY=1 so numba-cuda skips tests that need CTK executables. Binary-generated tests self-skip via `@unittest.skipIf(not NUMBA_CUDA_TEST_BIN_DIR, ...)`, so no `make -j` step is needed — mirrors mlir, which has no Makefile. Drops the --import-mode=importlib workaround from the prior commit; pytest.ini's consider_namespace_packages does the job.
Member
Author
|
/ok to test 3bc4acc |
pytest's --ignore-glob="**/test_ipc.py" didn't match through the --pyargs-resolved site-packages path (5 tests still ran and failed with the CUipcMemHandle.reserved AttributeError on linux-64). Match by class name instead — the affected tests live in TestIpcMemory and TestIpcStaged, both starting with "TestIpc", so this is unambiguous.
Member
Author
|
/ok to test b53d482 |
brandon-b-miller
approved these changes
Sep 30, 2026
brandon-b-miller
left a comment
Contributor
There was a problem hiding this comment.
numba-cuda/numba-cuda-mlir related changes LGTM.
leofang
commented
Sep 30, 2026
…s PR" This reverts commit 4363950.
Member
Author
|
Since the CI was green, let me admin-merge to save resource.
|
leofang
marked this pull request as ready for review
September 30, 2026 14:30
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Cleanup pass for the recurring failures tracked in #2748, grouped by affected library.
numba-cudanumpy<2.5inci/tools/run-testsfornightly-numba-cuda.numba-cuda 0.30.4callsnp.row_stack, which NumPy 2.5 removed, so every nightly numba-cuda job crashed at collection. Tracked upstream in NVIDIA/numba-cuda#907.pytest<9as a secondpip installinci/tools/run-tests.numba-cuda'ssubTestusage breaks under pytest 9 + xdist (NVIDIA/numba-cuda#637). Split into a second install becausecuda_core'stest-cuXXgroup hard-pinspytest==9.1.0, which pip's resolver can't satisfy alongsidepytest<9in the same solve.python -m numba.runteststopytest, mirroring thenightly-numba-cuda-mlirstep:run-testsexposesNUMBA_CUDA_VER, the workflow checks outNVIDIA/numba-cudaatv${NUMBA_CUDA_VER}, and the test step runs fromnumba-cuda-released/testing/so upstream'stesting/pytest.ini(consider_namespace_packages = true+--pyargs numba.cuda.tests) drives discovery. Avoids the namespace-package/pytest interaction that was pushingnumba.runtestsin the first place.NUMBA_CUDA_TEST_WHEEL_ONLY=1in wheels mode so tests requiring CTK executables self-skip. Same env var upstream numba-cuda uses in its own wheels-mode CI; matches theNUMBA_CUDA_MLIR_TEST_WHEEL_ONLYgate already in the mlir step.-k "not TestIpc"—numba-cuda 0.30.4accessesCUipcMemHandle.reserved, which cuda-bindings intentionally removed; numba-cuda is EOL upstream so no fix is coming. Only fires on linux-64 x86_64 —TestIpcMemoryandTestIpcStagedare stacked with@linux_only/@skip_on_arm/@skip_on_wsl2and auto-skip elsewhere.numba-cuda-mlirccclextra tocuda-toolkitinci/tools/run-testsfornightly-numba-cuda-mlir. NVRTC was failing withcatastrophic error: cannot open source file "nv/target"while compilingcooperative_groups/details/info.handcurand_kernel.h; theccclextra pullsnvidia-cuda-cccl, which drops the missing header atnvidia/cu13/include/nv/target.--deselectTestCudaDeviceRecordWithRecord::test_device_record_copy. The fixture usesnp.recarray(uninitialized memory); when the float32 field happens to encode a NaN,np.testing.assert_equalfails onNaN != NaNeven though the copy round-trip is correct. Filed upstream as NVIDIA/numba-cuda-mlir#341.cuda-core--deselectfortest_frozen_driver_table_covers_all_curesult_memberson releasedcuda-core <= 1.2.1. Main cuda-bindings 13.4.1 exposes 3 newCUresultmembers (CUDA_ERROR_MULTICAST_RESOURCE_FULL,CUDA_ERROR_INSUFFICIENT_LOADER_VERSION,CUDA_ERROR_FABRIC_NOT_READY) not covered by the frozen fallback table those releases ship. The test was already reverted from main by [no-ci] Refreeze enum tables (revert 2383, add comment) #2792, so this deselect drops automatically once the next cuda-core release ships that revert. Mirrors the existingv1.0.1 NvlinkVersiondeselect pattern in the same step.