cuda.core: require a cuda-bindings floor and call the driver through its resolved pointers (#2783) - #2920
Conversation
cuda.core accepted any cuda-bindings of the right major at build time and at run time, and built against whatever cuda.h was on the include path. Builds succeeded for configurations we never test, and an older cuda-bindings at run time surfaced as an ImportError for a missing C function, a silently disabled feature, or a null-pointer crash (NVIDIA#2783). Each supported CUDA major now has a floor, the newest cuda-bindings release its CI source root can build: 12.9.8 for CUDA 12 and 13.4.1 for CUDA 13, in the import-free module cuda/core/_bindings_floor.py. Build time: the pip build requirement becomes `cuda-bindings>=<floor>,==<major>.*`, and because conda-forge, pixi and --no-build-isolation installs bypass it, the backend itself checks the imported cuda-bindings against the floor and requires the cuda.h it compiles against to have the same major.minor as that cuda-bindings (the header cuda-bindings was generated from), even when CUDA_CORE_BUILD_MAJOR is set. It records the header version and the floor in the generated, gitignored cuda/core/_build_info.py, shipped like _version.py. Run time: cuda/core/__init__.py reads that record from the selected build (the cu12/cu13 subpackage of the merged wheel, or the top level of a plain build) and requires the installed cuda-bindings to be of the build's major and at least the floor, or at least the header's minor when that is newer. The error names the version found, the version required, and the pip command that fixes it. Pre-release and dev builds of an accepted version pass. Packaging: the cu12/cu13 extras pin the floor; a test keeps them and the ci/versions.yml toolkit pins in step with the module. CI: the BINDINGS_SOURCE=published rows, which paired a new wheel with cuda-bindings 13.0 (no longer supported), become BINDINGS_SOURCE=floor: they install the floor bindings, read from the wheel under test by ci/tools/cuda_core_bindings_floor.py, and keep their older CTK libraries. The prior-major rows whose CTK minor differs from prev_build do the same. Docs: support policy section, install guide, 1.3.0 release note. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The two `#if CUDA_VERSION >= 130x0` fences in _cpp/rt/driver_api.* compiled cuDevSmResourceSplit and cuMemcpyWithAttributesAsync out of a source build against an older 13.x header, while the run-time gates, which looked at the bindings and the driver, still reported the features as available; the calls then failed with CUDA_ERROR_NOT_SUPPORTED (NVIDIA#2783, "guards disable working features"). Both features also went through void*-typed C++ shims and has_*() presence probes whose only purpose was to avoid an eager cimport of a cydriver function that cuda-bindings 13.0 did not export. With the cuda-bindings floor in place (13.4.1 for CUDA 13), the cimport is safe, so the shims, the presence probes and the two fenced pointers go away: _device_resources.pyx and _buffer.pyx call cydriver.cuDevSmResourceSplit and cydriver.cuMemcpyWithAttributesAsync directly under the existing `IF CUDA_CORE_BUILD_MAJOR >= 13`, gated on the driver version alone. build_hooks.py now defines CUDA_CORE_BUILD_MAJOR (until now a Cython-only compile-time constant) and CUDA_CORE_MIN_CUDA_VERSION (the floor's major.minor) for the C++ compiler. The new _cpp/rt/versions.hpp, the first include of the tree via types.hpp, re-checks cuda.h against both with #error, so a build that bypasses the backend still cannot compile against an unsupported header. tests/test_rt_layout.py enforces that versions.hpp is the only file under _cpp/ that names CUDA_VERSION. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…VIDIA#2783) The C++ under _cpp/rt/ used to call the driver through cuda-bindings' Cython wrappers, extracted from cydriver.__pyx_capi__ at import. A wrapper for a function the driver lacks raises a Python exception the C++ never sees, and a wrapper the installed cuda-bindings lacks made the pointer optional and probed for null at every use. driver_api.hpp now lists every driver function the C++ calls, with the CUDA version cuda-bindings requests it at. The table is filled on first use from cuda.bindings._internal.driver._inspect_function_pointers(), which holds the driver's own entry points, so `import cuda.core` never touches the driver. Calls go through DRIVER_CALL(name, args...), which fills the table if needed and, when a pointer is still null after the fill (a missing feature gate or a failed fill), reports once and returns an error status from a trampoline instead of dereferencing null. The fill rejects a driver older than the CUDA major series. pw_ deleter wrappers use the same path, so the C++ never null-checks a pointer; functions the driver may lack are gated on the driver version in Cython. NVRTC, NVVM and nvJitLink each have their own one-entry table, filled when a program or linker handle is created. Consequences in this commit: the null-check fallbacks to CUDA_ERROR_NOT_SUPPORTED are gone; green-context stream creation is gated in _stream.pyx on driver 12.5; deviceptr_import_ipc resolves the table before taking ipc_import_mutex and keeps raw calls under the lock (marked `// raw:`); _rt.pyx no longer imports cydriver/cynvrtc/cynvvm/ cynvjitlink or reads __pyx_capi__. test_rt_layout.py checks the table against cuda-bindings' loader and forbids raw p_ calls elsewhere. DESIGN.md and AGENTS.md describe the mechanism and settle the GIL contract for entry points.
…as a floor (NVIDIA#2783) Every cuda-bindings cuda.core accepts (12.9.8+, 13.4.1+) has the full API surface of its CUDA major, so a run-time check of the cuda-bindings version says nothing a build-major fence does not. Double checks (driver and bindings) become driver-only checks; checks of the bindings minor become `IF CUDA_CORE_BUILD_MAJOR` fences in Cython or a comparison with the new `cuda.core._utils.version.BUILD_CUDA_MAJOR` in Python; checks that every accepted cuda-bindings satisfies are deleted, along with the fallbacks they guarded. - Kernel argument info, green contexts, workqueues, graph node updates, conditional graph nodes: driver-only gates. - cuGraphNodeGetParams (13.2): a driver gate on the CUDA 13 build, and the call goes through cydriver instead of the Python driver layer. - Managed memory NUMA locations, virtual memory MANAGED type: build fence (plus the 13.0 driver where a driver entry point is involved). - Copy attribute enums (CUDA 12.8): unconditional. - NVVM, nvJitLink, NVRTC PCH bindings: always present; only the library load is probed. - Checkpoint: the CUDA 13 build only (CUcheckpointGpuPair is a CUDA 13 type; the CUDA 12 bindings never had it), stated plainly. - cuda.core.system: NVML through cuda.bindings unconditionally; the driver/runtime fallbacks are gone and CUDA_BINDINGS_NVML_IS_COMPATIBLE is a deprecated constant True. system.typing exports DeviceArch and FieldId unconditionally. - Error-enum explanations come from cuda-bindings' enum docstrings; the frozen 13.1.1 tables and their loader are removed. - cuda.bindings.utils.warn_if_cuda_major_version_mismatch (13.3+) is called on the CUDA 13 build only, instead of try/except ImportError. cy_binding_version() has no callers left and is removed; binding_version() stays. Tests follow the same rule set.
NVIDIA#2783) test_checkpoint.py probed cuda.core.checkpoint._REQUIRED_BINDING_ATTRS, which the gate collapse removed, so the whole session failed at collection. The helpers build CUcheckpointGpuPair, a CUDA 13 type; skip them on the CUDA 12 build instead.
…untime-floor # Conflicts: # cuda_core/tests/system/test_system_events.py
…ts (NVIDIA#2783) stubgen-pyx regenerated version.pyi (BUILD_CUDA_MAJOR), system/_system.pyi (the constant is now True) and system/_device.pyi (NVLink 6.0 mapping is unconditional). ci/tools/cuda_core_bindings_floor.py reads the CUDA_BINDINGS_FLOOR literal with ast instead of executing the module; its tests and test_build_hooks.py load modules with importlib instead of exec, and the stubbed configuration check takes any arguments.
|
Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually. Contributors can view more details about this message here. |
|
/ok to test af88337 |
|
…A#1824 namespace repair (NVIDIA#2783) The build-configuration check imported cuda.bindings directly. In an isolated PEP 517 build with an in-tree backend (the wheel-from-sdist CI job), the project's own cuda/ directory is the whole `cuda` namespace, so the cuda-bindings pip installed into the build environment is not importable (NVIDIA#1824) and the check failed with "requires cuda-bindings to build". Import it the way _import_get_cuda_path_or_home() imports cuda.pathfinder: add the site-packages cuda/ directory to the namespace path when the plain import fails. The pxd-path block uses the same helper. Package metadata is not a substitute: pip's in-process hook runner forwards find_distributions without the requested name, so it returns this project's own metadata for `cuda-bindings`. test_managed_ops.py asserted the old wording of the NUMA-host message ("cuda-bindings 13.0+"), which now names the CUDA 13 build; the four regexes follow.
|
/ok to test b4fedec |
mdboom
left a comment
There was a problem hiding this comment.
Quick drive-by review. Looks good in general, but there are a few places where the floor version is specified (pyproject.toml, _bindings_floor.py etc.), and we'll need to keep it in sync.
If there is a way to get it to a single canonical version that gets pulled in everywhere, that would be preferable. That isn't always possible, so a precommit hook to confirm everything is synced, and maybe an AGENTS.md describing what needs to bump when we make a new major.minor cuda_bindings release would be helpful.
…untime-floor # Conflicts: # cuda_core/cuda/core/_linker.pyx # cuda_core/tests/graph/test_graph_builder.py # cuda_core/tests/test_optional_dependency_imports.py
…he Python bindings (NVIDIA#2783) NVIDIA#2867 fetched linked LTOIR through the Python-level cuda.bindings.nvjitlink module, probing it with hasattr for get_linked_ltoir_size/get_linked_ltoir because the cuda-bindings in use might predate them, and left a TODO(NVIDIA#2783) to switch once a floor existed. Both floors (12.9.8, 13.4.1) have the cynvjitlink getters, so Linker.link("ltoir") now calls nvJitLinkGetLinkedLTOIRSize/nvJitLinkGetLinkedLTOIR directly, like the cubin and ptx branches. The nvJitLink library-version gate (13.3) stays: that is the library, not the bindings. The cached module object is gone; only the cached library version remains. Tests that faked the module probe or asserted the cached module follow.
…ras (NVIDIA#2783) The floor of each CUDA major was typed in three places: a table in cuda/core/_bindings_floor.py, the cu12/cu13 extras in pyproject.toml, and the support-policy table in the docs, kept in step by tests that only ran in the GPU test jobs. The extras are now the single declaration, in the form `cuda-bindings[all]>=<floor>,==<major>.*`, and everything derives from them: - _bindings_floor.py keeps the logic only: floors_from_extras() parses the [project.optional-dependencies] table strictly (a malformed extra fails instead of shifting the floor), plus the version helpers and the import-time check, which now takes the build's recorded floor. - build_hooks.py reads pyproject.toml (tomllib; tomli on 3.10, added to build-system.requires) for the build-time check, the dynamic build requirement and -DCUDA_CORE_MIN_CUDA_VERSION, and records the floor in the generated _build_info.py, which cuda/core/__init__.py already reads. __init__.py no longer hard-codes the supported majors. - ci/tools/cuda_core_bindings_floor.py reads the floor out of the wheel's _build_info.py (per major in the merged wheel) instead of the table. - docs/source/conf.py defines |cuda-bindings-floor-cu12| and |cuda-bindings-floor-cu13| from pyproject.toml; support.rst uses them and labels the row with |release|. The stale cuda-bindings minimums in api_nvml.rst are gone: cuda.core.system has no requirement beyond the floor. - toolshed/check_cuda_core_bindings_floor.py, a new pre-commit hook, checks the two constraints that cannot be derived: ci/versions.yml must build each major against a toolkit of the floor's major.minor, and no documentation page (release notes excepted) may spell a floor out by hand. tests/test_bindings_floor.py runs the same checks. - AGENTS.md documents the coupling and a "Bumping the cuda-bindings floor" checklist.
… header rule (NVIDIA#2783) The build required cuda.h to match the installed cuda-bindings' version string by major.minor, and the import-time check demanded a cuda-bindings version at least as new as the build's header. Both fail on a cuda-bindings built from main right after a toolkit minor bump: its version string is the previous release's (13.4.3.devN) while it was generated from the new header. Compare headers instead: cuda.bindings.driver.CUDA_VERSION, present at both floors, is the CUDA_VERSION macro of the header cuda-bindings was generated from. The build check and check_installed_bindings() use it; the floor stays a version-string comparison; required_minimum() is gone. Isolated builds now request `cuda-bindings>=<floor>,==<major>.*,<<major>.<header minor + 1>` when cuda.h is readable, so pip cannot pick a newer minor than the toolkit, while the development build above (old version string, new header) still resolves and reaches the header rule. The pre-commit hook lets the CI toolkit run ahead of the floor's minor (the toolkit-bump window) and rejects only a toolkit below it. AGENTS.md documents the bump order and the accepted red window for the rows that install the literal floor. __init__.py reads cuda.bindings.driver by module object: the `cuda` namespace package carries no `bindings` attribute when the submodule comes from sys.modules. Found by the new child-interpreter test that imports cuda.core against fake cuda-bindings and asserts each ImportError, which issue NVIDIA#2783 asked for. The wrong-major message no longer promises a build for the installed major, and the pip commands use double quotes so they work in cmd.exe.
… attach the reason (NVIDIA#2783) A fill of the driver function table that failed was retried by every later DRIVER_CALL, re-importing cuda-bindings and re-warning each time, and its Python calls ran with whatever exception the caller had pending: a deleter making the first driver call during unwinding clobbered the user's exception. The failure is now latched per table (transient causes such as KeyboardInterrupt, MemoryError and RecursionError excepted), a PendingExceptionGuard saves and restores the pending exception around the fill, and fn_table_error() copies its text under the mutex. The CUDAError raised for the trampoline's CUDA_ERROR_NOT_INITIALIZED used to explain that cuInit() was not called. report_unavailable_fn() now records the fill's reason as the error detail (note_driver_table_failure), which the Cython error path attaches to the exception as a note, on every affected call; the fill itself warns once, and a null entry in a filled table (a gate bug) warns once per table. format_cuda_error() no longer null-checks the pointers, deviceptr_import_ipc() returns an error when the fill failed, and the graphics entry's requested version is annotated. The raw-pointer lint in test_rt_layout.py now catches null checks too. tests/test_driver_table.py runs a child interpreter whose _inspect_function_pointers() returns a table with a null baseline entry or a missing entry, and asserts the note, the single warning and the latch.
…text sweep (NVIDIA#2783) The API check gates "Check job status", so the stub of CUDA_BINDINGS_NVML_IS_COMPATIBLE must keep the bare `bool` annotation it had: the constant is declared with the annotation and assigned in a runtime-only block, which the stub generator does not record. It is always True and no longer called deprecated. DeviceArch is a plain enum.IntEnum instead of subclassing cuda-bindings' private FastEnum. Remaining text that named cuda-bindings versions as feature requirements now names the driver and the CUDA 13 build: the copy-options messages and docstrings in _buffer.pyx and _copy_ops.pyx, the graph node update docstrings, api.rst, the checkpoint section and module docstring, install.rst on conda-forge, DESIGN.md's summary, AGENTS.md's path to it, and the test helpers. The hasattr/getattr probes for NVML members present at both floors are gone, as is the probe-derived allow-list in test_enum_coverage.py.
|
/ok to test a37c10a |
rwgk
left a comment
There was a problem hiding this comment.
codex gpt-6-sol medium had no actionable findings.
I found a couple small things that should be easy to address.
| that names what was found and what is required. Building against an older CUDA Toolkit than | ||
| the floor's minor is not supported. | ||
| - **The CUDA driver** is unaffected. Feature availability is decided by the driver alone: a | ||
| feature the installed driver lacks raises when it is used, as before. |
There was a problem hiding this comment.
| feature the installed driver lacks raises when it is used, as before. | |
| feature the installed driver lacks raises when it is used. |
"as before" will be unclear outside the context of this PR. I think it'll be best to simply delete that part.
| ``cuda-bindings`` was generated from. Any other configuration fails the build with a message | ||
| that names what was found and what is required. Building against an older CUDA Toolkit than | ||
| the floor's minor is not supported. | ||
| - **The CUDA driver** is unaffected. Feature availability is decided by the driver alone: a |
There was a problem hiding this comment.
| - **The CUDA driver** is unaffected. Feature availability is decided by the driver alone: a | |
| - **The CUDA driver** is unaffected by the floor. Feature availability is decided by the driver alone: a |
| cu12 = ["cuda-bindings[all]>=12.9.8,==12.*", "cuda-toolkit==12.*"] | ||
| cu13 = ["cuda-bindings[all]>=13.4.1,==13.*", "cuda-toolkit==13.*"] |
There was a problem hiding this comment.
Could we spell these as >=12.9.8,<13 and >=13.4.1,<14? I find that more immediately understandable than >=...,==12.*. The floor reader currently requires the latter form. I'd favor adjusting it to accept the usual lower and upper bounds, while still checking that each cuN extra has a floor in major N and excludes major N+1. It needn't parse or validate more than that.
…lled <next major> (NVIDIA#2783) - cuda_bindings/build_hooks.py: before cythonize, compare the toolkit's cuda.h CUDA_VERSION with the one the sources were generated from (cydriver.pxd). A mismatch fails with a message that names the resolved cuda.h, both versions, the run-time compatibility, and the two remedies. Tests in cuda_bindings/tests/test_build_hooks.py; install page and 13.5.0 release note. - cuda.core build-time and import-time messages: name the resolved cuda.h, state what is needed and what was found, say that the rule does not constrain the CUDA driver or toolkit at run time, and link the support policy. - Review follow-ups: the cu12/cu13 extras read `>=<floor>,<<next major>`; the floor reader accepts `<N+1` or `==N.*` and checks that the bounds confine cuda-bindings to major N; bindings_requirement() takes an upper bound; the pip commands in the messages use the same form. Two wording fixes in support.rst. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…IDIA#2783) A plain-language pass over the prose the PR introduces: one idea per sentence, active voice, no semicolons or parenthetical asides, one term per concept (cuda-bindings for the package, cuda.bindings for the module). Error messages keep what they state and change only how they say it; the tests that match them follow. The regenerated stubs follow their docstrings. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
|
The latest upload addresses @rwgk's feedback (thanks), adds build-time checks for cuda-bindings, and includes a full copyediting sweep of the touched docs, docstrings, and error messages. |
|
/ok to test 1f05f42 |
| if found != needed: | ||
| raise RuntimeError( | ||
| f"This cuda-bindings source tree needs CUDA {needed} headers, but {_cuda_h_path(cuda_path)} is " | ||
| f"CUDA {found}. This is a build-time requirement only: at run time cuda-bindings supports any " |
There was a problem hiding this comment.
| f"CUDA {found}. This is a build-time requirement only: at run time cuda-bindings supports any " | |
| f"CUDA {found}. This is a build-time requirement only: at run time this cuda-bindings build can be used with any " |
Main concern: "supports" seems a tad too strong, e.g. I'd want to stay clear of leading someone to think that we're somehow supporting features added in toolkits with a minor version newer than the cuda-bindings minor version.
Minor concern: make "this cuda-bindings build" specific.
…DIA#2783) Per review: "supports any CUDA 13.x toolkit" overstates the claim, and "this cuda-bindings build" names what the note is about. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…untime-floor Conflicts: the toolchain and Cython-cache helpers from NVIDIA#2903 and NVIDIA#2933 landed next to this branch's floor and header-check helpers in both build_hooks.py files and their tests. Both sets are kept. cuda.core's build now stamps the build configuration (main) and checks the cuda-bindings floor and header (this branch) at the same point; the cuda-bindings tests file merges main's toolchain tests with this branch's header-check tests.
mdboom
left a comment
There was a problem hiding this comment.
As part of this rollout, we need to loosen the cuda-core requirement in cuda-python, otherwise installing cuda-python could be unresolvable by an update to cuda-core. I think we should remove the version specifier on cuda-core in cuda_python/setup.py altogether -- that is specify the cuda-python version exactly, but leave cuda-core floating.
| TEST_CUDA_MINOR="$(cut -d '.' -f 2 <<< ${CUDA_VER})" | ||
| # CI builds the prior-major half of the cuda-core wheel against the prev_build | ||
| # toolkit in ci/versions.yml and the backport branch's bindings. | ||
| BUILD_PREV_CUDA_VER="$(sed -n '/prev_build:/,/version:/s/.*version: *"\([^"]*\)".*/\1/p' ci/versions.yml)" |
There was a problem hiding this comment.
Should we use yq rather than sed here? (Just to be more robust against yaml syntax).
There was a problem hiding this comment.
Will do. yq is not currently used in the Linux builds, but I can add it.
| floor = load_build_module("_bindings_floor", cuda_major) | ||
| info = load_build_module("_build_info", cuda_major) |
There was a problem hiding this comment.
IIUC, _bindings_floor should never fail to import so probably shouldn't be inside the try/except.
The message for when _build_info.py is missing is a bit misleading -- isn't that the case when the package hasn't been built (e.g. importing directly from a clean checkout)?
There was a problem hiding this comment.
The import fails at top level because _bindings_floor.py was kept only under cu12/ and cu13/. I'll add it to the merge tool's keep list so that it appears at the top level and then move this import out of the try block.
I'll also separate the messages for "no build for this CUDA major in this installation" and "this checkout or editable install has not been built."
| _PACKAGE_DIR = Path(__file__).parent / "cuda" / "core" | ||
|
|
||
| # Generated at build time by _write_build_info() and read by cuda/core/__init__.py. | ||
| _BUILD_INFO_PATH = _PACKAGE_DIR / "_build_info.py" |
There was a problem hiding this comment.
This will break if two builds are run in parallel (for example for 12.x and 13.x). Generated files should be placed in the build/lib.* directory (which is accessible somehow through the build_hooks definition), not in the source directory.
But I also wonder if it wouldn't be cleaner to have a new Cython module that would import the actual cuda.h in question and expose the CUDA version that way. It would eliminate generating files, parsing headers, making sure the include patch matches exactly what the real Cython files were built with. The primary downside of that is the ci/tools/cuda_core_bindings_floor.py script would need to actually import this extension module rather than just unzipping the wheel and reading a Python source file as it currently does. I don't know if that would create some hazard that I'm not thinking of. But that overall seems safer -- it's testing what we really want to test (what CUDA version was this cuda-core version built with) rather than trying to recreate that value by mimicing the behavior of the build harness and the compiler etc. (and potentially getting that wrong).
There was a problem hiding this comment.
I see a problem with moving _build_info.py. Editable installs import it from the source tree, so it would have to be written in-tree for them and under build/lib.* for wheels, two code paths with tests for each. And it would not enable parallel builds: build/lib.<platform>-<python> and build/temp.* are shared across majors, and setuptools-scm rewrites _version.py in-tree on every build. (CI already builds the two CUDA majors in sequence.) I don't think the move is worth it.
There was a problem hiding this comment.
Yeah, I think your idea of a Cython module to eliminate the duplication will work. Trying it now.
There was a problem hiding this comment.
I've adopted the Cython module. cuda/core/_build_info.pyx now exposes CUDA_VERSION from the cuda.h the compiler resolved, plus the major and floor from compile_time_env. The generated _build_info.py is gone. The Python read of cuda.h stays for the decisions that precede any compile (which major to build when CUDA_CORE_BUILD_MAJOR is unset, the isolated-build cap on cuda-bindings, and the early mismatch message), and versions.hpp now fails the compile unless the compiler's cuda.h agrees with it by major.minor, so the two can never silently differ. ci/tools/cuda_core_bindings_floor.py no longer needs anything from inside the wheel's package: it reads the floor from the wheel's METADATA (Requires-Dist: cuda-bindings[all]<14,>=13.4.1; extra == "cu13").
…thon read; review follow-ups (NVIDIA#2783) - cuda/core/_build_info.pyx replaces the generated _build_info.py. It exposes CUDA_VERSION from the cuda.h the compiler resolved and the major and floor from the Cython compile-time environment, so the record cannot disagree with the binaries. cuda/core/__init__.py reads the extension. - build_hooks.py passes the cuda-bindings header version as a third macro; versions.hpp fails the compile unless the resolved cuda.h has the same major.minor. The Python read of cuda.h stays for the decisions that precede any compile, and can no longer silently differ from the compiler's. - ci/tools/cuda_core_bindings_floor.py reads the floor from the wheel's METADATA (the cu<major> extra) instead of parsing a generated module. - __init__.py imports _bindings_floor directly; the merge tool keeps it at the top level. A missing _build_info now distinguishes "no build for this CUDA major" from "not built". - ci/tools/env-vars reads ci/versions.yml with yq. - cuda_python/setup.py no longer pins cuda-core. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Done: cuda-core is unpinned in |
…untime-floor Conflicts in _rt.pxd and _rt.pyx: NVIDIA#2966 added noexcept to the sm_resource_split and memcpy_with_attributes_async wrappers, which this branch removed. The removal stands; NVIDIA#2966's performance-hints directive and its other changes are kept.
ci/tools/env-vars reads ci/versions.yml with yq, and the ubuntu:24.04 test container has none, so every Linux test job stopped at the environment step. A composite action installs the pinned, checksummed mikefarah/yq release for amd64 or arm64; the Windows GPU runners already provide yq. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
That's a good call-out -- that doesn't feel like a good user experience. Is there a better solution? The obvious solution is to always make a new cuda-python release whenever either cuda-bindings /or/ cuda-core is released -- I was trying to avoid that, but maybe that's not possible. EDIT: My agent suggests that leaving this as-is is probably a better choice in most cases. If the user specifies |
| # DeviceArch takes its values from cuda.bindings.nvml at definition time. | ||
| # It is an IntEnum rather than a StrEnum because the order of the values is | ||
| # meaningful, e.g. Kepler "or later". | ||
| class DeviceArch(enum.IntEnum): |
There was a problem hiding this comment.
This is a regression from _FastEnum to enum.IntEnum.
| # A restore onto other GPUs uses CUcheckpointGpuPair, a CUDA 13 type that | ||
| # cuda-bindings 12.x does not have. | ||
| if _BUILD_CUDA_MAJOR < 13: | ||
| raise RuntimeError("CUDA checkpointing requires the CUDA 13 build of cuda.core. Install cuda-core[cu13].") |
There was a problem hiding this comment.
In support.rst it says "Feature availability is decided by the driver alone: a feature the installed driver lacks raises when it is used.", but this is gating based on the build of cuda.core. (I think theoretically you could have the latest driver and cuda.core[cu12] and this feature isn't available even though the driver supports it). Not suggesting that we change how this works, but maybe just add some nuance to support.rst?
There was a problem hiding this comment.
Yeah, that's phrased poorly. It's trying to convey the fact that feature gates no longer double-check both the driver and cuda-bindings version at runtime. But we still have other feature gates in the cuda-core source itself.
| // true when the exception says nothing about cuda-bindings or the driver: an | ||
| // interruption such as KeyboardInterrupt or SystemExit, or exhaustion such as | ||
| // MemoryError or RecursionError. The fill reports such a failure but does not | ||
| // latch it, so the next call tries again. |
There was a problem hiding this comment.
No, "latch." I've reworded to clarify.
| && (!PyErr_ExceptionMatches(PyExc_Exception) || PyErr_ExceptionMatches(PyExc_MemoryError) | ||
| || PyErr_ExceptionMatches(PyExc_RecursionError)); |
There was a problem hiding this comment.
Will this end up swallowing a KeyboardInterrupt and other signals, and is that desirable here?
There was a problem hiding this comment.
Reworked in 5417ea7 to let signals pass through.
| jobs run in a plain ubuntu container). ci/tools/env-vars reads ci/versions.yml | ||
| with yq. A no-op when yq is already on PATH. Needs wget. | ||
|
|
||
| inputs: |
There was a problem hiding this comment.
Can't we just install yq with apt? We already have a step in test-wheel-linux to install dependencies with apt. (And it looks like Windows already has or doesn't need yq?)
| # cuda-bindings that cuda.core accepts has them. Keyed by ``str``: under | ||
| # ``python_version = "3.10"`` mypy resolves StrEnum to the unstubbed backports | ||
| # shim and so infers the members as plain ``str``. StrEnum members are ``str`` | ||
| # instances, so this holds on every version. The values are wrapped in | ||
| # ``int()`` because the driver enums are untyped. |
There was a problem hiding this comment.
Nit: This level of detail about making mypy happy seems unnecessary.
| nvml.BrandType.BRAND_NVIDIA_DLA: "NVIDIA DLA", | ||
| nvml.BrandType.BRAND_NVIDIA_VGAMEDEV: "NVIDIA vGameDev", | ||
| nvml.BrandType.BRAND_NVIDIA_NPU: "NVIDIA NPU", | ||
| }) |
There was a problem hiding this comment.
Nit: These can just be rolled into the dict literal above.
|
|
||
| _MODULES.append(system_typing) | ||
| # Every ClocksEventReasons member is mapped: the floor cuda-bindings has them all. | ||
| _CLOCKS_EVENT_REASONS_STR_UNMAPPED = set() |
There was a problem hiding this comment.
Nit: Don't need to have this as an explicit constant -- can just use set() inline now (as we do for other enums).
|
|
||
| 1. Bump the toolkit. | ||
| 2. Release `cuda-bindings` for the new minor. | ||
| 3. Bump the floor here. |
There was a problem hiding this comment.
Maybe someday we automate these steps, but not today.
| # Unpinned: cuda-core releases on its own cadence and declares its own | ||
| # cuda-bindings floors. A pin here would make older cuda-python releases | ||
| # unresolvable after a cuda-core release. | ||
| "cuda-core", |
There was a problem hiding this comment.
As mentioned in another comment, I think my suggestion may not be correct after all. Let's restore this to what it was and add an AGENTS.md and other documentation that this needs to be bumped when a cuda-core release is made. (And maybe a note about making a cuda-python immediately after each cuda-core release, though that's less critical).
…pt from the table fill; yq via apt; nits (NVIDIA#2783) - cuda_python/setup.py pins cuda-core~=1.2.0 again. The bump at each cuda-core release, followed by a cuda-python release, is documented in .github/RELEASE-core.md, cuda_python/AGENTS.md and cuda_core/AGENTS.md. - A KeyboardInterrupt raised inside the function-table fill is no longer swallowed. The fill re-arms it with PyErr_SetInterrupt() after its warning, and report_message() does the same when an interrupt fires inside the warnings machinery, so the user sees the interrupt at the next bytecode boundary and the warning survives. Empty exception text falls back to the type name. Test: cuda_core/tests/test_driver_table.py, through the NVRTC table, CPU-only. - The Linux GPU test containers get yq from apt (the Python jq wrapper, which accepts the same command); the composite action is removed. - DeviceArch is a FastEnum again (present at both floors). - support.rst: the driver bullet states what the floor does not change and what an older driver means. Small comment and release-note fixes. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
|
All of @mdboom's review feedback is addressed in 5417ea7 (merged with
|
Resolves the conflicts with NVIDIA#2920 (driver calls through cuda-bindings' resolved pointers): - driver_api.hpp: the VMM entry points join the X-macro table with the CUDA versions cuda-bindings requests them at; driver_api.cpp and the pointer-initialization block of _rt.pyx take main's side, since the table is now filled from cuda-bindings. - virtual_memory.cpp: raw p_ calls become DRIVER_CALL, and the cuStreamGetCaptureInfo arity fence branches on CUDA_CORE_BUILD_MAJOR. - py_report.cpp: PendingExceptionGuard plus main's KeyboardInterrupt handling; the identical local guard in py_driver_fns.cpp is replaced by the shared one in py.hpp. - _virtual_memory_resource.py was deleted on this branch; its one change on main (BUILD_CUDA_MAJOR instead of binding_version()) is applied to the .pyx.
Implements the plan in #2783:
cuda.corerequires a per-majorcuda-bindingsfloor, and with that floor in place it stops working around oldercuda-bindingsreleases.What changes
1. A
cuda-bindingsfloor per CUDA major, at build and run time.cuda/core/_bindings_floor.pyrecords 12.9.8 for CUDA 12 and 13.4.1 for CUDA 13.build_hooks.pyrefuses to build against an oldercuda-bindings, and thecuda.core._build_infoextension records the header version the compiler compiled against.import cuda.corechecks the installedcuda-bindingsagainst the build's major and minimum before importing any extension, and fails with a message that names the version found, the version required, and thepipcommand that fixes it. Thecu12/cu13extras pin the floors.2. The header rule. A source build requires a
cuda.hof the same major.minor as itscuda-bindings(the headercuda-bindingswas generated from). Any other configuration fails the build with a clear message. This is what makes the pointer table below sound:cuda-bindingskeys its function table by the symbolcuda.hmaps each name to (cuStreamDestroy->cuStreamDestroy_v2), so both sides must see the same header.3. C++ fences on the CUDA major only.
build_hooks.pypasses-DCUDA_CORE_BUILD_MAJORand-DCUDA_CORE_MIN_CUDA_VERSION; the new_cpp/rt/versions.hppre-checks the header against both. Every#if CUDA_VERSION >= 130x0fence is gone;tests/test_rt_layout.pyenforces thatversions.hppis the only file under_cpp/that namesCUDA_VERSION.4. Driver calls through
cuda-bindings' resolved pointers. The C++ under_cpp/rt/used to call the driver throughcuda-bindings' Cython wrappers, extracted fromcydriver.__pyx_capi__. A wrapper for a function the driver lacks raises a Python exception the C++ never sees.driver_api.hppnow lists the 61 driver functions the C++ calls in an X-macro table, each with the CUDA versioncuda-bindingsrequests it at. The table is filled on first use fromcuda.bindings._internal.driver._inspect_function_pointers(), which holds the driver's own entry points, soimport cuda.corenever loads the driver. Calls go throughDRIVER_CALL(name, args...); a pointer still null after the fill (a missing feature gate, or a failed fill) is reported once through the warning path and returns an error status from a trampoline of the right signature, never a null dereference and never an exception. The fill rejects a driver older than the CUDA major series. Thepw_deleter wrappers use the same path, so the C++ never null-checks a pointer; functions the installed driver may lack are gated on the driver version in Cython. NVRTC, NVVM and nvJitLink each have their own one-entry table.5. Feature gates depend on the driver alone. Checks that also inspected the
cuda-bindingsversion are gone: double checks (driver + bindings) became driver-only checks, checks of the bindings minor became build-major fences, and checks every acceptedcuda-bindingssatisfies were deleted along with the fallbacks they guarded.cuGraphNodeGetParamsnow goes throughcydriverinstead of the Python driver layer. The frozen error-explanation tables, the NVML fallbacks incuda.core.system, and the NVVM, nvJitLink and NVRTC PCH probes are removed.User-visible changes
import cuda.corewith acuda-bindingsolder than the floor fails with an actionableImportError. Previously it imported and failed later (missing C function, silently disabled feature, or crash).cuda-bindingsversion now name a driver version.cuda.core.system.CUDA_BINDINGS_NVML_IS_COMPATIBLEis alwaysTrueand deprecated.cuda.core.checkpointstates that it requires the CUDA 13 build. It did in effect before: the CUDA 12cuda-bindingslackCUcheckpointGpuPair.cuda.core.system.typingexportsDeviceArchandFieldIdunconditionally.CUDA_ERROR_NOT_SUPPORTEDfallback in C++).Supported CUDA drivers and CUDA Toolkit libraries are unchanged. Release notes:
docs/source/release/1.3.0-notes.rst.CI
ci/tools/env-varsreplaces thepublishedbindings source withfloor: the prior-major test rows that do not match the backport branch's minor installcuda-bindings==<floor>(read from the wheel byci/tools/cuda_core_bindings_floor.py), so the floor itself is tested.tests/test_bindings_floor.pychecks thatpyproject.tomlandci/versions.ymlagree with_bindings_floor.py.Testing
cuda_coretest suite on a local 2-GPU system (H20, driver 615.71) againstcuda-bindings13.4.1: 4307 passed, 121 skipped, 5 failed. The five failures are test-infrastructure issues unrelated to this change:test_device_cpu_affinityenumerates NVML devices where the process can only see a subset of the GPUs (fixed on main by tests: enumerate CUDA devices for GPU-only system checks #2916, merged here), and threetest_object_code_load_rdc*cases whose fixture builds object files inside the checkout, where a file mirror on the test host removes them mid-build.import cuda.coredoes not maplibcuda;Device.set_current()does; all 61 table entries resolve to non-null driver entry points.test_bindings_floor.py,test_build_hooks.py(build configuration checks, requirement string, define macros),test_rt_layout.py(table agrees withcuda-bindings' loader; no rawp_calls outside the table machinery),ci/tools/tests/test_cuda_core_bindings_floor.py.Notes for review
_cpp/rt/DESIGN.mdandAGENTS.mddescribe the pointer table and settle the GIL contract for the C++ entry points: they work with or without the GIL, never take a C++ lock while acquiring it, and only the reporting wrappers and the one-time table fill acquire it.deviceptr_import_ipcresolves the table before takingipc_import_mutexand keeps raw pointer calls under the lock, marked// raw:and linted.TODO(#2783)in_linker.pyxbecomes a direct cimport. cuda.core: fix pool setup and builder teardown under stream capture #2838 touches_rt.pyx/graph.cppand needs a rebase after this lands.cuda-bindingsper the floors (out of this repo).Update (merge of main and review follow-ups)
main(feat(cuda.core): support nvJitLink incremental linking #2867, cuda.core: fix pool setup and builder teardown under stream capture #2838, Use default setuptools-scm tag parsing #2922, lock refreshes). feat(cuda.core): support nvJitLink incremental linking #2867's linked-LTOIR output went through the Pythoncuda.bindings.nvjitlinkmodule withhasattrprobes and aTODO(#2783); both floors have thecynvjitlinkgetters, soLinker.link("ltoir")now calls them directly like the cubin and ptx branches. The nvJitLink 13.3 library-version gate stays.cu12/cu13extras inpyproject.toml(cuda-bindings[all]>=<floor>,<<next major>), per the review.build_hooks.pyreads them (viatomllib;tomlion 3.10 is a build requirement) and passes the floor into the Cython compile-time environment; thecuda.core._build_infoextension records it next to theCUDA_VERSIONof thecuda.hthe compiler resolved, and the import-time check reads that record.cuda/core/_bindings_floor.pykeeps only the logic. The docs read the floors into|cuda-bindings-floor-cu12|/|cuda-bindings-floor-cu13|substitutions and label the support table with|release|;ci/tools/cuda_core_bindings_floor.pyreads the floor from the wheel's METADATA.cuda.hagainstcuda.bindings.driver.CUDA_VERSION, the header the installed cuda-bindings was generated from, and the import-time check does the same against the build's recorded header. A cuda-bindings built frommainright after a toolkit minor bump carries the previous release's version string but the new header, so both checks pass for it. The rows that install the literal floor stay red until a cuda-bindings release of the new minor exists; that window is accepted and documented inAGENTS.md. Isolated source builds now requestcuda-bindings>=<floor>,<<toolkit major>.<toolkit minor + 1>, so pip cannot pick a newer minor than the toolkit.CUDAErrorwith the reason attached as a note instead of a generic "not initialized" explanation; the fill preserves a pending Python exception (a deleter can make the first driver call while an exception propagates); the IPC import returns an error when the fill failed.tests/test_driver_table.pyexercises the failed-fill path in a child interpreter with a fake_inspect_function_pointers.CUDA_BINDINGS_NVML_IS_COMPATIBLEstub keeps its bareboolannotation (the value is assigned in a runtime-only block the stub generator does not record), so griffe reports no change. The constant is alwaysTrue; it is not called deprecated.check-cuda-core-bindings-floor(toolshed/check_cuda_core_bindings_floor.py), checks the two constraints that cannot be derived:ci/versions.ymlmust build each major against a toolkit of the floor's major.minor, and no documentation page (release notes excepted) may spell a floor out by hand.tests/test_bindings_floor.pyruns the same checks and adds a child-interpreter test that importscuda.coreagainst a fake below-floor cuda-bindings and asserts the message.AGENTS.mddocuments the coupling, a "Bumping the cuda-bindings floor" checklist, and the toolkit-bump order. Per-feature cuda-bindings minimums are gone fromapi.rst,api_nvml.rst,_buffer.pyx,_copy_ops.pyxand the test helpers;DeviceArchis a plainIntEnum.cuda_core/pixi.lockcu12 cannot satisfy the 12.9.8 floor until conda-forge ships it), older-driver CI rows, and the CUDA 13 floor bump to 13.4.2 (release-time item; the build environments still carry 13.4.1).🤖 Generated with Claude Code
cuda/core/_build_info.pyxexposesCUDA_VERSIONfrom thecuda.hthe compiler resolved, and the major and floor from the compile-time environment; the generated_build_info.pyis gone.versions.hppfails the compile unless the compiler'scuda.hmatches the header the installed cuda-bindings was generated from, so the Python read ofcuda.h(needed before any compile) can never silently differ from the compiler's._bindings_floorstays at the top level of the merged wheel and is imported directly; a missing build distinguishes "no build for this CUDA major" from "not built".ci/tools/env-varsreadsci/versions.ymlwithyq(installed from apt in the Linux test containers).cuda_python/setup.pykeeps itscuda-core~=X.Y.0pin; the bump at eachcuda-corerelease, followed by acuda-pythonrelease, is documented in.github/RELEASE-core.mdand theAGENTS.mdfiles. AKeyboardInterruptraised inside the function-table fill is re-delivered to the user instead of being swallowed by the noexcept reporting path.Error messages
Toolkit header does not match cuda-bindings (cuda.core build):
cuda-bindings older than the headers cuda.core was compiled against (import):
cuda-bindings below the floor (import):
cuda-bindings of another major (import):
Toolkit header does not match the generated sources (cuda-bindings build):
Closes #2783
Closes #2699
Supersedes #2700