feat(cuda.core): support nvJitLink incremental linking - #2867
Conversation
| link_time_optimization : bool, optional | ||
| Perform link time optimization. | ||
| Default: False. | ||
| relocatable : bool, optional |
There was a problem hiding this comment.
Q: Maybe we should call it relocatable_device_code to match what we have in ProgramOptions? Seems like they have the same semantics?
There was a problem hiding this comment.
The design doc notes the distinction:
ProgramOptions.relocatable_device_codeand the proposedLinkerOptions.relocatableare also distinct:
- The program option makes a compilation unit suitable for device linking.
- The linker option permits a link result to remain relocatable and to retain unresolved device references.
There was a problem hiding this comment.
I sort of think relocatable_device_code and relocatable are respectively different notions created one on the nvrtc and nvjitlink side. My prior discussion with folks suggests rdc is quite commonly used as a shorthand to refer to the nvrtc behavior. Coercing the linker side to overlap this terminology could introduce some confusion in the future.
There was a problem hiding this comment.
I think I was confused for a moment when reading even relocatable_device_code and relocatable 😂
There was a problem hiding this comment.
@isVoid good idea, but maybe we should rename it for clarity? Such as incremental_link?
There was a problem hiding this comment.
+1. Like incremental_link
Maybe the page didn't refresh when you reply, see #2867 (comment). Do you have specific preference on incremental_link over incremental? @lijinf2
There was a problem hiding this comment.
Yes I think page did not refresh on my side. incremental looks good too!
There was a problem hiding this comment.
Don't the driver docs prefer the term "relative" over "incremental" though? Diverging from that could be confusing in a different way.
There was a problem hiding this comment.
Don't the driver docs prefer the term "relative" over "incremental" though? Diverging from that could be confusing in a different way.
I can see the term relocatable is used across the board to refer specifically to the output of the linker with -r or compiler with -rdc. But when referring to the -r flag, the nvjitlink do mention both terms, and actually in some sense highlighting the incremental to emphasize the linker behavior.
Andy-Jost
left a comment
There was a problem hiding this comment.
One blocker: link("ltoir") silently drops inputs that carry no LTOIR (details inline, with a repro). Two untested rows of the design matrix, and a few small cleanups. The relocatable and LTOIR round trips themselves look right.
| link_time_optimization : bool, optional | ||
| Perform link time optimization. | ||
| Default: False. | ||
| relocatable : bool, optional |
There was a problem hiding this comment.
The design doc notes the distinction:
ProgramOptions.relocatable_device_codeand the proposedLinkerOptions.relocatableare also distinct:
- The program option makes a compilation unit suitable for device linking.
- The linker option permits a link result to remain relocatable and to retain unresolved device references.
There was a problem hiding this comment.
Nice. Let me write down my understanding. Before this PR, application code had to gather all compiled device inputs (PTX/LTOIR/cubin) before invoking a single Linker. After this PR, application code can do cascading linking — each stage produces a partial result that feeds directly into the next Linker.
| link_time_optimization : bool, optional | ||
| Perform link time optimization. | ||
| Default: False. | ||
| relocatable : bool, optional |
There was a problem hiding this comment.
I think I was confused for a moment when reading even relocatable_device_code and relocatable 😂
|
Left a question here: #2867 (comment) |
|
/ok to test 3cc270d |
lijinf2
left a comment
There was a problem hiding this comment.
My comments all addressed. Thanks. PR looks good to me!
This comment has been minimized.
This comment has been minimized.
|
/ok to test fc8c850 |
This comment has been minimized.
This comment has been minimized.
1 similar comment
|
…he Python bindings (NVIDIA#2783) NVIDIA#2867 fetched linked LTOIR through the Python-level cuda.bindings.nvjitlink module, probing it with hasattr for get_linked_ltoir_size/get_linked_ltoir because the cuda-bindings in use might predate them, and left a TODO(NVIDIA#2783) to switch once a floor existed. Both floors (12.9.8, 13.4.1) have the cynvjitlink getters, so Linker.link("ltoir") now calls nvJitLinkGetLinkedLTOIRSize/nvJitLinkGetLinkedLTOIR directly, like the cubin and ptx branches. The nvJitLink library-version gate (13.3) stays: that is the library, not the bindings. The cached module object is gone; only the cached library version remains. Tests that faked the module probe or asserted the cached module follow.
…its resolved pointers (#2783) (#2920) * cuda.core: require a per-major cuda-bindings floor at build and run time cuda.core accepted any cuda-bindings of the right major at build time and at run time, and built against whatever cuda.h was on the include path. Builds succeeded for configurations we never test, and an older cuda-bindings at run time surfaced as an ImportError for a missing C function, a silently disabled feature, or a null-pointer crash (#2783). Each supported CUDA major now has a floor, the newest cuda-bindings release its CI source root can build: 12.9.8 for CUDA 12 and 13.4.1 for CUDA 13, in the import-free module cuda/core/_bindings_floor.py. Build time: the pip build requirement becomes `cuda-bindings>=<floor>,==<major>.*`, and because conda-forge, pixi and --no-build-isolation installs bypass it, the backend itself checks the imported cuda-bindings against the floor and requires the cuda.h it compiles against to have the same major.minor as that cuda-bindings (the header cuda-bindings was generated from), even when CUDA_CORE_BUILD_MAJOR is set. It records the header version and the floor in the generated, gitignored cuda/core/_build_info.py, shipped like _version.py. Run time: cuda/core/__init__.py reads that record from the selected build (the cu12/cu13 subpackage of the merged wheel, or the top level of a plain build) and requires the installed cuda-bindings to be of the build's major and at least the floor, or at least the header's minor when that is newer. The error names the version found, the version required, and the pip command that fixes it. Pre-release and dev builds of an accepted version pass. Packaging: the cu12/cu13 extras pin the floor; a test keeps them and the ci/versions.yml toolkit pins in step with the module. CI: the BINDINGS_SOURCE=published rows, which paired a new wheel with cuda-bindings 13.0 (no longer supported), become BINDINGS_SOURCE=floor: they install the floor bindings, read from the wheel under test by ci/tools/cuda_core_bindings_floor.py, and keep their older CTK libraries. The prior-major rows whose CTK minor differs from prev_build do the same. Docs: support policy section, install guide, 1.3.0 release note. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * cuda.core: fence C++ on the CUDA major only, never on CUDA_VERSION The two `#if CUDA_VERSION >= 130x0` fences in _cpp/rt/driver_api.* compiled cuDevSmResourceSplit and cuMemcpyWithAttributesAsync out of a source build against an older 13.x header, while the run-time gates, which looked at the bindings and the driver, still reported the features as available; the calls then failed with CUDA_ERROR_NOT_SUPPORTED (#2783, "guards disable working features"). Both features also went through void*-typed C++ shims and has_*() presence probes whose only purpose was to avoid an eager cimport of a cydriver function that cuda-bindings 13.0 did not export. With the cuda-bindings floor in place (13.4.1 for CUDA 13), the cimport is safe, so the shims, the presence probes and the two fenced pointers go away: _device_resources.pyx and _buffer.pyx call cydriver.cuDevSmResourceSplit and cydriver.cuMemcpyWithAttributesAsync directly under the existing `IF CUDA_CORE_BUILD_MAJOR >= 13`, gated on the driver version alone. build_hooks.py now defines CUDA_CORE_BUILD_MAJOR (until now a Cython-only compile-time constant) and CUDA_CORE_MIN_CUDA_VERSION (the floor's major.minor) for the C++ compiler. The new _cpp/rt/versions.hpp, the first include of the tree via types.hpp, re-checks cuda.h against both with #error, so a build that bypasses the backend still cannot compile against an unsupported header. tests/test_rt_layout.py enforces that versions.hpp is the only file under _cpp/ that names CUDA_VERSION. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * cuda.core: call the driver through cuda-bindings' resolved pointers (#2783) The C++ under _cpp/rt/ used to call the driver through cuda-bindings' Cython wrappers, extracted from cydriver.__pyx_capi__ at import. A wrapper for a function the driver lacks raises a Python exception the C++ never sees, and a wrapper the installed cuda-bindings lacks made the pointer optional and probed for null at every use. driver_api.hpp now lists every driver function the C++ calls, with the CUDA version cuda-bindings requests it at. The table is filled on first use from cuda.bindings._internal.driver._inspect_function_pointers(), which holds the driver's own entry points, so `import cuda.core` never touches the driver. Calls go through DRIVER_CALL(name, args...), which fills the table if needed and, when a pointer is still null after the fill (a missing feature gate or a failed fill), reports once and returns an error status from a trampoline instead of dereferencing null. The fill rejects a driver older than the CUDA major series. pw_ deleter wrappers use the same path, so the C++ never null-checks a pointer; functions the driver may lack are gated on the driver version in Cython. NVRTC, NVVM and nvJitLink each have their own one-entry table, filled when a program or linker handle is created. Consequences in this commit: the null-check fallbacks to CUDA_ERROR_NOT_SUPPORTED are gone; green-context stream creation is gated in _stream.pyx on driver 12.5; deviceptr_import_ipc resolves the table before taking ipc_import_mutex and keeps raw calls under the lock (marked `// raw:`); _rt.pyx no longer imports cydriver/cynvrtc/cynvvm/ cynvjitlink or reads __pyx_capi__. test_rt_layout.py checks the table against cuda-bindings' loader and forbids raw p_ calls elsewhere. DESIGN.md and AGENTS.md describe the mechanism and settle the GIL contract for entry points. * cuda.core: gate features on the driver alone now that cuda-bindings has a floor (#2783) Every cuda-bindings cuda.core accepts (12.9.8+, 13.4.1+) has the full API surface of its CUDA major, so a run-time check of the cuda-bindings version says nothing a build-major fence does not. Double checks (driver and bindings) become driver-only checks; checks of the bindings minor become `IF CUDA_CORE_BUILD_MAJOR` fences in Cython or a comparison with the new `cuda.core._utils.version.BUILD_CUDA_MAJOR` in Python; checks that every accepted cuda-bindings satisfies are deleted, along with the fallbacks they guarded. - Kernel argument info, green contexts, workqueues, graph node updates, conditional graph nodes: driver-only gates. - cuGraphNodeGetParams (13.2): a driver gate on the CUDA 13 build, and the call goes through cydriver instead of the Python driver layer. - Managed memory NUMA locations, virtual memory MANAGED type: build fence (plus the 13.0 driver where a driver entry point is involved). - Copy attribute enums (CUDA 12.8): unconditional. - NVVM, nvJitLink, NVRTC PCH bindings: always present; only the library load is probed. - Checkpoint: the CUDA 13 build only (CUcheckpointGpuPair is a CUDA 13 type; the CUDA 12 bindings never had it), stated plainly. - cuda.core.system: NVML through cuda.bindings unconditionally; the driver/runtime fallbacks are gone and CUDA_BINDINGS_NVML_IS_COMPATIBLE is a deprecated constant True. system.typing exports DeviceArch and FieldId unconditionally. - Error-enum explanations come from cuda-bindings' enum docstrings; the frozen 13.1.1 tables and their loader are removed. - cuda.bindings.utils.warn_if_cuda_major_version_mismatch (13.3+) is called on the CUDA 13 build only, instead of try/except ImportError. cy_binding_version() has no callers left and is removed; binding_version() stays. Tests follow the same rule set. * cuda.core tests: gate the checkpoint helper tests on the CUDA 13 build (#2783) test_checkpoint.py probed cuda.core.checkpoint._REQUIRED_BINDING_ATTRS, which the gate collapse removed, so the whole session failed at collection. The helpers build CUcheckpointGpuPair, a CUDA 13 type; skip them on the CUDA 12 build instead. * cuda.core: regenerate stubs and drop exec from the floor tool and tests (#2783) stubgen-pyx regenerated version.pyi (BUILD_CUDA_MAJOR), system/_system.pyi (the constant is now True) and system/_device.pyi (NVLink 6.0 mapping is unconditional). ci/tools/cuda_core_bindings_floor.py reads the CUDA_BINDINGS_FLOOR literal with ast instead of executing the module; its tests and test_build_hooks.py load modules with importlib instead of exec, and the stubbed configuration check takes any arguments. * cuda.core: import cuda.bindings for the build check through the #1824 namespace repair (#2783) The build-configuration check imported cuda.bindings directly. In an isolated PEP 517 build with an in-tree backend (the wheel-from-sdist CI job), the project's own cuda/ directory is the whole `cuda` namespace, so the cuda-bindings pip installed into the build environment is not importable (#1824) and the check failed with "requires cuda-bindings to build". Import it the way _import_get_cuda_path_or_home() imports cuda.pathfinder: add the site-packages cuda/ directory to the namespace path when the plain import fails. The pxd-path block uses the same helper. Package metadata is not a substitute: pip's in-process hook runner forwards find_distributions without the requested name, so it returns this project's own metadata for `cuda-bindings`. test_managed_ops.py asserted the old wording of the NUMA-host message ("cuda-bindings 13.0+"), which now names the CUDA 13 build; the four regexes follow. * cuda.core: read linked LTOIR through cynvjitlink instead of probing the Python bindings (#2783) #2867 fetched linked LTOIR through the Python-level cuda.bindings.nvjitlink module, probing it with hasattr for get_linked_ltoir_size/get_linked_ltoir because the cuda-bindings in use might predate them, and left a TODO(#2783) to switch once a floor existed. Both floors (12.9.8, 13.4.1) have the cynvjitlink getters, so Linker.link("ltoir") now calls nvJitLinkGetLinkedLTOIRSize/nvJitLinkGetLinkedLTOIR directly, like the cubin and ptx branches. The nvJitLink library-version gate (13.3) stays: that is the library, not the bindings. The cached module object is gone; only the cached library version remains. Tests that faked the module probe or asserted the cached module follow. * cuda.core: declare the cuda-bindings floor once, in the pyproject extras (#2783) The floor of each CUDA major was typed in three places: a table in cuda/core/_bindings_floor.py, the cu12/cu13 extras in pyproject.toml, and the support-policy table in the docs, kept in step by tests that only ran in the GPU test jobs. The extras are now the single declaration, in the form `cuda-bindings[all]>=<floor>,==<major>.*`, and everything derives from them: - _bindings_floor.py keeps the logic only: floors_from_extras() parses the [project.optional-dependencies] table strictly (a malformed extra fails instead of shifting the floor), plus the version helpers and the import-time check, which now takes the build's recorded floor. - build_hooks.py reads pyproject.toml (tomllib; tomli on 3.10, added to build-system.requires) for the build-time check, the dynamic build requirement and -DCUDA_CORE_MIN_CUDA_VERSION, and records the floor in the generated _build_info.py, which cuda/core/__init__.py already reads. __init__.py no longer hard-codes the supported majors. - ci/tools/cuda_core_bindings_floor.py reads the floor out of the wheel's _build_info.py (per major in the merged wheel) instead of the table. - docs/source/conf.py defines |cuda-bindings-floor-cu12| and |cuda-bindings-floor-cu13| from pyproject.toml; support.rst uses them and labels the row with |release|. The stale cuda-bindings minimums in api_nvml.rst are gone: cuda.core.system has no requirement beyond the floor. - toolshed/check_cuda_core_bindings_floor.py, a new pre-commit hook, checks the two constraints that cannot be derived: ci/versions.yml must build each major against a toolkit of the floor's major.minor, and no documentation page (release notes excepted) may spell a floor out by hand. tests/test_bindings_floor.py runs the same checks. - AGENTS.md documents the coupling and a "Bumping the cuda-bindings floor" checklist. * cuda.core: compare headers, not version strings, in the cuda-bindings header rule (#2783) The build required cuda.h to match the installed cuda-bindings' version string by major.minor, and the import-time check demanded a cuda-bindings version at least as new as the build's header. Both fail on a cuda-bindings built from main right after a toolkit minor bump: its version string is the previous release's (13.4.3.devN) while it was generated from the new header. Compare headers instead: cuda.bindings.driver.CUDA_VERSION, present at both floors, is the CUDA_VERSION macro of the header cuda-bindings was generated from. The build check and check_installed_bindings() use it; the floor stays a version-string comparison; required_minimum() is gone. Isolated builds now request `cuda-bindings>=<floor>,==<major>.*,<<major>.<header minor + 1>` when cuda.h is readable, so pip cannot pick a newer minor than the toolkit, while the development build above (old version string, new header) still resolves and reaches the header rule. The pre-commit hook lets the CI toolkit run ahead of the floor's minor (the toolkit-bump window) and rejects only a toolkit below it. AGENTS.md documents the bump order and the accepted red window for the rows that install the literal floor. __init__.py reads cuda.bindings.driver by module object: the `cuda` namespace package carries no `bindings` attribute when the submodule comes from sys.modules. Found by the new child-interpreter test that imports cuda.core against fake cuda-bindings and asserts each ImportError, which issue #2783 asked for. The wrong-major message no longer promises a build for the installed major, and the pip commands use double quotes so they work in cmd.exe. * cuda.core: latch a failed driver-table fill, keep pending exceptions, attach the reason (#2783) A fill of the driver function table that failed was retried by every later DRIVER_CALL, re-importing cuda-bindings and re-warning each time, and its Python calls ran with whatever exception the caller had pending: a deleter making the first driver call during unwinding clobbered the user's exception. The failure is now latched per table (transient causes such as KeyboardInterrupt, MemoryError and RecursionError excepted), a PendingExceptionGuard saves and restores the pending exception around the fill, and fn_table_error() copies its text under the mutex. The CUDAError raised for the trampoline's CUDA_ERROR_NOT_INITIALIZED used to explain that cuInit() was not called. report_unavailable_fn() now records the fill's reason as the error detail (note_driver_table_failure), which the Cython error path attaches to the exception as a note, on every affected call; the fill itself warns once, and a null entry in a filled table (a gate bug) warns once per table. format_cuda_error() no longer null-checks the pointers, deviceptr_import_ipc() returns an error when the fill failed, and the graphics entry's requested version is annotated. The raw-pointer lint in test_rt_layout.py now catches null checks too. tests/test_driver_table.py runs a child interpreter whose _inspect_function_pointers() returns a table with a null baseline entry or a missing entry, and asserts the note, the single warning and the latch. * cuda.core: keep the NVML constant's stub annotation-only; finish the text sweep (#2783) The API check gates "Check job status", so the stub of CUDA_BINDINGS_NVML_IS_COMPATIBLE must keep the bare `bool` annotation it had: the constant is declared with the annotation and assigned in a runtime-only block, which the stub generator does not record. It is always True and no longer called deprecated. DeviceArch is a plain enum.IntEnum instead of subclassing cuda-bindings' private FastEnum. Remaining text that named cuda-bindings versions as feature requirements now names the driver and the CUDA 13 build: the copy-options messages and docstrings in _buffer.pyx and _copy_ops.pyx, the graph node update docstrings, api.rst, the checkpoint section and module docstring, install.rst on conda-forge, DESIGN.md's summary, AGENTS.md's path to it, and the test helpers. The hasattr/getattr probes for NVML members present at both floors are gone, as is the probe-derived allow-list in test_enum_coverage.py. * cuda-bindings header check at build; final floor messages; bounds spelled <next major> (#2783) - cuda_bindings/build_hooks.py: before cythonize, compare the toolkit's cuda.h CUDA_VERSION with the one the sources were generated from (cydriver.pxd). A mismatch fails with a message that names the resolved cuda.h, both versions, the run-time compatibility, and the two remedies. Tests in cuda_bindings/tests/test_build_hooks.py; install page and 13.5.0 release note. - cuda.core build-time and import-time messages: name the resolved cuda.h, state what is needed and what was found, say that the rule does not constrain the CUDA driver or toolkit at run time, and link the support policy. - Review follow-ups: the cu12/cu13 extras read `>=<floor>,<<next major>`; the floor reader accepts `<N+1` or `==N.*` and checks that the bounds confine cuda-bindings to major N; bindings_requirement() takes an upper bound; the pip commands in the messages use the same form. Two wording fixes in support.rst. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * Copyedit the docs, comments, docstrings and messages this PR adds (#2783) A plain-language pass over the prose the PR introduces: one idea per sentence, active voice, no semicolons or parenthetical asides, one term per concept (cuda-bindings for the package, cuda.bindings for the module). Error messages keep what they state and change only how they say it; the tests that match them follow. The regenerated stubs follow their docstrings. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * cuda-bindings header check: review wording for the run-time note (#2783) Per review: "supports any CUDA 13.x toolkit" overstates the claim, and "this cuda-bindings build" names what the note is about. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * Compile the build record; verify the compiler's cuda.h against the Python read; review follow-ups (#2783) - cuda/core/_build_info.pyx replaces the generated _build_info.py. It exposes CUDA_VERSION from the cuda.h the compiler resolved and the major and floor from the Cython compile-time environment, so the record cannot disagree with the binaries. cuda/core/__init__.py reads the extension. - build_hooks.py passes the cuda-bindings header version as a third macro; versions.hpp fails the compile unless the resolved cuda.h has the same major.minor. The Python read of cuda.h stays for the decisions that precede any compile, and can no longer silently differ from the compiler's. - ci/tools/cuda_core_bindings_floor.py reads the floor from the wheel's METADATA (the cu<major> extra) instead of parsing a generated module. - __init__.py imports _bindings_floor directly; the merge tool keeps it at the top level. A missing _build_info now distinguishes "no build for this CUDA major" from "not built". - ci/tools/env-vars reads ci/versions.yml with yq. - cuda_python/setup.py no longer pins cuda-core. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * ci: install yq in the Linux GPU test containers ci/tools/env-vars reads ci/versions.yml with yq, and the ubuntu:24.04 test container has none, so every Linux test job stopped at the environment step. A composite action installs the pinned, checksummed mikefarah/yq release for amd64 or arm64; the Windows GPU runners already provide yq. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * Address review: restore the cuda-core pin; re-deliver KeyboardInterrupt from the table fill; yq via apt; nits (#2783) - cuda_python/setup.py pins cuda-core~=1.2.0 again. The bump at each cuda-core release, followed by a cuda-python release, is documented in .github/RELEASE-core.md, cuda_python/AGENTS.md and cuda_core/AGENTS.md. - A KeyboardInterrupt raised inside the function-table fill is no longer swallowed. The fill re-arms it with PyErr_SetInterrupt() after its warning, and report_message() does the same when an interrupt fires inside the warnings machinery, so the user sees the interrupt at the next bytecode boundary and the warning survives. Empty exception text falls back to the type name. Test: cuda_core/tests/test_driver_table.py, through the NVRTC table, CPU-only. - The Linux GPU test containers get yq from apt (the Python jq wrapper, which accepts the same command); the composite action is removed. - DeviceArch is a FastEnum again (present at both floors). - support.rst: the driver bullet states what the floor does not change and what an older driver means. Small comment and release-note fixes. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Summary
LinkerOptions(incremental=True), gated on nvJitLink 13.2+ObjectCode(code_type="cubin")so they round-trip directly into anotherLinkerLinker.link("ltoir")on nvJitLink 13.3+ so incremental LTO chains can retain IR until final code generationcode_type="object"means a host object containing device codeDesign choices
"cubin"; ELFET_RELdescribes link state, not a distinct cuda.core transport format.target_typeaccepts"cubin", or"ltoir"when LTO is enabled and the getter is available."ptx"is rejected.Linker."ltoir"; requesting"cubin"creates a machine-code boundary that later LTO cannot optimize through.Testing
Closes #2369