Skip to content

feat(cuda.core): support nvJitLink incremental linking - #2867

Merged
leofang merged 10 commits into
NVIDIA:mainfrom
isVoid:fea-2369-draft
Sep 19, 2026
Merged

leofang merged 10 commits into
NVIDIA:mainfrom
isVoid:fea-2369-draft

Conversation

@isVoid

@isVoid isVoid commented Sep 15, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • expose nvJitLink incremental linking as LinkerOptions(incremental=True), gated on nvJitLink 13.2+
  • return native partial results as ObjectCode(code_type="cubin") so they round-trip directly into another Linker
  • add Linker.link("ltoir") on nvJitLink 13.3+ so incremental LTO chains can retain IR until final code generation
  • reject unsupported driver-backend and partial-PTX combinations with explicit errors
  • clarify that code_type="object" means a host object containing device code

Design choices

  • A native incremental-link result remains "cubin"; ELF ET_REL describes link state, not a distinct cuda.core transport format.
  • In incremental mode, target_type accepts "cubin", or "ltoir" when LTO is enabled and the getter is available. "ptx" is rejected.
  • Native CUBIN and linked-LTOIR outputs need no wrapper or conversion before being supplied to a later Linker.
  • Intermediate LTO stages should request "ltoir"; requesting "cubin" creates a machine-code boundary that later LTO cannot optimize through.
  • Linked-LTOIR retrieval uses the Python binding dynamically so cuda-core continues to build with supported cuda-bindings 12.x releases whose Cython declarations predate the getter.

Testing

  • CUDA 13.4 linker suite: 82 passed, 3 skipped, including executable native-CUBIN and LTOIR three-stage round trips
  • CUDA 12.9 / published cuda-bindings 12.9.7 compatibility build and linker suite: 74 passed, 6 capability skips
  • Sphinx documentation build
  • repository pre-commit checks for all changed files (Ruff, mypy, Cython lint, generated stubs, SPDX, RST, memory-pool hygiene)

Closes #2369

@copy-pr-bot

copy-pr-bot Bot commented Sep 15, 2026

Copy link
Copy Markdown
Contributor

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

@github-actions github-actions Bot added the cuda.core Everything related to the cuda.core module label Sep 15, 2026
@leofang

leofang commented Sep 15, 2026 •

Copy link
Copy Markdown
Member

nit: @isVoid could you post your design doc as a comment in #2369 for posterity, in case your personal branch is removed in the future?

@leofang
leofang self-requested a review September 15, 2026 19:36
@isVoid

isVoid commented Sep 17, 2026

Copy link
Copy Markdown
Contributor Author

nit: @isVoid could you post your design doc as a comment in #2369 for posterity, in case your personal branch is removed in the future?

This is now posted.

@isVoid
isVoid marked this pull request as ready for review September 17, 2026 20:09
Comment thread cuda_core/cuda/core/_linker.pyi Outdated
link_time_optimization : bool, optional
Perform link time optimization.
Default: False.
relocatable : bool, optional

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Q: Maybe we should call it relocatable_device_code to match what we have in ProgramOptions? Seems like they have the same semantics?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The design doc notes the distinction:

ProgramOptions.relocatable_device_code and the proposed LinkerOptions.relocatable are also distinct:

  • The program option makes a compilation unit suitable for device linking.
  • The linker option permits a link result to remain relocatable and to retain unresolved device references.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I sort of think relocatable_device_code and relocatable are respectively different notions created one on the nvrtc and nvjitlink side. My prior discussion with folks suggests rdc is quite commonly used as a shorthand to refer to the nvrtc behavior. Coercing the linker side to overlap this terminology could introduce some confusion in the future.

ref:
nvjitlink linker options
nvrtc options

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think I was confused for a moment when reading even relocatable_device_code and relocatable 😂

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@isVoid good idea, but maybe we should rename it for clarity? Such as incremental_link?

@isVoid isVoid Sep 18, 2026 •

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

+1. Like incremental_link

Maybe the page didn't refresh when you reply, see #2867 (comment). Do you have specific preference on incremental_link over incremental? @lijinf2

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes I think page did not refresh on my side. incremental looks good too!

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Don't the driver docs prefer the term "relative" over "incremental" though? Diverging from that could be confusing in a different way.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Don't the driver docs prefer the term "relative" over "incremental" though? Diverging from that could be confusing in a different way.

I can see the term relocatable is used across the board to refer specifically to the output of the linker with -r or compiler with -rdc. But when referring to the -r flag, the nvjitlink do mention both terms, and actually in some sense highlighting the incremental to emphasize the linker behavior.

@Andy-Jost Andy-Jost left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One blocker: link("ltoir") silently drops inputs that carry no LTOIR (details inline, with a repro). Two untested rows of the design matrix, and a few small cleanups. The relocatable and LTOIR round trips themselves look right.

Comment thread cuda_core/cuda/core/_linker.pyx
Comment thread cuda_core/tests/test_linker.py Outdated
Comment thread cuda_core/cuda/core/_linker.pyx Outdated
Comment thread cuda_core/cuda/core/_linker.pyx Outdated
Comment thread cuda_core/cuda/core/_linker.pyx
Comment thread cuda_core/cuda/core/_linker.pyi Outdated
link_time_optimization : bool, optional
Perform link time optimization.
Default: False.
relocatable : bool, optional

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The design doc notes the distinction:

ProgramOptions.relocatable_device_code and the proposed LinkerOptions.relocatable are also distinct:

  • The program option makes a compilation unit suitable for device linking.
  • The linker option permits a link result to remain relocatable and to retain unresolved device references.

@lijinf2 lijinf2 left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nice. Let me write down my understanding. Before this PR, application code had to gather all compiled device inputs (PTX/LTOIR/cubin) before invoking a single Linker. After this PR, application code can do cascading linking — each stage produces a partial result that feeds directly into the next Linker.

Comment thread cuda_core/tests/test_linker.py Outdated
Comment thread cuda_core/tests/test_linker.py
Comment thread cuda_core/cuda/core/_linker.pyx Outdated
Comment thread cuda_core/cuda/core/_linker.pyi Outdated
link_time_optimization : bool, optional
Perform link time optimization.
Default: False.
relocatable : bool, optional

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think I was confused for a moment when reading even relocatable_device_code and relocatable 😂

@isVoid isVoid added the feature New feature or request label Sep 18, 2026
@isVoid isVoid added this to the cuda.core 1.3.0 milestone Sep 18, 2026
@isVoid isVoid added the P1 Medium priority - Should do label Sep 18, 2026
@isVoid
isVoid requested a review from Andy-Jost September 18, 2026 06:02

@Andy-Jost Andy-Jost left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

Comment thread cuda_core/cuda/core/LINKER_INCREMENTAL_LINKING_DESIGN.md Outdated
@leofang

leofang commented Sep 18, 2026

Copy link
Copy Markdown
Member

Left a question here: #2867 (comment)

@isVoid

isVoid commented Sep 18, 2026

Copy link
Copy Markdown
Contributor Author

/ok to test 3cc270d

@lijinf2 lijinf2 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

My comments all addressed. Thanks. PR looks good to me!

@github-actions

This comment has been minimized.

@isVoid

isVoid commented Sep 19, 2026

Copy link
Copy Markdown
Contributor Author

/ok to test fc8c850

@isVoid
isVoid enabled auto-merge (squash) September 19, 2026 00:46
@isVoid
isVoid disabled auto-merge September 19, 2026 00:46
@isVoid
isVoid requested a review from leofang September 19, 2026 00:46
@leofang
leofang merged commit 8b6e9f5 into NVIDIA:main Sep 19, 2026
115 checks passed
@github-actions

This comment has been minimized.

1 similar comment
@github-actions

Copy link
Copy Markdown
Contributor
Doc Preview CI
Preview removed because the pull request was closed or merged.

Andy-Jost added a commit to Andy-Jost/cuda-python that referenced this pull request Sep 22, 2026
…he Python bindings (NVIDIA#2783)

NVIDIA#2867 fetched linked LTOIR through the Python-level cuda.bindings.nvjitlink
module, probing it with hasattr for get_linked_ltoir_size/get_linked_ltoir
because the cuda-bindings in use might predate them, and left a
TODO(NVIDIA#2783) to switch once a floor existed. Both floors (12.9.8, 13.4.1)
have the cynvjitlink getters, so Linker.link("ltoir") now calls
nvJitLinkGetLinkedLTOIRSize/nvJitLinkGetLinkedLTOIR directly, like the
cubin and ptx branches. The nvJitLink library-version gate (13.3) stays:
that is the library, not the bindings. The cached module object is gone;
only the cached library version remains. Tests that faked the module
probe or asserted the cached module follow.
Andy-Jost added a commit that referenced this pull request Oct 1, 2026
…its resolved pointers (#2783) (#2920)

* cuda.core: require a per-major cuda-bindings floor at build and run time

cuda.core accepted any cuda-bindings of the right major at build time and at
run time, and built against whatever cuda.h was on the include path. Builds
succeeded for configurations we never test, and an older cuda-bindings at run
time surfaced as an ImportError for a missing C function, a silently disabled
feature, or a null-pointer crash (#2783).

Each supported CUDA major now has a floor, the newest cuda-bindings release its
CI source root can build: 12.9.8 for CUDA 12 and 13.4.1 for CUDA 13, in the
import-free module cuda/core/_bindings_floor.py.

Build time: the pip build requirement becomes `cuda-bindings>=<floor>,==<major>.*`,
and because conda-forge, pixi and --no-build-isolation installs bypass it, the
backend itself checks the imported cuda-bindings against the floor and requires
the cuda.h it compiles against to have the same major.minor as that
cuda-bindings (the header cuda-bindings was generated from), even when
CUDA_CORE_BUILD_MAJOR is set. It records the header version and the floor in
the generated, gitignored cuda/core/_build_info.py, shipped like _version.py.

Run time: cuda/core/__init__.py reads that record from the selected build (the
cu12/cu13 subpackage of the merged wheel, or the top level of a plain build) and
requires the installed cuda-bindings to be of the build's major and at least the
floor, or at least the header's minor when that is newer. The error names the
version found, the version required, and the pip command that fixes it.
Pre-release and dev builds of an accepted version pass.

Packaging: the cu12/cu13 extras pin the floor; a test keeps them and the
ci/versions.yml toolkit pins in step with the module.

CI: the BINDINGS_SOURCE=published rows, which paired a new wheel with
cuda-bindings 13.0 (no longer supported), become BINDINGS_SOURCE=floor: they
install the floor bindings, read from the wheel under test by
ci/tools/cuda_core_bindings_floor.py, and keep their older CTK libraries. The
prior-major rows whose CTK minor differs from prev_build do the same.

Docs: support policy section, install guide, 1.3.0 release note.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* cuda.core: fence C++ on the CUDA major only, never on CUDA_VERSION

The two `#if CUDA_VERSION >= 130x0` fences in _cpp/rt/driver_api.* compiled
cuDevSmResourceSplit and cuMemcpyWithAttributesAsync out of a source build
against an older 13.x header, while the run-time gates, which looked at the
bindings and the driver, still reported the features as available; the calls
then failed with CUDA_ERROR_NOT_SUPPORTED (#2783, "guards disable working
features"). Both features also went through void*-typed C++ shims and has_*()
presence probes whose only purpose was to avoid an eager cimport of a cydriver
function that cuda-bindings 13.0 did not export.

With the cuda-bindings floor in place (13.4.1 for CUDA 13), the cimport is
safe, so the shims, the presence probes and the two fenced pointers go away:
_device_resources.pyx and _buffer.pyx call cydriver.cuDevSmResourceSplit and
cydriver.cuMemcpyWithAttributesAsync directly under the existing
`IF CUDA_CORE_BUILD_MAJOR >= 13`, gated on the driver version alone.

build_hooks.py now defines CUDA_CORE_BUILD_MAJOR (until now a Cython-only
compile-time constant) and CUDA_CORE_MIN_CUDA_VERSION (the floor's
major.minor) for the C++ compiler. The new _cpp/rt/versions.hpp, the first
include of the tree via types.hpp, re-checks cuda.h against both with #error,
so a build that bypasses the backend still cannot compile against an
unsupported header. tests/test_rt_layout.py enforces that versions.hpp is the
only file under _cpp/ that names CUDA_VERSION.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* cuda.core: call the driver through cuda-bindings' resolved pointers (#2783)

The C++ under _cpp/rt/ used to call the driver through cuda-bindings'
Cython wrappers, extracted from cydriver.__pyx_capi__ at import. A
wrapper for a function the driver lacks raises a Python exception the
C++ never sees, and a wrapper the installed cuda-bindings lacks made the
pointer optional and probed for null at every use.

driver_api.hpp now lists every driver function the C++ calls, with the
CUDA version cuda-bindings requests it at. The table is filled on first
use from cuda.bindings._internal.driver._inspect_function_pointers(),
which holds the driver's own entry points, so `import cuda.core` never
touches the driver. Calls go through DRIVER_CALL(name, args...), which
fills the table if needed and, when a pointer is still null after the
fill (a missing feature gate or a failed fill), reports once and returns
an error status from a trampoline instead of dereferencing null. The fill
rejects a driver older than the CUDA major series. pw_ deleter wrappers
use the same path, so the C++ never null-checks a pointer; functions the
driver may lack are gated on the driver version in Cython. NVRTC, NVVM
and nvJitLink each have their own one-entry table, filled when a program
or linker handle is created.

Consequences in this commit: the null-check fallbacks to
CUDA_ERROR_NOT_SUPPORTED are gone; green-context stream creation is
gated in _stream.pyx on driver 12.5; deviceptr_import_ipc resolves the
table before taking ipc_import_mutex and keeps raw calls under the lock
(marked `// raw:`); _rt.pyx no longer imports cydriver/cynvrtc/cynvvm/
cynvjitlink or reads __pyx_capi__. test_rt_layout.py checks the table
against cuda-bindings' loader and forbids raw p_ calls elsewhere.
DESIGN.md and AGENTS.md describe the mechanism and settle the GIL
contract for entry points.

* cuda.core: gate features on the driver alone now that cuda-bindings has a floor (#2783)

Every cuda-bindings cuda.core accepts (12.9.8+, 13.4.1+) has the full
API surface of its CUDA major, so a run-time check of the cuda-bindings
version says nothing a build-major fence does not. Double checks
(driver and bindings) become driver-only checks; checks of the bindings
minor become `IF CUDA_CORE_BUILD_MAJOR` fences in Cython or a comparison
with the new `cuda.core._utils.version.BUILD_CUDA_MAJOR` in Python;
checks that every accepted cuda-bindings satisfies are deleted, along
with the fallbacks they guarded.

- Kernel argument info, green contexts, workqueues, graph node updates,
  conditional graph nodes: driver-only gates.
- cuGraphNodeGetParams (13.2): a driver gate on the CUDA 13 build, and
  the call goes through cydriver instead of the Python driver layer.
- Managed memory NUMA locations, virtual memory MANAGED type: build
  fence (plus the 13.0 driver where a driver entry point is involved).
- Copy attribute enums (CUDA 12.8): unconditional.
- NVVM, nvJitLink, NVRTC PCH bindings: always present; only the library
  load is probed.
- Checkpoint: the CUDA 13 build only (CUcheckpointGpuPair is a CUDA 13
  type; the CUDA 12 bindings never had it), stated plainly.
- cuda.core.system: NVML through cuda.bindings unconditionally; the
  driver/runtime fallbacks are gone and CUDA_BINDINGS_NVML_IS_COMPATIBLE
  is a deprecated constant True. system.typing exports DeviceArch and
  FieldId unconditionally.
- Error-enum explanations come from cuda-bindings' enum docstrings; the
  frozen 13.1.1 tables and their loader are removed.
- cuda.bindings.utils.warn_if_cuda_major_version_mismatch (13.3+) is
  called on the CUDA 13 build only, instead of try/except ImportError.

cy_binding_version() has no callers left and is removed; binding_version()
stays. Tests follow the same rule set.

* cuda.core tests: gate the checkpoint helper tests on the CUDA 13 build (#2783)

test_checkpoint.py probed cuda.core.checkpoint._REQUIRED_BINDING_ATTRS, which
the gate collapse removed, so the whole session failed at collection. The
helpers build CUcheckpointGpuPair, a CUDA 13 type; skip them on the CUDA 12
build instead.

* cuda.core: regenerate stubs and drop exec from the floor tool and tests (#2783)

stubgen-pyx regenerated version.pyi (BUILD_CUDA_MAJOR), system/_system.pyi
(the constant is now True) and system/_device.pyi (NVLink 6.0 mapping is
unconditional). ci/tools/cuda_core_bindings_floor.py reads the
CUDA_BINDINGS_FLOOR literal with ast instead of executing the module; its
tests and test_build_hooks.py load modules with importlib instead of exec,
and the stubbed configuration check takes any arguments.

* cuda.core: import cuda.bindings for the build check through the #1824 namespace repair (#2783)

The build-configuration check imported cuda.bindings directly. In an
isolated PEP 517 build with an in-tree backend (the wheel-from-sdist CI
job), the project's own cuda/ directory is the whole `cuda` namespace,
so the cuda-bindings pip installed into the build environment is not
importable (#1824) and the check failed with "requires cuda-bindings to
build". Import it the way _import_get_cuda_path_or_home() imports
cuda.pathfinder: add the site-packages cuda/ directory to the namespace
path when the plain import fails. The pxd-path block uses the same
helper. Package metadata is not a substitute: pip's in-process hook
runner forwards find_distributions without the requested name, so it
returns this project's own metadata for `cuda-bindings`.

test_managed_ops.py asserted the old wording of the NUMA-host message
("cuda-bindings 13.0+"), which now names the CUDA 13 build; the four
regexes follow.

* cuda.core: read linked LTOIR through cynvjitlink instead of probing the Python bindings (#2783)

#2867 fetched linked LTOIR through the Python-level cuda.bindings.nvjitlink
module, probing it with hasattr for get_linked_ltoir_size/get_linked_ltoir
because the cuda-bindings in use might predate them, and left a
TODO(#2783) to switch once a floor existed. Both floors (12.9.8, 13.4.1)
have the cynvjitlink getters, so Linker.link("ltoir") now calls
nvJitLinkGetLinkedLTOIRSize/nvJitLinkGetLinkedLTOIR directly, like the
cubin and ptx branches. The nvJitLink library-version gate (13.3) stays:
that is the library, not the bindings. The cached module object is gone;
only the cached library version remains. Tests that faked the module
probe or asserted the cached module follow.

* cuda.core: declare the cuda-bindings floor once, in the pyproject extras (#2783)

The floor of each CUDA major was typed in three places: a table in
cuda/core/_bindings_floor.py, the cu12/cu13 extras in pyproject.toml, and
the support-policy table in the docs, kept in step by tests that only ran
in the GPU test jobs. The extras are now the single declaration, in the
form `cuda-bindings[all]>=<floor>,==<major>.*`, and everything derives
from them:

- _bindings_floor.py keeps the logic only: floors_from_extras() parses
  the [project.optional-dependencies] table strictly (a malformed extra
  fails instead of shifting the floor), plus the version helpers and the
  import-time check, which now takes the build's recorded floor.
- build_hooks.py reads pyproject.toml (tomllib; tomli on 3.10, added to
  build-system.requires) for the build-time check, the dynamic build
  requirement and -DCUDA_CORE_MIN_CUDA_VERSION, and records the floor in
  the generated _build_info.py, which cuda/core/__init__.py already reads.
  __init__.py no longer hard-codes the supported majors.
- ci/tools/cuda_core_bindings_floor.py reads the floor out of the wheel's
  _build_info.py (per major in the merged wheel) instead of the table.
- docs/source/conf.py defines |cuda-bindings-floor-cu12| and
  |cuda-bindings-floor-cu13| from pyproject.toml; support.rst uses them
  and labels the row with |release|. The stale cuda-bindings minimums in
  api_nvml.rst are gone: cuda.core.system has no requirement beyond the
  floor.
- toolshed/check_cuda_core_bindings_floor.py, a new pre-commit hook, checks
  the two constraints that cannot be derived: ci/versions.yml must build
  each major against a toolkit of the floor's major.minor, and no
  documentation page (release notes excepted) may spell a floor out by
  hand. tests/test_bindings_floor.py runs the same checks.
- AGENTS.md documents the coupling and a "Bumping the cuda-bindings floor"
  checklist.

* cuda.core: compare headers, not version strings, in the cuda-bindings header rule (#2783)

The build required cuda.h to match the installed cuda-bindings' version
string by major.minor, and the import-time check demanded a cuda-bindings
version at least as new as the build's header. Both fail on a cuda-bindings
built from main right after a toolkit minor bump: its version string is the
previous release's (13.4.3.devN) while it was generated from the new header.
Compare headers instead: cuda.bindings.driver.CUDA_VERSION, present at both
floors, is the CUDA_VERSION macro of the header cuda-bindings was generated
from. The build check and check_installed_bindings() use it; the floor stays
a version-string comparison; required_minimum() is gone.

Isolated builds now request `cuda-bindings>=<floor>,==<major>.*,<<major>.<header minor + 1>`
when cuda.h is readable, so pip cannot pick a newer minor than the toolkit,
while the development build above (old version string, new header) still
resolves and reaches the header rule. The pre-commit hook lets the CI toolkit
run ahead of the floor's minor (the toolkit-bump window) and rejects only a
toolkit below it. AGENTS.md documents the bump order and the accepted red
window for the rows that install the literal floor.

__init__.py reads cuda.bindings.driver by module object: the `cuda`
namespace package carries no `bindings` attribute when the submodule comes
from sys.modules. Found by the new child-interpreter test that imports
cuda.core against fake cuda-bindings and asserts each ImportError, which
issue #2783 asked for. The wrong-major message no longer promises a build
for the installed major, and the pip commands use double quotes so they
work in cmd.exe.

* cuda.core: latch a failed driver-table fill, keep pending exceptions, attach the reason (#2783)

A fill of the driver function table that failed was retried by every later
DRIVER_CALL, re-importing cuda-bindings and re-warning each time, and its
Python calls ran with whatever exception the caller had pending: a deleter
making the first driver call during unwinding clobbered the user's exception.
The failure is now latched per table (transient causes such as
KeyboardInterrupt, MemoryError and RecursionError excepted), a
PendingExceptionGuard saves and restores the pending exception around the
fill, and fn_table_error() copies its text under the mutex.

The CUDAError raised for the trampoline's CUDA_ERROR_NOT_INITIALIZED used to
explain that cuInit() was not called. report_unavailable_fn() now records the
fill's reason as the error detail (note_driver_table_failure), which the
Cython error path attaches to the exception as a note, on every affected
call; the fill itself warns once, and a null entry in a filled table (a gate
bug) warns once per table. format_cuda_error() no longer null-checks the
pointers, deviceptr_import_ipc() returns an error when the fill failed, and
the graphics entry's requested version is annotated. The raw-pointer lint in
test_rt_layout.py now catches null checks too.

tests/test_driver_table.py runs a child interpreter whose
_inspect_function_pointers() returns a table with a null baseline entry or a
missing entry, and asserts the note, the single warning and the latch.

* cuda.core: keep the NVML constant's stub annotation-only; finish the text sweep (#2783)

The API check gates "Check job status", so the stub of
CUDA_BINDINGS_NVML_IS_COMPATIBLE must keep the bare `bool` annotation it had:
the constant is declared with the annotation and assigned in a runtime-only
block, which the stub generator does not record. It is always True and no
longer called deprecated. DeviceArch is a plain enum.IntEnum instead of
subclassing cuda-bindings' private FastEnum.

Remaining text that named cuda-bindings versions as feature requirements now
names the driver and the CUDA 13 build: the copy-options messages and
docstrings in _buffer.pyx and _copy_ops.pyx, the graph node update
docstrings, api.rst, the checkpoint section and module docstring, install.rst
on conda-forge, DESIGN.md's summary, AGENTS.md's path to it, and the test
helpers. The hasattr/getattr probes for NVML members present at both floors
are gone, as is the probe-derived allow-list in test_enum_coverage.py.

* cuda-bindings header check at build; final floor messages; bounds spelled <next major> (#2783)

- cuda_bindings/build_hooks.py: before cythonize, compare the toolkit's
  cuda.h CUDA_VERSION with the one the sources were generated from
  (cydriver.pxd). A mismatch fails with a message that names the resolved
  cuda.h, both versions, the run-time compatibility, and the two remedies.
  Tests in cuda_bindings/tests/test_build_hooks.py; install page and 13.5.0
  release note.
- cuda.core build-time and import-time messages: name the resolved cuda.h,
  state what is needed and what was found, say that the rule does not
  constrain the CUDA driver or toolkit at run time, and link the support
  policy.
- Review follow-ups: the cu12/cu13 extras read `>=<floor>,<<next major>`;
  the floor reader accepts `<N+1` or `==N.*` and checks that the bounds
  confine cuda-bindings to major N; bindings_requirement() takes an upper
  bound; the pip commands in the messages use the same form. Two wording
  fixes in support.rst.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* Copyedit the docs, comments, docstrings and messages this PR adds (#2783)

A plain-language pass over the prose the PR introduces: one idea per
sentence, active voice, no semicolons or parenthetical asides, one term per
concept (cuda-bindings for the package, cuda.bindings for the module).
Error messages keep what they state and change only how they say it; the
tests that match them follow. The regenerated stubs follow their docstrings.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* cuda-bindings header check: review wording for the run-time note (#2783)

Per review: "supports any CUDA 13.x toolkit" overstates the claim, and
"this cuda-bindings build" names what the note is about.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* Compile the build record; verify the compiler's cuda.h against the Python read; review follow-ups (#2783)

- cuda/core/_build_info.pyx replaces the generated _build_info.py. It exposes
  CUDA_VERSION from the cuda.h the compiler resolved and the major and floor
  from the Cython compile-time environment, so the record cannot disagree with
  the binaries. cuda/core/__init__.py reads the extension.
- build_hooks.py passes the cuda-bindings header version as a third macro;
  versions.hpp fails the compile unless the resolved cuda.h has the same
  major.minor. The Python read of cuda.h stays for the decisions that precede
  any compile, and can no longer silently differ from the compiler's.
- ci/tools/cuda_core_bindings_floor.py reads the floor from the wheel's
  METADATA (the cu<major> extra) instead of parsing a generated module.
- __init__.py imports _bindings_floor directly; the merge tool keeps it at
  the top level. A missing _build_info now distinguishes "no build for this
  CUDA major" from "not built".
- ci/tools/env-vars reads ci/versions.yml with yq.
- cuda_python/setup.py no longer pins cuda-core.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* ci: install yq in the Linux GPU test containers

ci/tools/env-vars reads ci/versions.yml with yq, and the ubuntu:24.04 test
container has none, so every Linux test job stopped at the environment step.
A composite action installs the pinned, checksummed mikefarah/yq release for
amd64 or arm64; the Windows GPU runners already provide yq.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* Address review: restore the cuda-core pin; re-deliver KeyboardInterrupt from the table fill; yq via apt; nits (#2783)

- cuda_python/setup.py pins cuda-core~=1.2.0 again. The bump at each cuda-core
  release, followed by a cuda-python release, is documented in
  .github/RELEASE-core.md, cuda_python/AGENTS.md and cuda_core/AGENTS.md.
- A KeyboardInterrupt raised inside the function-table fill is no longer
  swallowed. The fill re-arms it with PyErr_SetInterrupt() after its warning,
  and report_message() does the same when an interrupt fires inside the
  warnings machinery, so the user sees the interrupt at the next bytecode
  boundary and the warning survives. Empty exception text falls back to the
  type name. Test: cuda_core/tests/test_driver_table.py, through the NVRTC
  table, CPU-only.
- The Linux GPU test containers get yq from apt (the Python jq wrapper, which
  accepts the same command); the composite action is removed.
- DeviceArch is a FastEnum again (present at both floors).
- support.rst: the driver bullet states what the floor does not change and
  what an older driver means. Small comment and release-note fixes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

cuda.core Everything related to the cuda.core module feature New feature or request P1 Medium priority - Should do

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[FEA] cuda.core: support nvJitLink partial/incremental linking (CUDA 13.2)

4 participants