Conversation
Introduces cuda_bindings/_toolchain_shared.py as the single source of truth for the CUDA_PYTHON_TOOLCHAIN resolution helpers, with cuda_core mirroring it via a symlink. Both build_hooks.py files now import the shared helpers instead of carrying a byte-for-byte duplicated block, following up on the approval condition in review comment #pullrequestreview-5355844534. - New cuda_bindings/_toolchain_shared.py (canonical: constants, _resolve_toolchain_name, _apply_toolchain_env, _check_toolchain_available). - cuda_core/_toolchain_shared.py is a symlink to ../cuda_bindings/. Both packages' pyproject.toml already sets backend-path = ["."], so `from _toolchain_shared import ...` at the top of each build_hooks.py resolves via each backend's own directory. - Both MANIFEST.in include _toolchain_shared.py so sdists carry the resolved content: setuptools' make_release_tree copies through symlinks, so an extracted sdist gets a real 3099-byte file rather than a dangling link. - Tests pre-load _toolchain_shared into sys.modules before executing build_hooks via importlib, so the shared import resolves without adding the package directory to sys.path (which would shadow the installed cuda.bindings / cuda.core for the rest of the test session). - toolshed/check_build_hooks_sync.py and its pre-commit entry are gone; a single source of truth needs no drift check. Windows contributors outside of WSL now need Git to be configured for symlinks before cloning. A new "Development on Windows" section in CONTRIBUTING.md covers Developer Mode + core.symlinks=true (and folds in the existing lychee workaround). cuda_bindings and cuda_core install docs cross-link to it from "Installing from Source".
Match the CuPy contribution guide's phrasing and target URL so Windows contributors land on the same official Microsoft page.
…_path into shared file Approach A follow-up to the toolchain-block extraction: pull in the remaining verbatim duplicates between the two build_hooks.py files, and rename the shared file from _toolchain_shared.py to _build_shared.py to reflect its broader scope. cuda_core/_build_shared.py remains a symlink to the canonical cuda_bindings/_build_shared.py. Moved into _build_shared.py: - _import_get_cuda_path_or_home (PEP 517 pathfinder-shadowing workaround, the second helper the earlier PR flagged as prior art). - _get_cuda_path (@functools.cache wrapper around the above). - _BUILD_DIR and _abi_stamp_path (extension-ABI-scoped stamp helper). _BUILD_DIR relies on Python not resolving symlinks in __file__: when cuda_core imports the shared module through its symlink, Path(__file__) points at cuda_core/, so _BUILD_DIR is cuda_core/build/; imported from cuda_bindings, it's cuda_bindings/build/. Each package's stamps still land under its own build/ directory. Both build_hooks.py files now import these names from _build_shared and drop their local copies plus the now-unused import shutil (core), import functools + from pathlib import Path (bindings), and import sysconfig (core). Tests patch shutil / sysconfig at their module level instead of via build_hooks.shutil / build_hooks.sysconfig, since neither is needed on build_hooks after the extraction. Both suites still pass (15 bindings, 49 core). ruff.toml gets a T201 exemption for **/_build_shared.py so the print diagnostics inherited from the previous build_hooks.py copies keep working (build_hooks.py already has this exemption).
Resolves conflicts with NVIDIA#2933 (Cython generated-source cache). Ralf's _cython_cache_path and _stable_cython_alias helpers were vendored in both build_hooks.py files behind the # --- begin/end shared build helpers markers his PR renamed. Moved both into cuda_bindings/_build_shared.py alongside the toolchain block; both build_hooks.py now import them via `from _build_shared import ...` (no vendoring, no sync check). Conflict resolution: - cuda_bindings/build_hooks.py, cuda_core/build_hooks.py: keep the from-_build_shared import structure; extend the import list with _cython_cache_path and _stable_cython_alias. - cuda_bindings/_build_shared.py: add the two helpers plus the imports they need (contextlib, hashlib, uuid, warn). - cuda_bindings/build_hooks.py: drop hashlib (moved), keep contextlib (still used for contextlib.suppress), re-add Path from pathlib for the Cython.__file__ path handling in _build_cuda_bindings. - cuda_core/build_hooks.py: drop contextlib, hashlib, uuid, warn (all moved). - .pre-commit-config.yaml: keep my removal of the check-build-hooks-sync entry (script is deleted). - toolshed/check_build_hooks_sync.py: keep my deletion (single source of truth needs no drift check). - Auto-merged test files needed no manual touch; both suites now include Ralf's cython_cache tests and pass (33 bindings, 71 core). Verified after resolution: - ruff check + format: clean. - SPDX: clean. - Sdists for both packages contain the resolved 8579-byte _build_shared.py matching the canonical.
Two follow-ups to the merge commit, both aimed at removing remaining
duplication between the two backends.
*1. Key-based stamp helpers in _build_shared.py*
Adds `check_build_key(stamp, get_key)` and `record_build_key(stamp,
get_key)` to the shared module. Each backend now supplies a package-
specific `get_key` callable and keeps a thin wrapper:
- cuda_bindings: key is the toolchain name (re-derived from env for
`record_build_toolchain` via `_current_toolchain_key`); check gets
`lambda: toolchain` since the caller already has it.
- cuda_core: key is the composite `cu{major}-{toolchain}-{opt|debug}
[-cov]` string; both check and record wrap the value in `lambda: key`
so the shared helper's signature stays uniform.
`force_build_ext` stays a module-level bool in each `build_hooks.py`:
the shared helper just returns True/False, and each wrapper flips its
own package's `force_build_ext`. That preserves `build_hooks.force_build_ext`
attribute access from setup.py without any `_build_shared` re-export
tricks.
*2. Shared-helper test mixins*
New file `cuda_python_test_helpers/build_shared.py` mirrors Ralf's
existing `cython_cache.py` mixin pattern (class attr `build_hooks = None`,
subclass sets it). Three mixins:
- `ResolveToolchainSharedMixin` — 5 tests: `test_default_does_not_touch_env`,
`test_default_preserves_existing_cc`, `test_case_insensitive`,
`test_invalid_value_raises`, `test_llvm_overrides_external_cc` (the
latter was previously in cuda_core only; both packages exercise it now).
- `CheckToolchainAvailableSharedMixin` — 3 tests: default_is_noop,
llvm_missing_tool_lists_install_hint, llvm_present_passes.
- `AbiStampPathMixin` — 1 test: stamp path scoped to `EXT_SUFFIX`, run
against a fixed `.build-test` stem so it doesn't care whether the
package stems its stamps as `.build-toolchain` or `.build-config`.
Left in each package's `test_build_hooks.py`:
- `test_llvm_sets_env_and_flags` / `test_gnu_sets_env_and_flags` /
`test_gnu_keeps_gcc_only_flags` — flag-set assertions genuinely differ
between bindings and core.
- Stamp-bookkeeping tests (`TestBuildToolchainStamp`, `TestBuildConfigStamp`)
— bindings stamps just the toolchain name, core stamps the composite key.
Net: -~250 duplicated test lines, one added shared file. Bindings gains
one test (llvm_overrides_external_cc), so pass counts are 34 (bindings)
and 71 (core), up from 33 and 71.
All verifications green: ruff check + format clean, SPDX clean, both
sdists still resolve `_build_shared.py` to real content matching the
canonical.
…d_shared Following on from the key-gen refactor: remove the thin `_check_build_toolchain` / `record_build_toolchain` / `_check_build_config` / `record_build_config` skeletons entirely so call sites use the shared `check_build_key(stamp, get_key)` / `record_build_key(stamp, get_key)` directly. To make that safe: - `force_build_ext` moves out of the two `build_hooks.py` files into `_build_shared.py`. `check_build_key` mutates it directly on a detected change. - Each `build_hooks.py` grows a module-level `__getattr__` that re-exports `force_build_ext` from the shared module, so setup.py's existing `build_hooks.force_build_ext` attribute read keeps working unchanged. - `cuda_bindings/setup.py`'s `build_ext.build_extensions` replaces `build_hooks.record_build_toolchain()` with `build_hooks.record_build_key(build_hooks._BUILD_TOOLCHAIN_STAMP, build_hooks._current_toolchain_key)`. - `cuda_core/build_hooks.py`'s `_build_cuda_core` inlines what `_check_build_config` used to do: compute cuda_major and the config key, then call `check_build_key(_BUILD_CONFIG_STAMP, lambda: key)`. `build_wheel` / `build_editable` call `record_build_key(...)` directly instead of `record_build_config(key)`. - Tests: the `stamp` fixtures now reset the shared force flag via `sys.modules["_build_shared"]`; `TestBuildToolchainStamp` and `TestBuildConfigStamp` call `check_build_key` / `record_build_key` directly; `TestBuildHookStamping` monkey-patches `build_hooks.record_build_key` (2-arg signature) instead of the removed named wrappers. Pass counts unchanged (34 bindings, 71 core); ruff and SPDX clean; sdists still resolve `_build_shared.py` to real content.
Per offline discussion: if a happy-path CI wheel build passes, the build-system code it exercised is by definition working, so tests asserting that same happy path are redundant. Keep only tests that assert behavior a green wheel build would not surface — error paths, cross-configuration transitions, cross-Python-ABI scoping, failure paths, and data variants beyond the one CI uses. Removed (behavior exercised by any green wheel build): - Shared mixin: `test_default_does_not_touch_env`, `test_default_preserves_existing_cc`, `test_default_is_noop`. - Per-package `test_gnu_sets_env_and_flags` (bindings and core). - `test_cuda_path_is_resolved_before_importing_bindings` — namespace ordering; wheel build fails at import if wrong. - `TestBuildToolchainStamp` / `TestBuildConfigStamp` `test_missing_stamp_forces_rebuild` — every clean-checkout CI build triggers the missing-stamp path. - `TestBuildHookStamping.test_wheel_records_exact_prepared_config_after_success` + `test_editable_records_default_config_after_patch` — every green wheel/editable exercises the record path (failure-path guards below it are kept, since CI does not fail on purpose). - `TestGeneratedSourceDirIsKeyed.test_dir_is_anchored_not_relative_to_cwd`, `TestSetuptoolsSourcePaths.test_absolute_sources_are_made_relative`, `TestExtensionDepends.test_headers_under_module_directories_only`, `TestForceReachesBuildExt.test_flag_set_forces_rebuild` — all implicit in every green wheel build. Kept (not surfaced by any single successful CI wheel): - Toolchain error paths (`test_case_insensitive`, `test_invalid_value_raises`, `test_llvm_missing_tool_lists_install_hint`). - `test_llvm_present_passes` and `test_llvm_overrides_external_cc` (no llvm CI). - All `test_changed_*_forces_rebuild` / `test_same_*_does_not_force` transitions. - `test_stamp_path_is_scoped_to_extension_abi` (cross-Python-ABI). - `test_failed_wheel_build_does_not_record_config`, `test_failed_editable_patch_does_not_record_config` (failure paths). - Parametrized `TestGetCudaMajorVersion` data variants and error paths. - `test_majors_use_different_dirs` (cross-CUDA-major). - `TestExtensionSources` edge cases, `TestParallelSourceCompilation` hook edge cases. - `test_gnu_keeps_gcc_only_flags` (bindings): regression guard for the `_is_clang` removal in NVIDIA#2903, catches a specific stale-flag scenario wheel builds would not. Pass counts: bindings 34 → 29, core 71 → 59; 17 fewer test bodies total, 213 lines removed.
Move the `from _build_shared import ...` block up next to the other imports and register `_build_shared` as first-party in ruff.toml's isort config so it sorts after third-party (setuptools, Cython) as "ours". The `# noqa: E402` opt-out is no longer needed since the import is no longer preceded by non-import statements.
Closes the audit in NVIDIA#1882. cuda_bindings' and cuda_core's ``_resolve_toolchain`` framework was identical (call _resolve_toolchain_name, msvc-debug error, _apply_toolchain_env, tuple return); only the ``extra_compile_args``/``extra_link_args`` choices differed, and NVIDIA#1882 traced that difference to legacy drift, not intent. Both are now the same function in _build_shared. Consolidations: - New ``resolve_toolchain(debug, compile_for_coverage)`` in ``_build_shared.py`` handles the whole workflow (name resolution, msvc-debug guard, flag assembly, env application, return tuple). - New ``_build_flags(name, debug, compile_for_coverage)`` in ``_build_shared.py`` carries the unified flag set. Both backends drop their local copies. - Both ``build_hooks.py`` import ``resolve_toolchain`` and call it with just the two switches; no per-package callable is passed. Flag choices are now the single set: - Linux ``-std=c++17`` (bindings used ``c++14``) - MSVC ``/std:c++17`` (bindings had no ``/std:`` at all) - Non-debug ``-g0 -O2`` (bindings used ``-O3``) - Debug ``-g -O0 -D _GLIBCXX_ASSERTIONS`` - llvm ``-fuse-ld=lld`` - Coverage ``-DCYTHON_TRACE_NOGIL=1 -DCYTHON_USE_SYS_MONITORING=0`` - Non-debug ``-Wl,--strip-all`` Bindings loses ``-Wno-deprecated-declarations``, ``-fpermissive``, and ``-fno-var-tracking-assignments`` -- per NVIDIA#1882, these were legacy leftover with no active justification. CI will surface any compilation regressions if the Cython-generated C++ actually needs ``-fpermissive``. Tests: the shared ``ResolveToolchainSharedMixin`` gains ``test_llvm_sets_env_and_flags`` (both backends now assert the same flags). Per-package ``TestResolveToolchain`` classes are gone; ``test_gnu_keeps_gcc_only_flags`` (bindings) is deleted because the flags it guarded are gone. Pass counts: 28 bindings (was 29), 59 core (unchanged).
|
/ok to test 9828a2f |
|
@juenglin could you check if this PR works on your system? |
…lf-referencing anchor Windows GH runners don't inherit ``core.symlinks=true`` from the OS, so ``actions/checkout`` materializes ``cuda_core/_build_shared.py`` as a text stub containing ``../cuda_bindings/_build_shared.py``. Any subsequent Python import of the file syntax-errors, which took out every ``Test win-64`` job on this PR. Set ``git config --global core.symlinks true`` before the checkout in each workflow that runs on Windows: - ``.github/workflows/test-wheel-windows.yml`` (Windows-only job, no if). - ``.github/workflows/test-sdist-windows.yml`` (Windows-only, no if). - ``.github/workflows/build-wheel.yml`` (cross-platform; gated with ``if: startsWith(inputs.host-platform, 'win')`` so win-64 and win-arm64 both pick it up, per the existing prefix-matching convention in that file). Docs' lychee "Check rendered docs links" fails because ``install.rst`` cross-links ``https://github.com/NVIDIA/cuda-python/blob/main/CONTRIBUTING.md#development-on-windows`` and the anchor doesn't exist on ``main`` yet (only on this branch). The error is self-resolving after merge, but until then, exclude the exact URL in ``.github/workflows/build-docs.yml``'s lychee args. Safe to remove the exclude after this PR lands. -- Leo's bot
Replace the accretion of ``--exclude`` CLI flags in ``.github/workflows/build-docs.yml`` with a repo-root ``lychee.toml``. The action still points at the file explicitly via ``--config`` so no implicit discovery order matters. Each exclude gets a comment; the ``#development-on-windows`` anchor exclude added earlier in this PR is marked as removable once main carries that section.
The four comment lines above the ``args:`` block in ``.github/workflows/build-docs.yml`` (PR-preview canonical URLs, cuda-bindings docutils #id anchors + a TODO, Preferred Networks crawler rejection) explained the excludes they sat above. Now that the excludes live in ``lychee.toml``, move the comments in beside each rule.
|
/ok to test 727f231 |
|
Andy-Jost
left a comment
There was a problem hiding this comment.
Requesting changes for the two inline comments on cuda_core/tests/test_build_hooks.py: the stamp tests are order-dependent, and the force=True path through __getattr__ has no test. The comments on the shared mixin and CONTRIBUTING are small follow-ups.
One note with no single line to anchor: dropping -Wno-deprecated-declarations adds 38 warnings to the bindings build. In the linux-64 py3.12 build log: 18 each in _internal/runtime.cpp and _internal/runtime_ptds.cpp, one each in runtime.cpp and _internal/nvrtc.cpp, all cudaMemcpy*Array* and cudaGetDriverEntryPoint. Harmless today. #2966 adds -Werror in CI, and if that reaches the shared _build_flags, bindings stops building. Keep the flag for bindings, or account for it in #2966.
resolve_toolchain gains a required keyword-only ``cxx_std`` argument and an optional ``tweak`` post-hook. The shared ``_build_flags`` no longer picks a default C++ standard — the two backends legitimately differ (bindings stays on c++14 to avoid a c++17 variadic-template regression on kernel-launch paths; core is on c++17), and there is no defensible shared default. See NVIDIA#1882 for the audit. - ``cuda_bindings/build_hooks.py`` adds ``_tweak_flags`` which layers ``-Wno-deprecated-declarations`` on Linux for the 38 warnings ``cudaMemcpy*Array*`` / ``cudaGetDriverEntryPoint`` produce in 13.4 headers. Without it, once NVIDIA#2966 turns on ``-Werror``, bindings would stop building. - Call site passes ``cxx_std=14, tweak=_tweak_flags``. - ``cuda_core/build_hooks.py`` passes ``cxx_std=17`` (no tweak). Also addresses Andy-Jost's line-anchored review comments: - ``TestForceReachesBuildExt._finalized_build_ext`` now patches ``sys.modules["_build_shared"]`` instead of ``build_hooks``. Because ``build_hooks.force_build_ext`` is a module-level ``__getattr__`` fallback, monkey-patching it on ``build_hooks`` creates a real attribute that shadows the fallback and persists across teardown — a poison state that made ``TestBuildConfigStamp`` order-dependent (reproduces under ``pytest -p randomly --randomly-seed=1``). - Restores ``test_flag_set_forces_rebuild`` — the only direct check that ``force_build_ext=True`` reaches ``build_ext.force`` through the ``__getattr__`` re-export. Without it, a later ``from _build_shared import force_build_ext`` would silently snapshot ``False`` and forced rebuilds would stop with no failing test. - Restores ``test_default_preserves_existing_cc`` in ``ResolveToolchainSharedMixin``. A green wheel build wouldn't catch a regression here — if the default toolchain overwrote ``CC``, sccache would silently stop and ``sccache-summary`` only warns. - ``CONTRIBUTING.md`` recovery recipe now runs ``git config core.symlinks true`` (no ``--global``) inside the existing clone. ``git clone`` probes symlink support and writes ``core.symlinks=false`` into the *repo-local* config on failure; repo-local overrides ``--global``, so the previous recipe couldn't fix an existing clone. Test bookkeeping: - Per-package ``TestResolveToolchain.test_llvm_sets_env_and_flags`` restored in both packages — bindings asserts ``-std=c++14`` + ``-Wno-deprecated-declarations``, core asserts ``-std=c++17`` and the absence of ``-Wno-deprecated-declarations``. Removed from the shared mixin because it now asserts package-specific flag content. - Other shared-mixin resolve tests just pick ``cxx_std=17`` as a placeholder — they test env/name/error mechanics, not flag content. Pass counts: bindings 29 (was 28), core 61 (was 59). Randomized-order runs (``-p randomly --randomly-seed={1,2,4,42}``) all green.
…ws guide
Per review:
- Shared ``_build_flags`` / ``resolve_toolchain`` docstrings no longer
explain the c++14/c++17 choice — those live in the call sites in
``cuda_bindings/build_hooks.py`` (why c++14: the launch_{256,512}_args
variadic-template regression) and ``cuda_core/build_hooks.py`` (why
c++17: c++17 features under ``cuda/core/_cpp/``). Reading the shared
file now leads you to the call site for context, instead of asking
the shared file to know about both consumers.
- ``CONTRIBUTING.md`` ends the Windows recovery recipe with a note that
re-cloning after the Developer Mode + ``core.symlinks=true`` setup is
usually simpler than patching an existing broken clone.
…-floor context
Per review: the 3% geomean and the 15% launch_{256,512}_args outlier are
both real but mean different things; the comment now names both so a
future reader isn't left wondering which one drove the choice.
Modern setuptools' MSVCCompiler no longer hard-codes /Ox on release builds, so previously the effective Windows opt level was whatever default cl.exe / distutils inferred. Emit /O2 from the shared _build_flags for symmetry with Linux -O2 and to keep the effective Windows flag set legible in one place. Adds a shared-mixin regression guard (test_msvc_sets_std_and_opt, skipped off-Windows) so a future refactor that dropped /O2 wouldn't pass every Windows wheel build silently.
|
/ok to test 4949f56 |
The auto-merge conflict was in cuda_core/build_hooks.py: HEAD had deleted the whole in-file toolchain block (moved to _build_shared.py), while main had added a -Werror / /WX branch inside that same block via NVIDIA#2966. Resolved by taking HEAD's structural side and hoisting the toggle into the shared code: - _build_shared._build_flags and resolve_toolchain grow a warnings_as_errors=False keyword; the flag emission sits next to -O2/debug for a single legible flag path. - cuda_core/build_hooks.py keeps its WARNINGS_AS_ERRORS constant (reads CUDA_PYTHON_WERROR) and threads it through the resolve_toolchain call site. No per-project _tweak_flags needed for this. - Shared-mixin tests (test_warnings_as_errors_off_by_default, test_warnings_as_errors_on) guard both the default-off case and the exact tokens emitted per platform, so a future refactor can't silently disable CI's -Werror. cuda_bindings does not currently pass warnings_as_errors=True; its existing -Wno-deprecated-declarations tweak is order-independent from -Werror and composes cleanly when it opts in later.
|
Closing this in favor of three smaller PRs that together re-implement it. I volunteered to take it over, and splitting it keeps each piece easier to review:
Andy's review feedback on this PR is carried over in #3000: the stamp tests patch Thanks @leofang for the original work; the commits here were the starting point for the three PRs above. |
Summary
Follow-up to #2903 addressing the second condition of its approval (review comment):
Also closes the audit in #1882 — the two backends' flag sets had drifted for legacy reasons; this PR unifies them, with a per-project injection point for the small handful of genuinely-different bits.
The two
build_hooks.pyfiles now share a single canonical file (cuda_bindings/_build_shared.py) via a symlink atcuda_core/_build_shared.py, and each backend's ownbuild_hooks.pyis a thin shell that declares only what is genuinely per-package. All the mechanics — CUDA path resolution, toolchain resolution, flag set, stamp read/compare/write, Cython cache — live in one place.What's in
_build_shared.py_import_get_cuda_path_or_home,_get_cuda_path(the PEP 517 pathfinder-shadowing workaround from Known issue:cuda.pathfindercannot be imported in PEP 517 in-tree build backends without a workaround #1824)._TOOLCHAINS_LINUX/_WINDOWS,_TOOLCHAIN_COMPILERS),_resolve_toolchain_name,_apply_toolchain_env,_check_toolchain_available(the CUDA_PYTHON_TOOLCHAIN block from build: add CUDA_PYTHON_TOOLCHAIN override for compiler/linker selection #2903),resolve_toolchain(*, cxx_std, debug=False, compile_for_coverage=False, warnings_as_errors=False, tweak=None)(framework), and_build_flags(name, cxx_std, debug, coverage, warnings_as_errors)(shared flag set)._BUILD_DIR,_abi_stamp_path(usesEXT_SUFFIXso 3.10 / 3.11 / freethreaded builds don't stomp each other),check_build_key(stamp, get_key)andrecord_build_key(stamp, get_key)(mechanics),force_build_ext(module-level bool, mutated bycheck_build_key)._cython_cache_path,_stable_cython_alias.Path(__file__).parentinside the shared file resolves per-package because Python does not dereference symlinks in__file__, so_BUILD_DIRand the alias location land under each package's own directory when loaded from its symlink (verified empirically before committing).What each
build_hooks.pystill ownsGenuinely per-package:
cuda_bindings/build_hooks.py:_BUILD_TOOLCHAIN_STAMP = _abi_stamp_path(".build-toolchain"),_current_toolchain_key()(re-derives from env for setup.py's post-build stamping),_tweak_flags(adds-Wno-deprecated-declarationson Linux, see below),_rename_architecture_specific_files,_prep_extensions,_build_cuda_bindings.cuda_core/build_hooks.py:_BUILD_CONFIG_STAMP = _abi_stamp_path(".build-config"),_build_config_key(cuda_major, toolchain, debug, coverage),WARNINGS_AS_ERRORS(readsCUDA_PYTHON_WERROR; from cuda.core: eliminate build warnings; set option to build with warnings as errors #2966),_determine_cuda_major_version,_relativize_extension_sources,_extension_sources,_extension_depends,_build_cuda_core,_add_cython_include_paths_to_pth(from Cython .pth file support for pixi path dependencies #1562), plusget_requires_for_build_*that pins the cuda-bindings runtime dep.Both files have a module-level
__getattr__that re-exports_build_shared.force_build_exttransparently, sosetup.py's existingbuild_hooks.force_build_extattribute read keeps working unchanged.Per-project knobs on
resolve_toolchainThree knobs let each backend pick its own defaults without forking the shared code:
cxx_std(required kwarg) — each backend picks its own C++ standard. No shared default; the two backends have picked deliberately different values (see below), and forcing them onto a common default would either regress bindings or block core.warnings_as_errors=False— opts into-Werror(Linux) or/WXwith C4551/C4244 exemptions for Cython's utility code (MSVC). Each backend gates this on its own env var so the two backends' source cleanups stay on independent timelines.tweak(optional callable) — post-hook that gets(name, extra_compile_args, extra_link_args)and returns the layered pair. Used for the small handful of flags a single backend needs but that don't belong in the shared set.Current usage:
resolve_toolchain(cxx_std=14, ..., tweak=_tweak_flags).cxx_std=14because raising to c++17 costs a measured ~15% onlaunch.launch_{256,512}_argsfrom gcc's c++17 variadic-template expansion (the aggregate ~3% geomean is inside the pyperf--fastnoise floor, but the outlier is real; measured on origin/main vs. this PR under identical CTK 13.4 / conda-forge toolchain)._tweak_flagsadds-Wno-deprecated-declarationson Linux:cudaMemcpy*Array*andcudaGetDriverEntryPointare deprecated but still supported, and the ~38 resulting warnings would otherwise break a future-Werrorbuild (the flag is order-independent since it disables the warning class entirely, so it composes cleanly withwarnings_as_errors=Truewhen bindings eventually opts in). Bindings does not currently passwarnings_as_errors=True.resolve_toolchain(cxx_std=17, ..., warnings_as_errors=WARNINGS_AS_ERRORS). c++17 is required by helper code undercuda/core/_cpp/(structured bindings,if constexpr, etc.).WARNINGS_AS_ERRORSreadsCUDA_PYTHON_WERRORat module load; CI's wheel builds set it (see cuda.core: eliminate build warnings; set option to build with warnings as errors #2966). Notweak.Shared flag set
Per #1882, bindings and core carried different Linux and MSVC flags for legacy reasons. Both now go through
_build_shared._build_flags:-std=c++{cxx_std},-g0 -O2(opt) or-g -O0 -D _GLIBCXX_ASSERTIONS(debug)-fuse-ld=lld(llvm),-Wl,--strip-all(opt)/std:c++{cxx_std},/O2(modern setuptools' MSVCCompiler no longer forces/Ox, so we set the opt level explicitly for symmetry with Linux-O2)warnings_as_errors=True):-Werroron Linux;/WX /wd4551 /wd4244on MSVC. The exemptions cover the two warning classes Cython's utility code emits in every module (C4551 "function call missing argument list" and C4244 narrowing from@cython.overflowcheck(True)helpers).-DCYTHON_TRACE_NOGIL=1 -DCYTHON_USE_SYS_MONITORING=0Bindings loses
-fpermissive,-fno-var-tracking-assignments,-O3(and keeps-Wno-deprecated-declarationsvia itstweak). CI will surface any compilation regression if-fpermissivewas actually load-bearing for the Cython-generated C++.Windows contributors
The symlink now matters at clone time on Windows outside of WSL: without the right git config, the "symlink" lands as a text stub. Handled in two places:
test-wheel-windows.yml,test-sdist-windows.yml, andbuild-wheel.ymlgain agit config --global core.symlinks truestep before checkout on Windows runners.CONTRIBUTING.md: new top-level Development on Windows section documenting (1) Activate Developer Mode (matching CuPy's contribution guide), and (2)git config --global core.symlinks true. Includes an existing-clone recovery recipe, with a note that re-cloning is generally cleaner than trying to fix symlinks in place. The prior Pre-commit on Windows / Pre-commit lychee workaround subsections are folded in so all Windows-only setup lives in one place.cuda_bindings/docs/source/install.rstandcuda_core/docs/source/install.rstcross-link the new section from their Installing from Source notes.Not touched:
contribute.rstin either package (per @leofang, they're being refactored separately), andcuda_pathfinder/(pure Python, nobuild_hooks.py).What else changes
toolshed/check_build_hooks_sync.pyand its.pre-commit-config.yamlentry — deleted. A single source of truth needs no drift check.ruff.toml— one line: adds_build_sharedtoknown-first-partyso isort groups the import withcuda.*after third-party.include _build_shared.pyso sdists carry the resolved content (setuptools'make_release_treecopies through symlinks).lychee.toml(new, at repo root) — the four link-checker excludes previously inlined in.github/workflows/build-docs.ymlmove to a proper config file; the workflow now just points at--config ./lychee.toml. Includes the self-referencing#development-on-windowsanchor exclusion needed because the section only lands with this PR.cuda_python_test_helpers/build_shared.py(new) — shared test mixins forresolve_toolchain,_check_toolchain_available, and_abi_stamp_path. Both packages'test_build_hooks.pymix them in via subclass, so each shared assertion runs once in each package's env without duplicating the test source. Flag-content assertions (-std=c++14vsc++17, presence of-Wno-deprecated-declarations) stay in the per-package test files since they encode the per-project choice.main: absorbs#2966(CUDA_PYTHON_WERROR=1toggle for CI wheel builds). The auto-merge conflict lived incuda_core/build_hooks.py— HEAD had deleted the whole in-file toolchain block, main had added a-Werrorbranch inside it. Resolved by hoisting the toggle into the shared flag set (see Per-project knobs above) so bindings can adopt it later with a one-line flip.test_wheel_records_exact_prepared_config_after_success, etc.). Kept error paths, cross-configuration transitions, cross-Python-ABI scoping, failure paths, parametrized data variants, and llvm/edge cases wheel builds don't cover. Bindings 34 → 31, core 71 → 63.Test plan
pytest cuda_bindings/tests/test_build_hooks.py --noconftest— 31 pass / 2 skip locally.pytest cuda_core/tests/test_build_hooks.py --noconftest— 63 pass / 2 skip locally.pytest-randomlyrun of both suites — no order-dependent failures (this PR fixes one that had been latent: monkey-patchingbuild_hooks.force_build_extcreated a real attribute that shadowed the__getattr__fallback and poisoned later stamp tests; fixed by patchingsys.modules["_build_shared"]instead).ruff check+ruff format --checkclean on all touched Python files.toolshed/check_spdx.pyclean.python -m build --sdistfor both packages: the extracted tarball contains_build_shared.pybyte-identical to the canonical, and a from-scratch PEP 517 backend load (freshsys.path) resolves every shared helper's__module__to_build_shared.-O3): confirmed-O3→-O2is not the source of the ~15%launch_{256,512}_argsoutlier;-std=c++14→c++17is. That's why bindings stays oncxx_std=14./ok to test) — the flag change (drop-fpermissiveetc.) makes CI the authoritative check.Marked draft until CI confirms; ready for review otherwise.
-- Leo's bot