Is this a duplicate?
Type of Bug
Silent Failure
Component
cuda.core
Describe the bug
_prepare_nvvm_options_impl in cuda_core/cuda/core/_program.pyx raises CUDAError for 31 ProgramOptions fields that libNVVM can't use. Twelve other fields, which the NVRTC path does emit, are neither emitted nor rejected on the NVVM path. Setting one of them has no effect, and there's no error or warning:
no_cache, fdevice_time_trace, device_float128, frandom_seed, ofast_compile, pch, create_pch, use_pch, pch_dir, pch_verbose, pch_messages, instantiate_templates_in_pch
All of them except no_cache are documented as "(NVRTC only)" in the ProgramOptions docstring. device_int128 is rejected on NVVM, but device_float128 compiles without a word.
The rejection list walks the fields in almost the same order as _prepare_nvrtc_options_impl and stops at minimal, which is the field just before no_cache in the NVRTC builder. My guess is the fields after that point, and device_float128, were never added. I haven't found anything saying it was deliberate.
Two other places assume NVVM rejects some of these:
cuda_core/cuda/core/utils/_program_cache/_keys.py says NVVM "explicitly rejects all three at compile time" about create_pch, time and fdevice_time_trace, and gates the side-effect check on NVRTC for that reason. It also says "NVVM rejects them" about the external-content options (include_path, pre_include, pch, use_pch, pch_dir). Those two comments cover eight options. NVVM rejects three of them (time, include_path, pre_include) and silently drops the other five. With create_pch or fdevice_time_trace set, make_program_cache_key(code_type="nvvm", ...) returns a key, while code_type="c++" raises ValueError. I don't think that produces a wrong cache hit, since the compile drops the option as well.
test_nvvm_options_reject_each_unsupported_flag in cuda_core/tests/test_program.py says its table "mirrors _prepare_nvvm_options_impl's rejection list one-for-one", so a field missing from both can't fail it.
How to Reproduce
import dataclasses
import os
import tempfile
from cuda.bindings import nvvm
from cuda.core import CUDAError, Device, Program, ProgramOptions
# 1. Each field set on its own, then as_bytes("nvvm") compared with arch alone.
values = {
"no_cache": True, "fdevice_time_trace": "trace.json",
"device_float128": True, "frandom_seed": "1", "ofast_compile": "max", "pch": True,
"create_pch": "out.pch", "use_pch": "in.pch", "pch_dir": "pch-cache", "pch_verbose": True,
"pch_messages": True, "instantiate_templates_in_pch": True,
}
base = ProgramOptions(arch="sm_80").as_bytes("nvvm")
for f in dataclasses.fields(ProgramOptions):
if f.name not in values:
continue
opts = ProgramOptions(arch="sm_80", **{f.name: values[f.name]})
nvrtc = set(opts.as_bytes("nvrtc")) - set(ProgramOptions(arch="sm_80").as_bytes("nvrtc"))
print(f"{f.name:30} nvvm unchanged={opts.as_bytes('nvvm') == base} nvrtc adds {sorted(nvrtc)}")
# 2. End to end through Program.
Device().set_current()
major, minor, dmajor, dminor = nvvm.ir_version()
ir = f"""target triple = "nvptx64-unknown-cuda"
target datalayout = "e-p:64:64:64-i1:8:8-i8:8:8-i16:16:16-i32:32:32-i64:64:64-i128:128:128-f32:32:32-f64:64:64-v16:16:16-v32:32:32-v64:64:64-v128:128:128-n16:32:64"
define void @k(i32* %p) {{
store i32 1, i32* %p, align 4
ret void
}}
!nvvm.annotations = !{{!0}}
!0 = !{{void (i32*)* @k, !"kernel", i32 1}}
!nvvmir.version = !{{!1}}
!1 = !{{i32 {major}, i32 {minor}, i32 {dmajor}, i32 {dminor}}}
"""
for kw in ({"device_int128": True}, {"device_float128": True}):
try:
Program(ir, "nvvm", ProgramOptions(arch="sm_80", **kw)).compile("ptx")
print(kw, "compiled")
except CUDAError as e:
print(kw, "CUDAError:", e)
with tempfile.TemporaryDirectory() as d:
pch = os.path.join(d, "out.pch")
Program(ir, "nvvm", ProgramOptions(arch="sm_80", create_pch=pch)).compile("ptx")
print("nvvm create_pch compiled, pch written:", os.path.exists(pch))
Output with cuda.core built from main at f9ed2bd:
no_cache nvvm unchanged=True nvrtc adds [b'--no-cache']
fdevice_time_trace nvvm unchanged=True nvrtc adds [b'--fdevice-time-trace=trace.json']
device_float128 nvvm unchanged=True nvrtc adds [b'--device-float128']
frandom_seed nvvm unchanged=True nvrtc adds [b'--frandom-seed=1']
ofast_compile nvvm unchanged=True nvrtc adds [b'--Ofast-compile=max']
pch nvvm unchanged=True nvrtc adds [b'--pch']
create_pch nvvm unchanged=True nvrtc adds [b'--create-pch=out.pch']
use_pch nvvm unchanged=True nvrtc adds [b'--use-pch=in.pch']
pch_dir nvvm unchanged=True nvrtc adds [b'--pch-dir=pch-cache']
pch_verbose nvvm unchanged=True nvrtc adds [b'--pch-verbose=true']
pch_messages nvvm unchanged=True nvrtc adds [b'--pch-messages=true']
instantiate_templates_in_pch nvvm unchanged=True nvrtc adds [b'--instantiate-templates-in-pch=true']
{'device_int128': True} CUDAError: The following options are not supported by NVVM backend: device_int128
{'device_float128': True} compiled
nvvm create_pch compiled, pch written: False
Expected behavior
Each of these should either reach libNVVM or tell the user it was dropped. Which of the two is your call. The existing list raises CUDAError. On #2573, though, the review preferred a UserWarning over a new error in 1.x, following #2658, and the same reasoning looks like it applies here.
ofast_compile may be one to pass through rather than reject, though I'm not sure how useful it is there. Calling libNVVM directly on the IR above (nvidia-nvvm 13.4.92, nvvm.version() returns (2, 0)), -Ofast-compile=0 compiles. =min, =mid and =max each return ERROR_COMPILATION (9) with the log parse Can't read textual IR with a Context that discards named Values. So only 0 worked on this textual IR, but unlike -pch and -device-float128, which return ERROR_INVALID_OPTION (7), it isn't rejected as an unknown option.
link_time_optimization is also neither emitted nor rejected on NVVM. I left it off the list because test_nvvm_program_options passes it to NVVM on purpose, and target_type="ltoir" adds -gen-lto anyway.
Operating System
Ubuntu 24.04.4 LTS
nvidia-smi output
NVIDIA GeForce RTX 5090, driver 610.43.02. cuda.core built from source at f9ed2bd against cuda-bindings 13.4.3 and the cuda-toolkit 13.4.2 wheels, Python 3.12.13.
Is this a duplicate?
Type of Bug
Silent Failure
Component
cuda.core
Describe the bug
_prepare_nvvm_options_implincuda_core/cuda/core/_program.pyxraisesCUDAErrorfor 31ProgramOptionsfields that libNVVM can't use. Twelve other fields, which the NVRTC path does emit, are neither emitted nor rejected on the NVVM path. Setting one of them has no effect, and there's no error or warning:no_cache,fdevice_time_trace,device_float128,frandom_seed,ofast_compile,pch,create_pch,use_pch,pch_dir,pch_verbose,pch_messages,instantiate_templates_in_pchAll of them except
no_cacheare documented as "(NVRTC only)" in theProgramOptionsdocstring.device_int128is rejected on NVVM, butdevice_float128compiles without a word.The rejection list walks the fields in almost the same order as
_prepare_nvrtc_options_impland stops atminimal, which is the field just beforeno_cachein the NVRTC builder. My guess is the fields after that point, anddevice_float128, were never added. I haven't found anything saying it was deliberate.Two other places assume NVVM rejects some of these:
cuda_core/cuda/core/utils/_program_cache/_keys.pysays NVVM "explicitly rejects all three at compile time" aboutcreate_pch,timeandfdevice_time_trace, and gates the side-effect check on NVRTC for that reason. It also says "NVVM rejects them" about the external-content options (include_path,pre_include,pch,use_pch,pch_dir). Those two comments cover eight options. NVVM rejects three of them (time,include_path,pre_include) and silently drops the other five. Withcreate_pchorfdevice_time_traceset,make_program_cache_key(code_type="nvvm", ...)returns a key, whilecode_type="c++"raisesValueError. I don't think that produces a wrong cache hit, since the compile drops the option as well.test_nvvm_options_reject_each_unsupported_flagincuda_core/tests/test_program.pysays its table "mirrors _prepare_nvvm_options_impl's rejection list one-for-one", so a field missing from both can't fail it.How to Reproduce
Output with
cuda.corebuilt frommainat f9ed2bd:Expected behavior
Each of these should either reach libNVVM or tell the user it was dropped. Which of the two is your call. The existing list raises
CUDAError. On #2573, though, the review preferred aUserWarningover a new error in 1.x, following #2658, and the same reasoning looks like it applies here.ofast_compilemay be one to pass through rather than reject, though I'm not sure how useful it is there. Calling libNVVM directly on the IR above (nvidia-nvvm13.4.92,nvvm.version()returns(2, 0)),-Ofast-compile=0compiles.=min,=midand=maxeach returnERROR_COMPILATION (9)with the logparse Can't read textual IR with a Context that discards named Values. So only0worked on this textual IR, but unlike-pchand-device-float128, which returnERROR_INVALID_OPTION (7), it isn't rejected as an unknown option.link_time_optimizationis also neither emitted nor rejected on NVVM. I left it off the list becausetest_nvvm_program_optionspasses it to NVVM on purpose, andtarget_type="ltoir"adds-gen-ltoanyway.Operating System
Ubuntu 24.04.4 LTS
nvidia-smi output
NVIDIA GeForce RTX 5090, driver 610.43.02.
cuda.corebuilt from source at f9ed2bd against cuda-bindings 13.4.3 and the cuda-toolkit 13.4.2 wheels, Python 3.12.13.