Skip to content

[BUG]: NVVM backend silently ignores 12 ProgramOptions fields that NVRTC emits #2954

Description

@Arthur031221

Is this a duplicate?

Type of Bug

Silent Failure

Component

cuda.core

Describe the bug

_prepare_nvvm_options_impl in cuda_core/cuda/core/_program.pyx raises CUDAError for 31 ProgramOptions fields that libNVVM can't use. Twelve other fields, which the NVRTC path does emit, are neither emitted nor rejected on the NVVM path. Setting one of them has no effect, and there's no error or warning:

no_cache, fdevice_time_trace, device_float128, frandom_seed, ofast_compile, pch, create_pch, use_pch, pch_dir, pch_verbose, pch_messages, instantiate_templates_in_pch

All of them except no_cache are documented as "(NVRTC only)" in the ProgramOptions docstring. device_int128 is rejected on NVVM, but device_float128 compiles without a word.

The rejection list walks the fields in almost the same order as _prepare_nvrtc_options_impl and stops at minimal, which is the field just before no_cache in the NVRTC builder. My guess is the fields after that point, and device_float128, were never added. I haven't found anything saying it was deliberate.

Two other places assume NVVM rejects some of these:

  • cuda_core/cuda/core/utils/_program_cache/_keys.py says NVVM "explicitly rejects all three at compile time" about create_pch, time and fdevice_time_trace, and gates the side-effect check on NVRTC for that reason. It also says "NVVM rejects them" about the external-content options (include_path, pre_include, pch, use_pch, pch_dir). Those two comments cover eight options. NVVM rejects three of them (time, include_path, pre_include) and silently drops the other five. With create_pch or fdevice_time_trace set, make_program_cache_key(code_type="nvvm", ...) returns a key, while code_type="c++" raises ValueError. I don't think that produces a wrong cache hit, since the compile drops the option as well.
  • test_nvvm_options_reject_each_unsupported_flag in cuda_core/tests/test_program.py says its table "mirrors _prepare_nvvm_options_impl's rejection list one-for-one", so a field missing from both can't fail it.

How to Reproduce

import dataclasses
import os
import tempfile

from cuda.bindings import nvvm
from cuda.core import CUDAError, Device, Program, ProgramOptions

# 1. Each field set on its own, then as_bytes("nvvm") compared with arch alone.
values = {
    "no_cache": True, "fdevice_time_trace": "trace.json",
    "device_float128": True, "frandom_seed": "1", "ofast_compile": "max", "pch": True,
    "create_pch": "out.pch", "use_pch": "in.pch", "pch_dir": "pch-cache", "pch_verbose": True,
    "pch_messages": True, "instantiate_templates_in_pch": True,
}
base = ProgramOptions(arch="sm_80").as_bytes("nvvm")
for f in dataclasses.fields(ProgramOptions):
    if f.name not in values:
        continue
    opts = ProgramOptions(arch="sm_80", **{f.name: values[f.name]})
    nvrtc = set(opts.as_bytes("nvrtc")) - set(ProgramOptions(arch="sm_80").as_bytes("nvrtc"))
    print(f"{f.name:30} nvvm unchanged={opts.as_bytes('nvvm') == base}  nvrtc adds {sorted(nvrtc)}")

# 2. End to end through Program.
Device().set_current()
major, minor, dmajor, dminor = nvvm.ir_version()
ir = f"""target triple = "nvptx64-unknown-cuda"
target datalayout = "e-p:64:64:64-i1:8:8-i8:8:8-i16:16:16-i32:32:32-i64:64:64-i128:128:128-f32:32:32-f64:64:64-v16:16:16-v32:32:32-v64:64:64-v128:128:128-n16:32:64"
define void @k(i32* %p) {{
  store i32 1, i32* %p, align 4
  ret void
}}
!nvvm.annotations = !{{!0}}
!0 = !{{void (i32*)* @k, !"kernel", i32 1}}
!nvvmir.version = !{{!1}}
!1 = !{{i32 {major}, i32 {minor}, i32 {dmajor}, i32 {dminor}}}
"""
for kw in ({"device_int128": True}, {"device_float128": True}):
    try:
        Program(ir, "nvvm", ProgramOptions(arch="sm_80", **kw)).compile("ptx")
        print(kw, "compiled")
    except CUDAError as e:
        print(kw, "CUDAError:", e)

with tempfile.TemporaryDirectory() as d:
    pch = os.path.join(d, "out.pch")
    Program(ir, "nvvm", ProgramOptions(arch="sm_80", create_pch=pch)).compile("ptx")
    print("nvvm create_pch compiled, pch written:", os.path.exists(pch))

Output with cuda.core built from main at f9ed2bd:

no_cache                       nvvm unchanged=True  nvrtc adds [b'--no-cache']
fdevice_time_trace             nvvm unchanged=True  nvrtc adds [b'--fdevice-time-trace=trace.json']
device_float128                nvvm unchanged=True  nvrtc adds [b'--device-float128']
frandom_seed                   nvvm unchanged=True  nvrtc adds [b'--frandom-seed=1']
ofast_compile                  nvvm unchanged=True  nvrtc adds [b'--Ofast-compile=max']
pch                            nvvm unchanged=True  nvrtc adds [b'--pch']
create_pch                     nvvm unchanged=True  nvrtc adds [b'--create-pch=out.pch']
use_pch                        nvvm unchanged=True  nvrtc adds [b'--use-pch=in.pch']
pch_dir                        nvvm unchanged=True  nvrtc adds [b'--pch-dir=pch-cache']
pch_verbose                    nvvm unchanged=True  nvrtc adds [b'--pch-verbose=true']
pch_messages                   nvvm unchanged=True  nvrtc adds [b'--pch-messages=true']
instantiate_templates_in_pch   nvvm unchanged=True  nvrtc adds [b'--instantiate-templates-in-pch=true']
{'device_int128': True} CUDAError: The following options are not supported by NVVM backend: device_int128
{'device_float128': True} compiled
nvvm create_pch compiled, pch written: False

Expected behavior

Each of these should either reach libNVVM or tell the user it was dropped. Which of the two is your call. The existing list raises CUDAError. On #2573, though, the review preferred a UserWarning over a new error in 1.x, following #2658, and the same reasoning looks like it applies here.

ofast_compile may be one to pass through rather than reject, though I'm not sure how useful it is there. Calling libNVVM directly on the IR above (nvidia-nvvm 13.4.92, nvvm.version() returns (2, 0)), -Ofast-compile=0 compiles. =min, =mid and =max each return ERROR_COMPILATION (9) with the log parse Can't read textual IR with a Context that discards named Values. So only 0 worked on this textual IR, but unlike -pch and -device-float128, which return ERROR_INVALID_OPTION (7), it isn't rejected as an unknown option.

link_time_optimization is also neither emitted nor rejected on NVVM. I left it off the list because test_nvvm_program_options passes it to NVVM on purpose, and target_type="ltoir" adds -gen-lto anyway.

Operating System

Ubuntu 24.04.4 LTS

nvidia-smi output

NVIDIA GeForce RTX 5090, driver 610.43.02. cuda.core built from source at f9ed2bd against cuda-bindings 13.4.3 and the cuda-toolkit 13.4.2 wheels, Python 3.12.13.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    P2Low priority - Nice to havebugSomething isn't workingcuda.coreEverything related to the cuda.core module

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions