Skip to content

BigQuery: create() throws NullPointerException on a duplicate job id outside the US multi-region (java-bigquery#3034 not fixed by #3035) #14025

Description

@laughingman7743

Environment details

  1. API: BigQuery (java-bigquery)
  2. OS type and version: macOS 15 / Linux (not OS-specific)
  3. Java version: 17
  4. Version: google-cloud-bigquery 2.68.0, and present on current main
    (java-bigquery/google-cloud-bigquery/src/main/java/com/google/cloud/bigquery/BigQueryImpl.java
    lines 609–611)

Summary

BigQuery.create(JobInfo) with an explicit job id throws
NullPointerException: Cannot invoke "com.google.cloud.bigquery.Job.getStatistics()" because "job" is null
when the id already exists and the job lives outside the US multi-region while the JobId carries
no location — masking the real Already Exists: Job project:location.id error.

This is the failure googleapis/java-bigquery#3034 reported. That issue was closed by
googleapis/java-bigquery#3035, which does not cover this path:

The NPE therefore still reproduces, deterministically, on 2.68.0 — which contains #3035's change.

Root cause

jobs.get resolves a JobId that names no location against the US multi-region only. When the
duplicated job's location was inferred from the statement (a query or load touching a regional
dataset), the duplicate-id handler's re-fetch misses it, getJob returns null, and
job.getStatistics() throws. The idRandom branch a few lines below already handles exactly this
case with if (job == null) { throw createException; }; the fixed-id branch lacks the same guard.

Steps to reproduce

Run the code below against any project with a dataset outside the US multi-region (reproduced
against a us-central1 dataset). No location is set anywhere; the region comes from the query.

BigQuery bq = BigQueryOptions.newBuilder().setProjectId(project).build().getService();
String sql = "SELECT COUNT(*) FROM `" + project + "." + regionalDataset + ".INFORMATION_SCHEMA.TABLES`";
String id = "npe_repro_" + System.currentTimeMillis();

bq.create(JobInfo.of(JobId.newBuilder().setProject(project).setJob(id).build(),
        QueryJobConfiguration.of(sql)));                     // ok — server infers us-central1
bq.create(JobInfo.of(JobId.newBuilder().setProject(project).setJob(id).build(),
        QueryJobConfiguration.of(sql)));                     // NullPointerException

Control arm: setting the location on the second create's JobId makes the handler find the job
and return it — the intended duplicate-id behaviour — which isolates the defect to the
location-less re-fetch.

Stack trace

java.lang.NullPointerException: Cannot invoke "com.google.cloud.bigquery.Job.getStatistics()" because "job" is null
	at com.google.cloud.bigquery.BigQueryImpl.create(BigQueryImpl.java:611)
	at com.google.cloud.bigquery.BigQueryImpl.create(BigQueryImpl.java:511)

Expected behaviour and fix

When the re-fetch returns null, fall through to the original exception, as the idRandom branch
does. Verified against BigQuery with the guard patched in: the NPE becomes
BigQueryException: Already Exists: Job flink-gcp:us-central1.npe_repro_a_…, whose message names
the job's actual location. I can send the PR.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions