Environment details
- API: BigQuery (
java-bigquery)
- OS type and version: macOS 15 / Linux (not OS-specific)
- Java version: 17
- Version: google-cloud-bigquery 2.68.0, and present on current
main
(java-bigquery/google-cloud-bigquery/src/main/java/com/google/cloud/bigquery/BigQueryImpl.java
lines 609–611)
Summary
BigQuery.create(JobInfo) with an explicit job id throws
NullPointerException: Cannot invoke "com.google.cloud.bigquery.Job.getStatistics()" because "job" is null
when the id already exists and the job lives outside the US multi-region while the JobId carries
no location — masking the real Already Exists: Job project:location.id error.
This is the failure googleapis/java-bigquery#3034 reported. That issue was closed by
googleapis/java-bigquery#3035, which does not cover this path:
The NPE therefore still reproduces, deterministically, on 2.68.0 — which contains #3035's change.
Root cause
jobs.get resolves a JobId that names no location against the US multi-region only. When the
duplicated job's location was inferred from the statement (a query or load touching a regional
dataset), the duplicate-id handler's re-fetch misses it, getJob returns null, and
job.getStatistics() throws. The idRandom branch a few lines below already handles exactly this
case with if (job == null) { throw createException; }; the fixed-id branch lacks the same guard.
Steps to reproduce
Run the code below against any project with a dataset outside the US multi-region (reproduced
against a us-central1 dataset). No location is set anywhere; the region comes from the query.
BigQuery bq = BigQueryOptions.newBuilder().setProjectId(project).build().getService();
String sql = "SELECT COUNT(*) FROM `" + project + "." + regionalDataset + ".INFORMATION_SCHEMA.TABLES`";
String id = "npe_repro_" + System.currentTimeMillis();
bq.create(JobInfo.of(JobId.newBuilder().setProject(project).setJob(id).build(),
QueryJobConfiguration.of(sql))); // ok — server infers us-central1
bq.create(JobInfo.of(JobId.newBuilder().setProject(project).setJob(id).build(),
QueryJobConfiguration.of(sql))); // NullPointerException
Control arm: setting the location on the second create's JobId makes the handler find the job
and return it — the intended duplicate-id behaviour — which isolates the defect to the
location-less re-fetch.
Stack trace
java.lang.NullPointerException: Cannot invoke "com.google.cloud.bigquery.Job.getStatistics()" because "job" is null
at com.google.cloud.bigquery.BigQueryImpl.create(BigQueryImpl.java:611)
at com.google.cloud.bigquery.BigQueryImpl.create(BigQueryImpl.java:511)
Expected behaviour and fix
When the re-fetch returns null, fall through to the original exception, as the idRandom branch
does. Verified against BigQuery with the guard patched in: the NPE becomes
BigQueryException: Already Exists: Job flink-gcp:us-central1.npe_repro_a_…, whose message names
the job's actual location. I can send the PR.
Environment details
java-bigquery)main(
java-bigquery/google-cloud-bigquery/src/main/java/com/google/cloud/bigquery/BigQueryImpl.javalines 609–611)
Summary
BigQuery.create(JobInfo)with an explicit job id throwsNullPointerException: Cannot invoke "com.google.cloud.bigquery.Job.getStatistics()" because "job" is nullwhen the id already exists and the job lives outside the US multi-region while the
JobIdcarriesno location — masking the real
Already Exists: Job project:location.iderror.This is the failure googleapis/java-bigquery#3034 reported. That issue was closed by
googleapis/java-bigquery#3035, which does not cover this path:
getJob(jobId)togetJob(jobId, JobOption.fields(JobField.STATISTICS)). Narrowing the requested fields cannotturn a null return into a job, and add vision_v1p2beta1 #3034's stack trace shows the receiver
jobitself was null.testCreateJobTryGetNotRandom) mocks the RPCgetJobreturning a job,so only the found-job path is covered; the null return has no test.
The NPE therefore still reproduces, deterministically, on 2.68.0 — which contains #3035's change.
Root cause
jobs.getresolves aJobIdthat names no location against the US multi-region only. When theduplicated job's location was inferred from the statement (a query or load touching a regional
dataset), the duplicate-id handler's re-fetch misses it,
getJobreturns null, andjob.getStatistics()throws. TheidRandombranch a few lines below already handles exactly thiscase with
if (job == null) { throw createException; }; the fixed-id branch lacks the same guard.Steps to reproduce
Run the code below against any project with a dataset outside the US multi-region (reproduced
against a
us-central1dataset). No location is set anywhere; the region comes from the query.Control arm: setting the location on the second create's
JobIdmakes the handler find the joband return it — the intended duplicate-id behaviour — which isolates the defect to the
location-less re-fetch.
Stack trace
Expected behaviour and fix
When the re-fetch returns null, fall through to the original exception, as the
idRandombranchdoes. Verified against BigQuery with the guard patched in: the NPE becomes
BigQueryException: Already Exists: Job flink-gcp:us-central1.npe_repro_a_…, whose message namesthe job's actual location. I can send the PR.