Every batch run on Impala is addressed to a job. This is the one place where Impala’s API differs from OpenAI’s, so it’s worth understanding before your first run.
What a job is
A job is a model plus a configuration, provisioned for your account and identified by a job_id that looks like job-xxxxxxxx.
The configuration is the part that matters for throughput: which model weights are loaded, how they’re quantized, what hardware shape they run on, and how the scheduler treats the workload. Impala provisions a job for you rather than having you assemble those choices per request, which is how a batch run reaches its cost per token.
Practically, this means:
- You don’t pick a model and a machine type at request time. You pick a job.
- Two jobs can serve the same model with different configurations.
- A model that isn’t bound to one of your jobs can’t be batched yet. Ask your Impala contact to provision one.
Finding your job IDs
The Run tab in the Impala console lists every job ID available to your workspace, with one-click copy. That list is the definitive answer to “what can we run” — not the model catalog in the Playground, which exists for exploration and does not imply batch availability.
Using a job ID
Pass job_id when you create the batch:
The OpenAI SDK has no job_id parameter, so pass it through extra_body:
The model has to match
Every line of your JSONL carries a model field, and it must name the same model the job is bound to.
If body.model and the job’s model disagree, the request fails. Two habits that prevent it:
- Generate the JSONL from a single constant rather than typing the model name per line.
- Keep one job ID per pipeline, so the pairing lives in one place in your config.
url in each line must also match the endpoint you passed when creating the batch.
Need help?
Reach out to your Impala contact directly.