Skip to main content
Impala serves an OpenAI-compatible /chat/completions endpoint. The change is the base URL, the key, and the model name — see Quickstart for those three values.

The request

model and messages are the only required fields.
messages is the standard OpenAI chat format. A system message sets behavior, user carries the prompt, and assistant carries prior model turns when you’re replaying a conversation:
base_url routes your requests to Impala. Pass it exactly as provided — endpoints differ between accounts. model names the model your endpoint is provisioned for.

Timeouts and retries

Two settings matter more here than at an interactive provider. Raise your timeouts. A call takes seconds, and minutes under load. SDK defaults are tuned for interactive providers and will abort a request that would have succeeded. Reduce automatic retries. The SDK can’t tell a slow request from a failed one, so its default retries fire on timeouts for requests that are still running — adding load without shortening your wait. Lower the count and raise the timeout instead. Set it to 0 if your own harness handles retries.
For an agent, total time is per-step time multiplied by the number of steps, so turn count is usually a bigger lever than any individual call. See Run async and open source.

Keep the key out of your source

The OpenAI SDK reads these environment variables automatically:
Which lets your code drop the explicit arguments:
Don’t commit your key or share it outside your team. You can rotate it under API Keys on the platform.

Other parameters

The endpoint accepts the standard OpenAI chat completion fields, and everything else works as it does with the OpenAI SDK.

Next

Run a batch

File in, results out, at the lowest cost per token.

Anthropic SDKs and Claude Code

The same endpoint in Anthropic format.