/chat/completions endpoint. The change is the base URL, the key, and the model name — see Quickstart for those three values.
The request
model and messages are the only required fields.
messages is the standard OpenAI chat format. A system message sets behavior, user carries the prompt, and assistant carries prior model turns when you’re replaying a conversation:
base_url routes your requests to Impala. Pass it exactly as provided — endpoints differ between accounts. model names the model your endpoint is provisioned for.
Timeouts and retries
Two settings matter more here than at an interactive provider. Raise your timeouts. A call takes seconds, and minutes under load. SDK defaults are tuned for interactive providers and will abort a request that would have succeeded. Reduce automatic retries. The SDK can’t tell a slow request from a failed one, so its default retries fire on timeouts for requests that are still running — adding load without shortening your wait. Lower the count and raise the timeout instead. Set it to0 if your own harness handles retries.
Keep the key out of your source
The OpenAI SDK reads these environment variables automatically:Other parameters
The endpoint accepts the standard OpenAI chat completion fields, and everything else works as it does with the OpenAI SDK.Next
Run a batch
File in, results out, at the lowest cost per token.
Anthropic SDKs and Claude Code
The same endpoint in Anthropic format.

