Skip to main content
Impala serves models fine-tuned with LoRA (low-rank adaptation). You bring the adapter, we deploy it with its base model, and you send requests to it through serverless or batch.

What to send us

Either form works:
  • A separate adapter, alongside the name of the base model it was trained on.
  • A merged model, with the adapter weights already folded into the base model.
Contact your Impala account team to share it.

How it’s served

The adapter is loaded with the model when it’s deployed, so every request to that deployment uses it. Switching adapters per request isn’t supported yet.

Sending requests

Send requests exactly as you would. See Chat completions for serverless and Run a batch for batch. If you retrain full weights often, Impala Leap swaps new checkpoints onto a running deployment without a redeploy.