Start here
Start running today
Sign up on the Impala platform, get a serverless endpoint and an API key, and send your first request.
Quickstart
Endpoint, key, one call. The fastest way to start running.
Run a batch
Upload a JSONL file, create a batch, collect the results.
Run async and open source
Why our metric is tasks per dollar, and what high latency buys you.
Deployments
Serverless in our cloud, or bring your own cloud.
How it works
Impala is built to adapt to your workloads in real time through async adaptive scheduling. It treats inference as a high-performance computing problem. It runs in Impala’s cloud, or inside your own VPC. Impala is vertically integrated across the entire stack — optimizing end-to-end from kernels to orchestration.Use cases
Agents and multi-step pipelines, where each step is its own request:- Background coding agents — repo-wide refactors, dependency migrations, test generation, automated PR review
- Research agents that plan, search and synthesize across dozens of turns
- Agentic pipelines that plan, execute, evaluate and retry until the task is done
- MCP tool orchestration across many servers and long-running sessions
- Subagent fan-out, where one parent spawns dozens of children over a shared context
- Evaluation sweeps and rollouts — the same harness over a fixed set, run to completion
- Corpus-scale extraction and summarization across document stores
- Classification and enrichment over millions of records
- Training-data curation and labeling, including synthetic generation
- Scheduled reporting — compliance and risk analysis, AI-enriched nightly ETL

