Skip to main content
Open-weight models are at the frontier, so spending less no longer means running something weaker. What costs money is the volume of tokens you push through them — and most of that volume runs with nobody watching: coding agents grinding through a repo overnight, research agents that take four minutes to come back with an answer, nightly jobs classifying a few million records. The deadline applies to the finished task, so paying interactive prices for a fast first token buys you nothing. That is the traffic Impala is built for: async inference, meaning model calls where no human reads the tokens as they arrive. Latency you aren’t using is the one thing you can trade for a much lower price on the same model — cheap enough that you stop rationing turns and context.

The metric is tasks per dollar

For async work the question is: for a fixed budget, how many tasks finish? That’s tasks per dollar — the cost of one completed task. Two numbers feed it: throughput, how many tokens the fleet produces, and task completion time, how long one task takes end to end. Throughput and task completion time aren’t a trade against each other: faster and cheaper, not faster or cheaper. A single request may sit longer here than at an interactive provider, while the task it belongs to finishes sooner and costs less — because the scheduler works the whole trace rather than each call in isolation. A GPU costs the same per hour saturated or idle, so cost per token is a throughput problem: double throughput on the same hardware and cost falls by roughly half. The argument with the numbers is in Run It Hot; the engine that delivers it is Impala Herd.
The Usage view on the Impala platform, showing throughput in requests per minute and latency at p50, p90 and p95 measured in seconds