(01)

Solution

Inference

Serve models across regions with one routing table, and see p95 latency per model, per region, by the minute.

Inference line drawing

212 ms

median p95 across customer fleets

Serving

What it does

Each model gets a route, a budget and a latency target. Overhead shifts traffic between regions to hold the target and tells you when it can’t.

What you set

Targets, regions and the order to shed load in. Changes are reviewed like code.

Next step

Try it on your own cluster.

Read-only for 30 days, connected with one of our engineers.

Create a free website with Framer, the website builder loved by startups, designers and agencies.