(01)
Customer story
Tessara serves 38 models from one routing table.
Latency targets used to live in a spreadsheet. Now they are rules the system keeps.

38
models live
212 ms
p95 held
3
regions
(02)
The story
Before
Each model had its own deployment scripts and its own on-call rota.
After
Every model has a route, a budget and a target. Overhead shifts traffic between regions to hold them.
In their words
Latency targets used to live in a spreadsheet. Now they are rules the system keeps.
Ines Moreau
Serving Engineer
Tessara Labs
(03)
Customers
Teams that stopped babysitting GPUs.
Next step
Get results like these.
We connect one cluster with you, read-only, for 30 days.

