(01)
About
We built the screen we wanted at 3 a.m.

SITE 02 · BASIN B
18,560 GPUS · 94% BUSY
SRTM terrain · load plate
(02)
Why
Too many screens, too little control.
Every team we worked on ran AI across a mix of clouds, rented racks and their own hardware. Each part had its own dashboard, and none of them agreed on what was running.
Overhead is the one screen above all of them: one inventory, one queue, one set of rules and one record of what happened.
Under management, right now
GPUs across 6 regions, 2 clouds and 31 colocation cages, on one inventory.
run-7731 train-70b pool-b 62%
run-7735 eval-suite pool-a 18%
run-7740 embed-docs pool-c 91%
job-2210 index-shard-04 edge
job-2211 index-shard-05 edge
run-7742 HELD policy-41
run-7744 finetune-vision pool-b
job-2214 nightly-dedupe pool-c
run-7746 distill-8b pool-a
job-2218 cache-warm us-east-2
run-7749 rlhf-pass-3 pool-b
job-2221 export-audit csv
(03)
Regions
Six regions. One inventory.

+
EU-NORTH-1
Nominal
Northern Sweden
Power
48 MW
GPUs
6,144
Cooling
Liquid

+
EU-WEST-2
Nominal
Western Ireland
Power
32 MW
GPUs
4,096
Cooling
Air + rear door

+
US-EAST-2
Nominal
Ohio valley
Power
60 MW
GPUs
8,192
Cooling
Liquid

+
US-WEST-1
Maintenance window
High desert, Oregon
Power
40 MW
GPUs
5,120
Cooling
Evaporative

+
AP-SOUTH-1
Nominal
Western India
Power
24 MW
GPUs
3,072
Cooling
Liquid

+
ME-CENTRAL-1
Nominal
Arabian plateau
Power
36 MW
GPUs
4,608
Cooling
Liquid
(04)
How we work
Four rules we keep.
[01]
Show the real number
If a value is estimated, the screen says so.
[02]
Never kill a job silently
Hold it, say why, and name who can let it through.
[03]
Your data stays yours
Metadata only, and self-hosting when you need it.
[04]
Write it down
Every incident gets a public write-up on the changelog.
Next step
Want to work on this?
We hire infrastructure and product engineers who have carried a pager.