Overhead OS · v4.2
Overhead is the control layer for teams running AI on their own GPUs. Clusters, jobs, models and spend on one screen, with policies that hold.
Running on Overhead
Wire · 2,140 teams · 31 countries
+
Corvane Bio
2.1M jobs / month
+
Northwake
640 GPUs
+
Tessara Labs
38 models live
+
Oxbow Robotics
11 regions
+
Lumenfold
4.2 PB indexed
+
Parallax
0 policy misses
(01)
The problem
Learn more
Platform
→
2 days
median setup
41
policy templates
Self-hosted
or our cloud



(02)
How it’s built
Three layers. One picture of everything you run.
1.0
Compute
2.0
Orchestration
3.0
Control
(03)
How it works
Four verbs. Every job goes through all of them.

[04] / [04]
Every job
[01]
Observe
Reads every node, job and bill from your clusters and clouds into one inventory, every few seconds.
[02]
Schedule
Places each job where it fits best, and moves it when a node fails or a cheaper slot opens.
[03]
Enforce
Checks your policies before a job starts and while it runs. Breaking a rule holds the job, it does not kill it.
[04]
Audit
Writes down who ran what, where, for how long and at what cost. Export it for any auditor in one click.
(04)
The Terminal
Six panes. Everything that’s running, right now.
[F1]
Queue
Live
[F2]
GPU pools
Live
[F3]
Latency p95
Live
p95 212 ms
target 250 ms
[F4]
Events
Live
14:52
Node 41 down · run 7731 moved
14:47
Policy 41 held job 7742
14:31
Pool C scaled +64 GPUs
14:12
Spend cap 71% · team vision
13:58
Checkpoint saved · run 7731
13:40
Region us-west-1 maintenance
[F5]
Spend today
Live
[F6]
Regions
Live

EU-N
US-E
AP-S
[Q]
Queue
[P]
Pools
[R]
Regions
[/]
Search
[?]
Shortcuts

(05)
Fleet utilisation
Load, drawn as terrain.
Peaks are clusters near their limit. Valleys are capacity you are paying for and not using. The plate redraws every minute.
00:00
06:00
12:00
18:00
23:59 UTC
Now · 14:52 UTC
Peak right now
US-EAST-2 · 97.4%
Pool B · B200 · 1,536 GPUs
Clusters, jobs, models, spend and rules, on the same screen with the same clock.
(06)
Solutions
What teams use it for.

Long runs
Training
Multi-week training runs that keep their place when a node fails, with checkpoints you can find again.
0
runs lost to a node failure last quarter

Serving
Inference
Serve models across regions with one routing table, and see p95 latency per model, per region, by the minute.
212 ms
median p95 across customer fleets

Inventory
Fleet and regions
Every GPU, node and site in one inventory, including the ones you rent and the ones in your own racks.
6.1%
average idle capacity found in week one

Policy
Governance
Rules for who can run what, where and at what cost, written in plain files and enforced on every job.
41
ready-made policy templates

Spend
Cost control
Daily spend caps per team and per project, with forecasts that update as jobs start and finish.
−23%
GPU spend in the first 90 days (median)

Location
Data residency
Keep data and the jobs that touch it inside the regions you choose, with proof for your auditors.
100%
jobs placed in an allowed region
(07)
Regions
Six regions. One inventory.

+
EU-NORTH-1
Nominal
Northern Sweden
Power
48 MW
GPUs
6,144
Cooling
Liquid

+
EU-WEST-2
Nominal
Western Ireland
Power
32 MW
GPUs
4,096
Cooling
Air + rear door

+
US-EAST-2
Nominal
Ohio valley
Power
60 MW
GPUs
8,192
Cooling
Liquid

+
US-WEST-1
Maintenance window
High desert, Oregon
Power
40 MW
GPUs
5,120
Cooling
Evaporative

+
AP-SOUTH-1
Nominal
Western India
Power
24 MW
GPUs
3,072
Cooling
Liquid

+
ME-CENTRAL-1
Nominal
Arabian plateau
Power
36 MW
GPUs
4,608
Cooling
Liquid
Under management, right now
GPUs across 6 regions, 2 clouds and 31 colocation cages, on one inventory.
run-7731 train-70b pool-b 62%
run-7735 eval-suite pool-a 18%
run-7740 embed-docs pool-c 91%
job-2210 index-shard-04 edge
job-2211 index-shard-05 edge
run-7742 HELD policy-41
run-7744 finetune-vision pool-b
job-2214 nightly-dedupe pool-c
run-7746 distill-8b pool-a
job-2218 cache-warm us-east-2
run-7749 rlhf-pass-3 pool-b
job-2221 export-audit csv
(08)
Security and control
Rules you can read. A record you can prove.
policies/eu-residency.yaml
policy: eu-residency
applies_to: datasets.tag == “eu”
allow_regions:
- eu-north-1
- eu-west-2
on_violation: hold
notify: owner, #platform
approver: security-oncall
Audit log · last 6
14:52:08
moved run-7731 → pool-b
14:47:31
held run-7742 (policy 41)
14:31:02
scaled pool-c +64
14:12:44
cap warning team-vision
13:58:19
checkpoint run-7731
13:40:00
maintenance us-west-1
Self-hosted option
Run the whole control plane inside your own network.
Metadata only
Weights and datasets never leave your clusters.
One-click rollback
Every policy change can be undone, and the undo is logged.
(09)
Customers
Teams that stopped babysitting GPUs.
(10)
Pricing
Priced by GPUs under management.
Team
$0
up to 64 GPUs
For one team getting its first cluster under control.
+
Inventory and queue for one cluster
+
10 policy templates
+
7-day audit log
+
Community support
+
Self-serve setup in an afternoon
Start free
→
Fleet
$2,400
per month, up to 2,048 GPUs
For companies running AI across clouds and their own racks.
+
Every cluster, cloud and site
+
All 41 policy templates
+
Spend caps and forecasts
+
1-year audit log
+
Named support engineer
Request access
→
Sovereign
Custom
self-hosted, any size
For teams that must run everything inside their own walls.
+
Self-hosted control plane
+
Residency proofs for auditors
+
Unlimited audit retention
+
24/7 support with a 15-minute response
+
Installed by our engineers on site
Talk to us
→
(11)
Changelog
What shipped, what broke, what we fixed.
All changes
→
Release
v4.2
Spend caps per project
Caps can now be set per project as well as per team. Jobs that would cross a cap are held, and the owner is asked.
Fix
v4.1.3
Faster resume after node loss
Training runs now resume from the last checkpoint in under 40 seconds on most pools, down from about three minutes.
Policy
Policy 41
EU residency policy pack
A ready-made set of rules that keeps datasets and the jobs that read them inside EU regions.
Release
v4.1
Region view in the Terminal
A new pane shows each region's power, GPUs and queue depth side by side. Press R to open it.
Does Overhead replace our scheduler?
+
Where does our data go?
+
Can we run it ourselves?
+
How long does setup take?
+
What happens when a policy blocks a job?
+
Which hardware do you support?
+
How is pricing counted?
+
Can we try it on real workloads?
+
Request access
Put your whole stack on one screen.
We connect one cluster with you, read-only, for 30 days. No card needed.



