SITE 01 · NORTH ARRAY

212 MW · 99.998% UP

SITE 02 · BASIN B

18,560 GPUS · 94% BUSY

SITE 03 · EAST GRID

140 MW · COOLING OK

SITE 04 · RELAY 7

3.8 TB/S · 0 DROPS

SITE 05 · SOUTH ARRAY

POLICY 41 · ENFORCED

Sentinel-2 L2A · 10 m/px · 28 Sep 2026 · scan pass SWIR/NIR

Operating layer for AI infrastructure

Run every system

from one screen.

SITE 01 · NORTH ARRAY

212 MW · 99.998% UP

SITE 02 · BASIN B

18,560 GPUS · 94% BUSY

SITE 03 · EAST GRID

140 MW · COOLING OK

SITE 04 · RELAY 7

3.8 TB/S · 0 DROPS

SITE 05 · SOUTH ARRAY

POLICY 41 · ENFORCED

Sentinel-2 L2A · 10 m/px · 28 Sep 2026 · scan pass SWIR/NIR

Operating layer for AI infrastructure

Run every system

from one screen.

SITE 01 · NORTH ARRAY

212 MW · 99.998% UP

SITE 02 · BASIN B

18,560 GPUS · 94% BUSY

SITE 03 · EAST GRID

140 MW · COOLING OK

SITE 04 · RELAY 7

3.8 TB/S · 0 DROPS

SITE 05 · SOUTH ARRAY

POLICY 41 · ENFORCED

Sentinel-2 L2A · 10 m/px · 28 Sep 2026 · scan pass SWIR/NIR

Operating layer for AI infrastructure

Run every system

from one screen.

Overhead OS · v4.2

Overhead is the control layer for teams running AI on their own GPUs. Clusters, jobs, models and spend on one screen, with policies that hold.

System statusUTC 09:41:00
01.Ingest18,432 ev/s
02.Scheduler1,204 jobs
03.GPU pool94.6 % busy
04.Network3.82 Tb/s
05.Storage61.2 PB
06.Inference212 ms p95
07.Policy checks9,861 /min
08.Audit log100 % kept
All systems nominal · 08/08

Explore the platform

↗

Running on Overhead

Wire · 2,140 teams · 31 countries

+

Corvane Bio

2.1M jobs / month

+

Northwake

640 GPUs

+

Tessara Labs

38 models live

+

Oxbow Robotics

11 regions

+

Lumenfold

4.2 PB indexed

+

Parallax

0 policy misses

(01)

The problem

Your models, clusters, jobs and bills live in nine different tools. Overhead puts them on one screen you can act on, with rules that still hold when nobody is watching.

Your models, clusters, jobs and bills live in nine different tools. Overhead puts them on one screen you can act on, with rules that still hold when nobody is watching.

Your models, clusters, jobs and bills live in nine different tools. Overhead puts them on one screen you can act on, with rules that still hold when nobody is watching.

Learn more

Platform

→

2 days

median setup

41

policy templates

Self-hosted

or our cloud

(02)

How it’s built

Three layers. One picture of everything you run.

1.0

Compute

Every GPU, node and cluster you run, in one inventory. See idle capacity and who is using what, by the minute.

Every GPU, node and cluster you run, in one inventory. See idle capacity and who is using what, by the minute.

2.0

Orchestration

Jobs queue, schedule and move between sites on their own. A training run keeps its place when a node drops.

Jobs queue, schedule and move between sites on their own. A training run keeps its place when a node drops.

3.0

Control

Policies decide who can run what, where and at what cost. Every change is logged and can be rolled back.

Policies decide who can run what, where and at what cost. Every change is logged and can be rolled back.

NODE POOL A

2,048 H200 · 6.1% IDLE

NODE POOL B

1,536 B200 · 2.4% IDLE

QUEUE

1,204 JOBS · 38 MOVED TODAY

RUN 7731

CHECKPOINT KEPT · NODE 41 DOWN

POLICY 41

EU DATA STAYS IN EU-NORTH

SPEND CAP

$18K / DAY · 71% USED

Layer source: Sentinel-2 L2A (1.0, 2.0) · SRTM terrain (3.0)

(03)

How it works

Four verbs. Every job goes through all of them.

Overhead sits above the schedulers and clouds you already use. It does four things, in order, for every job you run.

Overhead sits above the schedulers and clouds you already use. It does four things, in order, for every job you run.

[04] / [04]

Every job

[01]

Observe

Reads every node, job and bill from your clusters and clouds into one inventory, every few seconds.

[02]

Schedule

Places each job where it fits best, and moves it when a node fails or a cheaper slot opens.

[03]

Enforce

Checks your policies before a job starts and while it runs. Breaking a rule holds the job, it does not kill it.

[04]

Audit

Writes down who ran what, where, for how long and at what cost. Export it for any auditor in one click.

(04)

The Terminal

Six panes. Everything that’s running, right now.

The Terminal is Overhead’s main screen. Every value on it is live, and every pane opens with one key.

The Terminal is Overhead’s main screen. Every value on it is live, and every pane opens with one key.

[F1]

Queue

Live

Queue
01.Run 7731 · train-70b62.4 % done
02.Run 7735 · eval-suite18 % done
03.Run 7740 · embed-docs91.2 % done
04.Batch 2210 · index1,204 jobs
05.Held · policy 413 jobs
1,204 queued · 38 moved today

[F2]

GPU pools

Live

GPU pools
01.Pool A · H20094.6 % busy
02.Pool B · B20097.6 % busy
03.Pool C · L40S61.0 % busy
04.Edge · mixed38.2 % busy
05.Idle found6.1 %
18,432 GPUs · 6 regions

[F3]

Latency p95

Live

p95 212 ms

target 250 ms

[F4]

Events

Live

14:52

Node 41 down · run 7731 moved

14:47

Policy 41 held job 7742

14:31

Pool C scaled +64 GPUs

14:12

Spend cap 71% · team vision

13:58

Checkpoint saved · run 7731

13:40

Region us-west-1 maintenance

[F5]

Spend today

Live

Spend today
01.Team research$9,240
02.Team vision$12,780
03.Team serving$6,410
04.Team data$2,130
05.Forecast$31,900
Caps on 9 teams · 0 breaches

[F6]

Regions

Live

EU-N

US-E

AP-S

[Q]

Queue

[P]

Pools

[R]

Regions

[/]

Search

[?]

Shortcuts

(05)

Fleet utilisation

Load, drawn as terrain.

Peaks are clusters near their limit. Valleys are capacity you are paying for and not using. The plate redraws every minute.

100%

75%

50%

25%

0% UTIL

00:00

06:00

12:00

18:00

23:59 UTC

Now · 14:52 UTC

Peak right now

US-EAST-2 · 97.4%

Pool B · B200 · 1,536 GPUs

Idle

Saturated

One

One

One

surface

surface

surface

for everything.

for everything.

for everything.

Clusters, jobs, models, spend and rules, on the same screen with the same clock.

(07)

Regions

Six regions. One inventory.

Every site reports power, GPUs, cooling and status to the same screen. Swap these for your own sites in the CMS.

Every site reports power, GPUs, cooling and status to the same screen. Swap these for your own sites in the CMS.

EU-NORTH-1 data plate

+

EU-NORTH-1

Nominal

Northern Sweden

Power

48 MW

GPUs

6,144

Cooling

Liquid

EU-WEST-2 data plate

+

EU-WEST-2

Nominal

Western Ireland

Power

32 MW

GPUs

4,096

Cooling

Air + rear door

US-EAST-2 data plate

+

US-EAST-2

Nominal

Ohio valley

Power

60 MW

GPUs

8,192

Cooling

Liquid

US-WEST-1 data plate

+

US-WEST-1

Maintenance window

High desert, Oregon

Power

40 MW

GPUs

5,120

Cooling

Evaporative

AP-SOUTH-1 data plate

+

AP-SOUTH-1

Nominal

Western India

Power

24 MW

GPUs

3,072

Cooling

Liquid

ME-CENTRAL-1 data plate

+

ME-CENTRAL-1

Nominal

Arabian plateau

Power

36 MW

GPUs

4,608

Cooling

Liquid

Under management, right now

18,432

18,432

18,432

GPUs across 6 regions, 2 clouds and 31 colocation cages, on one inventory.

run-7731 train-70b pool-b 62%

run-7735 eval-suite pool-a 18%

run-7740 embed-docs pool-c 91%

job-2210 index-shard-04 edge

job-2211 index-shard-05 edge

run-7742 HELD policy-41

run-7744 finetune-vision pool-b

job-2214 nightly-dedupe pool-c

run-7746 distill-8b pool-a

job-2218 cache-warm us-east-2

run-7749 rlhf-pass-3 pool-b

job-2221 export-audit csv

(08)

Security and control

Rules you can read. A record you can prove.

Policies live as plain files in your repository. Every decision Overhead makes is written to an audit log you own.

Policies live as plain files in your repository. Every decision Overhead makes is written to an audit log you own.

policies/eu-residency.yaml

policy: eu-residency

applies_to: datasets.tag == “eu”

allow_regions:

- eu-north-1

- eu-west-2

on_violation: hold

notify: owner, #platform

approver: security-oncall

Audit log · last 6

14:52:08

moved run-7731 → pool-b

14:47:31

held run-7742 (policy 41)

14:31:02

scaled pool-c +64

14:12:44

cap warning team-vision

13:58:19

checkpoint run-7731

13:40:00

maintenance us-west-1

Self-hosted option

Run the whole control plane inside your own network.

Metadata only

Weights and datasets never leave your clusters.

One-click rollback

Every policy change can be undone, and the undo is logged.

(10)

Pricing

Priced by GPUs under management.

Counted at the daily peak. No per-seat fees, no charge for the audit log.

Counted at the daily peak. No per-seat fees, no charge for the audit log.

Team

$0

up to 64 GPUs

For one team getting its first cluster under control.

+

Inventory and queue for one cluster

+

10 policy templates

+

7-day audit log

+

Community support

+

Self-serve setup in an afternoon

Start free

→

Fleet

$2,400

per month, up to 2,048 GPUs

For companies running AI across clouds and their own racks.

+

Every cluster, cloud and site

+

All 41 policy templates

+

Spend caps and forecasts

+

1-year audit log

+

Named support engineer

Request access

→

Sovereign

Custom

self-hosted, any size

For teams that must run everything inside their own walls.

+

Self-hosted control plane

+

Residency proofs for auditors

+

Unlimited audit retention

+

24/7 support with a 15-minute response

+

Installed by our engineers on site

Talk to us

→

(11)

Changelog

What shipped, what broke, what we fixed.

All changes

→

Release

v4.2

Spend caps per project

Caps can now be set per project as well as per team. Jobs that would cross a cap are held, and the owner is asked.

Fix

v4.1.3

Faster resume after node loss

Training runs now resume from the last checkpoint in under 40 seconds on most pools, down from about three minutes.

Policy

Policy 41

EU residency policy pack

A ready-made set of rules that keeps datasets and the jobs that read them inside EU regions.

Release

v4.1

Region view in the Terminal

A new pane shows each region's power, GPUs and queue depth side by side. Press R to open it.

(12)

Questions

Asked before every pilot.

Ask us something else

→

Does Overhead replace our scheduler?

+

Where does our data go?

+

Can we run it ourselves?

+

How long does setup take?

+

What happens when a policy blocks a job?

+

Which hardware do you support?

+

How is pricing counted?

+

Can we try it on real workloads?

+

Request access

Put your whole stack on one screen.

Work email

We connect one cluster with you, read-only, for 30 days. No card needed.

Create a free website with Framer, the website builder loved by startups, designers and agencies.