(01)

Platform

One control layer above every cluster you run.

Overhead reads your clusters, clouds and colocation cages, schedules work across them and enforces your rules. It runs alongside what you have; nothing gets ripped out.

Overhead reads your clusters, clouds and colocation cages, schedules work across them and enforces your rules. It runs alongside what you have; nothing gets ripped out.

SITE 02 · BASIN B

18,560 GPUS · 94% BUSY

POLICY 41

EU DATA STAYS IN EU-NORTH

Sentinel-2 L2A · false-colour pass

(02)

How it works

Four verbs. Every job goes through all of them.

Overhead sits above the schedulers and clouds you already use. It does four things, in order, for every job you run.

Overhead sits above the schedulers and clouds you already use. It does four things, in order, for every job you run.

[04] / [04]

Every job

[01]

Observe

Reads every node, job and bill from your clusters and clouds into one inventory, every few seconds.

[02]

Schedule

Places each job where it fits best, and moves it when a node fails or a cheaper slot opens.

[03]

Enforce

Checks your policies before a job starts and while it runs. Breaking a rule holds the job, it does not kill it.

[04]

Audit

Writes down who ran what, where, for how long and at what cost. Export it for any auditor in one click.

(03)

The Terminal

Six panes. Everything that’s running, right now.

The Terminal is Overhead’s main screen. Every value on it is live, and every pane opens with one key.

The Terminal is Overhead’s main screen. Every value on it is live, and every pane opens with one key.

[F1]

Queue

Live

Queue
01.Run 7731 · train-70b62.4 % done
02.Run 7735 · eval-suite18 % done
03.Run 7740 · embed-docs91.2 % done
04.Batch 2210 · index1,204 jobs
05.Held · policy 413 jobs
1,204 queued · 38 moved today

[F2]

GPU pools

Live

GPU pools
01.Pool A · H20094.6 % busy
02.Pool B · B20097.6 % busy
03.Pool C · L40S61.0 % busy
04.Edge · mixed38.2 % busy
05.Idle found6.1 %
18,432 GPUs · 6 regions

[F3]

Latency p95

Live

p95 212 ms

target 250 ms

[F4]

Events

Live

14:52

Node 41 down · run 7731 moved

14:47

Policy 41 held job 7742

14:31

Pool C scaled +64 GPUs

14:12

Spend cap 71% · team vision

13:58

Checkpoint saved · run 7731

13:40

Region us-west-1 maintenance

[F5]

Spend today

Live

Spend today
01.Team research$9,240
02.Team vision$12,780
03.Team serving$6,410
04.Team data$2,130
05.Forecast$31,900
Caps on 9 teams · 0 breaches

[F6]

Regions

Live

EU-N

US-E

AP-S

[Q]

Queue

[P]

Pools

[R]

Regions

[/]

Search

[?]

Shortcuts

(04)

Fleet utilisation

Load, drawn as terrain.

Peaks are clusters near their limit. Valleys are capacity you are paying for and not using. The plate redraws every minute.

100%

75%

50%

25%

0% UTIL

00:00

06:00

12:00

18:00

23:59 UTC

Now · 14:52 UTC

Peak right now

US-EAST-2 · 97.4%

Pool B · B200 · 1,536 GPUs

Idle

Saturated

(05)

Capabilities

Everything in the control layer.

Six parts, one data model. Each part has an API and a key in the Terminal.

Six parts, one data model. Each part has an API and a key in the Terminal.

[I]

Inventory

Every GPU, node, cage and cloud account, refreshed every few seconds.

[S]

Scheduler

Places and moves jobs across sites. Keeps checkpoints when nodes fail.

[P]

Policy engine

Plain files in your repository, checked before and during every job.

[$]

Spend

Live cost per job, team and project, with caps and forecasts.

[A]

Audit log

Every decision, kept for as long as you choose and exportable in one click.

[>]

API and CLI

Everything on the screen is also an endpoint and a command.

(06)

Security and control

Rules you can read. A record you can prove.

Policies live as plain files in your repository. Every decision Overhead makes is written to an audit log you own.

Policies live as plain files in your repository. Every decision Overhead makes is written to an audit log you own.

policies/eu-residency.yaml

policy: eu-residency

applies_to: datasets.tag == “eu”

allow_regions:

- eu-north-1

- eu-west-2

on_violation: hold

notify: owner, #platform

approver: security-oncall

Audit log · last 6

14:52:08

moved run-7731 → pool-b

14:47:31

held run-7742 (policy 41)

14:31:02

scaled pool-c +64

14:12:44

cap warning team-vision

13:58:19

checkpoint run-7731

13:40:00

maintenance us-west-1

Self-hosted option

Run the whole control plane inside your own network.

Metadata only

Weights and datasets never leave your clusters.

One-click rollback

Every policy change can be undone, and the undo is logged.

Next step

See your own fleet on one screen.

Connect one cluster, read-only, for 30 days.

Create a free website with Framer, the website builder loved by startups, designers and agencies.