(01)
Platform
One control layer above every cluster you run.

SITE 02 · BASIN B
18,560 GPUS · 94% BUSY
Sentinel-2 L2A · false-colour pass
(02)
How it works
Four verbs. Every job goes through all of them.

[04] / [04]
Every job
[01]
Observe
Reads every node, job and bill from your clusters and clouds into one inventory, every few seconds.
[02]
Schedule
Places each job where it fits best, and moves it when a node fails or a cheaper slot opens.
[03]
Enforce
Checks your policies before a job starts and while it runs. Breaking a rule holds the job, it does not kill it.
[04]
Audit
Writes down who ran what, where, for how long and at what cost. Export it for any auditor in one click.
(03)
The Terminal
Six panes. Everything that’s running, right now.
[F1]
Queue
Live
[F2]
GPU pools
Live
[F3]
Latency p95
Live
p95 212 ms
target 250 ms
[F4]
Events
Live
14:52
Node 41 down · run 7731 moved
14:47
Policy 41 held job 7742
14:31
Pool C scaled +64 GPUs
14:12
Spend cap 71% · team vision
13:58
Checkpoint saved · run 7731
13:40
Region us-west-1 maintenance
[F5]
Spend today
Live
[F6]
Regions
Live

EU-N
US-E
AP-S
[Q]
Queue
[P]
Pools
[R]
Regions
[/]
Search
[?]
Shortcuts

(04)
Fleet utilisation
Load, drawn as terrain.
Peaks are clusters near their limit. Valleys are capacity you are paying for and not using. The plate redraws every minute.
00:00
06:00
12:00
18:00
23:59 UTC
Now · 14:52 UTC
Peak right now
US-EAST-2 · 97.4%
Pool B · B200 · 1,536 GPUs
(05)
Capabilities
Everything in the control layer.
[I]
Inventory
Every GPU, node, cage and cloud account, refreshed every few seconds.
[S]
Scheduler
Places and moves jobs across sites. Keeps checkpoints when nodes fail.
[P]
Policy engine
Plain files in your repository, checked before and during every job.
[$]
Spend
Live cost per job, team and project, with caps and forecasts.
[A]
Audit log
Every decision, kept for as long as you choose and exportable in one click.
[>]
API and CLI
Everything on the screen is also an endpoint and a command.
(06)
Security and control
Rules you can read. A record you can prove.
policies/eu-residency.yaml
policy: eu-residency
applies_to: datasets.tag == “eu”
allow_regions:
- eu-north-1
- eu-west-2
on_violation: hold
notify: owner, #platform
approver: security-oncall
Audit log · last 6
14:52:08
moved run-7731 → pool-b
14:47:31
held run-7742 (policy 41)
14:31:02
scaled pool-c +64
14:12:44
cap warning team-vision
13:58:19
checkpoint run-7731
13:40:00
maintenance us-west-1
Self-hosted option
Run the whole control plane inside your own network.
Metadata only
Weights and datasets never leave your clusters.
One-click rollback
Every policy change can be undone, and the undo is logged.
Next step
See your own fleet on one screen.
Connect one cluster, read-only, for 30 days.