train-resnet-sweep — one MI300X, steady at ninety percent. It should run all night. You stop thinking about it. That's the deal.
Nine words from the scheduler. On most clusters, that is the entire explanation you will ever get — and it's waiting for you behind an SSH session and four thousand lines of stderr.
Not a guess from the exit code — it reads the job's own output and telemetry, classifies the failure, and drafts the exact fix. While you sleep.
Nothing writes to the cluster until a human says so. Not our agent, not anyone's. The rest of this page is behind the same gate the cluster is.
remediation · relaunch @ 00:05:00 · awaiting approval
a real decision — both buttons work
✓ EXECUTED · 07:02
relaunched as job 1871 · timeLimit 00:05:00
submitted under your credential — not the agent's
✕ REJECTED · 07:02
proposal discarded · cluster untouched
that's not a failure mode — that's the product.
The relaunch carries the corrected limit and runs as you — your quota, your accounting, your name. The lineage is on the record: parent, proposal, decision, child.
Measured on our live pilot: scheduler reports the job dead → classified diagnosis with a proposed relaunch in the approval queue. The wait for your tap is real and it stays — that's the gate, not a bug.
We'll kill a real run on a real MI300X cluster, watch it get diagnosed live, and hand you the approve button. It's more fun than a demo video.
your move · routed straight to the founders
✓ FILED
a founder will reply from founders@usechamber.io
usually same-day — it's booth season.