22:00 · Friday night

You submit the sweep and go home.

train-resnet-sweep — one MI300X, steady at ninety percent. It should run all night. You stop thinking about it. That's the deal.

train-resnet-sweep · katie · 1×MI300X GPU UTIL · LIVE
epoch 3 util 91% mem 142G temp 61°C
02:14 · four hours later

Slurm kills it.
You're asleep.


    

Nine words from the scheduler. On most clusters, that is the entire explanation you will ever get — and it's waiting for you behind an SSH session and four thousand lines of stderr.

02:14:09 · nine seconds later

Chamber has already read the logs.

Not a guess from the exit code — it reads the job's own output and telemetry, classifies the failure, and drafts the exact fix. While you sleep.

Timed out Needs you diagnosis · 02:14:09
Classification walltime-exceeded
What happened Throughput was steady and utilisation held near 90% until the scheduler killed the job at its one-minute limit. No stall, no crash — it simply needed more wall time than it was given.
Proposed fix · relaunch
timeLimit 00:01:00 00:05:00
02:14 → 07:02 · the part we refuse to automate

The agent cannot press this button.

Nothing writes to the cluster until a human says so. Not our agent, not anyone's. The rest of this page is behind the same gate the cluster is.

remediation · relaunch @ 00:05:00 · awaiting approval

a real decision — both buttons work

✓ EXECUTED · 07:02
relaunched as job 1871 · timeLimit 00:05:00
submitted under your credential — not the agent's

✕ REJECTED · 07:02
proposal discarded · cluster untouched
that's not a failure mode — that's the product.

07:02 · one tap later

Back to ninety percent.

The relaunch carries the corrected limit and runs as you — your quota, your accounting, your name. The lineage is on the record: parent, proposal, decision, child.

train-resnet-sweep · relaunch GPU UTIL · RESUMED
timed out job 1860 · parent 02:14
succeeded job 1871 · child · +4m walltime 07:41
08:00 · what the night cost you

You read zero log files.

<10s
failure → proposed fix
1
tap to recover
0
SSH sessions at 2am
100%
writes that need a human

Measured on our live pilot: scheduler reports the job dead → classified diagnosis with a proposed relaunch in the approval queue. The wait for your tap is real and it stays — that's the gate, not a bug.

console — fleet, runs, approvals cli — chamber runs efficiency mcp — Claude Code / Cursor tools chat — “why did my sweep die?” gpus — AMD MI300X + NVIDIA scheduler — Slurm now, seam for k8s
at the booth

Come break one of our jobs.

We'll kill a real run on a real MI300X cluster, watch it get diagnosed live, and hand you the approve button. It's more fun than a demo video.

your move · routed straight to the founders

no list, no drip — a founder replies

✓ FILED
a founder will reply from founders@usechamber.io
usually same-day — it's booth season.

↻ replay the night