App hla-law-app, region fra, machine 784e1eefe2ee48 (single machine, shared-cpu-1x, 512 MB, volume vol_r1jw5p5n7m1kme9r). Config: auto_stop_machines = "suspend", min_machines_running = 0.
Timeline (UTC, 2026-10-06):
- 20:47:40 proxy
cordon, thensuspensionstarted - 20:48:18 flyd:
suspended - 20:52:57 proxy
start→ machine enteredstartingand never left it
The machine stayed in starting until ~12:05 UTC on 2026-10-07. During that time, proxy logs showed repeated machine failed to start, currently starting and rate limit exceeded, then could not find a good candidate within 40 attempts at load balancing. All requests timed out.
Recovery attempts failed:
fly machine stop:unable to stop machine, current state invalid, starting(Request ID01M4B3X1M8V4N0JM2DCVHMGTHY-ams)fly machine start:machine failed to start, currently starting(01M4B3XZV5S13N3QZFA68D2JEM-ams)fly machine restart --force:internal: internal server error(01M4B3Y4G884M8G3X7DM0ZJPEA-ams)
I had to destroy the machine and redeploy. There was no deploy or config change around the failure; the image had been running since 2026-09-24.
Questions: why did the resume hang, and why didn’t it fall back to a cold boot? Is suspend currently safe for a single-machine app with a volume?