Summary
- Main issue. While a Sprite is paused, an HTTP POST to
<sprite-url>/mcp(https://<sprite>-<orgid>.sprites.app) gets no response until the client times out. That happened in 6 of 6 tries: 90 s timeout ×4, 45 s ×2, with both Pythonhttp.clientand curl. The POST doesn’t seem to resume the VM. The service only processes it when something else resumes the VM (an exec session). If the client is still connected at that point, it gets the response: one POST returned its 401 after 30.29 s, 285 ms after we opened an exec session. A GET/healthto the same paused Sprite is answered (200 in 0.15 s). But a POST sent 0.15–0.2 s after that GET still hangs. A POST only answers quickly (0.18–0.22 s) while an exec session is open. - Secondary issue / docs gap.
GET /v1/sprites/{name}flips betweenwarmandcoldduring idle. When it sayscold,last_running_at/last_warming_atarenull. Once it saidwarmand thencold119 ms later. Meanwhile the VM’s boot (btime, boot_id prefix) never changed and the Service PIDs stayed the same. We’ve read the staff reply in When is a Sprite actually cold? Reported-cold Sprite woke with its original process running (“cold” is broad and includes suspended Sprites whose memory is intact), so we’re asking for docs and a clear signal, not reporting a bug.
Environment
- Sprites (Fly.io), one Sprite (
<sprite>) in a single org. Region and org are left out here and can be shared privately. - URL settings:
auth: sprite(private, org token required at the edge).private_access: admins. - One registered Service with
http_port: 8080(the only HTTP service). Plain HTTP/1.1 server. Responses carryContent-Length(no chunked encoding, no streaming). - Endpoint under test:
POST /mcp, 41-byte JSON body{"jsonrpc":"2.0","id":1,"method":"ping"}withContent-Type: application/jsonandContent-Length. We send the org token asAuthorization: Bearerfor the edge and no app-level token, so the expected answer is401 {"error":"unauthorized"}. (Observed earlier, not captured: the edge stripsAuthorizationbefore forwarding, which is why the app uses its own header.) - Clients: Python
http.clientwith a fresh TLS connection per request andConnection: close. curl as a cross-check. Both run from an off-Fly Linux host. - Responses carry
server: Fly/75174e2d02 (2026-10-08)and afly-request-idheader.
Minimal repro
Placeholders: $ORG_TOKEN = a Sprites org token; <sprite> = Sprite name; <sprite-url> = the Sprite’s URL.
-
Create a Sprite with one Service on port 8080 that answers
POST /mcpwith a small401JSON body (withContent-Length) and logs one line per request. Register it with an HTTP port:sprite exec -s <sprite> -- sprite-env services create web --cmd python3 --args /home/sprite/server.py --http-port 8080Keep URL auth at the default
sprite. -
Baseline with the VM held by an exec session. In shell A run
sprite exec -s <sprite> -- sleep 20. Within that window, in shell B:curl -sS -m 30 -w '\nPOST %{http_code} %{time_total}s\n' \ -H "Authorization: Bearer $ORG_TOKEN" -H 'Content-Type: application/json' \ -H 'Accept: application/json' \ --data-binary '{"jsonrpc":"2.0","id":1,"method":"ping"}' <sprite-url>/mcp→
401in ~0.2 s. -
Close all exec/console sessions and send nothing to the URL for at least 90 s.
-
Optionally check status (control plane only; this doesn’t wake the VM):
curl -sS -H "Authorization: Bearer $ORG_TOKEN" \ https://api.sprites.dev/v1/sprites/<sprite> | jq '{status,last_running_at,last_warming_at}' -
Send the POST from step 2 with
-m 90.
→ No response. curl exits 28 (“Operation timed out … with 0 bytes received”). -
Idle again for 90 s or more. Send the POST with
-m 120, and about 30 s later, in another shell, runsprite exec -s <sprite> -- true.
→ The POST completes with401right after the exec session starts. -
Idle again for 90 s or more, then
curl -sS -o /dev/null -w '%{http_code} %{time_total}s\n' -H "Authorization: Bearer $ORG_TOKEN" <sprite-url>/health.
→200in ~0.15 s. -
Immediately after step 7 (no exec), send the POST again with
-m 45.
→ Still no response until the timeout. -
Through exec, read the service’s request log line count and the log file’s mtime before and after a timed-out POST.
→ The line for the timed-out POST is written when the exec resumes the VM, not when the POST was sent.
Expected vs. actual
| Expected | Actual (this capture) | |
|---|---|---|
| POST to a paused Sprite | The request wakes the VM like any other request (“The next request resumes it in 100–500ms”, lifecycle docs) | No response until client timeout (6/6). The request is processed only when something else resumes the VM |
| POST 0.15–0.2 s after a successful GET (no exec) | fast | No response until client timeout (2/2: one Python, one curl) |
| POST while an exec session is open | fast | 401 in 0.176–0.216 s |
GET /health to a paused Sprite |
100–500 ms warm / 1–2 s cold | 200 in 0.150–0.153 s |
status while the VM resumes warm |
stable warm (or a documented meaning for cold) |
flips warm ↔ cold, null timestamps when cold, including a warm → cold flip 119 ms apart |
Timings (UTC, 2026-10-08, one capture)
| Step | Sent (UTC) | Result | First byte / total |
|---|---|---|---|
Exec (status cold just before) |
20:27:49.774 | rc 0 | 0.48 s |
| GET /health, right after exec | 20:27:50.251 | 200 | 0.153 s |
| POST #1, 0.15 s after that GET | 20:27:50.405 | client timeout, 0 bytes | 90.176 s |
| POST #2, immediately after #1 | 20:29:20.584 | client timeout | 90.161 s |
POST after 100 s idle (application/json) |
20:32:31.060 | client timeout | 90.174 s |
| POST after 100 s idle, 120 s timeout | 20:35:41.658 | 401 | 30.293 s, completed at 20:36:11.952 |
| … exec opened 30 s after that POST | 20:36:11.667 | rc 0 | 0.32 s. The POST completed 285 ms after exec start |
GET /health after 100 s idle (status cold just before) |
20:37:52.308 | 200 | 0.150 s |
POST after 100 s idle (text/event-stream) |
20:39:32.737 | client timeout | 90.130 s |
POST ×2 while an exec session (sleep 20) was open |
20:41:45.014 / 20:41:45.231 | 401 / 401 | 0.216 s / 0.176 s |
| GET /health in the same exec window | 20:41:45.409 | 200 | 0.219 s |
| curl GET /health (≈5 s after that exec ended) | 20:42:06.361 | 200 | 0.201 s |
| curl POST, 0.2 s after that GET | 20:42:06.571 | curl rc 28, 0 bytes | 45.002 s |
| POST after 100 s idle, 45 s timeout | 20:44:32.266 | client timeout | 45.122 s |
When the service saw the requests (the request log has no timestamps, so we read the line count and file mtime through exec):
- 20:41:41, before the two exec-held POSTs: 12 lines, mtime 20:41:03.224. 20:42:51: 14 lines, mtime 20:41:45.353, which matches the second exec-held POST. So the held-window POSTs were logged right away.
- The previous mtime, 20:41:03.224, matches an exec at 20:41:03.039 that came right after the
text/event-streamPOST timed out on the client (sent 20:39:32.737, client gave up 20:41:02.867). That request was logged only when the exec resumed the VM. - Before the last 45 s POST: 14 lines. 10 s after the client gave up (20:45:17.388), an exec at 20:45:27.549 read 15 lines, mtime 20:45:27.760. Again, a POST was logged only at the exec-triggered resume. Across the curl POST and the last 45 s POST together, only one new line appeared.
VM continuity across the whole capture: btime 1791486213 at the start and end. uptime 5056.93 s → 6114.90 s, which is continuous (+1057.97 s over 1058.0 s of wall time). The same Service PIDs throughout, restarts 0. The boot_id prefix at the end matches the one recorded earlier the same day. No cold boot happened.
Status API during the capture (status last_running_at last_warming_at):
20:27:49.773 cold null null (exec then ran at 20:27:49.774)
20:31:20.937 cold null null
20:31:51.037 warm 2026-10-08T20:27:50Z 2026-10-08T20:27:50Z
20:32:30.939 cold null null
20:35:01.523 warm 2026-10-08T20:27:50Z 2026-10-08T20:27:50Z
20:35:31.667 cold null null
20:36:42.195 warm 2026-10-08T20:27:50Z 2026-10-08T20:36:11Z (= the exec time)
20:37:42.605 cold null null
20:38:22.970 warm 2026-10-08T20:27:50Z 2026-10-08T20:37:52Z (= the GET /health time)
20:42:01.360 running 2026-10-08T20:41:41Z null (during/just after exec; 20:41:41 = exec start)
20:44:32.146 warm 2026-10-08T20:41:41Z 2026-10-08T20:42:51Z
20:44:32.265 cold null null (119 ms later)
last_running_at moved only for exec sessions (20:41:41). last_warming_at moved after the exec (20:36:11) and after the GET /health (20:37:52), but never after any POST, so POSTs don’t appear to resume the VM at all. running showed up only while an exec session was open. (In the second capture, with a Tasks API hold active and POSTs answering in ~0.19 s, the status API still reported cold with null timestamps at 20:50:18.831, 20:50:28.708 and 20:51:29.054 UTC, in between running readings.)
Observed earlier, not captured:
- In a separate earlier poll the same day, status said
coldcontinuously from 19:45:24Z to at least 19:57:26Z (and reportedly until about 20:20Z) while the VM was warm-resumable. - Over about 22 minutes idle the Sprite never went truly cold, and there is no API to force it.
- On localhost inside the Sprite the POST answers in under 1 ms.
- A checkpoint’s
create_timeseemed to show the layer start time rather than the snapshot time.
Impact
Any client that talks to a Sprite over HTTP POST breaks once the Sprite has paused. MCP over Streamable HTTP is the clearest case: every call is a POST with a JSON body. After an idle pause the first tool call gets nothing until the client times out. The request is then processed late, whenever something else wakes the VM, so a retry can make the server run the same call twice. The same applies to webhooks and JSON-RPC/REST APIs. GET wakes the VM but POST doesn’t, which undercuts the “Services + wake-on-request = server that costs nothing while idle” pattern in the Services docs.
Workarounds we tried or are considering
- GET warm-up before POST: tested, doesn’t work on its own. The GET is answered in ~0.15 s, but a POST sent 0.15–0.2 s later still hung (2/2: one Python, one curl). Any GET warm-up would also have to hold the VM, not just touch it.
- Hold the Sprite active with the Tasks API: tested, works (second capture, 2026-10-08 20:48:47–20:55:18 UTC; internal ref
notes-ops-test/taskhold-20261008T204847Z.log).- We created a task via exec on the in-Sprite socket (
POST http://sprite/v1/tasks {"name":…,"expire":"10m"}on/.sprite/api.sock, which returned 201) and then closed all exec sessions (the session list showed 0). - After 100 s idle, 3 POSTs spaced 60 s apart all returned
401in 0.191 s, 0.181 s and 0.194 s. - We deleted the task (204; a later GET returned 404 and the task list was empty). After another 100 s idle, a control POST hung again until its 30 s timeout.
- Boot unchanged throughout (same btime and boot_id prefix).
- Cost while held (public rates at https://fly.io/sprites: $0.0385/CPU-hour on cpu.stat CPU time, $0.021875/GB-hour on actual memory, $0.000683/GB-hour hot storage). Measured on our Sprite during the hold: ~0.024 CPUs average, cgroup
memory.current≈ 2.29 GB, ~2.9 GB disk used. That works out to about $0.053 per hour held, or about $1.27/day and $39/month if held 24/7. About 95% of that is memory, becausememory.currentincludes page cache. - The cost is usage-based. If the Sprite pegged its full 8 vCPU / 8 GiB it would be about $0.50/hour.
- A task also doesn’t survive a cold boot (Services docs: “No, it’s a hold, not a process”), so the client side would still need a heartbeat or a re-create.
- We created a task via exec on the in-Sprite socket (
Questions for Fly
- Is it intended that a POST through the sprite URL doesn’t resume a paused Sprite, while a GET does? The POST seems to be queued and delivered on the next resume, which can be long after the client gave up.
- Is there a recommended pattern for request/response HTTP APIs (MCP, webhooks) on Sprites, other than holding a task the whole time?
- Status API: could
status(or a new field) tell “suspended, memory intact” apart from “stopped, memory discarded”? Couldlast_running_at/last_warming_atstay populated instead of goingnullwhen the status readscold? Could the lifecycle docs, which say cold means “the VM is fully stopped and in-memory state is dropped”, reflect the broader meaning?
We can share the Sprite name, org, fly-request-ids and exact timestamps privately.