An HTTP POST through the Sprite URL doesn't wake a paused Sprite (GET does)

Summary

  1. Main issue. While a Sprite is paused, an HTTP POST to <sprite-url>/mcp (https://<sprite>-<orgid>.sprites.app) gets no response until the client times out. That happened in 6 of 6 tries: 90 s timeout ×4, 45 s ×2, with both Python http.client and curl. The POST doesn’t seem to resume the VM. The service only processes it when something else resumes the VM (an exec session). If the client is still connected at that point, it gets the response: one POST returned its 401 after 30.29 s, 285 ms after we opened an exec session. A GET /health to the same paused Sprite is answered (200 in 0.15 s). But a POST sent 0.15–0.2 s after that GET still hangs. A POST only answers quickly (0.18–0.22 s) while an exec session is open.
  2. Secondary issue / docs gap. GET /v1/sprites/{name} flips between warm and cold during idle. When it says cold, last_running_at/last_warming_at are null. Once it said warm and then cold 119 ms later. Meanwhile the VM’s boot (btime, boot_id prefix) never changed and the Service PIDs stayed the same. We’ve read the staff reply in When is a Sprite actually cold? Reported-cold Sprite woke with its original process running (“cold” is broad and includes suspended Sprites whose memory is intact), so we’re asking for docs and a clear signal, not reporting a bug.

Environment

  • Sprites (Fly.io), one Sprite (<sprite>) in a single org. Region and org are left out here and can be shared privately.
  • URL settings: auth: sprite (private, org token required at the edge). private_access: admins.
  • One registered Service with http_port: 8080 (the only HTTP service). Plain HTTP/1.1 server. Responses carry Content-Length (no chunked encoding, no streaming).
  • Endpoint under test: POST /mcp, 41-byte JSON body {"jsonrpc":"2.0","id":1,"method":"ping"} with Content-Type: application/json and Content-Length. We send the org token as Authorization: Bearer for the edge and no app-level token, so the expected answer is 401 {"error":"unauthorized"}. (Observed earlier, not captured: the edge strips Authorization before forwarding, which is why the app uses its own header.)
  • Clients: Python http.client with a fresh TLS connection per request and Connection: close. curl as a cross-check. Both run from an off-Fly Linux host.
  • Responses carry server: Fly/75174e2d02 (2026-10-08) and a fly-request-id header.

Minimal repro

Placeholders: $ORG_TOKEN = a Sprites org token; <sprite> = Sprite name; <sprite-url> = the Sprite’s URL.

  1. Create a Sprite with one Service on port 8080 that answers POST /mcp with a small 401 JSON body (with Content-Length) and logs one line per request. Register it with an HTTP port:

    sprite exec -s <sprite> -- sprite-env services create web --cmd python3 --args /home/sprite/server.py --http-port 8080
    

    Keep URL auth at the default sprite.

  2. Baseline with the VM held by an exec session. In shell A run sprite exec -s <sprite> -- sleep 20. Within that window, in shell B:

    curl -sS -m 30 -w '\nPOST %{http_code} %{time_total}s\n' \
      -H "Authorization: Bearer $ORG_TOKEN" -H 'Content-Type: application/json' \
      -H 'Accept: application/json' \
      --data-binary '{"jsonrpc":"2.0","id":1,"method":"ping"}' <sprite-url>/mcp
    

    → 401 in ~0.2 s.

  3. Close all exec/console sessions and send nothing to the URL for at least 90 s.

  4. Optionally check status (control plane only; this doesn’t wake the VM):

    curl -sS -H "Authorization: Bearer $ORG_TOKEN" \
      https://api.sprites.dev/v1/sprites/<sprite> | jq '{status,last_running_at,last_warming_at}'
    
  5. Send the POST from step 2 with -m 90.
    → No response. curl exits 28 (“Operation timed out … with 0 bytes received”).

  6. Idle again for 90 s or more. Send the POST with -m 120, and about 30 s later, in another shell, run sprite exec -s <sprite> -- true.
    → The POST completes with 401 right after the exec session starts.

  7. Idle again for 90 s or more, then curl -sS -o /dev/null -w '%{http_code} %{time_total}s\n' -H "Authorization: Bearer $ORG_TOKEN" <sprite-url>/health.
    → 200 in ~0.15 s.

  8. Immediately after step 7 (no exec), send the POST again with -m 45.
    → Still no response until the timeout.

  9. Through exec, read the service’s request log line count and the log file’s mtime before and after a timed-out POST.
    → The line for the timed-out POST is written when the exec resumes the VM, not when the POST was sent.

Expected vs. actual

Expected Actual (this capture)
POST to a paused Sprite The request wakes the VM like any other request (“The next request resumes it in 100–500ms”, lifecycle docs) No response until client timeout (6/6). The request is processed only when something else resumes the VM
POST 0.15–0.2 s after a successful GET (no exec) fast No response until client timeout (2/2: one Python, one curl)
POST while an exec session is open fast 401 in 0.176–0.216 s
GET /health to a paused Sprite 100–500 ms warm / 1–2 s cold 200 in 0.150–0.153 s
status while the VM resumes warm stable warm (or a documented meaning for cold) flips warm ↔ cold, null timestamps when cold, including a warm → cold flip 119 ms apart

Timings (UTC, 2026-10-08, one capture)

Step Sent (UTC) Result First byte / total
Exec (status cold just before) 20:27:49.774 rc 0 0.48 s
GET /health, right after exec 20:27:50.251 200 0.153 s
POST #1, 0.15 s after that GET 20:27:50.405 client timeout, 0 bytes 90.176 s
POST #2, immediately after #1 20:29:20.584 client timeout 90.161 s
POST after 100 s idle (application/json) 20:32:31.060 client timeout 90.174 s
POST after 100 s idle, 120 s timeout 20:35:41.658 401 30.293 s, completed at 20:36:11.952
… exec opened 30 s after that POST 20:36:11.667 rc 0 0.32 s. The POST completed 285 ms after exec start
GET /health after 100 s idle (status cold just before) 20:37:52.308 200 0.150 s
POST after 100 s idle (text/event-stream) 20:39:32.737 client timeout 90.130 s
POST ×2 while an exec session (sleep 20) was open 20:41:45.014 / 20:41:45.231 401 / 401 0.216 s / 0.176 s
GET /health in the same exec window 20:41:45.409 200 0.219 s
curl GET /health (≈5 s after that exec ended) 20:42:06.361 200 0.201 s
curl POST, 0.2 s after that GET 20:42:06.571 curl rc 28, 0 bytes 45.002 s
POST after 100 s idle, 45 s timeout 20:44:32.266 client timeout 45.122 s

When the service saw the requests (the request log has no timestamps, so we read the line count and file mtime through exec):

  • 20:41:41, before the two exec-held POSTs: 12 lines, mtime 20:41:03.224. 20:42:51: 14 lines, mtime 20:41:45.353, which matches the second exec-held POST. So the held-window POSTs were logged right away.
  • The previous mtime, 20:41:03.224, matches an exec at 20:41:03.039 that came right after the text/event-stream POST timed out on the client (sent 20:39:32.737, client gave up 20:41:02.867). That request was logged only when the exec resumed the VM.
  • Before the last 45 s POST: 14 lines. 10 s after the client gave up (20:45:17.388), an exec at 20:45:27.549 read 15 lines, mtime 20:45:27.760. Again, a POST was logged only at the exec-triggered resume. Across the curl POST and the last 45 s POST together, only one new line appeared.

VM continuity across the whole capture: btime 1791486213 at the start and end. uptime 5056.93 s → 6114.90 s, which is continuous (+1057.97 s over 1058.0 s of wall time). The same Service PIDs throughout, restarts 0. The boot_id prefix at the end matches the one recorded earlier the same day. No cold boot happened.

Status API during the capture (status last_running_at last_warming_at):

20:27:49.773 cold    null                 null                  (exec then ran at 20:27:49.774)
20:31:20.937 cold    null                 null
20:31:51.037 warm    2026-10-08T20:27:50Z 2026-10-08T20:27:50Z
20:32:30.939 cold    null                 null
20:35:01.523 warm    2026-10-08T20:27:50Z 2026-10-08T20:27:50Z
20:35:31.667 cold    null                 null
20:36:42.195 warm    2026-10-08T20:27:50Z 2026-10-08T20:36:11Z  (= the exec time)
20:37:42.605 cold    null                 null
20:38:22.970 warm    2026-10-08T20:27:50Z 2026-10-08T20:37:52Z  (= the GET /health time)
20:42:01.360 running 2026-10-08T20:41:41Z null                  (during/just after exec; 20:41:41 = exec start)
20:44:32.146 warm    2026-10-08T20:41:41Z 2026-10-08T20:42:51Z
20:44:32.265 cold    null                 null                  (119 ms later)

last_running_at moved only for exec sessions (20:41:41). last_warming_at moved after the exec (20:36:11) and after the GET /health (20:37:52), but never after any POST, so POSTs don’t appear to resume the VM at all. running showed up only while an exec session was open. (In the second capture, with a Tasks API hold active and POSTs answering in ~0.19 s, the status API still reported cold with null timestamps at 20:50:18.831, 20:50:28.708 and 20:51:29.054 UTC, in between running readings.)

Observed earlier, not captured:

  • In a separate earlier poll the same day, status said cold continuously from 19:45:24Z to at least 19:57:26Z (and reportedly until about 20:20Z) while the VM was warm-resumable.
  • Over about 22 minutes idle the Sprite never went truly cold, and there is no API to force it.
  • On localhost inside the Sprite the POST answers in under 1 ms.
  • A checkpoint’s create_time seemed to show the layer start time rather than the snapshot time.

Impact

Any client that talks to a Sprite over HTTP POST breaks once the Sprite has paused. MCP over Streamable HTTP is the clearest case: every call is a POST with a JSON body. After an idle pause the first tool call gets nothing until the client times out. The request is then processed late, whenever something else wakes the VM, so a retry can make the server run the same call twice. The same applies to webhooks and JSON-RPC/REST APIs. GET wakes the VM but POST doesn’t, which undercuts the “Services + wake-on-request = server that costs nothing while idle” pattern in the Services docs.

Workarounds we tried or are considering

  1. GET warm-up before POST: tested, doesn’t work on its own. The GET is answered in ~0.15 s, but a POST sent 0.15–0.2 s later still hung (2/2: one Python, one curl). Any GET warm-up would also have to hold the VM, not just touch it.
  2. Hold the Sprite active with the Tasks API: tested, works (second capture, 2026-10-08 20:48:47–20:55:18 UTC; internal ref notes-ops-test/taskhold-20261008T204847Z.log).
    • We created a task via exec on the in-Sprite socket (POST http://sprite/v1/tasks {"name":…,"expire":"10m"} on /.sprite/api.sock, which returned 201) and then closed all exec sessions (the session list showed 0).
    • After 100 s idle, 3 POSTs spaced 60 s apart all returned 401 in 0.191 s, 0.181 s and 0.194 s.
    • We deleted the task (204; a later GET returned 404 and the task list was empty). After another 100 s idle, a control POST hung again until its 30 s timeout.
    • Boot unchanged throughout (same btime and boot_id prefix).
    • Cost while held (public rates at https://fly.io/sprites: $0.0385/CPU-hour on cpu.stat CPU time, $0.021875/GB-hour on actual memory, $0.000683/GB-hour hot storage). Measured on our Sprite during the hold: ~0.024 CPUs average, cgroup memory.current ≈ 2.29 GB, ~2.9 GB disk used. That works out to about $0.053 per hour held, or about $1.27/day and $39/month if held 24/7. About 95% of that is memory, because memory.current includes page cache.
    • The cost is usage-based. If the Sprite pegged its full 8 vCPU / 8 GiB it would be about $0.50/hour.
    • A task also doesn’t survive a cold boot (Services docs: “No, it’s a hold, not a process”), so the client side would still need a heartbeat or a re-create.

Questions for Fly

  1. Is it intended that a POST through the sprite URL doesn’t resume a paused Sprite, while a GET does? The POST seems to be queued and delivered on the next resume, which can be long after the client gave up.
  2. Is there a recommended pattern for request/response HTTP APIs (MCP, webhooks) on Sprites, other than holding a task the whole time?
  3. Status API: could status (or a new field) tell “suspended, memory intact” apart from “stopped, memory discarded”? Could last_running_at/last_warming_at stay populated instead of going null when the status reads cold? Could the lifecycle docs, which say cold means “the VM is fully stopped and in-memory state is dropped”, reflect the broader meaning?

We can share the Sprite name, org, fly-request-ids and exact timestamps privately.

Hi, a POST to a paused Sprite should wake it just like a GET does, so this isn’t intended. Can you share the sprite name, organization, and a few fly-request-ids if you have them, for the timed-out POSTs?

Until then, holding the Sprite with a task is the right workaround. You’re also right about the status API: some reads wrongly come back cold with null timestamps. I’ll look into fixing that and updating the docs.

Thanks for confirming. Details:

- Sprite: `notes-a`

- Org: `admin-mischief-dev`

- Region: iad

The POSTs that hung never got a response, so I have no fly-request-id for those. They queued and only ran once a GET or exec woke the sprite. One example: POSTs sent at about 15:10:10Z on 2026-10-09 hung and were processed together at 15:10:39Z, right after an exec woke the sprite.

We also saw HTTP 502 from the edge on a DELETE (~15:58:55Z) and on a POST (~16:14Z) that never reached the app. I didn’t capture request ids for those either.

I’m now logging fly-request-id and timestamps on every request and will post the next ones here. Two from just now, both of which woke the sprite fine: `01M4H9MJWVDQHC6BBRZHW70HJG-iad`, `01M4H9QB6HHTHR8R23JD5NF54Q-iad`.

Hi, so I tried reproducing this on a sprite of mine but it wakes up reliably with both a POST and a GET. The main difference is that my sprite is public, doesn’t use token authentication like yours.

IF you have a chance and could try this on another sprite, that’d be helpful:

  1. create new sprite, set it up with a listening service, make it public, test it wakes on POST and GET
  2. set it back to auth-based, and test the behavior: in here I’d expect it to wake on GET but not on POST.

I’ll also try to repro on my side when time permits.

Thanks!

Thanks for testing. I ran your steps on a fresh sprite in the same org (iad) with a minimal HTTP service on 8080. With it public, POST woke it from cold every time (01M4HEJ8PY9RBB8APRE9Q5P1FA-iad was a GET from cold, 01M4HFFETTMV7PZ1SR44HENSXH-iad a POST from cold). After switching back to sprite auth, a POST from cold also woke it, in about 1.1 s (01M4HJQND6TFWW1S9EV3XDFQ74-iad). So I can’t reproduce it on a clean sprite either way. The hangs we saw were on notes-a, a sprite running several services where the one on 8080 depends on three others. I’ll capture request IDs if it happens again there. Separately, the status API reported “cold” for several seconds after the sprite had clearly woken.