SIN edge latency and PU02 multihop timeouts to LAX since Sep 21

App: dreamteam-match-server; only app Machine runs in lax.

Philippine player API latency rose abruptly on Sep 21 around 09:47 UTC, six hours before our first deployment that day. In Fly Grafana, sin edge HTTP p95 rose from 480 ms at 09:47 to 977 ms at 09:48 and 2,676 ms at 09:49. lax ingress stayed around 170–185 ms; application handler p95 was 183 ms and CPU 13%. SIN response rate dropped from 53 to 42/s across those minutes. Both SIN edge hosts, b323 and c11b, slowed together on successful 2xx responses; the same Sunday clock points were stable around 460–480 ms. Fly’s SIN response-code counter first shows 502s at 09:49; none appeared in the matching Sunday window. Later that Monday, our Fly logs recorded 123 PU02 fly-proxy-p2p/.../http-multihop timeouts or resets in 16 minutes.

This persists on Sep 25. A successful direct /health request to our idle dev app (dreamteam-match-server-dev) took 2,506 ms via sin and LAX worker-lsh-lax1-042e (01M3C4NA4Z9WDHYYKP9Z5KK3N1-sin). A direct production request seconds earlier from the same Philippine probe took 217 ms via sin and LAX edge-dp-lax2-8582 (01M3C4N7A1WHFXRHN4Y0RQTC2V-sin). Neither used Cloudflare or the database; the dev Machine was already started. Across repeated probes on eight Philippine networks, the worker-lsh multihop class is much slower than the edge-dp class, but we cannot infer which physical hop is responsible.

A separate public-host problem appeared through Cloudflare Hong Kong→Fly Tokyo: two Sep 25 12:19 UTC /health requests waited 2.37 and 9.23 seconds to first byte (01M3C830EVKWPAQHYRGWWNSTP9-nrt, 01M3C830JZP1786AKZ2NTEMMKS-nrt), while direct Fly from the same probes took about 0.2 seconds. Cloudflare’s own edge processing was 7/12 ms and its origin fetch was 2.30/9.17 seconds. This could be the Cloudflare-to-Fly handoff or a rare NRT proxy stall; overall Fly NRT edge p95 stayed near 0.46 seconds. Cloudflare HKG origin-fetch p95 across Philippine API traffic was 3.15 seconds in that half-hour versus 0.33 seconds in the matching Sunday half-hour.

The Hong Kong slowdown was already present Monday: Sep 21 12:30–12:45 UTC HKG origin p95 reached 5.04 seconds versus 0.34 seconds on Sunday, while app-handler p95 stayed below 0.20 seconds. Fly SIN edge p95 reached 4.30 seconds during that Monday quarter, but Fly NRT edge p95 stayed below 0.45 seconds. We do not have the Monday per-request ingress labels for HKG traffic.

Across the complete Sep 21 12:30–13:00 UTC half-hour, Philippine API origin p95 was 2.02 seconds on about 115,000 estimated requests versus 0.47 seconds on about 142,000 Sunday. HKG origin p95 was 4.26 seconds versus 0.36 seconds Sunday; LAX remained about 0.20 seconds. Friday’s same half-hour was still slow at 1.47 seconds Philippine p95 on about 101,000 requests, while US-country p95 was 0.31 seconds.

Could Fly inspect these request IDs and the Sep 21 SIN-to-LAX proxy/backhaul transition? Which hop changed or failed, and is there a supported way to avoid the affected path while keeping the Machine in LAX? Could you also trace the HKG→NRT origin outliers? We can provide additional request IDs privately if useful.

Update, Sep 25 13:00–13:15 UTC: Philippine /api/* origin p95 reached 3.64 s on about 55,600 estimated requests, versus 0.50 s on about 82,300 the same Sunday quarter. Cloudflare HKG origin p95 was 6.01 s versus 0.36 s Sunday; LAX was 0.26 s. HKG header-receive p95 was 5.68 s, TCP/TLS handshake p95 0 ms. Our app-handler one-minute p95 never exceeded 0.44 s; Fly sin successful-response edge p95 reached 3.70 s. Fly nrt edge p95 stayed below 0.49 s, although its p99 reached 7.40 s (Sunday’s comparable p99 also reached 6.00 s, so this alone is not a diagnosis).

Two new Sep 25 13:19 UTC /health requests through Cloudflare HKG → Fly nrt are traceable:

  • Calamba Converge: 01M3CBGX891B0EGDS56FNTPKBC-nrt, 4,264 ms first byte; cfOrigin=4,213 ms, cfEdge=6 ms; LAX worker-lsh-lax1-6857. Direct Fly production from the same probe in the next sequential test: 211 ms.
  • Manila Scloud: 01M3CBGX1SRSVRREA1T4AWXYMH-nrt, 2,050 ms first byte; cfOrigin=2,025 ms, cfEdge=5 ms; LAX edge-dp-lax2-d732. Direct Fly production next: 683 ms.

In the same round, a direct request from Scloud to our idle dev app bypassed Cloudflare and took 3,493 ms through Fly sin → LAX worker-lsh-lax1-042e: 01M3CBH761KR9PHEQCZ7M1GJ68-sin. No database call. Could you trace the two HKG→NRT requests to separate Cloudflare→Fly transit from Fly proxy delay, and check whether the direct SIN delay has the same underlying backhaul issue?

Another trace candidate from Sep 25 13:39 UTC: a Philippine Manila Scloud GET /health through Cloudflare HKG → Fly nrt took 4,451 ms to first byte; Cloudflare reports 4,371 ms origin fetch and 59 ms edge work. Fly request ID: 01M3CCNJD1E4HY99YX4DHKSFBH-nrt; debug named LAX edge-dp-lax2-8582. The same probe’s next sequential direct production .fly.dev request took 216 ms, and the dev LAX tunnel took 513 ms. This reproduces the HKG/NRT tail even when the LAX target is edge-dp rather than worker-lsh.

It affects the actual game action, not just health checks: for Philippine POST /api/rpc/roll_player_skill during 13:00–14:00 UTC, Cloudflare HKG origin p95 was 4,184 ms on about 21 estimated requests; the all-Philippines p95 for this route at the same Sunday hour was 566 ms. The samples are small, and this origin measure includes any application/database time.

One timing question for Fly: the first ten Base32 characters of the earlier 9,229 ms HKG/NRT request ID (01M3C830JZP1786AKZ2NTEMMKS-nrt) decode to 12:19:21.311 UTC, while its response Date header is 12:19:30 UTC. Does that prefix represent when the NRT proxy received the request? If so, could you inspect where those ~9 seconds went before the LAX app returned headers?

Potential external trigger: Cloudflare’s official Asia-Pacific network incident says multiple subsea cable outages have caused Tokyo–Singapore congestion since Sep 21, 02:20 UTC; they rerouted traffic, and the network remains degraded today: Network Performance Degradation — Asia-Pacific - Cloudflare Status . Our measured game slowdown began later that Monday at 09:47 UTC. On matched 13:30–14:00 UTC Philippine API windows, HKG became ~4% of estimated requests Sunday versus ~37% Friday, and HKG origin p95 rose from 0.38 s to 6.62 s despite lower total game request volume. This confirms a regional Cloudflare incident but does not establish that Fly’s independent direct SIN→LAX delays or PU02s use those specific affected cables. Could Fly check whether its SIN/NRT→LAX proxy paths share this capacity issue, and whether there is a supported reroute? The request IDs above should help distinguish the paths.

One specific fallback-routing question: Fly’s own announcement says SIN edges can try another region when they cannot connect directly to a distant Machine, and flyio-debug reports that intermediate node in fbn: Smarter fly-proxy routing is now available in all regions . Two successful but slow direct Philippine GET /health requests on Sep 25 both had fbn: null: (1) idle dev app, Angeles City, 13:59 UTC, 2,634 ms first byte, 01M3CDT8QS5SX3J7RR9XM7V7R8-sin, edge-dp-sin1-c11b → worker-lsh-lax1-042e; (2) production app, Manila, 14:19 UTC, 2,228 ms, 01M3CEYV5QFJPHS2CG4SRCXY7C-sin, same SIN edge → worker-lsh-lax1-6857. Fly’s SIN successful-response edge p95 was 2.6–3.3 s during 14:05–14:30 UTC while NRT was ~0.45 s and our handler ~0.36 s. Can you locate where these requests waited and explain why fallback did not engage? Is there a supported temporary routing adjustment for SIN→LAX? Our game also uses WebSockets, so please say whether that adjustment would cover them.

Hi! So yes, we have rerouting mechanisms when network links go bad, but that one announced 2 years ago isn’t really working that well anymore. This kind of cross-continent network problems are to be expected, unfortunately, as we currently do not have a private link between all of our regions.

There is work to improve this though! There is an “upgraded” version that’s being applied to 6PN (you’ll see a more detailed write-up about it soon), and that same reroute logic will be used by fly-proxy as well for its request routing purposes. The new rerouting logic should be much smarter in deciding which links to use; we’ve also recently had one upstream provider that allows us to use their private links between most regions, and combined with this new rerouting logic everything shoulld be much better.

At the moment though before this is fully shipped, we would recommend that when possible, run multiple machines close to the client; if that’s not possible, you could also leverage the upgraded 6PN for now by running a “proxy” app in the source region that forwards traffic over 6PN. This will hopefully no longer be necessary soon.

Thanks, that helps. We need to keep authoritative match state and the database near our LAX Machine, so I will test a small SIN-region reverse proxy over 6PN against the idle dev app before changing player traffic. Is the upgraded 6PN rerouting already active for SIN↔LAX, including long-lived WebSockets? Could you also trace the two direct SIN request IDs in post #5 and the 9.2 s HKG→NRT ID in post #1 to identify whether each delay was in Fly’s inter-region path or before Fly received it? That would tell us which workaround covers both failures.

Yes, but keep in mind that this is still experimental and we do not guarantee that we can always pick the best route (yet). We’ll have an official Fresh Produce once we’re more confident about what it does :slight_smile:

Additional impact data for Philippine successful /api/* requests, comparing the same 14:00–15:00 UTC window: estimated origin fetches over 1 second were 4,507/262,548 (1.7%) on Sunday Sep 20; 22,381/212,254 (10.5%) on Monday Sep 21; and 20,849/197,629 (10.5%) on Friday Sep 25. Monday’s window ended before our first 15:49 UTC deployment. On Monday the >1s share was HKG 60.9%, MNL 15.2%, CGY 15.7%, SIN 14.1%, versus LAX 0.6% and SJC 0.8%. On Friday HKG was 24.1% and represented about 69% of slow Philippine requests. These are Cloudflare Adaptive sample estimates, not player counts.

I tested the suggested 6PN path in development only: SIN/NRT entry Machines and an LAX relay supported a 25-second WebSocket heartbeat, but two HKG→Fly NRT→SIN-entry health requests still took 1.3–1.5 seconds. A dev LAX Cloudflare tunnel took about 0.54 seconds in that comparison. The test Machines are stopped; production is unchanged.

Could Fly trace the direct SIN request IDs in post #5 and the HKG→NRT IDs in posts #1/#3? We need to know whether each request waited before Fly ingress or within Fly’s inter-region proxy to choose a production route.

One high-volume game endpoint confirms the Monday onset before any deploy. For Philippine GET /api/stadium/live-matches (Cloudflare edge status 200–499), origin-fetch p95 on Sep 21 was 371 ms at 09:46 UTC (~1,120 estimated requests), jumped to 1,284 ms at 09:47 (~689), then 658 ms at 09:48 (~1,123). The same Sunday minutes were 380/395/361 ms. In the full 14:00–15:00 UTC hour, this endpoint’s p95 was 484 ms Sunday, 1,936 ms Monday, and 1,794 ms Friday. Estimated >1 s counts were 709/82,192 (0.9%) Sunday, 5,555/54,437 (10.2%) Monday, and 7,623/76,764 (9.9%) Friday. These are adaptively sampled estimates; traffic fell on Monday.

Fly’s one-minute SIN edge p95 went 0.48 s at 09:47 → 0.98 s at 09:48 → 2.68 s at 09:49, while the LAX app-handler p95 stayed 0.18–0.20 s and LAX ingress 0.17–0.19 s. This is a real authenticated game request, not just /health. Could you trace the direct SIN and HKG→NRT request IDs in posts #1/#3/#5 to identify where the delay occurred before it reached the LAX handler? That is the missing piece for choosing a safe route change.

Additional Fly-native evidence at that onset: SIN TLS handshake p95 stayed about 49 ms at 09:47–09:49 Monday, matching Sunday, while SIN edge HTTP p95 rose from 0.48 to 2.68 s. App-handler p99 was about 0.47 s at 09:49. fly_edge_error_count became nonzero on both SIN proxy hosts (b323 and c11b) at 09:49; neither had errors in Sunday’s comparison window. This points to delay after the Fly TLS handshake, in proxy forwarding or the return path, but aggregate metrics cannot identify the physical link. Could Fly trace the request IDs cited above and name the failing hop?

A route-selection clue from direct Fly /health probes on eight Philippine networks today: when SIN ingress named a LAX worker-lsh multihop, production first-byte median was 632 ms (14/21 over 500 ms), versus 220 ms (0/47 over 500 ms) when it named LAX edge-dp. The quiet dev app showed the same split: 735 ms (18/20) versus 227 ms (0/35). Every sampled network encountered both node classes. This is an association, not proof the named worker is faulty. Could you inspect whether these two SIN→LAX paths use different proxy hops or underlay links? The request IDs and node names in #5 include both apps.

One more split to help trace the slow path. I regrouped direct Fly GET /health probes from eight Philippine networks through 10:39 UTC Sep 25 by exact SIN ingress host, app, and LAX multihop target. Both SIN hosts served both target types in both apps:

  • Production, edge-dp-sin1-b323: LAX worker-lsh median 677 ms (8/13 over 500 ms); LAX edge-dp median 220 ms (0/24).
  • Production, edge-dp-sin1-c11b: worker-lsh 591 ms (6/8); edge-dp 218 ms (0/23).
  • Quiet dev, b323: worker-lsh 559 ms (7/9); edge-dp 232 ms (0/19).
  • Quiet dev, c11b: worker-lsh 805 ms (11/11); edge-dp 219 ms (0/16).

Within app + Philippine network + exact SIN host groups that encountered both target classes, worker-lsh was slower in 23/24. Samples were sequential and selected, so this is an association, not proof a named worker is faulty. It does show the same ingress hosts can produce both fast and slow paths. Could Fly trace the request IDs in #5 and determine what differs in the SIN→LAX path to worker-lsh versus edge-dp? Is there a temporary way to steer this app away from the slow path, including WebSockets?

A close same-network example on the idle dev app: COMFAC Angeles City through the same edge-dp-sin1-b323 ingress took 944 ms via worker-lsh-lax1-042e (01M3BTCBCMNBPVNKPY5JN5ZRPT-sin, 08:19 UTC), then 211 ms via LAX edge-dp (01M3BTV9HCK1118H2QC0QF8PG4-sin, 08:27 UTC). The eight-minute gap prevents a causal per-request subtraction, but these IDs give both routes for your proxy logs.

Additional control: within the same measurement batch, app, and exact SIN ingress host, the LAX worker-lsh target had a higher median first-byte time than LAX edge-dp in all 22 groups containing both. Each batch used sequential requests from multiple Philippine networks, so this controls the measurement window and ingress host, not client network or the underlying carrier.

The same responses show a higher Fly flyio-debug mrtt value on the slow path: production median 268 ms for worker-lsh (n=21) versus 169 ms for edge-dp (n=47); quiet dev 244 ms (n=20) versus 169 ms (n=35). Median per-request first-byte time minus mrtt was also higher, 280 versus 50 ms in production. Can you confirm what mrtt measures and whether a proxy wait or retransmission beyond that reported round trip accounts for the rest?