PSA: Postmortem(s) rollup

A locking mystery on the following Sunday…

incid.
date
pub.
date
description
04/12 04/20 High edge CPU usage resulting in high latency in ORD

This wasn’t a reprise of the classic 0xffffffffffffffff but maybe something from the depths of SQLite instead.

A bug in the Linux kernel, which is an unusual conclusion for an Infra Log entry…

incid.
date
pub.
date
description
04/14 04/21 WireGuard wg0 one-way host connectivity

It was affecting egress IPs and the Fly Proxy (e.g., its load balancing), among other things.

The anticipated write-up of the widely noticed certificates incident on the subsequent Friday…

incid.
date
pub.
date
description
04/17 04/24 Vault outage broke TLS certificate lookups

Moving to a different storage arrangement was confirmed as being the plan, albeit in the “longer term”.

Further WireGuard wobbles, to start off the new week…

incid.
date
pub.
date
description
04/20 04/27 Duplicate Wireguard Mesh IPs Wreaking Havoc

This time it was a userland bug, however, not in the kernel.


Aside: There was also a status-page-only The status-page incident in Singapore on that day had the same underlying cause. (See @PeterCxy’s comment below for more details.)

Small side note: this was actually the same incident as the one in infra-log. The increased latency was caused by… duplicate wg addresses trashing one of our edges in sin rendering it mostly useless for a while :grimacing:

Ah… That does make sense. (And sin was specifically used as an example in the Log entry, too.)

Thanks for the correction!

500s on the dashboard and with GraphQL, as another Thursday rolled around…

incid.
date
pub.
date
description
04/23 04/30 Extension provider polling overloaded Postgres

It doesn’t sound like it was MPG that was overloaded, but rather an internal database of Fly.io’s own.


Addendum: The second incident on that day (April 23) was written up a bit later, as can be seen below.

On the following Monday, the Postgres storm clouds did move over MPG…

That first one apparently caused 6-hour outages for certain operations.

As the revived Infra Log’s second month drew to a close, several users reported odd breakage in deploys…

incid.
date
pub.
date
description
04/28 05/07 Machines API bug caused fly deploy to create duplicate Machines

Not only were extra Machines created, but existing ones weren’t updated to the new image.

This persisted slightly, half an hour, or so, into the following day (April 29).

A small graphical overview of the previous month, now that it’s complete in the Log…

April 2026
½▪ ½▪ ▪ ½▪ ▪ ½│ ½▪ SYD×2, GraphQL, dashboard, metrics, ORD, NRT
½▪ │ ▪ ─ ORD, SYD, WireGuard, certs
▪ ½▪ ½▪ ½▪ WireGuard, SIN, dashboard, GraphQL, IAD
│ ─ ─ deploys, MPG

See the earlier March grid for a description of the annotations.

The first four days of April (corresponding to the top row) were clear of incidents, which was certainly a nice way to start things off…

The wide red mark on April 17 was the Vault certificates store (again); this is one of the few remaining services from the era of using Raft-based clusters for global metadata/configuration (as I understand it). In the longer term, there are plans for replacing it, and a note in the companion forum thread mentioned the decentralized PetSem as the probable substitute.

The wide red stroke on April 28, eleven days later, was a global failure of deploys, due to the Machines API erroneously returning an empty list when asked about existing Machines. This event slightly straddled midnight (00:00 UTC), which is why there are two bars, two outgoing links, etc.


Aside: Four incidents didn’t make it into the Infra Log, per se. (Possibly just because there was no further commentary that could be added.) In those spots, the cell in the table links either to the real-time status page’s archives or to a post in the present forum thread, depending on what else was in the air that day.

A new month begins in the Infra Log…

incid.
date
pub.
date
description
05/05 05/12 Petsem primary host lost networking (IAD)
05/06 05/14 Machines API hitting failed hosts in SIN

That first one briefly affected attempts to mutate secrets, but did not stop reads (which are distributed).


Addenda: There was also a forum-only incident with FRA networking on the bottom row’s day (May 6). The recent Fresh Produce on NATing outgoing IPv6 may be the de facto postmortem for that one.

In a similar vein, the following date’s (May 7) real-time status page reported relatively brief incidents in BOM and SJC, compiled here for ease of reference in the next summary grid.

The next week, an intriguing interplay between heavy log traffic, hypervisor variants, and Linux kernel upgrades:

Most people don’t have Cloud Hypervisor underlying their own Machines (on Fly.io); that’s only needed for GPUs and Upstash’s backstage servers. Still, the popularity of the Upstash Redis extension resulted in considerable notice in the forum…


Aside: The real-time status page also mentioned a glitch in the Grafana logs on the top row’s day (May 11) as well as a reoccurrence of Redis on May 12.

Aside2: The Oban incident may have extended several hours into the following day (May 13).

Depends what time zone you’re in :wink:

More certificate wobbles, moving into the weekend:

incid.
date
pub.
date
description
05/15 05/24 Bad TLS cert update broke Consul (SSH, OIDC)
05/16 05/25 FRA managed Postgres control-plane outage (MPGv1)

The first row was most noticed for its short yet baleful effect on SSH connection attempts, but apparently it was a problem with Consul fundamentally.

The second was remarked upon even more, due to FRA users’ commendable (and characteristic) vigilance in the forum. The Log’s retrospective account of it describes a sticky problem with Kubernetes.

A considerable crop of entries for the subsequent mid-week…

The second row’s ended with a sentence on a possible future refurb of the infrastructure for logs and metrics. Fly.io has been mentioning wanting to change the underpinnings of those for a year or more, since at least Feb 2025.

In a parallel furrow, it looks like there is news about the storage side, reduced retention windows, etc., coming in the next few days.

A surprising but brief return of certificate tangles:

incid.
date
pub.
date
description
05/27 06/03 App creation timeouts from petsem-certs disk full

The petsem-certs in the entry’s title is the (upcoming) Vault replacement mentioned earlier, although it wasn’t intended to be in a position to cause any real errors yet…

Rounding out the final full week of May…

incid.
date
pub.
date
description
05/28 06/04 DNS cache was broken for CNAME’d domains
05/28 06/04 West coast edge proxies overloaded
05/30 06/06 Deploys blocked by billing error

There are excellent write-ups in all three of these, particularly the second one, for those who have been asking where, apart from the main blog, they might learn more about the platform’s internals, etc.

A small graphical overview of the previous month, now that May 2026 is complete in the Log…

See the earlier March grid for a description of the annotations.

Making the grid itself be clickable started to get a little unwieldy, so, instead, the <details> elements below can be expanded, to get each row’s links.

  • Week of May 03: Grafana, secrets, SIN, FRA, BOM, SJC, certs

    incid.
    date
    sourcessymbolbox description
    05/04s.f.n,
    forum
    Grafana logs
    05/05infra,
    s.f.n
    secrets
    05/06infra,
    s.f.n
    ½▪SIN listing Machines
    05/06forum,
    forum′
    FRA IPv6
    05/07s.f.n½▪BOM Machine creates/updates
    05/07s.f.nSJC networking
    05/08s.f.nLet's Encrypt outage
  • Week of May 10: Redis×2, Grafana, billing, SSH, FRA MPG

    incid.
    date
    sourcessymbolbox description
    05/11infra,
    s.f.n,
    forum
    Redis
    05/11s.f.nGrafana logs
    05/12infra,
    forum,
    forum′
    Redis (again)
    05/12infra½│billing lag
    05/13infra½│billing lag (continued)
    05/15infra,
    s.f.n,
    forum
    SSH & OIDC due to Consul
    05/16infra,
    s.f.n,
    forum
    FRA MPG
  • Week of May 17: SIN×4, IAD, dashboard, SYD, ORD, BOM, SJC

    incid.
    date
    sourcessymbolbox description
    05/19infra,
    s.f.n
    SIN Fly Proxy
    05/19infra,
    s.f.n
    IAD logs & metrics
    05/19infra,
    s.f.n
    dashboard
    05/20infra,
    s.f.n
    dashboard (continued)
    05/20infra,
    s.f.n
    SYD egress IPs
    05/20s.f.nSIN networking
    05/20s.f.n½│SIN high latency (segue)
    05/21s.f.nORD IPv6
    05/21s.f.nSIN networking (again)
    05/22s.f.nSIN IPv6
    05/23s.f.n½▪BOM networking
    05/23s.f.n½▪SJC networking
  • Week of May 24: EWR, app creates, SYD×2, DNS, SJC, LAX, ORD×3, deploys

    incid.
    date
    sourcessymbolbox description
    05/26s.f.n½▪EWR capacity
    05/27infra,
    s.f.n
    app creates
    05/27s.f.n½▪SYD 6PN
    05/28infraDNS cache
    05/28infra,
    s.f.n,
    forum,
    forum′
    SJC edge proxies
    05/28infra,
    s.f.n,
    forum,
    forum′
    LAX edge proxies
    05/29infra,
    forum,
    forum′
    LAX edge proxies (continued)
    05/29s.f.n½▪ORD networking
    05/29s.f.n½▪ORD networking (again)
    05/29s.f.n½▪ORD networking (again, again)
    05/30infra,
    s.f.n
    deploys
    05/30s.f.n½▪SYD 6PN
  • Week of May 31: ORD

    incid.
    date
    sourcessymbolbox description
    05/31s.f.nORD IPv6

(In the expanded tables’ second columns, “s.f.n” is status.flyio.net, the real-time status page, which serves as a second-tier source in this context.)

May followed the (lately) typical pattern of a handful of heftier boxes within an expansive speckling of relatively minor ones. The SVG pipeline that constructed the above enforces a minimum width and height, otherwise some of these would actually barely even be visible. Since there are more pixels to work with overall now, the minimums are roughly half of what they were in earlier renderings.

Of the more memorable cases…

   

The widely used Redis extension went down May 11–12, due to a mismatch between Linux kernel and hypervisor versions. Pathological behavior was triggered by virtualization guest log traffic.


    

The essential secrets and/or certificates features glitched on May 5 and May 27, although each time for only half an hour. These were the PetSem servicecodebase, including its growing pains in expanded roles.

[Edit: see @lillian’s clarification below.]


    

The West Coast proxy overloads on May 28 and a bit of May 29 affected a lot of people, due to the prominence of that part of the world in things generally Internet, but fortunately the durations of those were mainly in the 2 hour range (albeit with after-shocks).

Another new month begins in the Infra Log…

incid.
date
pub.
date
description
06/04 06/11 Stale 6PN mappings wreaking havoc

The underlying bug recounted in this first entry of June was likely also the cause of several mysterious 6PN failures in the past, where users found that .internal glitches could be fixed simply by destroying and then re-creating an unreachable Machine.

worth clarifying here: we run separate instances of petsem for secrets and for certificates, the same codebase but operated by different teams.