Onward to the first incidents of August…
| incid. date |
pub. date |
description |
|---|---|---|
| 08/01 | 08/11 | IAD MPGv2 Patroni etcd lease-expiration storm |
| 08/03 | 08/11 | BGP misconfiguration dropped ingress traffic |
| 08/04 | 08/11 | Certificate issuance delays from Sidekiq overload |
This was much appreciated detail, all around!
Aside: Fly.io mentioned deep in a forum thread recently that they had stealthily formed a group in June(-ish) to keep tabs on reliability. The possibility of the long-awaited automated system statistics appearing as a consequence was hinted at shortly afterward.
(There’s long been a rough consensus among both users and Fly.io staff that there’s a gap in the current incidents / uptime-reporting setup. The 2024 post proposed “detailed, unfiltered metrics” published online, without the need for manual intervention or judgment calls each time. This would really fill in the missing span of the spectrum between handcrafted status page updates and the retrospective Infra Log,
…)