I see there has been an incident regarding delayed metrics for about a full day now. There have been several updates indicating that service is recovering and processing a large backlog of data, but I’ve personally not seen any improvement all day. Metrics are anywhere from 30 minutes to an hour delayed (depending on which cluster instance I’m hitting, I assume).
Is there an estimated time to full recovery based on the amount of backlogged data and rate of processing? As an aside, I don’t really care about historical metrics for my use case (scale-by-metrics), so is it possible to just drop backlogged metrics data for my org and skip to the present? (I’m guessing this is probably not something you can do easily on a per-org basis, so the answer is probably “no”, but I thought I’d ask.)
As a small side note… It’s best to use the Questions / Help category when you’re hoping to get a reply from Fly.io themselves. There’s automation set up to create an entry in Fly Support’s formal ticketing system whenever a new post appears in (or is retroactively move into) that section of the forum.
(Otherwise, it has to wait for someone to come browsing through at the end of their day, etc.; some questions even fall through the cracks entirely.)
[Also, there was a brief update to the status-page incident right after you posted, saying “continuing to see gradual improvement”, just in case you haven’t polled it in the interim.]
Hey @tjhorner, I dont have an ETR for your at the moment but our engineers has been working with VictoriaMetrics folks to solve this, the fixes that have been put in place have solved some portion of the backlog issues for particular hosts in our network (about 40%). Which is why you see some inconsistency between where the metrics would’ve originated.
We’re looking at our options and we hope to have this fixed asap, unfortunately I don’t think dropping your particular org’s metrics will give you up to date metrics.
Thanks for the update, appreciate it. For what it’s worth, I’m seeing dramatically better results since my original post. Some of my machines’ metrics seem completely fixed while others are still spotty (they are on different hosts, makes sense).
I’ll keep an eye on the status page and scale manually in the meantime