Oh man. One of my old companies, the Devs would always blame the network. Even after we spent a year upgrading and removing all SPOFs. They’d blame the network…..
“Your application is somehow producing 2 billion packets per second and your SQL queries are returning 5GB of data”…. “See! The network is too slow and it has problems”
I always view the source of websites like this and this is one of the worst I've seen. 217 lines of code (including inline Javascript?!) and a Google tag for some reason, all to put the word YES in green on black.
These things happen when a skinflint company contracts out network setup for a decade, gets acquired by another skinflint company who axes the contractors and doesn't hire on-site network personnel, gradually builds out infra on top of the unsupported foundation, and then hires c suite buddies who want to bring in their own people to further muddy the waters.
I KNEW IT. It feels good to have my suspicions validated like this. The biggest companies are the ones most hyped over useless AI, and it's going to destroy them.
Much of this stuff is automatic - I've worked with such contracted services where uptime is guaranteed. The contracts dictate the terms and conditions for refunds, we see them on a monthly basis when uptime is missed and it's not done by a person.
I imagine many companies have already seen refunds for outage time, and Amazon scrambled to stop the automation around this.
They'll have little to stand on in court for something this visible and extensive, and could easily lose their shirt with fines and penalties when a big company sues over breech when they choose to not renew.
Just cause they're big doesn't mean all their clients are small or don't have legal teams of their own.
These contracts do not stipulate reimbursement for lost revenue. The “uptime guarantee” just gets you a partial discount or service refund for the impacted services.
It is on the customer to architect their environment for high availability (use multiple regions or even multiple hyperscalers, depending on the uptime need).
Source: I work at an enterprise that is bound by one of these agreements (although not with AWS).
SLA contracts can have a plethora of stipulations, including fines and damages for missing SLO. It really depends on how big and important the customer is. For example, you can imagine government contracts probably include hefty fines for causing downtime or data loss, although I am not involved with or familiar with public sector/ government contracts or their terms.
You can imagine that a customer that is big enough to contract a cloud provider to build new locations and install a bunch of new hardware just for them, would also be big enough to leverage contract terms that include fines and compensation for extended downtime or missing SLO.
I work at a data center for a major cloud provider, also not AWS
Depends on who we're talking about. Companies like finance orgs are all about legal contracts and would be able to hold their feet to the fire.
You don't want to go to court against a finance company or any very large org where contract law is their bread and butter (basically any large/multinational corp).
Good luck arguing that a missed config counts as an 'unforeseen issue'. If they go that route, people will be all over them for not being SOC compliant wrt change control.
They can try to argue that latency issue and the stale state were an unknown / unanticipated problem. Like when half of Canadas Rogers network went down affecting most debit payment systems. Testing of routing showed it OK, realworld flip went haywire.
99% uptime in a year gives you 3.65 days of downtime, which I think would still be within SLA (assuming nothing else happened this year). Though, once you get to 1 9 reliability (99.9%), you've got a shift and change you can be down before you breach SLA.
If their reliability metrics are monthly, 99% gets you less than a shift of down time, so they'd be out of SLA and could probably yell to get money back.
Oh yea, other companies will sue them, and when amazon completely fails they will be bailed out with consumers' tax money. Or did we already forget that's what happens?
The problem is that the current internet is structured in a way that creates high risk systems that can cause a massive outage. We went from having thousands of independent companies to a handful of massive ones. A mistake by a single company shouldn't be able to black out half the internet.
There was never any evidence to even suggest that AI was the cause, but as you're on lemmy I'm sure you know that AI is currently blamed for pretty much everything.
They rely on AWS due to favourable contract in hosting it, and also proving the proof of concept that they can be hosted securely on a hostile provider, without the provider having any clues at all in what data is being sent between the parties.
sure, proving to the audience that you can kick yourself in the nuts over and over while maintaining the privacy of your testicle's innards is impressive from a biological standpoint but it still looks stupid to a normal person. I don't hate signal, I will continue using it but this and their crypto scam makes me doubt some of their choices and how they'll operate in the future
This is purely anecdotal, but I have been running into a lot of DNS issues over the past couple months where I work. 3 of the computers and even one of the laptops for remote work were having DNS issues that needed to be fixed. One even needed Windows reinstalled after fixing the DNS issue (Which was probably unrelated, but worth mentioning)
I'm honestly starting to think that the internet in general might be imploding. Not sure why, but replacing so many developers and programmers with AI might be responsible. Who knows, but it's definitely very strange.
But but Bezos has to pay for another rocket and yacht and he just got married!!!! Think about his quarterly statement! My god are you heartless!!!!!!!!
A huge problem are developers who lack a fundamental understanding of how the internet even works. I've had to explain how short, unqualified names resolve vs how fqdns resolve. Or why even you may not be able to reach another node in your proverbial cluster, because they are on different subnets. Or, why using GUIDs as hostnames is a generally bad idea, and will cause things to fail in unpredictable ways, especially with deeply nested subdomains.
Why the fuck would anyone use a guid as a hostname?
My favorite I've seen in the category was when they had hostnames that were basically the IP address decorated with some bullshit. Like yeeeeeeeeah, that totally makes fucking sense. 😆
I'm glad these things happen... it keeps everyone aware that cloud is fragile and Plan B should be considered for mission critical tasks.
I'm also hoping that it will improve cloud resiliency because a complete / partial restart of cloud systems needs a whole different approach than maintaining a running system.
The issue here isn't DNS. The issue here is a large portion of the internet relying on a single data centre on the US East coast. Ideally, a lot of competing hosting companies would exist so if one goes down, it's just one service and very few people notice.
Why is Signal hosted in one location on AWS, for example? That's the sort of thing that should be in multiple places around the world with automatic fail over.
112 Comments
IsoKiero@sopuli.xyz · 204 pts · 331d
So it is always DNS
mhzawadi@lemmy.horwood.cloud · 123 pts · 331d
can confirm, its always DNS. Even when it looks like a network issue, its DNS
aarRJaay@lemmy.world · 28 pts · 331d
Spotted the Network guy
ramble81@lemmy.zip · 29 pts · 331d
Oh man. One of my old companies, the Devs would always blame the network. Even after we spent a year upgrading and removing all SPOFs. They’d blame the network…..
“Your application is somehow producing 2 billion packets per second and your SQL queries are returning 5GB of data”…. “See! The network is too slow and it has problems”
rumba@lemmy.zip · 17 pts · 331d
Dev: My app's getting a 400 hitting the server. Your firewall changes broke it.
Me: You're getting to the server, it's giving you back a malformed request error. Most likely it's a problem in your client.
Dev: it worked fine until you made that change in QA.
Me: Your server is in production.
After that, I just get too busy to look at it for a while.... They figure it out eventually.
fushuan@lemmy.blahaj.zone · 4 pts · 331d
They might be referring to their brain network being to slow and having problems.
lka1988@lemmy.dbzer0.com · 1 pts · 330d
Ah, klugerblickdummkopf
ijhoo@lemmy.ml · 41 pts · 331d
https://isitdns.com/
NickwithaC@lemmy.world · 75 pts · 331d
I always view the source of websites like this and this is one of the worst I've seen. 217 lines of code (including inline Javascript?!) and a Google tag for some reason, all to put the word YES in green on black.
Cyber@feddit.uk · 30 pts · 331d
Agreed, could be static HTML and a GIF.
Thanks, I won't click that link.
Xylight@lemdro.id · 21 pts · 331d
this made me mad so i made a single, ultra minimal html page in 5 minutes that you can just paste in your url box
source code:
aBundleOfFerrets@sh.itjust.works · 2 pts · 330d
Your website no longer uses DNS invalidating its use as a diagnostic tool lmao
Xylight@lemdro.id · 1 pts · 329d
ijhoo@lemmy.ml · 12 pts · 331d
Did not think of doing that.
I guess i never expected anyone to have a fcking JavaScript on a simple page as that
Randelung@lemmy.world · 5 pts · 331d
How else would you center a div??
NickwithaC@lemmy.world · 4 pts · 330d
http://howtocenterincss.com/
rumba@lemmy.zip · 5 pts · 331d
I just did the same f'ing thing and came here to write your comment!
well done.
hexagonwin@lemmy.sdf.org · 3 pts · 331d
lmao, considering some of the meaningless comments there i'm starting to think it's "vibe coded".
rumba@lemmy.zip · 4 pts · 331d
There have been 209 versions of that site
https://web.archive.org/web/20250331043558/https://www.isitdns.com/
it predated AI, but likely seems to have had some AI cleanup.
If it was truly just vibecoded, the comments would usually be on every element.
floofloof@lemmy.ca · 1 pts · 331d
Dubiousx99@lemmy.world · 33 pts · 331d
It’s always DNS
AtariDump@lemmy.world · 18 pts · 330d
magic_lobster_party@fedia.io · 200 pts · 331d
It’s not DNS
There’s no way it’s DNS
It was DNS
evidences@lemmy.world · 137 pts · 331d
possiblylinux127@lemmy.zip · 35 pts · 331d
That and BGP
MelodiousFunk@slrpnk.net · 27 pts · 331d
If I had a nickel for every time clearing the ARP tables fixed a problem, I'd have a shitload of nickels.
possiblylinux127@lemmy.zip · 18 pts · 331d
If clearing the ARP tables fixes the issue you have bigger problems
MelodiousFunk@slrpnk.net · 30 pts · 331d
These things happen when a skinflint company contracts out network setup for a decade, gets acquired by another skinflint company who axes the contractors and doesn't hire on-site network personnel, gradually builds out infra on top of the unsupported foundation, and then hires c suite buddies who want to bring in their own people to further muddy the waters.
sleepmode@lemmy.world · 3 pts · 330d
Like every MSP ever. When your CEO that started the company in college suddenly shows up in a green Lamborghini it is time to spruce up the resume.
the_q@lemmy.zip · 30 pts · 331d
GreenKnight23@lemmy.world · 128 pts · 331d
oh sure, when they fuck up DNS it's a "race condition".
when I fuck up DNS it's a "fireable offense".
sommerset@thelemmy.club · 13 pts · 331d
It's funny aws report didn't mention 40% sysops were replaced by AI. https://blog.stackademic.com/aws-just-fired-40-of-its-devops-team-then-let-ai-take-their-jobs-d9db9d298bfa
StopSpazzing@lemmy.world · 2 pts · 330d
Wasnt that source from a year ago?
sommerset@thelemmy.club · 0 pts · 330d
No
StopSpazzing@lemmy.world · 5 pts · 330d
You are right, was from july and there was no other confirmed layouts from credible sources since.
sommerset@thelemmy.club · 2 pts · 329d
Do you mean are you saying that you believe America has fair and open media that would publish some of this again bezos?
ZILtoid1991@lemmy.world · 1 pts · 330d
They need to uphold the AI hype, at any cost possible.
finitebanjo@lemmy.world · 1 pts · 330d
I KNEW IT. It feels good to have my suspicions validated like this. The biggest companies are the ones most hyped over useless AI, and it's going to destroy them.
falseWhite@lemmy.world · 102 pts · 331d
otacon239@lemmy.world · 77 pts · 331d
Consequences? For Amazon?
lol… lmao even
falseWhite@lemmy.world · 37 pts · 331d
Onomatopoeia@lemmy.cafe · 27 pts · 331d
Much of this stuff is automatic - I've worked with such contracted services where uptime is guaranteed. The contracts dictate the terms and conditions for refunds, we see them on a monthly basis when uptime is missed and it's not done by a person.
I imagine many companies have already seen refunds for outage time, and Amazon scrambled to stop the automation around this.
They'll have little to stand on in court for something this visible and extensive, and could easily lose their shirt with fines and penalties when a big company sues over breech when they choose to not renew.
Just cause they're big doesn't mean all their clients are small or don't have legal teams of their own.
WASTECH@lemmy.world · 8 pts · 331d
These contracts do not stipulate reimbursement for lost revenue. The “uptime guarantee” just gets you a partial discount or service refund for the impacted services.
It is on the customer to architect their environment for high availability (use multiple regions or even multiple hyperscalers, depending on the uptime need).
Source: I work at an enterprise that is bound by one of these agreements (although not with AWS).
CheezyWeezle@lemmy.world · 9 pts · 331d
SLA contracts can have a plethora of stipulations, including fines and damages for missing SLO. It really depends on how big and important the customer is. For example, you can imagine government contracts probably include hefty fines for causing downtime or data loss, although I am not involved with or familiar with public sector/ government contracts or their terms.
You can imagine that a customer that is big enough to contract a cloud provider to build new locations and install a bunch of new hardware just for them, would also be big enough to leverage contract terms that include fines and compensation for extended downtime or missing SLO.
I work at a data center for a major cloud provider, also not AWS
village604@adultswim.fan · 4 pts · 331d
It's not at all uncommon for fines to be built into an SLA
BakerBagel@midwest.social · 7 pts · 331d
Amazon has more money than most countries. They can outlast any company in court, or just ban you from their services in the future.
Onomatopoeia@lemmy.cafe · 11 pts · 331d
Depends on who we're talking about. Companies like finance orgs are all about legal contracts and would be able to hold their feet to the fire.
You don't want to go to court against a finance company or any very large org where contract law is their bread and butter (basically any large/multinational corp).
Amazon's not hosting just small operations.
fushuan@lemmy.blahaj.zone · 3 pts · 331d
Most banks have their data on Amazon/Azure. You don't want to enrage banks.
BCsven@lemmy.ca · 6 pts · 331d
Most services have a clause that they are not liable for unforseen issues.. Depends how good the lawyers were when formalizing the contracts.
Passerby6497@lemmy.world · 4 pts · 331d
Good luck arguing that a missed config counts as an 'unforeseen issue'. If they go that route, people will be all over them for not being SOC compliant wrt change control.
BCsven@lemmy.ca · 1 pts · 331d
They can try to argue that latency issue and the stale state were an unknown / unanticipated problem. Like when half of Canadas Rogers network went down affecting most debit payment systems. Testing of routing showed it OK, realworld flip went haywire.
Passerby6497@lemmy.world · 4 pts · 331d
99% uptime in a year gives you 3.65 days of downtime, which I think would still be within SLA (assuming nothing else happened this year). Though, once you get to 1 9 reliability (99.9%), you've got a shift and change you can be down before you breach SLA.
If their reliability metrics are monthly, 99% gets you less than a shift of down time, so they'd be out of SLA and could probably yell to get money back.
phoenixz@lemmy.ca · 10 pts · 331d
I worked at a datacenter that sold clients 99.99% uptime.
Fun times with a maximum of about one hour of downtime per year for hundreds of servers
87Six@lemmy.zip · 3 pts · 330d
Oh yea, other companies will sue them, and when amazon completely fails they will be bailed out with consumers' tax money. Or did we already forget that's what happens?
SeeMarkFly@lemmy.ml · 7 pts · 331d
They have ORANGE ass makeup on their lips. How did THAT get there???
bigboitricky@lemmy.world · 18 pts · 331d
Oops! All slop!
possiblylinux127@lemmy.zip · 17 pts · 331d
Mistakes happen with or without AI
The problem is that the current internet is structured in a way that creates high risk systems that can cause a massive outage. We went from having thousands of independent companies to a handful of massive ones. A mistake by a single company shouldn't be able to black out half the internet.
phoenixz@lemmy.ca · 13 pts · 331d
Was it proven that AI wa the cause?
In not saying it wasn't, just that if it really was, I'd like a source for that claim
jaybone@lemmy.zip · 7 pts · 331d
There was an article in my lemmy all feed yesterday claiming so. But it was a super questionable shady site, which people were calling out.
Serinus@lemmy.world · -2 pts · 331d
No, but it clearly wasn't the solution. They likely could have used some of those people they fired for that.
FreedomAdvocate@lemmy.net.au · -3 pts · 331d
There was never any evidence to even suggest that AI was the cause, but as you're on lemmy I'm sure you know that AI is currently blamed for pretty much everything.
phoenixz@lemmy.ca · 1 pts · 330d
Just because this may NOT have been caused by AI doesn't mean that AI in 99% of places isn't absolute horse shit
FreedomAdvocate@lemmy.net.au · 1 pts · 330d
Saying it was caused by AI despite zero evidence of AI causing it is dumb. It wasn’t AI, it was a DNS change made by a person.
The whole thing has nothing to do with AI, other than people who hate AI trying to make it about AI.
Auli@lemmy.ca · 4 pts · 331d
Silly peon rich people don't suffer consequences.
Zwuzelmaus@feddit.org · 2 pts · 331d
OK but then... what happens when their boss jerk fires hundreds of thousands?
https://lemmy.ca/post/53821900
joeldebruijn@lemmy.ml · 80 pts · 331d
Laser@feddit.org · 35 pts · 331d
Luckily, it's not the entire Internet, just the unfun part.
amino@lemmy.blahaj.zone · 11 pts · 331d
Signal is definitely part of the fun internet, they just decided to rely on AWS due to techbro culture I assume?
dubyakay@lemmy.ca · 4 pts · 331d
They rely on AWS due to favourable contract in hosting it, and also proving the proof of concept that they can be hosted securely on a hostile provider, without the provider having any clues at all in what data is being sent between the parties.
amino@lemmy.blahaj.zone · 5 pts · 331d
sure, proving to the audience that you can kick yourself in the nuts over and over while maintaining the privacy of your testicle's innards is impressive from a biological standpoint but it still looks stupid to a normal person. I don't hate signal, I will continue using it but this and their crypto scam makes me doubt some of their choices and how they'll operate in the future
dubyakay@lemmy.ca · 1 pts · 331d
Huh? What crypto scam?
amino@lemmy.blahaj.zone · 5 pts · 331d
Moxie works for MobileCoin and implemented a crypto wallet for it inside of Signal which could be an attempt to sneakily monetize the Signal userbase
amino@lemmy.blahaj.zone · 3 pts · 331d
after looking further into it, he might've also been involved in the MobileCoin pump and dump
slothrop@lemmy.ca · 69 pts · 331d
I DNS see that coming.
regedit@lemmy.zip · 48 pts · 330d
Unbelievable, racism even exists in networking!
StopSpazzing@lemmy.world · 5 pts · 330d
Beat me to it!
Zron@lemmy.world · 2 pts · 330d
Those damn ones
WhatsHerBucket@lemmy.world · 47 pts · 331d
It was the best race anyone has ever seen 🫲🍊🫱
BrianTheeBiscuiteer@lemmy.world · 10 pts · 331d
Let's be honest, not all races are equal 🫲🍊🫱
Kolanaki@pawb.social · 3 pts · 331d
Worst Race: Daytona 500.
Best Race: Kentucky Derby.
sommerset@thelemmy.club · 46 pts · 331d
It's funny aws report didn't mention 40% of aws sysops people were replaced by AI right prior https://blog.stackademic.com/aws-just-fired-40-of-its-devops-team-then-let-ai-take-their-jobs-d9db9d298bfa
aBundleOfFerrets@sh.itjust.works · 1 pts · 330d
this is unconfirmed and unlikely
sommerset@thelemmy.club · 3 pts · 330d
"Leave a billion dollar company alone, leave it alone" Bro it's most likeliest thing ever
pokexpert30@jlai.lu · 37 pts · 330d
Just one more layer bro, just one more automated planning system bro and this time it will be entirely faultless please bro one more layer
HurlingDurling@lemmy.world · 5 pts · 330d
I know a dude that talks like this... Like I hear his voice when I read this.
TommySoda@lemmy.world · 33 pts · 331d
This is purely anecdotal, but I have been running into a lot of DNS issues over the past couple months where I work. 3 of the computers and even one of the laptops for remote work were having DNS issues that needed to be fixed. One even needed Windows reinstalled after fixing the DNS issue (Which was probably unrelated, but worth mentioning)
I'm honestly starting to think that the internet in general might be imploding. Not sure why, but replacing so many developers and programmers with AI might be responsible. Who knows, but it's definitely very strange.
possiblylinux127@lemmy.zip · 58 pts · 331d
The biggest issue is how centralized the internet has become. It went from a bunch of local servers to a handful of cloud providers.
We need to spread things out again
metaStatic@kbin.earth · 7 pts · 331d
That's not how capitalism works though
Canopyflyer@lemmy.world · 1 pts · 330d
But but Bezos has to pay for another rocket and yacht and he just got married!!!! Think about his quarterly statement! My god are you heartless!!!!!!!!
/s
(just in case it's not obvious)
ubergeek@lemmy.today · 27 pts · 331d
A huge problem are developers who lack a fundamental understanding of how the internet even works. I've had to explain how short, unqualified names resolve vs how fqdns resolve. Or why even you may not be able to reach another node in your proverbial cluster, because they are on different subnets. Or, why using GUIDs as hostnames is a generally bad idea, and will cause things to fail in unpredictable ways, especially with deeply nested subdomains.
GreenKnight23@lemmy.world · 13 pts · 331d
I have worked with too many devs that didn't even know what the 7 layers/OSI are or why they exist.
they didn't know what a network port was used for and why it's important to not expose 3306 to the internet.
they couldn't understand that fragmentation of a message bus occurs when you don't dedupe the contents.
you know, morons.
metaStatic@kbin.earth · 14 pts · 331d
Ah, the common clay of the new Web
Appoxo@lemmy.dbzer0.com · 5 pts · 331d
GUIDs?
Could you expand on that topic? :)
ubergeek@lemmy.today · 2 pts · 331d
guids like these: https://guidgenerator.com/
aesthelete@lemmy.world · 3 pts · 331d
Why the fuck would anyone use a guid as a hostname?
My favorite I've seen in the category was when they had hostnames that were basically the IP address decorated with some bullshit. Like yeeeeeeeeah, that totally makes fucking sense. 😆
Appoxo@lemmy.dbzer0.com · 2 pts · 331d
I've seen those with public routing servers.
Example: IP-127.0.0.1.dtag.de
Makes sense there or for webservers.
But anywhere else? Lol not really
Appoxo@lemmy.dbzer0.com · 2 pts · 331d
Why would someone want that as their hostname???
I'd understand mountpoint but that?
ReedReads@lemmy.zip · 33 pts · 331d
Ironically, my pihole is blocking that link. So here’s a clean one: https://www.theregister.com/2025/10/23/amazon_outage_postmortem/
HeartyOfGlass@piefed.social · 26 pts · 331d
Racist DNS!
Cyber@feddit.uk · 14 pts · 331d
I'm glad these things happen... it keeps everyone aware that cloud is fragile and Plan B should be considered for mission critical tasks.
I'm also hoping that it will improve cloud resiliency because a complete / partial restart of cloud systems needs a whole different approach than maintaining a running system.
possiblylinux127@lemmy.zip · 5 pts · 331d
Many different companies abruptly realized they need a DR plan for cloud outages
Flax_vert@feddit.uk · 11 pts · 331d
Makes sense. DNS is quite a single point of failure
possiblylinux127@lemmy.zip · 5 pts · 331d
It is designed to not to be. The RFC literally warns against single points of failure
non_burglar@lemmy.world · 2 pts · 331d
Its true.
It comes up at work, it comes up in discussions on Linux podcasts I listen to, it comes up here...
We have a big, dangerous impending problem in DNS.
Flax_vert@feddit.uk · 10 pts · 331d
The issue here isn't DNS. The issue here is a large portion of the internet relying on a single data centre on the US East coast. Ideally, a lot of competing hosting companies would exist so if one goes down, it's just one service and very few people notice.
Onomatopoeia@lemmy.cafe · 7 pts · 331d
So much this.
Why is Signal hosted in one location on AWS, for example? That's the sort of thing that should be in multiple places around the world with automatic fail over.
Flax_vert@feddit.uk · 2 pts · 331d
I prefer end to end encrypted xmpp
chisel@piefed.social · 4 pts · 331d
I prefer face to face communication, speaking in code and whispering in eachother's ears so nobody else can hear.
aBundleOfFerrets@sh.itjust.works · 1 pts · 330d
Get a little tongue in there, maybe.
victorz@lemmy.world · 2 pts · 331d
I hope they work towards mitigating this risk from now on.
non_burglar@lemmy.world · 1 pts · 331d
Yes, that's true, I guess it's a separate issue. But the way DNS currently runs is a problem waiting to happen.
theoriginalcows@lemmings.world · 8 pts · 330d
popcornpizza@lemmy.blahaj.zone · 5 pts · 331d
oeuf@slrpnk.net · 4 pts · 331d
They should check out YUNOhost.
kossa@feddit.org · 2 pts · 328d
Yeah, I don't get why they don't just put a RasPi in some corner, put PiHole on it and call it a day.
Geez, I mean, they could even charge extra for it, as they now block ads for their customers as well.
Like, imma gonna sell my advice to Amazon now, so they can clean up their act.
MadMadBunny@lemmy.ca · 1 pts · 331d
They got off sync.