About outage

[QUOTE=Stephen]
All IPs are up, at the point Prakash posted many IPs were up and alive, but some 30-40% of IPs were not routing fully.

This was scheduled maintenance gone bad, it was not bad on our part but upstream there were to be some changes with the providers and the changes were, as sour as can be. They were reverted within 20 minutes and the problems continued.
In fact in this whole process no changes were made on our part, it was all on the upstream providers. We were told to expect a 30 second connection blip, that turned into 3 hours for some IP addresses. I am personally unhappy about the situation and will not be allowing such changes in the future without knowing exactly what commands they will be running to evaluate if it will affect us or not.
[/QUOTE]

Stephen, could you please provide more information:

  1. What upstream(s) provider(s) has/ve this problem? Or was it a outstream(internap) as Yash said?

  2. What series of IPs were affected? you said that this was around 3 hours and some of my customers in certain areas noted this while others do not detect anything at all.

  3. What part of the world affected this?

  4. Did this affected other datacenters?

Thanks!

  1. What upstream(s) provider(s) has/ve this problem? Or was it a outstream(internap) as Yash said?

Well there is no direct answer for it there really. It was very odd, in one place jodohost.com would open in another not. I am on Verizon DSL in Texas, jodohost.com opened and was super speedy, everything that is on 64.187.xxx.xxx ip worked for me.
Anything on 204.14.xxx.xxx or 204.10.xxx.xxx did not work for me, that included the forums and therefore why I was unable to post to the people that could see(it seems that the traffic was exactly 55% of what it normally is on a Saturday), so I’d say 55% could access one range, and 55% could access the other.
This is by no means a scientific figure, but going by traffic patterns and what I saw with what was happening. Our India staff actually came for all IPs long before I did, and Prakash gave the update because of that.

It was all caused by route announcement problems with internap not accepting the route, so really all stemmed from that.

  1. What series of IPs were affected? you said that this was around 3 hours and some of my customers in certain areas noted this while others do not detect anything at all.

It was as I clarified above, pretty much split among all IPs, it seems if one ISP could not access oen rang,e it could access the other, and you may call a friend down the street on a different ISP and have them check and it work, while the other IPs didn’t. It was a mess and a disaster for sure, we know that and we will request to approve all commands in the future, as they surely made a mistake in the commands to have such issues.

  1. What part of the world affected this?

I only had people from India and the US available at the time it happened, and we were evenly split on what was accessible. So it seems most of the world was somewhat split probably near evenly.

  1. Did this affected other datacenters?

Not other datacenter, and likley not even other clients in NAP, it was internap changes specifically for our cage that has some other clients in it as well, they were optimizing some routes and well, it simply did not optimize, it de-optimized :slight_smile:

I am trying to stick to the facts and not let my displeasure our upstream with the situation make me show through in this reply, as I am not any any way angry at you asking this question, I have this and many harder ones for them!