Multiple down times tonight to multiple servers

Hello,

I have been working all day and several time in the last hour I have totally lost connectivity to at least 3 servers at Jodo and the Jodo site was down. I think the longest was only a few minutes, but it has happened several times.

I have dont traceroutes and the failure seems to be right up to Jodo and not somewhere in the middle. Has Jodo detected some more problems?

Thanks

We have not detected any problems, but if you are reporting a loss of connectivity I am finding out. Sensors haven’t picked anything up, not has siteuptime or hyperspin

OK, Yogesh, one of our techs got disconnected for a few minutes from one of our servers too, in the last hour. So there may have been a brief loss of connectivity but everything seems absolutely normal now. I am imquiring from the NOC

[QUOTE=Yash]
We have not detected any problems, but if you are reporting a loss of connectivity I am finding out. Sensors haven’t picked anything up, not has siteuptime or hyperspin
[/QUOTE]

I even went to jodopulse.com and it came up with a timeout to mysql3 (i think it was mysql3). Regardless it timed out as well.

Which tells me that jodopulse lost it as well. Don’t know how it would report anything since it couldn’t get ty mysql3.

If I couldn’t get to jodopulse then it should have never given me the error connecting to the database. right?

[QUOTE=Yash]
OK, Yogesh, one of our techs got disconnected for a few minutes from one of our servers too, in the last hour. So there may have been a brief loss of connectivity but everything seems absolutely normal now. I am imquiring from the NOC
[/QUOTE]

Ok… good. The time was a couple of minutes and it happened at least 2 times in the last hour or two.

I am also experiencing some lost packets and right at the last verio hop before jodo server.

The ping rates are floating around 80ms or 90ms at the server where the last hop before the server is around 50ms average.

And now it looks like part of verio is running a little slower than usual.

Anyway… something is up and wanted to make sure someone was on it and checking things out.

Thanks

I experienced approx 5-10 min of downtime from Central IL to JH. Approx 9:23 was the start. Siteuptime reported to me as well…

That was in the morning. We experienced a brief DDOS attack. That issue was resolved. gsaunders is reporting some fresh outages in the last hour (not mor than 2 minutes each). We are investigating that.

We talked to the NOC and they said they experienced one brief 2minute outage. They are unaware of a second outage, we haven’t noticed a second outage either. However they on the phone with InterNap enquiring about the same

[09:30] Scott Mueller: internap is telling me
[09:31] Scott Mueller: that the circuit that handles you guys and just a few other smaller customers jumped more than 432mbps 2 hours ago for about 10 minutes inbound
[09:31] Scott Mueller: jumped 390mbps about 30 mins ago for about 10mins also inbound
[09:31] Scott Mueller: they are saying these clearly look like massive but very short dos attacks
[09:32] Scott Mueller: it’s so hard to figure these things out when they’re so short unfortunately.
[09:33] Scott Mueller: investigating… need to get rid of this problem asap

No, it was tonight…

This is an automated message from SiteUptime.

Alert Type: Site Not Available
Result: Failed
Time: November 27, 2005 19:23:05 PST
HostName/URL: illianaoncall.com
Monitor Name: Illiana On Call
Service: http

Alert Type: Site is Available
Result: Ok
Time: November 27, 2005 19:53:06 PST
HostName/URL: illianaoncall.com
Monitor Name: Illiana On Call
Service: http

We had issues through the weekend as well…

This is an automated message from SiteUptime.

Alert Type: Site Not Available
Result: Failed
Time: November 26, 2005 10:49:07 PST
HostName/URL: illianaoncall.com
Monitor Name: Illiana On Call
Service: http

This is an automated message from SiteUptime.

Alert Type: Site is Available
Result: Ok
Time: November 26, 2005 11:19:09 PST
HostName/URL: illianaoncall.com
Monitor Name: Illiana On Call
Service: http

I suggest you open a ticket about this. Your issue seems to be specific with your site, not the network.

We had some very brief outages tonight, for less than 2 minutes. No where close to 30min

OK, InterNap says that the reason for the very short outages tonight were a short but massive DDOS attaack they are trying to track down. They limit our bandwidth to 400mbps to protect us against DDOS attacks (4x our average usage) but today this massive attack was hitting us right upto that limit causing those brief outages. We are working to completely eliminate the DDOS attacks and customers shouldn’t experience this again

I hope this gets resolved soon as I am VERY concerned. This is twice now in the last 2 days and I already have upset clients.

Again what concerns me is it seems to be a challenge to get Jodohost to admit to these problems when we report them. We need to have a bit more proactive response to these issues other than “it must be your ISP, everything is fine here” as now twice it turned out not to be fine.

This DDOS attack issue is already resolved. We have not had any fresh attacks since..

There were a total of 3 short DDOS attacks yesterday, from morning to evening. The first attack bringing us down for 10 minutes in the morning, the attacks in the evening brought us down for 2 minutes each.

schooner, the issue you are facing is specific to your ISP and I have already advised you to contact your ISP. It is beyond our control

[QUOTE=schooner]

We need to have a bit more proactive response to these issues other than “it must be your ISP, everything is fine here” as now twice it turned out not to be fine.
[/QUOTE]

We rely on a number of sensors and reporting tools to monitor our network. Sometimes we do not pick up an issue immediately and our customers pick it up first. Sometimes an issue a customer is noticing is not being experienced by us. Hence we cannot always give a proactive response or immediate diagnosis

Yash - I am not trying to be difficult but my issue is that both problems yesterday were reported by customers. When the issue was raised I and the other person I beleive was told that everything was fine on your end but it turned out not to be. It just concerns me when I get the old “its all fine here” response as it seems to be a typical ISP/hoster response to issues.

I realize that some of my issues are ISP related, however the first issue I reported was not is all and it was not detected by Jodohost until I pointed it out down to the node in question.

I just hope things are now corrected and we can move forward with solid hosting.

As for the issue being beyond your control, jsut remeber this imapcts more than me. Any site hosted on Jodohost using this same route will have issues, and it is not my ISP, it is a major carrier in the middle that seems to be the issue.

schooner, the first DDOS outage in the morning was immediately picked up by our sensors and our most senior tech on shift was already on the phone with the NOC working out what was going on. I came to know about the issue over a phone call and I immediately came to the forum and answered the thread you created. Sometimes, a senior admin may be working on an issue that the support team hasn’t yet been informed about. If you were informed over chat that nothing was wrong, it only means that the support team isn’t aware of an issue that may be under investigation by senior techs on shift at that time. You can of course let them know of the problem or wait for more information to be available.

During the second and 3rd attack in the evening, I was not aware of the issue and hence I incorrectly informed you that it could be an ISP issue. But Yogesh, the senior admin on shift then had already picked up the issue when he was disconnected from servers and was trying to look into it. It wasn’t immediately apparent to us that it was a DDOS attack or it was a packet loss issue or a uniform network slowdown. I was informed a short while later that there was an issue and I immediately corrected my post.

Also, you don’t realise how difficult these DDOS attacks can be. I’ve seen hosting companies going down for a day or more due to similar attacks we face on a regular basis. We have a superb network team at the NOC that takes care of these issues very quickly. and we are paying a premium tier-1 ISP - InterNap which provides us excellent support in these issues. Many hosts have placed limits on their network ports so when a large 400mbps or 1000mbps attack comes in, they are completely overwhelmed and their servers cannot respond. However, for a good portion of this large 300 to 400mbpish attack, our network continued to respond because we had excess bandwisth available… something we are paying for.. We use excellent cisco equipment that allows network techs to use netflow to track down the source of the attack and what they are attacking. This is expensive equipment and the nature of the attack involved hundreds of source IPs

Ok thanks Yash, I seem to have hit a nerver with you and you are on the defensive about it. All I know is that the sites were down, I contacted support and they said it was fine, then after 10 mins or more finally said there was an issue and they were looking into it. At that same time I had clients emailing me that were not happy.

Consider the case closed. I’ll base what I do on the performance of the hosting at this point, I don’t want to argue about this further as it seems to be going no where. Just remeber who the customer is, and don’t try and make excuses for issues regardless of what the problem is.

[QUOTE=schooner]
Ok thanks Yash, I seem to have hit a nerver with you and you are on the defensive about it. All I know is that the sites were down, I contacted support and they said it was fine, then after 10 mins or more finally said there was an issue and they were looking into it. At that same time I had clients emailing me that were not happy.

Consider the case closed. I’ll base what I do on the performance of the hosting at this point, I don’t want to argue about this further as it seems to be going no where. Just remeber who the customer is, and don’t try and make excuses for issues regardless of what the problem is.
[/QUOTE]

Yash, I think all that is being said here is that when we call of an issue and we take steps on our end to isolate it down to your NOC or your network that we would prefer not to here it must by your ISP or we don’t have a problem. Just because it hasn’t been detected yet doesn’t mean it isn’t happening.

Perhaps a better response would be… “we have not detected anything at the moment, so I need you to do the following steps to provide any traces or ping statistics so we can better isolate what may be going on while Jodo checks things”

I know you probably get many calls where someone just freaks out and says I can’t get to Jodo and it really is due to their ISP or something else out of your control, but I think having some set tools you can point folks to run to help isolate the problem would be helpful. I have all the tools I need if I am having the issue or somewhere down the line, but others may not.

Anyway, just wanted to shed some more light from a customer.

Thanks for your hard work!

I have always said it my responses that we haven’t noticed any problems and please double-check if there is an ISP issue at your end. If we become aware of an issue that is affecting the network, we ALWAYS update the customer with correct information.

schooner submitted a ticket after the issue was fixed, but still being investigatd. The support team wasn’t updated yet. Supported tested the site, and it was obvious there was no issue. The support team was shortly later informed about the issue, and they accordingly started updating tickets and i updated the forum.

We have always updated customers in this manner. There are instances where we are not aware of an issue immediately and we use our experience and what we know at that time to give a response. Any hosting company does that. If you are looking for a correct update immediately, we unfortunately cannot provide that and would not be able to do so in the future either.