Web 14 and 15 down

It was good to see Web15 again! Thanks for your hard work!

When you get a chance (after some rest!) can you provide some info on the future of the hardware housing web14/15?

Based on how you had to fix the problems today, do you foresee or plan to have some down time to do any more maintenance?

After today’s problems it would be REALLY nice if we could plan ahead, notify people…whatever we can do to avoid surprise down time!

No, no more maintenance on this one. It is settled.

It isn’t the hardware I WANTED, but it is fine (hotswap chassis not possible with what I had to get)

I will explain a bit more here:

The server, was acting funny on web14, I was on way in about 10 minutes (9:15am I was planning to come here), but I was trying to finish my shopping (online) for new servers.

at 9:05 about this is with web14, then at about 9:10 it goes to being web15 also acting up. We try to issue reboot remote, and it hangs up.

I get here and hard reboot it (power button hold down), it comes up…all seems good for 15 minutes, BAM kernel panic.

this happens about 2 more times, before I say, this isn’t just software, but something hardware. So I pull out the almost identical hardware (2.66 ghz instead of 2.83 Ghz), but it would have worked in a pinch), it seems to work well, till I put one RAID card in to interface the hotswap chassis. Anyway, it actually was up about 35-45 minutes (sorry timing is HAYWIRE due to the day)

Then we bring network up, and it died, basically exactly the same as the first one, even with a totally different RAID card (even larger model)

I try to get it to limp along even, with Tanmaya’s help to buy a little time for a swap. no dice. I decide it is time to hit the local tech stores and get some parts. I come back with a quad core AMD Phenom II, quality mobo and some RAM, not what I wanted, but it will get us going well.

And it is working well now, the server is not having hotswap anymore, which is a bit of a pain if a drive dies we have to shut it down, but we’ve had this before even in rackmounts without hotswap, it happens.

We’ve got quite a few server towers here in the new facility, the flexibility for putting 8-12 hotswaps in a machine with 16-24 CPU cores and 64+GB of RAM is just phenomenal. We’ll be using some of these for some mesh storage solutions…in the future :wink:

A big thank you for all your hard work!

and you got two of everything, right? :slight_smile:

Only had one, I wanted 2+, but the good thing is with what we have done in a ‘semi cloud’ is that we can portabliy move it between brands, hardware, storage platforms etc quite well.

my shortfall here today was that we just had a hardware lapse due to another recent problem. I’ll be working to ensure that even won’t happen again!

Thanks for getting things up and running, and thank you very much for the detailed reply!

You might appreciate this client call this morning (6:00am local time…)

Client: I saw they got things up and running last night.
Me: Yep. (I need more coffee)
Client: Well…it’s been down now for about 25 minutes.
Me: WHAT?!?!
Client: Just kidding, it’s all fine!

The client is also a friend - with a bad sense of humor. I should probably fire him as a client…

On the positive side, I don’t need any more coffee this morning.

:smiley: Scared me for a sec, but really I knew if some issue more than 5 minutes I’d have been awoken (I got a good night sleep, finally cooling off outside and really nice open window weather!)

The server seems to be running really well, low loads and many ways so far, better than before(and I mean at best of times) so running well, so far :slight_smile: