failing drives, raid sets and downtime

Hi Jodohost staff,
I see a lot of expansion going on at Jodohost, only this year already a number of new servers are taken into production.

Wouldn’t it be a great idea to move a big step forward regarding disk-storage?

My advice would be to stop using SATA/SCSI/SAS drives behind all types of RAID controllers which do not provide any data-integrity guarantees, and move forward to centralized storage.

Nowadays iSCSI is a commonly used infrastructure for datacenters which can be very cost-effectively.

And next to the advantage of loosing all kind of different non-reliable RAID-controllers and other DAS (direct attached storage) systems there are some extra reasons why Jodohost should consider the investment:

  1. centralized storage gives much more functionality to prevent tape-restores and very costly diskdrive-reconstructs. If using snapshot-technology you can allways go back in time without a single tape!
  2. it will finally end up with a lower TCO (total cost of ownership)
  3. it’s much more easier to provide disaster-recovery services with centralized storage, for example you can offer hosting from 2 physical locations with less or no downtime.

I look forward to the next step Jodohost will take on this subject!

Regards,
Jan-Pieter
The Netherlands

jpveen,

its a good arguement, and we have looked, but in the end you have to consider the options there as well, there are a number of hosts that have gone to this, and in the ends they are suffering from the worst downtimes in their history.

Since we are in customer only i will over the next days work to get you some links to their forums, but there are multiple examples of 4+ hour downtimes caused by failing central storage.

In addition we have done some calculations after looking at our servers and the IO per second would have to be insanely huge, the SAN/iSCSI would have to be 1 per 3 servers, at a cost of over 18k per device for covering 3 servers with say a 2TB array, the amount of funds it quite a lot(its worth argueing that the cost of our failed drives is worth a LOT as well, as it is)

However, in recent days we have moved to RAID10 arrays, while certainly not fail-proof, it is fast and it is in many cases multiple redundant.

We have also looked at the extremely high end fiber channel side needing a HBA on each server, these provided better IO but we would still need almost 1 SAN per every 9 servers, at a cost of over 80k per SAN, plus over 1.5k per server for the interconnects.

We have also looked into using it for mail servers only, which still may be the best long term for them, but the RAID10 is working out well to this point, we will know more on it in the coming 6-18 months.

Unix2, it was time for it to be replaced, the drives on it lasted 3 years in 24/7 production, that is a good life, and important to not it never failed, nor did web6 that we migrated some time back, it was just on those servers better to migrate off instead of dealing with 3+ year old drives(google HD study comes to mind), or possibly failing raid card that was still working at the time but showing signs of problems.