I am posting this here and not as a ticket because I do not have what I think is enough info to narrow down a problem, and would like general advice.
I got an email from a web15 client this morning. When I looked and saw there was an issue with web14 lately, I thought that must be it, so it was being resolved… but then looked and saw these people were on web15 not web14. The general issue is sporatic performance on that server.
From my customer:
I’ve been having a lot of trouble getting onto my [name1] site.
Yesterday it was very slow all day as I tried to make updates [via wordpress admin pages]. And this morning (Friday) I cannot seem to get on at all.
Can’t get onto [domain 2, also on web15] site either, with makes me think it’s the server again
I sent out a big [marketing] email yesterday [expecting to generate traffic].
Not good.
And this was not the first incident since they have been on that server… and generally these people are not complainers. I give their reports a lot of credibility.
I think there used to be some performance monitoring on the various servers that we could access. I can’t find it now, does anything like that exist? I would like to let this customer know that these problems are rare, and be able to prove it… or I guess I want to justify moving to another server, and if so, I want to see some stats there before I do.
We have performance monitoring on them still, and get alerts when needed.
There are two things at play. 1. PHP is now isolated per site not run with mod_php like before and 2. that uses more RAM
There had been a few times the RAM gets high usage(and we already run with more RAM than before, but working on adding even more) on the servers making some swapping. This is something we know about, and working on
Web14 and web15 will be having a maintenance in the next 1.5 weeks (sorry don’t know exact time) and we should be able to help things out quite a bit with that. Among other things when the node went down we had to move to one with 12GB or RAM not 16GB, due to available stock of servers. It wasn’t an issue for a while after the hardware swap, but has become so recently
Thanks for the response, but this is really hard to sell to a customer.
I was hopeful that after the migration, jodo in general would have at least a honeymoon period where everything was at a nice over-engineered level, and not close to capacity. But instead we are at a place where it feels like we are barely treading water, and my customers and I don’t like that feeling.
I would like to feel like jodo understood the value in under-promising and over-delivering performance. It makes people happier over the long term, and investments in hardware capacity pay for themselves very quickly in time NOT spent on customer service complaints, improved customer satisfaction, and customer retention.
I don’t know what to tell my customer when I feel like I am asking them to stick it out with me here when I am not really satisfied either.
The hardware web15/web14 is on, it good, but remember it went down and we had to scramble a replacement.
Right now we have a major issue in that we have no more power. I have been on the electricians for all week now calling daily, emailing etc.
I also have a replacement server up for web14/15 but it is holiday in India, I don’t have all the answers here because I have had the server up for 2 weeks now(being that, it is using power already so this is not an excuse). I don’t know what the delay is but will push for why.
I do a lot around here, but I cannot do everything, this has been in the books since almost after the hardware went down in the first place, but has not been completed.
I am sorry, but I don’t have all answers. I can say I am doing what I can, and WILL GET answers.
Also let me add, the hardware is a minor part of the equation, with the migration we make some significant changes in the system for the way all works. notably PHP. In addition we upgraded apache on many of the servers. This has introduced some quirks at times, and we are working to tweak various means to accomodate that.
Almost all of the hardware is new and far better than prior. There is a small amount of reused hardware that was already higher end and Intel VT or AMD SVM capable, but all of it has new hard drives, new upgrade RAM.
Every server instance has more RAM than it had before, across the board, some of it doubled.
Just reading over this again. Sometimes I think I am ‘too honest’ but in reality, there is nothing to hide, and being honest is best in all cases
Also just as a side note, we’ve updated our About pages with some new pics of our datacenter, and we’ll be getting more on the site as well. We are in a private facility now that has a long proven track record.
I appreciate your efforts Stephen, as always. And I know you don’t have the power (electrically or politically) to do everything you want to do as fast as you would like… but all I can do as a reseller/customer of yours is tell you that we notice when things are not right, and it matters!
Right now I went to do a random check on that same domain, and all I get is:
“Error establishing a database connection”
Getting critical here. I need to do something for this customer. I will try to see more on this current issue and send a ticket if I can’t figure it out.
I am not seeing an issue on web15 right now all is loading well. Submit a ticket if needed, they can check this out.
(I also don’t see any on mysql servers at either)
Web14, at least, seems to suffer from regular overuse/abuse. It would be better to find the source than upgrade the hardware to overcome it. I’m still surprised how hard it seems to be to pinpoint the source of resource usage on the servers. Those of us with a vested interest in the workings of Jodohost understand the problems and dedication of the staff, but our clients don’t really give a flying f… they just want 100% reliability and no excuses.
I give a very reasonable rate for technical proofreading/re-writing, by the way, by an actual professional English technical writer (not me.) I know everyone thinks they can write, but most of us can’t. The About page could probably use it…
“All this ensures you data is safer than with us than your hard disk.” Oh dear…
I gave suggestions on that and it didn’t happen I will see that is fixed. In fact I plan to add a lot more there, but I don’t directly edit, just made the suggestions. I CAN access the server, and think I will go fix this one though
now about web14, it is happening due to backups + high traffic + google scans + emails going out at the same time upgrading the disk array is the ONLY thing that is going to fix it.
The high traffic and google scans,and to some extent the emails, happen all the time. I wonder if the problem is caused by more than one user having large automated jobs, most likely backups, set to run at exactly 12 midnight.
Personally, I try to have any automated jobs set to run at more obscure times when the server isn’t likely to be so busy, and try to avoid running them ‘on the hour’ at all. I usually choose some time between 2am and 4 am EST.
When is the server actually least busy?
I notice there is maintenance scheduled for web14 tonight.
Will that have any impact on web15?
Also, what has been the problem with MySQL8? I mentioned an incident in this thread, I saw another incident and sent a ticket yesterday (response:try it now). And I see a thread in “outages” that says there was at least one other incident this week when it was restarted. What is going on, and what is the real solution? This is really bad.
OK, Stephen, don’t take this personally, I know you are doing what you can. But if there is a reason tonight will “help” web15, please give me something to hope about… why would it help?
And MySQL8 seems to be the worst part of current issues for my clients. Like I said, not your fault personally, but I have run out of excuses and dance steps trying to keep my clients, so any real concrete info would help.
These issues have had REAL consequences.
I have had a retail site lose its googe adwords completely because of poor server response/timeouts on their pages when tested by google. They launched a new product, and could not get ads posted at all. “Server Suspended” from google’s system. It took phone calls and begging to get their ad campaign working again after 3 days… if we had not pushed or noticed the problem, it would have been well over a week, waiting for google to re-evaluate, and if they happened to hit another “bad day” forget ads again for another week maybe.
And today from another client:
I just received an email from one of my “email message” subscribers who clicked on a link on my web site and got only the header. Assuming the page was empty she quit.
I know from browsing myself that if a page doesn’t load for me quickly, I will move on…
well, web14 and 15 are on a shared node, web14 moves off…that leave just web15…that means web15 has 6.8GB more RAM and 2 more cores we can allocate..
MySQL8 is already being worked on to help with it.
ok, that is some hope.
So this will be a permanent divorce for web14 and web15?
And will it take long for web15 to have those additional, leftover resources allocated?