Hostmaster and ResellerClub issues
This week, several domain-related issues occurred: ResellerClub, the largest international domain registrar and hosting provider, experienced issues with its control panel and API, and Hostmaster, the administrator of the .UA domain and other Ukrainian domains, experienced an outage in its UAEPP domain registration system.
Using a brief example of how two companies interact with their partners and registrars, I'd like to show you this small but significant difference in doing business.
ResellerClub
Social media posts:
@ResellerClub: We're experiencing some issues that affect the Control Panel, SuperSite & API & they should be resolved in 2 hrs. Thanks for your patience
TeamResellerClub: We're experiencing some issues with our system that affect the Control Panel, SuperSite & the API. We're working our hardest to fix it at the moment and will have an update for you very soon. You can also follow the issue on our forum - http://bit.ly/W12vjY
And after the problem was resolved, a detailed report arrived detailing what had actually happened, what measures had been taken to ensure it didn't happen again, and that the company deeply regretted causing inconvenience to all its customers:
Yesterday, there was an emergency maintenance that was carried out on our platform servers due to which you may have not been able to access your Control Panel. The downtime lasted for a few hours and we can confirm that none of your orders or data was lost during the downtime. To help you understand this downtime, here is a summary of what happened.
At 10:49 UTC, our systems operators executed a command line that overwrote a few config files within our platform database. While all data remained secure, the files that were overwritten caused the database to stop functioning. This affected all platform related services (Control Panel, API & Whois).
Immediately, our monitoring systems detected the problem and we shutdown the platform to avoid the replication of the broken database structure and avoid any inconsistency issues. Following this, we started inspecting our database consistency and switched over to a standby database when necessary.
The verification of our database was a fairly lengthy process where our systems, software engineering and management teams started verifying each and every database and table to ensure that all transactions that were applied on the primary database were also applied on the standby database. This took a while as we had to bring up the primary database first and restore the postgres system data files to identify all platform related databases.
At 16:37 UTC, all data was perfectly verified, tested and subsequently made live. We will be reviewing this incident and come up with standard operating procedures to ensure that such a downtime does not reoccur.
We have added layers of security to all server config directories. This will make sure that we do not modify or overwrite server config files at any point of time. We will also be setting up new SOPs to check data consistency across our slave databases in a much faster manner.
We deeply regret the inconvenience caused during the downtime. Our teams worked diligently during the restore process and have now restored a fully functioning platform to our standby servers. We will need to switch back to our primary servers for which we have scheduled a one hour maintenance window on Sunday, 27th January, 2013 at 04:30 UTC.
Hostmaster
I would like to point out that I have omitted the many letters, questions and problems that arose due to the failure and to which, unfortunately, no answers were given.
Messages in the registrar mailing list:
14:24 (one of the registrars writes)
Hello.
I should probably have received a notification email by now that UAEPP has gone offline... But I haven't... Strange...
https://epp.hostmaster.ua/auth/ says "Service is temporarily unavailable"
The EPP server is also down...
14:28 (written by a representative of Hostmaster)
Good afternoon
EPP service is temporarily unavailable.
Recovery time is 10 minutes.
We apologize for the forced shutdown.
And a "detailed" explanation of the problem after the failure was fixed:
Due to a technical glitch in the UAEPP domain name registration system, registrars and registrants in public domains in Poltava, Ivano-Frankivsk, and Kherson may have received an erroneous notification about the expiration of their domain name registration.
All information about domain names in the specified public domains has been restored from a backup copy.
We apologize for the inconvenience.
The consequences of the failure have been eliminated. Operations can now resume as normal.
Due to a glitch in the UAEPP domain registration system, registrar clients received false notifications about their domains' expiration dates, and it's not Hostmaster that will be held accountable for this, but the registrars themselves and their support team.
I sincerely hope that the level of service and responsibility in our country would be at least a tenth of that of some local and international companies.
P.S. All messages quoted in this note were sent to general mailing list recipients without a note that they should not be published.