Service outage Nov. 23, 2010 (resolved, updated)

Our primary data center had another power interruption this morning at 7:28 am (Pacific time). All of our servers lost power and then had it restored, thus rebooting them. All customer web sites were unavailable during this time. Incoming email would have simply been delayed during the downtime, not lost. When the servers came back online e-mail may have seemed sluggish to some customers for a while but this should also be fixed now.

This incident follows another power incident the previous Saturday night. We are working with the data center to get more details, including an estimate of when they will have replaced any faulty equipment. We will update this post as more information becomes available.

Update Nov. 29: The final data center report is that on the night of November 20, lightning strikes damaged both of the redundant UPS systems, interrupting data center power for a few seconds. The UPS manufacturer scheduled replacements for November 23, but another PG&E utility power interruption lasting a few seconds occurred that morning before it was finished. The UPS manufacturer has since replaced all damaged parts, restoring full redundancy. In addition, the UPS manufacturer has overhauled each unit, replacing and upgrading other parts to increase robustness. We take this very seriously — it’s at the core of what we do — and we will continue to work with the data center to ensure that their infrastructure meets our high standards.

Service outage Nov. 20, 2010 (resolved)

A major power failure at our primary data center in Fremont, California, caused a complete outage for nearly all services beginning at 8:32 PM Pacific time Saturday night. It lasted between six and 13 minutes, depending on the server. Only our blog and redundant DNS infrastructure was unaffected.

All services are now fully operational; please don’t hesitate to contact us if you have any questions. We sincerely apologize for the inconvenience this caused our customers.

Read the rest of this entry »

Having trouble sending mail to aol.com? (resolved)

The aol.com mail servers have been having problems for the last 24 hours, according to the AOL blog and the AOL Twitter feed.

Read the rest of this entry »

Brief scheduled maintenance on flexo server (completed)

At approximately 10:00 PM Pacific time tonight, October 23, the “flexo” Web server will be restarted.

As a result, for customers on the “flexo” server (only), Web site service and the ability to read incoming e-mail will be unavailable for approximately five minutes. Customers on other servers will not be affected.

Read the rest of this entry »

Brief scheduled maintenance Friday, September 24 (completed)

Between 11:00 PM and 11:30 PM Pacific time this Friday, September 24, some of our hosting servers will be restarted. As a result, some customers will find that Web site service and the ability to read incoming e-mail will be unavailable for approximately five minutes at some point during this maintenance “window”.

Read the rest of this entry »

High load on some servers (resolved)

Three of our Web hosting servers (amy, flexo, and leela) experienced high load earlier today that caused some customers to see “503 errors” on their Web sites for a few minutes.

This was caused by an upgrade to the eAccelerator PHP caching system that removed all the cached files at once, which doesn’t normally happen.

The problem has been permanently resolved and will not recur.

Read the rest of this entry »

Brief scheduled maintenance Saturday, August 28 (completed)

Between 10:00 PM and 11:59 PM Pacific time this Saturday, August 28, all our hosting servers will be restarted. As a result, Web site service and the ability to read incoming e-mail will be unavailable for approximately five minutes at some point during this maintenance “window”.

Read the rest of this entry »

Comcast network problems August 12 (resolved)

Our monitoring systems are showing that some people who reach our servers via an “Internet backbone” company called Global Crossing, including some Comcast cable customers, have been intermittently unable to connect over the last hour or so.

This isn’t an outage on our end; these visitors are also unable to reach other sites that Comcast routes through Global Crossing (and not related to us), such as www.globalcrossing.com. It’s something Comcast and Global Crossing need to address.

We’ll continue to monitor this issue closely and post an update when we’re confident that it’s been resolved.

By the way, if you ever find that you’re unable to connect to our servers (or anyone else’s), a very useful site is CheckSite.us. It shows you whether the destination servers are down, or whether the problem is just a local routing problem that isn’t affecting most other people.

Update 9 AM PDT August 13: According to our monitoring systems, Comcast resolved this shortly after our post, and the problem has not recurred in the ten hours since then.

Brief scheduled maintenance Monday, August 2 on some servers (completed)

Between 11:00 PM and 11:59 PM Pacific time tonight (Monday August 2), several of our hosting servers will be restarted: bender, elzar, farnsworth, lrrr, mom, and seymour.

As a result, Web site service and the ability to read incoming e-mail for some customers will be unavailable for approximately five minutes at some point during this maintenance “window”.

Read the rest of this entry »

Brief maintenance on calculon server (completed)

The “calculon” Web server will be restarted at 9 PM Pacific time tonight (July 5). This will cause a five-minute interruption of Web and e-mail service for customers on that server.

Other servers will not be affected, and incoming mail will only be delayed, not lost.

Read the rest of this entry »