Computer Room Emergency: Only a Matter of Time

"Gideon T. Rasmussen, CISSP, CISA, CISM, CFSO, SCSA" <[email protected]> Thu, 02 Dec 2004 16:54:10 -0500
Newsgroups gmane.comp.security.papers
Message-ID <[email protected]>
http://www.cyberguard.com/news_room/news_newsletter_112304emergency.cfm

Computer Room Emergency =96 Only a Matter of Time
Gideon T. Rasmussen - CISSP, CISM, CFSO, SCSA

It's an infrastructure manager's worst nightmare: The computer room is=20
down. There are several events that can make this scenario a reality. A=20
hurricane knocks out power for several days. Building management=20
disrupts power for scheduled maintenance. Construction workers sever an=20
underground power line.

Small and medium sized organizations may not have adequate UPS and=20
generator systems. In that case, it is only a matter of time before=20
power is disrupted and the computer room must be shut down. Enterprise=20
class computer rooms with an absolute requirement for 24/7 uptime may=20
still be disrupted by an emergency. For example, a liquid spill such as=20
sprinkler discharge or a glycol spill from a broken AC pipe may=20
necessitate an emergency shutdown.

Preparation

Every six months completely power down the computer room to prepare for=20
the inevitable. Hold a meeting to plan the exercise. The goal is to=20
efficiently stop and start mission critical systems as quickly as=20
possible. Assign tasks to each team member. The meeting is an ideal time=20
to brainstorm.

Start by documenting the order in which systems will be shut down.=20
Address all critical systems. In addition to servers, this includes=20
networking gear, telecommunications equipment, UPS and AC units. Halt=20
systems in order of criticality. This helps minimize damage in the event=20
that UPS or generator systems fail. Shut down systems that are prone to=20
data loss or are painful to restore early on. Also consider dependencies=20
between systems. For example, the infrastructure supporting a=20
three-tiered application must be shut down in order. In most=20
organizations the development systems will be last on the list.=20
Carefully consider the order for starting systems as well. The start=20
order will be slightly different, as there is no concern for failing=20
power systems.

Create an operations guide for each system. Each OPS guide should be a=20
single point of reference. Detail stop/start procedures, where the=20
system is located and how to confirm it is providing services (versus=20
merely running from an operating system perspective). Keep in mind that=20
the guide may be used by a technologist who has little or no experience=20
with the system. Include a revision date at the bottom of each page.

Policies and procedures should ensure current administrative passwords=20
are available and appropriately safeguarded. Maintain a recall roster so=20
that the infrastructure team can be contacted in the event of an emergenc=
y.

Label systems and racks for easy identification (front and back). If a=20
keyboard, video, mouse (KVM) device is in use, label it with the systems=20
it is connected to and the key sequence required to switch between them.

Hardware may fail once powered down. Ensure tech support contracts are=20
current and support phone numbers are documented. Current backups and=20
installation media must also be on hand at the time of the exercise.

Consider whether the computer room UPS system can handle the current=20
load. Have new systems been added in the past six months? It might make=20
sense to have a UPS technician on-site and test UPS capacity and system=20
health.

Print the OPS guides and staple them separately. Separate guides enable=20
personnel to work without sharing documentation. Upon completion of a=20
task, they can return to the team lead to address any remaining systems.=20
This also helps track progress and makes efficient use of resources.

Meet again before the exercise and conduct a dry run-through. Take note=20
of any issues and fine tune the documentation.

Plan Execution

A senior team member should direct and monitor the progress of the=20
exercise. Coordinate and reassign resources as they become available.=20
Make use of available personnel and system keyboards. Take note of=20
elapsed time, discrepancies in documentation and issues as they arise.

Document functionality testing to ensure that once systems are powered=20
up they are providing the services required. Turn off internal=20
monitoring systems as late as possible. A shutdown exercise is the=20
perfect opportunity to test monitoring. Document notification from=20
external monitoring services as well.

Lessons Learned

At the conclusion of the exercise, the time required to shut down and=20
restart the enterprise systems will be known. The preparation required=20
keeps documentation current. The exercise itself provides valuable=20
on-the-job training. This continuity helps eliminate single points of=20
failure.

Provide senior management with a formal report detailing the results of=20
the exercise. Powering down the computer room is one of the first steps=20
of taking ownership of the organization=92s infrastructure. In my=20
experience, many things fall out of this exercise. It is better to learn=20
about them during a maintenance window rather than complicate an=20
emergency situation.