Computer Room Emergency: Only a Matter of Time
"Gideon T. Rasmussen, CISSP, CISA, CISM, CFSO, SCSA" <[email protected]> Thu, 02 Dec 2004 16:54:10 -0500
| Newsgroups | gmane.comp.security.papers |
|---|---|
| Message-ID | <[email protected]> |
http://www.cyberguard.com/news_room/news_newsletter_112304emergency.cfm Computer Room Emergency =96 Only a Matter of Time Gideon T. Rasmussen - CISSP, CISM, CFSO, SCSA It's an infrastructure manager's worst nightmare: The computer room is=20 down. There are several events that can make this scenario a reality. A=20 hurricane knocks out power for several days. Building management=20 disrupts power for scheduled maintenance. Construction workers sever an=20 underground power line. Small and medium sized organizations may not have adequate UPS and=20 generator systems. In that case, it is only a matter of time before=20 power is disrupted and the computer room must be shut down. Enterprise=20 class computer rooms with an absolute requirement for 24/7 uptime may=20 still be disrupted by an emergency. For example, a liquid spill such as=20 sprinkler discharge or a glycol spill from a broken AC pipe may=20 necessitate an emergency shutdown. Preparation Every six months completely power down the computer room to prepare for=20 the inevitable. Hold a meeting to plan the exercise. The goal is to=20 efficiently stop and start mission critical systems as quickly as=20 possible. Assign tasks to each team member. The meeting is an ideal time=20 to brainstorm. Start by documenting the order in which systems will be shut down.=20 Address all critical systems. In addition to servers, this includes=20 networking gear, telecommunications equipment, UPS and AC units. Halt=20 systems in order of criticality. This helps minimize damage in the event=20 that UPS or generator systems fail. Shut down systems that are prone to=20 data loss or are painful to restore early on. Also consider dependencies=20 between systems. For example, the infrastructure supporting a=20 three-tiered application must be shut down in order. In most=20 organizations the development systems will be last on the list.=20 Carefully consider the order for starting systems as well. The start=20 order will be slightly different, as there is no concern for failing=20 power systems. Create an operations guide for each system. Each OPS guide should be a=20 single point of reference. Detail stop/start procedures, where the=20 system is located and how to confirm it is providing services (versus=20 merely running from an operating system perspective). Keep in mind that=20 the guide may be used by a technologist who has little or no experience=20 with the system. Include a revision date at the bottom of each page. Policies and procedures should ensure current administrative passwords=20 are available and appropriately safeguarded. Maintain a recall roster so=20 that the infrastructure team can be contacted in the event of an emergenc= y. Label systems and racks for easy identification (front and back). If a=20 keyboard, video, mouse (KVM) device is in use, label it with the systems=20 it is connected to and the key sequence required to switch between them. Hardware may fail once powered down. Ensure tech support contracts are=20 current and support phone numbers are documented. Current backups and=20 installation media must also be on hand at the time of the exercise. Consider whether the computer room UPS system can handle the current=20 load. Have new systems been added in the past six months? It might make=20 sense to have a UPS technician on-site and test UPS capacity and system=20 health. Print the OPS guides and staple them separately. Separate guides enable=20 personnel to work without sharing documentation. Upon completion of a=20 task, they can return to the team lead to address any remaining systems.=20 This also helps track progress and makes efficient use of resources. Meet again before the exercise and conduct a dry run-through. Take note=20 of any issues and fine tune the documentation. Plan Execution A senior team member should direct and monitor the progress of the=20 exercise. Coordinate and reassign resources as they become available.=20 Make use of available personnel and system keyboards. Take note of=20 elapsed time, discrepancies in documentation and issues as they arise. Document functionality testing to ensure that once systems are powered=20 up they are providing the services required. Turn off internal=20 monitoring systems as late as possible. A shutdown exercise is the=20 perfect opportunity to test monitoring. Document notification from=20 external monitoring services as well. Lessons Learned At the conclusion of the exercise, the time required to shut down and=20 restart the enterprise systems will be known. The preparation required=20 keeps documentation current. The exercise itself provides valuable=20 on-the-job training. This continuity helps eliminate single points of=20 failure. Provide senior management with a formal report detailing the results of=20 the exercise. Powering down the computer room is one of the first steps=20 of taking ownership of the organization=92s infrastructure. In my=20 experience, many things fall out of this exercise. It is better to learn=20 about them during a maintenance window rather than complicate an=20 emergency situation.