Data Centers & Critical Infrastructure
Coming December 2026
Data Center Core
Course 1: Introduction to Data Centers & Critical Infrastructure
Prerequisites: None. This course is the entry point for the Data Centers & Critical Infrastructure library and assumes no prior data center or critical facility experience.
Description: This course introduces participants to data centers and other critical facilities: what they are, what they support, and why they are designed and operated differently from ordinary commercial buildings. The course describes the major areas of a facility, including the white space and support spaces, and surveys the major infrastructure systems that keep the critical load operating: electrical power, cooling and mechanical systems, fire protection and life safety, and monitoring and controls. The course also introduces critical load, availability, redundancy, and resilience at an introductory level, identifies the skilled-trades roles involved in building, operating, and maintaining critical facilities, and establishes the mission-critical operating mindset that is reinforced throughout the series.
Objectives:
- Explain the basic purpose and function of a data center and other critical facilities
- Identify the major areas of a critical facility, including the white space and support spaces
- Identify the major infrastructure systems in a critical facility and explain what each one does
- Explain the relationship between IT equipment and the electrical, cooling, and controls systems supporting it
- Describe critical load, availability, redundancy, and resilience at an introductory level
- Explain why equipment and systems in a critical facility must be viewed as interconnected
- Identify the major skilled-trades roles involved in constructing, operating, and maintaining critical facilities
- Describe the minimum life-safety behaviors expected of anyone working in a critical facility
- Explain why procedures, communication, documentation, and situational awareness are especially important in mission-critical environments
- Recognize that designs and procedures vary by facility and that technicians must follow site-specific requirements
Course 2: Critical Infrastructure Reliability, Redundancy & Resilience
Prerequisites: This course is designed for participants who have completed Introduction to Data Centers & Critical Infrastructure and are familiar with the major areas and infrastructure systems of a critical facility.
Description: This course explains why critical facilities are designed and operated around availability and introduces the architecture vocabulary used throughout the rest of the series. The course describes N, N+1, 2N, and A/B path configurations, reserve capacity, and active/standby concepts, and explains the difference between redundancy and resilience. The course also covers single points of failure, fault domains, common-mode failure, dependency chains, and failure containment; explains concurrent maintainability and how planned maintenance reduces available redundancy; and shows how a locally correct action can create system-level risk through upstream and downstream impacts. This course does not teach formal reliability engineering, availability mathematics, protection design, or facility design certification.
Objectives:
- Explain why critical facilities are designed and operated differently from ordinary commercial facilities
- Describe N, N+1, and 2N redundancy configurations and A/B path configurations, and identify examples of each
- Explain reserve capacity and active/standby concepts
- Distinguish between redundancy and resilience
- Identify single points of failure, fault domains, common-mode failures, and dependency chains
- Explain failure containment and why fault domains are kept separate
- Explain concurrent maintainability and the relationship between current configuration, planned maintenance, and available redundancy
- Describe how upstream and downstream impacts can turn a locally correct action into a system-level risk
- Recognize when a facility is operating outside its normal resilient state
Course 3: Critical Facility Electrical Power Systems
Prerequisites: This course is designed for participants who have completed Introduction to Data Centers & Critical Infrastructure and Critical Infrastructure Reliability, Redundancy & Resilience, and who are familiar with basic electrical principles and electrical safety.
Description: This course surveys how power moves from the utility source to the critical load in a data center or other critical facility. The course describes the major elements of the critical power chain, including utility service, generators, switchgear, UPS systems, power distribution units, and rack-level distribution, and explains how each element supports availability. The course also explains how the redundancy concepts from the previous course are applied to critical power architecture, and shows how to read a basic one-line diagram and trace an A/B power path. Detailed operation, switching, and maintenance of individual power equipment are covered in the Technical courses that follow.
Objectives:
- Describe the path power takes from the utility source to the critical load
- Identify the major elements of the critical power chain and explain the function of each
- Define critical load and distinguish it from other facility loads
- Explain how N+1, 2N, and A/B architecture are applied to critical power systems
- Read a basic one-line diagram and identify the major equipment and connections shown
- Trace an A/B power path from source to rack
- Explain how a failure or maintenance action at one point in the power chain affects equipment upstream and downstream
- Describe the electrical safety boundaries that apply to technicians working around critical power equipment
Course 4: Critical Facility Cooling & Mechanical Systems
Prerequisites: This course is designed for participants who have completed Introduction to Data Centers & Critical Infrastructure and Critical Infrastructure Reliability, Redundancy & Resilience, and who are familiar with basic HVAC and refrigeration principles.
Description: This course surveys how heat is removed from a data center or other critical facility and how the major cooling and mechanical systems work together to protect the critical load. The course explains the heat-removal path from the IT equipment to the outdoors and describes the major components along that path, including computer room air handlers and air conditioners, chilled-water systems, chillers, pumps, cooling towers, economizers, and air management devices. The course also explains airflow management, hot- and cold-aisle concepts, and temperature and humidity setpoints, and shows how redundancy and failure-impact concepts apply to cooling systems. High-density and liquid cooling are introduced at survey depth only and are covered in detail in Advanced Cooling & Liquid Cooling.
Objectives:
- Explain why heat removal is critical to the operation of a data center or critical facility
- Trace the heat-removal path from the IT equipment to the facility's heat-rejection system or other heat sink
- Identify the major cooling and mechanical components and explain the function of each
- Describe how air-based and liquid-based heat transport differ
- Explain airflow management, hot- and cold-aisle arrangements, and containment
- Describe the role of temperature, humidity, and dew-point control targets in protecting the IT load
- Explain how redundancy concepts apply to cooling and mechanical systems
- Describe how a cooling failure or maintenance action affects the critical load and the other facility systems
Course 5: Fire Detection, Suppression & Life Safety
Prerequisites: This course is designed for participants who have completed the Introduction, Reliability, Electrical Power Systems, and Cooling & Mechanical Systems courses in this series and are familiar with the electrical and mechanical systems that fire-protection interlocks act upon.
Description: This course explains how fire is detected and suppressed in an operating data center or critical facility and what technicians are and are not responsible for. The course describes detection methods, including spot smoke, aspirating and very-early-warning, heat, linear, and beam detection, and the effect of high airflow on detection. The course also covers preaction sprinkler concepts, clean-agent suppression awareness, and discharge implications; explains fire alarm control panel states, zones, and addressability; and describes the interlocks that shut down air handlers, close dampers, release doors, and control elevators and emergency power off functions. Technician responsibilities for impairment notification, clearances, hot and dusty work coordination, fire watch, and evacuation are covered in detail. This course does not qualify participants to design, install, inspect, test, or service fire-protection systems and does not replace code-required, employer-required, or site-specific training.
Objectives:
- Identify the major fire detection methods used in critical facilities and explain how high airflow affects them
- Describe preaction sprinkler systems and clean-agent suppression systems at awareness depth
- Explain the implications of a suppression discharge for the protected space and the equipment in it
- Identify alarm, trouble, and supervisory states on a fire alarm control panel
- Explain zones and addressability and how fire alarm conditions may be interfaced to the BMS and other facility monitoring systems
- Describe common fire-protection interlocks and interfaces that may affect air handlers, dampers, doors, elevators, and emergency power-off functions
- Explain technician responsibilities for impairment notification, maintaining clearances, and preventing obstruction
- Describe the coordination required for hot work, dusty work, and fire watch
- Explain why unauthorized silencing or bypassing of fire-protection systems is not permitted
- Describe the expected evacuation behavior in a critical facility
Course 6: Monitoring, Controls & Infrastructure Systems
Prerequisites: This course is designed for participants who have completed Courses 1 through 5 of this series and are familiar with the power, cooling, and fire-protection systems whose conditions are monitored.
Description: This course surveys how sensors, controls, alarms, and monitoring systems provide visibility into a data center or critical facility and support its operation. The course describes the common sensors and field devices used to measure electrical, thermal, environmental, and life-safety conditions; explains how control loops, setpoints, and sequences of operation govern facility equipment; and introduces alarm philosophy, including alarm priorities and the difference between an alarm and a normal state change. The course also explains the purpose of building management systems, data center infrastructure management systems, and related monitoring platforms, and how they relate to the systems they supervise. Practical operating depth on these platforms is covered in BMS, DCIM & Critical Facility Monitoring.
Objectives:
- Identify the common sensors and field devices used to monitor electrical, thermal, environmental, and life-safety conditions
- Explain how sensors, controllers, and controlled equipment work together in a control loop
- Describe setpoints and sequences of operation at an introductory level
- Explain alarm philosophy, alarm priorities, and the difference between an alarm and a normal state change
- Describe the purpose of a building management system and a data center infrastructure management system
- Explain how power, cooling, environmental, and fire conditions are presented as monitored facility conditions
- Explain why field verification should be performed, when safe and practical, before acting on a monitored value or alarm
- Describe how monitoring supports situational awareness and configuration awareness
Course 7: Critical Facility Operations, Procedures & Maintenance
Prerequisites: This course is designed for participants who have completed Courses 1 through 6 of this series and are familiar with all of the major systems of a critical facility.
Description: This course describes the normal operating discipline expected of technicians in a data center or other critical facility and closes the Core wave of the series. The course explains the purpose and use of standard operating procedures and methods of procedure, work authorization and permitting, and change control; describes configuration awareness and how planned maintenance affects available redundancy; and covers communication, documentation, shift turnover, and escalation. The course also explains how planned maintenance is coordinated so that work on one system does not create risk to another. Abnormal and emergency operations are deliberately excluded from this course and are covered in Critical Incident Response, Emergency Operations & Recovery.
Objectives:
- Explain why disciplined procedures are especially important in a critical facility compared with an ordinary commercial facility
- Describe the purpose and use of standard operating procedures and methods of procedure
- Explain work authorization, permitting, and change control and why required approvals must be obtained before work begins
- Describe configuration awareness and how to confirm the current state of the facility before starting work
- Explain how planned maintenance reduces available redundancy and how that risk is managed
- Describe how maintenance is coordinated across systems and crafts
- Explain the communication, documentation, and logging expectations for critical facility work
- Conduct an effective shift turnover
- Describe when and how to escalate a condition, and why escalation is never a failure
Data Center Technical
Course 8: Electrical Distribution & Switching
Prerequisites: This course is designed for participants who have completed Critical Facility Electrical Power Systems and Critical Facility Operations, Procedures & Maintenance, and who are familiar with AC characteristics, three-phase circuits, fuses and circuit breakers, electrical safety, and electrical print reading.
Description: This course describes the electrical distribution equipment found in a data center or critical facility and the switching discipline used to operate it. The course covers switchgear, switchboards, transformers, power distribution units, remote power panels, and busway; explains the function of protective devices, including circuit breakers, fuses, and protective relays at awareness depth; and shows how A/B paths and redundant distribution are laid out on facility one-line diagrams. The course also explains switching boundaries, isolation points, and the discipline required for approved switching operations. This course does not qualify participants for energized electrical work and does not replace lockout/tagout, arc flash, or qualified-worker training.
Objectives:
- Identify switchgear, switchboards, transformers, PDUs, RPPs, and busway and explain the function of each
- Describe the function of circuit breakers, fuses, and protective relays in facility distribution
- Explain the purpose of protective-device coordination and how coordination can help isolate a fault while minimizing unnecessary interruption of downstream loads
- Read a facility one-line diagram and identify A/B paths, tie points, and isolation boundaries
- Explain how a switching operation changes the configuration and available redundancy of the facility
- Describe the discipline required for an approved switching operation, including the switching plan, verification, and communication
- Identify the isolation points and boundaries relevant to a maintenance task
- Recognize the qualification and authorization limits that apply to switching and energized work
Course 9: Utility Service, Medium Voltage & Power Quality
Prerequisites: This course is designed for participants who have completed Critical Facility Electrical Power Systems, Critical Facility Operations, Procedures & Maintenance, and Electrical Distribution & Switching, and who are familiar with AC characteristics and three-phase circuits.
Description: This course connects the external utility source to the internal distribution architecture of a data center or critical facility. The course describes how utility service enters the site and transitions into facility distribution, and covers the service entrance, substations, medium-voltage switchgear, and transformers at technician awareness depth. The course also explains normal and alternate sources, metering, basic grounding and bonding context, and power-quality monitoring; describes voltage sag, swell, interruption, transients, harmonics, imbalance, frequency variation, and power factor as they affect critical loads; and explains load shedding and why power quality can be a systems problem rather than an isolated equipment problem. This course does not teach utility-system design, protective-relay engineering, or medium-voltage maintenance, and does not qualify participants for medium-voltage operation.
Objectives:
- Describe how utility service enters a site and transitions into facility distribution
- Identify the service entrance, substation, medium-voltage switchgear, and transformers and explain the function of each
- Explain normal and alternate sources and how the facility transitions between them
- Describe the purpose of utility and facility metering
- Explain basic grounding and bonding context as it applies to critical facility distribution
- Identify voltage sag, swell, interruption, transients, harmonics, imbalance, and frequency variation and describe their effects on critical loads
- Explain power factor and why it matters to the facility and the utility
- Describe load shedding and source limitations
- Recognize the boundaries of technician responsibility around medium-voltage equipment and know when to escalate
Course 10: UPS, Batteries & Energy Storage
Prerequisites: This course is designed for participants who have completed Critical Facility Electrical Power Systems, Electrical Distribution & Switching, and Utility Service, Medium Voltage & Power Quality, and who are familiar with AC/DC theory and electrical safety.
Description: This course explains how uninterruptible power supply systems and stored energy protect the critical load during source disturbances and transfers. The course describes UPS topologies and operating modes, including normal, battery, bypass, and maintenance bypass, and explains the relationship between the UPS and the distribution and source architecture already established. The course also covers battery and energy-storage technologies, including valve-regulated lead-acid and lithium-ion batteries and other stored-energy systems; describes battery monitoring, alarms, and common failure modes; and explains the maintenance considerations and safety hazards associated with stored energy. This course does not qualify participants to service UPS or battery systems and does not replace manufacturer or qualified-worker training.
Objectives:
- Explain the purpose of a UPS in the critical power chain
- Identify common UPS topologies and describe how they protect the load
- Describe normal, battery, static bypass, and maintenance bypass modes and explain when each is used
- Explain how the UPS interacts with upstream distribution, downstream loads, and the standby source
- Identify the battery and energy-storage technologies used in critical facilities and describe their characteristics
- Describe battery monitoring and the alarms that indicate degraded stored-energy capacity
- Identify common UPS and battery failure modes and their effect on available redundancy
- Describe the maintenance considerations and safety hazards associated with UPS and stored-energy systems
- Recognize the qualification and authorization limits that apply to UPS and battery work
Course 11: Generators & Transfer Systems
Prerequisites: This course is designed for participants who have completed Critical Facility Electrical Power Systems, Electrical Distribution & Switching, Utility Service, Medium Voltage & Power Quality, and UPS, Batteries & Energy Storage.
Description: This course describes how standby generation and transfer systems restore power to a data center or critical facility when the normal source is lost. The course covers standby generator components and support systems, including engine starting, fuel, cooling, exhaust, and controls; explains automatic transfer switches and paralleling switchgear; and describes the transfer sequence from loss of source through generator start, transfer, and retransfer. The course also explains how transfer timing interacts with UPS ride-through, describes generator testing and load-bank exercises, and covers the operating response expected of technicians during a transfer event. This course does not qualify participants to service generators or transfer equipment.
Objectives:
- Explain the role of standby generation in the critical power chain
- Identify the major components of a standby generator set and explain the function of each
- Describe the engine starting, fuel, cooling, and exhaust support systems and the common failures associated with each
- Explain how automatic transfer switches and paralleling switchgear work
- Describe the transfer sequence from loss of normal source through generator start, transfer, and retransfer
- Explain how transfer timing interacts with UPS ride-through and stored-energy capacity
- Describe generator testing, including no-load, building-load, and load-bank tests, and why each is performed
- Describe the operating response expected of technicians during a transfer event
- Explain how a generator or transfer failure affects facility redundancy and when to escalate
Course 12: Advanced Cooling & Liquid Cooling
Prerequisites: This course is designed for participants who have completed Critical Facility Cooling & Mechanical Systems and Critical Facility Operations, Procedures & Maintenance, and who are familiar with basic HVAC, refrigeration, and pump principles.
Description: This course builds on the cooling survey to explain how high-density and liquid-cooled environments are cooled and maintained. The course describes why AI and high-performance computing workloads exceed the limits of conventional air cooling and covers rear-door heat exchangers, in-row and overhead cooling, direct-to-chip cooling, and immersion cooling at awareness depth. The course also explains coolant distribution units, primary and secondary liquid loops, coolant and water-quality requirements, leak detection, and the controls that govern liquid cooling, and describes how hybrid air and liquid environments operate together. Fast-changing content in this course should be verified against current authoritative sources during technical review.
Objectives:
- Explain why high-density workloads can exceed the practical capability of conventional air cooling
- Identify rear-door heat exchangers, in-row and overhead cooling, direct-to-chip cooling, and immersion cooling and describe how each removes heat
- Describe the function of a coolant distribution unit
- Describe common primary and secondary liquid-loop architectures and trace the applicable cooling path from IT equipment to the facility heat-rejection system
- Explain coolant and water-quality requirements and why they matter
- Describe leak detection methods and the response expected when a leak is detected
- Explain how liquid cooling controls maintain flow, temperature, and pressure
- Describe how hybrid air and liquid environments are operated and maintained together
- Explain how a liquid cooling failure affects the critical load and available redundancy
Course 13: BMS, DCIM & Critical Facility Monitoring
Prerequisites: This course is designed for participants who have completed Monitoring, Controls & Infrastructure Systems and the Technical courses on electrical distribution, utility service, UPS, generators, and advanced cooling.
Description: This course provides practical operating depth on the building management and data center infrastructure management systems used to monitor and support operation of a critical facility. The course explains how to read and interpret dashboards, alarms, trends, and setpoints; describes sequences of operation and how to recognize when a system is not following its sequence; and covers capacity indicators for power, cooling, and space. The course also explains alarm handling, including acknowledgment, prioritization, and escalation, and reinforces field verification of monitored values before action is taken. Detailed programming and configuration of BMS or DCIM platforms are not covered.
Objectives:
- Navigate a BMS or DCIM dashboard and identify the systems and points displayed
- Interpret alarms, alarm priorities, and alarm states and describe the correct handling of each
- Read a trend and use it to identify a developing condition
- Explain setpoints and identify when a system is operating outside its setpoint
- Recognize when a system is not following its sequence of operation
- Interpret power, cooling, and space capacity indicators
- Explain the field verification required before acting on a monitored value or alarm
- Describe how monitoring data supports maintenance coordination, configuration awareness, and escalation
Course 14: Commissioning & Integrated Systems Testing
Prerequisites: This course is designed for participants who have completed Critical Infrastructure Reliability, Redundancy & Resilience and Courses 5 through 13 of this series, and who are familiar with the component systems, interlocks, procedures, and monitoring of a critical facility.
Description: This course teaches how technicians verify that individual systems, and the facility as a whole, work as intended under planned test conditions. The course explains the purpose of commissioning, the roles involved, and the structure of test plans, prerequisites, acceptance criteria, and documentation. The course also covers startup and turnover, baseline readings, sequence, alarm, and interlock verification, and field confirmation; describes functional performance testing of individual systems and integrated systems testing across power, cooling, controls, and life-safety interfaces; and explains planned failure-mode tests, abort conditions, punch lists, deficiency correction, retesting, and configuration restoration. This course does not credential a commissioning authority and does not authorize participants to create test procedures or defeat safeguards.
Objectives:
- Explain the purpose of commissioning and identify the roles involved
- Describe the elements of a test plan, including prerequisites, acceptance criteria, and documentation
- Explain startup and turnover and the purpose of baseline readings
- Execute assigned sequence, alarm, and interlock verification activities under an approved test plan
- Distinguish between functional performance testing and integrated systems testing
- Describe how integrated systems testing exercises the interfaces between power, cooling, controls, and life-safety systems
- Explain planned failure-mode tests, including what is changed, what should happen, what must be monitored, and what constitutes an abort condition
- Describe punch lists, deficiency correction, retesting, and recordkeeping
- Explain why configuration must be restored and verified after testing
- Describe technician responsibilities during test witnessing, communication, and safe execution of approved tests
Course 15: Critical Incident Response, Emergency Operations & Recovery
Prerequisites: This course is designed for participants who have completed Critical Infrastructure Reliability, Redundancy & Resilience and Courses 5 through 14 of this series. It is the final course in the series and draws on the terminology and system behavior of the entire library.
Description: This course teaches disciplined technician response to abnormal and emergency conditions in a data center or critical facility. The course explains how to recognize abnormal conditions, prioritize alarms, and distinguish degraded operation from an immediate emergency; describes the use of emergency operating procedures, communication, escalation, role clarity, and decision logging; and covers representative scenarios, including utility failure, UPS and battery alarms, generator and transfer failure, cooling loss, leak detection, fire and life-safety events, equipment trips, and multiple simultaneous alarms. The course also explains the stabilize-first principle, restoration and return-to-normal criteria, temporary configurations, handoff, evidence preservation, and post-event review. This course does not replace site emergency operating procedures, incident command plans, emergency-services direction, or qualified-worker requirements.
Objectives:
- Recognize abnormal conditions and distinguish degraded operation from an immediate emergency
- Prioritize alarms and identify a loss of redundancy
- Explain the purpose and use of emergency operating procedures
- Describe the communication, escalation, role clarity, and decision logging expected during an incident
- Describe the expected technician response to utility failure, UPS and battery alarms, generator and transfer failure, cooling loss, leaks, fire and life-safety events, and equipment trips
- Explain how to respond when multiple alarms occur at the same time
- Apply the stabilize-first principle: protect life and safety, preserve remaining capacity, avoid compounding failures, and verify system state before taking additional action when conditions permit
- Describe restoration, return-to-normal criteria, and the management of temporary configurations
- Explain handoff, evidence preservation, and the purpose of post-event review