PSU Error Alert - Troubleshooting Guide
Purpose
This
document provides a standardized approach for handling and troubleshooting the PSU
Error alert, ensuring quick identification of power supply issues and
maintaining system power redundancy.
Alert Name
PSU Error
Alert Description
This alert is triggered when a power supply unit (PSU)
error or AC power loss is detected.
It indicates a fault in one or more PSUs, which are critical
for maintaining redundant power in the system.
Loss of a PSU increases the risk of:
- Node
outage
- System
instability
- Complete power failure if redundancy is lost
Severity
P2 - High Impact
Possible Causes
PSU
hardware failure
AC
power cable disconnection
- Power
source failure
- Loose
PSU connection
- Faulty
power feed or PDU issue
- Redundancy configuration issue
L1 Engineer Actions
Step 1 – Log the Alert
- Log
in to the monitoring system
- Review
alert summary and affected node
Step 2 – Create Ticket
- Inform L2 team
- Create an internal ticket including:
- Alert name
- Hostname
- PSU details (if available)
- Timestamp
- Observations
Step 3 – Escalation (If required)
- Share alert details with L2 for further investigation
Step 4 – Documentation
- After confirmation with L2:
- Update Issue Master Sheet
- Ensure proper tracking of the issue
L2 Engineer Actions
Step 1 – Verify via BMC
- Log
in to BMC/IPMI interface
- Check
hardware health status
- Identify
PSU alerts or failures
Step 2 – Identify Failed PSU
- Determine:
- PSU
unit (Side A or Side B)
- Fault type (failure / not present / AC loss
Step 3 – Check in different rack and row
Step 4 - OEM Escalation
- Inform
OEM/vendor with:
- Hardware
logs (BMC screenshots/logs)
- Server
details
- PSU
failure details
Step 5 – Physical Verification
- Reseat
the PSU if loose
- Check
power cable connections
- Replace PSU if faulty
Step 6 - Check Power Feed Redundancy
- Verify
both power supplies are connected to:
- Separate
power sources (PDU/UPS)
- Ensure redundancy is intact
Resolution
The issue is considered resolved when:
- All
PSUs are operational
- Power
redundancy is restored
- No
hardware errors are reported in BMC
- Alert is cleared
Escalation
- L1
→ L2: For validation and troubleshooting
- L2
→ OEM/Vendor: For PSU replacement or hardware issues
Related Articles
Network Down Extern (P1 Critical Alert) - Troubleshooting Guide
Purpose This document provides a standardized approach for handling and troubleshooting the Network Down - Extern alert, including L1 and L2 actions, escalation procedures, and resolution criteria for critical incidents. Alert Name Network Down - ...
ZFS Pool Capacity Critical Alert - Troubleshooting Guide
Purpose This document provides a standardized approach for handling and troubleshooting the ZFS Pool Capacity Critical alert, enabling rapid response to prevent storage exhaustion, performance degradation, and potential service outages. Alert Name ...
Network Down Boot (P1 Critical Alert)- Troubleshooting Guide
Purpose This document provides a standardized approach for handling and troubleshooting the Network Down - Boot alert, ensuring quick response to critical incidents and maintaining cluster control and system accessibility. Alert Name Network Down - ...
ZFS Pool Capacity Warning Alert - Troubleshooting Guide
Purpose This document provides a standardized approach for handling and troubleshooting the ZFS Pool Capacity Warning alert, helping L1 and L2 engineers prevent storage exhaustion and maintain system stability. Alert Name ZFS Pool Capacity Warning ...
ZFS Pool Error Alert - Troubleshooting Guide
Purpose This document provides a standardized approach for handling and troubleshooting the ZFS Pool Error alert, enabling L1 and L2 engineers to quickly identify disk-related issues, prevent data loss, and maintain storage reliability. Alert Name ...