Network Down Boot (P1 Critical Alert)- Troubleshooting Guide
Purpose
This document provides a standardized approach for handling and troubleshooting the Network Down - Boot alert, ensuring quick response to critical incidents and maintaining cluster control and system accessibility.
Alert Name
Network Down - Boot
Alert Description
This alert is triggered when the boot network interface on a node is down or unavailable.
This interface is critical for:
-
Management communication
- Cluster control connectivity
- External access to the system
Loss of this interface can isolate nodes from cluster control and prevent user/system access.
Severity
P1 - Critical Impact
Possible Causes
- Network cable disconnected or faulty
- SFP/transceiver failure
- Switch port down or misconfigured
- Gateway unreachable
- NIC hardware issue
L1 Engineer Actions
Step 1 – Log the Alert
- Log the alert in the monitoring system
- Capture:
- Alert name
- Node/interface details
- Timestamp
- Current status
Step 2 – Create Ticket
- Monitor the alert for 15 minutes
- If the alert persists:
- Create an internal ticket
- Include all captured details
Step 3 – Escalation (If required)
- Inform L2 engineer immediately (P1 priority)
- Share complete alert information
Step 4 – Documentation
- After confirmation with L2:
- Update Issue Master Sheet/ Issue Tracker Sheet
L2 Engineer Actions
Step 1 – Network Reachability
- Ping the default gateway from the affected node
- Check for packet loss or connectivity failure
Step 2 – Interface & Link Check
- Verify NIC/interface status (UP/DOWN)
- Check for:
- Errors
- Packet drops
Step 3 – Switch Validation
- Verify switch port status
- Review switch logs for:
- Link flaps
- Port errors
- Shutdown events
Step 4 – Physical Verification
- Check cable connectivity or replace if required
- Verify SFP/transceiver status
- Replace faulty components if required
Step 5 – Isolation Testing
- Test with:
- Alternate switch port
- Known working cable/SFP
Step 6 – OEM Escalation (If required)
- If unresolved:
- Raise a case with OEM/vendor
- Share logs and troubleshooting details
Resolution
The issue is considered resolved when:
-
Network interface is UP
- Node is reachable (ping/SSH)
- No packet loss observed
- Alert is cleared from monitoring
Escalation
- L1 → L2: Immediate escalation for P1 alerts
- L2 → OEM/Vendor: If issue persists after troubleshooting
Related Articles
Network Down Extern (P1 Critical Alert) - Troubleshooting Guide
Purpose This document provides a standardized approach for handling and troubleshooting the Network Down - Extern alert, including L1 and L2 actions, escalation procedures, and resolution criteria for critical incidents. Alert Name Network Down - ...
Network Degraded (Boot) Alert - Troubleshooting Guide
Purpose This document provides a standardized approach for handling and troubleshooting the Network Degraded - Boot alert, enabling L1 and L2 engineers to ensure timely detection, prevent full network outage, and maintain service availability. Alert ...
Network Down IB Alert - Troubleshooting Guide
Purpose This document provides a standardized approach for handling and troubleshooting the Network Down – IB alert, including L1 and L2 actions, escalation process, and resolution criteria. Alert Name Network Down – IB Alert Description This alert ...
ZFS Pool Capacity Critical Alert - Troubleshooting Guide
Purpose This document provides a standardized approach for handling and troubleshooting the ZFS Pool Capacity Critical alert, enabling rapid response to prevent storage exhaustion, performance degradation, and potential service outages. Alert Name ...
PSU Error Alert - Troubleshooting Guide
Purpose This document provides a standardized approach for handling and troubleshooting the PSU Error alert, ensuring quick identification of power supply issues and maintaining system power redundancy. Alert Name PSU Error Alert Description This ...