This document provides a standardized approach for handling and troubleshooting the Network Down - Boot alert, ensuring quick response to critical incidents and maintaining cluster control and system accessibility.
This alert is triggered when the boot network interface on a node is down or unavailable.
This interface is critical for:
-
Management communication
- Cluster control connectivity
- External access to the system
Loss of this interface can isolate nodes from cluster control and prevent user/system access.
Severity
P1 - Critical Impact
Possible Causes
- Network cable disconnected or faulty
- SFP/transceiver failure
- Switch port down or misconfigured
- Gateway unreachable
- NIC hardware issue
L1 Engineer Actions
Step 1 – Log the Alert
- Log the alert in the monitoring system
- Capture:
- Alert name
- Node/interface details
- Timestamp
- Current status
Step 2 – Create Ticket
- Monitor the alert for 15 minutes
- If the alert persists:
- Create an internal ticket
- Include all captured details
Step 3 – Escalation (If required)
- Inform L2 engineer immediately (P1 priority)
- Share complete alert information
Step 4 – Documentation
- After confirmation with L2:
- Update Issue Master Sheet/ Issue Tracker Sheet
L2 Engineer Actions
Step 1 – Network Reachability
- Ping the default gateway from the affected node
- Check for packet loss or connectivity failure
Step 2 – Interface & Link Check
- Verify NIC/interface status (UP/DOWN)
- Check for:
- Errors
- Packet drops
Step 3 – Switch Validation
- Verify switch port status
- Review switch logs for:
- Link flaps
- Port errors
- Shutdown events
Step 4 – Physical Verification
- Check cable connectivity or replace if required
- Verify SFP/transceiver status
- Replace faulty components if required
Step 5 – Isolation Testing
- Test with:
- Alternate switch port
- Known working cable/SFP
Step 6 – OEM Escalation (If required)
- If unresolved:
- Raise a case with OEM/vendor
- Share logs and troubleshooting details
Resolution
The issue is considered resolved when:
-
Network interface is UP
- Node is reachable (ping/SSH)
- No packet loss observed
- Alert is cleared from monitoring
Escalation
- L1 → L2: Immediate escalation for P1 alerts
- L2 → OEM/Vendor: If issue persists after troubleshooting