Purpose
This document provides a standardized approach for handling and troubleshooting the Network Down - Extern alert, including L1 and L2 actions, escalation procedures, and resolution criteria for critical incidents.
Alert Name
Network Down - Extern
Alert Description
This alert is triggered when the external network interface on a node is down or unavailable.
This interface is critical for:
- External communication and access
- User/application connectivity
- Cluster interaction with external systems
- Failure of this interface can isolate nodes from users, external services, and partially from cluster operations.
Severity
P1 – Critical Impact
Possible Causes
- Faulty or disconnected network cable
- SFP/transceiver issue
- Switch port failure or shutdown
- Gateway or network path unreachable
- Network service failure on server
- NIC hardware issue
L1 Engineer Actions
Step 1 – Log the Alert
Log the alert in the monitoring tool. Capture the parameters as below:
- Alert name
- Node/interface details
- Timestamp
- Current status
Step 2 – Create Ticket
- Monitor the alert for 15 minutes
- If the alert persists:
- Create an internal ticket
- Include all captured details
Step 3 – Escalation (If required)
- Inform L2 engineer immediately (P1 priority)
- Share complete alert information
Step 4 – Documentation
After confirmation with L2, update the Issue Master Sheet/ Issue Tracking Sheet.
L2 Engineer Actions
Step 1 – Network Reachability
- Ping the default gateway from the affected node
- Identify packet loss or connectivity failure
Step 2 – Interface & Link Status
- Check NIC/interface status (UP/DOWN)
- Review for errors and packet drops.
Step 3 – Switch-Level Checks
- Verify switch port status
- Analyze switch logs for:
- Link flaps
- Errors
- Port shutdown events
Step 4 – Physical Layer Validation
- Verify cable connections
- Check SFP/transceiver health
- Replace faulty components if required
Step 5 – Network Service Validation
- Restart network service on the affected node
- Recheck connectivity after restart
Step 6 – OEM Escalation (If required)
- If issue persists:
- Raise a case with OEM/vendor
- Share logs, actions performed, and timestamps
Resolution
The issue is considered resolved when:
- External interface status is UP
- Node is reachable externally
- No packet loss is observed
- Alert is cleared in the monitoring system
Escalation Matrix
Level | Responsibility |
L1 | Monitoring & logging |
L2 | Troubleshooting & resolution |
OEM/Vendor | Advanced/network or hardware issues |
L1 → L2: Immediate escalation for P1 alerts
L2 → OEM/Vendor: If issue persists after troubleshooting