Network Down IB Alert - Troubleshooting Guide
Purpose
This document provides a standardized approach for handling and troubleshooting the Network Down – IB alert, including L1 and L2 actions, escalation process, and resolution criteria.
Alert Name
Network Down – IB
Alert Description
This alert is triggered when the InfiniBand (IB) interface on the server is down or unavailable.
It directly impacts:
- MPI (Message Passing Interface) communication
- Storage performance
- GPU-direct workloads
Severity
P2 – High Impact
Possible Causes
- InfiniBand interface down or misconfigured
- Faulty IB cable or loose connection
- Defective SFP/transceiver
- Switch port issues or misconfiguration
- IB fabric failure
- Hardware-level faults
L1 Engineer Actions
Step 1 – Log the Alert
- Log in to the affected node
- Validate alert summary and details
Step 2 – Create Ticket
- Monitor the alert for 20 minutes.
- If issue persists:
- Create an internal ticket
- Include all observations and alert details
Step 3 – Escalation (If required)
- Inform L2 team immediately if unresolved
- Share all collected details
Step 4 – Documentation
After confirmation with L2, update Issue Master Sheet/ Issue Tracker Sheet.
L2 Engineer Actions
Step 1 – Verify InfiniBand Interface Status
ibstat
- Check IB interface status and link state
- Verify link speed and availability
Step 2 – Check IB Fabric Health
- Validate overall InfiniBand fabric
-
Identify any node or link failures
Step 3 – Inspect Transceivers (SFPs)
- Check both- Server side and Switch side.
- Ensure proper seating and compatibility
Step 4 – Switch Side Validation (If Applicable)
- Perform port flapping- Shutdown/ No Shutdown
- Validate switch configuration for IB ports
Step 5 – Hardware Validation
- Reseat or replace faulty components:
- IB Cable
- SFP / transceiver
Step 6 – OEM Escalation (If Required)
- If issue is not resolved:
- Raise a ticket with OEM/vendor
- Share logs and troubleshooting details
Resolution
The issue is considered resolved when:
-
IB interface is up and operational
-
Link status and speed are normal
-
Cluster communication is restored
-
Alert is cleared
Escalation
L1 → L2: If not resolved within 20 minutes
L2 → OEM/Vendor: If issue persists after troubleshooting
Related Articles
Network Speed IB Alert - Troubleshooting Guide
Purpose This document provides a standardized approach for handling and troubleshooting the Network Speed – IB alert, including L1 and L2 actions, escalation process, and resolution guidelines. Alert Name Network Speed IB Alert Description This alert ...
GPU Memory Mismatch – Troubleshooting Guide
Alert Information Alert Name: GPU Memory Mismatch Severity: P3 – Medium Impact Alert Description: This alert is generated when the detected GPU memory size on a node does not match the expected configuration. Alert Summary The monitoring system has ...
Network Down Extern (P1 Critical Alert) - Troubleshooting Guide
Purpose This document provides a standardized approach for handling and troubleshooting the Network Down - Extern alert, including L1 and L2 actions, escalation procedures, and resolution criteria for critical incidents. Alert Name Network Down - ...
Network Down Boot (P1 Critical Alert)- Troubleshooting Guide
Purpose This document provides a standardized approach for handling and troubleshooting the Network Down - Boot alert, ensuring quick response to critical incidents and maintaining cluster control and system accessibility. Alert Name Network Down - ...
Execution of NVIDIA Field Diagnostic (FD) Tool and Collection of Diagnostic Logs, and troubleshooting of GPU(s) not detected
1. Objective To provide a standardized procedure for executing the NVIDIA Field Diagnostic (FD) tool, verifying GPU status, collecting the required diagnostic logs, and documenting the findings for further analysis. 2. Scope This procedure applies to ...