Network Speed IB Alert - Troubleshooting Guide
Purpose
This document provides a standardized approach for handling and troubleshooting the Network Speed – IB alert, including L1 and L2 actions, escalation process, and resolution guidelines.
Alert Name
Network Speed IB
Alert Description
This alert is triggered when the InfiniBand (IB) link operates below the expected bandwidth. It indicates a mismatch or degradation in IB link speed on the head node, which can impact high-performance computing (HPC) environments.
Severity
P2 – High Impact
Possible Causes
- InfiniBand link speed mismatch
- Faulty or incompatible IB cables
- Defective SFP/transceiver modules
- Switch port configuration issues
- Hardware limitations or degradation
- Issues in IB fabric health
L1 Engineer Actions
Step 1 – Log the Alert
- Log in to the affected node
- Review alert summary and details
Step 2 – Create Ticket
- Monitor the alert for 20 minutes for auto-resolution
- If issue persists:
- Create an internal ticket
- Include observations and alert details
Step 3 – Escalation (If required)
- Inform L2 team after 20 minutes if unresolved
- Share all relevant findings
Step 4 – Documentation
After confirmation with L2, update Issue Master Sheet/ Issue Tracker Sheet.
L2 Engineer Actions
Step 1 – Verify InfiniBand Link Speed
ibstat
- Check current IB link speed
- Compare with expected bandwidth
Step 2 – Check IB Fabric Health
- Validate overall InfiniBand fabric status
- Identify any inconsistencies or degraded links
Step 3 – Inspect Transceivers (SFPs)
- Check on both- Server side and Switch side
- Ensure proper seating and compatibility
Step 4 – Switch Side Validation (If Applicable)
- Perform port flapping- Shutdown/ No Shutdown
- Validate switch configuration for IB ports
Step 5 – Hardware Validation
- Reseat or replace faulty components:
- IB Cable
- SFP/ Transceiver
Step 6 – OEM Escalation (If Required)
- If issue persists after troubleshooting:
- Raise a ticket with OEM/vendor
- Share observations and test results
Resolution
The issue is considered resolved when:
- IB link speed matches expected hardware specifications
- No degradation observed in IB performance
- Cluster communication latency is normal
- Alert is cleared
Escalation
L1 → L2: If not resolved within 20 minutes
L2 → OEM/Vendor: If hardware or configuration issue persists
Related Articles
Network Down IB Alert - Troubleshooting Guide
Purpose This document provides a standardized approach for handling and troubleshooting the Network Down – IB alert, including L1 and L2 actions, escalation process, and resolution criteria. Alert Name Network Down – IB Alert Description This alert ...
Network Down Extern (P1 Critical Alert) - Troubleshooting Guide
Purpose This document provides a standardized approach for handling and troubleshooting the Network Down - Extern alert, including L1 and L2 actions, escalation procedures, and resolution criteria for critical incidents. Alert Name Network Down - ...
Network Down Boot (P1 Critical Alert)- Troubleshooting Guide
Purpose This document provides a standardized approach for handling and troubleshooting the Network Down - Boot alert, ensuring quick response to critical incidents and maintaining cluster control and system accessibility. Alert Name Network Down - ...
ZFS Pool Capacity Warning Alert - Troubleshooting Guide
Purpose This document provides a standardized approach for handling and troubleshooting the ZFS Pool Capacity Warning alert, helping L1 and L2 engineers prevent storage exhaustion and maintain system stability. Alert Name ZFS Pool Capacity Warning ...
Mapping mlx5_x Names to Actual InfiniBand Ports
Purpose In large GPU/AI clusters, each compute node often has multiple InfiniBand (IB) Host Channel Adapters (HCAs). On Linux, these HCAs appear as mlx5_x devices. To troubleshoot connectivity or performance issues, it is essential to map: mlx5_x ...