Network Speed IB Alert - Troubleshooting Guide

Network Speed IB Alert - Troubleshooting Guide

Purpose

This document provides a standardized approach for handling and troubleshooting the Network Speed – IB alert, including L1 and L2 actions, escalation process, and resolution guidelines.

Alert Name

Network Speed IB

Alert Description

This alert is triggered when the InfiniBand (IB) link operates below the expected bandwidth. It indicates a mismatch or degradation in IB link speed on the head node, which can impact high-performance computing (HPC) environments.

Severity

P2 – High Impact

Possible Causes

  1. InfiniBand link speed mismatch
  2. Faulty or incompatible IB cables
  3. Defective SFP/transceiver modules
  4. Switch port configuration issues
  5. Hardware limitations or degradation
  6. Issues in IB fabric health

L1 Engineer Actions

Step 1 – Log the Alert

  1. Log in to the affected node
  2. Review alert summary and details

Step 2 – Create Ticket

  1. Monitor the alert for 20 minutes for auto-resolution
  2. If issue persists:
    1. Create an internal ticket
    2. Include observations and alert details

Step 3 – Escalation (If required)

  1. Inform L2 team after 20 minutes if unresolved
  2. Share all relevant findings

Step 4 – Documentation

After confirmation with L2, update Issue Master Sheet/ Issue Tracker Sheet.

L2 Engineer Actions

ibstat

  1. Check current IB link speed
  2. Compare with expected bandwidth

Step 2 – Check IB Fabric Health

  1. Validate overall InfiniBand fabric status
  2. Identify any inconsistencies or degraded links

Step 3 – Inspect Transceivers (SFPs)

  1. Check on both- Server side and Switch side
  2. Ensure proper seating and compatibility

Step 4 – Switch Side Validation (If Applicable)

  1. Perform port flapping- Shutdown/ No Shutdown
  2. Validate switch configuration for IB ports

Step 5 – Hardware Validation

  1. Reseat or replace faulty components:
    1. IB Cable
    2. SFP/ Transceiver

Step 6 – OEM Escalation (If Required)

  1. If issue persists after troubleshooting:
    1. Raise a ticket with OEM/vendor
    2. Share observations and test results

Resolution

The issue is considered resolved when:
  1. IB link speed matches expected hardware specifications
  2. No degradation observed in IB performance
  3. Cluster communication latency is normal
  4. Alert is cleared

Escalation

  • L1 → L2: If not resolved within 20 minutes
  • L2 → OEM/Vendor: If hardware or configuration issue persists
      • Related Articles

      • Network Down IB Alert - Troubleshooting Guide

        Purpose This document provides a standardized approach for handling and troubleshooting the Network Down – IB alert, including L1 and L2 actions, escalation process, and resolution criteria. Alert Name Network Down – IB Alert Description This alert ...
      • Network Down Extern (P1 Critical Alert) - Troubleshooting Guide

        Purpose This document provides a standardized approach for handling and troubleshooting the Network Down - Extern alert, including L1 and L2 actions, escalation procedures, and resolution criteria for critical incidents. Alert Name Network Down - ...
      • Network Down Boot (P1 Critical Alert)- Troubleshooting Guide

        Purpose This document provides a standardized approach for handling and troubleshooting the Network Down - Boot alert, ensuring quick response to critical incidents and maintaining cluster control and system accessibility. Alert Name Network Down - ...
      • ZFS Pool Capacity Warning Alert - Troubleshooting Guide

        Purpose This document provides a standardized approach for handling and troubleshooting the ZFS Pool Capacity Warning alert, helping L1 and L2 engineers prevent storage exhaustion and maintain system stability. Alert Name ZFS Pool Capacity Warning ...
      • Mapping mlx5_x Names to Actual InfiniBand Ports

        Purpose In large GPU/AI clusters, each compute node often has multiple InfiniBand (IB) Host Channel Adapters (HCAs). On Linux, these HCAs appear as mlx5_x devices. To troubleshoot connectivity or performance issues, it is essential to map: mlx5_x ...