Network Down IB Alert - Troubleshooting Guide

Network Down IB Alert - Troubleshooting Guide

Purpose

This document provides a standardized approach for handling and troubleshooting the Network Down – IB alert, including L1 and L2 actions, escalation process, and resolution criteria.

Alert Name

Network Down – IB

Alert Description

This alert is triggered when the InfiniBand (IB) interface on the server is down or unavailable.
It directly impacts:
  1. MPI (Message Passing Interface) communication
  2. Storage performance
  3. GPU-direct workloads

Severity

P2 – High Impact

Possible Causes

  1. InfiniBand interface down or misconfigured
  2. Faulty IB cable or loose connection
  3. Defective SFP/transceiver
  4. Switch port issues or misconfiguration
  5. IB fabric failure
  6. Hardware-level faults

L1 Engineer Actions

Step 1 – Log the Alert

  1. Log in to the affected node
  2. Validate alert summary and details

Step 2 – Create Ticket

  1. Monitor the alert for 20 minutes. 
  2. If issue persists:
    1. Create an internal ticket
    2. Include all observations and alert details

Step 3 – Escalation (If required)

  1. Inform L2 team immediately if unresolved
  2. Share all collected details

Step 4 – Documentation

After confirmation with L2, update Issue Master Sheet/ Issue Tracker Sheet.

L2 Engineer Actions

Step 1 – Verify InfiniBand Interface Status

ibstat
  1. Check IB interface status and link state
  2. Verify link speed and availability

Step 2 – Check IB Fabric Health

  • Validate overall InfiniBand fabric
  • Identify any node or link failures

Step 3 – Inspect Transceivers (SFPs)

  1. Check both- Server side and Switch side.
  2. Ensure proper seating and compatibility

Step 4 – Switch Side Validation (If Applicable)

  1. Perform port flapping- Shutdown/ No Shutdown
  2. Validate switch configuration for IB ports

Step 5 – Hardware Validation

  1.  Reseat or replace faulty components:
    1. IB Cable
    2. SFP / transceiver

Step 6 – OEM Escalation (If Required)

  1. If issue is not resolved:
    1. Raise a ticket with OEM/vendor
    2. Share logs and troubleshooting details

Resolution

The issue is considered resolved when:

  • IB interface is up and operational
  • Link status and speed are normal
  • Cluster communication is restored
  • Alert is cleared

Escalation

  • L1 → L2: If not resolved within 20 minutes
  • L2 → OEM/Vendor: If issue persists after troubleshooting