Network Down Boot (P1 Critical Alert)- Troubleshooting Guide

Network Down Boot (P1 Critical Alert)- Troubleshooting Guide

Purpose

This document provides a standardized approach for handling and troubleshooting the Network Down - Boot alert, ensuring quick response to critical incidents and maintaining cluster control and system accessibility.

Alert Name

Network Down - Boot

Alert Description

This alert is triggered when the boot network interface on a node is down or unavailable.

This interface is critical for:

  1. Management communication
  2. Cluster control connectivity
  3. External access to the system

Loss of this interface can isolate nodes from cluster control and prevent user/system access.

Severity

P1 - Critical Impact

Possible Causes

  1. Network cable disconnected or faulty
  2. SFP/transceiver failure
  3. Switch port down or misconfigured
  4. Gateway unreachable
  5. NIC hardware issue

L1 Engineer Actions

Step 1 – Log the Alert

  1. Log the alert in the monitoring system
  2. Capture:
    1. Alert name
    2. Node/interface details
    3. Timestamp
    4. Current status

Step 2 – Create Ticket

  1. Monitor the alert for 15 minutes
  2. If the alert persists:
    1. Create an internal ticket
    2. Include all captured details

Step 3 – Escalation (If required)

  1. Inform L2 engineer immediately (P1 priority)
  2. Share complete alert information

Step 4 – Documentation

  1. After confirmation with L2:
    1. Update Issue Master Sheet/ Issue Tracker Sheet

L2 Engineer Actions

Step 1 – Network Reachability

  1. Ping the default gateway from the affected node
  2. Check for packet loss or connectivity failure
  1. Verify NIC/interface status (UP/DOWN)
  2. Check for:
    1. Errors
    2. Packet drops

Step 3 – Switch Validation

  1. Verify switch port status
  2. Review switch logs for:
    1. Link flaps
    2. Port errors
    3. Shutdown events

Step 4 – Physical Verification

  1. Check cable connectivity or replace if required
  2. Verify SFP/transceiver status
  3. Replace faulty components if required

Step 5 – Isolation Testing

  1. Test with:
    1. Alternate switch port
    2. Known working cable/SFP

Step 6 – OEM Escalation (If required)

  1. If unresolved:
    1. Raise a case with OEM/vendor
    2. Share logs and troubleshooting details

Resolution

The issue is considered resolved when:
  1. Network interface is UP
  2. Node is reachable (ping/SSH)
  3. No packet loss observed
  4. Alert is cleared from monitoring

Escalation

  1. L1 → L2: Immediate escalation for P1 alerts
  2. L2 → OEM/Vendor: If issue persists after troubleshooting
    • Related Articles

    • Network Down Extern (P1 Critical Alert) - Troubleshooting Guide

      Purpose This document provides a standardized approach for handling and troubleshooting the Network Down - Extern alert, including L1 and L2 actions, escalation procedures, and resolution criteria for critical incidents. Alert Name Network Down - ...
    • Network Degraded (Boot) Alert - Troubleshooting Guide

      Purpose This document provides a standardized approach for handling and troubleshooting the Network Degraded - Boot alert, enabling L1 and L2 engineers to ensure timely detection, prevent full network outage, and maintain service availability. Alert ...
    • Network Down IB Alert - Troubleshooting Guide

      Purpose This document provides a standardized approach for handling and troubleshooting the Network Down – IB alert, including L1 and L2 actions, escalation process, and resolution criteria. Alert Name Network Down – IB Alert Description This alert ...
    • ZFS Pool Capacity Critical Alert - Troubleshooting Guide

      Purpose This document provides a standardized approach for handling and troubleshooting the ZFS Pool Capacity Critical alert, enabling rapid response to prevent storage exhaustion, performance degradation, and potential service outages. Alert Name ...
    • PSU Error Alert - Troubleshooting Guide

      Purpose This document provides a standardized approach for handling and troubleshooting the PSU Error alert, ensuring quick identification of power supply issues and maintaining system power redundancy. Alert Name PSU Error Alert Description This ...