Network Down Boot (P1 Critical Alert)- Troubleshooting Guide

KB 905774 - Network Down Boot (P1 Critical Alert)- Troubleshooting Guide

Purpose

This document provides a standardized approach for handling and troubleshooting the Network Down - Boot alert, ensuring quick response to critical incidents and maintaining cluster control and system accessibility.

Alert Name

Network Down - Boot

Alert Description

This alert is triggered when the boot network interface on a node is down or unavailable.

This interface is critical for:

  1. Management communication
  2. Cluster control connectivity
  3. External access to the system

Loss of this interface can isolate nodes from cluster control and prevent user/system access.

Severity

P1 - Critical Impact

Possible Causes

  1. Network cable disconnected or faulty
  2. SFP/transceiver failure
  3. Switch port down or misconfigured
  4. Gateway unreachable
  5. NIC hardware issue

L1 Engineer Actions

Step 1 – Log the Alert

  1. Log the alert in the monitoring system
  2. Capture:
    1. Alert name
    2. Node/interface details
    3. Timestamp
    4. Current status

Step 2 – Create Ticket

  1. Monitor the alert for 15 minutes
  2. If the alert persists:
    1. Create an internal ticket
    2. Include all captured details

Step 3 – Escalation (If required)

  1. Inform L2 engineer immediately (P1 priority)
  2. Share complete alert information

Step 4 – Documentation

  1. After confirmation with L2:
    1. Update Issue Master Sheet/ Issue Tracker Sheet

L2 Engineer Actions

Step 1 – Network Reachability

  1. Ping the default gateway from the affected node
  2. Check for packet loss or connectivity failure
  1. Verify NIC/interface status (UP/DOWN)
  2. Check for:
    1. Errors
    2. Packet drops

Step 3 – Switch Validation

  1. Verify switch port status
  2. Review switch logs for:
    1. Link flaps
    2. Port errors
    3. Shutdown events

Step 4 – Physical Verification

  1. Check cable connectivity or replace if required
  2. Verify SFP/transceiver status
  3. Replace faulty components if required

Step 5 – Isolation Testing

  1. Test with:
    1. Alternate switch port
    2. Known working cable/SFP

Step 6 – OEM Escalation (If required)

  1. If unresolved:
    1. Raise a case with OEM/vendor
    2. Share logs and troubleshooting details

Resolution

The issue is considered resolved when:
  1. Network interface is UP
  2. Node is reachable (ping/SSH)
  3. No packet loss observed
  4. Alert is cleared from monitoring

Escalation

  1. L1 → L2: Immediate escalation for P1 alerts
  2. L2 → OEM/Vendor: If issue persists after troubleshooting