Network Degraded (Boot) Alert - Troubleshooting Guide

KB 905711 - Network Degraded (Boot) Alert - Troubleshooting Guide

Purpose

This document provides a standardized approach for handling and troubleshooting the Network Degraded - Boot alert, enabling L1 and L2 engineers to ensure timely detection, prevent full network outage, and maintain service availability.

Alert Name

Network Degraded - Boot

Alert Description

This alert is triggered when one slave interface in a bonded network configuration goes down.

It indicates that the bonded interface has lost redundancy, leaving the system with a single active network path

Severity

P2 – High Impact

Possible Causes

  1. Network cable disconnection or fault
  2. Faulty NIC/interface
  3. Switch port issues or shutdown
  4. Link flapping or instability
  5. Configuration mismatch in bonding
  6. Hardware or transceiver issues

L1 Engineer Actions

Step 1 – Log the Alert

  1. Log in to the monitoring system
  2. Review:
    1. Alert summary
    2. Affected host
    3. Interface details

Step 2 – Create Ticket

  1. Monitor the alert for 15 minutes to check if transient
  2. If issue persists:
    1. Create an internal ticket
    2. Include:
    3. Hostname
    4. Interface details
    5. Timestamp
    6. Observations

Step 3 – Escalation (If required)

  1. Inform L2 team if alert persists beyond 20 minutes
  2. Share all relevant details

Step 4 – Documentation

  1. After confirmation with L2:
    1. Update Issue Master Sheet/ Issue Tracking Sheet

L2 Engineer Actions

Step 1 – Check Bond Status

cat /proc/net/bonding/*
  1. Identify:
    1. Active and failed slave interfaces
    2. Link status (UP/DOWN)
    3. MII status

Step 2 – Verify Network Interface

ip link show
  1. Check interface state
  2. Look for down or unstable interfaces

Step 3 – Switch Port Validation

  1. Verify corresponding switch port:
    1. Link status (UP/DOWN)
    2. Errors or flapping
    3. Coordinate with network team if required

Step 4 – Physical Inspection

  1. Check and reseat network cable and trans receiver.
  2. Replace cable if faulty

Step 5 – OEM Escalation (If required)

  1. If issue persists:
    1. Raise ticket with OEM
    2. Share logs and observations

Resolution

The issue is considered resolved when:

  • All bond slave interfaces are UP
  • Bonded interface regains redundancy
  • No link instability or flapping observed
  • Alert is cleared

Escalation

  1. L1 → L2: If not resolved within 20 minutes
  2. L2 → OEM/Vendor: If hardware/network issue persists