Network Down Extern (P1 Critical Alert) - Troubleshooting Guide

KB 905833 - Network Down Extern (P1 Critical Alert) - Troubleshooting Guide

Purpose

This document provides a standardized approach for handling and troubleshooting the Network Down - Extern alert, including L1 and L2 actions, escalation procedures, and resolution criteria for critical incidents.

Alert Name

Network Down - Extern

Alert Description

This alert is triggered when the external network interface on a node is down or unavailable.

This interface is critical for:
  1. External communication and access
  2. User/application connectivity
  3. Cluster interaction with external systems
  4. Failure of this interface can isolate nodes from users, external services, and partially from cluster operations.

Severity

P1 – Critical Impact

Possible Causes

  1. Faulty or disconnected network cable
  2. SFP/transceiver issue
  3. Switch port failure or shutdown
  4. Gateway or network path unreachable
  5. Network service failure on server
  6. NIC hardware issue

L1 Engineer Actions

Step 1 – Log the Alert

Log the alert in the monitoring tool. Capture the parameters as below:
  1. Alert name
  2. Node/interface details
  3. Timestamp
  4. Current status

Step 2 – Create Ticket

  1. Monitor the alert for 15 minutes
  2. If the alert persists:
    1. Create an internal ticket
    2. Include all captured details

Step 3 – Escalation (If required)

  1. Inform L2 engineer immediately (P1 priority)
  2. Share complete alert information

Step 4 – Documentation

After confirmation with L2, update the Issue Master Sheet/ Issue Tracking Sheet.

L2 Engineer Actions

Step 1 – Network Reachability

  1. Ping the default gateway from the affected node
  2. Identify packet loss or connectivity failure
  1. Check NIC/interface status (UP/DOWN)
  2. Review for errors and packet drops.

Step 3 – Switch-Level Checks

  1. Verify switch port status
  2. Analyze switch logs for:
    1. Link flaps
    2. Errors
    3. Port shutdown events

Step 4 – Physical Layer Validation

  1. Verify cable connections
  2. Check SFP/transceiver health
  3. Replace faulty components if required

Step 5 – Network Service Validation

  1. Restart network service on the affected node
  2. Recheck connectivity after restart

Step 6 – OEM Escalation (If required)

  1. If issue persists:
    1. Raise a case with OEM/vendor
    2. Share logs, actions performed, and timestamps

Resolution

The issue is considered resolved when:
  1. External interface status is UP
  2. Node is reachable externally
  3. No packet loss is observed
  4. Alert is cleared in the monitoring system

Escalation Matrix

Level

Responsibility

L1

Monitoring & logging

L2

Troubleshooting & resolution

OEM/Vendor

Advanced/network or hardware issues


  • L1 → L2: Immediate escalation for P1 alerts
  • L2 → OEM/Vendor: If issue persists after troubleshooting
      • Related Articles

      • KB 905774 - Network Down Boot (P1 Critical Alert)- Troubleshooting Guide

        Purpose This document provides a standardized approach for handling and troubleshooting the Network Down - Boot alert, ensuring quick response to critical incidents and maintaining cluster control and system accessibility. Alert Name Network Down - ...
      • KB 715377 - ZFS Pool Capacity Critical Alert - Troubleshooting Guide

        Purpose This document provides a standardized approach for handling and troubleshooting the ZFS Pool Capacity Critical alert, enabling rapid response to prevent storage exhaustion, performance degradation, and potential service outages. Alert Name ...
      • KB 021137 - Network Down IB Alert - Troubleshooting Guide

        Purpose This document provides a standardized approach for handling and troubleshooting the Network Down – IB alert, including L1 and L2 actions, escalation process, and resolution criteria. Alert Name Network Down – IB Alert Description This alert ...
      • KB 717596 - PSU Error Alert - Troubleshooting Guide

        Purpose This document provides a standardized approach for handling and troubleshooting the PSU Error alert, ensuring quick identification of power supply issues and maintaining system power redundancy. Alert Name PSU Error Alert Description This ...
      • KB 021273 - Network Speed IB Alert - Troubleshooting Guide

        Purpose This document provides a standardized approach for handling and troubleshooting the Network Speed – IB alert, including L1 and L2 actions, escalation process, and resolution guidelines. Alert Name Network Speed IB Alert Description This alert ...
      • Recent Articles

      • KB-630208 | How to use Find and Locate to search for files in Linux

        Purpose To provide Linux administrators and Field Engineers with guidance on using the find and locate utilities to efficiently search for files and directories within a Linux filesystem based on criteria such as name, type, size, timestamps, ...
      • KB-630164 | How to use rsync to Synchronize Files

        Purpose To provide a clear and practical procedure for using the rsync utility to synchronize files and directories between local systems and remote hosts, including push and pull operations, while preserving file attributes and optimizing data ...
      • KB-630107 | How to create a Linux swap file

        Purpose To provide Field Engineers and Linux administrators with a standardized procedure for creating, configuring, and validating a Linux swap file. This procedure helps ensure sufficient virtual memory is available when physical RAM is exhausted ...
      • MSI All-in-One (AIO) PC – Physical Damage Policy for Warranty Claims

        MSI All-in-One (AIO) PC – Physical Damage Policy for Warranty Claims MSI's standard warranty does not cover physical or accidental damage to an All-in-One (AIO) PC. If the damage is determined to be caused by external factors rather than a ...
      • KB 134833 - Troubleshooting NVSM Alert NV-CPU-XX – Unrecoverable CPU Internal Error

        Purpose This document provides a general troubleshooting procedure for the NVIDIA System Management (NVSM) alert NV-CPU-XX, which indicates that a CPU has reported an internal error. The article outlines how to verify whether the alert represents an ...
      • Popular Articles

      • CP Plus Camera and NVR Configuration

        NVR Configuration The CP Plus Pro Series of NVRs have been meticulously designed for providing you with upgraded performance and higher recording quality in your IP video surveillance solution. The robust processor that has been inculcated in this ...
      • KB 775235 - Kerberos Authentication – Overview

        What is Kerberos? Kerberos is a secure authentication method used in our Active Directory (AD) environment (mbuzztech.com). It allows users to: Access multiple systems without re-entering passwords (Single Sign-On – SSO) Log in once Where We Use It ...
      • How to Remove and Reinstall NVIDIA Drivers on Ubuntu

        This article provides step by step guide to completely remove existing NVIDIA drivers and reinstall specific version of the NVIDIA driver on the Ubuntu system Prerequisites Administrative (sudo) access to the Ubuntu system. Internet access to ...
      • KB 692001 - Personal Computers and Servers - Classification and Point of Contact

        We can classify the computers that MBUZZ handles based on their form-factor as below: Tower Workstations, Desktops, Gaming PCs and SFF (Small form factor) PCs fall under this category. These are computers people would use on a desk and rarely move. ...
      • KB 298031 - M.2 SSD Tier List

        The sequential read and write speeds, which are usually the most advertised number, are not a proper benchmark of real-world performance or the quality of an SSD. This article categorizes and tiers SSDs based on factors like the type of NAND flash, ...