Node Exporter Down - Troubleshooting Guide

KB 265063 - Node Exporter Down - Troubleshooting Guide

Purpose

This document provides a standardized approach for handling and troubleshooting the Node Exporter Down alert, including L1 and L2 actions, escalation flow, and resolution criteria.

Alert Name

Node Exporter Down

Alert Description

This alert is triggered when the Node Exporter service is not responding.
Node Exporter is a system-level monitoring agent used with Prometheus to collect OS and hardware metrics from a Linux server (node).

Severity

P2 – High Impact

Possible Causes

  1. Node is down or unreachable (network issue)
  2. Node Exporter service crashed or stopped
  3. High system load or resource exhaustion
  4. Firewall or port blocking issues
  5. Prometheus scraping failure

L1 Engineer Actions

Step 1 – Log the Alert

  1. Review alert details:
    1. Timestamp
    2. Node name
    3. Alert source

Step 2 – Create Ticket

  1. Monitor the alert for 20 minutes to check auto-resolution
  2. If issue persists:
    1. Create an internal ticket
    2. Include all alert details

Step 3 – Escalation (If Required)

  1. Inform L2 Engineer after 20 minutes if unresolved
  2. Share all relevant details

Step 4 – Documentation

After confirmation with L2, update the Issues Master Sheet/ Issues Tracker Sheet.

L2 Engineer Actions

Step 1 – Check Node Reachability

ping <node_ip>
ssh <user>@<node_ip>
  1. Verify node is reachable over network.
  2. Confirm SSH access

Step 2 – Check Node Exporter Service Status

systemctl status node_exporter
  1. Verify if service is running

Step 3 – Restart Node Exporter Service

systemctl restart node_exporter

Step 4 – Verify Service Status

systemctl status node_exporter
  1. Ensure service is active

Step 5 – Check Port Accessibility (Default: 9100)

netstat -tulnp | grep 9100
  1. Confirm port is listening

Step 6 – Verify Monitoring Tools

  1. Check Prometheus: Verify target status (UP/DOWN)
  2. Check Grafana: Look for missing metrics

Step 7 – Verify System Health

top
df -h
free -m
  1. Check CPU, Memory, and Disk usage

Step 8 – Log Analysis & Escalation

journalctl -u node_exporter
  1. Investigate logs for errors.
  2. If issue persists:
    1. Inform customer / create external ticket

Step 9 – Reboot Node (If Required)

reboot
  1. Perform only if:
    1. Node is unresponsive
    2. Service repeatedly fails

Resolution

The issue is considered resolved when:
  1. Node is reachable
  2. Node Exporter service is running
  3. Metrics are visible in Prometheus/Grafana
  4. Alert is cleared

Escalation

  • L1 → L2: If not resolved within 20 minutes
  • L2 → OEM: If infrastructure issue persists

      • Recent Articles

      • KB-630208 | How to use Find and Locate to search for files in Linux

        Purpose To provide Linux administrators and Field Engineers with guidance on using the find and locate utilities to efficiently search for files and directories within a Linux filesystem based on criteria such as name, type, size, timestamps, ...
      • KB-630164 | How to use rsync to Synchronize Files

        Purpose To provide a clear and practical procedure for using the rsync utility to synchronize files and directories between local systems and remote hosts, including push and pull operations, while preserving file attributes and optimizing data ...
      • KB-630107 | How to create a Linux swap file

        Purpose To provide Field Engineers and Linux administrators with a standardized procedure for creating, configuring, and validating a Linux swap file. This procedure helps ensure sufficient virtual memory is available when physical RAM is exhausted ...
      • MSI All-in-One (AIO) PC – Physical Damage Policy for Warranty Claims

        MSI All-in-One (AIO) PC – Physical Damage Policy for Warranty Claims MSI's standard warranty does not cover physical or accidental damage to an All-in-One (AIO) PC. If the damage is determined to be caused by external factors rather than a ...
      • KB 134833 - Troubleshooting NVSM Alert NV-CPU-XX – Unrecoverable CPU Internal Error

        Purpose This document provides a general troubleshooting procedure for the NVIDIA System Management (NVSM) alert NV-CPU-XX, which indicates that a CPU has reported an internal error. The article outlines how to verify whether the alert represents an ...
      • Popular Articles

      • CP Plus Camera and NVR Configuration

        NVR Configuration The CP Plus Pro Series of NVRs have been meticulously designed for providing you with upgraded performance and higher recording quality in your IP video surveillance solution. The robust processor that has been inculcated in this ...
      • KB 775235 - Kerberos Authentication – Overview

        What is Kerberos? Kerberos is a secure authentication method used in our Active Directory (AD) environment (mbuzztech.com). It allows users to: Access multiple systems without re-entering passwords (Single Sign-On – SSO) Log in once Where We Use It ...
      • How to Remove and Reinstall NVIDIA Drivers on Ubuntu

        This article provides step by step guide to completely remove existing NVIDIA drivers and reinstall specific version of the NVIDIA driver on the Ubuntu system Prerequisites Administrative (sudo) access to the Ubuntu system. Internet access to ...
      • KB 692001 - Personal Computers and Servers - Classification and Point of Contact

        We can classify the computers that MBUZZ handles based on their form-factor as below: Tower Workstations, Desktops, Gaming PCs and SFF (Small form factor) PCs fall under this category. These are computers people would use on a desk and rarely move. ...
      • KB 298031 - M.2 SSD Tier List

        The sequential read and write speeds, which are usually the most advertised number, are not a proper benchmark of real-world performance or the quality of an SSD. This article categorizes and tiers SSDs based on factors like the type of NAND flash, ...