Purpose
Provide a standard troubleshooting and resolution procedure when filesystem capacity on a Head Node reaches critical levels.
When to Use
Use this KB when the monitoring system triggers the alert:
Filesystem Capacity Critical – Head Node
This alert indicates that one or more filesystems on the head node are reaching critical capacity and may cause service disruption if not addressed.
Alert Summary
The alert is generated when the filesystem usage on the head node exceeds the defined threshold.
High filesystem usage may lead to:
-
Application failures
-
Logging issues
-
Job scheduling failures
-
System instability
Severity
Critical
Immediate investigation is required to prevent system disruption.
L1 Engineer Actions
When the alert is received, perform the following checks:
Step 1 – Login to Head Node
SSH into the affected head node.
Step 2 – Wait and Monitor
- Wait 25–30 minutes
Step 3 – Escalation (If Required)
If the issue appears significant:
Step 4 – Documentation
After confirmation with L2:
L2 Engineer Actions
Step 1 – Verify Filesystem and Storage Devices
Run:
Check:
-
filesystem utilization
-
mounted storage
-
memory availability
Step 2 – Identify High Usage Filesystem
Locate the filesystem with critical usage.
Example:
Step 3 – Cleanup Space
Perform cleanup activities such as:
Example locations:
Step 4 – Validate After Cleanup
Verify filesystem usage again:
Ensure utilization drops below the alert threshold.
Step 5 – OEM Escalation (If Required)
If usage is caused by:
Escalate to OEM / platform support team.
Related Articles
KB 579212 - Filesystem Capacity Warning in Head Node - Troubleshooting Guide
Purpose Provide a troubleshooting and response procedure when a Filesystem Capacity Warning alert is triggered on a Head Node. This KB helps engineers quickly identify the issue, perform basic checks, and escalate when required. Alert Name Filesystem ...
KB 715377 - ZFS Pool Capacity Critical Alert - Troubleshooting Guide
Purpose This document provides a standardized approach for handling and troubleshooting the ZFS Pool Capacity Critical alert, enabling rapid response to prevent storage exhaustion, performance degradation, and potential service outages. Alert Name ...
KB 905833 - Network Down Extern (P1 Critical Alert) - Troubleshooting Guide
Purpose This document provides a standardized approach for handling and troubleshooting the Network Down - Extern alert, including L1 and L2 actions, escalation procedures, and resolution criteria for critical incidents. Alert Name Network Down - ...
KB 019872 - Memory Filesystem Capacity Critical -- Troubleshooting guide
Purpose Provide a standard troubleshooting and resolution procedure when filesystem capacity on any HGX Node reaches critical levels. When to Use Use this KB when the monitoring system triggers the alert: Memory Filesystem Capacity Critical – For Any ...
KB 265063 - Node Exporter Down - Troubleshooting Guide
Purpose This document provides a standardized approach for handling and troubleshooting the Node Exporter Down alert, including L1 and L2 actions, escalation flow, and resolution criteria. Alert Name Node Exporter Down Alert Description This alert is ...
Recent Articles
KB-630208 | How to use Find and Locate to search for files in Linux
Purpose To provide Linux administrators and Field Engineers with guidance on using the find and locate utilities to efficiently search for files and directories within a Linux filesystem based on criteria such as name, type, size, timestamps, ...
KB-630164 | How to use rsync to Synchronize Files
Purpose To provide a clear and practical procedure for using the rsync utility to synchronize files and directories between local systems and remote hosts, including push and pull operations, while preserving file attributes and optimizing data ...
KB-630107 | How to create a Linux swap file
Purpose To provide Field Engineers and Linux administrators with a standardized procedure for creating, configuring, and validating a Linux swap file. This procedure helps ensure sufficient virtual memory is available when physical RAM is exhausted ...
MSI All-in-One (AIO) PC – Physical Damage Policy for Warranty Claims
MSI All-in-One (AIO) PC – Physical Damage Policy for Warranty Claims MSI's standard warranty does not cover physical or accidental damage to an All-in-One (AIO) PC. If the damage is determined to be caused by external factors rather than a ...
KB 134833 - Troubleshooting NVSM Alert NV-CPU-XX – Unrecoverable CPU Internal Error
Purpose This document provides a general troubleshooting procedure for the NVIDIA System Management (NVSM) alert NV-CPU-XX, which indicates that a CPU has reported an internal error. The article outlines how to verify whether the alert represents an ...