KB 510520 - BeeGFS Capacity Critical – Troubleshooting Guide

KB 510520 - BeeGFS Capacity Critical – Troubleshooting Guide

Purpose

Provide a troubleshooting reference for engineers when a BeeGFS Capacity Critical alert is triggered in the cluster.

This document outlines the meaning of the alert and the actions required by L1 and L2 engineers.


When to Use

Use this KB article when the monitoring system generates the following alert:

Alert Name: BeeGFS Capacity Critical


What the Alert Means

This alert indicates that BeeGFS storage capacity has reached a critical threshold.

This may occur due to:

  • High storage consumption

  • Large files written to the filesystem

  • Stuck or deleted files consuming space

  • Storage servers (OSS) being offline

  • Metadata or storage thread bottlenecks

If not addressed, this may result in:

  • Job failures

  • Storage write failures

  • Performance degradation


Severity

Critical

Immediate investigation is recommended to prevent storage exhaustion.


L1 Engineer Actions

Follow the steps below when the alert is received.

Step 1 – Acknowledge the Alert

  • Record the alert in the Alert Monitoring System

Step 2 – Verify Alert Details

Check the following:

  • Affected filesystem

  • Node generating the alert

  • Current storage usage

Step 3 – Initial Monitoring

  • Confirm whether the usage spike is temporary

  • Monitor storage usage for a short duration

Step 4 – Escalate if Required

If the alert persists:

  • Escalate to L2 Engineer

  • Provide alert details and timestamps


L2 Engineer Actions

Perform deeper investigation once the alert is escalated.

Step 1 – Check OSS Server Status

Verify whether any OSS server is down or offline.

Step 2 – Verify Metadata / Storage Threads

Check for issues with:

  • Metadata service

  • Storage threads

Step 3 – Check Filesystem Usage

Run the following command:

df -h

Check usage on the BeeGFS mount point.

Step 4 – Identify Large Directories

Locate large directories consuming space.

du -sh *

Step 5 – Check for Stuck or Deleted Files

Investigate files that may still be consuming storage after deletion.

Step 6 – Coordinate with Storage Team

If required:

  • Coordinate with NetApp team

  • Investigate backend storage issues

Step 7 – Remediation

Take corrective actions such as:

  • Cleaning unused data

  • Moving data

  • Expanding storage capacity


Expected Outcome

After corrective action:

  • BeeGFS storage usage should fall below the critical threshold

  • Alert should automatically clear in the monitoring system

    • Recent Articles

    • KB-630208 | How to use Find and Locate to search for files in Linux

      Purpose To provide Linux administrators and Field Engineers with guidance on using the find and locate utilities to efficiently search for files and directories within a Linux filesystem based on criteria such as name, type, size, timestamps, ...
    • KB-630164 | How to use rsync to Synchronize Files

      Purpose To provide a clear and practical procedure for using the rsync utility to synchronize files and directories between local systems and remote hosts, including push and pull operations, while preserving file attributes and optimizing data ...
    • KB-630107 | How to create a Linux swap file

      Purpose To provide Field Engineers and Linux administrators with a standardized procedure for creating, configuring, and validating a Linux swap file. This procedure helps ensure sufficient virtual memory is available when physical RAM is exhausted ...
    • MSI All-in-One (AIO) PC – Physical Damage Policy for Warranty Claims

      MSI All-in-One (AIO) PC – Physical Damage Policy for Warranty Claims MSI's standard warranty does not cover physical or accidental damage to an All-in-One (AIO) PC. If the damage is determined to be caused by external factors rather than a ...
    • KB 134833 - Troubleshooting NVSM Alert NV-CPU-XX – Unrecoverable CPU Internal Error

      Purpose This document provides a general troubleshooting procedure for the NVIDIA System Management (NVSM) alert NV-CPU-XX, which indicates that a CPU has reported an internal error. The article outlines how to verify whether the alert represents an ...
    • Popular Articles

    • CP Plus Camera and NVR Configuration

      NVR Configuration The CP Plus Pro Series of NVRs have been meticulously designed for providing you with upgraded performance and higher recording quality in your IP video surveillance solution. The robust processor that has been inculcated in this ...
    • KB 775235 - Kerberos Authentication – Overview

      What is Kerberos? Kerberos is a secure authentication method used in our Active Directory (AD) environment (mbuzztech.com). It allows users to: Access multiple systems without re-entering passwords (Single Sign-On – SSO) Log in once Where We Use It ...
    • How to Remove and Reinstall NVIDIA Drivers on Ubuntu

      This article provides step by step guide to completely remove existing NVIDIA drivers and reinstall specific version of the NVIDIA driver on the Ubuntu system Prerequisites Administrative (sudo) access to the Ubuntu system. Internet access to ...
    • KB 692001 - Personal Computers and Servers - Classification and Point of Contact

      We can classify the computers that MBUZZ handles based on their form-factor as below: Tower Workstations, Desktops, Gaming PCs and SFF (Small form factor) PCs fall under this category. These are computers people would use on a desk and rarely move. ...
    • KB 298031 - M.2 SSD Tier List

      The sequential read and write speeds, which are usually the most advertised number, are not a proper benchmark of real-world performance or the quality of an SSD. This article categorizes and tiers SSDs based on factors like the type of NAND flash, ...