BeeGFS Capacity Warning – Troubleshooting Guide

BeeGFS Capacity Warning – Troubleshooting Guide

Purpose

Provide a standard troubleshooting procedure for engineers when a BeeGFS Capacity Warning alert is triggered in the cluster.

This guide ensures engineers follow consistent investigation and escalation steps.


When to Use

Use this KB when the monitoring system generates the following alert:

Alert Name: BeeGFS Capacity Warning


What the Alert Is

This alert indicates that the BeeGFS filesystem capacity is approaching the configured warning threshold.

If the storage usage continues to increase, it may eventually lead to:

  • Storage exhaustion

  • Job failures

  • Write failures in the filesystem

  • Performance degradation


Summary

BeeGFS storage usage has crossed the warning threshold, indicating the filesystem is gradually filling up.

Early investigation helps prevent the situation from becoming critical.


Severity

Warning

This is not an immediate failure but requires monitoring and investigation.


L1 Engineer Actions

Follow these steps when the alert is received.

Step 1 – Log the Alert

  • Record the alert in the Alert Summary sheet

Step 2 – Create Ticket

  • Open an incident ticket for tracking

Step 3 – Escalation (If Required)

If the issue appears significant:

  • Create an internal ticket

  • Inform L2 Engineer

Step 4 – Documentation

After confirmation with L2:

  • Update the Issues Master Sheet


L2 Engineer Actions

Perform deeper investigation after escalation.

Step 1 – Check OSS Server Status

Verify if any BeeGFS OSS (Object Storage Server) is:

  • Down

  • Unreachable

  • Not contributing to storage


Step 2 – Verify Metadata / Storage Threads

Check whether there are issues with:

  • Metadata services

  • Storage thread count

  • BeeGFS services


Step 3 – Verify Current Free Space

Check the current filesystem usage.

Example command:

df -h

Verify the available free space percentage.


Step 4 – Identify Growth Pattern

Investigate if storage growth is due to:

  • Large jobs

  • Bulk data ingestion

  • Unexpected data accumulation


Step 5 – Coordinate With NetApp Team

If storage continues to fill up or backend storage issues are suspected:

  • Coordinate with NetApp storage team


Step 6 – Escalation

If the issue persists after investigation:

  • Escalate to customer or higher engineering support


Expected Outcome

After investigation and corrective action:

  • BeeGFS capacity usage should stabilize

  • Free space should return to safe threshold

  • The alert should auto-clear in the monitoring system

    • Related Articles

    • ZFS Pool Capacity Warning Alert - Troubleshooting Guide

      Purpose This document provides a standardized approach for handling and troubleshooting the ZFS Pool Capacity Warning alert, helping L1 and L2 engineers prevent storage exhaustion and maintain system stability. Alert Name ZFS Pool Capacity Warning ...
    • ZFS Pool Capacity Critical Alert - Troubleshooting Guide

      Purpose This document provides a standardized approach for handling and troubleshooting the ZFS Pool Capacity Critical alert, enabling rapid response to prevent storage exhaustion, performance degradation, and potential service outages. Alert Name ...
    • BeeGFS Capacity Critical – Troubleshooting Guide

      Purpose Provide a troubleshooting reference for engineers when a BeeGFS Capacity Critical alert is triggered in the cluster. This document outlines the meaning of the alert and the actions required by L1 and L2 engineers. When to Use Use this KB ...
    • Memory Filesystem Capacity Warning -- Troubleshooting guide

      Purpose Provide a standard troubleshooting and resolution procedure when filesystem capacity on any HGX Node reaches warning levels. When to Use Use this KB when the monitoring system triggers the alert: Memory Filesystem Capacity warning – For Any ...
    • Filesystem Capacity Warning in Head Node - Troubleshooting Guide

      Purpose Provide a troubleshooting and response procedure when a Filesystem Capacity Warning alert is triggered on a Head Node. This KB helps engineers quickly identify the issue, perform basic checks, and escalate when required. Alert Name Filesystem ...