KB 511966 - BeeGFS Capacity Warning – Troubleshooting Guide

KB 511966 - BeeGFS Capacity Warning – Troubleshooting Guide

Purpose

Provide a standard troubleshooting procedure for engineers when a BeeGFS Capacity Warning alert is triggered in the cluster.

This guide ensures engineers follow consistent investigation and escalation steps.


When to Use

Use this KB when the monitoring system generates the following alert:

Alert Name: BeeGFS Capacity Warning


What the Alert Is

This alert indicates that the BeeGFS filesystem capacity is approaching the configured warning threshold.

If the storage usage continues to increase, it may eventually lead to:

  • Storage exhaustion

  • Job failures

  • Write failures in the filesystem

  • Performance degradation


Summary

BeeGFS storage usage has crossed the warning threshold, indicating the filesystem is gradually filling up.

Early investigation helps prevent the situation from becoming critical.


Severity

Warning

This is not an immediate failure but requires monitoring and investigation.


L1 Engineer Actions

Follow these steps when the alert is received.

Step 1 – Log the Alert

  • Record the alert in the Alert Summary sheet

Step 2 – Create Ticket

  • Open an incident ticket for tracking

Step 3 – Escalation (If Required)

If the issue appears significant:

  • Create an internal ticket

  • Inform L2 Engineer

Step 4 – Documentation

After confirmation with L2:

  • Update the Issues Master Sheet


L2 Engineer Actions

Perform deeper investigation after escalation.

Step 1 – Check OSS Server Status

Verify if any BeeGFS OSS (Object Storage Server) is:

  • Down

  • Unreachable

  • Not contributing to storage


Step 2 – Verify Metadata / Storage Threads

Check whether there are issues with:

  • Metadata services

  • Storage thread count

  • BeeGFS services


Step 3 – Verify Current Free Space

Check the current filesystem usage.

Example command:

df -h

Verify the available free space percentage.


Step 4 – Identify Growth Pattern

Investigate if storage growth is due to:

  • Large jobs

  • Bulk data ingestion

  • Unexpected data accumulation


Step 5 – Coordinate With NetApp Team

If storage continues to fill up or backend storage issues are suspected:

  • Coordinate with NetApp storage team


Step 6 – Escalation

If the issue persists after investigation:

  • Escalate to customer or higher engineering support


Expected Outcome

After investigation and corrective action:

  • BeeGFS capacity usage should stabilize

  • Free space should return to safe threshold

  • The alert should auto-clear in the monitoring system