Provide a standard troubleshooting procedure for engineers when a BeeGFS Capacity Warning alert is triggered in the cluster.
This guide ensures engineers follow consistent investigation and escalation steps.
Use this KB when the monitoring system generates the following alert:
Alert Name: BeeGFS Capacity Warning
This alert indicates that the BeeGFS filesystem capacity is approaching the configured warning threshold.
If the storage usage continues to increase, it may eventually lead to:
Storage exhaustion
Job failures
Write failures in the filesystem
Performance degradation
BeeGFS storage usage has crossed the warning threshold, indicating the filesystem is gradually filling up.
Early investigation helps prevent the situation from becoming critical.
Warning
This is not an immediate failure but requires monitoring and investigation.
Follow these steps when the alert is received.
Record the alert in the Alert Summary sheet
Open an incident ticket for tracking
If the issue appears significant:
Create an internal ticket
Inform L2 Engineer
After confirmation with L2:
Update the Issues Master Sheet
Perform deeper investigation after escalation.
Verify if any BeeGFS OSS (Object Storage Server) is:
Down
Unreachable
Not contributing to storage
Check whether there are issues with:
Metadata services
Storage thread count
BeeGFS services
Check the current filesystem usage.
Example command:
df -h
Verify the available free space percentage.
Investigate if storage growth is due to:
Large jobs
Bulk data ingestion
Unexpected data accumulation
If storage continues to fill up or backend storage issues are suspected:
Coordinate with NetApp storage team
If the issue persists after investigation:
Escalate to customer or higher engineering support
After investigation and corrective action:
BeeGFS capacity usage should stabilize
Free space should return to safe threshold
The alert should auto-clear in the monitoring system