Purpose
This document provides a standardized approach for handling and troubleshooting the GPU Memory Temperature Critical alert, including L1 and L2 actions, escalation procedures, and resolution criteria.
Alert Name
GPU Memory Temperature Critical
Alert Description
This alert is triggered when abnormally high temperature is detected in GPU memory modules.
GPU memory overheating can be more critical than GPU core temperature and may indicate:
- Improper airflow in the chassis
- High GPU workload
- Cooling system malfunction
- Hardware degradation or failure
If unresolved, this may lead to:
- GPU throttling
- Workload failures
- Hardware damage
Severity
P3 - Medium Impact
Possible Causes
- High GPU utilization due to workloads
- Improper chassis airflow
- Fan malfunction
- Data center cooling issues
- Dust accumulation in GPU server
- Firmware or driver anomalies
- Hardware degradation
L1 Engineer Actions
Step 1 - Log the Alert
- Log in to the monitoring platform
- Review:
- Alert summary
- Affected node details
Step2 - Create Ticket
- Create an internal incident ticket including:
- Alert name
- Node hostname
- Timestamp
- GPU ID (if available)
- Screenshot or alert logs
Step 3 – Initial Verification
- Check whether the alert is:
- Transient OR
- Persistent
Step 4 – Escalation (If required)
- If alert persists beyond threshold:
- Escalate to L2 support
Step 5 – Documentation
- After confirmation with L2:
- Update Issues Tracker Sheet
- Add troubleshooting steps and observations in ticket
L2 Engineer Actions
Step 1 – Check GPU Temperature & Utilization
nvidia-smi
Verify:
- GPU memory temperature
- GPU utilization
- Running GPU processes
- Power consumption
Step 2 – Check System Memory Status
free -h
Verify:
- Available memory
- Memory pressure
Step 3 – Identify Running GPU Processes
top OR htop
ps -aux
Check for:
- Long-running jobs
- Unexpected GPU usage
- High CPU/ GPU consumers
Step 4 – Check Zombie Processes
ps aux | grep Z
- Identify zombie processes
- Investigate parent processes
Step 5 – Verify Server Airflow & Cooling
- Physically check:
- Server fans are operational
- No airflow obstruction
- Rack cooling is functioning
- No dust accumulation
- If issue found:
- Inform data center operations team
Step 6 – Verify System Memory Configuration
cat /proc/meminfo
Step 7 – Verify Installed Memory Hardware
sudo dmidecode -t memory
Confirm:
- Correct memory modules
- No hardware mismatch
- All DIMMs detected
Step 8 – OEM Escalation (If required)
- Escalate if:
- GPU memory temperature remains critical
- Hardware anomaly detected
- Cooling components malfunction
- Repeated thermal alerts
- Provide:
- nvidia-smi output
- System logs
- Hardware configuration
- Sever serial number
Resolution
The issue is considered resolved when:
- GPU memory temperature returns to normal range
- No recurring alerts in monitoring system
- Workloads run without thermal throttling
Escalation
- L1 → L2: If alert persists beyond threshold
- L2 → OEM/Vendor: If hardware or thermal issue persists