This article provides guidance for engineers to identify and respond to the GPU Temperature Critical alert in GPU compute nodes. It outlines the alert meaning, possible causes, and the required L1 and L2 troubleshooting steps.
GPU Temperature Critical
This alert is triggered when the GPU temperature exceeds 85°C.
High GPU temperature typically occurs due to heavy workloads, insufficient cooling, or hardware issues affecting airflow.
P3 – Medium Impact
The system may still be operational, but prolonged high temperature can lead to:
GPU thermal throttling
Performance degradation
Potential hardware damage if not resolved
High GPU workload or sustained compute jobs
Cooling system inefficiency
Fan malfunction or reduced airflow
Data center ambient temperature increase
Blocked air vents or dust accumulation
Thermal throttling conditions
Follow these steps when the alert is received.
Record the alert in the Alert Summary sheet
Open an incident ticket for tracking
If the issue appears significant:
Create an internal ticket
Inform L2 Engineer
After confirmation with L2:
Update the Issues Master Sheet
Perform deeper analysis and corrective actions:
Confirm whether GPU is throttling due to high temperature using:
nvidia-smi -q | grep -i throttle
Depending on the root cause:
Reduce GPU workload temporarily
Improve airflow or cooling
Replace faulty fans if detected
Clean air vents or filters
Adjust environmental cooling in the data center
Server reboot / Hard reset
Escalate to hardware support/vendor if:
GPU temperature remains critical despite normal workload
Cooling components are faulty
GPU shows repeated thermal throttling