Provide a standard troubleshooting procedure for engineers when a CPU Temperature Critical alert is received in the cluster.
This article helps engineers quickly identify the cause and perform initial remediation before escalation.
Use this article when monitoring systems generate the alert:
CPU Temperature Critical
Typical trigger conditions:
CPU temperature exceeds safe operating threshold
Node thermal protection mechanisms are activated
Cooling or airflow issues are detected
This alert indicates that the CPU temperature has exceeded the defined safe limit on a compute node.
Possible causes include:
High system workload
Cooling failure
Fan malfunction
Airflow blockage
Thermal throttling
Hardware sensor issues
If not resolved quickly, the system may:
Throttle CPU performance
Shut down automatically
Cause hardware damage
Critical
Immediate investigation is required to prevent hardware damage or node shutdown.
Follow these steps when the alert is received.
Check monitoring dashboard or alert summary and confirm:
Node hostname
Timestamp of alert
Current CPU temperature
Alert severity
Wait 25–30 minutes
After confirming with L2:
Update the Issues Master Sheet
If the issue persists after L1 checks, escalate to L2 and perform the following.
Access the server BMC interface and verify:
CPU temperature
System temperature sensors
Fan speeds
Hardware warnings
Example:
ipmitool sensor
Check if the CPU is throttling due to overheating.
Example:
dmesg | grep -i thermal
Check for data center environmental issues:
Rack cooling
Airflow obstruction
Hot aisle / cold aisle issues
Ensure the node is not under abnormal computational load.
Investigate:
GPU jobs
CPU intensive processes
runaway processes
If temperature remains critical after all checks:
Open a hardware support case with the OEM
Provide logs and sensor outputs
Include node serial number and alert timestamp