CPU Temperature Critical - Troubleshooting Guide

CPU Temperature Critical - Troubleshooting Guide

Purpose

Provide a standard troubleshooting procedure for engineers when a CPU Temperature Critical alert is received in the cluster.

This article helps engineers quickly identify the cause and perform initial remediation before escalation.


When to Use

Use this article when monitoring systems generate the alert:

CPU Temperature Critical

Typical trigger conditions:

  • CPU temperature exceeds safe operating threshold

  • Node thermal protection mechanisms are activated

  • Cooling or airflow issues are detected


Alert Summary

This alert indicates that the CPU temperature has exceeded the defined safe limit on a compute node.

Possible causes include:

  • High system workload

  • Cooling failure

  • Fan malfunction

  • Airflow blockage

  • Thermal throttling

  • Hardware sensor issues

If not resolved quickly, the system may:

  • Throttle CPU performance

  • Shut down automatically

  • Cause hardware damage


Severity

Critical

Immediate investigation is required to prevent hardware damage or node shutdown.


L1 Engineer Actions

Follow these steps when the alert is received.

Step 1 — Verify Alert Details

Check monitoring dashboard or alert summary and confirm:

  • Node hostname

  • Timestamp of alert

  • Current CPU temperature

  • Alert severity


Step 2 – Wait and Monitor

  • Wait 25–30 minutes


Step 3 – Escalate if Alert Persists

If the alert does not clear after monitoring:
  1. Create an internal ticket
  2. Inform L2 Engineer


Step 4 – Documentation

After confirming with L2:

  • Update the Issues Master Sheet


L2 Engineer Actions

If the issue persists after L1 checks, escalate to L2 and perform the following.

Step 1 — Check via BMC / IPMI Sensors

Access the server BMC interface and verify:

  • CPU temperature

  • System temperature sensors

  • Fan speeds

  • Hardware warnings

Example:

ipmitool sensor

Step 2 — Verify Thermal Throttling

Check if the CPU is throttling due to overheating.

Example:

dmesg | grep -i thermal

Step 3 — Verify Cooling Infrastructure

Check for data center environmental issues:

  • Rack cooling

  • Airflow obstruction

  • Hot aisle / cold aisle issues


Step 4 — Validate Workaround Behavior

Ensure the node is not under abnormal computational load.

Investigate:

  • GPU jobs

  • CPU intensive processes

  • runaway processes


Step 5 — Escalate to OEM

If temperature remains critical after all checks:

  • Open a hardware support case with the OEM

  • Provide logs and sensor outputs

  • Include node serial number and alert timestamp

    • Related Articles

    • GPU Memory Temperature Critical - Troubleshooting Guide

      Purpose This document provides a standardized approach for handling and troubleshooting the GPU Memory Temperature Critical alert, including L1 and L2 actions, escalation procedures, and resolution criteria. Alert Name GPU Memory Temperature Critical ...
    • CPU Temperature Warning - Troubleshooting Guide

      Purpose Provide a troubleshooting guide for engineers when a CPU Temperature Warning alert is triggered on a Node/Server. This KB helps engineers quickly identify the issue and perform initial remediation before escalation. When to Use Use this KB ...
    • Fan Speed Warning - Troubleshooting Guide

      Purpose Provide guidance for engineers to investigate and respond to Fan Speed Warning alerts generated by the monitoring system. This KB article ensures that engineers follow a consistent troubleshooting procedure when cooling-related alerts occur. ...
    • ZFS Pool Capacity Critical Alert - Troubleshooting Guide

      Purpose This document provides a standardized approach for handling and troubleshooting the ZFS Pool Capacity Critical alert, enabling rapid response to prevent storage exhaustion, performance degradation, and potential service outages. Alert Name ...
    • Network Down Extern (P1 Critical Alert) - Troubleshooting Guide

      Purpose This document provides a standardized approach for handling and troubleshooting the Network Down - Extern alert, including L1 and L2 actions, escalation procedures, and resolution criteria for critical incidents. Alert Name Network Down - ...