Fan Speed Critical - Troubleshooting Guide

Fan Speed Critical - Troubleshooting Guide

Purpose

Provide a troubleshooting guide for engineers when a Fan Speed Critical alert is generated on a server node.

This article helps engineers quickly identify cooling issues and perform corrective actions to prevent overheating and potential hardware failure.


When to Use

Use this KB article when the monitoring system triggers the alert:

Fan Speed Critical

This alert indicates that one or more system fans are operating below the required speed or have stopped functioning.


What the Alert Is

The Fan Speed Critical alert occurs when the server hardware monitoring system detects abnormal fan behavior such as:

  • Fan speed dropping below threshold

  • Fan failure

  • Fan not detected by system sensors

This condition can reduce cooling efficiency and lead to CPU or system overheating.


Alert Summary

When this alert occurs, the system identifies that cooling fans are not operating within the expected RPM range.

Possible causes include:

  • Fan hardware failure

  • Loose fan connection

  • Obstruction in fan blades

  • Dust buildup affecting airflow

  • BMC sensor error

Immediate investigation is required to maintain proper cooling.


Severity

Critical

Cooling system failure can lead to:

  • CPU overheating

  • System thermal throttling

  • Unexpected node shutdown

  • Hardware damage


L1 Engineer Actions

Follow these steps when the alert is received.

Step 1 – Verify the Alert

Check the monitoring system and confirm:

  • Node hostname

  • Timestamp of alert

  • Affected fan sensor

  • Current fan speed reading


Step 2 – Wait and Monitor

  1. Wait 25–30 minutes

Step 3 – Escalate if Alert Persists

If the alert does not clear after monitoring:
  1. Create an internal ticket
  2. Inform L2 Engineer

Step 4 - Documentation

After confirmation with L2:
  1. Update the Issues Master List

L2 Engineer Actions

If the issue persists after L1 troubleshooting, escalate to L2.

Step 1 – Verify via BMC

Access the BMC / IPMI interface and verify fan sensor readings.

Check:

  • Fan RPM values

  • Fan health status

  • Hardware logs

Example command:

Note: if we need to run "ipmitool" command check server connectivity from the source machine and ipmi executable file or ipmi package must be installed in the same source machine. 
ipmitool sdr type fan

Step 2 – Identify Failed Fan

Determine which fan module has failed.

Check:

  • Fan ID

  • RPM readings

  • Sensor status

Step 3 – Plan Preventive Maintenance

If the fan continues to run below optimal speed:

  • Schedule preventive maintenance

  • Clean the server or replace the fan module if needed

Step 4 – Check for Dust or Airflow Blockage

Inspect the physical server if required.

Verify:

  • Dust accumulation

  • Blocked air vents

  • Proper airflow inside the rack

Step 5 – Try Physical Reseating the Fan

If the server supports hot-swappable fans:

  1. Remove the faulty fan module

  2. Reinsert the fan properly

  3. Verify the fan is detected again

Step 6 – Replace Faulty Fan

If reseating does not resolve the issue, replace the fan module.

Step 7 – Escalate to OEM

If the issue persists:

  • Collect hardware logs

  • Record node serial number

  • Capture sensor output

Open a support case with the OEM vendor.

    • Related Articles

    • Fan Speed Warning - Troubleshooting Guide

      Purpose Provide guidance for engineers to investigate and respond to Fan Speed Warning alerts generated by the monitoring system. This KB article ensures that engineers follow a consistent troubleshooting procedure when cooling-related alerts occur. ...
    • CPU Temperature Warning - Troubleshooting Guide

      Purpose Provide a troubleshooting guide for engineers when a CPU Temperature Warning alert is triggered on a Node/Server. This KB helps engineers quickly identify the issue and perform initial remediation before escalation. When to Use Use this KB ...
    • ZFS Pool Capacity Critical Alert - Troubleshooting Guide

      Purpose This document provides a standardized approach for handling and troubleshooting the ZFS Pool Capacity Critical alert, enabling rapid response to prevent storage exhaustion, performance degradation, and potential service outages. Alert Name ...
    • Network Down Extern (P1 Critical Alert) - Troubleshooting Guide

      Purpose This document provides a standardized approach for handling and troubleshooting the Network Down - Extern alert, including L1 and L2 actions, escalation procedures, and resolution criteria for critical incidents. Alert Name Network Down - ...
    • Network Speed IB Alert - Troubleshooting Guide

      Purpose This document provides a standardized approach for handling and troubleshooting the Network Speed – IB alert, including L1 and L2 actions, escalation process, and resolution guidelines. Alert Name Network Speed IB Alert Description This alert ...