KB 831053 - GPU Temperature Critical alert - Troubleshooting Guide

KB 831053 - GPU Temperature Critical alert - Troubleshooting Guide

Purpose

This article provides guidance for engineers to identify and respond to the GPU Temperature Critical alert in GPU compute nodes. It outlines the alert meaning, possible causes, and the required L1 and L2 troubleshooting steps.



Alert Name

GPU Temperature Critical


Alert Description

This alert is triggered when the GPU temperature exceeds 85°C.

High GPU temperature typically occurs due to heavy workloads, insufficient cooling, or hardware issues affecting airflow.


Severity

P3 – Medium Impact

The system may still be operational, but prolonged high temperature can lead to:

  • GPU thermal throttling

  • Performance degradation

  • Potential hardware damage if not resolved


Possible Causes

  • High GPU workload or sustained compute jobs

  • Cooling system inefficiency

  • Fan malfunction or reduced airflow

  • Data center ambient temperature increase

  • Blocked air vents or dust accumulation

  • Thermal throttling conditions


L1 Engineer Actions

Follow these steps when the alert is received.

Step 1 – Log the Alert

  • Record the alert in the Alert Summary sheet


Step 2 – Create Ticket

  • Open an incident ticket for tracking


Step 3 – Escalation (If Required)

If the issue appears significant:

  • Create an internal ticket

  • Inform L2 Engineer


Step 4 – Documentation

After confirmation with L2:

  • Update the Issues Master Sheet


L2 Engineer Actions

Perform deeper analysis and corrective actions:

Step 1 - Detailed GPU and process check

Use:
nvidia-smi
free -h

Identify processes consuming excessive GPU memory.

Step 2 - Check airflow and cooling

Verify:
Fan functionality
Airflow obstruction
Rack cooling conditions

Step 3 - Check thermal throttling

Confirm whether GPU is throttling due to high temperature using:

nvidia-smi -q | grep -i throttle

Step 4 - Analyze system workload

  1. Identify abnormal or stuck GPU-intensive jobs.
  2. Consider stopping or redistributing workloads if required.

Step 5 - Reboot the Server

  1. Check the server by performing the power cycle.

  2. Also, if possible, check with the Hard Reset. 


Resolution

Depending on the root cause:

  • Reduce GPU workload temporarily

  • Improve airflow or cooling

  • Replace faulty fans if detected

  • Clean air vents or filters

  • Adjust environmental cooling in the data center

  • Server reboot / Hard reset


Escalation

Escalate to hardware support/vendor if:

  • GPU temperature remains critical despite normal workload

  • Cooling components are faulty

  • GPU shows repeated thermal throttling

    • Related Articles

    • KB 744001 - GPU Memory Temperature Critical - Troubleshooting Guide

      Purpose This document provides a standardized approach for handling and troubleshooting the GPU Memory Temperature Critical alert, including L1 and L2 actions, escalation procedures, and resolution criteria. Alert Name GPU Memory Temperature Critical ...
    • KB 579006 - CPU Temperature Critical - Troubleshooting Guide

      Purpose Provide a standard troubleshooting procedure for engineers when a CPU Temperature Critical alert is received in the cluster. This article helps engineers quickly identify the cause and perform initial remediation before escalation. When to ...
    • KB 831093 - GPU Temperature Warning alert- Troubleshooting Guide

      Purpose Provide guidance for engineers to investigate and resolve the GPU Temperature Warning alert in GPU compute nodes. When To Use Use this article when the monitoring system generates the alert: GPU Temperature Warning Alert Details This alert is ...
    • KB 579132 - Fan Speed Warning - Troubleshooting Guide

      Purpose Provide guidance for engineers to investigate and respond to Fan Speed Warning alerts generated by the monitoring system. This KB article ensures that engineers follow a consistent troubleshooting procedure when cooling-related alerts occur. ...
    • KB 905833 - Network Down Extern (P1 Critical Alert) - Troubleshooting Guide

      Purpose This document provides a standardized approach for handling and troubleshooting the Network Down - Extern alert, including L1 and L2 actions, escalation procedures, and resolution criteria for critical incidents. Alert Name Network Down - ...
    • Recent Articles

    • KB-630208 | How to use Find and Locate to search for files in Linux

      Purpose To provide Linux administrators and Field Engineers with guidance on using the find and locate utilities to efficiently search for files and directories within a Linux filesystem based on criteria such as name, type, size, timestamps, ...
    • KB-630164 | How to use rsync to Synchronize Files

      Purpose To provide a clear and practical procedure for using the rsync utility to synchronize files and directories between local systems and remote hosts, including push and pull operations, while preserving file attributes and optimizing data ...
    • KB-630107 | How to create a Linux swap file

      Purpose To provide Field Engineers and Linux administrators with a standardized procedure for creating, configuring, and validating a Linux swap file. This procedure helps ensure sufficient virtual memory is available when physical RAM is exhausted ...
    • MSI All-in-One (AIO) PC – Physical Damage Policy for Warranty Claims

      MSI All-in-One (AIO) PC – Physical Damage Policy for Warranty Claims MSI's standard warranty does not cover physical or accidental damage to an All-in-One (AIO) PC. If the damage is determined to be caused by external factors rather than a ...
    • KB 134833 - Troubleshooting NVSM Alert NV-CPU-XX – Unrecoverable CPU Internal Error

      Purpose This document provides a general troubleshooting procedure for the NVIDIA System Management (NVSM) alert NV-CPU-XX, which indicates that a CPU has reported an internal error. The article outlines how to verify whether the alert represents an ...
    • Popular Articles

    • CP Plus Camera and NVR Configuration

      NVR Configuration The CP Plus Pro Series of NVRs have been meticulously designed for providing you with upgraded performance and higher recording quality in your IP video surveillance solution. The robust processor that has been inculcated in this ...
    • KB 775235 - Kerberos Authentication – Overview

      What is Kerberos? Kerberos is a secure authentication method used in our Active Directory (AD) environment (mbuzztech.com). It allows users to: Access multiple systems without re-entering passwords (Single Sign-On – SSO) Log in once Where We Use It ...
    • How to Remove and Reinstall NVIDIA Drivers on Ubuntu

      This article provides step by step guide to completely remove existing NVIDIA drivers and reinstall specific version of the NVIDIA driver on the Ubuntu system Prerequisites Administrative (sudo) access to the Ubuntu system. Internet access to ...
    • KB 692001 - Personal Computers and Servers - Classification and Point of Contact

      We can classify the computers that MBUZZ handles based on their form-factor as below: Tower Workstations, Desktops, Gaming PCs and SFF (Small form factor) PCs fall under this category. These are computers people would use on a desk and rarely move. ...
    • KB 298031 - M.2 SSD Tier List

      The sequential read and write speeds, which are usually the most advertised number, are not a proper benchmark of real-world performance or the quality of an SSD. This article categorizes and tiers SSDs based on factors like the type of NAND flash, ...