KB 265938 - Nvidia-smi Tool Alert - Troubleshooting Guide
Purpose
This document provides a standardized approach for handling and troubleshooting the Nvidia-smi Tool alert, including L1 and L2 actions, escalation guidelines, and resolution criteria.
Alert Name
Nvidia-smi Tool
Alert Description
This alert is triggered when the nvidia-smi command fails to execute successfully on a node.
The alert indicates a potential issue with:
NVIDIA driver stack
GPU communication
System-level
GPU services or state
Failure of this command may result in:
-
Inability to monitor GPU health
- GPU workloads being impacted
- Loss of visibility into GPU utilization and errors
Severity
P3 – Medium Impact
- No immediate outage, but GPU observability or performance may be affected
-
Requires timely investigation to prevent escalation
Possible Causes
- NVIDIA drivers not installed or corrupted
- GPU not detected by OS
- NVIDIA persistence daemon not running
- Kernel module issues
- Hardware or PCIe communication failure
- Recent system changes or updates
L1 Engineer Actions
Step 1 – Log the Alert
- Log in to the affected node
- Validate the alert
- Wait for 15-20 minutes if alert gets auto resolved or false alarm.
Step 2 – Create Ticket
- Verify if the issue persists
- Capture error output (if any)
- Create an internal ticket with:
- Error output
- Observations
- Node details
Step 3 – Escalation (If required)
- Inform L2 team immediately
- Share all collected details and logs
Step 4 – Documentation
- Update Alert Summary with:
- Observations
- Error messages
- After L2 confirmation, update:
- Issue Master Sheet / Issue Tracking Sheet
L2 Engineer Actions
Step 1 – Check NVIDIA Persistence Service
systemctl status nvidia-persistenced
- Ensure the service is active (running)
- Restart if required
systemctl restart nvidia-persistenced
Step 2 – Verify NVIDIA Driver Status
nvidia-smi
- If still failing:
lsmod | grep nvidia
- Confirm NVIDIA kernel modules are loaded
- Check for driver mismatch or failure
Step 3 – Check System Logs
dmesg | grep -i nvidia
journalctl -xe | grep -i nvidia
- Identify driver errors or hardware-related issues.
Step 4 – Validate GPU Detection
lspci | grep -i nvidia
- Ensure GPU is visible at hardware level.
Step 5 – OEM Escalation (If required)
- If issue persists after troubleshooting:
- Create a ticket with OEM
- Share logs, output, and observations.
Resolution
The issue is considered resolved when:
-
nvidia-smi executes successfully
-
GPU details are displayed correctly
-
No related errors are observed in system logs
Escalation
- Internal Escalation: L1 → L2 with complete logs and observations
- External Escalation: L2 → OEM (if hardware/driver issue persists
Related Articles
KB 774807 - GPU Miscount Alert - Troubleshooting Guide
Purpose Provide troubleshooting steps when a GPU Miscount alert is triggered, indicating that the system detects an incorrect number of GPUs on a node. Alert Name GPU Miscount What the Alert Is This alert indicates that the number of GPUs detected on ...
KB 721021 - GPU Memory Mismatch – Troubleshooting Guide
Alert Information Alert Name: GPU Memory Mismatch Severity: P3 – Medium Impact Alert Description: This alert is generated when the detected GPU memory size on a node does not match the expected configuration. Alert Summary The monitoring system has ...
KB 829939 - Execution of NVIDIA Field Diagnostic (FD) Tool and Collection of Diagnostic Logs, and troubleshooting of GPU(s) not detected
1. Objective To provide a standardized procedure for executing the NVIDIA Field Diagnostic (FD) tool, verifying GPU status, collecting the required diagnostic logs, and documenting the findings for further analysis. 2. Scope This procedure applies to ...
KB 744001 - GPU Memory Temperature Critical - Troubleshooting Guide
Purpose This document provides a standardized approach for handling and troubleshooting the GPU Memory Temperature Critical alert, including L1 and L2 actions, escalation procedures, and resolution criteria. Alert Name GPU Memory Temperature Critical ...
KB 905833 - Network Down Extern (P1 Critical Alert) - Troubleshooting Guide
Purpose This document provides a standardized approach for handling and troubleshooting the Network Down - Extern alert, including L1 and L2 actions, escalation procedures, and resolution criteria for critical incidents. Alert Name Network Down - ...
Recent Articles
KB-630208 | How to use Find and Locate to search for files in Linux
Purpose To provide Linux administrators and Field Engineers with guidance on using the find and locate utilities to efficiently search for files and directories within a Linux filesystem based on criteria such as name, type, size, timestamps, ...
KB-630164 | How to use rsync to Synchronize Files
Purpose To provide a clear and practical procedure for using the rsync utility to synchronize files and directories between local systems and remote hosts, including push and pull operations, while preserving file attributes and optimizing data ...
KB-630107 | How to create a Linux swap file
Purpose To provide Field Engineers and Linux administrators with a standardized procedure for creating, configuring, and validating a Linux swap file. This procedure helps ensure sufficient virtual memory is available when physical RAM is exhausted ...
MSI All-in-One (AIO) PC – Physical Damage Policy for Warranty Claims
MSI All-in-One (AIO) PC – Physical Damage Policy for Warranty Claims MSI's standard warranty does not cover physical or accidental damage to an All-in-One (AIO) PC. If the damage is determined to be caused by external factors rather than a ...
KB 134833 - Troubleshooting NVSM Alert NV-CPU-XX – Unrecoverable CPU Internal Error
Purpose This document provides a general troubleshooting procedure for the NVIDIA System Management (NVSM) alert NV-CPU-XX, which indicates that a CPU has reported an internal error. The article outlines how to verify whether the alert represents an ...
Popular Articles
CP Plus Camera and NVR Configuration
NVR Configuration The CP Plus Pro Series of NVRs have been meticulously designed for providing you with upgraded performance and higher recording quality in your IP video surveillance solution. The robust processor that has been inculcated in this ...
KB 775235 - Kerberos Authentication – Overview
What is Kerberos? Kerberos is a secure authentication method used in our Active Directory (AD) environment (mbuzztech.com). It allows users to: Access multiple systems without re-entering passwords (Single Sign-On – SSO) Log in once Where We Use It ...
How to Remove and Reinstall NVIDIA Drivers on Ubuntu
This article provides step by step guide to completely remove existing NVIDIA drivers and reinstall specific version of the NVIDIA driver on the Ubuntu system Prerequisites Administrative (sudo) access to the Ubuntu system. Internet access to ...
KB 692001 - Personal Computers and Servers - Classification and Point of Contact
We can classify the computers that MBUZZ handles based on their form-factor as below: Tower Workstations, Desktops, Gaming PCs and SFF (Small form factor) PCs fall under this category. These are computers people would use on a desk and rarely move. ...
KB 298031 - M.2 SSD Tier List
The sequential read and write speeds, which are usually the most advertised number, are not a proper benchmark of real-world performance or the quality of an SSD. This article categorizes and tiers SSDs based on factors like the type of NAND flash, ...