Purpose
This document provides a standardized approach for handling and troubleshooting the Nvidia-smi Tool alert, including L1 and L2 actions, escalation guidelines, and resolution criteria.
Alert Name
Nvidia-smi Tool
Alert Description
This alert is triggered when the nvidia-smi command fails to execute successfully on a node.
The alert indicates a potential issue with:
NVIDIA driver stack
GPU communication
System-level
GPU services or state
Failure of this command may result in:
-
Inability to monitor GPU health
- GPU workloads being impacted
- Loss of visibility into GPU utilization and errors
Severity
P3 – Medium Impact
- No immediate outage, but GPU observability or performance may be affected
-
Requires timely investigation to prevent escalation
Possible Causes
- NVIDIA drivers not installed or corrupted
- GPU not detected by OS
- NVIDIA persistence daemon not running
- Kernel module issues
- Hardware or PCIe communication failure
- Recent system changes or updates
L1 Engineer Actions
Step 1 – Log the Alert
- Log in to the affected node
- Validate the alert
- Wait for 15-20 minutes if alert gets auto resolved or false alarm.
Step 2 – Create Ticket
- Verify if the issue persists
- Capture error output (if any)
- Create an internal ticket with:
- Error output
- Observations
- Node details
Step 3 – Escalation (If required)
- Inform L2 team immediately
- Share all collected details and logs
Step 4 – Documentation
- Update Alert Summary with:
- Observations
- Error messages
- After L2 confirmation, update:
- Issue Master Sheet / Issue Tracking Sheet
L2 Engineer Actions
Step 1 – Check NVIDIA Persistence Service
systemctl status nvidia-persistenced
- Ensure the service is active (running)
- Restart if required
systemctl restart nvidia-persistenced
Step 2 – Verify NVIDIA Driver Status
nvidia-smi
- If still failing:
lsmod | grep nvidia
- Confirm NVIDIA kernel modules are loaded
- Check for driver mismatch or failure
Step 3 – Check System Logs
dmesg | grep -i nvidia
journalctl -xe | grep -i nvidia
- Identify driver errors or hardware-related issues.
Step 4 – Validate GPU Detection
lspci | grep -i nvidia
- Ensure GPU is visible at hardware level.
Step 5 – OEM Escalation (If required)
- If issue persists after troubleshooting:
- Create a ticket with OEM
- Share logs, output, and observations.
Resolution
The issue is considered resolved when:
-
nvidia-smi executes successfully
-
GPU details are displayed correctly
-
No related errors are observed in system logs
Escalation
- Internal Escalation: L1 → L2 with complete logs and observations
- External Escalation: L2 → OEM (if hardware/driver issue persists