Execution of NVIDIA Field Diagnostic (FD) Tool and Collection of Diagnostic Logs, and troubleshooting of GPU(s) not detected

Execution of NVIDIA Field Diagnostic (FD) Tool and Collection of Diagnostic Logs, and troubleshooting of GPU(s) not detected

1. Objective

To provide a standardized procedure for executing the NVIDIA Field Diagnostic (FD) tool, verifying GPU status, collecting the required diagnostic logs, and documenting the findings for further analysis.

 

2. Scope

This procedure applies to field engineers performing GPU diagnostics on Supermicro GPU servers where GPU related issues have been reported and execution of the NVIDIA Field Diagnostic (FD) tool for health verification or troubleshooting.

 

3. Prerequisites

Before starting the activity, ensure the following:

  • Coordinate with the customer and confirm maintenance window.
  • Customer availability and site access confirmed.
  • Verify the server hostname, rack location, and serial number (if available).
  • Ensure the required NVIDIA Field Diagnostic (FD) package is available.
  • Carry a laptop with SSH access utilities.
  • Verify access credentials will be available during the activity.
  • Sufficient storage space is available to save the generated reports.

 

4. Required Tools

Hardware

  • Laptop
  • Network cable (if required)

Software

  • NVIDIA Field Diagnostic (FD) Tool
  • SSH Client
  • IPMI Utilities
  • NVIDIA CLI Utilities
  • Linux Shell Access

 

5. Activity Procedure

Step 1 – Customer Coordination

  1. Meet the customer representative.
  2. Verify the maintenance activity.
  3. Identify the affected server.
  4. Verify the server hostname, serial number, or asset tag.

Step 2 – Physical Verification

  1. Inspect the server for any visible hardware abnormalities.
  2. Check for amber fault LEDs on the server and GPU tray

 

Step 3 – Access Verification

  1. Obtain SSH or console access to the operating system.
  2. Verify login access.

 

Step 4 – Initial Diagnostics

Execute the NVIDIA Field Diagnostic (FD) Tool.

  • Run the complete Field Diagnostic.
  • Save the diagnostic report.
  • Record the execution time and outcome.

Step 5 – GPU Tray Re-seat (if required)

If instructed by the Supermicro support team:

 

GPU Tray

The GPU tray is located at the top of the chassis. It houses the system's GPUs, cooling components, and other supporting devices. The GPU tray is accessible from the front of the unit by sliding the upper tray out after disengaging a lever on the left and right side of the chassis.

 

Notes:

1.      Lift the heavy universal baseboard (UBB) plate by using two UBB handles.

2.      You should extend only one tray in a rack at a time - extending two trays simultaneously may cause the server to become unstable. Failure to stabilize the server can cause the server to tip over.

Warning! Tip over hazard when not in a rack!

 

1.      Gracefully shut down the server (if required).

2.      Remove the GPU tray.

3.      Inspect the tray for:

    • Physical damage
    • Bent pins
    • Loose connections

4.      Reinsert the GPU tray carefully.

  1. Ensure the tray is fully seated and locked.
  2. Power on the server.

Step 6 – Cold Reboot (if required)

If instructed by the Supermicro support team:

Perform a complete cold reboot.

After the system becomes operational:

  • Verify all GPUs are detected.
  • Allow sufficient time for the operating system and NVIDIA services to initialize.

 

Step 7 – Execute Field Diagnostic

Run the NVIDIA Field Diagnostic Tool.

Verify:

  • GPU detection status
  • Diagnostic result
  • Any reported failures

Save the updated diagnostic report.

 

Step 8 – If GPUs Are Still Missing

If one or more GPUs are not detected:

Document the following evidence:

  • GPU tray
  • Connector side
  • Pin side
  • Any damaged components
  • Fault LEDs (if present)

Screenshots

  • GPU detection failure
  • Diagnostic output
  • Error messages

 

Step 9 – Collect Diagnostic Logs

Execute the following commands and save the outputs.

nvidia-smi -q

nvidia-smi nvlink -s

nvidia-bug-report.sh

lspci -tv | grep -i nvidia

systemctl status nvidia-fabricmanager


For IPMI Logs

Prerequisite: ipmitool must be available on the system. If it is not installed in the system, use below link to download it.
(Select OS as per requirement)

After installing the ipmi tool in system, use below commands to collect the logs:
In Linux system:
./ipmitool -I lan -H <BMC_IP> -U <USERNAME> -P '<PASSWORD>'
ipmitool fru
ipmitool sel list

In Windows system:
ipmitool.exe -I lan -H <BMC_IP> -U <USERNAME> -P "<PASSWORD>" 

ipmitool fru

ipmitool sel list


Collect:

  • NVIDIA Bug Report
  • Field Diagnostic Report
  • Command outputs
  • Screenshots (if applicable)

 

Step 10 – Documentation

Document the following:

  • Server Serial Number
  • Rack Location
  • Date and Time of Activity
  • Number of GPUs Detected
  • Amber LED Status
  • Field Diagnostic Result
  • Actions Performed

Attach:

  • Diagnostic Reports
  • Log Files
  • Screenshots

6. Expected Outcome

  • NVIDIA Field Diagnostic executed successfully.
  • GPU detection verified.
  • Required logs collected.
  • Physical inspection completed.
  • Findings documented.

7. Escalation Criteria

Escalate the case to Supermicro Support if:

  • GPUs are still not detected after reseating.
  • NVIDIA Field Diagnostic continues to fail.
  • Physical damage is observed.
  • Hardware fault LEDs remain illuminated.

Provide:

  • FD Report
  • NVIDIA Bug Report
  • Command Outputs
  • Screenshots
  • Activity Summary

8. Post-Activity

  • Restore the server to its original operational state.
  • Confirm server accessibility with the customer.
  • Share all collected logs with the support team.
  • Inform the customer that diagnostics have been completed.

9. Rollback Plan

If any unexpected issue occurs during diagnostics:

  • Restore all disconnected hardware.
  • Ensure the GPU tray is properly installed.
  • Boot the server to its previous operational state.
  • Inform the customer of the current status.
  • Escalate immediately to Supermicro Support before performing any additional corrective actions.

 

References:

Server manual: https://www.supermicro.com/en/products/system/ai_training/4u/sys-421ge-tnhr2-lcc?mlg=0&utm


    • Related Articles

    • Nvidia-smi Tool Alert - Troubleshooting Guide

      Purpose This document provides a standardized approach for handling and troubleshooting the Nvidia-smi Tool alert, including L1 and L2 actions, escalation guidelines, and resolution criteria. Alert Name Nvidia-smi Tool Alert Description This alert is ...
    • Using dmesg and Kernel Module Checks to Troubleshoot NVIDIA GPU Issues

      Overview This article outlines how to use dmesg logs and kernel module commands to diagnose issues where the operating system fails to detect NVIDIA GPUs—even though they appear in the system's BMC (Baseboard Management Controller). When to Use This ...
    • GPU Memory Mismatch – Troubleshooting Guide

      Alert Information Alert Name: GPU Memory Mismatch Severity: P3 – Medium Impact Alert Description: This alert is generated when the detected GPU memory size on a node does not match the expected configuration. Alert Summary The monitoring system has ...
    • NVIDIA Field Diagnostics Test (FDIAG) Tool for Validating H100/H200

      FDIAG NVIDIA’s Field Diagnostics Test Tool (FDIAG)is an offline hardware test tool for GPU health test. Mainly used in RMA workflow and tool is not provided to public for download. They are distributed through OEMs or directly from NVIDIA support as ...
    • How to collect diagnostic logs using the NetApp Log Collection Script

      1. Purpose This document describes the procedure to collect diagnostic logs using the NetApp Log Collection Script in environments running: BeeGFS NetApp E-Series backend storage HA cluster using Pacemaker and Corosync This script is typically ...