Execution of NVIDIA Field Diagnostic (FD) Tool and Collection of Diagnostic Logs, and troubleshooting of GPU(s) not detected

KB 829939 - Execution of NVIDIA Field Diagnostic (FD) Tool and Collection of Diagnostic Logs, and troubleshooting of GPU(s) not detected

1. Objective

To provide a standardized procedure for executing the NVIDIA Field Diagnostic (FD) tool, verifying GPU status, collecting the required diagnostic logs, and documenting the findings for further analysis.

 

2. Scope

This procedure applies to field engineers performing GPU diagnostics on Supermicro GPU servers where GPU related issues have been reported and execution of the NVIDIA Field Diagnostic (FD) tool for health verification or troubleshooting.

 

3. Prerequisites

Before starting the activity, ensure the following:

  • Coordinate with the customer and confirm maintenance window.
  • Customer availability and site access confirmed.
  • Verify the server hostname, rack location, and serial number (if available).
  • Ensure the required NVIDIA Field Diagnostic (FD) package is available.
  • Carry a laptop with SSH access utilities.
  • Verify access credentials will be available during the activity.
  • Sufficient storage space is available to save the generated reports.

 

4. Required Tools

Hardware

  • Laptop
  • Network cable (if required)

Software

  • NVIDIA Field Diagnostic (FD) Tool
  • SSH Client
  • IPMI Utilities
  • NVIDIA CLI Utilities
  • Linux Shell Access

 

5. Activity Procedure

Step 1 – Customer Coordination

  1. Meet the customer representative.
  2. Verify the maintenance activity.
  3. Identify the affected server.
  4. Verify the server hostname, serial number, or asset tag.

Step 2 – Physical Verification

  1. Inspect the server for any visible hardware abnormalities.
  2. Check for amber fault LEDs on the server and GPU tray

 

Step 3 – Access Verification

  1. Obtain SSH or console access to the operating system.
  2. Verify login access.

 

Step 4 – Initial Diagnostics

Execute the NVIDIA Field Diagnostic (FD) Tool.

  • Run the complete Field Diagnostic.
  • Save the diagnostic report.
  • Record the execution time and outcome.

Step 5 – GPU Tray Re-seat (if required)

If instructed by the Supermicro support team:

 

GPU Tray

The GPU tray is located at the top of the chassis. It houses the system's GPUs, cooling components, and other supporting devices. The GPU tray is accessible from the front of the unit by sliding the upper tray out after disengaging a lever on the left and right side of the chassis.

 

Notes:

1.      Lift the heavy universal baseboard (UBB) plate by using two UBB handles.

2.      You should extend only one tray in a rack at a time - extending two trays simultaneously may cause the server to become unstable. Failure to stabilize the server can cause the server to tip over.

Warning! Tip over hazard when not in a rack!

 

1.      Gracefully shut down the server (if required).

2.      Remove the GPU tray.

3.      Inspect the tray for:

    • Physical damage
    • Bent pins
    • Loose connections

4.      Reinsert the GPU tray carefully.

  1. Ensure the tray is fully seated and locked.
  2. Power on the server.

Step 6 – Cold Reboot (if required)

If instructed by the Supermicro support team:

Perform a complete cold reboot.

After the system becomes operational:

  • Verify all GPUs are detected.
  • Allow sufficient time for the operating system and NVIDIA services to initialize.

 

Step 7 – Execute Field Diagnostic

Run the NVIDIA Field Diagnostic Tool.

Verify:

  • GPU detection status
  • Diagnostic result
  • Any reported failures

Save the updated diagnostic report.

 

Step 8 – If GPUs Are Still Missing

If one or more GPUs are not detected:

Document the following evidence:

  • GPU tray
  • Connector side
  • Pin side
  • Any damaged components
  • Fault LEDs (if present)

Screenshots

  • GPU detection failure
  • Diagnostic output
  • Error messages

 

Step 9 – Collect Diagnostic Logs

Execute the following commands and save the outputs.

nvidia-smi -q

nvidia-smi nvlink -s

nvidia-bug-report.sh

lspci -tv | grep -i nvidia

systemctl status nvidia-fabricmanager


For IPMI Logs

Prerequisite: ipmitool must be available on the system.

Use below commands to collect the logs:
In Linux system:
./ipmitool -I lan -H <BMC_IP> -U <USERNAME> -P '<PASSWORD>'
ipmitool fru
ipmitool sel list

In Windows system:
ipmitool.exe -I lan -H <BMC_IP> -U <USERNAME> -P "<PASSWORD>" 

ipmitool fru

ipmitool sel list


Collect:

  • NVIDIA Bug Report
  • Field Diagnostic Report
  • Command outputs
  • Screenshots (if applicable)

 

Step 10 – Documentation

Document the following:

  • Server Serial Number
  • Rack Location
  • Date and Time of Activity
  • Number of GPUs Detected
  • Amber LED Status
  • Field Diagnostic Result
  • Actions Performed

Attach:

  • Diagnostic Reports
  • Log Files
  • Screenshots

6. Expected Outcome

  • NVIDIA Field Diagnostic executed successfully.
  • GPU detection verified.
  • Required logs collected.
  • Physical inspection completed.
  • Findings documented.

7. Escalation Criteria

Escalate the case to Supermicro Support if:

  • GPUs are still not detected after reseating.
  • NVIDIA Field Diagnostic continues to fail.
  • Physical damage is observed.
  • Hardware fault LEDs remain illuminated.

Provide:

  • FD Report
  • NVIDIA Bug Report
  • Command Outputs
  • Screenshots
  • Activity Summary

8. Post-Activity

  • Restore the server to its original operational state.
  • Confirm server accessibility with the customer.
  • Share all collected logs with the support team.
  • Inform the customer that diagnostics have been completed.

9. Rollback Plan

If any unexpected issue occurs during diagnostics:

  • Restore all disconnected hardware.
  • Ensure the GPU tray is properly installed.
  • Boot the server to its previous operational state.
  • Inform the customer of the current status.
  • Escalate immediately to Supermicro Support before performing any additional corrective actions.

 

References:

Server manual link: https://mbuzztech.sharepoint.com/:b:/s/ITServices/IQAs_FOeVW5pRbj6-66JurYXAUgUXY_3VjtF87QkjcShA1I?e=FxngYQ


    • Recent Articles

    • KB-630208 | How to use Find and Locate to search for files in Linux

      Purpose To provide Linux administrators and Field Engineers with guidance on using the find and locate utilities to efficiently search for files and directories within a Linux filesystem based on criteria such as name, type, size, timestamps, ...
    • KB-630164 | How to use rsync to Synchronize Files

      Purpose To provide a clear and practical procedure for using the rsync utility to synchronize files and directories between local systems and remote hosts, including push and pull operations, while preserving file attributes and optimizing data ...
    • KB-630107 | How to create a Linux swap file

      Purpose To provide Field Engineers and Linux administrators with a standardized procedure for creating, configuring, and validating a Linux swap file. This procedure helps ensure sufficient virtual memory is available when physical RAM is exhausted ...
    • MSI All-in-One (AIO) PC – Physical Damage Policy for Warranty Claims

      MSI All-in-One (AIO) PC – Physical Damage Policy for Warranty Claims MSI's standard warranty does not cover physical or accidental damage to an All-in-One (AIO) PC. If the damage is determined to be caused by external factors rather than a ...
    • KB 134833 - Troubleshooting NVSM Alert NV-CPU-XX – Unrecoverable CPU Internal Error

      Purpose This document provides a general troubleshooting procedure for the NVIDIA System Management (NVSM) alert NV-CPU-XX, which indicates that a CPU has reported an internal error. The article outlines how to verify whether the alert represents an ...
    • Popular Articles

    • CP Plus Camera and NVR Configuration

      NVR Configuration The CP Plus Pro Series of NVRs have been meticulously designed for providing you with upgraded performance and higher recording quality in your IP video surveillance solution. The robust processor that has been inculcated in this ...
    • KB 775235 - Kerberos Authentication – Overview

      What is Kerberos? Kerberos is a secure authentication method used in our Active Directory (AD) environment (mbuzztech.com). It allows users to: Access multiple systems without re-entering passwords (Single Sign-On – SSO) Log in once Where We Use It ...
    • How to Remove and Reinstall NVIDIA Drivers on Ubuntu

      This article provides step by step guide to completely remove existing NVIDIA drivers and reinstall specific version of the NVIDIA driver on the Ubuntu system Prerequisites Administrative (sudo) access to the Ubuntu system. Internet access to ...
    • KB 692001 - Personal Computers and Servers - Classification and Point of Contact

      We can classify the computers that MBUZZ handles based on their form-factor as below: Tower Workstations, Desktops, Gaming PCs and SFF (Small form factor) PCs fall under this category. These are computers people would use on a desk and rarely move. ...
    • KB 298031 - M.2 SSD Tier List

      The sequential read and write speeds, which are usually the most advertised number, are not a proper benchmark of real-world performance or the quality of an SSD. This article categorizes and tiers SSDs based on factors like the type of NAND flash, ...