Execution of NVIDIA Field Diagnostic (FD) Tool and Collection of Diagnostic Logs, and troubleshooting of GPU(s) not detected

KB 829939 - Execution of NVIDIA Field Diagnostic (FD) Tool and Collection of Diagnostic Logs, and troubleshooting of GPU(s) not detected

1. Objective

To provide a standardized procedure for executing the NVIDIA Field Diagnostic (FD) tool, verifying GPU status, collecting the required diagnostic logs, and documenting the findings for further analysis.

 

2. Scope

This procedure applies to field engineers performing GPU diagnostics on Supermicro GPU servers where GPU related issues have been reported and execution of the NVIDIA Field Diagnostic (FD) tool for health verification or troubleshooting.

 

3. Prerequisites

Before starting the activity, ensure the following:

  • Coordinate with the customer and confirm maintenance window.
  • Customer availability and site access confirmed.
  • Verify the server hostname, rack location, and serial number (if available).
  • Ensure the required NVIDIA Field Diagnostic (FD) package is available.
  • Carry a laptop with SSH access utilities.
  • Verify access credentials will be available during the activity.
  • Sufficient storage space is available to save the generated reports.

 

4. Required Tools

Hardware

  • Laptop
  • Network cable (if required)

Software

  • NVIDIA Field Diagnostic (FD) Tool
  • SSH Client
  • IPMI Utilities
  • NVIDIA CLI Utilities
  • Linux Shell Access

 

5. Activity Procedure

Step 1 – Customer Coordination

  1. Meet the customer representative.
  2. Verify the maintenance activity.
  3. Identify the affected server.
  4. Verify the server hostname, serial number, or asset tag.

Step 2 – Physical Verification

  1. Inspect the server for any visible hardware abnormalities.
  2. Check for amber fault LEDs on the server and GPU tray

 

Step 3 – Access Verification

  1. Obtain SSH or console access to the operating system.
  2. Verify login access.

 

Step 4 – Initial Diagnostics

Execute the NVIDIA Field Diagnostic (FD) Tool.

  • Run the complete Field Diagnostic.
  • Save the diagnostic report.
  • Record the execution time and outcome.

Step 5 – GPU Tray Re-seat (if required)

If instructed by the Supermicro support team:

 

GPU Tray

The GPU tray is located at the top of the chassis. It houses the system's GPUs, cooling components, and other supporting devices. The GPU tray is accessible from the front of the unit by sliding the upper tray out after disengaging a lever on the left and right side of the chassis.

 

Notes:

1.      Lift the heavy universal baseboard (UBB) plate by using two UBB handles.

2.      You should extend only one tray in a rack at a time - extending two trays simultaneously may cause the server to become unstable. Failure to stabilize the server can cause the server to tip over.

Warning! Tip over hazard when not in a rack!

 

1.      Gracefully shut down the server (if required).

2.      Remove the GPU tray.

3.      Inspect the tray for:

    • Physical damage
    • Bent pins
    • Loose connections

4.      Reinsert the GPU tray carefully.

  1. Ensure the tray is fully seated and locked.
  2. Power on the server.

Step 6 – Cold Reboot (if required)

If instructed by the Supermicro support team:

Perform a complete cold reboot.

After the system becomes operational:

  • Verify all GPUs are detected.
  • Allow sufficient time for the operating system and NVIDIA services to initialize.

 

Step 7 – Execute Field Diagnostic

Run the NVIDIA Field Diagnostic Tool.

Verify:

  • GPU detection status
  • Diagnostic result
  • Any reported failures

Save the updated diagnostic report.

 

Step 8 – If GPUs Are Still Missing

If one or more GPUs are not detected:

Document the following evidence:

  • GPU tray
  • Connector side
  • Pin side
  • Any damaged components
  • Fault LEDs (if present)

Screenshots

  • GPU detection failure
  • Diagnostic output
  • Error messages

 

Step 9 – Collect Diagnostic Logs

Execute the following commands and save the outputs.

nvidia-smi -q

nvidia-smi nvlink -s

nvidia-bug-report.sh

lspci -tv | grep -i nvidia

systemctl status nvidia-fabricmanager


For IPMI Logs

Prerequisite: ipmitool must be available on the system.

Use below commands to collect the logs:
In Linux system:
./ipmitool -I lan -H <BMC_IP> -U <USERNAME> -P '<PASSWORD>'
ipmitool fru
ipmitool sel list

In Windows system:
ipmitool.exe -I lan -H <BMC_IP> -U <USERNAME> -P "<PASSWORD>" 

ipmitool fru

ipmitool sel list


Collect:

  • NVIDIA Bug Report
  • Field Diagnostic Report
  • Command outputs
  • Screenshots (if applicable)

 

Step 10 – Documentation

Document the following:

  • Server Serial Number
  • Rack Location
  • Date and Time of Activity
  • Number of GPUs Detected
  • Amber LED Status
  • Field Diagnostic Result
  • Actions Performed

Attach:

  • Diagnostic Reports
  • Log Files
  • Screenshots

6. Expected Outcome

  • NVIDIA Field Diagnostic executed successfully.
  • GPU detection verified.
  • Required logs collected.
  • Physical inspection completed.
  • Findings documented.

7. Escalation Criteria

Escalate the case to Supermicro Support if:

  • GPUs are still not detected after reseating.
  • NVIDIA Field Diagnostic continues to fail.
  • Physical damage is observed.
  • Hardware fault LEDs remain illuminated.

Provide:

  • FD Report
  • NVIDIA Bug Report
  • Command Outputs
  • Screenshots
  • Activity Summary

8. Post-Activity

  • Restore the server to its original operational state.
  • Confirm server accessibility with the customer.
  • Share all collected logs with the support team.
  • Inform the customer that diagnostics have been completed.

9. Rollback Plan

If any unexpected issue occurs during diagnostics:

  • Restore all disconnected hardware.
  • Ensure the GPU tray is properly installed.
  • Boot the server to its previous operational state.
  • Inform the customer of the current status.
  • Escalate immediately to Supermicro Support before performing any additional corrective actions.

 

References:

Server manual link: https://mbuzztech.sharepoint.com/:b:/s/ITServices/IQAs_FOeVW5pRbj6-66JurYXAUgUXY_3VjtF87QkjcShA1I?e=FxngYQ