1. Objective
To provide a standardized procedure for executing the NVIDIA
Field Diagnostic (FD) tool, verifying GPU status, collecting the required
diagnostic logs, and documenting the findings for further analysis.
2. Scope
This procedure applies to field engineers performing GPU
diagnostics on Supermicro GPU servers where GPU related issues have been
reported and execution of the NVIDIA Field Diagnostic (FD) tool for health
verification or troubleshooting.
3. Prerequisites
Before starting the activity, ensure the following:
- Coordinate
with the customer and confirm maintenance window.
- Customer
availability and site access confirmed.
- Verify
the server hostname, rack location, and serial number (if available).
- Ensure
the required NVIDIA Field Diagnostic (FD) package is available.
- Carry
a laptop with SSH access utilities.
- Verify
access credentials will be available during the activity.
- Sufficient
storage space is available to save the generated reports.
Hardware
- Laptop
- Network
cable (if required)
Software
- NVIDIA
Field Diagnostic (FD) Tool
- SSH
Client
- IPMI
Utilities
- NVIDIA
CLI Utilities
- Linux
Shell Access
5. Activity Procedure
Step 1 – Customer Coordination
- Meet
the customer representative.
- Verify
the maintenance activity.
- Identify
the affected server.
- Verify
the server hostname, serial number, or asset tag.
Step 2 – Physical Verification
- Inspect
the server for any visible hardware abnormalities.
- Check
for amber fault LEDs on the server and GPU tray
Step 3 – Access Verification
- Obtain
SSH or console access to the operating system.
- Verify
login access.
Step 4 – Initial Diagnostics
Execute the NVIDIA Field Diagnostic (FD) Tool.
- Run
the complete Field Diagnostic.
- Save
the diagnostic report.
- Record
the execution time and outcome.
Step 5 – GPU Tray Re-seat (if required)
If instructed by the Supermicro support team:
GPU Tray
The GPU tray is located at the top of the chassis. It houses
the system's GPUs, cooling components, and other supporting devices. The GPU
tray is accessible from the front of the unit by sliding the upper tray out
after disengaging a lever on the left and right side of the chassis.


Notes:
1. Lift
the heavy universal baseboard (UBB) plate by using two UBB handles.

2. You
should extend only one tray in a rack at a time - extending two trays
simultaneously may cause the server to become unstable. Failure to stabilize
the server can cause the server to tip over.

Warning! Tip over
hazard when not in a rack!
1. Gracefully shut down the server (if required).
2. Remove the GPU tray.
3. Inspect the tray for:
- Physical
damage
- Bent
pins
- Loose
connections
4. Reinsert the GPU tray carefully.
- Ensure the tray is fully
seated and locked.
- Power on the server.
Step 6 – Cold Reboot (if required)
If instructed by the Supermicro support team:
Perform a complete cold reboot.
After the system becomes operational:
- Verify
all GPUs are detected.
- Allow
sufficient time for the operating system and NVIDIA services to
initialize.
Step 7 – Execute Field
Diagnostic
Run the NVIDIA Field Diagnostic Tool.
Verify:
- GPU
detection status
- Diagnostic
result
- Any
reported failures
Save the updated diagnostic report.
Step 8 – If GPUs Are Still Missing
If one or more GPUs are not detected:
Document the following evidence:
- GPU
tray
- Connector
side
- Pin
side
- Any
damaged components
- Fault
LEDs (if present)
Screenshots
- GPU
detection failure
- Diagnostic
output
- Error
messages
Step 9 – Collect Diagnostic Logs
Execute the following commands and save the outputs.
nvidia-smi
-q
nvidia-smi
nvlink -s
nvidia-bug-report.sh
lspci
-tv | grep -i nvidia
systemctl
status nvidia-fabricmanager
For IPMI Logs
Prerequisite: ipmitool must be available on the system. If it is not installed in the system, use below link to download it.
(Select OS as per requirement)
After installing the ipmi tool in system, use below commands to collect the logs:
In Linux system:
./ipmitool -I lan -H <BMC_IP> -U <USERNAME> -P '<PASSWORD>'
ipmitool fru
ipmitool sel list
In Windows system:
ipmitool.exe -I lan -H <BMC_IP> -U <USERNAME> -P "<PASSWORD>"
ipmitool
fru
ipmitool
sel list
Collect:
- NVIDIA
Bug Report
- Field
Diagnostic Report
- Command
outputs
- Screenshots
(if applicable)
Step 10 – Documentation
Document the following:
- Server
Serial Number
- Rack
Location
- Date
and Time of Activity
- Number
of GPUs Detected
- Amber
LED Status
- Field
Diagnostic Result
- Actions
Performed
Attach:
- Diagnostic
Reports
- Log
Files
- Screenshots
6. Expected Outcome
- NVIDIA
Field Diagnostic executed successfully.
- GPU
detection verified.
- Required
logs collected.
- Physical
inspection completed.
- Findings
documented.
7. Escalation Criteria
Escalate the case to Supermicro Support if:
- GPUs
are still not detected after reseating.
- NVIDIA
Field Diagnostic continues to fail.
- Physical
damage is observed.
- Hardware
fault LEDs remain illuminated.
Provide:
- FD
Report
- NVIDIA
Bug Report
- Command
Outputs
- Screenshots
- Activity
Summary
8. Post-Activity
- Restore
the server to its original operational state.
- Confirm
server accessibility with the customer.
- Share
all collected logs with the support team.
- Inform
the customer that diagnostics have been completed.
9. Rollback Plan
If any unexpected issue occurs during diagnostics:
- Restore
all disconnected hardware.
- Ensure
the GPU tray is properly installed.
- Boot
the server to its previous operational state.
- Inform
the customer of the current status.
- Escalate
immediately to Supermicro Support before performing any additional
corrective actions.
References:
Server manual: https://www.supermicro.com/en/products/system/ai_training/4u/sys-421ge-tnhr2-lcc?mlg=0&utm