The objective of this Method of Procedure (MOP) is to safely replace the defective PCIe Gen5 Switch Board in the Supermicro server while minimizing system downtime and ensuring all PCIe devices, including GPUs, NICs, NVMe drives, and other connected devices, are properly detected after replacement.
The procedure also includes pre-maintenance validation, hardware replacement, post-maintenance verification, and rollback steps.
This procedure applies to:
Supermicro server
PCIe Gen5 Switch Board FRU replacement
Systems experiencing:
Missing GPUs
PCIe device enumeration failures
PCIe switch faults
PCIe Bus errors
Hardware failure confirmed by OEM Support
Faulty AOM
Model Number | Serial Number |
SYS-821GE-TNHR | A514359X5902573 |
SYS-821GE-TNHR | A514359X5A01822 |
SYS-821GE-TNHR | A514359X5A03335 |
This procedure shall be performed only during an approved maintenance window.
Before beginning the activity, ensure the following:
Obtain customer approval.
Confirm maintenance window.
Notify stakeholders of expected downtime.
Verify remote monitoring alerts are acknowledged.
Verify server model.
Verify Serial Number.
Verify replacement PCIe Gen5 Switch Board FRU.
Confirm replacement part is compatible with the server.
Collect the following before shutdown:
BIOS Version
BMC Firmware Version
CPLD Version
GPU Firmware (if applicable)
Ensure customer confirms application shutdown.
Confirm no active production workloads.
Backup critical logs if required.
Replacement PCIe Gen5 Switch Board
Phillips Screwdriver
Torque screwdriver (if applicable)
ESD Wrist Strap
ESD Mat
Flashlight
Labels for cable identification
BMC/IPMI Access
SSH Access
NVIDIA Field Diagnostic (if GPUs installed)
IPMITool
Linux Shell Access

Useful commands:
Contact customer.
Confirm approved maintenance window.
Verify workloads have been stopped.
Obtain approval to power down the server.
Inspect the server.
Verify:
No active alarms
No visible physical damage
Power supplies healthy
Cooling fans operational
Front and rear LEDs
Record any abnormalities.
Verify access to:
BMC
Operating System
SSH
Console
KVM
Ensure administrative credentials are available.
Before shutting down the server collect:
Verify:
Sensors
Event Logs
PCIe Errors
Hardware faults
Run:
lspci
Verify PCIe topology.
Run:
nvidia-smi
Verify GPU count.
Run:
dmesg
Look for:
PCIe errors
AER errors
Link failures
Enumeration failures
Collect:
BMC SEL
System Logs
GPU logs
If the issue is intermittent and recommended by OEM Support:
Remove GPU Tray.
Inspect connector.
Inspect guide rails.
Check for contamination.
Reinsert tray.
Verify locking mechanism.
Power on system.
Verify GPU detection.
If issue is resolved, no replacement is required.
Otherwise continue.
Perform complete AC Power Cycle.
Shutdown OS.
Power Off.
Disconnect AC power.
Wait 2 minutes.
Reconnect power.
Boot server.
Verify:
lspci
and
nvidia-smi
If issue remains proceed to hardware replacement.
Shutdown operating system gracefully.
Verify:
Power LED OFF
Fans stopped
Disconnect:
AC Power Cords
Wait approximately two minutes.
Wear ESD wrist strap.
Connect to grounded surface.
Place removed hardware on ESD mat.
Remove top cover.
Place screws safely.
Carefully label and disconnect all cables connected to the PCIe Gen5 Switch Board.
Typical connections may include:
Power cables
PCIe interconnect cables
Signal cables
Management cables
Do not pull cables by the wires.
Remove mounting screws.
Support board while removing.
Lift vertically.
Avoid damaging motherboard connectors.
Inspect:
Connector pins
Standoffs
Cable connectors
Install replacement PCIe Gen5 Switch Board.
Ensure:
Board is fully seated
No connector misalignment
Mounting holes aligned
Secure screws.
Reconnect all previously labeled cables.
Verify:
No loose cables
No pinched cables
Proper cable routing
Verify:
GPU trays seated
PCIe risers secured
No foreign objects
Fan cables connected
Air shrouds installed correctly
Install top cover.
Reconnect:
AC power
Network cables
Management cables
Power on server.
Monitor:
POST
BIOS
BMC
Verify no POST errors.
If GPUs remain undetected after replacement:
Check:
lspci
nvidia-smi
Review:
BIOS PCIe configuration
BMC inventory
PCIe link status
Perform:
GPU re-seat
PCIe cable inspection
OEM troubleshooting
If unresolved:
Escalate to Supermicro Engineering.
Collect:
SEL Log
Hardware Inventory
Sensor Logs
dmesg
journalctl
lspci -vv
nvidia-smi -q
If applicable:
Run NVIDIA Field Diagnostic.
Attach all logs to the support case.
Record:
Maintenance Start Time
Maintenance End Time
Server Serial Number
Removed FRU Serial Number
Installed FRU Serial Number
Engineer Name
Customer Representative
Test Results
Final System Health
Capture photographs of:
Installed board
Completed installation
After replacement:
PCIe Switch Board detected successfully
All GPUs detected
PCIe devices enumerated correctly
No BMC hardware faults
No PCIe AER errors
Server boots normally
Customer applications restored
System returned to production
Escalate to Supermicro Technical Support if:
PCIe Switch Board not detected
Multiple GPUs remain missing
PCIe Bus enumeration failure persists
Server fails POST
BMC reports hardware faults
BIOS cannot detect PCIe switch
Physical connector damage observed
Firmware incompatibility suspected
Provide:
Logs
Photographs
Serial numbers
Diagnostic reports
Error screenshots
Perform final validation.
Verify:
No warning LEDs
Fans operating normally
Temperature within limits
No active alarms
Sensor status Normal
Event log reviewed
Run:
lspci
Confirm expected PCIe topology.
Run:
nvidia-smi
Confirm all installed GPUs are detected and healthy.
Optionally execute:
NVIDIA Field Diagnostics
GPU stress test
Burn-in test
Customer application verification
Obtain customer confirmation before closing the maintenance activity.
If the replacement does not resolve the issue or introduces new faults:
Power down the server gracefully.
Disconnect all AC power sources.
Remove the replacement PCIe Gen5 Switch Board.
Reinstall the original PCIe Gen5 Switch Board (if confirmed functional and safe to reuse).
Reconnect all internal cables as originally labeled.
Reassemble the server and restore power.
Verify BIOS, BMC, and operating system functionality.
Confirm PCIe device enumeration and GPU detection.
If the issue persists after rollback, collect updated diagnostic logs and escalate to Supermicro Engineering with complete findings and evidence.