PCIe Gen5 Switch Board Replacement

PCIe Gen5 Switch Board Replacement


1. Objective

The objective of this Method of Procedure (MOP) is to safely replace the defective PCIe Gen5 Switch Board in the Supermicro server while minimizing system downtime and ensuring all PCIe devices, including GPUs, NICs, NVMe drives, and other connected devices, are properly detected after replacement.

The procedure also includes pre-maintenance validation, hardware replacement, post-maintenance verification, and rollback steps.


2. Scope

This procedure applies to:

  • Supermicro server

  • PCIe Gen5 Switch Board FRU replacement

  • Systems experiencing:

    • Missing GPUs

    • PCIe device enumeration failures

    • PCIe switch faults

    • PCIe Bus errors

    • Hardware failure confirmed by OEM Support

    • Faulty AOM


Model Number
Serial Number
SYS-821GE-TNHR
A514359X5902573
SYS-821GE-TNHR
A514359X5A01822
SYS-821GE-TNHR
A514359X5A03335


This procedure shall be performed only during an approved maintenance window.


3. Prerequisites

Before beginning the activity, ensure the following:

Customer Coordination

  • Obtain customer approval.

  • Confirm maintenance window.

  • Notify stakeholders of expected downtime.

  • Verify remote monitoring alerts are acknowledged.

Hardware Verification

  • Verify server model.

  • Verify Serial Number.

  • Verify replacement PCIe Gen5 Switch Board FRU.

  • Confirm replacement part is compatible with the server.

Software Verification

Collect the following before shutdown:

  • BIOS Version

  • BMC Firmware Version

  • CPLD Version

  • GPU Firmware (if applicable)

Backup

  • Ensure customer confirms application shutdown.

  • Confirm no active production workloads.

  • Backup critical logs if required.


4. Required Tools

Hardware

  • Replacement PCIe Gen5 Switch Board

  • Phillips Screwdriver

  • Torque screwdriver (if applicable)

  • ESD Wrist Strap

  • ESD Mat

  • Flashlight

  • Labels for cable identification

Software

  • BMC/IPMI Access

  • SSH Access

  • NVIDIA Field Diagnostic (if GPUs installed)

  • IPMITool

  • Linux Shell Access


Quote

Useful commands:

lspci
nvidia-smi
dmidecode
dmesg
journalctl
ipmitool

5. Activity Procedure


Step 1 – Customer Coordination

  • Contact customer.

  • Confirm approved maintenance window.

  • Verify workloads have been stopped.

  • Obtain approval to power down the server.


Step 2 – Physical Verification

Inspect the server.

Verify:

  • No active alarms

  • No visible physical damage

  • Power supplies healthy

  • Cooling fans operational

  • Front and rear LEDs

Record any abnormalities.


Step 3 – Access Verification

Verify access to:

  • BMC

  • Operating System

  • SSH

  • Console

  • KVM

Ensure administrative credentials are available.


Step 4 – Initial Diagnostics

Before shutting down the server collect:

BMC Health

Verify:

  • Sensors

  • Event Logs

  • PCIe Errors

  • Hardware faults

Linux Validation

Run:

lspci

Verify PCIe topology.

Run:

nvidia-smi

Verify GPU count.

Run:

dmesg

Look for:

  • PCIe errors

  • AER errors

  • Link failures

  • Enumeration failures

Collect:

  • BMC SEL

  • System Logs

  • GPU logs


Step 5 – GPU Tray Re-seat (if required)

If the issue is intermittent and recommended by OEM Support:

  • Remove GPU Tray.

  • Inspect connector.

  • Inspect guide rails.

  • Check for contamination.

  • Reinsert tray.

  • Verify locking mechanism.

Power on system.

Verify GPU detection.

If issue is resolved, no replacement is required.

Otherwise continue.


Step 6 – Cold Reboot (if required)

Perform complete AC Power Cycle.

  1. Shutdown OS.

  2. Power Off.

  3. Disconnect AC power.

  4. Wait 2 minutes.

  5. Reconnect power.

  6. Boot server.

Verify:

lspci

and

nvidia-smi

If issue remains proceed to hardware replacement.


Step 7 – PCIe Gen5 Switch Board Replacement

Power Down

Shutdown operating system gracefully.

Verify:

  • Power LED OFF

  • Fans stopped

Disconnect:

  • AC Power Cords

Wait approximately two minutes.


Motherboard Tray: Top View


ESD Protection

  • Wear ESD wrist strap.

  • Connect to grounded surface.

  • Place removed hardware on ESD mat.


Remove Server Cover

  • Remove top cover.

  • Place screws safely.


Disconnect Internal Cables

Carefully label and disconnect all cables connected to the PCIe Gen5 Switch Board.

Typical connections may include:

  • Power cables

  • PCIe interconnect cables

  • Signal cables

  • Management cables

Do not pull cables by the wires.


Remove Existing Switch Board

  • Remove mounting screws.

  • Support board while removing.

  • Lift vertically.

  • Avoid damaging motherboard connectors.

Inspect:

  • Connector pins

  • Standoffs

  • Cable connectors


Install Replacement Board

Install replacement PCIe Gen5 Switch Board.

Ensure:

  • Board is fully seated

  • No connector misalignment

  • Mounting holes aligned

Secure screws.

Reconnect all previously labeled cables.

Verify:

  • No loose cables

  • No pinched cables

  • Proper cable routing


Internal Inspection

Verify:

  • GPU trays seated

  • PCIe risers secured

  • No foreign objects

  • Fan cables connected

  • Air shrouds installed correctly


Reassemble Server

Install top cover.

Reconnect:

  • AC power

  • Network cables

  • Management cables


Power On

Power on server.

Monitor:

  • POST

  • BIOS

  • BMC

Verify no POST errors.


Step 8 – If GPUs Are Still Missing

If GPUs remain undetected after replacement:

Check:

lspci
nvidia-smi

Review:

  • BIOS PCIe configuration

  • BMC inventory

  • PCIe link status

Perform:

  • GPU re-seat

  • PCIe cable inspection

  • OEM troubleshooting

If unresolved:

Escalate to Supermicro Engineering.


Step 9 – Collect Diagnostic Logs

Collect:

BMC

  • SEL Log

  • Hardware Inventory

  • Sensor Logs

Operating System

dmesg
journalctl
lspci -vv
nvidia-smi -q

If applicable:

Run NVIDIA Field Diagnostic.

Attach all logs to the support case.


Step 10 – Documentation

Record:

  • Maintenance Start Time

  • Maintenance End Time

  • Server Serial Number

  • Removed FRU Serial Number

  • Installed FRU Serial Number

  • Engineer Name

  • Customer Representative

  • Test Results

  • Final System Health

Capture photographs of:

  • Installed board

  • Completed installation


6. Expected Outcome

After replacement:

  • PCIe Switch Board detected successfully

  • All GPUs detected

  • PCIe devices enumerated correctly

  • No BMC hardware faults

  • No PCIe AER errors

  • Server boots normally

  • Customer applications restored

  • System returned to production


7. Escalation Criteria

Escalate to Supermicro Technical Support if:

  • PCIe Switch Board not detected

  • Multiple GPUs remain missing

  • PCIe Bus enumeration failure persists

  • Server fails POST

  • BMC reports hardware faults

  • BIOS cannot detect PCIe switch

  • Physical connector damage observed

  • Firmware incompatibility suspected

Provide:

  • Logs

  • Photographs

  • Serial numbers

  • Diagnostic reports

  • Error screenshots


8. Post-Activity

Perform final validation.

Verify:

Hardware

  • No warning LEDs

  • Fans operating normally

  • Temperature within limits

BMC

  • No active alarms

  • Sensor status Normal

  • Event log reviewed

Operating System

Run:

lspci

Confirm expected PCIe topology.

Run:

nvidia-smi

Confirm all installed GPUs are detected and healthy.

Optionally execute:

  • NVIDIA Field Diagnostics

  • GPU stress test

  • Burn-in test

  • Customer application verification

Obtain customer confirmation before closing the maintenance activity.


9. Rollback Plan

If the replacement does not resolve the issue or introduces new faults:

  1. Power down the server gracefully.

  2. Disconnect all AC power sources.

  3. Remove the replacement PCIe Gen5 Switch Board.

  4. Reinstall the original PCIe Gen5 Switch Board (if confirmed functional and safe to reuse).

  5. Reconnect all internal cables as originally labeled.

  6. Reassemble the server and restore power.

  7. Verify BIOS, BMC, and operating system functionality.

  8. Confirm PCIe device enumeration and GPU detection.

  9. If the issue persists after rollback, collect updated diagnostic logs and escalate to Supermicro Engineering with complete findings and evidence.

    • Related Articles

    • Network Switches

      What is a Network Switch? A Switch is a network device that is used to segment the networks into different subnetworks called subnets or LAN segments. It is responsible for filtering and forwarding the packets between LAN segments based on the MAC ...
    • How to do a remote power cycle on NVIDIA QM9700 Switch?

      1. Purpose To perform a remote reboot of NVIDIA QM9700 switch using the NVIDIA's Web GUI. If the remote reboot does not resolve any issues occurred, a physical power-cycle should be carried out onsite as per OEM recommendations. 2. Scope This MOP ...
    • Configuring Management Interface 0 (eth0) on SONiC 4.5.0 Switch via SONiC CLI

      Purpose: This article provides step-by-step instructions to configure the management interface (Management 0 / eth0) on a SONiC 4.5.0 switch with a static IP and default gateway using the SONiC CLI. Scope: Applicable to Edgecore S-series switches ...
    • Cisco Nexus switch(9312) Unexpected Cluster Connectivity Loss

      1. Summary This article describes an incident where a faulty fan module in a Cisco Nexus leaf switch caused physical-layer instability, leading to intermittent connectivity loss for critical cluster nodes. The issue resulted in a cluster-wide access ...
    • Switch Radix

      What is a radix network? A radix network, also known as a butterfly network, is a type of switching network used in parallel computing. It's a non-blocking network that can connect multiple inputs to multiple outputs in a grid-like pattern without ...