Related Articles
Export Control for NVDIA GPUs
The Export Control Classification Number (ECCN) is part of the U.S. export control regulations that apply to certain advanced computing items, including some of NVIDIA’s products. This classification is used to identify items that require specific ...
Powering NVIDIA L40S GPUs in ASUS ESC8000A-E12P Servers
Powering NVIDIA L40S GPUs in ASUS ESC8000A-E12P Servers Overview When deploying NVIDIA L40S GPUs in ASUS ESC8000A-E12P servers, users may encounter a power connector mismatch issue. This KB article explains the root cause and provides the official ...
CPU vs GPU vs TPU
CPU vs GPU vs TPU CPU is the processing unit that works as the brains of a computer designed to be ideal for general-purpose programming. GPU is a performance accelerator that enhances computer graphics and AI workloads. TPUs are Google's ...
NVIDIA L20 vs. L40: A Deep Dive into Performance and Capabilities for AI and Professional Workloads
When it comes to choosing the right GPU for AI and professional workloads, NVIDIA's L20 and L40, based on the Ada Lovelace architecture, are both stellar options. As someone who is spent considerable time exploring these technologies, I have found ...
Execution of NVIDIA Field Diagnostic (FD) Tool and Collection of Diagnostic Logs, and troubleshooting of GPU(s) not detected
1. Objective To provide a standardized procedure for executing the NVIDIA Field Diagnostic (FD) tool, verifying GPU status, collecting the required diagnostic logs, and documenting the findings for further analysis. 2. Scope This procedure applies to ...
Recent Articles
Troubleshooting NVSM Alert NV-CPU-XX – Unrecoverable CPU Internal Error
Purpose This document provides a general troubleshooting procedure for the NVIDIA System Management (NVSM) alert NV-CPU-XX, which indicates that a CPU has reported an internal error. The article outlines how to verify whether the alert represents an ...
PCIe Gen5 Switch Board Replacement
1. Objective The objective of this Method of Procedure (MOP) is to safely replace the defective PCIe Gen5 Switch Board in the Supermicro server while minimizing system downtime and ensuring all PCIe devices, including GPUs, NICs, NVMe drives, and ...
Local Boot Support for DGX H200 with BCM 11
Overview This Knowledge Base (KB) article explains the supported method for deploying and managing a DGX H200 system using Bright Cluster Manager (BCM) 11 while booting the operating system from the node's local NVMe storage. To be managed by BCM, ...
Execution of NVIDIA Field Diagnostic (FD) Tool and Collection of Diagnostic Logs, and troubleshooting of GPU(s) not detected
1. Objective To provide a standardized procedure for executing the NVIDIA Field Diagnostic (FD) tool, verifying GPU status, collecting the required diagnostic logs, and documenting the findings for further analysis. 2. Scope This procedure applies to ...
Fiber Optic Bend Radius Measurement and Compliance
1. Purpose This article outlines the procedure for verifying that installed fiber optic cables comply with minimum bend radius requirements. Proper verification prevents signal degradation, ensures optimal optical performance, and protects the ...