Troubleshooting
Troubleshooting NVSM Alert NV-CPU-XX – Unrecoverable CPU Internal Error
Purpose This document provides a general troubleshooting procedure for the NVIDIA System Management (NVSM) alert NV-CPU-XX, which indicates that a CPU has reported an internal error. The article outlines how to verify whether the alert represents an ...
PCIe Gen5 Switch Board Replacement
1. Objective The objective of this Method of Procedure (MOP) is to safely replace the defective PCIe Gen5 Switch Board in the Supermicro server while minimizing system downtime and ensuring all PCIe devices, including GPUs, NICs, NVMe drives, and ...
Local Boot Support for DGX H200 with BCM 11
Overview This Knowledge Base (KB) article explains the supported method for deploying and managing a DGX H200 system using Bright Cluster Manager (BCM) 11 while booting the operating system from the node's local NVMe storage. To be managed by BCM, ...
Execution of NVIDIA Field Diagnostic (FD) Tool and Collection of Diagnostic Logs, and troubleshooting of GPU(s) not detected
1. Objective To provide a standardized procedure for executing the NVIDIA Field Diagnostic (FD) tool, verifying GPU status, collecting the required diagnostic logs, and documenting the findings for further analysis. 2. Scope This procedure applies to ...
Fiber Optic Bend Radius Measurement and Compliance
1. Purpose This article outlines the procedure for verifying that installed fiber optic cables comply with minimum bend radius requirements. Proper verification prevents signal degradation, ensures optimal optical performance, and protects the ...
Fix: Installation Errors for Unsloth Studio on DGX Spark (ARM64)
Problem: When installing Unsloth Studio on DGX Spark systems, the default installation process may fail or lead to ModuleNotFoundError during startup. This is primarily due to the ARM64 (aarch64) architecture of the DGX Spark, which can cause certain ...
Fix: DGX Spark Kernal Panic - OS Reinstall via System Recovery
The Issue : Kernel Panic: VFS Unable to Mount Root FS on Unknown-Block(0,0) This error is one of the more alarming things you can encounter on a Linux-based system. When the DGX Spark throws a kernel panic with the message VFS: Unable to mount root ...
IPMI User Management (Reset, Create, Delete, Privileges) Using ipmitool on Ubuntu
1. Prerequisites Ubuntu system with sudo access and ipmitool installed Local BMC access (/dev/ipmi0) or remote BMC IP Maintenance window recommended Confirm correct User ID before modifying 2. Install and Validate ipmitool 2.1 Install bash sudo apt ...
Kerberos Authentication – Overview
What is Kerberos? Kerberos is a secure authentication method used in our Active Directory (AD) environment (mbuzztech.com). It allows users to: Access multiple systems without re-entering passwords (Single Sign-On – SSO) Log in once Where We Use It ...
ZFS Pool Capacity Critical Alert - Troubleshooting Guide
Purpose This document provides a standardized approach for handling and troubleshooting the ZFS Pool Capacity Critical alert, enabling rapid response to prevent storage exhaustion, performance degradation, and potential service outages. Alert Name ...
ZFS Pool Capacity Warning Alert - Troubleshooting Guide
Purpose This document provides a standardized approach for handling and troubleshooting the ZFS Pool Capacity Warning alert, helping L1 and L2 engineers prevent storage exhaustion and maintain system stability. Alert Name ZFS Pool Capacity Warning ...
ZFS Pool Error Alert - Troubleshooting Guide
Purpose This document provides a standardized approach for handling and troubleshooting the ZFS Pool Error alert, enabling L1 and L2 engineers to quickly identify disk-related issues, prevent data loss, and maintain storage reliability. Alert Name ...
PSU Error Alert - Troubleshooting Guide
Purpose This document provides a standardized approach for handling and troubleshooting the PSU Error alert, ensuring quick identification of power supply issues and maintaining system power redundancy. Alert Name PSU Error Alert Description This ...
RMA FAQ
RMA FAQ What is RMA? Answer: RMA ( Return Merchandise Authorization) is a process that allows customers to return defective or faulty products for repair, replacement or refund as per warranty terms. When should I request RMA? Answer: You should ...
ASUS Server Testing Workflow – MBUZZ Service Center
This document outlines the mandatory steps to be completed before submitting any ASUS server to the QC department. All engineers must strictly follow these procedures to ensure consistency, reliability, and compliance with MBUZZ service standards. ...
Nvidia-smi Tool Alert - Troubleshooting Guide
Purpose This document provides a standardized approach for handling and troubleshooting the Nvidia-smi Tool alert, including L1 and L2 actions, escalation guidelines, and resolution criteria. Alert Name Nvidia-smi Tool Alert Description This alert is ...
Node Exporter Down - Troubleshooting Guide
Purpose This document provides a standardized approach for handling and troubleshooting the Node Exporter Down alert, including L1 and L2 actions, escalation flow, and resolution criteria. Alert Name Node Exporter Down Alert Description This alert is ...
Memory Filesystem Capacity Warning -- Troubleshooting guide
Purpose Provide a standard troubleshooting and resolution procedure when filesystem capacity on any HGX Node reaches warning levels. When to Use Use this KB when the monitoring system triggers the alert: Memory Filesystem Capacity warning – For Any ...
Memory Filesystem Capacity Critical -- Troubleshooting guide
Purpose Provide a standard troubleshooting and resolution procedure when filesystem capacity on any HGX Node reaches critical levels. When to Use Use this KB when the monitoring system triggers the alert: Memory Filesystem Capacity Critical – For Any ...
How to Add Device Images in PNETLab (Installed on Proxmox)?
1. Overview This article provides a comprehensive guide for adding device images (such as Cisco, Huawei, FortiGate, and other network operating systems) to PNETLab. Adding images is essential for creating virtual network labs. There are two primary ...
Key Planning Steps Before Building a Custom Loop Workstation
When making a custom liquid-cooling loop workstation, you need to carefully plan for things like heat load, case layout, component compatibility, and long-term maintenance. Based on advice from experts and real-life experience as a builder, here is a ...
Network Speed IB Alert - Troubleshooting Guide
Purpose This document provides a standardized approach for handling and troubleshooting the Network Speed – IB alert, including L1 and L2 actions, escalation process, and resolution guidelines. Alert Name Network Speed IB Alert Description This alert ...
Network Down IB Alert - Troubleshooting Guide
Purpose This document provides a standardized approach for handling and troubleshooting the Network Down – IB alert, including L1 and L2 actions, escalation process, and resolution criteria. Alert Name Network Down – IB Alert Description This alert ...
Network Down Extern (P1 Critical Alert) - Troubleshooting Guide
Purpose This document provides a standardized approach for handling and troubleshooting the Network Down - Extern alert, including L1 and L2 actions, escalation procedures, and resolution criteria for critical incidents. Alert Name Network Down - ...
Network Down Boot (P1 Critical Alert)- Troubleshooting Guide
Purpose This document provides a standardized approach for handling and troubleshooting the Network Down - Boot alert, ensuring quick response to critical incidents and maintaining cluster control and system accessibility. Alert Name Network Down - ...
Network Degraded (Boot) Alert - Troubleshooting Guide
Purpose This document provides a standardized approach for handling and troubleshooting the Network Degraded - Boot alert, enabling L1 and L2 engineers to ensure timely detection, prevent full network outage, and maintain service availability. Alert ...
Linux Software RAID Failure Alert - Troubleshooting Guide
Purpose This document provides a standardized approach for handling and troubleshooting the Linux Software RAID Failure alert, enabling L1 and L2 engineers to respond quickly, minimize risk, and maintain system stability. Alert Name Linux Software ...
GPU Temperature Warning alert- Troubleshooting Guide
Purpose Provide guidance for engineers to investigate and resolve the GPU Temperature Warning alert in GPU compute nodes. When To Use Use this article when the monitoring system generates the alert: GPU Temperature Warning Alert Details This alert is ...
GPU Temperature Critical alert - Troubleshooting Guide
Purpose This article provides guidance for engineers to identify and respond to the GPU Temperature Critical alert in GPU compute nodes. It outlines the alert meaning, possible causes, and the required L1 and L2 troubleshooting steps. Alert Name GPU ...
GPU Miscount Alert - Troubleshooting Guide
Purpose Provide troubleshooting steps when a GPU Miscount alert is triggered, indicating that the system detects an incorrect number of GPUs on a node. Alert Name GPU Miscount What the Alert Is This alert indicates that the number of GPUs detected on ...
GPU Memory Temperature Critical - Troubleshooting Guide
Purpose This document provides a standardized approach for handling and troubleshooting the GPU Memory Temperature Critical alert, including L1 and L2 actions, escalation procedures, and resolution criteria. Alert Name GPU Memory Temperature Critical ...
GPU Memory Mismatch – Troubleshooting Guide
Alert Information Alert Name: GPU Memory Mismatch Severity: P3 – Medium Impact Alert Description: This alert is generated when the detected GPU memory size on a node does not match the expected configuration. Alert Summary The monitoring system has ...
How to Diagnose Fan Module Failure on Cisco Catalyst & NX-OS Switches?
1.Overview This document provides a systematic guide to diagnosing fan module failures on Cisco Catalyst switches running IOS/XE and Cisco Nexus switches running NX-OS. Fan failures can lead to overheating and switch shutdown if not addressed ...
Filesystem Device Error - Troubleshooting Guide
Purpose Provide troubleshooting and resolution steps when a Filesystem Device Error alert is triggered on a server. This article helps engineers quickly identify disk or filesystem issues and take corrective action to prevent data corruption or ...
Filesystem Capacity Warning in Head Node - Troubleshooting Guide
Purpose Provide a troubleshooting and response procedure when a Filesystem Capacity Warning alert is triggered on a Head Node. This KB helps engineers quickly identify the issue, perform basic checks, and escalate when required. Alert Name Filesystem ...
Filesystem Capacity Critical in Head Node - Troubleshooting guide
Purpose Provide a standard troubleshooting and resolution procedure when filesystem capacity on a Head Node reaches critical levels. When to Use Use this KB when the monitoring system triggers the alert: Filesystem Capacity Critical – Head Node This ...
Fan Speed Warning - Troubleshooting Guide
Purpose Provide guidance for engineers to investigate and respond to Fan Speed Warning alerts generated by the monitoring system. This KB article ensures that engineers follow a consistent troubleshooting procedure when cooling-related alerts occur. ...
Fan Speed Critical - Troubleshooting Guide
Purpose Provide a troubleshooting guide for engineers when a Fan Speed Critical alert is generated on a server node. This article helps engineers quickly identify cooling issues and perform corrective actions to prevent overheating and potential ...
CPU Temperature Warning - Troubleshooting Guide
Purpose Provide a troubleshooting guide for engineers when a CPU Temperature Warning alert is triggered on a Node/Server. This KB helps engineers quickly identify the issue and perform initial remediation before escalation. When to Use Use this KB ...
CPU Temperature Critical - Troubleshooting Guide
Purpose Provide a standard troubleshooting procedure for engineers when a CPU Temperature Critical alert is received in the cluster. This article helps engineers quickly identify the cause and perform initial remediation before escalation. When to ...
Next page