KB 158059 - How to collect diagnostic logs using the NetApp Log Collection Script

KB 158059 - How to collect diagnostic logs using the NetApp Log Collection Script

1. Purpose

This document describes the procedure to collect diagnostic logs using the NetApp Log Collection Script in environments running:

  • BeeGFS

  • NetApp E-Series backend storage

  • HA cluster using Pacemaker and Corosync

This script is typically requested by NetApp Support for storage-side or HA-related troubleshooting.


2. When to Use

Use this procedure when:

  • Storage path failures occur

  • Multipath errors are detected

  • Pacemaker resource failures happen

  • Cluster failover behaves unexpectedly

  • NetApp explicitly requests the log bundle


3. What This Script Collects

  • /var/log/beegfs*

  • /var/log/messages*

  • pcs status

  • NVMe device details

  • Multipath configuration

  • IP configuration

  • dmesg output

  • Journald logs

  • Pacemaker & Corosync logs

This focuses on OS + cluster + storage diagnostics.


4. Script

Quote#!/bin/bash
# Set up variables
TS="$(date '+%F_%H-%M-%S-%Z')-$(hostname)"
WORKDIR="/tmp/${TS}-logbundle"
ARCHIVE="${TS}-support-bundle.tar.gz"

# Create working directory
mkdir -p "$WORKDIR"

# Collect log files
tar -zcvf "$WORKDIR/beegfs-logs.tar.gz" /var/log/beegfs* /var/log/messages* 2>/dev/null

# Collect pcs status
pcs status > "$WORKDIR/pcs_status.log"

# Collect NVMe devices
nvme list > "$WORKDIR/nvme_device.log"

# Collect NVMe connections
nvme list-subsys > "$WORKDIR/nvme_connections.log"

# Collect lsblk
lsblk > "$WORKDIR/lsblk.log"

# Collect multipath -ll
multipath -ll > "$WORKDIR/multipath.log"

# Collect IPs
ip a > "$WORKDIR/ip_address.log"

# Collect IP rules
ip rule show > "$WORKDIR/ip_rules.log"

# Collect dmesg output
dmesg -T > "$WORKDIR/dmesg.log"

# Collect system logs from journald
journalctl --no-pager -x > "$WORKDIR/journal.log"

# Collect pacemaker and corosync logs
tar -zcvf "$WORKDIR/pacemaker-logs.tar.gz" /var/log/pacemaker/pacemaker* /var/log/netapp/* /var/log/cluster /corosync* 2>/dev/null

# Bundle everything up
tar -zcvf "$ARCHIVE" -C "$WORKDIR" .

# Clean up
rm -rf "$WORKDIR"
echo "Support bundle created: $ARCHIVE"

5. Procedure

Step 1 – Create Script

Add this script file to the OSS node:

  1. You can manually create a file and paste this script into it
  2. Or you can use any tools like MobaXterm to upload a script file to any servers or nodes

Step 2 – Make Executable

chmod +x netapp-log-collect.sh


Step 3 – Execute

./netapp-log-collect.sh


6. Output

The script generates:

<timestamp>-support-bundle.tar.gz

Location: Current working directory

Example:

2026-02-17_22-45-01-support-bundle.tar.gz


7. Scope

Run on:

  • All affected OSS nodes

  • Metadata nodes (if involved)

  • All cluster nodes in HA setup

    • Recent Articles

    • KB-630208 | How to use Find and Locate to search for files in Linux

      Purpose To provide Linux administrators and Field Engineers with guidance on using the find and locate utilities to efficiently search for files and directories within a Linux filesystem based on criteria such as name, type, size, timestamps, ...
    • KB-630164 | How to use rsync to Synchronize Files

      Purpose To provide a clear and practical procedure for using the rsync utility to synchronize files and directories between local systems and remote hosts, including push and pull operations, while preserving file attributes and optimizing data ...
    • KB-630107 | How to create a Linux swap file

      Purpose To provide Field Engineers and Linux administrators with a standardized procedure for creating, configuring, and validating a Linux swap file. This procedure helps ensure sufficient virtual memory is available when physical RAM is exhausted ...
    • MSI All-in-One (AIO) PC – Physical Damage Policy for Warranty Claims

      MSI All-in-One (AIO) PC – Physical Damage Policy for Warranty Claims MSI's standard warranty does not cover physical or accidental damage to an All-in-One (AIO) PC. If the damage is determined to be caused by external factors rather than a ...
    • KB 134833 - Troubleshooting NVSM Alert NV-CPU-XX – Unrecoverable CPU Internal Error

      Purpose This document provides a general troubleshooting procedure for the NVIDIA System Management (NVSM) alert NV-CPU-XX, which indicates that a CPU has reported an internal error. The article outlines how to verify whether the alert represents an ...
    • Popular Articles

    • CP Plus Camera and NVR Configuration

      NVR Configuration The CP Plus Pro Series of NVRs have been meticulously designed for providing you with upgraded performance and higher recording quality in your IP video surveillance solution. The robust processor that has been inculcated in this ...
    • KB 775235 - Kerberos Authentication – Overview

      What is Kerberos? Kerberos is a secure authentication method used in our Active Directory (AD) environment (mbuzztech.com). It allows users to: Access multiple systems without re-entering passwords (Single Sign-On – SSO) Log in once Where We Use It ...
    • How to Remove and Reinstall NVIDIA Drivers on Ubuntu

      This article provides step by step guide to completely remove existing NVIDIA drivers and reinstall specific version of the NVIDIA driver on the Ubuntu system Prerequisites Administrative (sudo) access to the Ubuntu system. Internet access to ...
    • KB 692001 - Personal Computers and Servers - Classification and Point of Contact

      We can classify the computers that MBUZZ handles based on their form-factor as below: Tower Workstations, Desktops, Gaming PCs and SFF (Small form factor) PCs fall under this category. These are computers people would use on a desk and rarely move. ...
    • KB 298031 - M.2 SSD Tier List

      The sequential read and write speeds, which are usually the most advertised number, are not a proper benchmark of real-world performance or the quality of an SSD. This article categorizes and tiers SSDs based on factors like the type of NAND flash, ...