Skip to main content

Nvidia GPU

Nvidia GPU

Plugin: go.d.plugin Module: nvidia_smi

Maintained by Netdata

Overview​

This collector monitors GPUs performance metrics using the nvidia-smi CLI tool.

This collector is supported on all platforms.

This collector supports collecting metrics from multiple instances of this integration.

Nvidia GPU can be monitored further using the following other integrations:

Default Behavior​

Auto-Detection​

This integration doesn't support auto-detection.

Limits​

The default configuration for this integration does not impose any limits on data collection.

Performance Impact​

The default configuration for this integration is not expected to impose a significant performance impact on the system.

Setup​

You can configure the nvidia_smi collector in two ways:

MethodBest forHow to
UIFast setup without editing filesGo to Nodes → Configure this node → Collectors → Jobs, search for nvidia_smi, then click + to add a job.
FileIf you prefer configuring via file, or need to automate deployments (e.g., with Ansible)Edit go.d/nvidia_smi.conf and add a job.
important

UI configuration requires paid Netdata Cloud plan.

Prerequisites​

No action required.

Configuration​

Options​

The following options can be defined globally: update_every, autodetection_retry.

Config options
OptionDescriptionDefaultRequired
update_everyData collection frequency.10no
autodetection_retryAutodetection retry interval (seconds). Set 0 to disable.0no
binary_pathPath to the nvidia-smi binary.nvidia-smino
timeoutCommand execution timeout, in seconds. In loop mode, it limits the wait for each nvidia-smi sample instead.10no
loop_modeKeep nvidia-smi running and reading GPU data on a fixed interval (the -l option) instead of starting it for every collection.yesno
binary_path​

When it is empty or the file does not exist, the collector looks for nvidia-smi in the PATH directories and, on Windows, in the default NVIDIA install locations.

timeout​

In loop mode, the collector waits up to this timeout for the first sample after it starts. Afterwards, a sample that does not arrive within the loop interval plus this timeout leaves a gap in the charts, and nvidia-smi is restarted. On hosts with several GPUs, the first sample can take 5 to 10 seconds.

loop_mode​

The loop interval is the data collection interval (update_every), at most 5 seconds. Loop mode is not available on Windows: there, each collection runs nvidia-smi once.

via UI​

Configure the nvidia_smi collector from the Netdata web interface:

  1. Go to Nodes.
  2. Select the node where you want the nvidia_smi data-collection job to run and click the ⚙ (Configure this node). That node will run the data collection.
  3. The Collectors → Jobs view opens by default.
  4. In the Search box, type nvidia_smi (or scroll the list) to locate the nvidia_smi collector.
  5. Click the + next to the nvidia_smi collector to add a new job.
  6. Fill in the job fields, then click Test to verify the configuration and Submit to save.
    • Test validates the provided settings and checks the collector's startup prerequisites. Successful validation does not guarantee that every metric will be available during collection.
    • If it fails, an error message appears with details (for example, connection refused, timeout, or command execution errors), so you can adjust and retest.

via File​

The configuration file name for this integration is go.d/nvidia_smi.conf.

The file format is YAML. Generally, the structure is:

update_every: 1
autodetection_retry: 0
jobs:
- name: some_name1
- name: some_name2

You can edit the configuration file using the edit-config script from the Netdata config directory.

cd /etc/netdata 2>/dev/null || cd /opt/netdata/etc/netdata
sudo ./edit-config go.d/nvidia_smi.conf
Examples​
Custom binary path​

The executable is not in the directories specified in the PATH environment variable.

Config
jobs:
- name: nvidia_smi
binary_path: /usr/local/sbin/nvidia_smi

Alerts​

There are no alerts configured by default for this integration.

Metrics​

Metrics grouped by scope.

The scope defines the instance that the metric belongs to. An instance is uniquely identified by a set of labels.

Per gpu​

These metrics refer to the GPU.

Labels:

LabelDescription
uuidGPU uuid (e.g. GPU-27b94a00-ed54-5c24-b1fd-1054085de32a)
indexGPU index (nvidia_smi typically orders GPUs by PCI bus ID)
product_nameGPU product name (e.g. NVIDIA A100-SXM4-40GB)

Metrics:

MetricDescriptionDimensionsUnit
nvidia_smi.gpu_pcie_bandwidth_usagePCI Express Bandwidth Usagerx, txB/s
nvidia_smi.gpu_pcie_bandwidth_utilizationPCI Express Bandwidth Utilizationrx, tx%
nvidia_smi.gpu_fan_speed_percFan speedfan_speed%
nvidia_smi.gpu_utilizationGPU utilizationgpu%
nvidia_smi.gpu_memory_utilizationMemory utilizationmemory%
nvidia_smi.gpu_decoder_utilizationDecoder utilizationdecoder%
nvidia_smi.gpu_encoder_utilizationEncoder utilizationencoder%
nvidia_smi.gpu_frame_buffer_memory_usageFrame buffer memory usagefree, used, reservedB
nvidia_smi.gpu_bar1_memory_usageBAR1 memory usagefree, usedB
nvidia_smi.gpu_temperatureTemperaturetemperatureCelsius
nvidia_smi.gpu_voltageVoltagevoltageV
nvidia_smi.gpu_clock_freqClock current frequencygraphics, video, sm, memMHz
nvidia_smi.gpu_power_drawPower drawpower_drawWatts
nvidia_smi.gpu_performance_statePerformance stateP0-P15state
nvidia_smi.gpu_mig_mode_current_statusMIG current modeenabled, disabledstatus
nvidia_smi.gpu_mig_devices_countMIG devicesmigdevices

Per mig​

These metrics refer to the Multi-Instance GPU (MIG).

Labels:

LabelDescription
uuidGPU uuid (e.g. GPU-27b94a00-ed54-5c24-b1fd-1054085de32a)
product_nameGPU product name (e.g. NVIDIA A100-SXM4-40GB)
gpu_instance_idGPU instance id (e.g. 1)

Metrics:

MetricDescriptionDimensionsUnit
nvidia_smi.gpu_mig_frame_buffer_memory_usageFrame buffer memory usagefree, used, reservedB
nvidia_smi.gpu_mig_bar1_memory_usageBAR1 memory usagefree, usedB

Troubleshooting​

Diagnostics​

Debug Mode​

Important: Debug mode is not supported for data collection jobs created via the UI using the Dyncfg feature.

To troubleshoot issues with the nvidia_smi collector, run the go.d.plugin with the debug option enabled. The output should give you clues as to why the collector isn't working.

  • Navigate to the plugins.d directory, usually at /usr/libexec/netdata/plugins.d/. If that's not the case on your system, open netdata.conf and look for the plugins setting under [directories].

    cd /usr/libexec/netdata/plugins.d/
  • Switch to the netdata user.

    sudo -u netdata -s
  • Run the go.d.plugin to debug the collector:

    ./go.d.plugin -d -m nvidia_smi

    To debug a specific job:

    ./go.d.plugin -d -m nvidia_smi -j jobName

Getting Logs​

If you're encountering problems with the nvidia_smi collector, follow these steps to retrieve logs and identify potential issues:

  • Run the command specific to your system (systemd, non-systemd, or Docker container).
  • Examine the output for any warnings or error messages that might indicate issues. These messages should provide clues about the root cause of the problem.
System with systemd​

Use the following command to view logs generated since the last Netdata service restart:

journalctl _SYSTEMD_INVOCATION_ID="$(systemctl show --value --property=InvocationID netdata)" --namespace=netdata --grep nvidia_smi
System without systemd​

Locate the collector log file, typically at /var/log/netdata/collector.log, and use grep to filter for collector's name:

grep nvidia_smi /var/log/netdata/collector.log

Note: This method shows logs from all restarts. Focus on the latest entries for troubleshooting current issues.

Docker Container​

If your Netdata runs in a Docker container named "netdata" (replace if different), use this command:

docker logs netdata 2>&1 | grep nvidia_smi

Do you have any feedback for this page? If so, you can open a new issue on our netdata/learn repository.