Skip to main content

Monitor GPU usage (NVIDIA GPUs) and GPU-related metrics in Grafana


βœ… Overview of the Stack


πŸ›  Step-by-Step Setup

βœ… 1. Install NVIDIA DCGM Exporter

The DCGM Exporter (Data Center GPU Manager) exposes GPU metrics in Prometheus format.
Verify:
You should see Prometheus-format metrics like DCGM_FI_DEV_GPU_UTIL.

βœ… 2. Ensure Prometheus is Installed and Scraping

If using Prometheus Operator (from Helm):

Add scrape config to ServiceMonitor

If using kube-prometheus-stack, this is automatic with correct labels. To scrape manually, add this ServiceMonitor:

βœ… 3. Install Grafana Dashboard

Use NVIDIA’s official GPU dashboards:
  • Open Grafana
  • Go to Dashboards > Import
  • Use one of the following dashboard IDs:
You can also customize and clone these dashboards.

βœ… 4. Verify GPU Metrics in Prometheus

Run in Prometheus UI or via Grafana Explore:
These will show % utilization, memory used, etc.

πŸ”’ Optional: Node Exporter GPU Plugin (Advanced)

If you want host-level detail alongside GPU, you can also use node_exporter with custom GPU scripts, but dcgm-exporter is the preferred method for Kubernetes.

βœ… Summary


Would you like a Helm chart-based setup for Prometheus + Grafana + DCGM on EKS?