Grafana — Visualization & Observability Platform
Before you start: this course assumes you've read the Prometheus overview first (Grafana is almost always paired with it) and are comfortable with basic Kubernetes concepts. No prior visualization-tool experience is needed.
Grafana is the industry-standard open-source platform for monitoring visualization. It doesn't store data — it queries data sources (Prometheus, Loki, Elasticsearch, CloudWatch, and 100+ others) and renders dashboards, sets alerts, and provides a unified observability UI.
Analogy — Think of Grafana like a universal remote control, not a TV. A universal remote doesn't generate any picture itself — it sends commands to whatever device is actually plugged in (a cable box, a game console, a streaming stick) and shows you the result. Grafana works the same way: it doesn't generate or store any data itself, it just sends queries to whatever's actually holding the data (Prometheus, Loki, Elasticsearch, CloudWatch) and renders whatever comes back, in a consistent interface regardless of which "device" is behind it.
Grafana (the remote) --queries--> Prometheus / Loki / CloudWatch / ... (the actual devices)
│ │
└──────────────── renders the result back as a dashboard ─────────┘
What Grafana Provides
•Dashboards — Beautiful, interactive charts from any data source
•Alerting — Unified alert rules with routing to Slack, PagerDuty, email
•Explore — Ad-hoc query interface for debugging
•Annotations — Mark deployments/incidents on graphs
•Variables — Dynamic dashboards (filter by environment, pod, namespace)
•Plugins — 150+ data source plugins
Data Sources
bash
# Most common data sources:
# Prometheus — metrics (K8s, applications)
# Loki — logs (Kubernetes logs)
# Elasticsearch — logs and search
# CloudWatch — AWS metrics and logs
# Azure Monitor — Azure metrics
# InfluxDB — time-series metrics
# PostgreSQL — relational data for business metrics
# Jaeger/Tempo — distributed traces
Essential Dashboard Panels
Time series — Line charts for metrics over time (CPU, memory, requests)
Stat — Single big number with trend (current error rate, uptime)
Gauge — Dial showing value vs thresholds (disk usage %)
Bar chart — Comparisons across categories
Table — Detailed data with sorting and filtering
Logs — Log stream from Loki or Elasticsearch
Heatmap — Request latency distributions over time
Geomap — Geographic distribution of traffic
PromQL Dashboard Queries
promql
# CPU Usage by Pod — time series
rate(container_cpu_usage_seconds_total{
namespace="$namespace",
pod=~"$pod"
}[5m]) * 100
# Memory Usage — gauge (% of limit)
(
container_memory_working_set_bytes{namespace="$namespace", pod=~"$pod"}
/
container_spec_memory_limit_bytes{namespace="$namespace", pod=~"$pod"}
) * 100
# HTTP Request Rate — time series with status breakdown
sum by (status_code) (
rate(http_requests_total{job="$job"}[5m])
)
# P50/P95/P99 Latency — overlay on same graph
histogram_quantile(0.50, sum(rate(http_request_duration_seconds_bucket{job="$job"}[5m])) by (le))
histogram_quantile(0.95, sum(rate(http_request_duration_seconds_bucket{job="$job"}[5m])) by (le))
histogram_quantile(0.99, sum(rate(http_request_duration_seconds_bucket{job="$job"}[5m])) by (le))
# Error Rate as percentage — stat panel
(
sum(rate(http_requests_total{status=~"5.."}[5m]))
/
sum(rate(http_requests_total[5m]))
) * 100
# Pod Restarts in last 24h — table
sort_desc(
sum by (pod, namespace) (
increase(kube_pod_container_status_restarts_total[24h])
)
)
Variables — Dynamic Dashboards
# Create dashboard variables (Dashboard Settings → Variables)
# Variable: namespace
Type: Query
Query: label_values(kube_pod_info, namespace)
Multi-value: true
Include All: true
# Variable: pod (depends on namespace)
Type: Query
Query: label_values(kube_pod_info{namespace=~"$namespace"}, pod)
Regex: /^(.*)-[a-z0-9]+-[a-z0-9]+$/ # Strip pod hash
# Variable: environment
Type: Custom
Values: dev,staging,prod
# Use in queries:
container_cpu_usage_seconds_total{namespace="$namespace", pod=~"$pod"}
Alert Rules in Grafana
yaml
# Grafana managed alerts (via UI or provisioning)
# Create via Provisioning (rules as code)
# /etc/grafana/provisioning/alerting/rules.yaml
apiVersion: 1
groups:
- orgId: 1
name: Production Alerts
folder: Alerts
interval: 1m
rules:
- uid: high-error-rate
title: High Error Rate
condition: C
data:
- refId: A
datasourceUid: prometheus
model:
expr: |
sum(rate(http_requests_total{status=~"5.."}[5m]))
/ sum(rate(http_requests_total[5m])) * 100
- refId: C
datasourceUid: __expr__
model:
type: threshold
conditions:
- evaluator:
type: gt
params: [5] # Alert if >5% errors
query:
params: [A]
noDataState: NoData
execErrState: Error
for: 5m
annotations:
summary: "Error rate {{ $values.A.Value | printf \"%.2f\" }}% on {{ $labels.job }}"
labels:
severity: critical
isPaused: false
Grafana Alerting — Contact Points
yaml
# /etc/grafana/provisioning/alerting/contact_points.yaml
apiVersion: 1
contactPoints:
- orgId: 1
name: Slack Production
receivers:
- uid: slack-prod
type: slack
settings:
url: https://hooks.slack.com/services/...
channel: "#production-alerts"
title: |
{{ if eq .Status "firing" }}🔴{{ else }}✅{{ end }}
{{ .CommonAnnotations.summary }}
text: |
*Severity:* {{ .CommonLabels.severity }}
*Started:* {{ .StartsAt | timeZone "Asia/Kolkata" }}
- orgId: 1
name: PagerDuty Critical
receivers:
- uid: pagerduty
type: pagerduty
settings:
integrationKey: YOUR_PD_INTEGRATION_KEY
severity: critical
class: "{{ .CommonLabels.alertname }}"
component: "{{ .CommonLabels.job }}"
Loki — Log Queries in Grafana
Loki is Prometheus's sister project for logs — instead of PromQL, you query it with LogQL, a similar-looking language for filtering and aggregating log lines.
logql
# Basic log stream
{namespace="production", pod=~"myapp-.*"}
# Filter by content
{namespace="production"} |= "ERROR"
# Parse JSON logs and filter
{namespace="production"}
| json
| level="error"
| line_format "{{.timestamp}} {{.message}}"
# Count errors over time (metric from logs)
sum(rate({namespace="production"} |= "ERROR" [5m])) by (pod)
# Extract and visualize response times from logs
{job="nginx"}
| pattern '<_> <_> <_> "<method> <uri> <_>" <status> <bytes> <_> "<_>" <duration>'
| unwrap duration
| quantile_over_time(0.95, [5m]) by (uri)
Production Dashboard: Kubernetes Cluster Overview
Row 1: Cluster Health
- Stat: Node count (healthy/total)
- Stat: Pod count (running/total)
- Stat: CPU utilization (%)
- Stat: Memory utilization (%)
Row 2: Workloads
- Time series: Pod CPU usage by namespace
- Time series: Pod memory usage by namespace
- Table: Top 10 CPU consumers (pods)
- Table: Pods not Running
Row 3: Application Performance
- Time series: Request rate (RPS — requests per second) by service
- Time series: Error rate (%) by service
- Time series: P95 latency by service
- Heatmap: Latency distribution
Row 4: Infrastructure
- Heatmap: Node CPU per AZ
- Time series: Disk usage trend
- Time series: Network I/O
- Table: Node resource allocation
Import Community Dashboards
bash
# Best Grafana dashboard IDs to import immediately:
# 315 - Kubernetes cluster monitoring
# 6417 - Kubernetes pod monitoring
# 1860 - Node Exporter Full
# 12006 - Kubernetes Networking
# 7249 - Kubernetes Cluster (resource requests)
# 13659 - Kubernetes Deployments
# 14584 - ArgoCD
# 10991 - Prometheus stats
# Import via Grafana UI:
# + → Import → Enter dashboard ID → Load → Select data source → Import
Try It (2 Minutes)
1.docker run -d -p 3000:3000 grafana/grafana and open http://localhost:3000 (default login admin/admin, it'll prompt you to change it).
2.Go to Connections → Data Sources → Add data source → choose TestData DB (a built-in fake data source made exactly for this — no real Prometheus/Loki needed).
3.Create a new dashboard, add a panel, select the TestData source, and pick the "Random Walk" scenario — you'll immediately see a live-updating line chart, generated entirely by the data source, not by Grafana itself. Switch the panel's data source dropdown to a different one later (once you have Prometheus running, for example) and the exact same panel type renders completely different real data — that's the universal-remote idea directly: the interface stays the same, only which "device" it's pointed at changes.
Interview Questions
What is the difference between Grafana and Prometheus?
Prometheus is a time-series database — it scrapes, stores, and evaluates metrics. Grafana is a visualization layer — it has no storage, just queries Prometheus (and 100+ other sources) via PromQL and renders the results as dashboards. Think of Prometheus as the database and Grafana as the dashboard on top. You can use Grafana without Prometheus (with other data sources), and Prometheus without Grafana (via its built-in expression browser), but together they are the industry-standard monitoring stack.
How do you avoid alert fatigue?
Alert only on symptoms that affect users (SLO violations), not causes. Use for: 5m — only alert if condition sustained, not on brief spikes. Group related alerts. Set correct severity levels — critical means wake someone up at 3am. Use inhibition rules — if the cluster is down, suppress pod-level alerts. Do monthly alert reviews — delete alerts that haven't fired in 6 months or have never led to action. The best alert is one that requires action every time it fires.