Perses Perses Operator

Kubernetes Native Observability Dashboards with Perses

Without the YAML/JSON Hell



Jayapriya Pai

About me

  • Senior Software Engineer at Red Hat
  • OpenShift In-cluster monitoring
  • Maintainer: prometheus-operator · kube-prometheus · perses · metrics-server
  • Member: Kubernetes SIG Instrumentation
  • GitHub: slashpai
Perses

What is Perses?

  • CNCF Sandbox project for observability dashboards
  • Open, vendor-neutral dashboard & datasource specification
  • First class support for Prometheus data sources and PromQL
  • Author dashboards via UI, Go SDK, CUE SDK, or Kubernetes CRDs
Perses

Perses at a glance

Perses features

Multi Datasource: Full Observability Stack

Perses has a plugin architecture: Add new datasources without changing core:

  • Metrics: Prometheus, GreptimeDB
  • Logs: Loki, OpenSearch, Splunk, VictoriaLogs, ClickHouse
  • Traces: Tempo, Jaeger
  • Profiling: Pyroscope
  • Alerts: Alertmanager

18+ panel types: time series, stat, gauge, bar, pie, heatmap, table, flame chart, trace gantt …

This talk focuses on Prometheus (datasource + PromQL).

UI is great for exploration.

But how do you author, validate, and ship dashboards without handwritten YAML/JSON?

A familiar failure mode

PromQL lives as an unvalidated string in giant JSON.

  • Edit PromQL as a raw string buried in 1000 lines of JSON
  • PR diffs are noise where the PromQL change is buried, so reviewers can miss it
  • No CI check: broken PromQL reaches production unchecked
  • After deployment: No Data from a query that was never validated

What if dashboards were application code?

Treat them like services you ship every day

  • Authored in Go
  • Reviewed as small, meaningful diffs
  • Validated before merge (structure + PromQL)
  • Delivered as Kubernetes resources through GitOps or kubectl apply

"YAML is the artifact."

"Code is the source of truth."

The Dashboard-as-Code toolkit

Three Go libraries make this possible:

Library Role
Perses Go SDK Typed builders for dashboards, panels, queries
promql-builder Construct PromQL as a Go AST, not raw strings
community-mixins Reusable panel library, import as a Go module

Go SDK: Dashboard Builder

The Perses Go SDK builds the dashboard structure: datasources, variables, panels as typed Go.

dashboard package

dashboard.New, AddPanelGroup, AddDatasource, AddVariable, Duration

Go SDK: Panel Builder

panel package

panel.New, AddQuery, Description, Plugin, Title

pkg.go.dev/github.com/perses/perses/go-sdk

Not a custom DSL. Not a config language. Just Go.

How a dashboard is composed

Each layer is a typed Go function: compose, review, and validate like application code.

dashboard.New("node-exporter-overview", …)
 └── dashboard.AddPanelGroup("CPU (community-mixins)",
      └── panelgroup.AddPanel("CPU Usage",   // NodeCPUUsagePercentage
           ├── panel.Description("Shows CPU utilization percentage…")
           ├── timeSeriesPanel.Chart(...)
           └── panel.AddQuery(query.PromQL(
                 promql.SetLabelMatchersV2(
                   NodeExporterCommonPanelQueries["NodeExporterCPUUsagePercentage"],
                   labelMatchers,
                 ).Pretty(0), ...))

What is a Mixin

A Mixin is a reusable package of dashboards, alerts, and recording rules for a specific component.

Traditionally:

  • Written in Jsonnet → generates JSON dashboards + PrometheusRule YAML
  • Examples: kubernetes-mixin, node-exporter-mixin, etcd-mixin
  • In practice many teams skip Jsonnet, fetch the rendered YAML, and patch with Kustomize and reuse never sticks

Perses community-mixins: mixins, but in Go

Goal is not only ready-made YAML, it's a panel library you import and compose.

  • Dashboards built with the Perses Go SDK
  • Reusable panels: import what you need, extend with your own
  • Same patterns → consistency across teams and services
  • Customize label matchers for your environment via Go functions

go get github.com/perses/community-mixins, then import panels like any Go module.

Compose and extend community-mixins

import (
    communityPanels "github.com/perses/community-mixins/pkg/panels/node_exporter"
    mixinpromql     "github.com/perses/community-mixins/pkg/promql"
    gitpromql       "github.com/slashpai/perses-gitops-workflow/dashboards/promql"
)

communityPanels.SetNodeExporterLabelValue("node-exporter")
jobMatcher := &labels.Matcher{Name: "job", Type: labels.MatchEqual,
    Value: communityPanels.GetNodeExporterLabelValue()}
instanceMatcher := mixinpromql.InstanceVarV2

// Reuse community panels with kube-prometheus matchers
dashboard.AddPanelGroup("CPU (community-mixins)",
    communityPanels.NodeCPUUsagePercentage(datasource, jobMatcher, instanceMatcher),
    communityPanels.NodeAverage(datasource, jobMatcher, instanceMatcher),
),
dashboard.AddPanelGroup("Memory (community-mixins)",
    communityPanels.NodeMemoryUsageBytes(datasource, jobMatcher, instanceMatcher),
    communityPanels.NodeMemoryUsagePercentage(datasource, jobMatcher, instanceMatcher),
),

// Extend with your own panel alongside
dashboard.AddPanelGroup("Filesystem (custom)",
    panelgroup.AddPanel("Filesystem Used",
        panel.AddQuery(
            query.PromQL(gitpromql.FilesystemUsedRatio().Pretty(0), ...),
        ),
    ),
),

Dashboard built from scratch

No community-mixins panels? Build the whole dashboard from your own queries:

import (
    gitpromql "github.com/slashpai/perses-gitops-workflow/dashboards/promql"
)

dashboard.New("prometheus-operator-health",
    dashboard.Name("Prometheus Operator / Health"),
    dashboard.AddVariable("job", ...),       // Helm job names vary
    dashboard.AddVariable("namespace", ...),

    dashboard.AddPanelGroup("Reconciliation",
        panelgroup.AddPanel("Reconcile Rate",
            query.PromQL(gitpromql.ReconcileRate().Pretty(0), ...)),
        panelgroup.AddPanel("Reconcile Error Ratio",
            query.PromQL(gitpromql.ReconcileErrorRatio().Pretty(0), ...)),
        panelgroup.AddPanel("Reconcile Duration (p99 / p50)", ...),
    ),
    dashboard.AddPanelGroup("Triggers", ...),
    dashboard.AddPanelGroup("API Operations", ...),
    dashboard.AddPanelGroup("Status", ...),
)

same SDK, same CI, same GitOps but local PromQL helpers not community imports.

PromQL as an AST (not a fragile string)

Queries inside those panels use promql-builder: PromQL as a Go AST, not a raw string.

// ReconcileRate — prometheus-operator metrics
promqlbuilder.Sum(
    promqlbuilder.Rate(
        matrix.New(
            vector.New(vector.WithMetricName(
                "prometheus_operator_reconcile_operations_total")),
            matrix.WithRangeAsVariable("$__rate_interval"),
        ),
    ),
).By("controller", "namespace")
// then withOperatorMatchers → job=~"$job", namespace=~"$namespace"
// FilesystemUsedRatio — custom extend panel (dashboards/promql/queries.go)
promqlbuilder.Div(
    promqlbuilder.Sub(
        vector.New(vector.WithMetricName("node_filesystem_size_bytes"),
            vector.WithLabelMatchers(fstype!="", mountpoint!="")),
        vector.New(vector.WithMetricName("node_filesystem_avail_bytes"),
            vector.WithLabelMatchers(fstype!="", mountpoint!="")),
    ),
    vector.New(vector.WithMetricName("node_filesystem_size_bytes"), ...),
)
// then withNodeMatchers → job="node-exporter", instance=~"$instance"

promql-builder → construct queries as typed Go AST, validate at build / CI time.

PromQL validated before it ships

From community-mixins pkg/promql used by demo and mixin panels:

func SetLabelMatchersV2(query parser.Expr, matchers []*labels.Matcher) parser.Expr {
    copy := promqlbuilder.DeepCopyExpr(query)
    for _, l := range matchers {
        copy = labelsSetPromQLV2(copy, l.Type, l.Name, l.Value)
    }
    if err := promqlbuilder.Validate(copy); err != nil {
        panic(err)   // bad PromQL never renders
    }
    return copy
}

Every query passes through promqlbuilder.Validate: malformed PromQL is caught before YAML is generated.

go test ./... → validate all dashboards → render CRs → commit.

CI catches bad PromQL before merge

PR adds a Workqueue panel: checks failing

PR with Workqueue Adds query

promqlbuilder.Validate in CI

rate() needs a range vector missing [$__rate_interval]

CI validation error from promql-builder

Dashboards are validated. How do they reach the cluster now?

Perses Operator

Kubernetes-native dashboard & datasource lifecycle via CRDs:

CRD Role
Perses Server instance (Deploy/STS, Service)
PersesDashboard Dashboard → Perses project (namespace)
PersesDatasource Project-scoped datasource
PersesGlobalDatasource Cluster-scoped datasource

The controller syncs valid CRs to Perses and reports status back.

What you ship → Generated PersesDashboard YAML

apiVersion: perses.dev/v1alpha2
kind: PersesDashboard
metadata:
  name: node-exporter-overview
  namespace: perses-dev
spec:
  config:
    display:
      name: Node Exporter / Overview
    panels:
      # … generated from Go — not hand-edited …
  • Built with the Perses Go SDK
  • Queries via promql-builder
  • Rendered to a PersesDashboard CR

Don't hand-edit this YAML.

End-to-end flow

  dashboards/ (Go)
       │  make validate-dashboards   ← go test ./...
       │  make render-dashboards     ← go run ./cmd/render
       ▼
  manifests/dashboards/
    ├── node-exporter-overview.yaml
    └── prometheus-operator-health.yaml
       │  PR + CI drift check
       ▼
  Argo CD  →  perses-operator  →  Perses UI

Demo repo: slashpai/perses-gitops-workflow

GitOps: Argo CD Application

perses-dashboards · manifests/dashboards → perses-dev · Healthy / Synced

Argo CD Applications list

Perses Operator reconciles the CR

Resource tree: PersesDashboard / node-exporter-overview

Argo CD synced dashboard

Perses UI after first sync

Node Exporter / Overview

Perses dashboards — Overview only

Evolve via PR → new dashboard lands

Add Prometheus Operator / Health in Go → validate → render → merge → both CRs sync

Argo CD synced both dashboards

Perses UI after second sync

Overview + Prometheus Operator / Health: same GitOps pipeline, two authoring patterns

Perses dashboards — Overview and Health

What you see in Perses

Node Exporter / Overview: community-mixins + custom Filesystem panel, validated PromQL from Go

Perses Node Exporter / Overview dashboard

What you see in Perses

Prometheus Operator / Health: dashboard built from scratch

Prometheus Operator / Health dashboard

Validation layers

When What Tool
Build / CI Bad PromQL structure promqlbuilder.Validate
Build / CI Manifest drift git diff --exit-code manifests/dashboards/
Deploy Invalid dashboard spec Operator → Perses API validation

Takeaways

  1. Stop rebuilding the same dashboards: import a panel library, compose and extend
  2. Dashboards like app code: reviewed, validated in CI, delivered via GitOps
  3. Kubernetes-native: CRDs + operator manage lifecycle, not import scripts
  4. Two patterns: compose/extend community-mixins when panels exist; from scratch when they don't

YAML is how you ship. Go is how you author.

Perses now ships as an addon in kube-prometheus

  • Perses deployed out of the box alongside Prometheus, Alertmanager, and other kube-prometheus components
  • 23 community-mixins dashboards deployed as PersesDashboard CRs

Try it / go deeper

Perses

Interested in contributing or want to know more?

Perses is a CNCF Sandbox project: contributions are welcome!

Scan to explore Perses

Perses Website QR

Thank you!

Questions?

===================== ACT 1: CONTEXT =====================

right-aligned via scoped style below

===================== ACT 2: PROBLEM =====================

============= ACT 3: VISION + BIG PICTURE ===============

================ ACT 4: AUTHORING TOOLKIT ===============

=============== ACT 5: CODE DEEP-DIVE ===================

=============== ACT 6: PROMQL SAFETY ====================

========== ACT 7: DELIVERY — OPERATOR + GITOPS ==========

=============== ACT 8: VISUAL PROOF =====================

=================== ACT 9: CLOSE ========================