Rancher 2.15 Monitoring Migration: The Missing Practitioner Runbook

Share this article

Rancher 2.15 Splits Monitoring into Two Lifecycles

Rancher 2.15 fundamentally changes how Kubernetes monitoring is delivered. The monolithic, all-in-one rancher-monitoring chart gives way to an unbundled architecture: upstream kube-prometheus-stack supplies Prometheus Operator, Prometheus, Alertmanager, Grafana, exporters, and rules, while rancher-monitoring-dashboards adds Rancher's dashboards and the service proxies used by the Rancher UI.

In the official Rancher 2.15 documentation, the transition can look almost like a two-click replacement: install the runtime, then add the dashboard integration. In real clusters, the risks sit at the new ownership boundaries. Proxy failures remain silent until the UI returns HTTP 504, CRDs follow a lifecycle decoupled from the Helm release, and restrictive selectors produce scrape configurations without the expected targets.

In the monitoring proxy, the break may initially surface only as:

host not found in upstream
"kube-prometheus-stack-grafana.cattle-monitoring-system.svc.cluster.local"

Even after the expected Services exist, an already-running NGINX proxy can retain stale ClusterIPs: the workloads look healthy, yet Rancher still responds with HTTP 504. Independently, existing ServiceMonitor objects can be absent from the generated scrape configuration because of the default selectors.

This field-tested transition manual is written for platform teams upgrading existing monitoring installations or setting up new Rancher 2.15 clusters. It covers not only installation order but also CRD ownership, metrics history, storage, selectors, dashboards, alert routing, and recovery. Existing legacy installations remain supported, but the split stack is the official path forward for new installations.


What Actually Changed

The old chart was convenient because Rancher owned the whole monitoring release. It also coupled the lifecycle of an upstream observability stack to Rancher's release process. The new split puts the runtime back on the widely used prometheus-community Helm chart while keeping Rancher-specific UI assets separate.

Concern Legacy rancher-monitoring Rancher 2.15 split stack
Runtime Bundled, Rancher-maintained distribution Upstream kube-prometheus-stack
Rancher UI Bundled in the same release rancher-monitoring-dashboards and its proxy
Operator and CRDs Coupled to Rancher chart versions Pinned and upgraded with the upstream chart
Grafana dashboards Rancher defaults and runtime move together Dashboard ConfigMaps can change independently
Upgrade ownership One convenient but coarse lifecycle Two explicit, independently pinned lifecycles
Portability Rancher chart conventions Standard Prometheus Operator resources
Operational cost Fewer initial choices More values and version management, but clearer ownership

This is a sensible boundary. Prometheus, Grafana, and Alertmanager no longer need a Rancher-specific fork merely to appear inside Rancher. The trade-off is real: a one-click app becomes two releases plus a values file that the platform team must own.

The boundary also explains the naming contract. With the upstream release named kube-prometheus-stack, Rancher's proxy expects these Services:

kube-prometheus-stack-grafana
kube-prometheus-stack-prometheus
kube-prometheus-stack-alertmanager

It then exposes Rancher-facing Services with the established names:

rancher-monitoring-grafana
rancher-monitoring-prometheus
rancher-monitoring-alertmanager

Changing the Helm release name is possible only if the dashboard chart's upstream Service values are changed with it. Keeping the documented release name removes an unnecessary integration variable.


The Missing Manual: Five Ways a Naive Migration Breaks

1. CRDs outlive releases and have cluster-wide blast radius

Prometheus, Alertmanager, ServiceMonitor, PodMonitor, PrometheusRule, and AlertmanagerConfig are not namespaced implementation details of a Helm release. Their definitions are cluster-scoped APIs, and their objects may be used by teams in many namespaces.

Helm installs CRDs from a chart's crds/ directory before regular resources, skips a CRD that already exists, and deliberately does not upgrade or delete CRDs. That policy avoids implicit data loss, but it leaves version and ownership decisions with the operator. The behavior is documented in Helm's CRD best practices and is the reason kube-prometheus-stack has explicit CRD upgrade notes.

Never start by deleting *.monitoring.coreos.com CRDs. Deleting a CRD also deletes every custom resource stored behind that API. Inventory and export the resources first. There should be one active Prometheus Operator responsible for a given set of resources during the handover, not two operators racing over them.

There are two defensible CRD strategies:

  • Clean handover: export all monitoring resources, remove the legacy runtime, verify that no other release owns or needs its CRDs, update the CRDs to the version required by the new operator, and install the upstream stack.
  • Independent CRD lifecycle: manage the Prometheus Operator CRDs separately, install a compatible version before the new operator, and set crds.enabled: false in kube-prometheus-stack. This adds process, but it makes cluster API ownership explicit.

The chart's crds.upgradeJob can apply newer definitions, but in chart 88.3.0 that feature is marked preview and uses cluster-scoped RBAC. Treat it as a deliberate platform decision, not as a box to tick during an outage.

2. A retained PVC is not yet a metrics migration

Prometheus stores local TSDB blocks and a write-ahead log on its data volume. Its local storage is neither clustered nor replicated. The Prometheus project recommends snapshots for consistent backups and warns that copying a live TSDB directory can omit recent data or capture an incoherent state. The official storage documentation explains the block and WAL layout.

Before removing the old release, choose one of these outcomes:

  • Clean cutoff: start the new Prometheus with an empty PVC and retain the old volume for a defined rollback window.
  • Snapshot and restore: enable the TSDB admin API temporarily, create a snapshot, copy it using storage-native tooling, and test restoration into a compatible Prometheus version. This is a migration project of its own.
  • Continuous history outside the pod: use remote write or a long-term system such as Thanos, Mimir, or another compatible backend. This is the cleanest long-term boundary but not a last-minute migration toggle.

Do not mount the old PVC into the new StatefulSet merely because its filesystem is visible. Check Prometheus version compatibility, volume access mode, security context, StatefulSet claim naming, and rollback behavior first.

3. ServiceMonitors can exist while Prometheus ignores them

Prometheus Operator uses label and namespace selectors to associate a Prometheus object with ServiceMonitor and PodMonitor objects. The upstream chart normally narrows discovery to monitors carrying its Helm release labels. Rancher-provided monitors and application-owned monitors may not have those labels.

The failure is quiet: the ServiceMonitor exists, the target Service exists, and Prometheus still has no scrape job. Prometheus Operator's troubleshooting guide describes how to inspect the generated configuration.

For a cluster-wide platform Prometheus, we used:

prometheus:
  prometheusSpec:
    serviceMonitorSelectorNilUsesHelmValues: false
    podMonitorSelectorNilUsesHelmValues: false
    serviceMonitorSelector: {}
    serviceMonitorNamespaceSelector: {}
    podMonitorSelector: {}
    podMonitorNamespaceSelector: {}

An empty selector matches all relevant objects; a null selector has different semantics depending on the field. This broad discovery is appropriate only when the platform team controls who may create monitoring resources. In a multi-tenant cluster, use namespace labels and an explicit monitor label instead of accepting arbitrary scrape configurations cluster-wide.

4. Rancher dashboards are provisioned files, not Grafana plugins

The dashboard chart does not install a Grafana extension CRD. It creates dashboard ConfigMaps, normally in cattle-dashboards, labelled:

grafana_dashboard: "1"

The Grafana sidecar watches labelled ConfigMaps and Secrets across namespaces, writes their JSON payloads to /tmp/dashboards, and triggers Grafana's provisioning reload. Grafana then loads those files through a provisioning provider. The underlying behavior is standard Grafana dashboard provisioning.

That means custom dashboards can be GitOps-managed as Kubernetes objects. It also means imported Grafana.com JSON needs normalization. The interactive Grafana import flow substitutes datasource variables such as ${DS_PROMETHEUS}; file provisioning does not. Hard-coded datasource UIDs are equally brittle. Normalize every Prometheus datasource reference to the UID created by this chart, prometheus, and remove __inputs before putting the JSON in a ConfigMap.

5. Defaults are installation defaults, not capacity planning

Scrape interval, retention, sample cardinality, rule count, and Grafana concurrency determine actual resource use. The 1,750 MiB request and 2,500 MiB limit in Rancher's K3s guidance were a workable starting point for our small cluster, not a universal capacity figure.

The Prometheus admission webhook should remain enabled so malformed PrometheusRule resources are rejected before they can break rule loading. If its certificate patch job or webhook endpoint fails, fix that dependency. Permanently changing the failure policy to ignore merely turns an installation problem into a delayed monitoring problem.


Execution Runbook

The commands below assume a kubeconfig named rancher.yaml, context rancher, and namespace cattle-monitoring-system. Pin chart versions that you have tested. We used kube-prometheus-stack 88.3.0 with Prometheus Operator 0.93.0 and rancher-monitoring-dashboards 110.0.0+up0.1.2.

Phase 0: Freeze changes and capture the old state

Run this before uninstalling anything:

export KUBECONFIG="$PWD/rancher.yaml"
export MONITORING_NAMESPACE="cattle-monitoring-system"

kubectl config current-context
helm list -A | grep -E 'rancher-monitoring|kube-prometheus'

helm -n "$MONITORING_NAMESPACE" \
  get values rancher-monitoring --all > legacy-rancher-monitoring-values.yaml

kubectl get crd | grep monitoring.coreos.com
kubectl get prometheuses,alertmanagers,servicemonitors,podmonitors,prometheusrules \
  -A -o wide
kubectl -n "$MONITORING_NAMESPACE" get pvc
kubectl get pv -o custom-columns='NAME:.metadata.name,CLAIM_NAMESPACE:.spec.claimRef.namespace,CLAIM_NAME:.spec.claimRef.name,RECLAIM:.spec.persistentVolumeReclaimPolicy,CLASS:.spec.storageClassName'

Export the custom resources. Store the export securely; rules and endpoints may disclose internal topology.

for resource in \
  prometheuses.monitoring.coreos.com \
  alertmanagers.monitoring.coreos.com \
  servicemonitors.monitoring.coreos.com \
  podmonitors.monitoring.coreos.com \
  prometheusrules.monitoring.coreos.com \
  alertmanagerconfigs.monitoring.coreos.com
do
  kubectl get "$resource" -A -o yaml > "backup-${resource%%.*}.yaml"
done

Record the old and new Prometheus versions. Decide now whether history is being cut over, snapshotted, or retained remotely. If the old PV must survive deletion of its claim, confirm the storage driver's behavior and set the PV reclaim policy to Retain before removing the claim. A Kubernetes Retain policy protects the PV object from automatic backend deletion; it is not a backup.

Phase 1: Remove the legacy runtime without deleting its APIs blindly

Stop changes to monitoring resources during the handover. Remove the legacy rancher-monitoring App from Rancher Apps > Installed Apps, or use Helm if Helm directly owns that release:

helm -n "$MONITORING_NAMESPACE" uninstall rancher-monitoring

Do not automatically uninstall a separate CRD release and do not run kubectl delete crd. First verify that the legacy Prometheus, Alertmanager, operator Deployment, and webhooks are gone while the exported ServiceMonitor and PrometheusRule objects remain available.

kubectl -n "$MONITORING_NAMESPACE" get \
  deploy,statefulset,daemonset,job,pod,pvc
kubectl get validatingwebhookconfiguration,mutatingwebhookconfiguration \
  | grep -E 'prometheus|monitoring' || true

If old chart hooks, webhooks, ClusterRoles, or Services collide with the new release, identify their owning release from labels and annotations. Do not solve an ownership error by overwriting metadata until you know which controller is still using the object.

Before starting a newer operator against retained CRDs, review the actual schema diff and apply the CRDs shipped with the chart version you are about to install:

helm repo add prometheus-community \
  https://prometheus-community.github.io/helm-charts \
  --force-update
helm repo update

helm show crds prometheus-community/kube-prometheus-stack \
  --version 88.3.0 > kube-prometheus-stack-crds-88.3.0.yaml

kubectl diff -f kube-prometheus-stack-crds-88.3.0.yaml
kubectl apply --server-side \
  --field-manager=kube-prometheus-stack-crds \
  -f kube-prometheus-stack-crds-88.3.0.yaml

Stop if the server-side apply reports ownership conflicts. Resolve them against the controller that currently manages those fields; do not add --force-conflicts as an automatic migration flag. If CRDs have an independent lifecycle in your platform, commit this rendered artifact, review it like any other cluster API change, and use crds.enabled: false below.

Phase 2: Build an owned values.yaml

This is the compact values file we used on a single-node K3s cluster. It contains the Rancher integration contract and the fixes found during installation:

crds:
  enabled: true
  upgradeJob:
    enabled: false

grafana:
  grafana.ini:
    security:
      allow_embedding: true
    auth:
      disable_login_form: false
    auth.anonymous:
      enabled: true
      org_role: Viewer
    dashboards:
      default_home_dashboard_path: /tmp/dashboards/rancher-default-home.json
    users:
      auto_assign_org_role: Viewer

prometheus:
  prometheusSpec:
    serviceMonitorSelectorNilUsesHelmValues: false
    podMonitorSelectorNilUsesHelmValues: false
    serviceMonitorSelector: {}
    serviceMonitorNamespaceSelector: {}
    podMonitorSelector: {}
    podMonitorNamespaceSelector: {}
    scrapeInterval: 30s
    evaluationInterval: 30s
    retention: 15d
    resources:
      requests:
        cpu: 250m
        memory: 1750Mi
      limits:
        memory: 2500Mi

prometheusOperator:
  admissionWebhooks:
    enabled: true

# K3s needs Rancher's PushProx path for these control-plane components.
kubeEtcd:
  enabled: false
kubeControllerManager:
  enabled: false
kubeScheduler:
  enabled: false
kubeProxy:
  enabled: false

On a conventional Kubernetes distribution, review the four disabled component monitors rather than copying the K3s settings.

For production persistence, add a tested StorageClass and leave headroom between the PVC size and Prometheus retention size:

prometheus:
  prometheusSpec:
    retention: 15d
    retentionSize: 40GB
    storageSpec:
      volumeClaimTemplate:
        spec:
          storageClassName: longhorn
          accessModes:
            - ReadWriteOnce
          resources:
            requests:
              storage: 50Gi

grafana:
  persistence:
    enabled: true
    storageClassName: longhorn
    size: 5Gi

Replace longhorn with a StorageClass you actually operate and restore-test. Provisioned dashboards do not depend on Grafana's database, but users, UI edits, preferences, and other mutable state do.

Phase 3: Install the upstream runtime first

helm repo add prometheus-community \
  https://prometheus-community.github.io/helm-charts \
  --force-update
helm repo update

helm upgrade --install kube-prometheus-stack \
  prometheus-community/kube-prometheus-stack \
  --version 88.3.0 \
  --namespace "$MONITORING_NAMESPACE" \
  --create-namespace \
  --values kube-prometheus-stack-values.yaml \
  --wait \
  --timeout 15m

Wait for Prometheus, Alertmanager, Grafana, the operator, kube-state-metrics, and node-exporter. A Helm release stuck in pending-install is not a candidate for repeated blind retries. Inspect hooks, jobs, events, webhook certificates, and PVCs first.

helm -n "$MONITORING_NAMESPACE" status kube-prometheus-stack
kubectl -n "$MONITORING_NAMESPACE" get pods,pvc
kubectl -n "$MONITORING_NAMESPACE" get events \
  --sort-by=.lastTimestamp | tail -40

Phase 4: Install the Rancher dashboard integration

Install rancher-monitoring-dashboards from the Rancher 2.15 chart catalog only after the three upstream Services exist. The Rancher UI is useful here because it injects the local cluster ID, cluster name, Rancher URL, and System Project ID.

In Apps > Charts, select rancher-monitoring-dashboards, install it into cattle-monitoring-system, and keep the default upstream Service names when the runtime release is named kube-prometheus-stack. Keep k3sServer.enabled: true for K3s; disable it for other distributions as the chart describes.

If the dashboard chart was installed too early, upgrade or retry it after the runtime is ready. If the runtime was uninstalled and reinstalled, restart only the proxy so NGINX resolves the new Service ClusterIPs:

kubectl -n "$MONITORING_NAMESPACE" rollout restart \
  deployment/rancher-monitoring-dashboards-monitoring-proxy
kubectl -n "$MONITORING_NAMESPACE" rollout status \
  deployment/rancher-monitoring-dashboards-monitoring-proxy \
  --timeout=3m

That restart fixed the HTTP 504 we saw after recreating the upstream release.

Phase 5: Add dashboards as Kubernetes objects

For dashboard 17682, Kubernetes PersistentVolume Overview, we downloaded the JSON, removed __inputs, and normalized every Prometheus datasource UID to prometheus. The checked-in object follows this shape:

apiVersion: v1
kind: ConfigMap
metadata:
  name: grafana-dashboard-kubernetes-persistentvolume-overview
  namespace: cattle-dashboards
  labels:
    grafana_dashboard: "1"
    app.kubernetes.io/part-of: custom-grafana-dashboards
  annotations:
    grafana.com/dashboard-id: "17682"
    grafana.com/dashboard-revision: "1"
data:
  kubernetes-persistentvolume-overview.json: |-
    {
      "id": null,
      "uid": "peR80gTGk",
      "title": "Kubernetes PersistentVolume Overview",
      "schemaVersion": 38,
      "version": 1,
      "panels": [
        {
          "id": 1,
          "type": "timeseries",
          "title": "PersistentVolume usage",
          "datasource": {
            "type": "prometheus",
            "uid": "prometheus"
          },
          "targets": [
            {
              "refId": "A",
              "expr": "100 * kubelet_volume_stats_used_bytes / kubelet_volume_stats_capacity_bytes"
            }
          ]
        }
      ]
    }

The shortened JSON above documents the object shape; use the complete dashboard export in the applied manifest.

kubectl apply -f grafana-dashboard-kubernetes-persistentvolume-overview.yaml

For further dashboards, repeat the same workflow: pin a dashboard revision, keep a stable filename and UID, normalize datasource UIDs, remove import-only inputs, add grafana_dashboard: "1", and apply the ConfigMap. Review third-party PromQL before deploying it. A dashboard may assume metric names or labels your exporters do not provide.

Sidecar-provisioned dashboards are declarative. With allowUiUpdates: false, edit the ConfigMap source rather than Grafana UI; a later sidecar reconciliation will otherwise replace the UI change. A custom dashboard appears in Grafana, but it does not automatically become a new Rancher navigation item. Rancher's built-in links target known dashboard UIDs.

Phase 6: Route existing alerts with AlertmanagerConfig

An AlertmanagerConfig routes alerts already produced by PrometheusRule objects. It does not create an alert rule. We needed no custom rule for the n8n integration because the stack's existing rules already emitted the relevant alerts.

First configure the Alertmanager custom resource to select only explicitly labelled configs from the protected monitoring namespace:

alertmanager:
  alertmanagerSpec:
    alertmanagerConfigSelector:
      matchLabels:
        alertmanagerconfig: n8n-webhook
    alertmanagerConfigNamespaceSelector:
      matchLabels:
        kubernetes.io/metadata.name: cattle-monitoring-system
    alertmanagerConfigMatcherStrategy:
      type: None

OnNamespace is the operator's default matcher strategy. It adds namespace scoping so a config only processes alerts whose namespace label equals the config's namespace. A central config stored in cattle-monitoring-system would therefore miss most workload alerts. None permits cluster-wide matching, while the two selectors above limit who can provide that cluster-wide route. These semantics are defined in the Prometheus Operator API reference.

For an immediately understandable example, the placeholder URL is written directly in the AlertmanagerConfig. This keeps the routing logic visible without an additional Secret object:

apiVersion: monitoring.coreos.com/v1alpha1
kind: AlertmanagerConfig
metadata:
  name: n8n-webhook
  namespace: cattle-monitoring-system
  labels:
    alertmanagerconfig: n8n-webhook
spec:
  route:
    receiver: n8n-webhook
    groupBy:
      - namespace
      - alertname
    groupWait: 30s
    groupInterval: 5m
    repeatInterval: 4h
    matchers:
      - name: alertname
        matchType: "!~"
        value: "Watchdog|InfoInhibitor"
  receivers:
    - name: n8n-webhook
      webhookConfigs:
        - url: "https://n8n.example.com/webhook/..."
          sendResolved: true
          timeout: 10s

Apply the Helm values change and the object:

helm upgrade kube-prometheus-stack \
  prometheus-community/kube-prometheus-stack \
  --version 88.3.0 \
  --namespace "$MONITORING_NAMESPACE" \
  --values kube-prometheus-stack-values.yaml \
  --wait --timeout 15m

kubectl apply -f alertmanager-n8n-webhook.yaml

Testing this route with a real host alert—in our case the NodeClockNotSynchronising time-drift warning—and its subsequent resolved notification confirmed end-to-end delivery to n8n.

Treat the webhook path like an access token: in production, replace the placeholder during deployment through the established secret-management process, but do not commit the real URL to Git.


Verification Checklist

Do not call the migration complete merely because both Helm releases say deployed. These five checks cover the integration boundaries.

  • Workloads and storage: every long-running pod is Ready, Prometheus has the intended PVC, and no unexpected emptyDir replaced durable storage.

    helm -n "$MONITORING_NAMESPACE" list
    kubectl -n "$MONITORING_NAMESPACE" get pods,pvc
    
  • Rancher-facing proxies: Grafana, Prometheus, and Alertmanager answer through the proxy Services.

    kubectl -n "$MONITORING_NAMESPACE" exec \
      deployment/rancher-monitoring-dashboards-monitoring-proxy -- \
      wget -qO- http://rancher-monitoring-grafana/api/health
    
    kubectl -n "$MONITORING_NAMESPACE" exec \
      deployment/rancher-monitoring-dashboards-monitoring-proxy -- \
      wget -qO- http://rancher-monitoring-prometheus:9090/-/ready
    
    kubectl -n "$MONITORING_NAMESPACE" exec \
      deployment/rancher-monitoring-dashboards-monitoring-proxy -- \
      wget -qO- http://rancher-monitoring-alertmanager:9093/-/ready
    
  • Target discovery: the active-target API has no unexplained down targets, and application ServiceMonitors appear in the generated Prometheus configuration.

    kubectl -n "$MONITORING_NAMESPACE" port-forward \
      svc/kube-prometheus-stack-prometheus 9090:9090
    curl -fsS 'http://127.0.0.1:9090/api/v1/targets?state=active' \
      | jq -r '[.data.activeTargets[].health] | group_by(.)[] | "\(.[0]): \(length)"'
    
  • Dashboard provisioning and UI integration: Rancher's Monitoring menu opens Grafana, Prometheus Targets, and Alertmanager; Grafana search returns Rancher's dashboards and the custom UID peR80gTGk.

    kubectl -n "$MONITORING_NAMESPACE" port-forward \
      svc/kube-prometheus-stack-grafana 3000:80
    curl -fsS 'http://127.0.0.1:3000/api/search?query=Rancher' | jq length
    
  • Alert delivery: Alertmanager's rendered configuration contains the n8n receiver, its last reload succeeded, a controlled test alert arrives once, and the resolved notification arrives after clearing it. Check receiver failure counters before relying on the path.

Our final state had both releases deployed, every pod Ready, all 15 active Prometheus targets up, 16 Rancher dashboards loaded, and successful n8n delivery. Those counts describe this cluster, not a universal expected value.


What We Keep After the Migration

The visible result is a new dashboard. The lasting payoff is a cleaner ownership model.

We can upgrade Prometheus Operator after reading its upstream CRD and breaking-change notes without waiting for a Rancher fork. Rancher dashboards can move on their own schedule. Application teams keep using standard ServiceMonitor, PodMonitor, and PrometheusRule resources. Custom Grafana dashboards and alert routes live as reviewable Kubernetes manifests rather than UI-only state.

The platform team still owns the hard parts: tested chart pins, CRD upgrades, storage and retention, selector boundaries, webhook credentials, and restore drills. Splitting the charts does not remove that work. It puts each responsibility in a place where upstream documentation, tooling, and operational experience apply directly.

My senior advice is simple: rehearse the transition in a cluster with representative CRDs and storage, keep the old PV for a time-bounded rollback window, and test the Rancher proxy after every runtime recreation. Treat dashboards and alerts as code from the first day. The result is standard Prometheus infrastructure with a small Rancher integration layer, not a monitoring stack locked to a vendor-specific release train.

That ownership boundary is also where an experienced second review can prevent expensive mistakes. If you are planning a Rancher monitoring migration, CRD handover, or Prometheus storage design, get in touch; a short review before uninstalling the old chart is cheaper than reconstructing deleted monitoring APIs or unexplained metric history afterward.

Related Posts