Rancher 2.15 Monitoring Migration: The Missing Practitioner Runbook
A battle-tested migration guide from legacy rancher-monitoring to upstream kube-prometheus-stack and rancher-monitoring-dashboards, including CRDs, PVCs, custom dashboards, and Alertmanager.

Contents

A battle-tested migration guide from legacy rancher-monitoring to upstream kube-prometheus-stack and rancher-monitoring-dashboards, including CRDs, PVCs, custom dashboards, and Alertmanager.
Share this article
Rancher 2.15 Splits Monitoring into Two Lifecycles
Rancher 2.15 fundamentally changes how Kubernetes monitoring is delivered. The monolithic, all-in-one rancher-monitoring chart gives way to an unbundled architecture: upstream kube-prometheus-stack supplies Prometheus Operator, Prometheus, Alertmanager, Grafana, exporters, and rules, while rancher-monitoring-dashboards adds Rancher's dashboards and the service proxies used by the Rancher UI.
In the official Rancher 2.15 documentation, the transition can look almost like a two-click replacement: install the runtime, then add the dashboard integration. In real clusters, the risks sit at the new ownership boundaries. Proxy failures remain silent until the UI returns HTTP 504, CRDs follow a lifecycle decoupled from the Helm release, and restrictive selectors produce scrape configurations without the expected targets.
In the monitoring proxy, the break may initially surface only as:
host not found in upstream
"kube-prometheus-stack-grafana.cattle-monitoring-system.svc.cluster.local"
Even after the expected Services exist, an already-running NGINX proxy can retain stale ClusterIPs: the workloads look healthy, yet Rancher still responds with HTTP 504. Independently, existing ServiceMonitor objects can be absent from the generated scrape configuration because of the default selectors.
This field-tested transition manual is written for platform teams upgrading existing monitoring installations or setting up new Rancher 2.15 clusters. It covers not only installation order but also CRD ownership, metrics history, storage, selectors, dashboards, alert routing, and recovery. Existing legacy installations remain supported, but the split stack is the official path forward for new installations.
What Actually Changed
The old chart was convenient because Rancher owned the whole monitoring release. It also coupled the lifecycle of an upstream observability stack to Rancher's release process. The new split puts the runtime back on the widely used prometheus-community Helm chart while keeping Rancher-specific UI assets separate.
| Concern | Legacy rancher-monitoring |
Rancher 2.15 split stack |
|---|---|---|
| Runtime | Bundled, Rancher-maintained distribution | Upstream kube-prometheus-stack |
| Rancher UI | Bundled in the same release | rancher-monitoring-dashboards and its proxy |
| Operator and CRDs | Coupled to Rancher chart versions | Pinned and upgraded with the upstream chart |
| Grafana dashboards | Rancher defaults and runtime move together | Dashboard ConfigMaps can change independently |
| Upgrade ownership | One convenient but coarse lifecycle | Two explicit, independently pinned lifecycles |
| Portability | Rancher chart conventions | Standard Prometheus Operator resources |
| Operational cost | Fewer initial choices | More values and version management, but clearer ownership |
This is a sensible boundary. Prometheus, Grafana, and Alertmanager no longer need a Rancher-specific fork merely to appear inside Rancher. The trade-off is real: a one-click app becomes two releases plus a values file that the platform team must own.
The boundary also explains the naming contract. With the upstream release named kube-prometheus-stack, Rancher's proxy expects these Services:
kube-prometheus-stack-grafana
kube-prometheus-stack-prometheus
kube-prometheus-stack-alertmanager
It then exposes Rancher-facing Services with the established names:
rancher-monitoring-grafana
rancher-monitoring-prometheus
rancher-monitoring-alertmanager
Changing the Helm release name is possible only if the dashboard chart's upstream Service values are changed with it. Keeping the documented release name removes an unnecessary integration variable.
The Missing Manual: Five Ways a Naive Migration Breaks
1. CRDs outlive releases and have cluster-wide blast radius
Prometheus, Alertmanager, ServiceMonitor, PodMonitor, PrometheusRule, and AlertmanagerConfig are not namespaced implementation details of a Helm release. Their definitions are cluster-scoped APIs, and their objects may be used by teams in many namespaces.
Helm installs CRDs from a chart's crds/ directory before regular resources, skips a CRD that already exists, and deliberately does not upgrade or delete CRDs. That policy avoids implicit data loss, but it leaves version and ownership decisions with the operator. The behavior is documented in Helm's CRD best practices and is the reason kube-prometheus-stack has explicit CRD upgrade notes.
Never start by deleting *.monitoring.coreos.com CRDs. Deleting a CRD also deletes every custom resource stored behind that API. Inventory and export the resources first. There should be one active Prometheus Operator responsible for a given set of resources during the handover, not two operators racing over them.
There are two defensible CRD strategies:
- Clean handover: export all monitoring resources, remove the legacy runtime, verify that no other release owns or needs its CRDs, update the CRDs to the version required by the new operator, and install the upstream stack.
- Independent CRD lifecycle: manage the Prometheus Operator CRDs separately, install a compatible version before the new operator, and set
crds.enabled: falseinkube-prometheus-stack. This adds process, but it makes cluster API ownership explicit.
The chart's crds.upgradeJob can apply newer definitions, but in chart 88.3.0 that feature is marked preview and uses cluster-scoped RBAC. Treat it as a deliberate platform decision, not as a box to tick during an outage.
2. A retained PVC is not yet a metrics migration
Prometheus stores local TSDB blocks and a write-ahead log on its data volume. Its local storage is neither clustered nor replicated. The Prometheus project recommends snapshots for consistent backups and warns that copying a live TSDB directory can omit recent data or capture an incoherent state. The official storage documentation explains the block and WAL layout.
Before removing the old release, choose one of these outcomes:
- Clean cutoff: start the new Prometheus with an empty PVC and retain the old volume for a defined rollback window.
- Snapshot and restore: enable the TSDB admin API temporarily, create a snapshot, copy it using storage-native tooling, and test restoration into a compatible Prometheus version. This is a migration project of its own.
- Continuous history outside the pod: use remote write or a long-term system such as Thanos, Mimir, or another compatible backend. This is the cleanest long-term boundary but not a last-minute migration toggle.
Do not mount the old PVC into the new StatefulSet merely because its filesystem is visible. Check Prometheus version compatibility, volume access mode, security context, StatefulSet claim naming, and rollback behavior first.
3. ServiceMonitors can exist while Prometheus ignores them
Prometheus Operator uses label and namespace selectors to associate a Prometheus object with ServiceMonitor and PodMonitor objects. The upstream chart normally narrows discovery to monitors carrying its Helm release labels. Rancher-provided monitors and application-owned monitors may not have those labels.
The failure is quiet: the ServiceMonitor exists, the target Service exists, and Prometheus still has no scrape job. Prometheus Operator's troubleshooting guide describes how to inspect the generated configuration.
For a cluster-wide platform Prometheus, we used:
prometheus:
prometheusSpec:
serviceMonitorSelectorNilUsesHelmValues: false
podMonitorSelectorNilUsesHelmValues: false
serviceMonitorSelector: {}
serviceMonitorNamespaceSelector: {}
podMonitorSelector: {}
podMonitorNamespaceSelector: {}
An empty selector matches all relevant objects; a null selector has different semantics depending on the field. This broad discovery is appropriate only when the platform team controls who may create monitoring resources. In a multi-tenant cluster, use namespace labels and an explicit monitor label instead of accepting arbitrary scrape configurations cluster-wide.
4. Rancher dashboards are provisioned files, not Grafana plugins
The dashboard chart does not install a Grafana extension CRD. It creates dashboard ConfigMaps, normally in cattle-dashboards, labelled:
grafana_dashboard: "1"
The Grafana sidecar watches labelled ConfigMaps and Secrets across namespaces, writes their JSON payloads to /tmp/dashboards, and triggers Grafana's provisioning reload. Grafana then loads those files through a provisioning provider. The underlying behavior is standard Grafana dashboard provisioning.
That means custom dashboards can be GitOps-managed as Kubernetes objects. It also means imported Grafana.com JSON needs normalization. The interactive Grafana import flow substitutes datasource variables such as ${DS_PROMETHEUS}; file provisioning does not. Hard-coded datasource UIDs are equally brittle. Normalize every Prometheus datasource reference to the UID created by this chart, prometheus, and remove __inputs before putting the JSON in a ConfigMap.
5. Defaults are installation defaults, not capacity planning
Scrape interval, retention, sample cardinality, rule count, and Grafana concurrency determine actual resource use. The 1,750 MiB request and 2,500 MiB limit in Rancher's K3s guidance were a workable starting point for our small cluster, not a universal capacity figure.
The Prometheus admission webhook should remain enabled so malformed PrometheusRule resources are rejected before they can break rule loading. If its certificate patch job or webhook endpoint fails, fix that dependency. Permanently changing the failure policy to ignore merely turns an installation problem into a delayed monitoring problem.
Execution Runbook
The commands below assume a kubeconfig named rancher.yaml, context rancher, and namespace cattle-monitoring-system. Pin chart versions that you have tested. We used kube-prometheus-stack 88.3.0 with Prometheus Operator 0.93.0 and rancher-monitoring-dashboards 110.0.0+up0.1.2.
Phase 0: Freeze changes and capture the old state
Run this before uninstalling anything:
export KUBECONFIG="$PWD/rancher.yaml"
export MONITORING_NAMESPACE="cattle-monitoring-system"
kubectl config current-context
helm list -A | grep -E 'rancher-monitoring|kube-prometheus'
helm -n "$MONITORING_NAMESPACE" \
get values rancher-monitoring --all > legacy-rancher-monitoring-values.yaml
kubectl get crd | grep monitoring.coreos.com
kubectl get prometheuses,alertmanagers,servicemonitors,podmonitors,prometheusrules \
-A -o wide
kubectl -n "$MONITORING_NAMESPACE" get pvc
kubectl get pv -o custom-columns='NAME:.metadata.name,CLAIM_NAMESPACE:.spec.claimRef.namespace,CLAIM_NAME:.spec.claimRef.name,RECLAIM:.spec.persistentVolumeReclaimPolicy,CLASS:.spec.storageClassName'
Export the custom resources. Store the export securely; rules and endpoints may disclose internal topology.
for resource in \
prometheuses.monitoring.coreos.com \
alertmanagers.monitoring.coreos.com \
servicemonitors.monitoring.coreos.com \
podmonitors.monitoring.coreos.com \
prometheusrules.monitoring.coreos.com \
alertmanagerconfigs.monitoring.coreos.com
do
kubectl get "$resource" -A -o yaml > "backup-${resource%%.*}.yaml"
done
Record the old and new Prometheus versions. Decide now whether history is being cut over, snapshotted, or retained remotely. If the old PV must survive deletion of its claim, confirm the storage driver's behavior and set the PV reclaim policy to Retain before removing the claim. A Kubernetes Retain policy protects the PV object from automatic backend deletion; it is not a backup.
Phase 1: Remove the legacy runtime without deleting its APIs blindly
Stop changes to monitoring resources during the handover. Remove the legacy rancher-monitoring App from Rancher Apps > Installed Apps, or use Helm if Helm directly owns that release:
helm -n "$MONITORING_NAMESPACE" uninstall rancher-monitoring
Do not automatically uninstall a separate CRD release and do not run kubectl delete crd. First verify that the legacy Prometheus, Alertmanager, operator Deployment, and webhooks are gone while the exported ServiceMonitor and PrometheusRule objects remain available.
kubectl -n "$MONITORING_NAMESPACE" get \
deploy,statefulset,daemonset,job,pod,pvc
kubectl get validatingwebhookconfiguration,mutatingwebhookconfiguration \
| grep -E 'prometheus|monitoring' || true
If old chart hooks, webhooks, ClusterRoles, or Services collide with the new release, identify their owning release from labels and annotations. Do not solve an ownership error by overwriting metadata until you know which controller is still using the object.
Before starting a newer operator against retained CRDs, review the actual schema diff and apply the CRDs shipped with the chart version you are about to install:
helm repo add prometheus-community \
https://prometheus-community.github.io/helm-charts \
--force-update
helm repo update
helm show crds prometheus-community/kube-prometheus-stack \
--version 88.3.0 > kube-prometheus-stack-crds-88.3.0.yaml
kubectl diff -f kube-prometheus-stack-crds-88.3.0.yaml
kubectl apply --server-side \
--field-manager=kube-prometheus-stack-crds \
-f kube-prometheus-stack-crds-88.3.0.yaml
Stop if the server-side apply reports ownership conflicts. Resolve them against the controller that currently manages those fields; do not add --force-conflicts as an automatic migration flag. If CRDs have an independent lifecycle in your platform, commit this rendered artifact, review it like any other cluster API change, and use crds.enabled: false below.
Phase 2: Build an owned values.yaml
This is the compact values file we used on a single-node K3s cluster. It contains the Rancher integration contract and the fixes found during installation:
crds:
enabled: true
upgradeJob:
enabled: false
grafana:
grafana.ini:
security:
allow_embedding: true
auth:
disable_login_form: false
auth.anonymous:
enabled: true
org_role: Viewer
dashboards:
default_home_dashboard_path: /tmp/dashboards/rancher-default-home.json
users:
auto_assign_org_role: Viewer
prometheus:
prometheusSpec:
serviceMonitorSelectorNilUsesHelmValues: false
podMonitorSelectorNilUsesHelmValues: false
serviceMonitorSelector: {}
serviceMonitorNamespaceSelector: {}
podMonitorSelector: {}
podMonitorNamespaceSelector: {}
scrapeInterval: 30s
evaluationInterval: 30s
retention: 15d
resources:
requests:
cpu: 250m
memory: 1750Mi
limits:
memory: 2500Mi
prometheusOperator:
admissionWebhooks:
enabled: true
# K3s needs Rancher's PushProx path for these control-plane components.
kubeEtcd:
enabled: false
kubeControllerManager:
enabled: false
kubeScheduler:
enabled: false
kubeProxy:
enabled: false
On a conventional Kubernetes distribution, review the four disabled component monitors rather than copying the K3s settings.
For production persistence, add a tested StorageClass and leave headroom between the PVC size and Prometheus retention size:
prometheus:
prometheusSpec:
retention: 15d
retentionSize: 40GB
storageSpec:
volumeClaimTemplate:
spec:
storageClassName: longhorn
accessModes:
- ReadWriteOnce
resources:
requests:
storage: 50Gi
grafana:
persistence:
enabled: true
storageClassName: longhorn
size: 5Gi
Replace longhorn with a StorageClass you actually operate and restore-test. Provisioned dashboards do not depend on Grafana's database, but users, UI edits, preferences, and other mutable state do.
Phase 3: Install the upstream runtime first
helm repo add prometheus-community \
https://prometheus-community.github.io/helm-charts \
--force-update
helm repo update
helm upgrade --install kube-prometheus-stack \
prometheus-community/kube-prometheus-stack \
--version 88.3.0 \
--namespace "$MONITORING_NAMESPACE" \
--create-namespace \
--values kube-prometheus-stack-values.yaml \
--wait \
--timeout 15m
Wait for Prometheus, Alertmanager, Grafana, the operator, kube-state-metrics, and node-exporter. A Helm release stuck in pending-install is not a candidate for repeated blind retries. Inspect hooks, jobs, events, webhook certificates, and PVCs first.
helm -n "$MONITORING_NAMESPACE" status kube-prometheus-stack
kubectl -n "$MONITORING_NAMESPACE" get pods,pvc
kubectl -n "$MONITORING_NAMESPACE" get events \
--sort-by=.lastTimestamp | tail -40
Phase 4: Install the Rancher dashboard integration
Install rancher-monitoring-dashboards from the Rancher 2.15 chart catalog only after the three upstream Services exist. The Rancher UI is useful here because it injects the local cluster ID, cluster name, Rancher URL, and System Project ID.
In Apps > Charts, select rancher-monitoring-dashboards, install it into cattle-monitoring-system, and keep the default upstream Service names when the runtime release is named kube-prometheus-stack. Keep k3sServer.enabled: true for K3s; disable it for other distributions as the chart describes.
If the dashboard chart was installed too early, upgrade or retry it after the runtime is ready. If the runtime was uninstalled and reinstalled, restart only the proxy so NGINX resolves the new Service ClusterIPs:
kubectl -n "$MONITORING_NAMESPACE" rollout restart \
deployment/rancher-monitoring-dashboards-monitoring-proxy
kubectl -n "$MONITORING_NAMESPACE" rollout status \
deployment/rancher-monitoring-dashboards-monitoring-proxy \
--timeout=3m
That restart fixed the HTTP 504 we saw after recreating the upstream release.
Phase 5: Add dashboards as Kubernetes objects
For dashboard 17682, Kubernetes PersistentVolume Overview, we downloaded the JSON, removed __inputs, and normalized every Prometheus datasource UID to prometheus. The checked-in object follows this shape:
apiVersion: v1
kind: ConfigMap
metadata:
name: grafana-dashboard-kubernetes-persistentvolume-overview
namespace: cattle-dashboards
labels:
grafana_dashboard: "1"
app.kubernetes.io/part-of: custom-grafana-dashboards
annotations:
grafana.com/dashboard-id: "17682"
grafana.com/dashboard-revision: "1"
data:
kubernetes-persistentvolume-overview.json: |-
{
"id": null,
"uid": "peR80gTGk",
"title": "Kubernetes PersistentVolume Overview",
"schemaVersion": 38,
"version": 1,
"panels": [
{
"id": 1,
"type": "timeseries",
"title": "PersistentVolume usage",
"datasource": {
"type": "prometheus",
"uid": "prometheus"
},
"targets": [
{
"refId": "A",
"expr": "100 * kubelet_volume_stats_used_bytes / kubelet_volume_stats_capacity_bytes"
}
]
}
]
}
The shortened JSON above documents the object shape; use the complete dashboard export in the applied manifest.
kubectl apply -f grafana-dashboard-kubernetes-persistentvolume-overview.yaml
For further dashboards, repeat the same workflow: pin a dashboard revision, keep a stable filename and UID, normalize datasource UIDs, remove import-only inputs, add grafana_dashboard: "1", and apply the ConfigMap. Review third-party PromQL before deploying it. A dashboard may assume metric names or labels your exporters do not provide.
Sidecar-provisioned dashboards are declarative. With allowUiUpdates: false, edit the ConfigMap source rather than Grafana UI; a later sidecar reconciliation will otherwise replace the UI change. A custom dashboard appears in Grafana, but it does not automatically become a new Rancher navigation item. Rancher's built-in links target known dashboard UIDs.
Phase 6: Route existing alerts with AlertmanagerConfig
An AlertmanagerConfig routes alerts already produced by PrometheusRule objects. It does not create an alert rule. We needed no custom rule for the n8n integration because the stack's existing rules already emitted the relevant alerts.
First configure the Alertmanager custom resource to select only explicitly labelled configs from the protected monitoring namespace:
alertmanager:
alertmanagerSpec:
alertmanagerConfigSelector:
matchLabels:
alertmanagerconfig: n8n-webhook
alertmanagerConfigNamespaceSelector:
matchLabels:
kubernetes.io/metadata.name: cattle-monitoring-system
alertmanagerConfigMatcherStrategy:
type: None
OnNamespace is the operator's default matcher strategy. It adds namespace scoping so a config only processes alerts whose namespace label equals the config's namespace. A central config stored in cattle-monitoring-system would therefore miss most workload alerts. None permits cluster-wide matching, while the two selectors above limit who can provide that cluster-wide route. These semantics are defined in the Prometheus Operator API reference.
For an immediately understandable example, the placeholder URL is written directly in the AlertmanagerConfig. This keeps the routing logic visible without an additional Secret object:
apiVersion: monitoring.coreos.com/v1alpha1
kind: AlertmanagerConfig
metadata:
name: n8n-webhook
namespace: cattle-monitoring-system
labels:
alertmanagerconfig: n8n-webhook
spec:
route:
receiver: n8n-webhook
groupBy:
- namespace
- alertname
groupWait: 30s
groupInterval: 5m
repeatInterval: 4h
matchers:
- name: alertname
matchType: "!~"
value: "Watchdog|InfoInhibitor"
receivers:
- name: n8n-webhook
webhookConfigs:
- url: "https://n8n.example.com/webhook/..."
sendResolved: true
timeout: 10s
Apply the Helm values change and the object:
helm upgrade kube-prometheus-stack \
prometheus-community/kube-prometheus-stack \
--version 88.3.0 \
--namespace "$MONITORING_NAMESPACE" \
--values kube-prometheus-stack-values.yaml \
--wait --timeout 15m
kubectl apply -f alertmanager-n8n-webhook.yaml
Testing this route with a real host alert—in our case the NodeClockNotSynchronising time-drift warning—and its subsequent resolved notification confirmed end-to-end delivery to n8n.
Treat the webhook path like an access token: in production, replace the placeholder during deployment through the established secret-management process, but do not commit the real URL to Git.
Verification Checklist
Do not call the migration complete merely because both Helm releases say deployed. These five checks cover the integration boundaries.
-
Workloads and storage: every long-running pod is Ready, Prometheus has the intended PVC, and no unexpected
emptyDirreplaced durable storage.helm -n "$MONITORING_NAMESPACE" list kubectl -n "$MONITORING_NAMESPACE" get pods,pvc -
Rancher-facing proxies: Grafana, Prometheus, and Alertmanager answer through the proxy Services.
kubectl -n "$MONITORING_NAMESPACE" exec \ deployment/rancher-monitoring-dashboards-monitoring-proxy -- \ wget -qO- http://rancher-monitoring-grafana/api/health kubectl -n "$MONITORING_NAMESPACE" exec \ deployment/rancher-monitoring-dashboards-monitoring-proxy -- \ wget -qO- http://rancher-monitoring-prometheus:9090/-/ready kubectl -n "$MONITORING_NAMESPACE" exec \ deployment/rancher-monitoring-dashboards-monitoring-proxy -- \ wget -qO- http://rancher-monitoring-alertmanager:9093/-/ready -
Target discovery: the active-target API has no unexplained
downtargets, and application ServiceMonitors appear in the generated Prometheus configuration.kubectl -n "$MONITORING_NAMESPACE" port-forward \ svc/kube-prometheus-stack-prometheus 9090:9090 curl -fsS 'http://127.0.0.1:9090/api/v1/targets?state=active' \ | jq -r '[.data.activeTargets[].health] | group_by(.)[] | "\(.[0]): \(length)"' -
Dashboard provisioning and UI integration: Rancher's Monitoring menu opens Grafana, Prometheus Targets, and Alertmanager; Grafana search returns Rancher's dashboards and the custom UID
peR80gTGk.kubectl -n "$MONITORING_NAMESPACE" port-forward \ svc/kube-prometheus-stack-grafana 3000:80 curl -fsS 'http://127.0.0.1:3000/api/search?query=Rancher' | jq length -
Alert delivery: Alertmanager's rendered configuration contains the n8n receiver, its last reload succeeded, a controlled test alert arrives once, and the resolved notification arrives after clearing it. Check receiver failure counters before relying on the path.
Our final state had both releases deployed, every pod Ready, all 15 active Prometheus targets up, 16 Rancher dashboards loaded, and successful n8n delivery. Those counts describe this cluster, not a universal expected value.
What We Keep After the Migration
The visible result is a new dashboard. The lasting payoff is a cleaner ownership model.
We can upgrade Prometheus Operator after reading its upstream CRD and breaking-change notes without waiting for a Rancher fork. Rancher dashboards can move on their own schedule. Application teams keep using standard ServiceMonitor, PodMonitor, and PrometheusRule resources. Custom Grafana dashboards and alert routes live as reviewable Kubernetes manifests rather than UI-only state.
The platform team still owns the hard parts: tested chart pins, CRD upgrades, storage and retention, selector boundaries, webhook credentials, and restore drills. Splitting the charts does not remove that work. It puts each responsibility in a place where upstream documentation, tooling, and operational experience apply directly.
My senior advice is simple: rehearse the transition in a cluster with representative CRDs and storage, keep the old PV for a time-bounded rollback window, and test the Rancher proxy after every runtime recreation. Treat dashboards and alerts as code from the first day. The result is standard Prometheus infrastructure with a small Rancher integration layer, not a monitoring stack locked to a vendor-specific release train.
That ownership boundary is also where an experienced second review can prevent expensive mistakes. If you are planning a Rancher monitoring migration, CRD handover, or Prometheus storage design, get in touch; a short review before uninstalling the old chart is cheaper than reconstructing deleted monitoring APIs or unexplained metric history afterward.


