Creuto is now an OpenAI Select Partner Read More

Software Architecture & Technical

Multi-tenant Prometheus: let teams query their own metrics

A multi-tenant Prometheus pattern from Adobe: kube-rbac-proxy for identity, prom-label-proxy for namespace enforcement, and the three ways it still fails.

Multi-tenant Prometheus: let teams query their own metrics

A GPU ran at zero per cent utilisation for eleven consecutive days at Adobe — allocated, powered on, billed, and unnoticed. Every second of that utilisation had already been recorded. Multi-tenant Prometheus is the problem sitting behind that story: the data existed in a central Prometheus holding thousands of namespaces, and no team could safely be let near it.

Two Adobe engineers, Bingi Narasimha Karthik and Ramkumar Nagaraj, published the pattern they used to fix it on the CNCF blog in September 2026, along with the proxy itself under Apache 2.0. This post covers how the pattern works, the custom resource that makes it self-service, the per-tenant storage option, and the three ways it can still fail.

One note on sourcing. InfoQ's coverage frames this as teams getting a safer way to see their GPU metrics, which reads like a new Kubernetes GPU feature. It is not. GPUs are the expensive thing that made the gap visible; the pattern is generic tenant-aware access to Prometheus, and it applies identically to a cluster with no GPUs in it.

What multi-tenant Prometheus has to solve

On a shared cluster, you end up with one of three arrangements, and the first two are both bad.

ArrangementWhat it costsWhat it leaks
Nobody gets access; the platform team runs queries on requestA ticket queue, and an eleven-day idle GPUNothing, which is the point, and also the problem
Everyone gets PromQL against the central PrometheusUnbounded queries against shared infrastructureEvery other team's series, labels and metric names
One Prometheus per teamDuplicate scrapes, duplicate storage, duplicate operationsNothing, but it does not scale past a handful of teams

The third is worth dwelling on, because it is the default answer and it fails quietly. Running a Prometheus per team does not scale — you pay to scrape and store the same infrastructure series once per tenant, and the platform team inherits every one of those instances as an operational surface.

The pattern Adobe published takes the second arrangement and puts three things in front of it: authentication, authorisation, and label enforcement that happens below the query language rather than inside it.

The access path, layer by layer

A request from a team engineer traverses four stages before it reaches any data.

  1. NGINX terminates the request and routes it.
  2. kube-rbac-proxy establishes who is asking and whether they are allowed. It identifies the caller from a client TLS certificate or a bearer token, running a TokenReview against the Kubernetes API for tokens, then performs a SubjectAccessReview to confirm the user holds the required RBAC role. Identity is Kubernetes-native — the same RBAC that governs everything else in the cluster.
  3. The multi-tenant proxy resolves the tenant, discovers healthy Prometheus backends through the Kubernetes API, fans the query out across them, and aggregates what comes back.
  4. prom-label-proxy rewrites the PromQL so the namespace constraint is part of the query Prometheus actually evaluates.

What makes this worth copying is that no stage is asked to do a job it is bad at. RBAC decides identity. The label proxy decides scope. The central Prometheus keeps doing exactly what it did before, unaware that it is now multi-tenant.

Why enforcement has to happen below PromQL

The instinct is to filter queries — inspect the PromQL, reject anything that touches a namespace the caller does not own. That approach loses. PromQL has label matchers, regular expressions, subqueries, joins and functions over label sets, and a filter that parses queries is playing whack-a-mole against a language designed to be expressive.

prom-label-proxy takes the other route: it rewrites. Given an enforced label, it injects or replaces the matcher in every query it forwards, so http_requests_total{namespace=~"a.*"} becomes http_requests_total{namespace="b"} when tenant B asks. The matcher the caller supplied is not validated, it is overwritten. There is no expression to smuggle past, because the constraint is added after parsing rather than checked before it.

It covers the endpoints that matter, not just /api/v1/query: query_range, query_exemplars, series, rules, alerts, federate, the label and label-values APIs behind a flag, and several Alertmanager endpoints including silences. That last one matters more than it sounds — a tenant who can silence another tenant's alerts has done real damage without reading a single series.

MetricAccess: teams declare what they need

The self-service half of the pattern is a Kubernetes custom resource. A team creates a MetricAccess object in its own namespace naming the metrics it wants, by exact name, by regular expression, or by PromQL selector:

spec:
  source: my-app-namespace
  metricIsolation: true
  metrics:
    - "http_requests_total"
    - "http_.*"
    - '{job="my-app"}'

This is the design decision we would steal first. Access becomes a reviewable object in the cluster, created through the same pull request and RBAC path as everything else a team owns, rather than a ticket to the platform team or a row in a config file only the platform team can edit. Multiple resources per namespace are supported, so a team can separate its application metrics from its infrastructure ones and have them reviewed by different people.

It also gives the platform team something it never had: a written record of which metrics each tenant actually depends on. That list is what lets you retire a metric or change a label without a week of archaeology, and it is the same reason we push teams towards declared instrumentation standards rather than whatever each service happened to emit.

Per-tenant Prometheus instances, and the 97% that disappears

The second data path is optional and, for most teams, the more interesting one. Rather than querying through the proxy every time, a tenant can run its own small Prometheus that receives only its metrics by remote write, on an interval it sets.

With metricIsolation: true, the collection query itself is routed through prom-label-proxy so that the namespace filter is applied at collection rather than only at query time. Adobe's repository puts the difference at roughly 10,000 or more series stored per tenant without isolation, against around 300 with it — about 97% less storage, on data the tenant was never entitled to see in the first place.

That is the part worth internalising. The isolation control and the cost control are the same control. A tenant Prometheus that only ever ingests its own namespace is cheaper, faster to query, and incapable of leaking, and you do not have to choose which of those three you were buying. Cardinality discipline usually arrives as a separate initiative, the way bucket and label cardinality does; here it arrives as a side effect of getting the permissions right.

The failure modes this pattern does not remove

Three things will bite you, and the projects involved say so plainly.

The tenant identity is a header

Adobe's proxy identifies the tenant from an X-Tenant-Namespace HTTP header, with a namespace query parameter as a fallback. A header is forgeable by definition. prom-label-proxy makes the same point about itself: it performs no authentication or authorisation, that has to happen before the request reaches it, and you must trust whatever supplies the label value.

So the entire boundary rests on the layers in front stripping any client-supplied tenant header and setting it from the authenticated identity. If your ingress passes an inbound X-Tenant-Namespace straight through, you have built a system where reading another team's metrics is a curl flag. Test this explicitly, with a request that sets the header to someone else's namespace, before you call the rollout done.

The write path is out of scope

prom-label-proxy enforces on reads. Its README is explicit that write tenant isolation is outside the project's scope: if tenants control scrape configuration, or honor_labels is enabled, they can pollute another tenant's metrics. Adobe's own example configuration sets honorLabels: true on remote write, which is correct for that path and worth understanding rather than copying blindly.

Worth noting too that kube-rbac-proxy describes itself as alpha stage, with flags, configuration and behaviour subject to significant change, and warns that a component receiving a bearer token can impersonate the client with it — recommending mTLS and advising against passing highly privileged tokens. Neither caveat is a reason not to use it. Both are reasons to pin versions and read release notes, which is the ordinary cost of a pattern assembled from three projects instead of bought from one.

Noisy neighbours and cardinality

Label enforcement bounds what a query returns. It does not bound what a query costs. A tenant can still issue a range query across a year at high resolution and make the shared Prometheus unpleasant for everyone. The per-tenant remote-write path is the real answer — it moves routine querying off the shared instance entirely, and the tenant's own instance is small enough that a bad query only hurts the team that wrote it. Per-tenant collection intervals and curated metric sets do the rest.

When not to build this

If you have four teams on one cluster, this is over-engineered. Give each team a Grafana organisation with a datasource scoped by the same namespace label and move on. The pattern earns its complexity somewhere around the point where you stop being able to name every namespace owner, and where the central Prometheus has become expensive enough that someone is asking about it in a cost review.

If you are already running Grafana Mimir or Cortex, you have multi-tenancy in the product and should use it rather than assembling this in front of it. The Adobe pattern exists for the very common case of one operator-managed Prometheus that grew into a shared service without anybody deciding it should.

The question to take into your next platform review is narrower than "should we do multi-tenancy". It is: if a team engineer sets the tenant header to another team's namespace tomorrow, what stops them? If the honest answer is that nobody has tried, that is the afternoon's work — and it is the same audit we run at the start of any Kubernetes platform engagement, alongside the more familiar question of whether the GPUs you are paying for are doing anything at all.

Frequently asked questions

Put authentication and Kubernetes RBAC in front of Prometheus, then use prom-label-proxy to rewrite each incoming PromQL query so a namespace matcher is enforced below the query language. Optionally remote-write each tenant only its own series into a small per-tenant Prometheus instance.

prom-label-proxy is a Prometheus Community project that enforces a given label value in Prometheus and Alertmanager API requests. It rewrites incoming PromQL to inject or replace the label matcher, so a tenant cannot widen its own scope. It performs no authentication of its own.

Direct PromQL access to a shared Prometheus is not safe, because any query can read every namespace's series and label names. Access is safe once identity is established by RBAC ahead of the query and a namespace matcher is enforced by rewriting rather than by filtering.

MetricAccess is a Kubernetes custom resource in Adobe's open-source multi-tenant proxy. A team declares, in its own namespace, which metrics it needs by exact name, regular expression or PromQL selector, plus whether collection should be namespace-isolated and remote-written to its own Prometheus.

You do not need one, but it removes most of the noisy-neighbour risk. With namespace isolation enabled at collection, Adobe reports a tenant store dropping from roughly 10,000 series to around 300, so routine querying leaves the shared instance and a costly query only affects its author.

If you already run Grafana Mimir or Cortex, use their built-in multi-tenancy rather than assembling a proxy chain in front of them. The Adobe pattern targets the common case of a single operator-managed Prometheus that became a shared service without anyone designing it to be one.

Written by

Akash Mohapatra

Akash Mohapatra

Co Founder & Director

29 Sep 2026

·

9 min read

Share

LET'S CONNECT

Connect with Creuto!

Ready to take the first step towards unlocking opportunities, realizing goals, and embracing innovation? We're here and eager to connect.

We don't just aim to fit in – we strive to stand out. Experience the perfect blend of innovation, excellence, and trust that makes us truly unforgettable. Discover the difference with Creuto.

© 2026 Creuto All Rights Reserved