SysOps

Prometheus Metric Design: Type Selection and Cardinality Control

Choosing the wrong metric type makes rate() and alerting rules produce wrong conclusions, and a single user_id label can take down a Prometheus cluster. This guide gives decision rules for Counter, Gauge, Histogram and Summary, a method for estimating and bounding cardinality, naming and label conventions, and how to catch high-cardinality metrics during review.

By LaoHand Team·10 min read·Updated 2026-09-30

Choosing a type: Counter, Gauge, Histogram, Summary

The rule compresses to one line: monotonically accumulating, never decreasing, is a Counter (total HTTP requests, queue enqueues); values that swing in both directions where the current value matters are a Gauge (memory usage, connection count, queue length); when what you care about is the *distribution* of a value, use Histogram or Summary (request latency, payload size). Picking the wrong type is not a naming problem — it makes `rate()` and alerting rules compute the wrong thing.

The classic mistake is modelling "current open connections" as a Counter, which makes `rate()` yield negative or meaningless values and silently breaks capacity alerts. The mirror-image mistake is making "total requests" a Gauge: it does reset on restart, but under Gauge semantics you cannot distinguish a reset from counter wraparound, and the long-term graph shows inexplicable cliffs.

The decisive difference between Histogram and Summary is who computes quantiles. A Histogram buckets observations client-side into `_bucket`, `_sum`, and `_count` series, so quantiles can be computed and re-aggregated server-side at any time. A Summary uses a quantile reservoir client-side, exposing `_sum`, `_count`, and non-aggregatable `quantile` labels. If your series branch and must be summed — an aggregatable service, many instances combined — Histogram is mandatory.

Histogram cost is direct: `buckets` multiplies straight into cardinality, with N buckets producing N+1 series for `_bucket` (which carries an `le` label). Design buckets around your SLO boundaries instead of copying defaults. If the P99 objective is 200 ms, then 100/150/200/300/500 ms must exist and the rest can go. Summary instead costs client-side memory for the reservoir and cannot merge quantiles across instances.

# Counter:只增不减,进程重启才归零
http_requests_total{method="GET", status="200"} 10240
# Histogram:可跨实例再聚合
http_request_duration_seconds_bucket{le="0.2"} 942
http_request_duration_seconds_sum 187.3
http_request_duration_seconds_count 1024

Cardinality: compute it before you ship it

Cardinality estimation is simple: time series ≈ number of metric names × the product of all label value combinations. With 4 methods, 8 statuses, and 30 endpoints, that metric alone is 4 × 8 × 30 = 960 series. Every label dimension multiplies rather than adds — the most underestimated fact in this whole area.

Practical thresholds: a single Prometheus instance stays comfortable below roughly one million active series; past two million, query latency and memory pressure become obvious. One metric combining full HTTP status codes (hundreds of non-standard values) with dozens of endpoints can trivially reach hundreds of thousands.

Dangerous label sources fall into a fixed set: user IDs, request IDs, session IDs, full URL paths including query strings, timestamps, and full error message text. All of them are unbounded sets — the concept of an upper bound does not apply, so no static limit configuration can save you.

Relatively safe sources: method, status class (collapse statuses into 2xx/4xx/5xx), route templates (`/users/:id`, not `/users/12345`), and shard or instance identifiers. The test is "can the value set be enumerated and kept small" — if the answer is no, it does not belong in a label.

promtool tsdb analyze metrics.prom.gz
# 输出中看 Total series 与各 label 的基数贡献
# 逐指标查看基数
curl -s http://localhost:9090/api/v1/label/__name__/values | jq '.data | length'

Naming and label conventions

Follow `<namespace>_<subsystem>_<unit>`, using base units as a suffix: seconds as `_seconds`, not `_ms`; bytes as `_bytes`. Only then do `rate()` and unit conversion stay consistent — a metric named `foo_duration_ms` is guaranteed to have someone forget the conversion.

Counter names should express the cumulative meaning. Community convention is a plain noun: `http_requests_total`, `process_cpu_seconds_total`. The `_total` suffix is not mandatory, but it lets a reviewer identify a Counter at a glance and prevents type confusion. Histogram-derived metrics naturally produce `_bucket`, `_sum`, and `_count` under an unsuffixed base name.

Label names are lower snake_case and self-explanatory: `method`, `status`, `endpoint`, `instance`. Label names are themselves cardinality dimensions — five labels versus three is an order-of-magnitude difference in series count. Never split into two labels what one can express.

Express a given piece of meaning exactly once. A common incident is having both `status` and `status_class`, which can disagree and cannot be aligned during aggregation. The clean approach is to normalize to the class client-side and keep a single label.

# 规范示例
http_requests_total{method,code_class,endpoint}   # Counter,无 _total 之外的后缀
http_request_duration_seconds_bucket{le,endpoint}  # Histogram
process_resident_memory_bytes                          # Gauge,单位明确

Catching cardinality problems during review

The most effective guardrail is running `promtool check metrics` in CI, which rejects malformed metrics outright. Going further, treat the service `/metrics` output as a CI artifact and run `promtool tsdb analyze` over it with a threshold gate, failing the build on breach. This is far cheaper than discovering the problem on a dashboard.

At runtime, the TSDB status API is the direct entry point: `/api/v1/status/tsdb` returns the hottest series grouped by metric name, and `/api/v1/status/tsdb?limit=N` gives the highest-consuming label combinations. This is the fastest way to answer "which metric is eating the memory".

For an existing high-cardinality metric there are two ways to stop the bleeding. Add `metric_relabel_configs` to drop unwanted series at scrape time (cheapest, lives in Prometheus config), or fix the exporter (more expensive but thorough). Both require evaluating the effect on existing dashboards and alerting rules so that rules do not fail silently.

Long term, alerting on cardinality pays off: `prometheus_tsdb_head_series` is a built-in metric, and thresholding it catches the problem while it is still small. Most teams do not lack remediation knowledge — they discover too late. By the time queries time out, dozens of dashboards and rules already depend on those series, which makes removal extremely expensive.

promtool check metrics < app.txt
promtool tsdb analyze metrics.txt
curl -s "http://localhost:9090/api/v1/status/tsdb?limit=15"
# 抓取端丢弃高基数序列
metric_relabel_configs:
  - source_labels: [user_id]
    regex: ".*"
    action: drop

Official References

Each command links to its official documentation below, so you can verify the latest usage and read deeper.