Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Metered

Metered is a metric-state, composition, and schema/value collection library for Rust services. Core metered gives services readable metric state, typed metric trees, schema/value collection, and registry views that borrow through the real service graph at scrape time. Exposition formats live in sink crates such as metered-om.

Metered layers operation instrumentation. metered-tracing derives metric families from the tracing spans your code already emits. An instrumented method gets performance metrics without a second instrumentation. Everything else is plain core metric state that your code owns and updates directly.

This book does two things:

  1. Teach you to use Metered well.
  2. Teach you to build great metrics – and explain why Metered has its shape, so the choices feel inevitable rather than arbitrary.

If you have never instrumented a service before, that is fine: the Metrics & OpenMetrics primer starts from zero.

A first taste

#![allow(unused)]
fn main() {
use metered::entry::{counter, gauge};
use metered::{MetricTreeView, Unit};
use std::sync::atomic::AtomicU64;

struct Api {
    requests: AtomicU64,
    in_flight: AtomicU64,
}

fn metrics() -> MetricTreeView<'static, Api> {
    let mut view = MetricTreeView::with_prefix("api");
    view.register(counter("requests").select(|api: &Api| &api.requests).help("Requests"));
    view.register(gauge("in_flight").select(|api: &Api| &api.in_flight).help("In-flight requests").unit(Unit::Items));
    view
}
}

That view exposes the state the service already owns: request totals and current in-flight work. The default model starts with owned state and explicit composition. Tracing spans and metered-tracing add operation measurement.

To see the whole picture, run the demo:

cargo run -p order-service-demo

It prints the OpenMetrics document for a small e-commerce service. The demo turns spans into metrics, components own their layout, and a dynamic payment-rail fleet fans out by label. Demo App documents each pattern module by module.

How to read this book

  • Concepts explains the model and the reasoning behind it. Read Why Metered has this design early. It is the key to the rest.
  • Building Great Metrics is the practical craft: which metric type to reach for, how to keep labels safe, and how to keep wire names stable as code changes.
  • Instrumenting & Exposing covers the mechanics: tracing-derived operation metrics, registries, exemplars, and turning a schema into dashboards.
  • Reference covers the demo, migrating from older versions, and the feature flags.

What Metered deliberately avoids

  • Global registries and statics. Metrics are fields on your types.
  • serde on the default path. Exposition is native OpenMetrics text.
  • Reset/clear semantics. Counters and histograms are cumulative. The query engine computes rates and quantiles at query time, where they aggregate correctly across replicas.
  • Foreign types leaking through the public API. A metered upgrade does not drag serde or hdrhistogram along.

The next section explains why each of those is a feature, not a limitation.

Metrics & OpenMetrics primer

This page is for readers who have not worked with Prometheus / OpenMetrics before. If you already have, skim it for the vocabulary Metered uses and move on.

What a metric is

A metric is a named, numeric measurement of your running program, sampled over time. “Number of requests handled,” “current queue depth,” and “request latency” are all metrics. A monitoring system scrapes your process periodically, for example every 15 seconds. It reads the current numbers and stores them as a time series it can graph and alert on.

OpenMetrics is the standard text format for that exchange, the successor to the Prometheus exposition format. Metered produces it directly.

Pull, and cumulative

Two ideas underpin everything else:

  • Pull, not push. The monitoring system asks your process for its current numbers; your process does not send them anywhere. So a metric is just state you can read on demand, not an event you emit.
  • Cumulative, not reset. A counter only ever goes up (for the life of the process). You never reset it after a scrape. The monitoring system stores each sample and computes differences itself. This is what lets two replicas be summed correctly, and what lets a scrape that arrives late or twice not corrupt your data.

Keep these in mind. They explain why Metered has no flush or clear operation. They also explain why the query engine computes rates, not your process.

The metric types

OpenMetrics has a small set of types. Metered models each one.

Counter

A value that only increases: requests handled, errors returned, bytes written. On the wire, the encoder exposes a counter http_requests as http_requests_total.

You do not expose a rate. You expose the running total. The query rate(http_requests_total[5m]) turns it into “requests per second.” The query computes it over whatever window the dashboard chooses, and it sums correctly across instances.

Gauge

A gauge goes up and down and represents current state. Examples are queue depth, in-flight requests, connection-pool size, a temperature, and an on/off flag (1 or 0). The monitoring system reads a gauge as-is at scrape time.

The litmus test: if the right thing to graph is the value itself, it is a gauge. If the right thing to graph is how fast it grew, it is a counter.

Histogram

A distribution, for things like latency. A histogram does not store every observation. It counts how many fell into each of a fixed set of buckets with upper bounds (le, “less than or equal”). It also keeps a running sum and count. For a latency metric http_request_duration_seconds you get series like:

http_request_duration_seconds_bucket{le="0.005"} 24
http_request_duration_seconds_bucket{le="0.01"}  41
http_request_duration_seconds_bucket{le="+Inf"}  57
http_request_duration_seconds_sum               0.83
http_request_duration_seconds_count             57

Buckets are cumulative (le="0.01" includes everything le="0.005" counted). The query engine computes percentiles at query time with histogram_quantile(0.95, ...). Because the buckets are plain counters, histograms from many replicas add up, so a fleet-wide p95 is meaningful.

Summary

An older shape that ships pre-computed quantiles (for example quantile="0.95") plus sum/count. The catch: you cannot aggregate pre-computed quantiles – you cannot average two replicas’ p95s to get the fleet p95. Prefer histograms. Metered offers summaries only for backwards compatibility (see Migrating).

Info

Static key/value facts about the process – build version, commit, region – exposed as a constant 1 that carries the facts as labels:

build_info{version="0.10.0",commit="abc123"} 1

You join on it in queries to attach those facts to other metrics.

StateSet

A set of mutually exclusive boolean states – a lifecycle, say – where exactly one is 1 and the rest are 0:

service_lifecycle{service_lifecycle="starting"} 0
service_lifecycle{service_lifecycle="running"}  1
service_lifecycle{service_lifecycle="draining"} 0

Labels

You can split a metric along labels – key/value pairs that create one time series per combination:

http_requests_total{route="/orders",method="POST"} 12
http_requests_total{route="/orders",method="GET"}  87

Labels are powerful and dangerous. Every distinct combination of label values is a separate stored time series. A label with unbounded values, such as a user id, a request id, or a raw URL, creates unbounded series. This “cardinality explosion” can take down your monitoring system. The rule: keep labels bounded. Labels and Families covers this in depth.

Naming and units

Conventions that Metered follows and encourages:

  • Use a namespace_subsystem_name shape: http_request_duration_seconds.
  • Counters carry no rate in the name; the encoding adds the _total suffix. So name the field requests, not requests_total (or you get requests_total_total).
  • Put the unit in the name, and use base units: seconds, not milliseconds, and bytes, not kilobytes. Use _seconds and _bytes. Metered records durations as seconds for this reason.

The mental model to carry forward

A metric is state you expose, read at scrape time, cumulative where it counts. Rates and percentiles are the monitoring system’s job, not yours. The next section shows how Metered turns that model into an API.

Why Metered has this design

Metered makes a handful of strong, opinionated choices. Each one follows from the mental model in the primer: a metric is state you expose, pull-based and cumulative. This page explains the choices, because understanding them makes the API feel obvious.

Architecture at a glance

Three independent concerns meet at the metric state your service owns:

  1. Metric state is readable service state: counters, gauges, histograms, info, state sets, and families.
  2. Composition describes how to borrow that state at scrape time through MetricTree, Registry, and MetricTreeView.
  3. Instrumentation updates state from tracing spans, or directly from the code that owns it.

Core metered owns the first two concerns. It does not decide where service operations begin or how you categorize errors.

flowchart TD
    subgraph svc["Your service (owns all metric state)"]
        code["service code"]
        spans["tracing spans"]
        state["metric state<br/>Counter/Gauge implementors<br/>Histogram / Family / Info / StateSet"]
        code -->|"direct updates:<br/>incr / set / observe"| state
        spans -->|"metered-tracing:<br/>span metrics + exemplars"| state
    end

    subgraph expo["Exposition (at scrape time)"]
        tree["MetricTree<br/>composed by Registry / MetricTreeView"]
        schema["MetricSchema<br/>(the shape)"]
        values["MetricValues<br/>(the samples)"]
        render["MetricSink<br/>e.g. metered-om:<br/>OpenMetricsEncoder / OpenMetricsRender"]
        tree -->|describe| schema
        tree -->|collect| values
        schema --> render
        values --> render
    end

    state -.->|"borrowed at scrape, no Arc"| tree
    render --> text["OpenMetrics text"]
    text --> scraper["Prometheus / VictoriaMetrics"]
    schema -.->|dashboard_queries| promql["PromQL dashboard seeds"]

The rest of this page is why each of those pieces looks the way it does.

Metrics are state, not a shadow system

The central idea: a metric is a piece of your service’s state, owned by the component whose behavior it describes. A queue’s depth metric is the queue’s length. A pool’s “in use” gauge is the pool’s checked-out count.

So Metered has no global registry and no statics. You do not “register a metric with the metrics system” and then find it again by string name. You hold concrete metric state, for example an AtomicU64 counter or a histogram, as a field, exactly where the relevant state lives. You expose it at scrape time. This means:

  • No name-based lookups, no typos resolved at runtime, no init ordering.
  • No accidental sharing: two subsystems cannot clobber each other’s metric by using the same global name.
  • The borrow checker keeps instrumentation honest – a metric cannot outlive the thing it measures.

A consequence you notice in practice: prefer exposing existing state over maintaining a parallel counter. If the queue knows its length, read the queue. Do not increment a separate gauge on every push and pop and hope it never drifts.

No Arc, even across .await

Instrumentation must not force you to wrap your service in Arc, and must work in async code where the body holds &mut self across an .await.

Direct updates satisfy this trivially. incr, set, and observe take &self through interior mutability. They borrow the metric only for the instant of the update, never across the measured body. Span-derived measurement satisfies it structurally. The metered-tracing layer records the duration when the span closes. Nothing borrows your service, or self, while the operation runs.

For exposition, the same principle drives MetricTreeView: it stores selector closures, not metric references, and borrows the live service at scrape time. No Arc on every metric, no shared ownership just to encode text.

Record exactly once – even on panic or cancellation

Span-derived metrics inherit tracing’s guard semantics. The guard closes a span exactly once: on a normal return, a panic, an early return, or an async task cancellation. metered-tracing records on close. So span counters and duration histograms do not leak in-flight state or lose observations when the body exits abnormally.

OpenMetrics-native, no serde on the default path

Metered writes the OpenMetrics text format directly. It does not serialize metrics to a generic data model and then map field names to Prometheus conventions.

Why it matters:

  • Fidelity. Counter _total suffixes, histogram _bucket/_sum/_count, # TYPE/# HELP/# UNIT metadata, exemplars, and stateset semantics are first-class, not approximated by reshaping JSON.
  • Dependency hygiene. The default public API of metered has no foreign types, so a metered upgrade never forces a serde or hdrhistogram bump on your workspace. See Feature flags and stability.

Cumulative histograms over in-process summaries

Older metrics libraries and pre-computed HDR summaries compute quantiles in your process and expose them as a summary. The problem is aggregation: you cannot combine two replicas’ pre-computed p95s into a fleet p95.

Metered’s operation duration paths – metered-tracing span durations and direct observe_duration calls – record into cumulative bucket histograms. The query engine computes percentiles at query time with histogram_quantile, where they aggregate across replicas correctly. The bucket counters are also lock-free, so the hot path never blocks.

A lock-free, allocation-free hot path

Stock metrics back their state with atomics allocated once at construction. The recording path – incr, observe, gauge set – is a relaxed atomic operation with no allocation and no mutex. Per-bucket exemplars are lock-free too. Each bucket has its own swap slot, off the counting path. The slot changes only when your code supplies an exemplar.

The cost guidance follows. A plain counter is the cheap metric for the hottest paths. Duration histograms are the richer metrics you reserve for entry points.

Schema, values, and rendering are separate

Every metric tree can do two independent things: describe its schema – family names, types, units, label names – and collect its current values. The renderer combines a schema and a value set into OpenMetrics text.

This split buys a lot:

  • The schema can drive documentation and dashboard generation without a live scrape (Schema and Dashboards).
  • The encoder can emit values incrementally, with a budget, for very large metric sets (OpenMetrics Exposition).
  • Because a single MetricTree::encode is defined as describe + collect + render, a tree can never advertise one shape in its schema and emit another in its samples. The consistency is structural, not a convention.

A type for the leaf, a trait for the tree

Two traits, with one job each:

  • Metric: a single OpenMetrics family implements it (a Counter implementor, a Gauge implementor, a histogram). It couples the family’s type and its value collection in one place, so they cannot drift.
  • MetricTree: anything composed of families implements it, such as a #[derive(MetricTree)] struct or a Family. Leaves get it for free via a blanket implementation.

You implement Metric for a new leaf. You usually derive MetricTree for a composite. That is the whole extension story.

Evolvable on purpose

Metered expects to reach 1.0 without churning its callers:

  • Open enums like MetricType are #[non_exhaustive], so new OpenMetrics constructs can land without breaking matches.
  • The macros emit ::metered:: absolute paths, so generated code is immune to local name shadowing.
  • The procedural and derive macros are re-exported from metered, so downstream crates depend on metered alone and the two halves always move together.

With the “why” in hand, the Core model introduces the concrete types.

Two tiers: contract metrics vs diagnostics

Not all metrics are the same kind of thing. A team that treats them as one bucket gets broken dashboards and 4 different latency metrics for the same operation. The Metered model distinguishes two tiers, and serves each differently.

Tier 1 – contract or platform “API” metrics

Some metrics are an API: a stable shape that many components expose identically, that dashboards and alerts depend on, and that must not drift. Examples:

  • gRPC server metrics: request count, in-flight, latency histogram, error breakdown – the same families for every service and method, distinguished only by service / method / instance labels.
  • HTTP server metrics, the RED signals, the same for every route.
  • Tokio runtime metrics: worker count, busy ratio, queue depths, poll counts.

These share three properties:

  1. Defined once, reused everywhere. Every gRPC service should expose the same metric shape; you do not want each service inventing its own.
  2. A committed contract. Renaming or reshaping them breaks fleet-wide dashboards and alerts, so they should not change casually.
  3. Populated by infrastructure, not business code. A tower/tonic middleware or a runtime collector fills them in; the service author writes nothing.

The contract is a type

The elegant part: in Metered, the contract is just a typed MetricTree – no separate schema-assertion mechanism needed. An integration crate defines the shape once:

// in metered-tonic (illustrative)
#[derive(Default, MetricTree)]
#[metrics(prefix = "rpc_server")]
pub struct GrpcServerMetrics {
    #[metrics(tree, rename = "requests")]
    requests: Family<MethodKey, std::sync::atomic::AtomicU64>,
    #[metrics(tree)]
    in_flight: Family<MethodKey, std::sync::atomic::AtomicI64>,
    #[metrics(tree)]
    duration_seconds: Family<MethodKey, BucketHistogram>,
    #[metrics(flatten)]
    errors: Family<MethodErrorKey, std::sync::atomic::AtomicU64>,
}

The middleware records into a GrpcServerMetrics. The middleware can only record into that type, so the compiler enforces that every service emits exactly the contract shape. There is nothing to keep in sync, no runtime schema check, and no way to drift. The reusable struct is the contract, and describe() is its machine-readable schema for docs and dashboards.

This is why name shaping (#[metrics(rename/flatten)]) matters here: the wire contract stays fixed even as the integration crate refactors its internals.

Tier 2 – diagnostic, implementation-detail metrics

Other metrics are implementation details: service-specific counters and timers you add to understand this code. Examples are a retry count, a cache hit rate, and time spent in a specific phase. You expose them for troubleshooting, and they evolve with the code. If one disappears in a refactor, no fleet dashboard breaks.

These are exactly what owner-local primitives, Family, MetricTree, and MetricTreeView are for: cheap to put next to the behavior, owner-local, no central contract. For diagnostics on code that already carries tracing spans, metered-tracing derives the metrics from those spans, on top of the same primitives. Use name shaping when you want a particular diagnostic to stay stable, but the default expectation is that they track the code.

Choosing the tier

Contract, Tier 1Diagnostic, Tier 2
Who defines the shapean integration crate, oncethe service author, as needed
Who populates itmiddleware / collectorservice code via primitives / families / views; span-derived via metered-tracing
Stabilitycommitted; do not drifttracks the code
Breaking itbreaks fleet dashboardslow stakes, troubleshooting only
Cardinalitybounded by label contractbounded by author discipline

A healthy service exposes both: the platform contract metrics, so it shows up on the standard fleet dashboards for free, plus its own diagnostics. Metered serves Tier 1 through integration crates that provide the reusable typed tree and the collector. Examples are the runtime telemetry crates here – Tokio, process, and system – or a framework’s own RPC/HTTP contract trees built the same way. Metered serves Tier 2 through owner-local primitives, families, metric trees, and views. Both tiers encode into the same OpenMetrics document.

Core model

This page names the concrete types and how they fit together. It is the map. Later pages are the territory.

Three layers

  1. Metric state: counters, gauges, histograms, info, families.
  2. Composition: MetricTree, Registry, and MetricTreeView.
  3. Instrumentation: metrics derived from the tracing spans your code already emits (metered-tracing), or your own code updating the state directly.

Core metered is layers 1 and 2. It does not decide where service operations begin or how you categorize errors.

The layers are deliberately independent. You can expose state that Metered never “recorded,” for example an existing atomic or a queue length. Instrumentation can update metric state without owning exposition.

instrumentation               metric state              composition / exposition
---------------               ------------              -----------------------
tracing spans      ------->   counter/gauge values      Metric  (one family)
direct updates                Histogram / Family ---->  MetricTree (families)
(your code)                   Info / StateSet           Registry / MetricTreeView
                                                        MetricSchema + MetricValues
                                                        -> OpenMetrics text

Instrumentation: spans and direct updates

If code already runs inside tracing spans, metered-tracing exports span counters and duration histograms as metric state. Instrument a method once, and its performance shows up in both traces and metrics. There is no second call site to maintain. The span guard closes the span exactly once: on a normal return, a panic, or an async cancellation. The derived metrics never leak an observation.

Everywhere else, your code is the instrumentation, and it updates owned state. Increment a counter when the event happens. Set a gauge when the value changes. Time an operation with BucketHistogram::observe_duration. There is no measuring wrapper layer between your code and the metric.

Leaves: Metric

A [Metric] is one OpenMetrics family. It states its metric_type() and knows how to collect_metric(...) its current samples. Coupling the two in one trait means a metric’s declared type and its emitted values cannot disagree.

The stock leaf categories:

APIOpenMetrics typeRole
Counter implementors such as AtomicU64countera monotonic count you own
Gauge implementors such as AtomicI64 / AtomicU64gaugea current value you own
BucketHistogramhistograma distribution with classic le buckets
Infoinfostatic key/value facts
StateSetstatesetone-of-N lifecycle state

Standard-library atomics implement the relevant metric traits or Metric, so you can expose existing state directly.

Trees: MetricTree

A [MetricTree] is anything made of families. It can describe its schema and collect its values. The default text rendering combines the two. Leaves are trees automatically, through a blanket impl. You get a MetricTree from:

  • #[derive(MetricTree)] – a struct of metrics;
  • Family<L, M> – one metric per label set;
  • a hand-written impl for a custom composite.

How the traits and types relate:

flowchart TD
    counter["Counter/Gauge implementors<br/>Histogram / Info / StateSet / atomics"] -->|impl| metric["trait Metric<br/>(one family)"]
    metric -->|"blanket impl"| tree["trait MetricTree<br/>(a tree of families)"]
    der["derive(MetricTree) struct"] -->|impl| tree
    fam["Family&lt;L, M&gt;"] -->|impl| tree
    tree -->|"composed by"| registry["Registry (borrowed)<br/>MetricTreeView&lt;C&gt; (closures)"]

A leaf implements Metric. Everything else implements MetricTree directly. A Registry or MetricTreeView composes trees under a prefix and constant labels.

The upkeep path: housekeep

MetricTree carries a third pair of methods: needs_housekeep and housekeep. This is the seam for maintenance that must not run on the recording path: work that takes a lock, allocates, or rebuilds internal state. The DynamicExponentialHistogram downscale and its straggler drain live here. So does an interval-histogram swap behind with_housekeep.

The scrape pipeline drives it, not your code. MetricTree::encode runs housekeep first when needs_housekeep reports pending work. A Registry scrape does the same by default. metered_om::SnapshotCache drives it on its own refresh cycle.

The result is one simple contract. The observe path of every shipped instrument stays lock-free. Everything that is not lock-free waits for the upkeep pass, which runs once per scrape on the scraping task.

Hand-written MetricTree implementations that contain histograms must forward housekeep and needs_housekeep to their fields. A tree that forgets freezes its dynamic histograms at their saturation point. The derive forwards automatically.

Composing for a scrape: Registry and MetricTreeView

Both turn a set of trees into one OpenMetrics document under a shared prefix and constant labels. They differ in ownership:

  • Registry is a borrowed view: you hand it & references for the duration of one encode. Good for one-off snapshots and for adapter metrics.
  • MetricTreeView<C> stores selector closures over an app context C and borrows the live metrics at scrape time. Build it once, reuse it every scrape, with no shared ownership. This is the recommended service pattern.

Schema versus values

A MetricSchema is the static contract – family names, types, HELP, UNIT, label names. MetricValues is the sampled state at one instant. A MetricSink turns that pair into a wire format: the metered-om crate’s OpenMetricsEncoder renders text, and OpenMetricsRender does the same incrementally. The core crate has no encoder of its own, so a service picks or writes the sink it needs. Because the schema is independent of any scrape, it also feeds documentation and dashboards.

Choosing what to hold

A quick guide, expanded in Choosing Metric Types:

NeedReach for
Count eventsa Counter implementor such as AtomicU64
Current value you owna Gauge implementor such as AtomicI64 / AtomicU64
Current value something else ownsexpose that state directly
A distribution / latencyBucketHistogram or DynamicExponentialHistogram
One-of-N stateStateSet
Static factsInfo
A bounded extra dimensionFamily<L, M>

Counters and histograms are cumulative. Gauges and state sets are source-of-truth values. There is deliberately no generic “reset” – it would be meaningless for some of these and wrong for the rest.

Choosing metric types

Picking the right type is most of what makes metrics good. This page is a decision guide plus a reference for each type Metered offers.

The decision in one paragraph

For an event you count, use a counter. For a current value, use a gauge, but if another component already owns that value, expose that value instead of a copy. For a distribution, latency in particular, use a histogram. For one-of-N state, use a StateSet. For a static fact, use info. For an extra bounded dimension, add a Family.

The same decision as a flowchart:

flowchart TD
    start{"What are you<br/>measuring?"}
    start -->|"an event happened"| counter["Counter"]
    start -->|"a current value"| owned{"who owns<br/>the value?"}
    owned -->|"this metric"| gauge["Gauge"]
    owned -->|"something else"| expose["expose that state directly<br/>(view reader / adapter)"]
    start -->|"a distribution (latency)"| hist["Histogram"]
    start -->|"one-of-N state"| stateset["StateSet"]
    start -->|"a static fact"| info["Info"]
    counter --> dim{"need a bounded<br/>extra dimension?"}
    gauge --> dim
    hist --> dim
    dim -->|"yes"| family["wrap in Family⟨L, M⟩"]
    dim -->|"no"| done["done"]

All metric types at a glance

OpenMetrics typeInstrumentNotes
counterCounter / AtomicU64Monotonic; suffix _total.
gaugeGauge view or owned Gauge/AtomicI64 fieldLive pull or cached accumulator. See Gauge below.
histogramBucketHistogram, DynamicExponentialHistogramAggregates across replicas; prefer for latencies.
summarySummary<S> over a QuantileSourcePer-instance quantile values; does not aggregate.
gaugehistogramGaugeHistogram<S> over a GaugeHistogramSource, for example GaugeBucketsCurrent-value buckets; _gcount/_gsum.
statesetStateSetMutually exclusive boolean states.
infoInfoMetricStatic metadata; value always 1.
unknownpassthrough samplesForeign metrics of unknown semantics.

Counter

For values that only increase: requests, errors, retries, bytes, dropped messages. Counter is the trait. AtomicU64 is the usual concrete storage.

#![allow(unused)]
fn main() {
use metered::Counter;
use std::sync::atomic::AtomicU64;

let processed = AtomicU64::new(0);
processed.incr();
processed.incr_by(10);
assert_eq!(processed.get(), 11);
}

Make intent obvious through the field name – attempts, failures, cache_misses – or through a domain newtype over the counter. The storage is the same AtomicU64 either way.

Do not name a counter field *_total: the encoding adds _total, so requests becomes requests_total on the wire (a requests_total field becomes requests_total_total).

Gauge

For a current value that moves up and down: in-flight work, queue depth, pool size, a flag. Gauge is the trait. Standard atomics are the usual concrete storage.

#![allow(unused)]
fn main() {
use std::sync::atomic::{AtomicI64, Ordering};

let in_flight = AtomicI64::new(0);
in_flight.fetch_add(1, Ordering::Relaxed); // a request started
in_flight.fetch_sub(1, Ordering::Relaxed); // it finished
let flag = AtomicI64::new(0);
flag.store(1, Ordering::Relaxed);          // 1 / 0
}

Your code updates plain atomics. The exposition side reads them through the Gauge trait, which the standard atomics implement. The trait also has incr/decr/set helpers if you prefer them, but nothing requires a metrics call on the update path.

Gauge vs counter is the most common mistake. “Total requests,” “failed publishes,” “retries,” and “dropped messages” are counters, not gauges – you want their rate, not their instantaneous value. “In-flight requests,” “queue depth,” “cache entries,” and “enabled?” are gauges.

Prefer existing state. If another component already owns a value, expose that value. Do not maintain a parallel gauge that can drift. See Adding Metrics and the queue module in the demo.

Live vs cached. Both forms encode as a plain gauge because OpenMetrics has no UpDownCounter type. A live gauge is a view reader like gauge_value("queue_depth").read(|state: &State| state.queue.len() as i64). Each scrape reads it fresh from your domain state, with no cache, so it is always truthful, but the read must be cheap.

A cached accumulator is an owned Gauge/AtomicI64 you bump with incr/decr. It reads O(1) at scrape. Use it when the value is cheap to maintain incrementally but expensive or impossible to observe live. There is no per-metric value cache: metered-om’s SnapshotCache bounds live-read cost at the document level.

Histogram

For distributions – almost always latency. BucketHistogram records raw values with observe. It records durations with observe_duration, which converts them to seconds, the base unit. Histogram is the trait that abstracts over the bucket and exponential backends. BucketHistogram is the classic le-bucket implementation.

#![allow(unused)]
fn main() {
use metered::{BucketHistogram, Buckets};

// Choose buckets that bracket your expected range.
let sizes = BucketHistogram::new(Buckets::exponential(64.0, 2.0, 10));
sizes.observe(512.0);
}

Bucket presets, all in seconds for durations:

  • Buckets::seconds_default() – general request latencies, 5 ms to 10 s.
  • Buckets::fast_seconds() – services below 5 ms, down to 25 µs.
  • Buckets::slow_seconds() – DB-heavy / batch work, out to a minute.
  • Buckets::wide_seconds() – fine near a microsecond, coarse near multi-second timeouts: 1 µs to about 30 s, at most 30% relative error. For fast paths whose latency spans a very wide range.
  • Buckets::relative(min, max, max_relative_error) – exponential buckets sized to a target relative resolution. You give the range and the error you tolerate, and the builder solves for the count.
  • Buckets::exponential(start, factor, n) / exponential_range(min, max, n) – the explicit-count exponential builders.

Wide dynamic range: fine low, coarse high

When you care about microsecond-scale fast paths but only need rough numbers near timeouts, use relative resolution. A bucket near value v is about v * max_relative_error wide. The absolute resolution is then automatically fine at the bottom and coarse at the top.

#![allow(unused)]
fn main() {
use metered::{BucketHistogram, Buckets};

// <=10% relative from 1µs to 30s: ~0.1µs near the floor, ~1s near a 10s timeout.
let h = BucketHistogram::new(Buckets::relative(0.000_001, 30.0, 0.10));
}

Tighter error or a wider range means more buckets. Each boundary is a stored le series per label set, so this is a dial between resolution and cardinality. About 10% error over 1 µs to 30 s is about 180 buckets, and about 30%, from wide_seconds, is about 70. There is no cheap way to get fine absolute resolution across the whole range. Linear 10 µs buckets up to 10 s would be a million series. That is exactly why exponential relative-error histogram designs exist.

The quantile values come from histogram_quantile(...) at query time, so they aggregate across replicas. That is the whole reason to prefer a histogram over a summary.

Exponential histograms

For wide dynamic ranges, Metered also provides log-linear exponential histogram backends: FixedExponentialHistogram and DynamicExponentialHistogram. They record sparse exponential bucket snapshots as ExponentialSnapshot values. metered-om can encode them either as classic cumulative le buckets or as VictoriaMetrics vmrange buckets. The full comparison – who chooses the buckets, memory, observe cost, saturation behavior – is in Histograms in Depth.

This is separate from Prometheus native histogram protobuf exposition. If a protobuf sink is worth the effort later, it can encode from the same MetricValues model. Metric ownership and observation code do not change. A BucketHistogram can also attach exemplars to buckets via observe_with_exemplar.

Summary

A Summary<S> renders each pre-computed quantile, for example name{quantile="0.99"}, plus name_sum and name_count. It reads them from any QuantileSource. Its quantile values are per-instance and do not aggregate across replicas. Prefer a histogram unless you specifically need exact per-instance quantile values or legacy dashboard parity.

#![allow(unused)]
fn main() {
use metered::summary::BucketQuantiles;
use metered::{BucketHistogram, Buckets, Summary};

let latency = BucketHistogram::new(Buckets::seconds_default());
latency.observe(0.012);
// Any `Histogram` is a `QuantileSource` via `BucketQuantiles`.
let summary = Summary::new(BucketQuantiles::new(&latency));
}

BucketQuantiles adapts any Histogram. The DynamicExponentialHistogram is the recommended source: lock-free observe, bounded memory, bounded relative error. A bespoke streaming sketch, for example CKMS, is intentionally not provided. It cannot be lock-free, and it does not beat the exponential histogram on memory. It only offers a different accuracy model, based on rank error. If a project ever needs one, it drops in as a QuantileSource implementation behind the same seam.

Gauge histogram

A GaugeHistogram<S> renders a current value distribution, buckets that can decrease, as name_bucket{le} + name_gcount + name_gsum. Use it for a live population’s size distribution, not a cumulative count of events. Implement GaugeHistogramSource for your own state, or use the provided GaugeBuckets:

#![allow(unused)]
fn main() {
use metered::{GaugeBuckets, GaugeHistogram};

let sizes = GaugeBuckets::new([1.0, 10.0, 100.0]);
sizes.enter(5.0); // an item of size 5 is now held
sizes.leave(5.0); // it was released
let gh = GaugeHistogram::new(sizes);
}

StateSet

For mutually exclusive state – a lifecycle, a mode – where exactly one member is active:

#![allow(unused)]
fn main() {
use metered::StateSet;

let lifecycle = StateSet::new(["starting", "running", "draining"]);
lifecycle.set("running");
}

It emits one series per state, the active one 1 and the rest 0, with the state carried in a label named after the metric.

Info

For static facts about the process – version, commit, region – as a constant 1 carrying labels:

#![allow(unused)]
fn main() {
use metered::InfoMetric;

let build = InfoMetric::new([("version", "0.10.0"), ("commit", "abc123")]);
}

Use it to attach build context to dashboards by joining on *_info.

Family: a bounded extra dimension

When one metric needs a label dimension – per route, per method, per upstream – wrap it in a Family. See Labels and Families for the cardinality discipline that keeps this safe.

#![allow(unused)]
fn main() {
use metered::{Counter, Family};
use std::sync::atomic::AtomicU64;

let by_route: Family<Vec<(String, String)>, AtomicU64> =
    Family::with_label_names(["route"]);
by_route.with(&vec![("route".to_owned(), "/health".to_owned())], |c| c.incr());
}

Unknown: passthrough

MetricType::Unknown exists for passthrough or foreign metrics with unknown semantics. It renders a plain sample with no suffix. You don’t construct it directly. The passthrough sources that re-emit metrics from another system use it.

Serving a foreign metrics endpoint, including Metered 0.9

metered_om::TextSourceTree re-emits any classic Prometheus/OpenMetrics text as a MetricTree. This lets you serve a foreign producer on a 0.10 /metrics endpoint during migration. The producer can be a sidecar, another exporter, or a Metered 0.9 registry’s serde_prometheus output. Nothing about it is 0.9-specific, and Metered 0.9 is just one such producer. To run 0.9 in-process, depend on the published 0.9 crate (see Migrating From Older Versions).

use metered_om::TextSourceTree;
let legacy = TextSourceTree::new(|| old_0_9_registry.to_prometheus_text());
// mount `legacy` alongside your native 0.10 trees

The passthrough is a normalizing re-encode, not a byte copy. Names, labels, and shapes survive, so a 0.9 HDR summary stays a summary and existing dashboards keep working. Value tokens re-encode from their parsed form, and the re-encode drops sample timestamps. After you migrate a metric to a native 0.10 instrument, drop it from the passthrough source.

When none of these fit

Implement Metric for a custom leaf: its type plus how it reads its value. Implement MetricTree for a custom composite. This is rare. Reach for it only when you genuinely have a new OpenMetrics shape or an unusual source for the value.

Histograms in depth

Metered ships three recording engines for cumulative distributions: BucketHistogram, FixedExponentialHistogram, and DynamicExponentialHistogram. They share the wire model, the exemplar machinery, and the query story. They differ in who chooses the buckets, what memory costs, and what happens when your traffic surprises you. This section gives the full comparison. For the one-paragraph version, read Choosing Metric Types.

The classic bucket histogram

BucketHistogram records into boundaries you choose (le buckets), plus a running _sum and _count.

Benefits. The boundaries carry meaning. Put a bucket edge exactly at your latency objective, and the dashboard answers “how many requests beat the objective” with no interpolation error. Memory cost and encode cost stay fixed and small. Every scraper understands the output.

Costs. You must know the range before the first deploy. Boundaries that miss the real distribution give useless resolution, and changing them later breaks dashboard continuity. Each bucket is one series on the wire, so resolution multiplies cardinality.

Use it when the boundaries are part of the contract: objective edges, fixed size classes, or parity with an existing dashboard.

Exponential histograms

The exponential engines remove the boundary decision. Bucket i covers [base^i, base^(i+1)) where base = 2^(2^-schema). One integer, the schema, sets the relative resolution everywhere at once:

schemaRelative error per bucket
0one power of two, about 100%
3~9%
5~2.2%
8~0.27%

The error is relative, so one setting serves microseconds and minutes in the same histogram. There is no boundary list to choose, tune, or migrate. Non-positive values land in a dedicated zero bucket. The valid schema range is 0 to 20.

FixedExponentialHistogram: dense, bounded range

FixedExponentialHistogram::try_new(min, max, schema) allocates every bucket between min and max up front.

Benefits. The observe path is fully lock-free with no branches for table management. Memory is exact and known at construction. Construction fails closed: an impossible range or schema is an error, not a surprise later.

Costs. You are back to declaring a range. Values outside it clamp to the edge buckets. A wide range at a fine schema allocates many slots whether you hit them or not.

Use it for one hot histogram whose range you genuinely know.

DynamicExponentialHistogram: sparse, self-scaling

DynamicExponentialHistogram::new() starts at schema 5, about 2.2% resolution, with a 256-slot table. with_params(start_schema, capacity) tunes both. This is the engine the heavy-duty deployments run as their default, and the one the rest of the stack assumes.

Benefits. Buckets live in a fixed-capacity lock-free table, and memory tracks the buckets your values populate, not the range you configured. A fleet of mostly idle histograms stays cheap. The observe path is one log2 and one atomic increment, with a one-time compare-and-swap when a value claims a new bucket. Per-bucket exemplars are lock-free too.

Costs. The capacity is a budget. When the table saturates, the histogram downscales: it merges adjacent buckets, which halves the resolution. The counts already recorded coarsen with it. The schema only ever decreases.

A value range far wider than the capacity affords is not an error. The histogram converges to a coarse but correct summary of the range.

What saturation does not cost. The rebuild never runs on the observe path. An observation that finds the table full sets a flag. The merge happens on the upkeep path, in housekeep (see the upkeep path).

Be precise about the lock-freedom claim: it covers the observe path only. The upkeep pass is not lock-free. It allocates the replacement table and takes a mutex over the retired-table list. That is fine, because it runs on the scraping task, once per scrape, never on a recording thread. A compare-and-swap gate keeps concurrent scrapers honest: one performs the rebuild, the rest skip past it.

The swap uses read-copy-update: observers keep using the old table lock-free until the histogram publishes the coarser one. The drain then folds a straggler’s late increment into the live table exactly once. No observation is ever dropped or double-counted.

Trade-offs at a glance

BucketHistogramFixedExponentialHistogramDynamicExponentialHistogram
You chooseevery boundaryrange + schemaschema + slot budget
Memoryfixed, per boundaryfixed, whole rangetracks populated buckets
Observe costlock-free branch scanfully lock-free indexlock-free log2 + add
Surprise rangeclamps into edge bucketsclamps into edge bucketsdownscales, keeps counting
Resolution over timeconstantconstantcan coarsen, never below schema 0
Failure modewrong boundaries foreverconstruction errorcoarser buckets

Memory for common ranges

Per-bucket costs, from the struct layouts:

EngineBytes per bucketWhat a bucket holds
BucketHistogram~24 Bboundary f64 + count + exemplar slot
FixedExponentialHistogram8 Bcount only; this engine has no per-bucket exemplar slots
DynamicExponentialHistogram32 B per table slotindex + count + exemplar slot + window state

An exponential engine needs log2(max / min) x 2^schema buckets to span a range. For ranges that services actually meter:

RangeSpreadBuckets at schema 3 / 5 / 8
Cache operation: 1 µs to 10 ms10^4107 / 426 / 3,402
RPC latency: 100 µs to 10 s10^5133 / 532 / 4,252
Payload size: 64 B to 16 MiB2^18144 / 576 / 4,608

What that costs per engine, on the RPC latency range:

  • Classic, 14 hand-picked boundaries: ~360 B, and all 15 bucket series are on the wire at every scrape, populated or not.
  • Fixed at schema 5: 532 buckets x 8 B = ~4.3 KiB, allocated up front whether traffic hits them or not. At schema 8 that becomes ~34 KiB. Only populated buckets reach the wire.
  • Dynamic at the defaults: an 8 KiB table of 256 slots x 32 B, and that is the ceiling for any range. Only populated slots reach the wire.

The dynamic budget rule: capacity / 2^schema is how many powers of two can populate before a downscale. The defaults give 256 / 32 = 8 powers of two. That is a x256 spread at the full ~2.2% resolution. A healthy latency distribution concentrates well inside that. A uniform flood across the full x10^5 RPC range would settle at schema 3, ~133 populated buckets. That still resolves ~9% per bucket, from the same 8 KiB.

The fleet math is where the engines separate. One thousand mostly idle per-target histograms cost a fixed engine the full range each: ~4.3 MiB at schema 5 on the RPC range. Dynamic tables only fill as targets actually observe. The wire carries only what filled.

Choosing

  1. Boundaries are part of a contract, or a dashboard depends on exact edges: BucketHistogram.
  2. One histogram, hot path, known range, and you want zero table management: FixedExponentialHistogram.
  3. Everything else – and especially many instances, unknown ranges, or per-target families: DynamicExponentialHistogram. This is the default worth reaching for first.

Rendering and querying

Both exponential engines snapshot into the same ExponentialSnapshot, and metered-om encodes a histogram family in either of two forms:

  • Classic le buckets. Compatible with every Prometheus-style scraper and histogram_quantile().
  • VictoriaMetrics vmrange series. The native form for VictoriaMetrics; its query functions consume the ranges directly.

The family declares its encode intent and the sink resolves it against its own capability, so the same metric definition serves both fleets. Bucket exemplars attach identically in either encoding. The quantile values always come from the query engine, at read time, so they aggregate correctly across replicas – the whole reason to prefer histograms over summaries.

Adding metrics

This guide is for agents and humans who add observability to Rust services with Metered. Follow it before introducing a new metric.

The rule

Metrics belong to the object that owns the observed behavior or state. Avoid detached “metrics bags” that measure unrelated components. Prefer a small metric tree or view next to the service state that already owns the behavior.

Choose the shape

Start with the state the service owns:

  • A concrete Counter implementor, such as AtomicU64, for monotonic counts.
  • A concrete Gauge implementor, such as AtomicI64 or AtomicU64, for current values the service owns.
  • Histograms for distributions. Duration metric names should include units, usually _seconds.
  • Info for static build/version facts.
  • StateSet for enum-like lifecycle state.
  • Family<L, M> for bounded label dimensions.
  • A family view – family_view, or its one-string-key sugar family_by – when your state is already a keyed map of components. Expose the map you own instead of mirroring it into an owned Family.
  • Existing state directly when the value already exists, such as AtomicBool, an atomic depth cache, or a queue length read while holding the queue lock.

Then choose operation instrumentation only if you are measuring an operation boundary:

  • For an operation that already runs in a tracing span – an instrumented method, or a span-opening middleware – derive its metrics with metered-tracing: one instrumentation, no double bookkeeping.
  • For an operation that is not span-shaped, use plain core state on the owning component: a counter for attempts, a counter for failures, and a histogram observed with observe_duration.

Do not mirror mutable service state into a separate metric just to expose it. If a queue already owns its length, expose a QueueDepth metric that reads the queue length. If that lock is too expensive for scrapes, update a cached atomic depth during queue mutations and expose that cache through a Metric newtype.

Gauge guidance

Gauges are for current state, not for event bookkeeping.

Good gauge examples:

  • in-flight requests
  • current queue depth
  • on/off flag
  • active workers
  • cache entries

Bad gauge examples:

  • total requests handled
  • failed publishes
  • retries
  • dropped messages

Those are counters.

Labels

Labels must stay bounded and operationally useful: operation, result, route, upstream, mode. Never label with raw user/account/request ids, URLs, payloads, raw errors, or any open-ended input – that explodes cardinality. If the value set is open-ended, you need a different metric, a normalized category, or no label. See Labels and Families for the full treatment and for how to add a label dimension with Family.

Preferred service pattern

Expose a borrowed metric view over the service. The view stores no metrics and does not require shared ownership.

#![allow(unused)]
fn main() {
use std::sync::atomic::AtomicU64;
use metered::entry::{counter, gauge};
use metered::{MetricTreeView, Unit};

struct Service {
    processed: AtomicU64,
    queue_depth: AtomicU64,
}

let mut view = MetricTreeView::with_prefix("service");
view.register(
    counter("processed")
        .select(|service: &Service| &service.processed)
        .help("Processed jobs")
        .unit(Unit::Items),
);
view.register(
    gauge("queue_depth")
        .select(|service: &Service| &service.queue_depth)
        .help("Jobs waiting in the queue")
        .unit(Unit::Items),
);
}

For a stable struct of borrowed metrics, derive MetricTree:

#![allow(unused)]
fn main() {
use std::sync::atomic::{AtomicI64, AtomicU64, AtomicUsize};
use metered::MetricTree;

#[derive(MetricTree)]
struct ServiceMetrics<'a> {
    #[metric(gauge)]
    enabled: &'a AtomicI64,
    #[metric(gauge)]
    queue_depth: &'a AtomicUsize,
    #[metric(counter)]
    processed: &'a AtomicU64,
}
}

Use this when you want a named metric tree type that you can register or nest inside another tree. To keep wire names stable while you refactor such a struct, see Shaping Names.

Custom metrics

For one OpenMetrics family, implement Metric. This keeps the type and values in one place. A queue depth encoded as a gauge is a gauge. The implementation just decides where the current value comes from:

#![allow(unused)]
fn main() {
use metered::{Metric, MetricType, MetricValues};

struct QueueDepth<'a>(&'a std::sync::atomic::AtomicUsize);

impl Metric for QueueDepth<'_> {
    fn metric_type(&self) -> MetricType {
        MetricType::Gauge
    }

    fn collect_metric(&self, name: &str, labels: &[(&str, &str)], values: &mut MetricValues) {
        values.gauge(name, labels, self.0.load(std::sync::atomic::Ordering::Relaxed));
    }
}
}

For a tree that emits multiple families, implement or derive MetricTree.

To keep emitted names stable as you rename fields/methods or reorganize structs, use Shaping Names (#[metric(rename/flatten)]).

Checklist

Before you finish:

  • The metric owner is the service/object that owns the behavior or state.
  • Labels stay bounded and useful on dashboards.
  • Durations include units in the metric name.
  • Gauges represent current state.
  • Counters represent cumulative events.
  • You expose existing state directly or through a deliberate cached value.
  • The domain operation remains readable.
  • Tests cover emitted names, labels, classification, and schema when relevant.
  • You update dashboards or docs when names or labels change.

Labels and families

Labels turn one metric into many series – one per combination of label values. They are the most useful and the most dangerous feature in metrics. This page covers how to use them safely and the two ways Metered attaches them.

Two kinds of labels

  • Constant labels are the same on every series from a registry: service, instance, region, env. You set them once on the Registry / MetricTreeView, and every series carries them.
  • Dimensional labels vary per observation: route, method, result, upstream. Each distinct value is a separate time series. These are the ones that need discipline.
#![allow(unused)]
fn main() {
use metered::Registry;

let mut registry = Registry::with_prefix("orders");
registry.label("service", "orders");   // constant: on every series
registry.label("instance", "i-1");
}

Cardinality: the one rule

Every distinct combination of dimensional label values is a separate stored time series in your monitoring system. The cost is multiplicative: 5 routes × 4 methods × 3 results = 60 series for one metric. That is fine. But a label with unbounded values creates unbounded series and can overwhelm the backend. Examples are a user id, an order id, a request id, a raw URL with query strings, and a raw error message. This is “cardinality explosion.”

The rule: dimensional labels must stay bounded, with values you mostly know ahead of time.

Good dimensional labels: route from a fixed set of endpoints, method, result with values ok / error, upstream, mode, and error kind as an enum.

Bad: anything per-user, per-request, per-entity, or free-form. If you want one of those, you usually want one of three things. Use a different metric. Use a normalized category: status class 5xx instead of the exact code, or route template /orders/{id} instead of the concrete path. Or use an exemplar, which carries a trace id without a new series.

Adding a dimension with Family

A Family<L, M> keeps one metric M per label set L, creating them on first use – the analogue of a Prometheus metric family.

Dynamic label sets

For ad-hoc string labels, declare the label names up front so the schema does not depend on which values traffic happens to produce:

#![allow(unused)]
fn main() {
use metered::{Counter, Family};
use std::sync::atomic::AtomicU64;

let by_route: Family<Vec<(String, String)>, AtomicU64> =
    Family::with_label_names(["route"]);

by_route.with(&vec![("route".to_owned(), "/health".to_owned())], |c| c.incr());
}

Typed label keys, preferred

For a stable dimension, derive LabelSet on a key struct. Each field becomes a label. The type makes the dimension explicit and prevents typos:

#![allow(unused)]
fn main() {
use metered::{Counter, Family, LabelSet};
use std::sync::atomic::AtomicU64;

#[derive(Clone, PartialEq, Eq, Hash, LabelSet)]
struct RouteKey {
    route: String,
    method: String,
}

let requests: Family<RouteKey, AtomicU64> = Family::default();
requests.with(
    &RouteKey { route: "/orders".into(), method: "POST".into() },
    |c| c.incr(),
);
}

Family::default() reads the label names from the derived LabelSet, so its schema is correct before any traffic arrives. You can drop a stale series, such as a closed connection or a removed route, with Family::remove.

Keyed state as families

A family is a keyed subtree: one label set selects one member’s metrics. Metered gives that subtree two ownership modes. In the owned mode, a Family<L, M> stores the members inside the family. In the borrowed mode, a family view iterates members that your own state stores. The two modes give the same wire output for the same logical data.

Owned: whole metric structs per key

The member type M of a Family<L, M> is any MetricTree, not just a single metric. A derived metric struct works as-is, so one key can own a whole bundle of metrics:

#![allow(unused)]
fn main() {
use metered::{Counter, Family, Gauge, LabelSet, MetricTree};
use std::sync::atomic::{AtomicI64, AtomicU64};

#[derive(Clone, PartialEq, Eq, Hash, LabelSet)]
struct RailLabels {
    rail: String,
}

#[derive(Default, MetricTree)]
struct RailMetrics {
    #[metric(counter)]
    sent: AtomicU64,
    #[metric(gauge)]
    queue_depth: AtomicI64,
}

let rails: Family<RailLabels, RailMetrics> = Family::default();
rails.with(&RailLabels { rail: "sepa".to_owned() }, |m| {
    Counter::incr(&m.sent);
    Gauge::set(&m.queue_depth, 3);
});
}

The owned mode carries the machinery with it. A MetricConstructor builds each member on first use. Metered sorts the series by label pairs. When key values come from external input, bound them with BoundedValues: it interns values up to a cap and maps the rest to one overflow value.

Borrowed: family views over your own state

Often the keyed state already exists in your service: a map of rails, remotes, or shards. The metrics live inside the members. Do not mirror that map into an owned Family. Expose it borrowed instead, through MetricTreeView::family_view, which keys members by a typed LabelSet:

#![allow(unused)]
fn main() {
use metered::entry::counter;
use metered::{LabelSet, MetricTreeView, MetricsView};
use std::collections::HashMap;
use std::sync::RwLock;
use std::sync::atomic::AtomicU64;

#[derive(Clone, PartialEq, Eq, Hash, LabelSet)]
struct RailLabels {
    rail: String,
    direction: String,
}

struct Rail {
    sent: AtomicU64,
}

impl MetricsView for Rail {
    fn metrics_view() -> MetricTreeView<'static, Self> {
        let mut view = MetricTreeView::new();
        view.register(
            counter("sent")
                .select(|rail: &Rail| &rail.sent)
                .help("Payments sent on this rail"),
        );
        view
    }
}

struct Rails {
    map: RwLock<HashMap<RailLabels, Rail>>,
}

let mut view = MetricTreeView::with_prefix("rails");
view.family_view(Rail::metrics_view(), |rails: &Rails, out| {
    for (key, rail) in rails.map.read().unwrap().iter() {
        out.emit(key, rail);
    }
});
}

There is one borrowed-family primitive, family_view, and one layer of sugar. MetricTreeView::family_by is family_view for the common one-string-key case. It takes a label name and a closure that emits (key, member) pairs. Internally, both forms feed the same per-member emission seam. The string form stamps its one (label, key) pair borrowed from the caller’s &str, without a copy. And Family is the same contract with Metered-owned storage: one keyed group of members rendered as one labeled family, with the storage inverted.

The element shape comes from a context-free element view. Metered declares the schema once, from Rail::metrics_view() alone. The schema declares the key’s label names without values. An empty group still advertises its families. Membership churn appears automatically at the next scrape: your map’s inserts and removes are the lifecycle. For family_view, the key type must declare its label names statically – #[derive(LabelSet)] keys do, Vec<(String, String)> does not.

Two disciplines transfer to you in the borrowed mode. Cardinality control belongs to whatever admits entries into your map. The iterate closure runs on the scrape path, so keep its lock scope small. Emission order is caller-driven: the document lists members in the order you emit them. Sort in iterate if you want the same sorted order as a Family.

Upkeep also forwards through the group. housekeep reaches every emitted member, so dynamic histograms inside members keep rescaling.

Which mode to use

QuestionOwned Family<L, M>Borrowed family view
Storage ownerThe family stores the members in its own map.Your map or state stores the members.
Key typeA typed LabelSet, or Vec<(String, String)> with declared label names.A typed LabelSet that declares its names with family_view, or one string label with family_by.
Cardinality controlIntern external-input keys with BoundedValues.Whatever admits entries into your map.
Membership lifecycleFamily::with creates a member on first use; Family::remove drops one.Your map’s inserts and removes; the next scrape reflects them.
Lock disciplinewith read-locks the family while your closure runs; keep it short.iterate runs on the scrape path; keep its lock scope small.
OrderingSorted by label pairs.Caller-driven; sort in iterate for parity.

One story on the wire

For the same logical data, an owned Family<L, M> and a borrowed family_view write the same OpenMetrics document, byte for byte. The documents have the same families, the same label names and values, and the same samples. A test in the Metered repository asserts that byte equality. Pick the mode by who owns the storage, not by the output.

Where labels come from at exposition

A registered metric inherits the registry’s constant labels. A Family adds its dimensional labels on top. A StateSet adds a label named after the metric. An Info carries its facts as labels. They compose, so the encoded series carry the union – for example {service="orders",route="/orders",method="POST"}.

Shaping names

A series’ name comes from the path through the metric tree: the Registry prefix, the registered name, then one segment per nesting level. A nesting level is a struct field or a view entry. That makes the wire name a function of your code structure – so a refactor silently renames metrics:

  • rename a struct field, and its series name changes;
  • extract a few metrics into a sub-struct for tidiness, and they all gain a new segment.

Renamed metrics break dashboards and alerts. Name shaping decouples the wire name from the code so you can refactor freely.

Derive attributes

#[derive(MetricTree)] accepts two field attributes:

#![allow(unused)]
fn main() {
use metered::{Counter, Gauge, MetricTree};
use std::sync::atomic::{AtomicI64, AtomicU64};

#[derive(Default, MetricTree)]
struct PoolMetrics {
    #[metrics(counter)]
    acquired: AtomicU64,
    #[metrics(gauge)]
    idle: AtomicI64,
}

#[derive(Default, MetricTree)]
struct ApiMetrics {
    // The Rust field is `request_count`, but the wire segment stays `requests`.
    #[metrics(counter, rename = "requests")]
    request_count: AtomicU64,

    // Extracted into a sub-struct for organization, but flattened so the names
    // do not gain a `pool` segment: `api_acquired_total`, not
    // `api_pool_acquired_total`.
    #[metrics(flatten)]
    pool: PoolMetrics,
}
}
  • #[metrics(rename = "wire_name")] sets the segment a field contributes.
  • #[metrics(flatten)] drops the field’s segment so its children sit at the parent level – the metric-tree analogue of #[serde(flatten)].

A self-contained root: container prefix and label

A #[derive(MetricTree)] struct can carry its own name prefix and constant labels, so a root metric tree needs no hand-wired Registry:

#![allow(unused)]
fn main() {
use metered::MetricTree;
use metered_om::OpenMetricsExt;
use std::sync::atomic::{AtomicI64, AtomicU64, Ordering};

#[derive(Default, MetricTree)]
#[metrics(prefix = "app", label(service = "orders", region = "eu"))]
struct AppMetrics {
    #[metrics(counter)]
    requests: AtomicU64,
    #[metrics(gauge)]
    queue_depth: AtomicI64,
}

let metrics = AppMetrics::default();
metrics.requests.fetch_add(1, Ordering::Relaxed);

// Rendered directly -- prefix and labels are baked in.
let text = metrics.encode_to_string().unwrap();
// app_requests_total{service="orders",region="eu"} 1
}
  • #[metrics(prefix = "...")] joins a segment onto the inherited name.
  • #[metrics(label(key = "value", ...))] appends constant labels to every family.

Both compose when you nest the tree under another, exactly like a Registry prefix and label. MetricTreeExt provides schema / values to expose any self-contained tree’s schema and values. metered_om::OpenMetricsExt adds encode_to_string to encode it without a Registry. Reach for a Registry or MetricTreeView only when you compose several trees or apply the prefix and labels at the exposition site instead.

Programmatic adaptors

For hand-built trees and the Registry, the same control is available as the metered::shape adaptors Renamed and Flatten:

#![allow(unused)]
fn main() {
use metered::entry::metric;
use metered::{Counter, Registry};
use metered::shape::{Flatten, Renamed};
use std::sync::atomic::AtomicU64;

let hits = AtomicU64::new(0);
let pool_tree = AtomicU64::new(0);
let renamed = Renamed::new("requests", &hits); // contributes the `requests` segment
let flattened = Flatten::new(&pool_tree);       // contributes no segment

let mut registry = Registry::new();
registry.register(metric("api").source(&renamed).help("API requests")); // emits `api_requests_total`
}

The one invariant

Both the attributes and the adaptors apply the transform uniformly to describe and collect, and therefore to the default encode. The schema and the values can never disagree about a shaped name. There is no path where the # TYPE line says one thing and the samples say another.

Shaping works on segments within the path, not on absolute names, so a Registry prefix still composes correctly. A renamed or flattened subtree still gets the prefix like everything else.

OpenMetrics exposition

The OpenMetrics text exposition lives in its own crate, metered-om. The core metered crate only describes a MetricSchema and collects MetricValues through the MetricSink trait. A service depends on a sink crate and chooses it at the exposition site. Add both crates:

[dependencies]
metered = "0.10.0-rc.1"
metered-om = "0.10.0-rc.1"

Registry composes metric trees. The OpenMetricsRegistryExt trait adds the encode_to_string convenience:

#![allow(unused)]
fn main() {
use metered::entry::counter;
use metered::{Counter, Registry};
use metered_om::OpenMetricsRegistryExt;
use std::sync::atomic::AtomicU64;

let requests = AtomicU64::new(0);
requests.incr();

let mut registry = Registry::with_prefix("demo");
registry.label("service", "api");
registry.register(counter("requests").source(&requests).help("Total requests handled"));

let text = registry.encode_to_string().unwrap();
assert!(text.contains("# HELP demo_requests Total requests handled"));
assert!(text.contains("demo_requests_total{service=\"api\"} 1"));
}

For reusable buffers or streaming responses, drive the encoder (a MetricSink) over the schema/values directly:

#![allow(unused)]
fn main() {
use metered_om::OpenMetricsEncoder;

let schema = registry.schema();
let values = registry.values();

let mut text = String::new();
let mut encoder = OpenMetricsEncoder::new(&mut text);
encoder.encode_document(&schema, &values).unwrap();
encoder.finish().unwrap();
}

The encoder keys HELP and UNIT metadata to the exact metric family. Metadata registered for a composite tree does not leak onto its child families.

VictoriaMetrics vmrange buckets

The encoder receives each histogram whole, so the sink chooses the bucket rendering. The default is the classic cumulative le form. Switch the encoder to vmrange for VictoriaMetrics, an extension of the same OpenMetrics text format. Exponential histograms emit vmrange natively, as non-cumulative lo...hi ranges. Classic bucket histograms fall back to le:

#![allow(unused)]
fn main() {
use metered_om::{HistogramProfile, OpenMetricsEncoder};

let mut text = String::new();
let mut encoder =
    OpenMetricsEncoder::new(&mut text).histogram_profile(HistogramProfile::VmRange);
registry.encode(&mut encoder).unwrap();
encoder.finish().unwrap();
// latency_seconds_bucket{vmrange="5.000e-3...1.000e-2"} 3
}

Rendering large metric sets incrementally

encode_document renders a whole document in one synchronous call. For very large metric sets, OpenMetricsRender writes the same document in budget-bounded steps so you can yield between chunks. An item is a family declaration or a sample line. Each step emits at most budget items:

#![allow(unused)]
fn main() {
use metered_om::{OpenMetricsRender, RenderProgress};

let schema = registry.schema();
let values = registry.values();

let mut render = OpenMetricsRender::new(&schema, &values);
let mut text = String::new();
while render.step(&mut text, 256).unwrap() == RenderProgress::Pending {
    // hand control back to your loop/runtime between chunks
}
assert!(text.trim_end().ends_with("# EOF"));
}

The stepper is not async, so it works anywhere. Async callers can use RenderFuture, a dependency-free Future that writes one budget chunk per poll and yields back to the executor while more work remains:

#![allow(unused)]
fn main() {
async fn scrape(schema: &metered::MetricSchema, values: &metered::MetricValues) {
use metered_om::RenderFuture;

let text = RenderFuture::new(schema, values, 256).await.unwrap();
let _ = text;
}
}

For tests and tooling, parse text exposition back into a structural model:

#![allow(unused)]
fn main() {
use metered_om::OpenMetricsDocument;

let doc = OpenMetricsDocument::parse(&text).unwrap();
let requests = doc.family("demo_requests").unwrap();
assert_eq!(requests.help.as_deref(), Some("Total requests handled"));
}

When metrics already live inside an app context, use MetricTreeView<C>. It stores selector closures rather than metric references:

#![allow(unused)]
fn main() {
use metered::entry::counter;
use metered::MetricTreeView;
use metered_om::OpenMetricsViewExt;
use std::sync::atomic::AtomicU64;

struct App {
    requests: AtomicU64,
}

let app = App { requests: AtomicU64::new(0) };
let mut view = MetricTreeView::with_prefix("demo");
view.register(counter("requests").select(|app: &App| &app.requests).help("Total requests"));

let text = view.encode_to_string(&app).unwrap();
}

This keeps registry composition free of shared ownership: the app/context owns the metrics, and the view only describes how to borrow them.

For existing non-metric state, register a direct reader:

#![allow(unused)]
fn main() {
use metered::entry::gauge_value;
use metered::MetricTreeView;
struct App { enabled: bool }
let app = App { enabled: true };
let mut view = MetricTreeView::with_prefix("demo");
view.register(
    gauge_value("enabled")
        .read(|app: &App| app.enabled as i64)
        .help("Whether the app is enabled"),
);
}

Adapting plain state with Registry

On the borrowed Registry path, the adapter metrics expose state that is not itself a metric: an AtomicBool, a queue length, a running total. The adapter reads the state at encode time, so the exported value is always live:

#![allow(unused)]
fn main() {
use std::sync::atomic::{AtomicBool, Ordering};
use metered::adapter::{flag, CounterFn};
use metered::entry::metric;
use metered::Registry;
use metered_om::OpenMetricsRegistryExt;

let enabled = AtomicBool::new(true);
let processed = std::sync::atomic::AtomicU64::new(7);

let flag_metric = flag(|| enabled.load(Ordering::Relaxed));     // gauge 0/1
let processed_metric = CounterFn(|| processed.load(Ordering::Relaxed));

let mut registry = Registry::new();
registry.register(metric("enabled").source(&flag_metric).help("Enabled"));
registry.register(metric("processed").source(&processed_metric).help("Processed"));
let text = registry.encode_to_string().unwrap();
}

When you own the type, prefer to implement Metric for it, so the type and its value live in one place. The adapter metrics are for state you only want to read.

Serving foreign Prometheus text

Migrations are rarely all-or-nothing. Part of a process often still produces classic Prometheus text: an older metrics stack, a sidecar, a library you do not own. TextSourceTree keeps those metrics on the same scrape endpoint as your native metered trees. It is a MetricTree that parses its source’s Prometheus text and re-emits the samples. It changes no name, label, or value, so existing dashboards keep working unmodified while you migrate one subsystem at a time.

The parser is lenient by design: it accepts the dialect that serde_prometheus and similar producers emit. That dialect is looser than OpenMetrics: metadata lines are optional, the parser tolerates spaces around = inside label braces, and a counter can lack the _total suffix. Mount one TextSourceTree beside your native trees:

#![allow(unused)]
fn main() {
use metered::entry::{counter, metric};
use metered::{Counter, Registry};
use metered_om::prom_text::TextSourceTree;
use metered_om::OpenMetricsRegistryExt;
use std::sync::atomic::AtomicU64;

// A native metered counter...
let requests = AtomicU64::new(0);
requests.incr();

// ...beside a foreign producer that already emits classic Prometheus text.
let legacy = TextSourceTree::new(|| {
    "legacy_hit_count{method=\"GetOrder\"} 42\n\
     legacy_response_seconds{method=\"GetOrder\",quantile=\"0.95\"} 0.250\n"
        .to_owned()
});

let mut registry = Registry::with_prefix("demo");
registry.register(counter("requests").source(&requests).help("Native requests"));
registry.register(metric("legacy").source(&legacy));

let text = registry.encode_to_string().unwrap();
// Native metrics carry the registry prefix...
assert!(text.contains("demo_requests_total 1"));
// ...while foreign samples are re-emitted exactly as parsed: no `demo_` prefix,
// no `_total` normalization, no invented `# TYPE` line.
assert!(text.contains("legacy_hit_count{method=\"GetOrder\"} 42"));
}

The tree calls the closure passed to TextSourceTree::new on every scrape, so the exported values are always live. The tree emits each sample with the same name, the same labels, and the same integral or float rendering as the source. Only the label form changes: the output uses the canonical OpenMetrics form, k="v" with no spaces. The bytes can differ, but the samples stay the same. Foreign samples are deliberately untyped: classic Prometheus text carries no # TYPE metadata, and an invented one would change the exposition. For that reason the tree’s mount name does not prefix foreign samples, and the encoder writes no type line for them.

A scrape must never fail because a foreign source glitched. The parser drops each line that it cannot parse. The good lines survive, and the endpoint still returns 200.

If the foreign source needs its own per-scrape maintenance, for example a swap of an interval histogram, attach that work with with_housekeep. The tree’s housekeep drives the hook once per scrape cycle:

#![allow(unused)]
fn main() {
use metered_om::prom_text::TextSourceTree;

let legacy = TextSourceTree::new(|| produce_legacy_text())
    .with_housekeep(|| swap_interval_histograms());
let _ = legacy;
fn produce_legacy_text() -> String { String::new() }
fn swap_interval_histograms() {}
}

Use TextSourceTree only during a migration. When a subsystem moves to native metered families, drop the TextSourceTree. Expose the families directly so they carry full schema metadata. If you only need the parsed samples, for a test or a one-off transform, parse_prometheus_text returns them as RawSamples without the MetricTree wrapper.

Parsing exposition back

For tests and tooling, parse OpenMetrics text into a structural model rather than matching strings:

#![allow(unused)]
fn main() {
let text = "# TYPE demo_requests counter\ndemo_requests_total 1\n# EOF\n";
use metered_om::OpenMetricsDocument;

let doc = OpenMetricsDocument::parse(text).unwrap();
assert_eq!(doc.families.len(), 1);
assert_eq!(doc.sample("demo_requests_total").map(|s| s.value.as_str()), Some("1"));
}

Exemplars

A BucketHistogram can attach OpenMetrics exemplars to its buckets. Metered does not depend on tracing or OpenTelemetry: the service or an integration layer mints the Exemplar and hands it to the observation.

#![allow(unused)]
fn main() {
use metered::bucket_histogram::Exemplar;
use metered::{BucketHistogram, Buckets};

let latency = BucketHistogram::new(Buckets::fast_seconds());

let observed = 0.012; // seconds
latency.observe_with_exemplar(
    observed,
    Exemplar {
        labels: vec![("trace_id".to_owned(), "abc123".to_owned())],
        value: observed,
        timestamp_seconds: None,
    },
);
}

observe_with_exemplar counts exactly like observe and publishes the exemplar into the bucket the value lands in, with one lock-free swap. Each bucket keeps one exemplar, the most recent, as the OpenMetrics rule requires. The exposition prints it on the _bucket line. Set Exemplar::value to the observed value yourself – the histogram stores the exemplar as given.

observe and observe_with_exemplar both return the landing bucket’s index. A sampling layer can then decide cheaply, without a second lookup, whether an observation fell into an outlier bucket worth keeping a trace for.

Supplying exemplars: ExemplarSource

When exemplars come from ambient context rather than a call-site literal, implement ExemplarSource. This is a cheap, usually stateless type. It mints an exemplar, with trace labels and an optional timestamp, for the observation just recorded.

#![allow(unused)]
fn main() {
use metered::bucket_histogram::{Exemplar, ExemplarSource};

#[derive(Clone)]
struct TraceSource {
    trace_id: &'static str,
}

impl ExemplarSource for TraceSource {
    fn exemplar(&self) -> Option<Exemplar> {
        Some(Exemplar {
            labels: vec![("trace_id".to_owned(), self.trace_id.to_owned())],
            value: 0.0, // the caller fills in the observed value
            timestamp_seconds: None,
        })
    }
}
}

The default source, NoExemplars, never produces one.

Why exemplars instead of a label

An exemplar attaches a trace id to a single observation and creates no new time series. A trace id in a label would explode cardinality (see Labels and Families). An exemplar is the sanctioned way to move from a slow bucket to an example trace.

Ambient context

Wiring a trace_id through every call site is tedious. With the exemplar-context feature, a tracing layer can set an ambient exemplar for the current scope (set_current_exemplar / with_exemplar). Any observation point can read it back through the ThreadLocalExemplars source:

use metered::bucket_histogram::{with_exemplar, Exemplar, ExemplarSource, ThreadLocalExemplars};
use metered::BucketHistogram;

let latency = BucketHistogram::default();

let span_exemplar = Exemplar {
    labels: vec![("trace_id".to_owned(), "abc123".to_owned())],
    value: 0.0,
    timestamp_seconds: None,
};
with_exemplar(span_exemplar, || {
    // Inside the scope, read the ambient exemplar and attach it.
    let observed = 0.012;
    if let Some(mut exemplar) = ThreadLocalExemplars.exemplar() {
        exemplar.value = observed;
        latency.observe_with_exemplar(observed, exemplar);
    } else {
        latency.observe(observed);
    }
});

This keeps Metered free of any tracing/OpenTelemetry dependency: the feature is just a thread-local seam a tracing integration can drive.

The span-integrated path

For span-derived metrics, you do not write any of this by hand. The exemplar feature of metered-tracing drives the ambient context from the active span’s fields. It attaches trace and span-id exemplars to the duration histograms it records. See Tracing Integration.

Tracing Integration

metered-tracing turns tracing spans into metered metrics. The core crate does not depend on tracing, OpenTelemetry, or a service framework. The point is to measure once. Code that already carries spans – an #[tracing::instrument]-ed method, a span-opening middleware – gets counters and duration histograms derived from those spans. There is no second instrumentation to write or keep in sync.

Declare one SpanMetric per semantic span, for example an RPC server, a DB client, or a worker. Wire them into a TracingMetrics layer. Enable the crate’s exemplar feature to feed the ambient exemplar context of metered from tracing spans. Any observation point can consume that context through the ThreadLocalExemplars source (see Exemplars).

Service identity is intentionally not owned by this crate. A service framework should expose service metadata (service.name, version, deployment environment, commit, etc.) as a normal Info metric and/or shared registry labels.

Semantic span metrics

Spans are not folded into one generic span_* family. Each SpanMetric:

  • matches spans by name, for example rpc.server,
  • records a <name>_requests_total counter and a <name>_duration_seconds histogram on close,
  • labels both from the span’s semantic fields, with a per-label default.

A SpanMetric is a cheap handle that you can clone. Hand one clone to the layer, which records span closes. Place another clone in your metric view under the semantic name it should carry on the wire. The result is exporter-agnostic: register it with metered-om or another sink at the exposition site.

#![allow(unused)]
fn main() {
use metered::entry::metric;
use metered::Registry;
use metered_om::OpenMetricsRegistryExt;
use metered_tracing::{SpanMetric, TracingMetrics};
use tracing_subscriber::prelude::*;

let rpc = SpanMetric::for_span("rpc.server")
    .help("RPC server calls")
    .label("rpc_method", "rpc.method")
    .label_or("rpc_status", "rpc.grpc.status_code", "OK")
    .build();

let layer = TracingMetrics::builder().recorder(rpc.clone()).build();
let subscriber = tracing_subscriber::registry().with(layer);

tracing::subscriber::with_default(subscriber, || {
    let span = tracing::info_span!(
        "rpc.server",
        rpc.method = "CreateOrder",
        rpc.grpc.status_code = "OK"
    );
    let _entered = span.enter();
});

let mut registry = Registry::with_prefix("my_service");
registry.register(metric("rpc_server").source(&rpc));

let text = registry.encode_to_string().unwrap();
assert!(text.contains("# TYPE my_service_rpc_server_duration_seconds histogram"));
assert!(text.contains(
    "my_service_rpc_server_requests_total{rpc_method=\"CreateOrder\",rpc_status=\"OK\"} 1"
));
}

Registering the rpc handle under rpc_server emits:

  • rpc_server_requests_total
  • rpc_server_duration_seconds

both labeled by the configured fields (rpc_method, rpc_status) read from the span at close. Distinct span names feed distinct families, so the method lives in a label, not the metric name.

A label reads the final field value at close, so you can declare a field up front and record it later:

#![allow(unused)]
fn main() {
use metered_tracing::SpanMetric;
let _ = SpanMetric::for_span("rpc.server").label("rpc_status", "rpc.grpc.status_code").build();
let span = tracing::info_span!(
    "rpc.server",
    rpc.grpc.status_code = tracing::field::Empty
);
span.record("rpc.grpc.status_code", "INTERNAL");
}

Custom duration buckets are per span metric:

#![allow(unused)]
fn main() {
use metered_tracing::SpanMetric;
let db = SpanMetric::for_span("db.query")
    .duration_buckets(metered::Buckets::fast_seconds())
    .label("db_operation", "db.operation")
    .build();
}

Composing across crates

A span metric belongs to the component that emits the span, because only that component knows the span’s name and field names. The component therefore owns the SpanMetric as a field. A SpanMetric is an Arc-backed handle, so the component does two things with the one it owns. It mounts a clone in its own metric view: this is the exposition side, where the metric belongs in the tree. It hands a clone to the routing layer: this is the recording side. The component implements SpanMetricsSource for the latter:

#![allow(unused)]
fn main() {
use metered::{MetricTreeView, MetricsView};
use metered_tracing::{SpanMetric, SpanMetricsSource, SpanRecorder, TracingMetrics};
use std::sync::Arc;

// In the RPC framework crate: the layer owns its span metric.
struct RpcLayer {
    server: SpanMetric,
}

impl RpcLayer {
    fn new() -> Self {
        RpcLayer {
            server: SpanMetric::for_span("rpc.server")
                .label("rpc_method", "rpc.method")
                .build(),
        }
    }
}

// Hand the routing layer the recorders this component contributes (recording
// side). `span_recorders` takes `self: &Arc<Self>` so a component can project
// durations from its own shared handle; an owned `SpanMetric` is itself a
// `SpanRecorder`.
impl SpanMetricsSource for RpcLayer {
    fn span_recorders(self: &Arc<Self>) -> Vec<Box<dyn SpanRecorder>> {
        vec![Box::new(self.server.clone())]
    }
}

// ...and mount the same handle where it belongs (exposition side). Under the
// `rpc` field in the app, the `server` segment yields `..._rpc_server_*`.
impl MetricsView for RpcLayer {
    fn metrics_view() -> MetricTreeView<'static, Self> {
        let mut view = MetricTreeView::new();
        view.register(metered::entry::metric("server").select(|layer: &RpcLayer| &layer.server));
        view
    }
}

// In the service binary: assemble the routing layer from each component's owned
// metrics. Components are shared as `Arc`s (the same handles exported through
// their views), and `.source` takes `&Arc<T>`. The layer only writes; each
// component exports its own metric in its own view, so nothing is flattened
// into a separate telemetry blob.
let rpc = Arc::new(RpcLayer::new());
let layer = TracingMetrics::builder().source(&rpc).build();
let _ = layer;
}

The service’s own MetricTree mounts each component, such as rpc and db, under its field, so every span-derived metric sits with its owner. Adding a subsystem is one more field plus one more .source(&component). No top-level code needs to know its span names or labels.

Projecting durations from a component

A SpanMetric owns its families: it both counts and times. A component can already own its duration metric as a plain Family<L, DynamicExponentialHistogram> field, the very field its MetricsView exposes for scraping. In that case you do not want a second, separately owned copy of that histogram. SpanDurations::on adapts the span to the family the component already owns. It records the open-to-close duration, in seconds, on each matching span close.

The histogram backend is a generic parameter with the dynamic exponential histogram as its default: SpanDurations<C, L> means SpanDurations<C, L, DynamicExponentialHistogram>. When an alert contract requires fixed le bounds, project to a Family<L, BucketHistogram> built with your service-level objective buckets instead. The adapter accepts any backend that implements ObserveSampled.

The rule is Arc the context, project to the metric. The adapter holds an Arc of the containing component plus a projection to the family inside it. No Arc ever wraps an individual metric, so metrics stay plain struct fields. The same component handle serves both sides: its view exports it, and the recorder feeds it to the layer.

The typed label key comes from a #[derive(SpanLabels)] struct. The field names are the OpenMetrics labels, and #[span("otel.field")] maps each to the span field it reads at close. metered_info_span! opens the span with the same field names, so there is no drift between what the span carries and what keys the histogram.

#![allow(unused)]
fn main() {
use metered::{DynamicExponentialHistogram, Family, LabelSet};
use metered_tracing::{metered_info_span, SpanDurations, SpanLabels, TracingMetrics};
use std::sync::Arc;
use tracing_subscriber::prelude::*;

#[derive(Clone, PartialEq, Eq, Hash, LabelSet, SpanLabels)]
#[span(name = "db.query", help = "DB query duration")]
struct DbLabels {
    #[span("db.operation.name")]
    db_operation: String,
}

// The component owns the duration family as a plain field; no Arc wraps the
// metric. The same field is what `Db`'s MetricsView exposes for scraping.
struct Db {
    duration: Family<DbLabels, DynamicExponentialHistogram>,
}

let db = Arc::new(Db { duration: Family::default() });

// Arc the *component* (`&db`), project to the *metric* (`|db| &db.duration`).
let telemetry = TracingMetrics::builder()
    .recorder(SpanDurations::on(DbLabels::SPAN, &db, |db: &Db| &db.duration))
    .build();
let subscriber = tracing_subscriber::registry().with(telemetry);

tracing::subscriber::with_default(subscriber, || {
    metered_info_span!(DbLabels; db_operation = "insert".to_owned()).in_scope(|| {});
});
let _ = db;
}

SpanDurations is the recorder a SpanMetricsSource component returns when its metric is a bare family rather than a SpanMetric: span_recorders clones the component Arc into one SpanDurations::on(...) per timed span.

The layer captures span fields typed: a u64 span value never round-trips through a string. The typed key conversion is fallible. A captured value that does not convert to its declared label type skips the observation. The layer counts it under TracingMetrics::malformed_spans(), a bounded counter keyed by span name that you can mount in a metric view. The label never silently defaults. A custom label type implements FromFieldValue, typically with a parse of its text form.

Exemplars

Enable metered-tracing with the exemplar feature (pulls in metered’s exemplar-context). The layer sets the ambient exemplar from the fields of the active span. The example below consumes it on a plain core histogram observation. Prefer a single layer when you need both span metrics and exemplars:

metered-tracing = { version = "0.10.0-rc.1", features = ["exemplar"] }
#![allow(unused)]
fn main() {
use metered::bucket_histogram::{ExemplarSource, ThreadLocalExemplars};
use metered::BucketHistogram;
use metered_tracing::{FieldExemplarProvider, TracingMetrics};
use tracing_subscriber::prelude::*;

let latency = BucketHistogram::default();
let provider = FieldExemplarProvider::new(["trace_id", "span_id"]);
let tracing_metrics = TracingMetrics::builder().build().with_exemplar_provider(provider);
let subscriber = tracing_subscriber::registry().with(tracing_metrics);

tracing::subscriber::with_default(subscriber, || {
    let span = tracing::info_span!(
        "http.request",
        trace_id = "4bf92f3577b34da6a3ce929d0e0e4736",
        span_id = "00f067aa0ba902b7"
    );
    let _entered = span.enter();

    // Inside the span, the ambient context carries its trace/span ids.
    let observed = 0.012;
    if let Some(mut exemplar) = ThreadLocalExemplars.exemplar() {
        exemplar.value = observed;
        latency.observe_with_exemplar(observed, exemplar);
    } else {
        latency.observe(observed);
    }
});
}

For exemplars only, with no span counters or histograms, use TracingExemplarLayer or TracingMetrics::exemplar_only(provider).

Exemplars are not distributed-tracing-specific. An exemplar is any label set that points at a concrete observation, so FieldExemplarProvider lifts whatever fields you name. Name ["trace_id", "span_id"] to link to a trace. Or name a purely local identifier like ["order_id"] to jump from a latency bucket straight to the exact entity behind it, with no trace system required. Frameworks that own canonical trace context can implement ExemplarProvider directly for custom mapping.

Marking traces for retention

A histogram adopts an exemplar when the exemplar wins its sampling window and becomes the visible bucket exemplar. The trace behind an adopted exemplar is one a dashboard can jump to, so it is exactly the trace that tail-sampling should keep. on_exemplar_adopted is that seam: the builder fires the hook with each adopted exemplar, whose labels carry the trace id.

#![allow(unused)]
fn main() {
use metered_tracing::TracingMetrics;

let telemetry = TracingMetrics::builder()
    // ...add your span recorders with `.recorder(...)` / `.source(...)`...
    .on_exemplar_adopted(|exemplar| {
        // The exemplar won its bucket: mark its trace for retention so
        // tail-sampling keeps the trace behind this latency sample.
        if let Some((_, trace_id)) = exemplar.labels.iter().find(|(k, _)| k == "trace_id") {
            mark_for_retention(trace_id);
        }
    })
    .build();
let _ = telemetry;
fn mark_for_retention(_trace_id: &str) {}
}

When you later hand the built bundle an exemplar provider with with_exemplar_provider, the bundle keeps the hook. You can register the hook on the plain builder and still attach trace context afterwards.

Fitting with service context

A service framework can keep the three observability planes aligned without coupling them:

flowchart LR
    service["Service context<br/>name, version, env, commit"] --> info["metered Info<br/>service_info"]
    service --> traceid["trace-id generator"]
    tracing["tracing spans<br/>semconv attributes"] --> traces["OTLP traces"]
    tracing --> spanmetrics["metered-tracing<br/>span metrics"]
    tracing --> exemplars["Trace exemplars<br/>on histograms"]
    info --> vm["VictoriaMetrics"]
    spanmetrics --> vm
    exemplars --> vm
    traces --> tempo["Tempo"]

Schema and dashboards

Every metric tree can both describe its schema and collect current values. The registry accepts MetricTree values and exposes both halves separately:

#![allow(unused)]
fn main() {
let schema = registry.schema();
let values = registry.values();
}

A sink combines those two parts. The OpenMetrics text sink lives in the metered-om crate:

#![allow(unused)]
fn main() {
use metered_om::OpenMetricsEncoder;

let mut text = String::new();
let mut encoder = OpenMetricsEncoder::new(&mut text);
encoder.encode_document(&schema, &values).unwrap();
encoder.finish().unwrap();
}

encode_to_string() (from metered_om::OpenMetricsRegistryExt and the sibling extension traits) is a convenience over that schema/value/encode pipeline. Because the core only exposes schema() / values() through the MetricSink seam, a different exposition format is just a different sink crate.

For a single leaf metric, implement Metric instead. Metric couples the OpenMetrics type and sample encoding in one place, and Metered provides the MetricTree implementation from that single definition. Implement MetricTree directly for composite trees that emit multiple families.

The schema captures:

  • Family name.
  • Metric type.
  • HELP text.
  • UNIT.
  • Label names.

It also produces query seeds in Prometheus PromQL or in VictoriaMetrics MetricsQL. The dialect chooses the histogram bucket grouping, le or vmrange. It also chooses whether to wrap heatmaps in prometheus_buckets:

#![allow(unused)]
fn main() {
use metered::QueryDialect;

for query in schema.queries(QueryDialect::MetricsQl) {
    println!("{} => {}", query.title, query.expr);
}
}

These are not meant to replace a real dashboard author. They are a strong starting point: counter rates, gauge/state panels, histogram p50/p95/p99, and a heatmap seed with the right bucket grouping.

Some third-party values cannot implement MetricTree. For those, use Registry::register_opaque or Registry::register_opaque_with_unit. Declare the OpenMetrics type and label names explicitly. Pass a collect closure that pushes the current samples into MetricValues. Prefer to implement MetricTree when you own the type.

Why the schema is separate

Because describe does not need a live scrape, the schema is available at build time. You can generate documentation tables, dashboard templates, or review-time diffs of what this service exposes. That works in CI, before anything runs. And because the same tree defines encode as describe plus collect, the schema you document and the metrics you emit cannot drift apart.

Demo app

The demo in examples/order-service is a small e-commerce service that shows how Metered fits a real service. Spans become metrics. Components own their metric layout. A family_by view fans out a dynamic fleet of payment rails by label. Treat it as the executable companion to this book.

cargo run -p order-service-demo

It runs a 100-request workload with the span-metrics layer installed, then prints the OpenMetrics exposition for the whole service.

The modules

ModulePattern it teaches
appthe composition root: a custom ServiceIdentity implementing Info, and App deriving MetricTree to mount components in two flattened groups – Standard (shared names + service label) and Business (order_service_ prefix, no label). This module assembles the routing layer from the components’ own span metrics
telemetrythe cross-cutting policy: the naming translation between dotted semconv span fields and snake OpenMetrics labels, plus the shared exemplar provider. There is no telemetry bundle – components own their span metrics
rpca Tower-style metrics layer that owns its rpc.server SpanMetric but never records by hand: it only opens an rpc.server semconv span and mints trace context; the metrics derive from the span
ordersbusiness counters via a typed Family + #[derive(LabelSet)], the order cache exposed as a computed gauge, and the orders.create_order operation span
dbtotal schema control via a view over real internals: the connection pool’s raw atomics shaped into chosen names/types, plus a synthesized pool_utilization no field stores, alongside db.query span-derived latency
jobsa real job runner: a live queue exposed as a gauge, plus a jobs.run span with an outcome label
paymentsa dynamic sub-service fleet of payment rails: family_by walks a live map and emits every rail’s own metrics labeled by name – a rail can join at runtime

What to notice

The demo varies along two orthogonal axes. Keep them separate as you read:

  • Composition shape – how the document splices in a metric: subtree mounts a child under a name segment, flatten splices a sub-tree with no segment, and family_by fans out N dynamic members keyed by a label.
  • Where the numbers come from – what backs a metric: a pure-metrics struct + #[derive(MetricTree)], a hand-written view over real internals, or span-derived via SpanMetric.

Most components mix several. For example, orders does all three of the second axis.

  • Components own their metrics. Each component implements MetricsView; the app never restates a child’s metrics, it just splices them. New subsystem, one line.

  • Two naming conventions, one translation. Spans use OTel semconv (dotted: rpc.method, order.category); metrics use OpenMetrics (snake: rpc_method). The SpanMetric translates between them, labeling a curated low-cardinality subset; the rest stays span-only as trace/exporter context.

  • Spans are the metrics. The RPC layer, DB, orders, and jobs only open semconv spans; a metered-tracing layer turns them into rpc_server_*, db_client_*, orders_create_*, and jobs_run_* families, each labeled from span fields. Each component owns its SpanMetric and mounts it in its own view; the service assembles the routing layer from those handles with .source(&component). There is no telemetry blob to flatten.

  • Exemplars link metrics to traces. The RPC layer puts trace_id/span_id on its span; rpc_server_duration_seconds buckets carry the matching exemplar.

  • Derive for pure metrics, view for real internals. BusinessMetrics is metrics, so it uses #[derive(MetricTree)]. The DB connection pool is real operational state with no place for #[metrics] attributes, so its schema is a hand-shaped view over the raw atomics – which is also where you get total control: chosen names, gauge-vs-counter, and synthesized metrics like pool_utilization that no field holds.

  • Dynamic shape, not a central Family. payments keeps in_flight / settlements / failures inside each PaymentRail (the service’s real shape) and a family_by view fans them out as order_service_payments_settlements_total{rail="card"} over a live, dynamic map – a rail added at runtime shows up on the next scrape with no extra wiring.

  • Two naming tiers: label for shared, prefix for service-specific. Metrics split into two groups, composed as two flattened sub-trees on App:

    • Standard / cross-service (Standard: rpc, db): the family names are the shared convention (rpc_server_requests_total, db_client_duration_seconds), so the producer is a constant service="order-service" label (#[metrics(label(...))]), letting them aggregate across the fleet.
    • Service-specific (Business: orders, payments, jobs): only this service defines them, so they live under its own order_service_ prefix (#[metrics(prefix = "order_service")]) and carry no service label – the prefix is the identity.

    Richer build identity stays in the top-level service_info metric.

Each module’s top-of-file comment states the pattern and the reasoning, so reading the source top to bottom is itself a guided tour.

Migrating from older versions

This section is for users migrating from 0.9.0 and earlier.

Metered 0.10 is a new metric model, not a new version of the old API. There is no source-compatibility shim. The new model removes the method-level measuring macros and wrappers. It removes Clear: metrics are cumulative or source-of-truth state. It moves serde off the default path. Cumulative bucket and exponential histograms replace the HDR ResponseTime/Throughput summaries, and you compute their quantiles at query time.

The strategy: coexistence, not conversion

You do not port a service in one commit. Cargo treats metered 0.9 and 0.10 as distinct packages, so both can live in one binary while you migrate. Keep old modules on 0.9 through a renamed dependency, and let new and migrated code use 0.10:

[dependencies]
metered = "0.10.0-rc.1"
metered09 = { package = "metered", version = "0.9" }

Old code changes only its use paths, for example use metered09::.... Its metrics keep recording exactly as before. Migrate module by module. After you migrate the last 0.9 metric, delete metered09 and the bridge below.

One endpoint from day one: bridge the 0.9 metrics

Exposition unifies on the 0.10 side: a single OpenMetrics endpoint, served by a 0.10 Registry / MetricTreeView, carries both worlds. The 0.9 registries do not implement 0.10’s MetricTree, so you write a small bridge. The bridge is a hand-written MetricTree impl that holds the 0.9 registry/metric handles. It reads their current values at collect time and re-emits them through the 0.10 schema/values API.

The bridge is a recipe, not a shipped crate. You own it, and it is a few dozen lines. You delete it at the end of the migration. The sketch that follows is illustrative, not compiled, because 0.9 is not a dependency of this workspace:

use metered::{join_name, MetricSchema, MetricTree, MetricType, MetricValues};
use std::sync::Arc;

/// Bridges still-live 0.9 metrics into the 0.10 exposition.
struct Bridge09 {
    /// Your macro-generated 0.9 registry for the orders module.
    orders: Arc<OrderServiceMetrics>,
}

impl MetricTree for Bridge09 {
    fn describe(&self, name: &str, labels: &[(&str, &str)], schema: &mut MetricSchema) {
        schema.add_family(&join_name(name, "find_order_hits"), MetricType::Counter, labels);
    }

    fn collect(&self, name: &str, labels: &[(&str, &str)], values: &mut MetricValues) {
        // Read the 0.9 hit counter's current value (`.0` is its inner
        // `AtomicInt`) and re-emit it as a 0.10 counter sample --
        // `values.counter` adds the `_total` suffix.
        let hits = self.orders.find_order.hit_count.0.get();
        values.counter(&join_name(name, "find_order_hits"), labels, hits);
    }
}

Mount the bridge in your 0.10 registry or view like any other tree. Because you write describe and collect together, the bridged families get real # TYPE lines, prefixes, and constant labels. They are first-class 0.10 metrics whose storage happens to still be in the 0.9 crate. As each module migrates to native 0.10 state, delete its lines from the bridge.

If you would rather not name each metric, metered_om::TextSourceTree is the zero-effort alternative. It re-encodes a 0.9 registry’s serialized serde_prometheus output through the 0.10 endpoint. The re-encode normalizes. Names, labels, and shapes survive, so an HDR summary stays a summary. Value tokens re-encode from their parsed form, and the re-encode drops sample timestamps.

It gives you no schema, no type checking, and no name shaping. Use it as a stopgap, and use the bridge as the managed path.

Dashboards

A 0.9 HDR summary exposed pre-computed quantiles, for example name{quantile="0.99"}. A 0.10 histogram exposes _bucket/_sum/_count, and dashboards query histogram_quantile(0.99, ...) instead. Update the panels for a metric when you migrate its module – Registry::schema() and the Schema and Dashboards section generate the starting queries. Query-time quantiles aggregate correctly across replicas, which the pre-computed ones never did.

The endgame

The migration ends when metered09 disappears from Cargo.toml and you delete the bridge type. There is nothing else to unwind: the endpoint, names, and dashboards were on the 0.10 shape all along.

Feature flags and stability

Metered’s default build is lean and has no foreign types in its public API. The core crate, metered-core, is metric state, composition, and schema/value collection. Operation instrumentation comes through support crates or opt-in features.

metered-core features

FeatureDefaultWhat it adds
none✓The core model: readable metric state – Counter / Gauge implementors, histograms, Info, StateSet, Family – plus Registry / MetricTreeView, schema/value collection, MetricTree / LabelSet derives, and name shaping. Use metered-om for text rendering, incremental rendering, parser support, and VictoriaMetrics vmrange.
exemplar-context✗ThreadLocalExemplars + set_current_exemplar / with_exemplar: an ambient exemplar seam for tracing layers. See Exemplars.

The metered facade

metered is a facade crate. It re-exports the whole core model, metered-core, wholesale. It surfaces the support crates as feature-gated modules, so apps carry one dependency:

FeatureDefaultWhat it adds
om✗metered::om: OpenMetrics text exposition from metered-om.
tracing✗metered::tracing: tracing-subscriber layers from metered-tracing.
telemetry-tokio✗metered::telemetry_tokio: Tokio runtime/task telemetry from metered-telemetry-tokio.
telemetry-process✗metered::telemetry_process: process telemetry from metered-telemetry-process.
telemetry-system✗metered::telemetry_system: host system telemetry from metered-telemetry-system. Raises the required rustc to 1.95 for sysinfo. Every other crate and feature holds at 1.85.
exemplar-context✗Forwards metered-core/exemplar-context.
full✗All of the preceding features.

Libraries that want maximal stability can depend on metered-core directly. The facade re-exports the same types, so their metric trees compose into any app.

Support crates

The support crates keep integration dependencies out of metered itself:

  • metered-om: OpenMetrics text rendering, incremental rendering, snapshot caching, VictoriaMetrics vmrange rendering, and Hyper 1 helpers (behind its hyper-1 feature).
  • metered-tracing: tracing-subscriber layers that turn spans into semantic metric families (one SpanMetric per span kind), each a counter + duration histogram labeled from the span’s fields. Optional exemplar feature feeds trace/span IDs into metered’s ambient exemplar context.
  • metered-telemetry-tokio: Tokio task/runtime telemetry as metric trees.
  • metered-telemetry-process: standard process telemetry (CPU, memory, file descriptors, threads) under the canonical process_* names, sampled cross-platform on each scrape.
  • metered-telemetry-system: host telemetry (CPU, memory, swap, load average, uptime) as a metric tree; optional tokio feature samples on a background task so the scrape path stays non-blocking.

Stability and dependency policy

Metered lets you upgrade it, or parts of it, without an upgrade of your whole workspace:

  • No foreign types in the default public API. metered pulls neither serde nor hdrhistogram into its public API, so a metered bump never forces a serde or hdrhistogram bump on callers.
  • One direct dependency. The procedural and derive macros are re-exported from metered, so downstream crates depend on metered alone (never metered-macro); the two halves always move together.
  • Hygienic, relocatable macros. Generated code uses ::metered:: absolute paths, so it is immune to local name shadowing.
  • Evolvable surface. Open enums such as MetricType are #[non_exhaustive], so new OpenMetrics constructs can land without a breaking change – your matches just need a wildcard arm.

Versioning

The crate is on the 0.10 line, heading toward a 1.0 that stabilizes the API. Until then, minor releases may adjust unstable corners. The preceding principles – no leaked dependencies, a single dependency, hygienic macros – do not change. If you are coming from 0.9 or earlier, see Migrating From Older Versions.