Metered
Metered is a metric-state, composition, and schema/value collection library for
Rust services. Core metered gives services readable metric state, typed metric
trees, schema/value collection, and registry views that borrow through the real
service graph at scrape time. Exposition formats live in sink crates such as
metered-om.
Metered layers operation instrumentation. metered-tracing derives metric
families from the tracing spans your code already emits. An instrumented
method gets performance metrics without a second instrumentation. Everything
else is plain core metric state that your code owns and updates directly.
This book does two things:
- Teach you to use Metered well.
- Teach you to build great metrics – and explain why Metered has its shape, so the choices feel inevitable rather than arbitrary.
If you have never instrumented a service before, that is fine: the Metrics & OpenMetrics primer starts from zero.
A first taste
#![allow(unused)]
fn main() {
use metered::entry::{counter, gauge};
use metered::{MetricTreeView, Unit};
use std::sync::atomic::AtomicU64;
struct Api {
requests: AtomicU64,
in_flight: AtomicU64,
}
fn metrics() -> MetricTreeView<'static, Api> {
let mut view = MetricTreeView::with_prefix("api");
view.register(counter("requests").select(|api: &Api| &api.requests).help("Requests"));
view.register(gauge("in_flight").select(|api: &Api| &api.in_flight).help("In-flight requests").unit(Unit::Items));
view
}
}
That view exposes the state the service already owns: request totals and current
in-flight work. The default model starts with owned state and explicit
composition. Tracing spans and metered-tracing add operation measurement.
To see the whole picture, run the demo:
cargo run -p order-service-demo
It prints the OpenMetrics document for a small e-commerce service. The demo turns spans into metrics, components own their layout, and a dynamic payment-rail fleet fans out by label. Demo App documents each pattern module by module.
How to read this book
- Concepts explains the model and the reasoning behind it. Read Why Metered has this design early. It is the key to the rest.
- Building Great Metrics is the practical craft: which metric type to reach for, how to keep labels safe, and how to keep wire names stable as code changes.
- Instrumenting & Exposing covers the mechanics: tracing-derived operation metrics, registries, exemplars, and turning a schema into dashboards.
- Reference covers the demo, migrating from older versions, and the feature flags.
What Metered deliberately avoids
- Global registries and statics. Metrics are fields on your types.
serdeon the default path. Exposition is native OpenMetrics text.- Reset/clear semantics. Counters and histograms are cumulative. The query engine computes rates and quantiles at query time, where they aggregate correctly across replicas.
- Foreign types leaking through the public API. A
meteredupgrade does not dragserdeorhdrhistogramalong.
The next section explains why each of those is a feature, not a limitation.
Metrics & OpenMetrics primer
This page is for readers who have not worked with Prometheus / OpenMetrics before. If you already have, skim it for the vocabulary Metered uses and move on.
What a metric is
A metric is a named, numeric measurement of your running program, sampled over time. “Number of requests handled,” “current queue depth,” and “request latency” are all metrics. A monitoring system scrapes your process periodically, for example every 15 seconds. It reads the current numbers and stores them as a time series it can graph and alert on.
OpenMetrics is the standard text format for that exchange, the successor to the Prometheus exposition format. Metered produces it directly.
Pull, and cumulative
Two ideas underpin everything else:
- Pull, not push. The monitoring system asks your process for its current numbers; your process does not send them anywhere. So a metric is just state you can read on demand, not an event you emit.
- Cumulative, not reset. A counter only ever goes up (for the life of the process). You never reset it after a scrape. The monitoring system stores each sample and computes differences itself. This is what lets two replicas be summed correctly, and what lets a scrape that arrives late or twice not corrupt your data.
Keep these in mind. They explain why Metered has no flush or clear
operation. They also explain why the query engine computes rates, not your
process.
The metric types
OpenMetrics has a small set of types. Metered models each one.
Counter
A value that only increases: requests handled, errors returned, bytes written.
On the wire, the encoder exposes a counter http_requests as
http_requests_total.
You do not expose a rate. You expose the running total. The query
rate(http_requests_total[5m]) turns it into “requests per second.” The query
computes it over whatever window the dashboard chooses, and it sums correctly
across instances.
Gauge
A gauge goes up and down and represents current state. Examples are queue
depth, in-flight requests, connection-pool size, a temperature, and an on/off
flag (1 or 0). The monitoring system reads a gauge as-is at scrape time.
The litmus test: if the right thing to graph is the value itself, it is a gauge. If the right thing to graph is how fast it grew, it is a counter.
Histogram
A distribution, for things like latency. A histogram does not store every
observation. It counts how many fell into each of a fixed set of buckets
with upper bounds (le, “less than or equal”). It also keeps a running sum
and count. For
a latency metric http_request_duration_seconds you get series like:
http_request_duration_seconds_bucket{le="0.005"} 24
http_request_duration_seconds_bucket{le="0.01"} 41
http_request_duration_seconds_bucket{le="+Inf"} 57
http_request_duration_seconds_sum 0.83
http_request_duration_seconds_count 57
Buckets are cumulative (le="0.01" includes everything le="0.005"
counted). The query engine computes percentiles at query time with
histogram_quantile(0.95, ...). Because the buckets are plain counters, histograms from many replicas add
up, so a fleet-wide p95 is meaningful.
Summary
An older shape that ships pre-computed quantiles (for example
quantile="0.95") plus sum/count. The catch: you cannot aggregate
pre-computed quantiles –
you cannot average two replicas’ p95s to get the fleet p95. Prefer histograms.
Metered offers summaries only for backwards compatibility (see
Migrating).
Info
Static key/value facts about the process – build version, commit, region –
exposed as a constant 1 that carries the facts as labels:
build_info{version="0.10.0",commit="abc123"} 1
You join on it in queries to attach those facts to other metrics.
StateSet
A set of mutually exclusive boolean states – a lifecycle, say – where exactly
one is 1 and the rest are 0:
service_lifecycle{service_lifecycle="starting"} 0
service_lifecycle{service_lifecycle="running"} 1
service_lifecycle{service_lifecycle="draining"} 0
Labels
You can split a metric along labels – key/value pairs that create one time series per combination:
http_requests_total{route="/orders",method="POST"} 12
http_requests_total{route="/orders",method="GET"} 87
Labels are powerful and dangerous. Every distinct combination of label values is a separate stored time series. A label with unbounded values, such as a user id, a request id, or a raw URL, creates unbounded series. This “cardinality explosion” can take down your monitoring system. The rule: keep labels bounded. Labels and Families covers this in depth.
Naming and units
Conventions that Metered follows and encourages:
- Use a
namespace_subsystem_nameshape:http_request_duration_seconds. - Counters carry no rate in the name; the encoding adds the
_totalsuffix. So name the fieldrequests, notrequests_total(or you getrequests_total_total). - Put the unit in the name, and use base units: seconds, not
milliseconds, and bytes, not kilobytes. Use
_secondsand_bytes. Metered records durations as seconds for this reason.
The mental model to carry forward
A metric is state you expose, read at scrape time, cumulative where it counts. Rates and percentiles are the monitoring system’s job, not yours. The next section shows how Metered turns that model into an API.
Why Metered has this design
Metered makes a handful of strong, opinionated choices. Each one follows from the mental model in the primer: a metric is state you expose, pull-based and cumulative. This page explains the choices, because understanding them makes the API feel obvious.
Architecture at a glance
Three independent concerns meet at the metric state your service owns:
- Metric state is readable service state: counters, gauges, histograms, info, state sets, and families.
- Composition describes how to borrow that state at scrape time through
MetricTree,Registry, andMetricTreeView. - Instrumentation updates state from tracing spans, or directly from the code that owns it.
Core metered owns the first two concerns. It does not decide where service
operations begin or how you categorize errors.
flowchart TD
subgraph svc["Your service (owns all metric state)"]
code["service code"]
spans["tracing spans"]
state["metric state<br/>Counter/Gauge implementors<br/>Histogram / Family / Info / StateSet"]
code -->|"direct updates:<br/>incr / set / observe"| state
spans -->|"metered-tracing:<br/>span metrics + exemplars"| state
end
subgraph expo["Exposition (at scrape time)"]
tree["MetricTree<br/>composed by Registry / MetricTreeView"]
schema["MetricSchema<br/>(the shape)"]
values["MetricValues<br/>(the samples)"]
render["MetricSink<br/>e.g. metered-om:<br/>OpenMetricsEncoder / OpenMetricsRender"]
tree -->|describe| schema
tree -->|collect| values
schema --> render
values --> render
end
state -.->|"borrowed at scrape, no Arc"| tree
render --> text["OpenMetrics text"]
text --> scraper["Prometheus / VictoriaMetrics"]
schema -.->|dashboard_queries| promql["PromQL dashboard seeds"]
The rest of this page is why each of those pieces looks the way it does.
Metrics are state, not a shadow system
The central idea: a metric is a piece of your service’s state, owned by the component whose behavior it describes. A queue’s depth metric is the queue’s length. A pool’s “in use” gauge is the pool’s checked-out count.
So Metered has no global registry and no statics. You do not “register a
metric with the metrics system” and then find it again by string name. You hold
concrete metric state, for example an AtomicU64 counter or a histogram, as a
field, exactly where the relevant state lives. You expose it at scrape time.
This means:
- No name-based lookups, no typos resolved at runtime, no init ordering.
- No accidental sharing: two subsystems cannot clobber each other’s metric by using the same global name.
- The borrow checker keeps instrumentation honest – a metric cannot outlive the thing it measures.
A consequence you notice in practice: prefer exposing existing state over maintaining a parallel counter. If the queue knows its length, read the queue. Do not increment a separate gauge on every push and pop and hope it never drifts.
No Arc, even across .await
Instrumentation must not force you to wrap your service in Arc, and must work
in async code where the body holds &mut self across an .await.
Direct updates satisfy this trivially. incr, set, and observe take
&self through interior mutability. They borrow the metric only for the
instant of the update, never across the measured body. Span-derived measurement
satisfies it structurally. The metered-tracing layer records the duration
when the span closes. Nothing borrows your service, or self, while the
operation runs.
For exposition, the same principle drives MetricTreeView: it stores selector
closures, not metric references, and borrows the live service at scrape time. No
Arc on every metric, no shared ownership just to encode text.
Record exactly once – even on panic or cancellation
Span-derived metrics inherit tracing’s guard semantics. The guard closes a
span exactly once: on a normal return, a panic, an early return, or an
async task cancellation. metered-tracing records on close. So span counters and
duration histograms do not leak in-flight state or lose observations when the
body exits abnormally.
OpenMetrics-native, no serde on the default path
Metered writes the OpenMetrics text format directly. It does not serialize metrics to a generic data model and then map field names to Prometheus conventions.
Why it matters:
- Fidelity. Counter
_totalsuffixes, histogram_bucket/_sum/_count,# TYPE/# HELP/# UNITmetadata, exemplars, andstatesetsemantics are first-class, not approximated by reshaping JSON. - Dependency hygiene. The default public API of
meteredhas no foreign types, so ameteredupgrade never forces aserdeorhdrhistogrambump on your workspace. See Feature flags and stability.
Cumulative histograms over in-process summaries
Older metrics libraries and pre-computed HDR summaries compute quantiles in your process and expose them as a summary. The problem is aggregation: you cannot combine two replicas’ pre-computed p95s into a fleet p95.
Metered’s operation duration paths – metered-tracing span durations and
direct observe_duration calls – record into cumulative bucket histograms.
The query
engine computes percentiles at query time with histogram_quantile, where they
aggregate across replicas correctly. The bucket counters are also lock-free, so the hot path never
blocks.
A lock-free, allocation-free hot path
Stock metrics back their state with atomics allocated once at construction. The
recording path – incr, observe, gauge set – is a relaxed atomic operation
with no allocation and no mutex. Per-bucket exemplars are lock-free too. Each
bucket has its own swap slot, off the counting path. The slot changes only when
your code supplies an exemplar.
The cost guidance follows. A plain counter is the cheap metric for the hottest paths. Duration histograms are the richer metrics you reserve for entry points.
Schema, values, and rendering are separate
Every metric tree can do two independent things: describe its schema – family names, types, units, label names – and collect its current values. The renderer combines a schema and a value set into OpenMetrics text.
This split buys a lot:
- The schema can drive documentation and dashboard generation without a live scrape (Schema and Dashboards).
- The encoder can emit values incrementally, with a budget, for very large metric sets (OpenMetrics Exposition).
- Because a single
MetricTree::encodeis defined asdescribe + collect + render, a tree can never advertise one shape in its schema and emit another in its samples. The consistency is structural, not a convention.
A type for the leaf, a trait for the tree
Two traits, with one job each:
Metric: a single OpenMetrics family implements it (aCounterimplementor, aGaugeimplementor, a histogram). It couples the family’s type and its value collection in one place, so they cannot drift.MetricTree: anything composed of families implements it, such as a#[derive(MetricTree)]struct or aFamily. Leaves get it for free via a blanket implementation.
You implement Metric for a new leaf. You usually derive MetricTree for a
composite. That is the whole extension story.
Evolvable on purpose
Metered expects to reach 1.0 without churning its callers:
- Open enums like
MetricTypeare#[non_exhaustive], so new OpenMetrics constructs can land without breakingmatches. - The macros emit
::metered::absolute paths, so generated code is immune to local name shadowing. - The procedural and derive macros are re-exported from
metered, so downstream crates depend onmeteredalone and the two halves always move together.
With the “why” in hand, the Core model introduces the concrete types.
Two tiers: contract metrics vs diagnostics
Not all metrics are the same kind of thing. A team that treats them as one bucket gets broken dashboards and 4 different latency metrics for the same operation. The Metered model distinguishes two tiers, and serves each differently.
Tier 1 – contract or platform “API” metrics
Some metrics are an API: a stable shape that many components expose identically, that dashboards and alerts depend on, and that must not drift. Examples:
- gRPC server metrics: request count, in-flight, latency histogram, error
breakdown – the same families for every service and method, distinguished
only by
service/method/instancelabels. - HTTP server metrics, the RED signals, the same for every route.
- Tokio runtime metrics: worker count, busy ratio, queue depths, poll counts.
These share three properties:
- Defined once, reused everywhere. Every gRPC service should expose the same metric shape; you do not want each service inventing its own.
- A committed contract. Renaming or reshaping them breaks fleet-wide dashboards and alerts, so they should not change casually.
- Populated by infrastructure, not business code. A tower/tonic middleware or a runtime collector fills them in; the service author writes nothing.
The contract is a type
The elegant part: in Metered, the contract is just a typed MetricTree –
no separate schema-assertion mechanism needed. An integration crate defines the
shape once:
// in metered-tonic (illustrative)
#[derive(Default, MetricTree)]
#[metrics(prefix = "rpc_server")]
pub struct GrpcServerMetrics {
#[metrics(tree, rename = "requests")]
requests: Family<MethodKey, std::sync::atomic::AtomicU64>,
#[metrics(tree)]
in_flight: Family<MethodKey, std::sync::atomic::AtomicI64>,
#[metrics(tree)]
duration_seconds: Family<MethodKey, BucketHistogram>,
#[metrics(flatten)]
errors: Family<MethodErrorKey, std::sync::atomic::AtomicU64>,
}
The middleware records into a GrpcServerMetrics. The middleware can only
record into that type, so the compiler enforces that every service emits
exactly the contract shape. There is nothing to keep in sync, no runtime schema
check, and no way to drift. The reusable struct is the contract, and
describe() is its machine-readable schema for docs and dashboards.
This is why name shaping (#[metrics(rename/flatten)]) matters here: the wire
contract stays fixed even as the integration crate refactors its internals.
Tier 2 – diagnostic, implementation-detail metrics
Other metrics are implementation details: service-specific counters and timers you add to understand this code. Examples are a retry count, a cache hit rate, and time spent in a specific phase. You expose them for troubleshooting, and they evolve with the code. If one disappears in a refactor, no fleet dashboard breaks.
These are exactly what owner-local primitives, Family, MetricTree, and
MetricTreeView are for: cheap to put next to the behavior, owner-local, no central
contract. For diagnostics on code that already carries tracing spans,
metered-tracing derives the metrics from those spans, on top of the same
primitives. Use name
shaping when you want a particular diagnostic to stay stable, but the default
expectation is that they track the code.
Choosing the tier
| Contract, Tier 1 | Diagnostic, Tier 2 | |
|---|---|---|
| Who defines the shape | an integration crate, once | the service author, as needed |
| Who populates it | middleware / collector | service code via primitives / families / views; span-derived via metered-tracing |
| Stability | committed; do not drift | tracks the code |
| Breaking it | breaks fleet dashboards | low stakes, troubleshooting only |
| Cardinality | bounded by label contract | bounded by author discipline |
A healthy service exposes both: the platform contract metrics, so it shows up on the standard fleet dashboards for free, plus its own diagnostics. Metered serves Tier 1 through integration crates that provide the reusable typed tree and the collector. Examples are the runtime telemetry crates here – Tokio, process, and system – or a framework’s own RPC/HTTP contract trees built the same way. Metered serves Tier 2 through owner-local primitives, families, metric trees, and views. Both tiers encode into the same OpenMetrics document.
Core model
This page names the concrete types and how they fit together. It is the map. Later pages are the territory.
Three layers
- Metric state: counters, gauges, histograms, info, families.
- Composition:
MetricTree,Registry, andMetricTreeView. - Instrumentation: metrics derived from the
tracingspans your code already emits (metered-tracing), or your own code updating the state directly.
Core metered is layers 1 and 2. It does not decide where service operations
begin or how you categorize errors.
The layers are deliberately independent. You can expose state that Metered never “recorded,” for example an existing atomic or a queue length. Instrumentation can update metric state without owning exposition.
instrumentation metric state composition / exposition
--------------- ------------ -----------------------
tracing spans -------> counter/gauge values Metric (one family)
direct updates Histogram / Family ----> MetricTree (families)
(your code) Info / StateSet Registry / MetricTreeView
MetricSchema + MetricValues
-> OpenMetrics text
Instrumentation: spans and direct updates
If code already runs inside tracing spans, metered-tracing exports span
counters and duration histograms as metric state. Instrument a method once, and
its performance shows up in both traces and metrics. There is no second call
site to maintain. The span guard closes the span exactly once: on a normal
return, a panic, or an async cancellation. The derived metrics never leak
an observation.
Everywhere else, your code is the instrumentation, and it updates owned state.
Increment a counter when the event happens. Set a gauge when the value changes.
Time an operation with BucketHistogram::observe_duration. There is no
measuring wrapper layer between your code and the metric.
Leaves: Metric
A [Metric] is one OpenMetrics family. It states its metric_type() and knows
how to collect_metric(...) its current samples. Coupling the two in one trait
means a metric’s declared type and its emitted values cannot disagree.
The stock leaf categories:
| API | OpenMetrics type | Role |
|---|---|---|
Counter implementors such as AtomicU64 | counter | a monotonic count you own |
Gauge implementors such as AtomicI64 / AtomicU64 | gauge | a current value you own |
BucketHistogram | histogram | a distribution with classic le buckets |
Info | info | static key/value facts |
StateSet | stateset | one-of-N lifecycle state |
Standard-library atomics implement the relevant metric traits or Metric, so
you can expose existing state directly.
Trees: MetricTree
A [MetricTree] is anything made of families. It can describe its schema and
collect its values. The default text rendering combines the two.
Leaves are trees automatically, through a blanket impl. You get a MetricTree from:
#[derive(MetricTree)]– a struct of metrics;Family<L, M>– one metric per label set;- a hand-written
implfor a custom composite.
How the traits and types relate:
flowchart TD
counter["Counter/Gauge implementors<br/>Histogram / Info / StateSet / atomics"] -->|impl| metric["trait Metric<br/>(one family)"]
metric -->|"blanket impl"| tree["trait MetricTree<br/>(a tree of families)"]
der["derive(MetricTree) struct"] -->|impl| tree
fam["Family<L, M>"] -->|impl| tree
tree -->|"composed by"| registry["Registry (borrowed)<br/>MetricTreeView<C> (closures)"]
A leaf implements Metric. Everything else implements MetricTree directly. A
Registry or MetricTreeView composes trees under a prefix and constant labels.
The upkeep path: housekeep
MetricTree carries a third pair of methods: needs_housekeep and
housekeep. This is the seam for maintenance that must not run on the
recording path: work that takes a lock, allocates, or rebuilds internal state.
The DynamicExponentialHistogram downscale and its straggler drain live here.
So does an interval-histogram swap behind with_housekeep.
The scrape pipeline drives it, not your code. MetricTree::encode runs
housekeep first when needs_housekeep reports pending work. A Registry
scrape does the same by default. metered_om::SnapshotCache drives it on its
own refresh cycle.
The result is one simple contract. The observe path of every shipped instrument stays lock-free. Everything that is not lock-free waits for the upkeep pass, which runs once per scrape on the scraping task.
Hand-written MetricTree implementations that contain histograms must
forward housekeep and needs_housekeep to their fields. A tree that
forgets freezes its dynamic histograms at their saturation point. The derive
forwards automatically.
Composing for a scrape: Registry and MetricTreeView
Both turn a set of trees into one OpenMetrics document under a shared prefix and constant labels. They differ in ownership:
Registryis a borrowed view: you hand it&references for the duration of one encode. Good for one-off snapshots and foradaptermetrics.MetricTreeView<C>stores selector closures over an app contextCand borrows the live metrics at scrape time. Build it once, reuse it every scrape, with no shared ownership. This is the recommended service pattern.
Schema versus values
A MetricSchema is the static contract – family names, types, HELP, UNIT,
label names. MetricValues is the sampled state at one instant. A MetricSink
turns that pair into a wire format: the metered-om crate’s
OpenMetricsEncoder renders text, and OpenMetricsRender does the same
incrementally. The core crate has no encoder of its own, so a service picks or
writes the sink it needs. Because the schema is independent of any scrape, it
also feeds documentation and dashboards.
Choosing what to hold
A quick guide, expanded in Choosing Metric Types:
| Need | Reach for |
|---|---|
| Count events | a Counter implementor such as AtomicU64 |
| Current value you own | a Gauge implementor such as AtomicI64 / AtomicU64 |
| Current value something else owns | expose that state directly |
| A distribution / latency | BucketHistogram or DynamicExponentialHistogram |
| One-of-N state | StateSet |
| Static facts | Info |
| A bounded extra dimension | Family<L, M> |
Counters and histograms are cumulative. Gauges and state sets are source-of-truth values. There is deliberately no generic “reset” – it would be meaningless for some of these and wrong for the rest.
Choosing metric types
Picking the right type is most of what makes metrics good. This page is a decision guide plus a reference for each type Metered offers.
The decision in one paragraph
For an event you count, use a counter. For a current value, use a
gauge, but if another component already owns that value, expose that value
instead of a copy. For a distribution, latency in particular, use a
histogram. For one-of-N state, use a StateSet. For a static
fact, use info. For an extra bounded dimension, add a Family.
The same decision as a flowchart:
flowchart TD
start{"What are you<br/>measuring?"}
start -->|"an event happened"| counter["Counter"]
start -->|"a current value"| owned{"who owns<br/>the value?"}
owned -->|"this metric"| gauge["Gauge"]
owned -->|"something else"| expose["expose that state directly<br/>(view reader / adapter)"]
start -->|"a distribution (latency)"| hist["Histogram"]
start -->|"one-of-N state"| stateset["StateSet"]
start -->|"a static fact"| info["Info"]
counter --> dim{"need a bounded<br/>extra dimension?"}
gauge --> dim
hist --> dim
dim -->|"yes"| family["wrap in Family⟨L, M⟩"]
dim -->|"no"| done["done"]
All metric types at a glance
| OpenMetrics type | Instrument | Notes |
|---|---|---|
counter | Counter / AtomicU64 | Monotonic; suffix _total. |
gauge | Gauge view or owned Gauge/AtomicI64 field | Live pull or cached accumulator. See Gauge below. |
histogram | BucketHistogram, DynamicExponentialHistogram | Aggregates across replicas; prefer for latencies. |
summary | Summary<S> over a QuantileSource | Per-instance quantile values; does not aggregate. |
gaugehistogram | GaugeHistogram<S> over a GaugeHistogramSource, for example GaugeBuckets | Current-value buckets; _gcount/_gsum. |
stateset | StateSet | Mutually exclusive boolean states. |
info | InfoMetric | Static metadata; value always 1. |
unknown | passthrough samples | Foreign metrics of unknown semantics. |
Counter
For values that only increase: requests, errors, retries, bytes, dropped
messages. Counter is the trait. AtomicU64 is the usual concrete storage.
#![allow(unused)]
fn main() {
use metered::Counter;
use std::sync::atomic::AtomicU64;
let processed = AtomicU64::new(0);
processed.incr();
processed.incr_by(10);
assert_eq!(processed.get(), 11);
}
Make intent obvious through the field name – attempts, failures,
cache_misses – or through a domain newtype over the counter. The storage is
the same AtomicU64 either way.
Do not name a counter field *_total: the encoding adds _total, so requests
becomes requests_total on the wire (a requests_total field becomes
requests_total_total).
Gauge
For a current value that moves up and down: in-flight work, queue depth, pool
size, a flag. Gauge is the trait. Standard atomics are the usual concrete
storage.
#![allow(unused)]
fn main() {
use std::sync::atomic::{AtomicI64, Ordering};
let in_flight = AtomicI64::new(0);
in_flight.fetch_add(1, Ordering::Relaxed); // a request started
in_flight.fetch_sub(1, Ordering::Relaxed); // it finished
let flag = AtomicI64::new(0);
flag.store(1, Ordering::Relaxed); // 1 / 0
}
Your code updates plain atomics. The exposition side reads them through the
Gauge trait, which the standard atomics implement. The trait also has
incr/decr/set helpers if you prefer them, but nothing requires a metrics
call on the update path.
Gauge vs counter is the most common mistake. “Total requests,” “failed publishes,” “retries,” and “dropped messages” are counters, not gauges – you want their rate, not their instantaneous value. “In-flight requests,” “queue depth,” “cache entries,” and “enabled?” are gauges.
Prefer existing state. If another component already owns a value, expose
that value. Do not maintain a parallel gauge that can drift. See
Adding Metrics and the queue module in the demo.
Live vs cached. Both forms encode as a plain gauge because OpenMetrics has
no UpDownCounter type. A live gauge is a view reader like
gauge_value("queue_depth").read(|state: &State| state.queue.len() as i64).
Each scrape reads it fresh from your domain state, with no cache, so it is
always truthful, but the read must be cheap.
A cached accumulator is an owned Gauge/AtomicI64 you bump with
incr/decr. It reads O(1) at scrape. Use it when the value is cheap to
maintain incrementally but expensive or impossible to observe live. There is no
per-metric value cache: metered-om’s SnapshotCache bounds live-read cost at
the document level.
Histogram
For distributions – almost always latency. BucketHistogram records raw values
with observe. It records durations with observe_duration, which converts
them to seconds, the base unit. Histogram is the trait that abstracts over the
bucket and exponential backends. BucketHistogram is the classic le-bucket
implementation.
#![allow(unused)]
fn main() {
use metered::{BucketHistogram, Buckets};
// Choose buckets that bracket your expected range.
let sizes = BucketHistogram::new(Buckets::exponential(64.0, 2.0, 10));
sizes.observe(512.0);
}
Bucket presets, all in seconds for durations:
Buckets::seconds_default()– general request latencies, 5 ms to 10 s.Buckets::fast_seconds()– services below 5 ms, down to 25 µs.Buckets::slow_seconds()– DB-heavy / batch work, out to a minute.Buckets::wide_seconds()– fine near a microsecond, coarse near multi-second timeouts: 1 µs to about 30 s, at most 30% relative error. For fast paths whose latency spans a very wide range.Buckets::relative(min, max, max_relative_error)– exponential buckets sized to a target relative resolution. You give the range and the error you tolerate, and the builder solves for the count.Buckets::exponential(start, factor, n)/exponential_range(min, max, n)– the explicit-count exponential builders.
Wide dynamic range: fine low, coarse high
When you care about microsecond-scale fast paths but only need rough numbers
near timeouts, use relative resolution. A bucket near value v is about
v * max_relative_error wide. The absolute resolution is then automatically
fine at the bottom and coarse at the top.
#![allow(unused)]
fn main() {
use metered::{BucketHistogram, Buckets};
// <=10% relative from 1µs to 30s: ~0.1µs near the floor, ~1s near a 10s timeout.
let h = BucketHistogram::new(Buckets::relative(0.000_001, 30.0, 0.10));
}
Tighter error or a wider range means more buckets. Each boundary is a stored
le series per label set, so this is a dial between resolution and cardinality.
About 10% error over 1 µs to 30 s is about 180 buckets, and about 30%, from
wide_seconds, is about 70. There is no cheap way to get fine absolute
resolution across the whole range. Linear 10 µs buckets up to 10 s would be a
million series. That is exactly why exponential relative-error histogram designs
exist.
The quantile values come from histogram_quantile(...) at query time, so they aggregate
across replicas. That is the whole reason to prefer a histogram over a summary.
Exponential histograms
For wide dynamic ranges, Metered also provides log-linear exponential histogram
backends: FixedExponentialHistogram and DynamicExponentialHistogram. They
record sparse exponential bucket snapshots as ExponentialSnapshot values.
metered-om can encode them either as classic cumulative le buckets or as
VictoriaMetrics vmrange buckets. The full comparison – who chooses the
buckets, memory, observe cost, saturation behavior – is in
Histograms in Depth.
This is separate from Prometheus native histogram protobuf exposition. If a
protobuf sink is worth the effort later, it can encode from the same
MetricValues model. Metric ownership and observation code do not change.
A BucketHistogram can also attach exemplars to buckets via
observe_with_exemplar.
Summary
A Summary<S> renders each pre-computed quantile, for example
name{quantile="0.99"}, plus name_sum and name_count. It reads them from
any QuantileSource. Its quantile values are per-instance and do not
aggregate across replicas. Prefer a histogram unless you specifically need
exact per-instance quantile values or legacy dashboard parity.
#![allow(unused)]
fn main() {
use metered::summary::BucketQuantiles;
use metered::{BucketHistogram, Buckets, Summary};
let latency = BucketHistogram::new(Buckets::seconds_default());
latency.observe(0.012);
// Any `Histogram` is a `QuantileSource` via `BucketQuantiles`.
let summary = Summary::new(BucketQuantiles::new(&latency));
}
BucketQuantiles adapts any Histogram. The DynamicExponentialHistogram is
the recommended source: lock-free observe, bounded memory, bounded relative
error. A bespoke streaming sketch, for example CKMS, is intentionally not
provided. It cannot be lock-free, and it does not beat the exponential histogram
on memory. It only offers a different accuracy model, based on rank error. If a
project ever needs one, it drops in as a QuantileSource implementation behind
the same seam.
Gauge histogram
A GaugeHistogram<S> renders a current value distribution, buckets that can
decrease, as name_bucket{le} + name_gcount + name_gsum. Use it for a
live population’s size distribution, not a cumulative count of events. Implement
GaugeHistogramSource for your own state, or use the provided GaugeBuckets:
#![allow(unused)]
fn main() {
use metered::{GaugeBuckets, GaugeHistogram};
let sizes = GaugeBuckets::new([1.0, 10.0, 100.0]);
sizes.enter(5.0); // an item of size 5 is now held
sizes.leave(5.0); // it was released
let gh = GaugeHistogram::new(sizes);
}
StateSet
For mutually exclusive state – a lifecycle, a mode – where exactly one member is active:
#![allow(unused)]
fn main() {
use metered::StateSet;
let lifecycle = StateSet::new(["starting", "running", "draining"]);
lifecycle.set("running");
}
It emits one series per state, the active one 1 and the rest 0, with the
state carried in a label named after the metric.
Info
For static facts about the process – version, commit, region – as a constant
1 carrying labels:
#![allow(unused)]
fn main() {
use metered::InfoMetric;
let build = InfoMetric::new([("version", "0.10.0"), ("commit", "abc123")]);
}
Use it to attach build context to dashboards by joining on *_info.
Family: a bounded extra dimension
When one metric needs a label dimension – per route, per method, per upstream –
wrap it in a Family. See Labels and Families for the
cardinality discipline that keeps this safe.
#![allow(unused)]
fn main() {
use metered::{Counter, Family};
use std::sync::atomic::AtomicU64;
let by_route: Family<Vec<(String, String)>, AtomicU64> =
Family::with_label_names(["route"]);
by_route.with(&vec![("route".to_owned(), "/health".to_owned())], |c| c.incr());
}
Unknown: passthrough
MetricType::Unknown exists for passthrough or foreign metrics with unknown
semantics. It renders a plain sample with no suffix. You don’t construct it
directly. The passthrough sources that re-emit metrics from another system use
it.
Serving a foreign metrics endpoint, including Metered 0.9
metered_om::TextSourceTree re-emits any classic Prometheus/OpenMetrics text as
a MetricTree. This lets you serve a foreign producer on a 0.10 /metrics
endpoint during migration. The producer can be a sidecar, another exporter, or a
Metered 0.9 registry’s serde_prometheus output. Nothing about it is
0.9-specific, and Metered 0.9 is just one such producer. To run 0.9 in-process,
depend on the published 0.9 crate (see
Migrating From Older Versions).
use metered_om::TextSourceTree;
let legacy = TextSourceTree::new(|| old_0_9_registry.to_prometheus_text());
// mount `legacy` alongside your native 0.10 trees
The passthrough is a normalizing re-encode, not a byte copy. Names, labels, and shapes survive, so a 0.9 HDR summary stays a summary and existing dashboards keep working. Value tokens re-encode from their parsed form, and the re-encode drops sample timestamps. After you migrate a metric to a native 0.10 instrument, drop it from the passthrough source.
When none of these fit
Implement Metric for a custom leaf: its type plus how it reads
its value. Implement MetricTree for a custom composite. This is rare. Reach
for it only when you genuinely have a new OpenMetrics shape or an unusual source
for the value.
Histograms in depth
Metered ships three recording engines for cumulative distributions:
BucketHistogram, FixedExponentialHistogram, and
DynamicExponentialHistogram. They share the wire model, the exemplar
machinery, and the query story. They differ in who chooses the buckets, what
memory costs, and what happens when your traffic surprises you. This section
gives the full comparison. For the one-paragraph version, read
Choosing Metric Types.
The classic bucket histogram
BucketHistogram records into boundaries you choose (le buckets), plus a
running _sum and _count.
Benefits. The boundaries carry meaning. Put a bucket edge exactly at your latency objective, and the dashboard answers “how many requests beat the objective” with no interpolation error. Memory cost and encode cost stay fixed and small. Every scraper understands the output.
Costs. You must know the range before the first deploy. Boundaries that miss the real distribution give useless resolution, and changing them later breaks dashboard continuity. Each bucket is one series on the wire, so resolution multiplies cardinality.
Use it when the boundaries are part of the contract: objective edges, fixed size classes, or parity with an existing dashboard.
Exponential histograms
The exponential engines remove the boundary decision. Bucket i covers
[base^i, base^(i+1)) where base = 2^(2^-schema). One integer, the
schema, sets the relative resolution everywhere at once:
schema | Relative error per bucket |
|---|---|
| 0 | one power of two, about 100% |
| 3 | ~9% |
| 5 | ~2.2% |
| 8 | ~0.27% |
The error is relative, so one setting serves microseconds and minutes in
the same histogram. There is no boundary list to choose, tune, or migrate.
Non-positive values land in a dedicated zero bucket. The valid schema range
is 0 to 20.
FixedExponentialHistogram: dense, bounded range
FixedExponentialHistogram::try_new(min, max, schema) allocates every bucket
between min and max up front.
Benefits. The observe path is fully lock-free with no branches for table management. Memory is exact and known at construction. Construction fails closed: an impossible range or schema is an error, not a surprise later.
Costs. You are back to declaring a range. Values outside it clamp to the edge buckets. A wide range at a fine schema allocates many slots whether you hit them or not.
Use it for one hot histogram whose range you genuinely know.
DynamicExponentialHistogram: sparse, self-scaling
DynamicExponentialHistogram::new() starts at schema 5, about 2.2%
resolution, with a 256-slot table. with_params(start_schema, capacity) tunes both. This
is the engine the heavy-duty deployments run as their default, and the one
the rest of the stack assumes.
Benefits. Buckets live in a fixed-capacity lock-free table, and memory
tracks the buckets your values populate, not the range you configured. A
fleet of mostly idle histograms stays cheap. The observe path is one log2
and one atomic increment, with a one-time compare-and-swap when a value claims
a new bucket. Per-bucket exemplars are lock-free too.
Costs. The capacity is a budget. When the table saturates, the histogram downscales: it merges adjacent buckets, which halves the resolution. The counts already recorded coarsen with it. The schema only ever decreases.
A value range far wider than the capacity affords is not an error. The histogram converges to a coarse but correct summary of the range.
What saturation does not cost. The rebuild never runs on the observe path.
An observation that finds the table full sets a flag. The merge happens on
the upkeep path, in housekeep (see
the upkeep path).
Be precise about the lock-freedom claim: it covers the observe path only. The upkeep pass is not lock-free. It allocates the replacement table and takes a mutex over the retired-table list. That is fine, because it runs on the scraping task, once per scrape, never on a recording thread. A compare-and-swap gate keeps concurrent scrapers honest: one performs the rebuild, the rest skip past it.
The swap uses read-copy-update: observers keep using the old table lock-free until the histogram publishes the coarser one. The drain then folds a straggler’s late increment into the live table exactly once. No observation is ever dropped or double-counted.
Trade-offs at a glance
BucketHistogram | FixedExponentialHistogram | DynamicExponentialHistogram | |
|---|---|---|---|
| You choose | every boundary | range + schema | schema + slot budget |
| Memory | fixed, per boundary | fixed, whole range | tracks populated buckets |
| Observe cost | lock-free branch scan | fully lock-free index | lock-free log2 + add |
| Surprise range | clamps into edge buckets | clamps into edge buckets | downscales, keeps counting |
| Resolution over time | constant | constant | can coarsen, never below schema 0 |
| Failure mode | wrong boundaries forever | construction error | coarser buckets |
Memory for common ranges
Per-bucket costs, from the struct layouts:
| Engine | Bytes per bucket | What a bucket holds |
|---|---|---|
BucketHistogram | ~24 B | boundary f64 + count + exemplar slot |
FixedExponentialHistogram | 8 B | count only; this engine has no per-bucket exemplar slots |
DynamicExponentialHistogram | 32 B per table slot | index + count + exemplar slot + window state |
An exponential engine needs log2(max / min) x 2^schema buckets to span a
range. For ranges that services actually meter:
| Range | Spread | Buckets at schema 3 / 5 / 8 |
|---|---|---|
| Cache operation: 1 µs to 10 ms | 10^4 | 107 / 426 / 3,402 |
| RPC latency: 100 µs to 10 s | 10^5 | 133 / 532 / 4,252 |
| Payload size: 64 B to 16 MiB | 2^18 | 144 / 576 / 4,608 |
What that costs per engine, on the RPC latency range:
- Classic, 14 hand-picked boundaries: ~360 B, and all 15 bucket series are on the wire at every scrape, populated or not.
- Fixed at
schema5: 532 buckets x 8 B = ~4.3 KiB, allocated up front whether traffic hits them or not. Atschema8 that becomes ~34 KiB. Only populated buckets reach the wire. - Dynamic at the defaults: an 8 KiB table of 256 slots x 32 B, and that is the ceiling for any range. Only populated slots reach the wire.
The dynamic budget rule: capacity / 2^schema is how many powers of two
can populate before a downscale. The defaults give 256 / 32 = 8 powers of
two. That is a x256 spread at the full ~2.2% resolution. A healthy latency
distribution concentrates well inside that. A uniform flood across the full
x10^5 RPC range would settle at schema 3, ~133 populated buckets. That
still resolves ~9% per bucket, from the same 8 KiB.
The fleet math is where the engines separate. One thousand mostly idle
per-target histograms cost a fixed engine the full range each: ~4.3 MiB at
schema 5 on the RPC range. Dynamic tables only fill as targets actually
observe. The wire carries only what filled.
Choosing
- Boundaries are part of a contract, or a dashboard depends on exact edges:
BucketHistogram. - One histogram, hot path, known range, and you want zero table management:
FixedExponentialHistogram. - Everything else – and especially many instances, unknown ranges, or
per-target families:
DynamicExponentialHistogram. This is the default worth reaching for first.
Rendering and querying
Both exponential engines snapshot into the same ExponentialSnapshot, and
metered-om encodes a histogram family in either of two forms:
- Classic
lebuckets. Compatible with every Prometheus-style scraper andhistogram_quantile(). - VictoriaMetrics
vmrangeseries. The native form for VictoriaMetrics; its query functions consume the ranges directly.
The family declares its encode intent and the sink resolves it against its
own capability, so the same metric definition serves both fleets. Bucket
exemplars attach identically in either encoding. The quantile values always
come from the query engine, at read time, so they aggregate correctly across
replicas – the whole reason to prefer histograms over summaries.
Adding metrics
This guide is for agents and humans who add observability to Rust services with Metered. Follow it before introducing a new metric.
The rule
Metrics belong to the object that owns the observed behavior or state. Avoid detached “metrics bags” that measure unrelated components. Prefer a small metric tree or view next to the service state that already owns the behavior.
Choose the shape
Start with the state the service owns:
- A concrete
Counterimplementor, such asAtomicU64, for monotonic counts. - A concrete
Gaugeimplementor, such asAtomicI64orAtomicU64, for current values the service owns. - Histograms for distributions. Duration metric names should include units,
usually
_seconds. Infofor static build/version facts.StateSetfor enum-like lifecycle state.Family<L, M>for bounded label dimensions.- A family view –
family_view, or its one-string-key sugarfamily_by– when your state is already a keyed map of components. Expose the map you own instead of mirroring it into an ownedFamily. - Existing state directly when the value already exists, such as
AtomicBool, an atomic depth cache, or a queue length read while holding the queue lock.
Then choose operation instrumentation only if you are measuring an operation boundary:
- For an operation that already runs in a
tracingspan – an instrumented method, or a span-opening middleware – derive its metrics withmetered-tracing: one instrumentation, no double bookkeeping. - For an operation that is not span-shaped, use plain core state on the owning
component: a counter for attempts, a counter for failures, and a histogram
observed with
observe_duration.
Do not mirror mutable service state into a separate metric just to expose it. If
a queue already owns its length, expose a QueueDepth metric that reads the
queue length. If that lock is too expensive for scrapes, update a cached atomic
depth during queue mutations and expose that cache through a Metric newtype.
Gauge guidance
Gauges are for current state, not for event bookkeeping.
Good gauge examples:
- in-flight requests
- current queue depth
- on/off flag
- active workers
- cache entries
Bad gauge examples:
- total requests handled
- failed publishes
- retries
- dropped messages
Those are counters.
Labels
Labels must stay bounded and operationally useful: operation, result,
route, upstream, mode. Never label with raw user/account/request ids, URLs,
payloads, raw errors, or any open-ended input – that explodes cardinality. If
the value set is open-ended, you need a different metric, a normalized category,
or no label. See Labels and Families for the full
treatment and for how to add a label dimension with Family.
Preferred service pattern
Expose a borrowed metric view over the service. The view stores no metrics and does not require shared ownership.
#![allow(unused)]
fn main() {
use std::sync::atomic::AtomicU64;
use metered::entry::{counter, gauge};
use metered::{MetricTreeView, Unit};
struct Service {
processed: AtomicU64,
queue_depth: AtomicU64,
}
let mut view = MetricTreeView::with_prefix("service");
view.register(
counter("processed")
.select(|service: &Service| &service.processed)
.help("Processed jobs")
.unit(Unit::Items),
);
view.register(
gauge("queue_depth")
.select(|service: &Service| &service.queue_depth)
.help("Jobs waiting in the queue")
.unit(Unit::Items),
);
}
For a stable struct of borrowed metrics, derive MetricTree:
#![allow(unused)]
fn main() {
use std::sync::atomic::{AtomicI64, AtomicU64, AtomicUsize};
use metered::MetricTree;
#[derive(MetricTree)]
struct ServiceMetrics<'a> {
#[metric(gauge)]
enabled: &'a AtomicI64,
#[metric(gauge)]
queue_depth: &'a AtomicUsize,
#[metric(counter)]
processed: &'a AtomicU64,
}
}
Use this when you want a named metric tree type that you can register or nest inside another tree. To keep wire names stable while you refactor such a struct, see Shaping Names.
Custom metrics
For one OpenMetrics family, implement Metric. This keeps the type and values in
one place. A queue depth encoded as a gauge is a gauge. The implementation just
decides where the current value comes from:
#![allow(unused)]
fn main() {
use metered::{Metric, MetricType, MetricValues};
struct QueueDepth<'a>(&'a std::sync::atomic::AtomicUsize);
impl Metric for QueueDepth<'_> {
fn metric_type(&self) -> MetricType {
MetricType::Gauge
}
fn collect_metric(&self, name: &str, labels: &[(&str, &str)], values: &mut MetricValues) {
values.gauge(name, labels, self.0.load(std::sync::atomic::Ordering::Relaxed));
}
}
}
For a tree that emits multiple families, implement or derive MetricTree.
To keep emitted names stable as you rename fields/methods or reorganize structs,
use Shaping Names (#[metric(rename/flatten)]).
Checklist
Before you finish:
- The metric owner is the service/object that owns the behavior or state.
- Labels stay bounded and useful on dashboards.
- Durations include units in the metric name.
- Gauges represent current state.
- Counters represent cumulative events.
- You expose existing state directly or through a deliberate cached value.
- The domain operation remains readable.
- Tests cover emitted names, labels, classification, and schema when relevant.
- You update dashboards or docs when names or labels change.
Labels and families
Labels turn one metric into many series – one per combination of label values. They are the most useful and the most dangerous feature in metrics. This page covers how to use them safely and the two ways Metered attaches them.
Two kinds of labels
- Constant labels are the same on every series from a registry:
service,instance,region,env. You set them once on theRegistry/MetricTreeView, and every series carries them. - Dimensional labels vary per observation:
route,method,result,upstream. Each distinct value is a separate time series. These are the ones that need discipline.
#![allow(unused)]
fn main() {
use metered::Registry;
let mut registry = Registry::with_prefix("orders");
registry.label("service", "orders"); // constant: on every series
registry.label("instance", "i-1");
}
Cardinality: the one rule
Every distinct combination of dimensional label values is a separate stored time series in your monitoring system. The cost is multiplicative: 5 routes × 4 methods × 3 results = 60 series for one metric. That is fine. But a label with unbounded values creates unbounded series and can overwhelm the backend. Examples are a user id, an order id, a request id, a raw URL with query strings, and a raw error message. This is “cardinality explosion.”
The rule: dimensional labels must stay bounded, with values you mostly know ahead of time.
Good dimensional labels: route from a fixed set of endpoints, method,
result with values ok / error, upstream, mode, and error kind as an
enum.
Bad: anything per-user, per-request, per-entity, or free-form. If you want one
of those, you usually want one of three things. Use a different metric. Use a
normalized category: status class 5xx instead of the exact code, or route
template /orders/{id} instead of the concrete path. Or use an
exemplar, which carries a trace id without a new series.
Adding a dimension with Family
A Family<L, M> keeps one metric M per label set L, creating them on first
use – the analogue of a Prometheus metric family.
Dynamic label sets
For ad-hoc string labels, declare the label names up front so the schema does not depend on which values traffic happens to produce:
#![allow(unused)]
fn main() {
use metered::{Counter, Family};
use std::sync::atomic::AtomicU64;
let by_route: Family<Vec<(String, String)>, AtomicU64> =
Family::with_label_names(["route"]);
by_route.with(&vec![("route".to_owned(), "/health".to_owned())], |c| c.incr());
}
Typed label keys, preferred
For a stable dimension, derive LabelSet on a key struct. Each field becomes a
label. The type makes the dimension explicit and prevents typos:
#![allow(unused)]
fn main() {
use metered::{Counter, Family, LabelSet};
use std::sync::atomic::AtomicU64;
#[derive(Clone, PartialEq, Eq, Hash, LabelSet)]
struct RouteKey {
route: String,
method: String,
}
let requests: Family<RouteKey, AtomicU64> = Family::default();
requests.with(
&RouteKey { route: "/orders".into(), method: "POST".into() },
|c| c.incr(),
);
}
Family::default() reads the label names from the derived LabelSet, so its
schema is correct before any traffic arrives. You can drop a stale series, such
as a closed connection or a removed route, with Family::remove.
Keyed state as families
A family is a keyed subtree: one label set selects one member’s metrics.
Metered gives that subtree two ownership modes. In the owned mode, a
Family<L, M> stores the members inside the family. In the borrowed mode,
a family view iterates members that your own state stores. The two modes
give the same wire output for the same logical data.
Owned: whole metric structs per key
The member type M of a Family<L, M> is any MetricTree, not just a single
metric. A derived metric struct works as-is, so one key can own a whole bundle
of metrics:
#![allow(unused)]
fn main() {
use metered::{Counter, Family, Gauge, LabelSet, MetricTree};
use std::sync::atomic::{AtomicI64, AtomicU64};
#[derive(Clone, PartialEq, Eq, Hash, LabelSet)]
struct RailLabels {
rail: String,
}
#[derive(Default, MetricTree)]
struct RailMetrics {
#[metric(counter)]
sent: AtomicU64,
#[metric(gauge)]
queue_depth: AtomicI64,
}
let rails: Family<RailLabels, RailMetrics> = Family::default();
rails.with(&RailLabels { rail: "sepa".to_owned() }, |m| {
Counter::incr(&m.sent);
Gauge::set(&m.queue_depth, 3);
});
}
The owned mode carries the machinery with it. A MetricConstructor builds
each member on first use. Metered sorts the series by label pairs. When key values
come from external input, bound them with BoundedValues: it interns values
up to a cap and maps the rest to one overflow value.
Borrowed: family views over your own state
Often the keyed state already exists in your service: a map of rails, remotes,
or shards. The metrics live inside the members. Do not mirror that map into an
owned Family. Expose it borrowed instead, through
MetricTreeView::family_view, which keys members by a typed LabelSet:
#![allow(unused)]
fn main() {
use metered::entry::counter;
use metered::{LabelSet, MetricTreeView, MetricsView};
use std::collections::HashMap;
use std::sync::RwLock;
use std::sync::atomic::AtomicU64;
#[derive(Clone, PartialEq, Eq, Hash, LabelSet)]
struct RailLabels {
rail: String,
direction: String,
}
struct Rail {
sent: AtomicU64,
}
impl MetricsView for Rail {
fn metrics_view() -> MetricTreeView<'static, Self> {
let mut view = MetricTreeView::new();
view.register(
counter("sent")
.select(|rail: &Rail| &rail.sent)
.help("Payments sent on this rail"),
);
view
}
}
struct Rails {
map: RwLock<HashMap<RailLabels, Rail>>,
}
let mut view = MetricTreeView::with_prefix("rails");
view.family_view(Rail::metrics_view(), |rails: &Rails, out| {
for (key, rail) in rails.map.read().unwrap().iter() {
out.emit(key, rail);
}
});
}
There is one borrowed-family primitive, family_view, and one layer of
sugar. MetricTreeView::family_by is family_view for the common
one-string-key case. It takes a label name and a closure that emits
(key, member) pairs. Internally, both forms feed the same per-member
emission seam. The string form stamps its one (label, key) pair borrowed
from the caller’s &str, without a copy. And Family is the same contract
with Metered-owned storage: one keyed group of members rendered as one
labeled family, with the storage inverted.
The element shape comes from a context-free element view. Metered declares
the schema once, from Rail::metrics_view() alone. The schema declares the
key’s label names without values. An empty group still advertises its
families. Membership churn appears automatically at the next scrape: your
map’s inserts and removes are the lifecycle. For family_view, the key type
must declare its label names statically – #[derive(LabelSet)] keys do,
Vec<(String, String)> does not.
Two disciplines transfer to you in the borrowed mode. Cardinality control
belongs to whatever admits entries into your map. The iterate closure runs
on the scrape path, so keep its lock scope small. Emission order is
caller-driven: the document lists members in the order you emit them. Sort in
iterate if you want the same sorted order as a Family.
Upkeep also forwards through the group. housekeep reaches every emitted
member, so dynamic histograms inside members keep rescaling.
Which mode to use
| Question | Owned Family<L, M> | Borrowed family view |
|---|---|---|
| Storage owner | The family stores the members in its own map. | Your map or state stores the members. |
| Key type | A typed LabelSet, or Vec<(String, String)> with declared label names. | A typed LabelSet that declares its names with family_view, or one string label with family_by. |
| Cardinality control | Intern external-input keys with BoundedValues. | Whatever admits entries into your map. |
| Membership lifecycle | Family::with creates a member on first use; Family::remove drops one. | Your map’s inserts and removes; the next scrape reflects them. |
| Lock discipline | with read-locks the family while your closure runs; keep it short. | iterate runs on the scrape path; keep its lock scope small. |
| Ordering | Sorted by label pairs. | Caller-driven; sort in iterate for parity. |
One story on the wire
For the same logical data, an owned Family<L, M> and a borrowed
family_view write the same OpenMetrics document, byte for byte. The documents have
the same families, the same label names and values, and the same samples. A
test in the Metered repository asserts
that byte equality. Pick the mode by who owns the storage, not by the output.
Where labels come from at exposition
A registered metric inherits the registry’s constant labels. A Family adds its
dimensional labels on top. A StateSet adds a label named after the metric. An
Info carries its facts as labels. They compose, so the encoded series carry the
union – for example {service="orders",route="/orders",method="POST"}.
Shaping names
A series’ name comes from the path through the metric tree: the Registry
prefix, the registered name, then one segment per nesting level. A nesting level
is a struct field or a view entry. That makes the wire name a function of your
code structure – so a refactor silently renames metrics:
- rename a struct field, and its series name changes;
- extract a few metrics into a sub-struct for tidiness, and they all gain a new segment.
Renamed metrics break dashboards and alerts. Name shaping decouples the wire name from the code so you can refactor freely.
Derive attributes
#[derive(MetricTree)] accepts two field attributes:
#![allow(unused)]
fn main() {
use metered::{Counter, Gauge, MetricTree};
use std::sync::atomic::{AtomicI64, AtomicU64};
#[derive(Default, MetricTree)]
struct PoolMetrics {
#[metrics(counter)]
acquired: AtomicU64,
#[metrics(gauge)]
idle: AtomicI64,
}
#[derive(Default, MetricTree)]
struct ApiMetrics {
// The Rust field is `request_count`, but the wire segment stays `requests`.
#[metrics(counter, rename = "requests")]
request_count: AtomicU64,
// Extracted into a sub-struct for organization, but flattened so the names
// do not gain a `pool` segment: `api_acquired_total`, not
// `api_pool_acquired_total`.
#[metrics(flatten)]
pool: PoolMetrics,
}
}
#[metrics(rename = "wire_name")]sets the segment a field contributes.#[metrics(flatten)]drops the field’s segment so its children sit at the parent level – the metric-tree analogue of#[serde(flatten)].
A self-contained root: container prefix and label
A #[derive(MetricTree)] struct can carry its own name prefix and constant
labels, so a root metric tree needs no hand-wired Registry:
#![allow(unused)]
fn main() {
use metered::MetricTree;
use metered_om::OpenMetricsExt;
use std::sync::atomic::{AtomicI64, AtomicU64, Ordering};
#[derive(Default, MetricTree)]
#[metrics(prefix = "app", label(service = "orders", region = "eu"))]
struct AppMetrics {
#[metrics(counter)]
requests: AtomicU64,
#[metrics(gauge)]
queue_depth: AtomicI64,
}
let metrics = AppMetrics::default();
metrics.requests.fetch_add(1, Ordering::Relaxed);
// Rendered directly -- prefix and labels are baked in.
let text = metrics.encode_to_string().unwrap();
// app_requests_total{service="orders",region="eu"} 1
}
#[metrics(prefix = "...")]joins a segment onto the inherited name.#[metrics(label(key = "value", ...))]appends constant labels to every family.
Both compose when you nest the tree under another, exactly like a Registry
prefix and label. MetricTreeExt provides schema / values to expose any
self-contained tree’s schema and values. metered_om::OpenMetricsExt adds
encode_to_string to encode it without a Registry. Reach for a Registry or
MetricTreeView only when you compose several trees or apply the prefix and
labels at the exposition site instead.
Programmatic adaptors
For hand-built trees and the Registry, the same control is available as the
metered::shape adaptors Renamed and Flatten:
#![allow(unused)]
fn main() {
use metered::entry::metric;
use metered::{Counter, Registry};
use metered::shape::{Flatten, Renamed};
use std::sync::atomic::AtomicU64;
let hits = AtomicU64::new(0);
let pool_tree = AtomicU64::new(0);
let renamed = Renamed::new("requests", &hits); // contributes the `requests` segment
let flattened = Flatten::new(&pool_tree); // contributes no segment
let mut registry = Registry::new();
registry.register(metric("api").source(&renamed).help("API requests")); // emits `api_requests_total`
}
The one invariant
Both the attributes and the adaptors apply the transform uniformly to
describe and collect, and therefore to the default encode. The schema and
the values can never disagree about a shaped name. There is no path where the
# TYPE line says one thing and the samples say another.
Shaping works on segments within the path, not on absolute names, so a
Registry prefix still composes correctly. A renamed or flattened subtree still
gets the prefix like everything else.
OpenMetrics exposition
The OpenMetrics text exposition lives in its own crate, metered-om.
The core metered crate only describes a MetricSchema and collects
MetricValues through the MetricSink trait. A service depends on a sink
crate and chooses it at the exposition site. Add both crates:
[dependencies]
metered = "0.10.0-rc.1"
metered-om = "0.10.0-rc.1"
Registry composes metric trees. The OpenMetricsRegistryExt trait adds the
encode_to_string convenience:
#![allow(unused)]
fn main() {
use metered::entry::counter;
use metered::{Counter, Registry};
use metered_om::OpenMetricsRegistryExt;
use std::sync::atomic::AtomicU64;
let requests = AtomicU64::new(0);
requests.incr();
let mut registry = Registry::with_prefix("demo");
registry.label("service", "api");
registry.register(counter("requests").source(&requests).help("Total requests handled"));
let text = registry.encode_to_string().unwrap();
assert!(text.contains("# HELP demo_requests Total requests handled"));
assert!(text.contains("demo_requests_total{service=\"api\"} 1"));
}
For reusable buffers or streaming responses, drive the encoder (a MetricSink)
over the schema/values directly:
#![allow(unused)]
fn main() {
use metered_om::OpenMetricsEncoder;
let schema = registry.schema();
let values = registry.values();
let mut text = String::new();
let mut encoder = OpenMetricsEncoder::new(&mut text);
encoder.encode_document(&schema, &values).unwrap();
encoder.finish().unwrap();
}
The encoder keys HELP and UNIT metadata to the exact metric family.
Metadata registered for a composite tree does not leak onto its child families.
VictoriaMetrics vmrange buckets
The encoder receives each histogram whole, so the sink chooses the bucket
rendering. The default is the classic cumulative le form. Switch the encoder
to vmrange for VictoriaMetrics, an extension of the same OpenMetrics text
format. Exponential histograms emit vmrange natively, as non-cumulative
lo...hi ranges. Classic bucket histograms fall back to le:
#![allow(unused)]
fn main() {
use metered_om::{HistogramProfile, OpenMetricsEncoder};
let mut text = String::new();
let mut encoder =
OpenMetricsEncoder::new(&mut text).histogram_profile(HistogramProfile::VmRange);
registry.encode(&mut encoder).unwrap();
encoder.finish().unwrap();
// latency_seconds_bucket{vmrange="5.000e-3...1.000e-2"} 3
}
Rendering large metric sets incrementally
encode_document renders a whole document in one synchronous call. For very
large metric sets, OpenMetricsRender writes the same document in
budget-bounded steps so you can yield between chunks. An item is a family
declaration or a sample line. Each step emits at most budget items:
#![allow(unused)]
fn main() {
use metered_om::{OpenMetricsRender, RenderProgress};
let schema = registry.schema();
let values = registry.values();
let mut render = OpenMetricsRender::new(&schema, &values);
let mut text = String::new();
while render.step(&mut text, 256).unwrap() == RenderProgress::Pending {
// hand control back to your loop/runtime between chunks
}
assert!(text.trim_end().ends_with("# EOF"));
}
The stepper is not async, so it works anywhere. Async callers can use
RenderFuture, a dependency-free Future that writes one budget chunk per poll
and yields back to the executor while more work remains:
#![allow(unused)]
fn main() {
async fn scrape(schema: &metered::MetricSchema, values: &metered::MetricValues) {
use metered_om::RenderFuture;
let text = RenderFuture::new(schema, values, 256).await.unwrap();
let _ = text;
}
}
For tests and tooling, parse text exposition back into a structural model:
#![allow(unused)]
fn main() {
use metered_om::OpenMetricsDocument;
let doc = OpenMetricsDocument::parse(&text).unwrap();
let requests = doc.family("demo_requests").unwrap();
assert_eq!(requests.help.as_deref(), Some("Total requests handled"));
}
When metrics already live inside an app context, use MetricTreeView<C>.
It stores selector closures rather than metric references:
#![allow(unused)]
fn main() {
use metered::entry::counter;
use metered::MetricTreeView;
use metered_om::OpenMetricsViewExt;
use std::sync::atomic::AtomicU64;
struct App {
requests: AtomicU64,
}
let app = App { requests: AtomicU64::new(0) };
let mut view = MetricTreeView::with_prefix("demo");
view.register(counter("requests").select(|app: &App| &app.requests).help("Total requests"));
let text = view.encode_to_string(&app).unwrap();
}
This keeps registry composition free of shared ownership: the app/context owns the metrics, and the view only describes how to borrow them.
For existing non-metric state, register a direct reader:
#![allow(unused)]
fn main() {
use metered::entry::gauge_value;
use metered::MetricTreeView;
struct App { enabled: bool }
let app = App { enabled: true };
let mut view = MetricTreeView::with_prefix("demo");
view.register(
gauge_value("enabled")
.read(|app: &App| app.enabled as i64)
.help("Whether the app is enabled"),
);
}
Adapting plain state with Registry
On the borrowed Registry path, the adapter metrics expose state that is not
itself a metric: an AtomicBool, a queue length, a running total. The adapter
reads the state at encode time, so the exported value is always live:
#![allow(unused)]
fn main() {
use std::sync::atomic::{AtomicBool, Ordering};
use metered::adapter::{flag, CounterFn};
use metered::entry::metric;
use metered::Registry;
use metered_om::OpenMetricsRegistryExt;
let enabled = AtomicBool::new(true);
let processed = std::sync::atomic::AtomicU64::new(7);
let flag_metric = flag(|| enabled.load(Ordering::Relaxed)); // gauge 0/1
let processed_metric = CounterFn(|| processed.load(Ordering::Relaxed));
let mut registry = Registry::new();
registry.register(metric("enabled").source(&flag_metric).help("Enabled"));
registry.register(metric("processed").source(&processed_metric).help("Processed"));
let text = registry.encode_to_string().unwrap();
}
When you own the type, prefer to implement Metric for it, so the type and its
value live in one place. The adapter metrics are for state you only want to
read.
Serving foreign Prometheus text
Migrations are rarely all-or-nothing. Part of a process often still produces
classic Prometheus text: an older metrics stack, a sidecar, a library you do
not own. TextSourceTree keeps those metrics on the same scrape endpoint as
your native metered trees. It is a MetricTree that parses its source’s
Prometheus text and re-emits the samples. It changes no name, label, or
value, so existing dashboards keep working unmodified while you migrate one
subsystem at a time.
The parser is lenient by design: it accepts the dialect that serde_prometheus
and similar producers emit. That dialect is looser than OpenMetrics: metadata
lines are optional, the parser tolerates spaces around = inside label braces,
and a counter can lack the _total suffix. Mount one TextSourceTree beside
your native trees:
#![allow(unused)]
fn main() {
use metered::entry::{counter, metric};
use metered::{Counter, Registry};
use metered_om::prom_text::TextSourceTree;
use metered_om::OpenMetricsRegistryExt;
use std::sync::atomic::AtomicU64;
// A native metered counter...
let requests = AtomicU64::new(0);
requests.incr();
// ...beside a foreign producer that already emits classic Prometheus text.
let legacy = TextSourceTree::new(|| {
"legacy_hit_count{method=\"GetOrder\"} 42\n\
legacy_response_seconds{method=\"GetOrder\",quantile=\"0.95\"} 0.250\n"
.to_owned()
});
let mut registry = Registry::with_prefix("demo");
registry.register(counter("requests").source(&requests).help("Native requests"));
registry.register(metric("legacy").source(&legacy));
let text = registry.encode_to_string().unwrap();
// Native metrics carry the registry prefix...
assert!(text.contains("demo_requests_total 1"));
// ...while foreign samples are re-emitted exactly as parsed: no `demo_` prefix,
// no `_total` normalization, no invented `# TYPE` line.
assert!(text.contains("legacy_hit_count{method=\"GetOrder\"} 42"));
}
The tree calls the closure passed to TextSourceTree::new on every scrape, so
the exported values are always live. The tree emits each sample with the same
name, the same labels, and the same integral or float rendering as the source.
Only the label form changes: the output uses the canonical OpenMetrics form,
k="v" with no spaces. The bytes can differ, but the samples stay the same.
Foreign samples are deliberately untyped: classic Prometheus text carries no
# TYPE metadata, and an invented one would change the exposition. For that
reason the tree’s mount name does not prefix foreign samples, and the encoder
writes no type line for them.
A scrape must never fail because a foreign source glitched. The parser drops
each line that it cannot parse. The good lines survive, and the endpoint still
returns 200.
If the foreign source needs its own per-scrape maintenance, for example a swap
of an interval histogram, attach that work with with_housekeep. The tree’s
housekeep drives the hook once per scrape cycle:
#![allow(unused)]
fn main() {
use metered_om::prom_text::TextSourceTree;
let legacy = TextSourceTree::new(|| produce_legacy_text())
.with_housekeep(|| swap_interval_histograms());
let _ = legacy;
fn produce_legacy_text() -> String { String::new() }
fn swap_interval_histograms() {}
}
Use TextSourceTree only during a migration. When a subsystem moves to native
metered families, drop the TextSourceTree. Expose the families directly so
they carry full schema metadata. If you only need the parsed samples, for a test
or a one-off transform, parse_prometheus_text returns them as RawSamples
without the MetricTree wrapper.
Parsing exposition back
For tests and tooling, parse OpenMetrics text into a structural model rather than matching strings:
#![allow(unused)]
fn main() {
let text = "# TYPE demo_requests counter\ndemo_requests_total 1\n# EOF\n";
use metered_om::OpenMetricsDocument;
let doc = OpenMetricsDocument::parse(text).unwrap();
assert_eq!(doc.families.len(), 1);
assert_eq!(doc.sample("demo_requests_total").map(|s| s.value.as_str()), Some("1"));
}
Exemplars
A BucketHistogram can attach OpenMetrics exemplars to its buckets. Metered
does not depend on tracing or OpenTelemetry: the service or an integration
layer mints the Exemplar and hands it to the observation.
#![allow(unused)]
fn main() {
use metered::bucket_histogram::Exemplar;
use metered::{BucketHistogram, Buckets};
let latency = BucketHistogram::new(Buckets::fast_seconds());
let observed = 0.012; // seconds
latency.observe_with_exemplar(
observed,
Exemplar {
labels: vec![("trace_id".to_owned(), "abc123".to_owned())],
value: observed,
timestamp_seconds: None,
},
);
}
observe_with_exemplar counts exactly like observe and publishes the
exemplar into the bucket the value lands in, with one lock-free swap. Each
bucket keeps one exemplar, the most recent, as the OpenMetrics rule requires.
The exposition prints it on the _bucket line. Set Exemplar::value to
the observed value yourself – the histogram stores the exemplar as given.
observe and observe_with_exemplar both return the landing bucket’s index. A
sampling layer can then decide cheaply, without a second lookup, whether an
observation fell into an outlier bucket worth keeping a trace for.
Supplying exemplars: ExemplarSource
When exemplars come from ambient context rather than a call-site literal,
implement ExemplarSource. This is a cheap, usually stateless type. It mints
an exemplar, with trace labels and an optional timestamp, for the observation
just recorded.
#![allow(unused)]
fn main() {
use metered::bucket_histogram::{Exemplar, ExemplarSource};
#[derive(Clone)]
struct TraceSource {
trace_id: &'static str,
}
impl ExemplarSource for TraceSource {
fn exemplar(&self) -> Option<Exemplar> {
Some(Exemplar {
labels: vec![("trace_id".to_owned(), self.trace_id.to_owned())],
value: 0.0, // the caller fills in the observed value
timestamp_seconds: None,
})
}
}
}
The default source, NoExemplars, never produces one.
Why exemplars instead of a label
An exemplar attaches a trace id to a single observation and creates no new time series. A trace id in a label would explode cardinality (see Labels and Families). An exemplar is the sanctioned way to move from a slow bucket to an example trace.
Ambient context
Wiring a trace_id through every call site is tedious. With the
exemplar-context feature, a tracing layer can set an ambient exemplar for the
current scope (set_current_exemplar / with_exemplar). Any observation
point can read it back through the ThreadLocalExemplars source:
use metered::bucket_histogram::{with_exemplar, Exemplar, ExemplarSource, ThreadLocalExemplars};
use metered::BucketHistogram;
let latency = BucketHistogram::default();
let span_exemplar = Exemplar {
labels: vec![("trace_id".to_owned(), "abc123".to_owned())],
value: 0.0,
timestamp_seconds: None,
};
with_exemplar(span_exemplar, || {
// Inside the scope, read the ambient exemplar and attach it.
let observed = 0.012;
if let Some(mut exemplar) = ThreadLocalExemplars.exemplar() {
exemplar.value = observed;
latency.observe_with_exemplar(observed, exemplar);
} else {
latency.observe(observed);
}
});
This keeps Metered free of any tracing/OpenTelemetry dependency: the feature is just a thread-local seam a tracing integration can drive.
The span-integrated path
For span-derived metrics, you do not write any of this by hand. The
exemplar feature of metered-tracing drives the ambient context from the
active span’s fields. It attaches trace and span-id exemplars to the duration
histograms it records. See Tracing Integration.
Tracing Integration
metered-tracing turns tracing spans into
metered metrics. The core crate does not depend on tracing, OpenTelemetry,
or a service framework. The point is to measure once. Code that already
carries spans – an #[tracing::instrument]-ed method, a span-opening
middleware – gets counters and duration histograms derived from those spans.
There is no second instrumentation to write or keep in sync.
Declare one SpanMetric per semantic span, for example an RPC server, a
DB client, or a worker. Wire them into a TracingMetrics layer. Enable the
crate’s exemplar feature to feed the ambient exemplar context of metered
from tracing spans. Any observation point can consume that context through the
ThreadLocalExemplars source (see Exemplars).
Service identity is intentionally not owned by this crate. A service framework
should expose service metadata (service.name, version, deployment environment,
commit, etc.) as a normal Info metric and/or shared registry labels.
Semantic span metrics
Spans are not folded into one generic span_* family. Each SpanMetric:
- matches spans by name, for example
rpc.server, - records a
<name>_requests_totalcounter and a<name>_duration_secondshistogram on close, - labels both from the span’s semantic fields, with a per-label default.
A SpanMetric is a cheap handle that you can clone. Hand one clone to the
layer, which records span closes. Place another clone in your metric view
under the semantic name it should carry on the wire. The result is
exporter-agnostic: register it with metered-om or another sink at the
exposition site.
#![allow(unused)]
fn main() {
use metered::entry::metric;
use metered::Registry;
use metered_om::OpenMetricsRegistryExt;
use metered_tracing::{SpanMetric, TracingMetrics};
use tracing_subscriber::prelude::*;
let rpc = SpanMetric::for_span("rpc.server")
.help("RPC server calls")
.label("rpc_method", "rpc.method")
.label_or("rpc_status", "rpc.grpc.status_code", "OK")
.build();
let layer = TracingMetrics::builder().recorder(rpc.clone()).build();
let subscriber = tracing_subscriber::registry().with(layer);
tracing::subscriber::with_default(subscriber, || {
let span = tracing::info_span!(
"rpc.server",
rpc.method = "CreateOrder",
rpc.grpc.status_code = "OK"
);
let _entered = span.enter();
});
let mut registry = Registry::with_prefix("my_service");
registry.register(metric("rpc_server").source(&rpc));
let text = registry.encode_to_string().unwrap();
assert!(text.contains("# TYPE my_service_rpc_server_duration_seconds histogram"));
assert!(text.contains(
"my_service_rpc_server_requests_total{rpc_method=\"CreateOrder\",rpc_status=\"OK\"} 1"
));
}
Registering the rpc handle under rpc_server emits:
rpc_server_requests_totalrpc_server_duration_seconds
both labeled by the configured fields (rpc_method, rpc_status) read from the
span at close. Distinct span names feed distinct families, so the method lives in
a label, not the metric name.
A label reads the final field value at close, so you can declare a field up front and record it later:
#![allow(unused)]
fn main() {
use metered_tracing::SpanMetric;
let _ = SpanMetric::for_span("rpc.server").label("rpc_status", "rpc.grpc.status_code").build();
let span = tracing::info_span!(
"rpc.server",
rpc.grpc.status_code = tracing::field::Empty
);
span.record("rpc.grpc.status_code", "INTERNAL");
}
Custom duration buckets are per span metric:
#![allow(unused)]
fn main() {
use metered_tracing::SpanMetric;
let db = SpanMetric::for_span("db.query")
.duration_buckets(metered::Buckets::fast_seconds())
.label("db_operation", "db.operation")
.build();
}
Composing across crates
A span metric belongs to the component that emits the span, because only
that component knows the span’s name and field names. The component therefore
owns the SpanMetric as a field. A SpanMetric is an Arc-backed handle,
so the component does two things with the one it owns. It mounts a clone in
its own metric view: this is the exposition side, where the metric belongs in
the tree. It hands a clone to the routing layer: this is the recording side.
The component implements SpanMetricsSource for the latter:
#![allow(unused)]
fn main() {
use metered::{MetricTreeView, MetricsView};
use metered_tracing::{SpanMetric, SpanMetricsSource, SpanRecorder, TracingMetrics};
use std::sync::Arc;
// In the RPC framework crate: the layer owns its span metric.
struct RpcLayer {
server: SpanMetric,
}
impl RpcLayer {
fn new() -> Self {
RpcLayer {
server: SpanMetric::for_span("rpc.server")
.label("rpc_method", "rpc.method")
.build(),
}
}
}
// Hand the routing layer the recorders this component contributes (recording
// side). `span_recorders` takes `self: &Arc<Self>` so a component can project
// durations from its own shared handle; an owned `SpanMetric` is itself a
// `SpanRecorder`.
impl SpanMetricsSource for RpcLayer {
fn span_recorders(self: &Arc<Self>) -> Vec<Box<dyn SpanRecorder>> {
vec![Box::new(self.server.clone())]
}
}
// ...and mount the same handle where it belongs (exposition side). Under the
// `rpc` field in the app, the `server` segment yields `..._rpc_server_*`.
impl MetricsView for RpcLayer {
fn metrics_view() -> MetricTreeView<'static, Self> {
let mut view = MetricTreeView::new();
view.register(metered::entry::metric("server").select(|layer: &RpcLayer| &layer.server));
view
}
}
// In the service binary: assemble the routing layer from each component's owned
// metrics. Components are shared as `Arc`s (the same handles exported through
// their views), and `.source` takes `&Arc<T>`. The layer only writes; each
// component exports its own metric in its own view, so nothing is flattened
// into a separate telemetry blob.
let rpc = Arc::new(RpcLayer::new());
let layer = TracingMetrics::builder().source(&rpc).build();
let _ = layer;
}
The service’s own MetricTree mounts each component, such as rpc and db,
under its field, so every span-derived metric sits with its owner. Adding a
subsystem is one more field plus one more .source(&component). No top-level
code needs to know its span names or labels.
Projecting durations from a component
A SpanMetric owns its families: it both counts and times. A component can
already own its duration metric as a plain
Family<L, DynamicExponentialHistogram> field, the very field its
MetricsView exposes for scraping. In that case you do not want a second,
separately owned copy of that histogram. SpanDurations::on adapts the span
to the family the component already owns. It records the open-to-close
duration, in seconds, on each matching span close.
The histogram backend is a generic parameter with the dynamic exponential
histogram as its default: SpanDurations<C, L> means
SpanDurations<C, L, DynamicExponentialHistogram>. When an alert contract
requires fixed le bounds, project to a Family<L, BucketHistogram> built
with your service-level objective buckets instead. The adapter accepts
any backend that implements ObserveSampled.
The rule is Arc the context, project to the metric. The adapter holds an
Arc of the containing component plus a projection to the family inside it.
No Arc ever wraps an individual metric, so metrics stay plain struct fields.
The same component handle serves both sides: its view exports it, and the
recorder feeds it to the layer.
The typed label key comes from a #[derive(SpanLabels)] struct. The field
names are the OpenMetrics labels, and #[span("otel.field")] maps each to the
span field it reads at close. metered_info_span! opens the span with the
same field names, so there is no drift between what the span carries and what
keys the histogram.
#![allow(unused)]
fn main() {
use metered::{DynamicExponentialHistogram, Family, LabelSet};
use metered_tracing::{metered_info_span, SpanDurations, SpanLabels, TracingMetrics};
use std::sync::Arc;
use tracing_subscriber::prelude::*;
#[derive(Clone, PartialEq, Eq, Hash, LabelSet, SpanLabels)]
#[span(name = "db.query", help = "DB query duration")]
struct DbLabels {
#[span("db.operation.name")]
db_operation: String,
}
// The component owns the duration family as a plain field; no Arc wraps the
// metric. The same field is what `Db`'s MetricsView exposes for scraping.
struct Db {
duration: Family<DbLabels, DynamicExponentialHistogram>,
}
let db = Arc::new(Db { duration: Family::default() });
// Arc the *component* (`&db`), project to the *metric* (`|db| &db.duration`).
let telemetry = TracingMetrics::builder()
.recorder(SpanDurations::on(DbLabels::SPAN, &db, |db: &Db| &db.duration))
.build();
let subscriber = tracing_subscriber::registry().with(telemetry);
tracing::subscriber::with_default(subscriber, || {
metered_info_span!(DbLabels; db_operation = "insert".to_owned()).in_scope(|| {});
});
let _ = db;
}
SpanDurations is the recorder a SpanMetricsSource component returns when its
metric is a bare family rather than a SpanMetric: span_recorders clones the
component Arc into one SpanDurations::on(...) per timed span.
The layer captures span fields typed: a u64 span value never round-trips
through a string. The typed key conversion is fallible. A captured value that
does not convert to its declared label type skips the observation. The layer
counts it under TracingMetrics::malformed_spans(), a bounded counter keyed
by span name that you can mount in a metric view. The label never silently
defaults. A custom label type implements FromFieldValue, typically with a
parse of its text form.
Exemplars
Enable metered-tracing with the exemplar feature (pulls in metered’s
exemplar-context). The layer sets the ambient exemplar from the fields of the
active span. The example below consumes it on a plain core histogram
observation. Prefer a single layer when you need both span metrics and
exemplars:
metered-tracing = { version = "0.10.0-rc.1", features = ["exemplar"] }
#![allow(unused)]
fn main() {
use metered::bucket_histogram::{ExemplarSource, ThreadLocalExemplars};
use metered::BucketHistogram;
use metered_tracing::{FieldExemplarProvider, TracingMetrics};
use tracing_subscriber::prelude::*;
let latency = BucketHistogram::default();
let provider = FieldExemplarProvider::new(["trace_id", "span_id"]);
let tracing_metrics = TracingMetrics::builder().build().with_exemplar_provider(provider);
let subscriber = tracing_subscriber::registry().with(tracing_metrics);
tracing::subscriber::with_default(subscriber, || {
let span = tracing::info_span!(
"http.request",
trace_id = "4bf92f3577b34da6a3ce929d0e0e4736",
span_id = "00f067aa0ba902b7"
);
let _entered = span.enter();
// Inside the span, the ambient context carries its trace/span ids.
let observed = 0.012;
if let Some(mut exemplar) = ThreadLocalExemplars.exemplar() {
exemplar.value = observed;
latency.observe_with_exemplar(observed, exemplar);
} else {
latency.observe(observed);
}
});
}
For exemplars only, with no span counters or histograms, use
TracingExemplarLayer or TracingMetrics::exemplar_only(provider).
Exemplars are not distributed-tracing-specific. An exemplar is any label set
that points at a concrete observation, so FieldExemplarProvider lifts
whatever fields you name. Name ["trace_id", "span_id"] to link to a trace.
Or name a purely local identifier like ["order_id"] to jump from a latency
bucket straight to the exact entity behind it, with no trace system required.
Frameworks that own canonical trace context can implement ExemplarProvider
directly for custom mapping.
Marking traces for retention
A histogram adopts an exemplar when the exemplar wins its sampling window
and becomes the visible bucket exemplar. The trace behind an adopted exemplar
is one a dashboard can jump to, so it is exactly the trace that tail-sampling
should keep. on_exemplar_adopted is that seam: the builder fires the hook
with each adopted exemplar, whose labels carry the trace id.
#![allow(unused)]
fn main() {
use metered_tracing::TracingMetrics;
let telemetry = TracingMetrics::builder()
// ...add your span recorders with `.recorder(...)` / `.source(...)`...
.on_exemplar_adopted(|exemplar| {
// The exemplar won its bucket: mark its trace for retention so
// tail-sampling keeps the trace behind this latency sample.
if let Some((_, trace_id)) = exemplar.labels.iter().find(|(k, _)| k == "trace_id") {
mark_for_retention(trace_id);
}
})
.build();
let _ = telemetry;
fn mark_for_retention(_trace_id: &str) {}
}
When you later hand the built bundle an exemplar provider with
with_exemplar_provider, the bundle keeps the hook. You can register the hook
on the plain builder and still attach trace context afterwards.
Fitting with service context
A service framework can keep the three observability planes aligned without coupling them:
flowchart LR
service["Service context<br/>name, version, env, commit"] --> info["metered Info<br/>service_info"]
service --> traceid["trace-id generator"]
tracing["tracing spans<br/>semconv attributes"] --> traces["OTLP traces"]
tracing --> spanmetrics["metered-tracing<br/>span metrics"]
tracing --> exemplars["Trace exemplars<br/>on histograms"]
info --> vm["VictoriaMetrics"]
spanmetrics --> vm
exemplars --> vm
traces --> tempo["Tempo"]
Schema and dashboards
Every metric tree can both describe its schema and collect current values. The
registry accepts MetricTree values and exposes both halves separately:
#![allow(unused)]
fn main() {
let schema = registry.schema();
let values = registry.values();
}
A sink combines those two parts. The OpenMetrics text sink lives in the
metered-om crate:
#![allow(unused)]
fn main() {
use metered_om::OpenMetricsEncoder;
let mut text = String::new();
let mut encoder = OpenMetricsEncoder::new(&mut text);
encoder.encode_document(&schema, &values).unwrap();
encoder.finish().unwrap();
}
encode_to_string() (from metered_om::OpenMetricsRegistryExt and the
sibling extension traits) is a convenience over that schema/value/encode
pipeline. Because the core only exposes schema() / values() through the
MetricSink seam, a different exposition format is just a different sink crate.
For a single leaf metric, implement Metric instead. Metric couples the
OpenMetrics type and sample encoding in one place, and Metered provides the
MetricTree implementation from that single definition. Implement
MetricTree directly for composite trees that emit multiple families.
The schema captures:
- Family name.
- Metric type.
HELPtext.UNIT.- Label names.
It also produces query seeds in Prometheus PromQL or in VictoriaMetrics
MetricsQL. The dialect chooses the histogram bucket grouping, le or
vmrange. It also chooses whether to wrap heatmaps in prometheus_buckets:
#![allow(unused)]
fn main() {
use metered::QueryDialect;
for query in schema.queries(QueryDialect::MetricsQl) {
println!("{} => {}", query.title, query.expr);
}
}
These are not meant to replace a real dashboard author. They are a strong starting point: counter rates, gauge/state panels, histogram p50/p95/p99, and a heatmap seed with the right bucket grouping.
Some third-party values cannot implement MetricTree. For those, use
Registry::register_opaque or Registry::register_opaque_with_unit. Declare
the OpenMetrics type and label names explicitly. Pass a collect closure that
pushes the current samples into MetricValues. Prefer to implement
MetricTree when you own the type.
Why the schema is separate
Because describe does not need a live scrape, the schema is available at
build time. You can generate documentation tables, dashboard templates, or
review-time diffs of what this service exposes. That works in CI, before
anything runs. And because the same tree defines encode as describe plus
collect, the schema you document and the metrics you emit cannot drift apart.
Demo app
The demo in examples/order-service is a small e-commerce service that shows
how Metered fits a real service. Spans become metrics. Components own their
metric layout. A family_by view fans out a dynamic fleet of payment rails by
label. Treat it as the executable companion to this book.
cargo run -p order-service-demo
It runs a 100-request workload with the span-metrics layer installed, then prints the OpenMetrics exposition for the whole service.
The modules
| Module | Pattern it teaches |
|---|---|
app | the composition root: a custom ServiceIdentity implementing Info, and App deriving MetricTree to mount components in two flattened groups – Standard (shared names + service label) and Business (order_service_ prefix, no label). This module assembles the routing layer from the components’ own span metrics |
telemetry | the cross-cutting policy: the naming translation between dotted semconv span fields and snake OpenMetrics labels, plus the shared exemplar provider. There is no telemetry bundle – components own their span metrics |
rpc | a Tower-style metrics layer that owns its rpc.server SpanMetric but never records by hand: it only opens an rpc.server semconv span and mints trace context; the metrics derive from the span |
orders | business counters via a typed Family + #[derive(LabelSet)], the order cache exposed as a computed gauge, and the orders.create_order operation span |
db | total schema control via a view over real internals: the connection pool’s raw atomics shaped into chosen names/types, plus a synthesized pool_utilization no field stores, alongside db.query span-derived latency |
jobs | a real job runner: a live queue exposed as a gauge, plus a jobs.run span with an outcome label |
payments | a dynamic sub-service fleet of payment rails: family_by walks a live map and emits every rail’s own metrics labeled by name – a rail can join at runtime |
What to notice
The demo varies along two orthogonal axes. Keep them separate as you read:
- Composition shape – how the document splices in a metric:
subtreemounts a child under a name segment,flattensplices a sub-tree with no segment, andfamily_byfans out N dynamic members keyed by a label. - Where the numbers come from – what backs a metric: a pure-metrics
struct +
#[derive(MetricTree)], a hand-written view over real internals, or span-derived viaSpanMetric.
Most components mix several. For example, orders does all three of the
second axis.
-
Components own their metrics. Each component implements
MetricsView; the app never restates a child’s metrics, it just splices them. New subsystem, one line. -
Two naming conventions, one translation. Spans use OTel semconv (dotted:
rpc.method,order.category); metrics use OpenMetrics (snake:rpc_method). TheSpanMetrictranslates between them, labeling a curated low-cardinality subset; the rest stays span-only as trace/exporter context. -
Spans are the metrics. The RPC layer, DB, orders, and jobs only open semconv spans; a
metered-tracinglayer turns them intorpc_server_*,db_client_*,orders_create_*, andjobs_run_*families, each labeled from span fields. Each component owns itsSpanMetricand mounts it in its own view; the service assembles the routing layer from those handles with.source(&component). There is no telemetry blob to flatten. -
Exemplars link metrics to traces. The RPC layer puts
trace_id/span_idon its span;rpc_server_duration_secondsbuckets carry the matching exemplar. -
Derive for pure metrics, view for real internals.
BusinessMetricsis metrics, so it uses#[derive(MetricTree)]. The DB connection pool is real operational state with no place for#[metrics]attributes, so its schema is a hand-shaped view over the raw atomics – which is also where you get total control: chosen names, gauge-vs-counter, and synthesized metrics likepool_utilizationthat no field holds. -
Dynamic shape, not a central Family.
paymentskeepsin_flight/settlements/failuresinside eachPaymentRail(the service’s real shape) and afamily_byview fans them out asorder_service_payments_settlements_total{rail="card"}over a live, dynamic map – a rail added at runtime shows up on the next scrape with no extra wiring. -
Two naming tiers: label for shared, prefix for service-specific. Metrics split into two groups, composed as two
flattened sub-trees onApp:- Standard / cross-service (
Standard:rpc,db): the family names are the shared convention (rpc_server_requests_total,db_client_duration_seconds), so the producer is a constantservice="order-service"label (#[metrics(label(...))]), letting them aggregate across the fleet. - Service-specific (
Business:orders,payments,jobs): only this service defines them, so they live under its ownorder_service_prefix (#[metrics(prefix = "order_service")]) and carry noservicelabel – the prefix is the identity.
Richer build identity stays in the top-level
service_infometric. - Standard / cross-service (
Each module’s top-of-file comment states the pattern and the reasoning, so reading the source top to bottom is itself a guided tour.
Migrating from older versions
This section is for users migrating from 0.9.0 and earlier.
Metered 0.10 is a new metric model, not a new version of the old API. There
is no source-compatibility shim. The new model removes the method-level
measuring macros and wrappers. It removes Clear: metrics are cumulative or
source-of-truth state. It moves serde off the default path. Cumulative bucket
and exponential histograms replace the HDR ResponseTime/Throughput
summaries, and you compute their quantiles at query time.
The strategy: coexistence, not conversion
You do not port a service in one commit. Cargo treats metered 0.9 and 0.10 as
distinct packages, so both can live in one binary while you migrate. Keep
old modules on 0.9 through a renamed dependency, and let new and migrated code
use 0.10:
[dependencies]
metered = "0.10.0-rc.1"
metered09 = { package = "metered", version = "0.9" }
Old code changes only its use paths, for example use metered09::.... Its metrics keep
recording exactly as before. Migrate module by module. After you migrate the
last 0.9 metric, delete metered09 and the bridge below.
One endpoint from day one: bridge the 0.9 metrics
Exposition unifies on the 0.10 side: a single OpenMetrics endpoint, served by a
0.10 Registry / MetricTreeView, carries both worlds. The 0.9 registries do
not implement 0.10’s MetricTree, so you write a small bridge. The bridge
is a hand-written MetricTree impl that holds the 0.9 registry/metric
handles. It reads their current values at collect time and re-emits them
through the 0.10 schema/values API.
The bridge is a recipe, not a shipped crate. You own it, and it is a few dozen lines. You delete it at the end of the migration. The sketch that follows is illustrative, not compiled, because 0.9 is not a dependency of this workspace:
use metered::{join_name, MetricSchema, MetricTree, MetricType, MetricValues};
use std::sync::Arc;
/// Bridges still-live 0.9 metrics into the 0.10 exposition.
struct Bridge09 {
/// Your macro-generated 0.9 registry for the orders module.
orders: Arc<OrderServiceMetrics>,
}
impl MetricTree for Bridge09 {
fn describe(&self, name: &str, labels: &[(&str, &str)], schema: &mut MetricSchema) {
schema.add_family(&join_name(name, "find_order_hits"), MetricType::Counter, labels);
}
fn collect(&self, name: &str, labels: &[(&str, &str)], values: &mut MetricValues) {
// Read the 0.9 hit counter's current value (`.0` is its inner
// `AtomicInt`) and re-emit it as a 0.10 counter sample --
// `values.counter` adds the `_total` suffix.
let hits = self.orders.find_order.hit_count.0.get();
values.counter(&join_name(name, "find_order_hits"), labels, hits);
}
}
Mount the bridge in your 0.10 registry or view like any other tree. Because
you write describe and collect together, the bridged families get real
# TYPE lines, prefixes, and constant labels. They are first-class 0.10
metrics whose storage happens to still be in the 0.9 crate. As each module
migrates to native 0.10 state, delete its lines from the bridge.
If you would rather not name each metric, metered_om::TextSourceTree is the
zero-effort alternative. It re-encodes a 0.9 registry’s serialized
serde_prometheus output through the 0.10 endpoint. The re-encode
normalizes. Names, labels, and shapes survive, so an HDR summary stays a
summary. Value tokens re-encode from their parsed form, and the re-encode
drops sample timestamps.
It gives you no schema, no type checking, and no name shaping. Use it as a stopgap, and use the bridge as the managed path.
Dashboards
A 0.9 HDR summary exposed pre-computed quantiles, for example
name{quantile="0.99"}. A
0.10 histogram exposes _bucket/_sum/_count, and dashboards query
histogram_quantile(0.99, ...) instead. Update the panels for a metric when
you migrate its module – Registry::schema() and the
Schema and Dashboards section generate the starting
queries. Query-time quantiles aggregate correctly across replicas, which the
pre-computed ones never did.
The endgame
The migration ends when metered09 disappears from Cargo.toml and you
delete the bridge type. There is nothing else to unwind: the endpoint, names,
and dashboards were on the 0.10 shape all along.
Feature flags and stability
Metered’s default build is lean and has no foreign types in its public API. The
core crate, metered-core, is metric state, composition, and schema/value
collection. Operation instrumentation comes through support crates or opt-in
features.
metered-core features
| Feature | Default | What it adds |
|---|---|---|
| none | ✓ | The core model: readable metric state – Counter / Gauge implementors, histograms, Info, StateSet, Family – plus Registry / MetricTreeView, schema/value collection, MetricTree / LabelSet derives, and name shaping. Use metered-om for text rendering, incremental rendering, parser support, and VictoriaMetrics vmrange. |
exemplar-context | ✗ | ThreadLocalExemplars + set_current_exemplar / with_exemplar: an ambient exemplar seam for tracing layers. See Exemplars. |
The metered facade
metered is a facade crate. It re-exports the whole core model,
metered-core, wholesale. It surfaces the support crates as feature-gated
modules, so apps carry one dependency:
| Feature | Default | What it adds |
|---|---|---|
om | ✗ | metered::om: OpenMetrics text exposition from metered-om. |
tracing | ✗ | metered::tracing: tracing-subscriber layers from metered-tracing. |
telemetry-tokio | ✗ | metered::telemetry_tokio: Tokio runtime/task telemetry from metered-telemetry-tokio. |
telemetry-process | ✗ | metered::telemetry_process: process telemetry from metered-telemetry-process. |
telemetry-system | ✗ | metered::telemetry_system: host system telemetry from metered-telemetry-system. Raises the required rustc to 1.95 for sysinfo. Every other crate and feature holds at 1.85. |
exemplar-context | ✗ | Forwards metered-core/exemplar-context. |
full | ✗ | All of the preceding features. |
Libraries that want maximal stability can depend on metered-core directly.
The facade re-exports the same types, so their metric trees compose into any
app.
Support crates
The support crates keep integration dependencies out of metered itself:
metered-om: OpenMetrics text rendering, incremental rendering, snapshot caching, VictoriaMetricsvmrangerendering, and Hyper 1 helpers (behind itshyper-1feature).metered-tracing:tracing-subscriberlayers that turn spans into semantic metric families (oneSpanMetricper span kind), each a counter + duration histogram labeled from the span’s fields. Optionalexemplarfeature feeds trace/span IDs intometered’s ambient exemplar context.metered-telemetry-tokio: Tokio task/runtime telemetry as metric trees.metered-telemetry-process: standard process telemetry (CPU, memory, file descriptors, threads) under the canonicalprocess_*names, sampled cross-platform on each scrape.metered-telemetry-system: host telemetry (CPU, memory, swap, load average, uptime) as a metric tree; optionaltokiofeature samples on a background task so the scrape path stays non-blocking.
Stability and dependency policy
Metered lets you upgrade it, or parts of it, without an upgrade of your whole workspace:
- No foreign types in the default public API.
meteredpulls neitherserdenorhdrhistograminto its public API, so ameteredbump never forces aserdeorhdrhistogrambump on callers. - One direct dependency. The procedural and derive macros are re-exported
from
metered, so downstream crates depend onmeteredalone (nevermetered-macro); the two halves always move together. - Hygienic, relocatable macros. Generated code uses
::metered::absolute paths, so it is immune to local name shadowing. - Evolvable surface. Open enums such as
MetricTypeare#[non_exhaustive], so new OpenMetrics constructs can land without a breaking change – yourmatches just need a wildcard arm.
Versioning
The crate is on the 0.10 line, heading toward a 1.0 that stabilizes the API.
Until then, minor releases may adjust unstable corners. The preceding
principles – no leaked dependencies, a single dependency, hygienic macros –
do not change. If
you are coming from 0.9 or earlier, see
Migrating From Older Versions.