Micrometer & Observability: Custom Metrics, Timers, Counters
Learn to implement custom metrics, timers, and counters with Micrometer in Spring Boot for Prometheus-based observability and production monitoring.
Learn to implement custom metrics, timers, and counters with Micrometer in Spring Boot for Prometheus-based observability and production monitoring. The guide uses practical examples to explain what is micrometer and why observability matters, meterregistry: the foundation of custom metrics and shows how to apply the ideas in a Spring Boot project.
Micrometer & Observability: Custom Metrics, Timers, Counters, Prometheus
Observability is the cornerstone of running reliable production systems. When something breaks at 3 AM, you need more than logs telling you something went wrong. You need metrics that reveal exactly where, when, and how frequently problems occur. Micrometer provides that capability in Spring Boot applications, acting as a facade that normalizes metric collection across different monitoring backends.
This post shows how to implement custom observability with Micrometer in Spring Boot 3.5, wire metrics to Prometheus, and visualize them in Grafana. For Prometheus and Grafana setup details, see Prometheus and Grafana integration.
What is Micrometer and Why Observability Matters
Introduction
Micrometer gives Spring Boot applications a vendor-neutral way to record metrics about requests, dependencies, and business behavior. This guide explains meters, tags, timers, counters, registries, and export to monitoring backends so teams can turn application signals into actionable observability.
MeterRegistry: The Foundation of Custom Metrics
Spring Boot auto-configures a MeterRegistry instance when you add the Actuator dependency. This registry handles all meter registration and is your main integration point for custom metrics.
Auto-Configured Registry
Spring Boot creates a CompositeMeterRegistry containing backends for each detected monitoring system. Add the micrometer-registry-prometheus dependency and Prometheus export activates automatically.
The registry bean is available for injection throughout your application:
import io.micrometer.core.instrument.MeterRegistry;
import io.micrometer.core.instrument.Tags;
import org.springframework.stereotype.Component;
@Component
public class OrderMetrics {
private final MeterRegistry registry;
public OrderMetrics(MeterRegistry registry) {
this.registry = registry;
}
}
Creating Custom Meters
Meters are created through builder patterns on the registry:
// Register a counter
Counter orderCounter = Counter.builder("shop.orders")
.description("Total number of orders processed")
.tags("region", "us-east", "environment", "production")
.register(registry);
// Increment the counter
orderCounter.increment();
// Register a timer
Timer orderTimer = Timer.builder("shop.order.duration")
.description("Time to process an order")
.publishPercentiles(0.5, 0.95, 0.99)
.register(registry);
// Record duration
orderTimer.record(Duration.ofMillis(150));
Using MeterBinder for Bean-Dependent Metrics
When metrics depend on other Spring beans, use MeterBinder to ensure proper initialization order:
import io.micrometer.core.instrument.Gauge;
import io.micrometer.core.instrument.binder.MeterBinder;
import org.springframework.context.annotation.Bean;
import org.springframework.stereotype.Component;
@Component
public class QueueMetricsBinder {
@Bean
public MeterBinder queueSizeMetrics(ProcessingQueue queue) {
return registry -> Gauge.builder("processing.queue.size", queue::size)
.description("Current size of the processing queue")
.register(registry);
}
}
This pattern keeps metric registration lazy. The queue size is measured only when Prometheus scrapes the endpoint, not at application startup.
Counter and Gauge: Measuring What Changes
Micrometer provides several meter types. Picking the right one matters.
Counter: Tracking Invocations
A counter tracks how many times an event occurs. Counters always increment by a positive value and never decrease. They are ideal for measuring request counts, error tallies, and event frequencies.
import io.micrometer.core.instrument.Counter;
import io.micrometer.core.instrument.MeterRegistry;
import io.micrometer.core.instrument.Tags;
@Service
public class PaymentService {
private final Counter successCounter;
private final Counter failureCounter;
public PaymentService(MeterRegistry registry) {
this.successCounter = Counter.builder("payment.success")
.description("Number of successful payments")
.tags("service", "stripe")
.register(registry);
this.failureCounter = Counter.builder("payment.failure")
.description("Number of failed payments")
.tags("service", "stripe")
.register(registry);
}
public void processPayment(Payment payment) {
try {
// Payment processing logic
successCounter.increment();
} catch (PaymentException e) {
failureCounter.increment();
throw e;
}
}
}
Counters automatically handle concurrent access safely. Multiple threads can increment simultaneously without synchronization issues.
Gauge: Measuring Current Values
A gauge tracks a value that can go up or down. Unlike counters, gauges reflect the current state of a measurement. Queue length, cache size, and memory usage are typical gauge candidates.
@Service
public class CacheMetrics {
public CacheMetrics(MeterRegistry registry) {
// Measure cache size directly
Gauge.builder("cache.entries", this::getCacheSize)
.description("Number of entries in the cache")
.tags("cache", "user-sessions")
.register(registry);
// Measure cache hit ratio
Gauge.builder("cache.hitRatio", this::calculateHitRatio)
.description("Cache hit ratio percentage")
.tags("cache", "user-sessions")
.register(registry);
}
private double getCacheSize() {
return userSessionCache.size();
}
private double calculateHitRatio() {
long hits = cacheHits.get();
long misses = cacheMisses.get();
return (double) hits / (hits + misses);
}
}
Gauges only record values when scraped. The measurement function runs on-demand during each Prometheus scrape interval.
Timer: Measuring Duration Distributions
Timers measure how long things take and record the distribution over time. They calculate count, total duration, and percentile distributions automatically.
Basic Timer Usage
import io.micrometer.core.instrument.Timer;
import io.micrometer.core.instrument.MeterRegistry;
import java.time.Duration;
@Service
public class ReportGenerationService {
private final Timer reportTimer;
public ReportGenerationService(MeterRegistry registry) {
this.reportTimer = Timer.builder("report.generation")
.description("Time to generate reports")
.publishPercentiles(0.5, 0.95, 0.99)
.publishPercentileHistogram()
.minimumExpectedValue(Duration.ofMillis(10))
.maximumExpectedValue(Duration.ofSeconds(30))
.register(registry);
}
public Report generateReport(ReportRequest request) {
return reportTimer.record(() -> {
// Report generation logic
return buildReport(request);
});
}
}
The publishPercentileHistogram() flag enables Prometheus histogram buckets, which are essential for accurate percentile calculation.
Timed Annotation: Automatic Timing
The @Timed annotation eliminates repetitive timer recording code. Spring Boot AOP automatically wraps annotated methods with timer logic:
import io.micrometer.core.annotation.Timed;
@RestController
@RequestMapping("/api/users")
public class UserController {
@Timed(value = "api.users.list", longTask = true)
@GetMapping
public List<User> listUsers() {
return userService.findAll();
}
@Timed(value = "api.users.create", description = "Time to create a user")
@PostMapping
public User createUser(@RequestBody UserRequest request) {
return userService.create(request);
}
}
The longTask = true parameter enables long task timers, which track duration even when requests take extended periods. This is valuable for monitoring batch operations and background jobs.
Programmatic Meter Registration
For dynamic meter creation based on runtime conditions, register meters programmatically:
@Component
public class DynamicMeterRegistry {
private final MeterRegistry registry;
private final ConcurrentHashMap<String, Counter> customerCounters = new ConcurrentHashMap<>();
public DynamicMeterRegistry(MeterRegistry registry) {
this.registry = registry;
}
public Counter getCustomerCounter(String customerId) {
return customerCounters.computeIfAbsent(customerId, id ->
Counter.builder("customer.requests")
.description("Requests per customer")
.tag("customerId", id)
.register(registry)
);
}
}
Be cautious with dynamic meters. Each unique tag combination creates a new time series in Prometheus. High cardinality tag values generate millions of potential series, overwhelming storage and query performance.
When to Use Each Meter Type
Pick the correct meter type to avoid instrumentation bugs and get meaningful data.
Use Counter when:
- Tracking events that occur a finite number of times
- Measuring request counts, error counts, or conversion events
- The metric should never decrease
Use Gauge when:
- Measuring a value that fluctuates bidirectionally
- Tracking current queue depth, cache size, or connection pool utilization
- The measurement is expensive and should only occur during scrapes
Use Timer when:
- Measuring the duration of operations
- Percentile distributions matter
- Tracking latency across many requests
Mermaid Diagram: Observability Pipeline
Here is the complete observability flow from your application to visualization:
graph TD
A[Spring Boot Application] --> B[Micrometer API]
B --> C[MeterRegistry]
C --> D[Prometheus Scrape Endpoint]
D --> E[Prometheus Server]
E --> F[Grafana Dashboards]
A --> G[ObservationRegistry]
G --> H[OpenTelemetry Collector]
H --> I[Distributed Traces]
J[Actuator Endpoints] --> D
K[JVM Metrics] --> C
L[Spring MVC Metrics] --> C
Spring Boot instruments JVM metrics, HTTP requests, and database connections automatically. Custom meters flow through the same registry and export pipeline.
Implementation Snippets
Custom Counter with Tags
import io.micrometer.core.instrument.Counter;
import io.micrometer.core.instrument.MeterRegistry;
import io.micrometer.core.instrument.Tags;
@Component
public class ApiMetrics {
private final MeterRegistry registry;
public ApiMetrics(MeterRegistry registry) {
this.registry = registry;
}
public void recordApiCall(String endpoint, int statusCode, long durationMs) {
Tags tags = Tags.of(
"endpoint", endpoint,
"status", String.valueOf(statusCode),
"duration_bucket", getDurationBucket(durationMs)
);
registry.counter("api.calls", tags).increment();
}
private String getDurationBucket(long ms) {
if (ms < 100) return "fast";
if (ms < 500) return "normal";
if (ms < 1000) return "slow";
return "timeout";
}
}
Timed Annotation with @Timed
import io.micrometer.core.annotation.Timed;
import org.springframework.stereotype.Service;
@Service
public class OrderService {
@Timed(value = "order.process", longTask = false)
public Order processOrder(OrderRequest request) {
validateRequest(request);
calculatePricing(request);
reserveInventory(request);
initiatePayment(request);
return createOrder(request);
}
@Timed(value = "order.validation", description = "Order validation time")
private void validateRequest(OrderRequest request) {
// Validation logic
}
}
Programmatic Meter Registration with MeterBinder
import io.micrometer.core.instrument.MeterRegistry;
import io.micrometer.core.instrument.binder.MeterBinder;
import org.springframework.context.annotation.Bean;
import org.springframework.stereotype.Component;
@Component
public class DatabaseMetricsConfig {
@Bean
public MeterBinder databasePoolMetrics(DataSource dataSource) {
return registry -> {
// Register HikariCP-specific metrics
HikariDataSource hikariDS = (HikariDataSource) dataSource;
Gauge.builder("db.pool.connections.active", hikariDS::getActiveConnections)
.description("Active database connections")
.register(registry);
Gauge.builder("db.pool.connections.idle", hikariDS::getIdleConnections)
.description("Idle database connections")
.register(registry);
Gauge.builder("db.pool.connections.total", hikariDS::getTotalConnections)
.description("Total database connections in pool")
.register(registry);
Gauge.builder("db.pool.connections.pending", hikariDS::getThreadsAwaitingConnection)
.description("Threads waiting for connection")
.register(registry);
};
}
}
Failure Scenarios
These failure modes cause production incidents when instrumentation itself breaks.
High Cardinality Tags
Creating meters with unbounded tag values generates millions of time series:
// PROBLEMATIC: High cardinality user ID
Counter.builder("user.request")
.tag("userId", userId) // Never do this with user IDs
.register(registry);
// BETTER: Use low-cardinality groupings
Counter.builder("user.request")
.tag("userTier", user.getTier()) // free, premium, enterprise
.tag("region", user.getRegion()) // limited to a few values
.register(registry);
Prometheus handles high cardinality poorly. Each unique tag combination is a separate time series requiring memory and computation. At scale, this causes OOM kills and slow queries.
Clock Drift in Timers
Timers rely on monotonic clock sources when available. However, virtualized environments sometimes report negative durations when CPU scheduling causes clock anomalies:
// Use distribution summaries for environments with clock issues
DistributionSummary.builder("request.size")
.description("Request size distribution")
.publishPercentiles(0.5, 0.95, 0.99)
.register(registry);
// Or filter outlier measurements
Timer.builder("operation.duration")
.publishPercentiles(0.5, 0.95, 0.99)
.maximumExpectedValue(Duration.ofMinutes(5)) // Reject outliers
.register(registry);
Metric Prefix Collisions
Micrometer prepends the application name automatically. Ensure metric names do not collide with existing metrics:
# application.yml
spring.application.name: payment-service
management.metrics.enable: true
management.metrics.tags:
application: ${spring.application.name}
Prometheus Scrape Failures
When Prometheus cannot reach your endpoint, gaps appear in time series data. Configure scrape timeouts and ensure your /actuator/prometheus endpoint responds within limits:
# Prometheus scrape config
scrape_configs:
- job_name: "payment-service"
scrape_interval: 15s
scrape_timeout: 10s
metrics_path: "/actuator/prometheus"
static_configs:
- targets: ["payment-service:8080"]
Production Failure Scenarios
Metric instrumentation failures are particularly painful because you discover them during an outage when you most need the data.
High Cardinality Tags Causing Prometheus Storage Overload
Each unique combination of tag values creates a separate time series in Prometheus. A metric tagged with userId against 50,000 users generates 50,000 distinct series. Prometheus handles this poorly: query performance degrades, disk usage spikes, and eventually the OOM killer terminates the Prometheus server. Your dashboards go blank during the incident you were trying to debug. The fix is tag cardinality discipline: only use low-cardinality tag values. Use userTier (free, premium, enterprise) not userId. For individual user tracking, use traces or logs.
Metric Registry Memory Leak from Unbounded Meter Registration
Dynamically creating meters in a loop without cleanup causes memory to grow unboundedly. A common mistake: creating a new timer or counter per incoming request keyed by a unique value. Each unique key permanently allocates a meter object that is never garbage collected. Monitor your MeterRegistry size in staging with GET /actuator/metrics/jvm.gc.pause during load tests. Use MeterFilter to reject or limit high-cardinality meters at registration time.
Missing Percentile Histograms Breaking SLO Alerts
Timers configured without publishPercentileHistogram() do not store the data needed to calculate percentiles. Your p95 latency alert always returns NaN because Prometheus has no histogram buckets to compute the quantile from. This silently breaks during an incident. Always enable percentile histograms on timers for user-facing endpoints, or at minimum on your critical path operations.
Scrape Interval Too Long to Catch Latency Spikes
Prometheus scrapes at a fixed interval. If your scrape_interval is 15 seconds but a latency spike lasts 8 seconds, the spike is invisible in your metrics. Averaging over 15-second intervals masks spikes that users feel but the metrics do not show. For volatile metrics, use shorter scrape intervals or push metrics via the Prometheus PushGateway for time-sensitive measurements.
Gauge Metrics Reporting Stale State
Gauges reflect the last recorded value, not the current state. If a component stops updating a gauge due to a crash or exception, the last value remains in Prometheus indefinitely. An operator looking at a queue depth gauge sees the stale value and makes decisions based on outdated information. Combine gauges with metrics that track when values were last updated, or use push-based metrics for rapidly changing state.
Trade-off Table: Micrometer vs Dropwizard vs Java Flight Recorder
| Aspect | Micrometer | Dropwizard Metrics | Java Flight Recorder |
|---|---|---|---|
| Backend Support | Prometheus, Datadog, CloudWatch, InfluxDB, Graphite | Datadog, Graphite, Ganglia | Flight Recorder only |
| API Design | Fluent builder pattern | Annotations and registry | Automatic through JVM |
| Spring Integration | First-class support | Requires bridge | Requires JMX bridge |
| Histogram Support | Percentile and percentile histogram | Settable percentiles | Only latency histograms |
| Cardinality Control | Tag-based, supports filtering | Counter per metric name | Automatic, no filtering |
| Performance Overhead | Minimal, lazy registration | Low | Minimal, native |
| Learning Curve | Low to moderate | Low | Steep for custom events |
Micrometer wins for multi-backend scenarios and Spring Boot applications. Dropwizard Metrics works if you are already committed to Graphite. JFR is excellent for JVM-level diagnostics but demands significant expertise for application-level custom events.
Observability Checklist
Use these approaches when designing your metric strategy. For comprehensive monitoring setup, see metrics monitoring and alerting.
RED Method (Rate, Errors, Duration)
Apply RED to every service endpoint:
Rate: Requests per second
api_requests_total{endpoint="/api/orders", method="POST"}
Errors: Error rate as a percentage
rate(api_requests_total{status=~"5.."}[5m])
/ rate(api_requests_total[5m]) * 100
Duration: Response time percentiles
histogram_quantile(0.95, rate(api_request_duration_seconds_bucket[5m]))
USE Method (Utilization, Saturation, Errors)
Apply USE to every resource:
Utilization: Percentage of time resource is busy
jvm_memory_used_bytes{area="heap"} / jvm_memory_max_bytes{area="heap"}
Saturation: How full the queue is
processing_queue_size / processing_queue_capacity
Errors: Error count for the resource
db_connections_failed_total
Pre-Production Checklist
- Identify all external dependencies (databases, queues, downstream APIs)
- Define SLIs (Service Level Indicators) for each endpoint
- Instrument counters for all business transactions
- Add timers to all operations exceeding 100ms
- Configure percentile histograms for latency metrics
- Set up gauges for queue depths and connection pool utilization
- Validate metric names follow naming conventions
- Test that Prometheus scrapes all required endpoints
- Verify Grafana dashboards display meaningful data
- Create alerts for SLO violations
Security Notes
The metrics endpoint needs protection. For a complete actuator security configuration, see the Spring Boot Actuator deep dive.
Exposing /actuator/prometheus Securely
Do not expose the Prometheus endpoint publicly. Restrict access to internal networks and service accounts:
# application.yml
management:
endpoints:
web:
exposure:
include: health,info,prometheus,metrics
base-path: /actuator
endpoint:
prometheus:
enabled: true
health:
show-details: when_authorized
prometheus:
metrics:
export:
enabled: true
# Network policy (Kubernetes)
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: restrict-metrics-access
spec:
podSelector:
matchLabels:
app: payment-service
policyTypes:
- Ingress
ingress:
- from:
- namespaceSelector:
matchLabels:
name: monitoring
ports:
- protocol: TCP
port: 8080
Sensitive Metrics
Some metrics may expose sensitive information. Filter them from Prometheus exports:
@Configuration
public class MetricsSecurityConfig {
@Bean
public MeterFilter sensitiveMetricsFilter() {
return MeterFilter.deny(id -> {
String name = id.getName();
return name.contains("password") ||
name.contains("secret") ||
name.contains("token");
});
}
}
Authentication for Metrics Endpoints
In environments requiring strict access control, add authentication to actuator endpoints:
management:
endpoints:
web:
authentication: basic
auth:
enabled: true
Common Pitfalls / Anti-Patterns
These mistakes degrade metric data quality.
Forgetting to Close Resources
While Micrometer meters are generally lifecycle-managed, ensure you clean up resources when dynamically creating meters:
// Auto-closed when application stops
Timer timer = Timer.builder("batch.job.duration")
.register(registry);
// Manual cleanup for dynamically created meters
public void removeCustomerMetrics(String customerId) {
registry.remove(MeterIdRequest.of("customer.requests",
Tags.of("customerId", customerId)));
}
Nanoseconds vs Milliseconds
Timers record in the base time unit of the clock. Ensure consistency:
// CORRECT: Using Duration handles unit conversion
timer.record(Duration.ofMillis(100));
// DANGEROUS: Assuming nanosecond precision
timer.record(100_000_000L); // This is actually 100ms in nanoseconds
Percentiles with Small Sample Sizes
Percentiles calculated from fewer than 1000 samples are statistically meaningless:
// Wait for sufficient data
Timer.builder("rare.operation")
.publishPercentiles(0.5, 0.95, 0.99)
.minimumExpectedSamples(1000) // Wait for enough data
.register(registry);
Wrong Aggregation Type
Summing response times across requests produces garbage. Always use histograms for percentiles:
// WRONG: Averaging latencies loses distribution information
avg(http_server_requests_seconds_sum / http_server_requests_seconds_count)
// CORRECT: Using histogram preserves distribution
histogram_quantile(0.95, rate(http_request_duration_seconds_bucket[5m]))
Quick Recap Checklist
- Micrometer gives you a vendor-neutral API for metrics
- MeterRegistry is the central hub for registering meters
- Counters track event occurrences (only increment, never decrease)
- Gauges measure values that go up and down (sampled when scraped)
- Timers measure duration with automatic percentile calculation
- Tags add dimensions but avoid high cardinality values
- @Timed annotation automates timing for method executions
- MeterBinder handles bean-dependent metrics with correct initialization order
- Enable publishPercentileHistogram() for Prometheus histogram buckets
- Secure the /actuator/prometheus endpoint with network policies
- Apply RED method to services, USE method to resources
- Test metrics in non-production first
Interview Questions
A Counter tracks how many times an event occurs and always increments by a positive value. It is ideal for measuring request counts, error tallies, and events that happen a finite number of times. Counters never decrease.
A Gauge measures a value that can go up or down at any time. Gauges reflect the current state of something, such as queue depth, cache size, or memory utilization. Unlike counters, gauges are sampled when Prometheus scrapes the endpoint, not continuously recorded.
High cardinality occurs when tag values have unbounded possibilities, like user IDs or request IDs. Each unique combination of tags creates a separate time series in Prometheus.
Prevent this by only using low-cardinality tag values such as region codes, environment names, or tier levels. For user-specific tracking, aggregate metrics using histograms or store granular data in application logs instead of metrics.
Additionally, configure maximumExpectedTags() and tagCardinalityLimit() to reject metrics exceeding cardinality thresholds.
The @Timed annotation from io.micrometer.core.annotation automatically records the execution time of annotated methods. Spring Boot's AOP infrastructure wraps method execution in timer recording logic.
The annotation supports parameters including value for the metric name, description for human-readable explanation, and longTask for enabling long task timers that track extended operations separately.
When applied to controller endpoints, all requests automatically record timing metrics with HTTP method, URI, and status code as tags.
The RED method defines three metric types for every service endpoint:
- Rate: The number of requests per second. Use
counter_totalwithrate()in Prometheus to calculate throughput. - Errors: The error rate as a percentage. Filter by 5xx status codes and divide by total requests.
- Duration: Response time distribution using percentiles. The p50 shows median latency, p95 shows 95th percentile, and p99 shows near-worst-case.
These three metrics provide sufficient observability to detect most service degradation patterns.
Never expose metrics endpoints publicly. Use multiple layers of protection:
Network policies restrict access to only the monitoring namespace using Kubernetes NetworkPolicy or cloud security groups.
Authentication adds basic auth or mTLS for service-to-service verification.
Filtering using MeterFilter.deny() removes sensitive metrics like passwords or tokens before export.
Separate service accounts with minimal permissions ensure Prometheus can scrape without broader cluster access.
publishPercentileHistogram() enables Prometheus histogram buckets for your timer. Without this, percentile calculations return NaN because Prometheus has no histogram data to compute quantiles from.
When enabled, Micrometer records measurements into configurable latency buckets. Prometheus then uses histogram_quantile() to calculate p50, p95, p99 from the bucket data. Always enable this for user-facing endpoints where percentiles matter.
MeterBinder defers meter registration until all Spring beans are initialized. This solves the circular dependency problem where metrics require beans that have not yet been created.
Without MeterBinder, injecting a ProcessingQueue bean into a component that registers gauges at construction time fails because the queue bean does not exist yet. MeterBinder receives the fully-initialized beans as method parameters, guaranteeing safe access.
A Timer automatically converts durations to the base time unit and calculates statistics like count, total time, and percentiles. Timers are purpose-built for measuring operation duration.
A DistributionSummary records arbitrary counts or sizes without time unit conversion. Use it for request sizes, response byte counts, or any non-duration measurement where percentile distributions still matter.
User IDs and session IDs have extremely high cardinality. Each unique ID creates a separate time series in Prometheus. With millions of users, you generate millions of series that overwhelm Prometheus storage and slow down queries.
Instead, aggregate users into low-cardinality groupings like userTier (free/premium/enterprise), region, or accountAge bucket. For individual user tracking, use distributed traces or application logs.
The USE method (Utilization, Saturation, Errors) applies to infrastructure resources like CPU, memory, disks, and network interfaces. It answers whether a resource is busy, how full the queue is, and whether errors are occurring.
RED method (Rate, Errors, Duration) applies to services. It measures how many requests arrive, what fraction fail, and how long they take. USE and RED together provide complete coverage: RED for services, USE for underlying resources.
CompositeMeterRegistry holds multiple delegate registries simultaneously. When you register a meter, it broadcasts to all delegates. Spring Boot auto-configures a composite containing whichever monitoring backends it detects (Prometheus, Datadog, CloudWatch, etc.).
This means a single instrumentation call populates all backends at once, allowing you to switch monitoring systems without changing application code.
Timers use monotonic clock sources when available, but virtualized environments can report negative durations when CPU scheduling anomalies occur. The guest virtual machine may be paused or throttled, causing clock jumps.
Solutions include using DistributionSummary instead of Timer for environments with unreliable clocks, or setting maximumExpectedValue() to reject outlier measurements that exceed reasonable bounds.
Track jvm.gc.pause during load tests to detect MeterRegistry bloat. Monitor the total meter count via GET /actuator/metrics/micrometer.registry if exposed.
Use MeterFilter to reject or limit high-cardinality meters at registration time rather than after they accumulate. Set cardinality limits with tagCardinalityLimit() to automatically reject problematic meters.
longTask = true enables long task timers that track duration even when requests take extended periods. Standard timers complete when the request ends, but long task timers can track operations that outlive a single request.
Use this for batch jobs, background processing, and scheduled tasks where you want to monitor in-progress duration separately from completed request metrics.
Micrometer integrates with OpenTelemetry through the micrometer-tracing-bridge-otel module. This bridge connects Micrometer metrics with OpenTelemetry traces, allowing correlation between the two signals.
When both are configured, you can link a specific trace ID from a slow request directly to the p95 latency metric that alerted you, bridging the gap between "something is slow" and "this specific operation caused it."
Each dynamically registered meter allocates objects that remain in memory indefinitely. Unlike request-scoped objects, meters live for the application lifetime and accumulate.
Common mistake: creating a new timer per incoming request keyed by unique values. This permanently allocates meter objects that are never garbage collected. Always reuse meters or implement proper cleanup via registry.remove().
Percentiles represent the distribution shape, not the average. avg(latency_sum / latency_count) produces a single number that conceals bimodal distributions, outliers, and correlation between latency and error rate.
Histograms preserve distribution shape across time. histogram_quantile(0.95, rate(latency_bucket[5m])) correctly calculates the 95th percentile from the actual distribution stored in histogram buckets.
Pull: Prometheus scrapes your /actuator/prometheus endpoint at fixed intervals. Simple to operate, scales well, but may miss short-lived events between scrapes.
Push: Your application pushes metrics to PushGateway or a collector. Better for batch jobs, ephemeral tasks, or metrics that change faster than scrape intervals. Adds complexity and potential backpressure.
Micrometer Timers always assume nanosecond precision when recording raw long values. timer.record(100_000_000L) records 100ms, not 100 seconds, because the value is interpreted as nanoseconds.
Always use Duration objects: timer.record(Duration.ofMillis(100)). This handles unit conversion correctly and makes your intent explicit, preventing the nanosecond assumption bug.
minimumExpectedSamples() tells Micrometer the minimum sample count before publishing percentile values. Below this threshold, percentiles are statistically meaningless.
Setting minimumExpectedSamples(1000) prevents publishing p95 values based on 10 samples, which would be wildly inaccurate. This helps avoid misleading alerts on low-traffic endpoints.
Further Reading
- Prometheus and Grafana Integration - Complete setup guide for the observability stack
- Spring Boot Actuator Deep Dive - Security and configuration for actuator endpoints
- Observability Engineering Guide - Broader observability concepts and signal correlation
- Metrics Monitoring and Alerting - Setting up effective alerts from your metrics
- Micrometer Documentation - Official library reference
- Prometheus Querying Basics - Learning PromQL for effective metric queries
Conclusion
Micrometer provides the foundation for meaningful observability in Spring Boot applications, but the value comes from instrumenting the right things with discipline. The metrics you collect should answer the questions you actually ask during incidents: Is the service healthy? Where is latency coming from? Which operations are failing and how often? Without clear answers to these questions in your metrics, you have data but not observability.
The most important principles to internalize: tag cardinality discipline prevents Prometheus from becoming a memory problem, percentile histograms enable meaningful latency analysis rather than misleading averages, and gauges complement counters and timers by tracking current state. Always test your metrics in non-production first, and validate that the dashboards and alerts you build actually surface the information you need during incidents.
The investment in observability infrastructure pays back during the first production incident where you can answer questions quickly rather than debate theories. Micrometer’s vendor-neutral design means your instrumentation survives monitoring system changes, making it worth doing correctly from the start.
Category
Related Posts
Spring Boot Build Tools: Maven & Gradle
Configure Maven and Gradle for Spring Boot projects—plugins, dependency management, packaging JARs and WARs, and build automation essentials.
Embedded Web Servers in Spring Boot: Tomcat, Jetty, Undertow
Configure embedded servers in Spring Boot: compare Tomcat, Jetty, and Undertow, tune thread pools, enable access logs, and switch implementations.
JUnit 5 & Jupiter: Lifecycle, Nested & Parameterized Tests
Explore JUnit 5 Jupiter features: master test lifecycle annotations, organize tests with @Nested, and parameterize tests with @CsvSource and @MethodSource.