Metrics

Observability is crucial when operating a SaaS system because it’s not possible to debug it live. Alongside structured logging and distributed tracing, metrics are one of the pillars of observability.

Microbus supports both push and pull models:

  • To push metrics to an OpenTelemetry collector, set the OTEL_EXPORTER_OTLP_METRICS_ENDPOINT or OTEL_EXPORTER_OTLP_ENDPOINT environment variable appropriately
  • Alternatively, set the MICROBUS_PROMETHEUS_EXPORTER environment variable and configure Prometheus to scrape metrics from the metrics core microservice

When pushing over OTLP, all exporters in one executable that target the same collector share a single connection across the traces, metrics and logs signals. svc.MeterProvider() returns the underlying OpenTelemetry meter provider — a no-op provider (never nil) when metrics are disabled — so application code can instrument third-party libraries against the same metrics pipeline and resource as the framework.

Standard Metrics

By default, all microservices produce a standard set of metrics:

  • The microservice’s uptime
  • Histogram of the execution time of callbacks such as OnStartup, OnShutdown, tickers, etc.
  • Histogram of the processing time of incoming requests
  • Histogram of the size of the response to incoming requests
  • Count of outgoing requests, along with the error and status code
  • Histogram of time to receive an acknowledgement from a downstream microservice
  • Count of log messages recorded
  • Count of distributed cache operations, including hit and miss stats
  • Memory usage of the distributed cache

Server Histogram Labels

The two server-side histograms — microbus_server_request_duration_seconds and microbus_server_response_body_bytes — carry name, port, method, code, type and error, on top of the service and id labels every Microbus instrument gets.

The endpoint’s route and its canonical form are deliberately not labels. Together, service, port and name already identify an endpoint completely, so a path label would only restate them — while multiplying the number of series a Prometheus instance has to hold. Group by those three when writing queries against these histograms.

Third-Party Instrument Names

Libraries the framework embeds export under their own prefixes rather than microbus_, and their names follow that library’s conventions:

  • The SQL layer exports sequel_*, including the connection-pool series sequel_pool_waits_total and sequel_pool_wait_duration_seconds_total. Both are counters, so a rate query is rate(sequel_pool_waits_total[5m]).
  • The embedded workflow engine exports dwarf_*. A few of its gauges are read from the shared database and reported identically by every replica, so they aggregate with max rather than sumProduction Readiness names which.

Custom Metrics

Custom metrics are defined using the Connector’s DescribeCounter, DescribeGauge or DescribeHistogram. Metrics are incremented or observed using IncrementCounter, RecordGauge or RecordHistogram, depending on their type.

The coding agent can assist in the definition of metrics.

Hey Claude, create a metric that counts the number of likes per post ID.

IncrementCounterLikes (or something similar) will be created by the coding agent in intermediate.go.

func (svc *Intermediate) IncrementCounterLikes(ctx context.Context, num int, postId string) error {
	// ...
}

It can then be used to count the number of likes from anywhere in the microservice’s code.

func (svc *Service) Like(ctx context.Context, postId string) error {
	// ...

	err := svc.IncrementCounterLikes(ctx, 1, postId)
	if err != nil {
		return errors.Trace(err)
	}
	return nil
}

See Also

  • Calculator example — pairs a counter (UsedOperators, incremented inline in every arithmetic call) with an observable gauge (SumOperations, populated by an OnObserveSumOperations JIT observer reading per-operator atomic.Int64 accumulators). The canonical “running total” pattern.