Observability
Metrics
Prometheus metrics are served on the admin listener (:9090/metrics by default;
see metrics.enabled and metrics.path). They are never served on the public port.
Every metric has the nauthera_ prefix:
| Metric | Type | What it tells you |
|---|---|---|
nauthera_http_requests_total | counter | Public HTTP requests by operation and status |
nauthera_http_request_duration_seconds | histogram | Public HTTP latency |
nauthera_http_in_flight_requests | gauge | Requests being served now |
nauthera_http_cache_total | counter | Per-operation response-cache hits and misses |
nauthera_grpc_handled_total | counter | AdminService calls by method and code |
nauthera_grpc_handling_seconds | histogram | AdminService latency |
nauthera_grpc_in_flight_requests | gauge | AdminService calls in flight |
nauthera_cache_operations_total | counter | Redis cache operations |
nauthera_cache_operation_seconds | histogram | Redis cache latency |
nauthera_auth_logins_total | counter | Sign-in outcomes |
nauthera_auth_account_lockouts_total | counter | Accounts soft-locked |
nauthera_audit_events_total | counter | Audit events recorded |
nauthera_audit_events_dropped_total | counter | Audit events dropped on overflow |
nauthera_audit_buffer_depth | gauge | Audit queue depth |
nauthera_audit_flushes_total, nauthera_audit_flush_duration_seconds | counter, histogram | Audit batch writes |
nauthera_build_info | gauge | Version, commit and build date |
With the Helm chart, metrics.serviceMonitor and metrics.prometheusRule create a
ServiceMonitor and alerting rules for the Prometheus Operator.
Tracing
OpenTelemetry tracing is opt-in and exports over OTLP/gRPC:
tracing:
enabled: true
endpoint: tempo:4317
insecure: true # no TLS to the collector
sample_ratio: 1.0 # head-based sampling, 0.0–1.0
environment: production # deployment.environment resource attributeThe public HTTP handler, the gRPC server and cache calls are instrumented.
Logs
Structured logs come from zap, as logging.format: json (the default) or console.
When tracing is on, every log line written inside a request carries trace_id and
span_id, so you can go from a log line to its trace. DSNs, driver errors and secrets
are logged on the server and never returned to clients.
Health
| Probe | Path / mechanism |
|---|---|
| Liveness | GET /healthz on :9090 |
| Readiness | GET /readyz on :9090. Fails while the database is unavailable and during shutdown. |
| Startup | GET /startupz on :9090 |
| gRPC | grpc.health.v1 on :9091, for Kubernetes' native gRPC probe |
| Inside the image | nauthera-server healthcheck --url http://127.0.0.1:9090/readyz exits 0 or 1 |
The Redis cache is checked as a soft dependency: if it is down, the server reports degraded, but the replica is not taken out of service.
Dashboards
The Docker Compose stack in the repository provisions Grafana with:
- a server overview dashboard backed by Prometheus and Tempo, which can jump from a metric exemplar to its trace, and
- an audit-log explorer (
/d/nauthera-audit) that readsaudit_eventsfrom PostgreSQL.
Prometheus in the same stack loads a set of alerting rules from
deploy/compose/alerts.yml.