Monitoring
This guide introduces how to observe an Axual Platform: the logs, metrics and traces its components emit, and the stack that collects them. Axual does not ship that stack, so this group covers what the platform produces and how to point an existing Prometheus, Grafana, log and tracing setup at it. The end-user metrics inside Self-Service are a different subject, covered in Client Metrics in the Self-Service Guide.
Type |
Reference |
Goal |
Find how to collect and read one of the three signals the platform emits, and how to alert on them. |
Audience |
Platform Operators, working with whoever runs the organisation’s central monitoring stack. |
When to use |
Stage 7 of The installation order. It can run alongside stages 3 to 5, and belongs before anyone relies on the platform. |
Observability means understanding a running system from the outside, asking questions of it without knowing its internals, so a problem nobody anticipated can still be diagnosed. That depends on the system emitting the signals to ask those questions of: logs, metrics and traces, the three pillars of observability.
The platform emits all three. The pages in this section cover them in the order they are usually needed:
-
Installing the Monitoring Stack covers standing up or connecting the collection stack itself.
-
How to Configure Logging covers setting a component’s log level and pattern, and which components set theirs elsewhere.
-
How to Ship Logs to a Central Stack covers emitting JSON so a collector can index the fields.
-
How to Search Logs in Kibana covers querying the stack once logs arrive in it.
-
How to Alert on a Log Pattern in Kibana covers turning a filter into a Watcher alert.
-
Monitoring & Metrics covers what the components expose, the Prometheus resources they ship, and the Grafana dashboards Axual publishes.
-
How to Enable Distributed Tracing covers enabling OpenTelemetry traces and exporting them to a backend.
-
How to Import the Axual Grafana Dashboards covers getting the published dashboards and alerting rules into Grafana and Prometheus.
Collecting the signals is only half of it. Alerts fire on metrics, and Acting on Alerts covers what to do when one fires.