Skip to main content

Cluster monitoring

A Feldera instance regularly monitors its three primary components (API server, compiler server, and runner) and stores these cluster monitor events in the internal database. Only a limited number of cluster monitor events are retained in the database, notably at most 1000 and with a time limit of 72 hours (whichever comes first). With this, it is possible to access both the latest health check of the cluster and its health in the recent past. The events are accessible through the API.

The resources monitoring feature can be deactivated by setting in the Helm chart disableClusterMonitorResources to true.

API usage​

The cluster monitor events can be retrieved via two endpoints:

  • GET /v0/cluster/events: retrieves all cluster monitor events stored in the database, sorted from the latest to the earliest. It returns only the status fields to limit its response size, and further individual event details can be retrieved via the individual endpoint.

  • GET /v0/cluster/events/[latest|event-id]?selector=[all|status]: retrieves the details of the latest recorded event, or that of a specific event. Specify ?selector=all to retrieve more detailed status information for each service including a human-readable description of the status reported by the services themselves and of the Kubernetes resources that back them.

  • GET /v0/cluster_healthz: reports the health of the cluster derived from the latest event. It answers 200 when every service is healthy and the data behind the report is fresh, and 503 otherwise.

Stale monitoring data​

The monitor is the only writer of these events. In the enterprise edition it runs within the runner process, so when the runner dies the newest event keeps describing a cluster that no longer exists. Because a service cannot report its own death, the reader checks how old that event is instead.

The monitor writes at least every 10 minutes. Once the latest event is older than three such intervals, that is 30 minutes, it carries stale: true. Only latest reports the field, since older events are old by design. The all_healthy field of an event keeps describing that event, so read stale alongside it.

GET /v0/cluster_healthz answers the question "is the cluster healthy now", so there staleness sets all_healthy to false and the response code to 503: monitoring that has died cannot vouch for anything. Its response then carries the last recorded statuses rather than a description of the cluster now.

This affects retries. When a request fails with 502, the Rust and Python clients call this endpoint to decide what to do next. They retry immediately if the cluster is healthy, and wait a fixed pause if it is not, 90 seconds by default in Python. Stale data now gives the second answer, so retries slow down even while every service is serving. Restart the monitor to clear this, or shorten the pause in the client's retry settings.

Examples​

All events​

Request

curl -X GET 'http://127.0.0.1:8080/v0/cluster/events' | jq

Response

[
{
"id": "019afe45-ec1f-7de0-9cd1-3a6a4350b5e9",
"recorded_at": "2025-12-08T14:03:06.655736Z",
"all_healthy": true,
"api_status": "Healthy",
"compiler_status": "Healthy",
"runner_status": "Healthy"
},
{
"id": "019afe45-850c-7c82-8271-96a81e843dea",
"recorded_at": "2025-12-08T14:02:40.268247Z",
"all_healthy": true,
"api_status": "Healthy",
"compiler_status": "Healthy",
"runner_status": "Healthy"
},
(...)
]

Latest event​

Request

curl -X GET 'http://127.0.0.1:8080/v0/cluster/events/latest' | jq

Response

{
"id": "019afe45-ec1f-7de0-9cd1-3a6a4350b5e9",
"recorded_at": "2025-12-08T14:03:06.655736Z",
"stale": false,
"all_healthy": true,
"api_status": "Healthy",
"compiler_status": "Healthy",
"runner_status": "Healthy"
}

Specific event with all details​

Request

curl -X GET 'http://127.0.0.1:8080/v0/cluster/events/019afe45-ec1f-7de0-9cd1-3a6a4350b5e9?selector=all' | jq

Response

{
"id": "019afe45-ec1f-7de0-9cd1-3a6a4350b5e9",
"recorded_at": "2025-12-08T14:03:06.655736Z",
"all_healthy": true,
"api_status": "Healthy",
"api_self_info": "(...)",
"api_resources_info": "(...)",
"compiler_status": "Healthy",
"compiler_self_info": "(...)",
"compiler_resources_info": "(...)",
"runner_status": "Healthy",
"runner_self_info": "(...)",
"runner_resources_info": "(...)"
}

... with "(...)" representing human-readable status explanations (omitted for brevity).