Skip to main content

Check Cluster Health

Required role: read or higher.

Determine the latest cluster health via the latest cluster monitor event. Each service's unchanged_since reports the approximate time it last transitioned between healthy and unhealthy, bounded by event retention.

The cluster monitor is the only writer of these events. In the enterprise edition it runs within the runner process, so when the runner dies the newest event keeps describing a cluster that no longer exists; elsewhere the monitor is a task in the process that answers this request. Either way a service cannot report its own death, so this endpoint additionally checks how old that event is. A stale response repeats the last recorded statuses instead of describing the cluster now, and counts as unhealthy: monitoring that has died cannot vouch for anything.

Responses
200

All services healthy

Schema — OPTIONAL
all_healthy boolean

Whether every service is healthy and the monitoring data behind this report is fresh.

api object

Health of a single service derived from the retained cluster monitor events.

checked_at date-time

Timestamp of the most recent cluster monitor event.

healthy boolean

Whether the service passed its most recent health check.

message string

Human-readable report from the most recent health check.

unchanged_since date-time

Approximate time the service last transitioned between healthy and unhealthy: the timestamp of the oldest retained consecutive cluster monitor event with the same healthy conclusion. Bounded by event retention.

compiler object

Health of a single service derived from the retained cluster monitor events.

checked_at date-time

Timestamp of the most recent cluster monitor event.

healthy boolean

Whether the service passed its most recent health check.

message string

Human-readable report from the most recent health check.

unchanged_since date-time

Approximate time the service last transitioned between healthy and unhealthy: the timestamp of the oldest retained consecutive cluster monitor event with the same healthy conclusion. Bounded by event retention.

runner object

Health of a single service derived from the retained cluster monitor events.

checked_at date-time

Timestamp of the most recent cluster monitor event.

healthy boolean

Whether the service passed its most recent health check.

message string

Human-readable report from the most recent health check.

unchanged_since date-time

Approximate time the service last transitioned between healthy and unhealthy: the timestamp of the oldest retained consecutive cluster monitor event with the same healthy conclusion. Bounded by event retention.

stale boolean

Whether the cluster monitor stopped writing events, which makes the statuses below the last recorded ones rather than current ones.

stale_after_seconds int64

Age at which monitoring data counts as stale.

503

One or more services unhealthy

Schema — OPTIONAL
all_healthy boolean

Whether every service is healthy and the monitoring data behind this report is fresh.

api object

Health of a single service derived from the retained cluster monitor events.

checked_at date-time

Timestamp of the most recent cluster monitor event.

healthy boolean

Whether the service passed its most recent health check.

message string

Human-readable report from the most recent health check.

unchanged_since date-time

Approximate time the service last transitioned between healthy and unhealthy: the timestamp of the oldest retained consecutive cluster monitor event with the same healthy conclusion. Bounded by event retention.

compiler object

Health of a single service derived from the retained cluster monitor events.

checked_at date-time

Timestamp of the most recent cluster monitor event.

healthy boolean

Whether the service passed its most recent health check.

message string

Human-readable report from the most recent health check.

unchanged_since date-time

Approximate time the service last transitioned between healthy and unhealthy: the timestamp of the oldest retained consecutive cluster monitor event with the same healthy conclusion. Bounded by event retention.

runner object

Health of a single service derived from the retained cluster monitor events.

checked_at date-time

Timestamp of the most recent cluster monitor event.

healthy boolean

Whether the service passed its most recent health check.

message string

Human-readable report from the most recent health check.

unchanged_since date-time

Approximate time the service last transitioned between healthy and unhealthy: the timestamp of the oldest retained consecutive cluster monitor event with the same healthy conclusion. Bounded by event retention.

stale boolean

Whether the cluster monitor stopped writing events, which makes the statuses below the last recorded ones rather than current ones.

stale_after_seconds int64

Age at which monitoring data counts as stale.