Skip to content
Telemetry
For platform teams investigating restarts, readiness loss, and workload instability

Kubernetes reliability monitoring with SQL

Query Kubernetes workload changes alongside application events. See whether restarts are part of a planned rollout and whether users are seeing failures.

Reviewed by the Telemetry product team on . We checked the event fields, suggested queries, and data to exclude. Who reviews this page

Why this works
  • Group by workload ownership instead of short-lived pod names.
  • Keep restart transitions, cumulative counters, and readiness samples as distinct signals.
  • Compare cluster problems with application releases and failed user requests.
How to test this use case

Measure Kubernetes reliability monitoring with SQL and check the results

To measure kubernetes reliability monitoring with sql, choose one workflow and its owner. Define the events, test them with known inputs, and write a query that answers a specific question.

  1. 1

    Choose when to log

    Choose the clusters, namespaces, and customer-facing workloads in scope.

  2. 2

    Capture the outcome

    Begin with kubernetes_workload_sampled, kubernetes_container_restarted, kubernetes_rollout_observed and document the grain of each event.

  3. 3

    Check the stored rows

    Review sustained thresholds with the workload owner before paging.

  4. 4

    Make the decision

    Which workloads restart repeatedly instead of only during rollout?

Agent prompt

Paste this into your coding agent

Replace YOUR_API_KEY after signup, then ask the agent to run the product flow and verify the first events.

agent prompt

Kubernetes reliability monitoring with SQL setup prompt

text
Instrument Kubernetes workload reliability with Telemetry.

Use /skill.md and this Telemetry API key: YOUR_API_KEY

Emit controlled workload events with cluster, namespace, workload, pod, event_name, restart_count, ready, release, and environment. Add a bounded reason category only when it is already available and approved.

Create SQL and a dashboard for restart transitions, readiness loss, rollout changes, and related application errors by workload. Alert only on repeated restarts plus sustained readiness or product impact.

Do not send Kubernetes secrets, environment variable values, full manifests, raw logs, container arguments, or customer payloads.

Setup steps

  1. 1Choose the clusters, namespaces, and customer-facing workloads in scope.
  2. 2Emit restart, readiness, rollout, and workload events with a fixed set of status and reason values.
  3. 3Exercise a safe synthetic restart and a planned replacement.
  4. 4Review sustained thresholds with the workload owner before paging.

Events to capture

kubernetes_workload_sampledkubernetes_container_restartedkubernetes_rollout_observedservice_request_completed

Questions you can answer

  • Which workloads restart repeatedly instead of only during rollout?
  • Where does readiness loss coincide with failed user requests or jobs?
  • Did instability begin with a release, node pool, or cluster change?

Example event schemas

Check what each event records, when to send it, and which field types it needs. Review the example payload and privacy checklist before using it in production.

Use these queries in Telemetry

Learn about Alerts

Add a threshold and recipients to your reliability query to get alerts.

Related SQL recipes

More SQL recipes

Run the query using this workflow's event fields and check the example result. Save the result to a dashboard or set up an alert.

Browse all recipes

Next step

Create the API key your agent will use

The free plan is enough to run the prompt, send test events, and review the first dashboard.

Related pages