Skip to content
Dashboard example

Background job reliability dashboard

Separate final job outcomes from attempts, then compare retries, queue wait, and duration by stable job name.

Reviewed by the Telemetry product team on . We checked what each event and metric represents, the SQL and sample results, and the assumptions and data freshness needed to interpret them. Who reviews this page

What one row represents

One terminal outcome per logical job.

Decision this helps you make

Choose whether to inspect capacity, retry policy, job code, or a dependency.

Metric definitions

Put the denominator and units beside the chart

Use your own field names, but keep the same definition of one row. Check the linked schema before mapping your production events.

  • Completed jobs
  • Terminal failure rate
  • Retry rate
  • p95 duration
Inspect the event schema
Complete DataFusion SQL

Review the query before adapting the fields

Adapt this query to your events before using it in production. Set a time range and environment filter, decide how to handle incomplete time buckets, and require enough events for a useful comparison.

SELECT
  job_name,
  COUNT(*) AS completed_jobs,
  ROUND(
    100.0 * SUM(CASE WHEN status = 'error' THEN 1 ELSE 0 END)
    / NULLIF(COUNT(*), 0),
    2
  ) AS terminal_failure_rate_pct,
  ROUND(AVG(attempt_count), 2) AS average_attempts,
  approx_percentile_cont(duration_ms, 0.95) AS p95_duration_ms
FROM job_runs
GROUP BY job_name
ORDER BY terminal_failure_rate_pct DESC, completed_jobs DESC;
Data assumptions

Check these before publishing the result

  • A stable job identifier connects attempts without exposing job payloads.
  • Attempt count includes the terminal attempt.
  • Queue wait and execution duration use separate millisecond fields.
Query review

Check the query and metric definitions

  • Use attempt-level data only when the denominator is explicitly attempts.
  • Add queue wait percentiles beside execution duration.
  • Use start and final lifecycle events to detect vanished workers.

What the results mean

What this result cannot prove by itself

Count one final outcome per job. Attempt-level logs inflate volume and can make a recovered retry look like a failed customer workflow.

Related dashboard examples

Validate the definition with known data

Test the query with known successes, failures, missing values, duplicates, and values at its limits. Check the results before using it in a production dashboard or alert.