What one row represents
One terminal outcome per logical job.
Decision this helps you make
Choose whether to inspect capacity, retry policy, job code, or a dependency.
Metric definitions
Put the denominator and units beside the chart
Use your own field names, but keep the same definition of one row. Check the linked schema before mapping your production events.
- Completed jobs
- Terminal failure rate
- Retry rate
- p95 duration
failure rate %
Demonstration data, not a customer result or benchmark
Review the query before adapting the fields
Adapt this query to your events before using it in production. Set a time range and environment filter, decide how to handle incomplete time buckets, and require enough events for a useful comparison.
SELECT
job_name,
COUNT(*) AS completed_jobs,
ROUND(
100.0 * SUM(CASE WHEN status = 'error' THEN 1 ELSE 0 END)
/ NULLIF(COUNT(*), 0),
2
) AS terminal_failure_rate_pct,
ROUND(AVG(attempt_count), 2) AS average_attempts,
approx_percentile_cont(duration_ms, 0.95) AS p95_duration_ms
FROM job_runs
GROUP BY job_name
ORDER BY terminal_failure_rate_pct DESC, completed_jobs DESC;Check these before publishing the result
- A stable job identifier connects attempts without exposing job payloads.
- Attempt count includes the terminal attempt.
- Queue wait and execution duration use separate millisecond fields.
Check the query and metric definitions
- Use attempt-level data only when the denominator is explicitly attempts.
- Add queue wait percentiles beside execution duration.
- Use start and final lifecycle events to detect vanished workers.
What the results mean
What this result cannot prove by itself
Count one final outcome per job. Attempt-level logs inflate volume and can make a recovered retry look like a failed customer workflow.
Related dashboard examples
SaaS health overview
Did account adoption, application reliability, or represented revenue change enough to require a deeper review?
Inspect exampleAPI reliability and latency
Which endpoints have rising error rates or p95 latency, and enough requests to trust the comparison?
Inspect exampleIncident customer impact
Which accounts and operations were directly represented in a declared incident window?
Inspect exampleValidate the definition with known data
Test the query with known successes, failures, missing values, duplicates, and values at its limits. Check the results before using it in a production dashboard or alert.