Skip to content
Dashboard example

API reliability and latency dashboard

Compare traffic, error rate, and tail latency by endpoint while retaining enough volume to judge whether a percentile is stable.

Reviewed by the Telemetry product team on . We checked what each event and metric represents, the SQL and sample results, and the assumptions and data freshness needed to interpret them. Who reviews this page

What one row represents

One completed request before grouping by route template.

Decision this helps you make

Prioritize an endpoint for release comparison, dependency inspection, or rollback review.

Metric definitions

Put the denominator and units beside the chart

Use your own field names, but keep the same definition of one row. Check the linked schema before mapping your production events.

  • Request volume
  • Error rate
  • p50 latency
  • p95 latency
Inspect the event schema
Complete DataFusion SQL

Review the query before adapting the fields

Adapt this query to your events before using it in production. Set a time range and environment filter, decide how to handle incomplete time buckets, and require enough events for a useful comparison.

SELECT
  endpoint,
  COUNT(*) AS requests,
  ROUND(
    100.0 * SUM(CASE WHEN status = 'error' THEN 1 ELSE 0 END)
    / NULLIF(COUNT(*), 0),
    2
  ) AS error_rate_pct,
  approx_percentile_cont(latency_ms, 0.50) AS p50_ms,
  approx_percentile_cont(latency_ms, 0.95) AS p95_ms
FROM api_requests
GROUP BY endpoint
ORDER BY error_rate_pct DESC, requests DESC;
Data assumptions

Check these before publishing the result

  • Endpoint values use route templates, such as /users/:id, instead of raw URLs.
  • Latency uses one end-to-end millisecond definition across producers.
  • Operational charts omit the newest incomplete time bucket.
Query review

Check the query and metric definitions

  • Set a start and end time before using this query in production.
  • Keep request volume beside every percentile and rate.
  • Segment by release and error type after detecting a regression.

What the results mean

What this result cannot prove by itself

Always show request count beside a percentile. A high p95 from a tiny group is not equivalent to broad customer impact.

Related dashboard examples

Validate the definition with known data

Test the query with known successes, failures, missing values, duplicates, and values at its limits. Check the results before using it in a production dashboard or alert.