Skip to content
Telemetry
For teams running queues, cron jobs, imports, and async workers

Background job monitoring

Use structured events to see job throughput, retries, failures, dead letters, queue health, and p95 duration.

Reviewed by the Telemetry product team on . We checked the event fields, suggested queries, and data to exclude. Who reviews this page

Why this works
  • Find repeated failures by job name, queue, worker, and customer account.
  • Find stalled jobs and retry loops, and see which customers they affect.
  • Give support and engineering one queryable history of async work.
How to test this use case

Measure Background job monitoring and check the results

To measure background job monitoring, choose one workflow and its owner. Define the events, test them with known inputs, and write a query that answers a specific question.

  1. 1

    Choose when to log

    Log job_started, job_completed, job_failed, and job_retried.

  2. 2

    Capture the outcome

    Begin with job_started, job_completed, job_failed and document the grain of each event.

  3. 3

    Check the stored rows

    Add alerts for stalled jobs, repeated failures, and dead-letter creation.

  4. 4

    Make the decision

    Which jobs fail or retry most often?

Use case versus template

Choose what to measure

Use this guide to choose what to measure and when to log it. For a shorter setup prompt, open the matching template.

Open Background job failure monitor

Agent prompt

Paste this into your coding agent

Replace YOUR_API_KEY after signup, then ask the agent to run the product flow and verify the first events.

agent prompt

Background job monitoring setup prompt

text
Instrument background jobs and async workers with Telemetry.

Use /skill.md and this Telemetry API key: YOUR_API_KEY

Please log job_started, job_completed, and job_failed events with job_name, queue_name, attempt, status, duration_ms, scheduled_at, started_at, completed_at, item_count, retry_count, and error_type.

Create a dashboard showing throughput by job, failures by job, p95 duration, retry volume, oldest pending job, and dead-letter events. Add alerts for stalled jobs and repeated failures.

Setup steps

  1. 1Log job_started, job_completed, job_failed, and job_retried.
  2. 2Include queue, job name, attempt, duration, item count, and error type.
  3. 3Build a dashboard for throughput, failures, retries, and p95 duration.
  4. 4Add alerts for stalled jobs, repeated failures, and dead-letter creation.

Events to capture

job_startedjob_completedjob_failedjob_retrieddead_letter_createdcron_tick_failed

Questions you can answer

  • Which jobs fail or retry most often?
  • Which queues are falling behind?
  • Which customers are affected by failed background work?

Example event schemas

Check what each event records, when to send it, and which field types it needs. Review the example payload and privacy checklist before using it in production.

Use these queries in Telemetry

Learn about Alerts

Add a threshold and recipients to your reliability query to get alerts.

Related SQL recipes

More SQL recipes

Run the query using this workflow's event fields and check the example result. Save the result to a dashboard or set up an alert.

Browse all recipes
Recipe collectionsBackground jobs SQL

Customer evidence

Related customer stories

Next step

Create the API key your agent will use

The free plan is enough to run the prompt, send test events, and review the first dashboard.

Related pages