Telemetry
Copy-paste instrumentation template

Background Job Failure Monitor

Capture queue, cron, import, billing sync, and webhook job health from the first run through retries and failures.

Reviewed by the Telemetry product team on . Event names, recommended fields, analysis questions, and privacy boundaries. Review standards and ownership

Questions this unlocks
  • Which jobs are failing repeatedly?
  • Which queues are falling behind?
  • Which job names have the worst p95 duration?
Template evidence path

Background Job Failure Monitor: implementation to decision

Treat the prompt as an implementation brief. The useful artifact is not copied code alone, but a reviewed event contract that produces a trustworthy answer.

  1. 1

    Select the boundary

    Instrument the point where job_started becomes final.

  2. 2

    Create the contract

    Start with job_started, job_completed, job_failed and keep every field typed, bounded, and privacy-reviewed.

  3. 3

    Run a fixture

    Exercise known success, failure, retry, and empty-result cases before relying on aggregate results.

  4. 4

    Answer the question

    Which jobs are failing repeatedly?

Template versus use case

This page is the implementation brief

Copy this template when the measurement goal is already clear. Use the matching use-case guide to review event boundaries, success definitions, and the decisions the resulting SQL should support.

Read Background Job Monitoring

Template

Paste this into your coding agent

Replace YOUR_API_KEY, run the flow locally, then verify the generated events and dashboards.

job-failure-monitor

Background Job Failure Monitor

text
Instrument background jobs with Telemetry.

Use /skill.md and this Telemetry API key: YOUR_API_KEY

Log job_started, job_completed, job_failed, job_retried, and dead_letter_created with:
job_name, queue_name, attempt, status, duration_ms, scheduled_at, started_at, completed_at, item_count, retry_count, worker_name, and error_type.

Create queries and a dashboard for throughput by job, failures by job, p95 duration, retry volume, stalled jobs, and newest dead-letter events.

Do not log raw job payloads, credentials, webhook bodies, or customer content.

Events to capture

job_startedjob_completedjob_failedjob_retrieddead_letter_created

Verification checklist

What a complete instrumentation pass leaves behind

Events

Synthetic events reach the intended table with stable names and field types.

Queries

The first SQL queries return plausible rows with an explicit time window.

Views

A dashboard uses the real fields and includes enough context to explain a change.

Safety

Prompts, bodies, credentials, signatures, and private content were checked for redaction.

Event schema starting points

Review the row grain, emit boundary, required types, privacy classes, example payload, and validation checklist before adapting a query or snippet to production.

Related product capability

Continue this workflow in Alerts

Promote the reviewed reliability query into an owned threshold and response workflow.

Related SQL recipes

Answer the next question with SQL

Run the query against the structured fields from this workflow, inspect the example result, and turn a useful answer into a dashboard or alert.

Browse all recipes
Complete collectionsBackground jobs SQL

More templates