Telemetry
For AI products with model usage, token spend, and latency risk

OpenAI Cost Monitoring

Monitor OpenAI and LLM spend by model, feature, customer, latency, error rate, and outcome with SQL dashboards and budget alerts.

Reviewed by the Telemetry product team on . Event contract, recommended analysis, and privacy boundaries. Review standards and ownership

Why this works
  • Model spend by feature and customer account.
  • Latency and failure trends before they become support tickets.
  • Acceptance or save rate for generated outputs where the app exposes it.
Measurement path

Turn token usage into reviewable unit economics

Provider usage becomes more useful when a versioned estimate is joined to retries and the downstream outcome the application actually values.

  1. 1

    Model request

    Capture provider, model, feature, tokens, latency, cache, and retry context.

  2. 2

    Cost estimate

    Apply a dated price source and retain its version beside the estimate.

  3. 3

    Product outcome

    Connect the request to accepted, resolved, saved, escalated, or discarded work.

  4. 4

    Unit economics

    Compare cost per reviewed outcome, not token or request volume alone.

Use case versus template

This page explains what to measure and why

Use the use-case guide to choose outcomes, event boundaries, and analysis questions. Open the matching template when you are ready for a shorter copy-paste implementation brief.

Open LLM Cost Tracker

Agent prompt

Paste this into your coding agent

Replace YOUR_API_KEY after signup, then ask the agent to run the product flow and verify the first events.

agent prompt

OpenAI Cost Monitoring setup prompt

text
Instrument this project with Telemetry so we can understand OpenAI and LLM usage.

Use /skill.md and this Telemetry API key: YOUR_API_KEY

Please log:
1. Every model request with model, provider, route, feature, input_tokens, output_tokens, total_tokens, estimated_cost_usd, latency_ms, status, and error_type when relevant.
2. Tool calls made by the agent or assistant, including tool_name, status, latency_ms, and result_category.
3. User-facing AI workflow outcomes, including feature, status, retry_count, and whether the user accepted, copied, saved, or discarded the result.
4. A dashboard with daily cost, cost by feature, failures by model, p95 latency, and accepted output rate.

Keep prompts and raw completions out of telemetry unless I explicitly approve storing them.

Setup steps

  1. 1Log each model request without storing raw prompts by default.
  2. 2Record tokens, estimated cost, latency, provider, model, and workflow.
  3. 3Connect product outcome events like saved, copied, accepted, or retried.
  4. 4Create dashboards for margin, reliability, and user value.

Events to capture

llm_request_completedllm_request_failedassistant_tool_calledai_output_acceptedai_output_discardedai_feature_retained

Questions unlocked

  • Which AI features cost the most per activated user?
  • Which model has the best accepted-output rate per dollar?
  • Where are retries or timeouts damaging conversion?

LLM cost and unit economics

Normalize provider usage, then measure cost per useful outcome

Token spend becomes actionable when it stays connected to the feature, account, retry path, and reviewed user outcome that created it. Use the calculator to test assumptions, then run the same analysis against the multi-table SQL Lab.

Illustrative monthly cost model

The price inputs are examples, not current provider prices. Replace them with the rates and billable token categories from your provider agreement.

Provider attempts

105,000

Estimated monthly cost

$304.50

Estimated retry cost

$14.50

Accepted outputs

60,000

Cost per accepted output

$0.0051

This browser-only estimate is not sent to Telemetry and is not a provider invoice.

A provider-neutral event contract

Preserve the provider response fields you need, but normalize the analysis surface so model and pricing changes do not require a new dashboard.

Normalized fieldsWhy they belong together
provider, model, featureAttribute usage to the provider and product workflow.
input_tokens, output_tokensPreserve the provider-reported usage components.
cached_input_tokens, reasoning_tokensKeep optional billable categories separate when available.
estimated_cost_usd, pricing_versionMake the analytical estimate reproducible after prices change.
retry_count, cache_hit, latency_msExplain cost and reliability changes inside the execution path.
accepted, saved, discarded, human_handoffConnect provider consumption to a reviewed product outcome.

Keep the billing claim bounded

Use provider invoices as billing truth. Event-level cost is an analytical estimate for attribution, product decisions, margin investigation, and anomaly detection.

Read the OpenAI API cost tracking implementation guide

Event schema starting points

Review the row grain, emit boundary, required types, privacy classes, example payload, and validation checklist before adapting a query or snippet to production.

Related product capability

Continue this workflow in AI agent monitoring

Connect agent runs, tool use, model cost, quality, and product outcomes with reviewable SQL.

Related SQL recipes

Answer the next question with SQL

Run the query against the structured fields from this workflow, inspect the example result, and turn a useful answer into a dashboard or alert.

Browse all recipes

Next step

Create the API key your agent will use

The free plan is enough to run the prompt, send test events, and review the first dashboard.

Related pages