Event schema
Fields the query expects
| Field | Type | Why it exists |
|---|---|---|
| timestamp_utc | Timestamp | Request completion time in UTC. |
| feature_flag | Utf8 | Controlled feature-flag name. |
| cohort | Utf8 | Rollout or control assignment. |
| release | Utf8 | Application release identifier. |
| status | Utf8 | Terminal success or failed status. |
| latency_ms | Int64 | End-to-end request latency in milliseconds. |
| environment | Utf8 | Deployment environment. |
Copy the query
SELECT
feature_flag,
cohort,
release,
COUNT(*) AS requests,
SUM(CASE WHEN status = 'failed' THEN 1 ELSE 0 END) AS errors,
100.0 * SUM(CASE WHEN status = 'failed' THEN 1 ELSE 0 END)
/ NULLIF(COUNT(*), 0) AS error_rate_pct,
AVG(latency_ms) AS avg_latency_ms
FROM feature_rollout_events
WHERE timestamp_utc >= now() - INTERVAL '24 hours'
AND environment = 'production'
GROUP BY feature_flag, cohort, release
ORDER BY error_rate_pct DESC, cohort;This read-only query is planned and executed against an empty typed table with Apache DataFusion 45.2.0. We review the synthetic sample output separately. Check field types, thresholds, and counting rules against your own data. Read the testing methodology.
Query result
Rollout error rate by cohort
The synthetic rollout cohort has a 20% failure rate while the control has none.
| feature_flag | cohort | release | requests | errors | error_rate_pct | avg_latency_ms |
|---|---|---|---|---|---|---|
| new-checkout | rollout | api-88 | 5 | 1 | 20 | 300 |
| new-checkout | control | api-88 | 5 | 0 | 0 | 200 |
Synthetic example output. Run the query against your own event schema and thresholds before using it for operational decisions.
Reproduce the example
Download the sample data
The JSON bundle includes the event schema with field types, reproducible input rows, exact SQL, expected output, review notes, and engine version. The CSV contains the displayed result.
How the SQL works
- 1Record the feature flag, cohort, and release so someone else can repeat the rollout comparison.
- 2Request volume beside error rate exposes comparisons that are too small for a confident decision.
- 3Latency remains in the same result because a rollout can regress experience without increasing errors.
Edge cases to check
- Assignment must be stable enough that the same account does not switch cohorts inside the comparison window.
- Compare groups with similar traffic and customer mixes. If they differ, split the results by an approved field such as plan or region.
- Use confidence intervals and a longer window for irreversible product conclusions.
Recommended dashboard
- Bars: error_rate_pct by cohort
- Trend: requests, errors, and latency by rollout percentage
- Table: feature flags by release and minimum sample status
Alert guidance
Pause or roll back only after the rollout meets a reviewed minimum volume and exceeds its approved error or latency guardrail.
Read alert setupSet up the events this query needs
Related instrumentation and guides
Define the source data
Event schemas for this analysis
Continue the analysis
Calculate API error rate by route
Rank API routes by 5xx error rate. Exclude routes with fewer than 20 requests so one failure does not dominate the results.
Open recipeCalculate p50, p95, and p99 API latency
Compare median and tail latency by endpoint with DataFusion-compatible percentile SQL.
Open recipeCalculate API error-budget burn rate
Turn hourly request failures into an SLO burn-rate series that shows how quickly the allowed error budget is being consumed.
Open recipeRun it on your events
Create a table, adapt the fields, and save the result
Start free, send structured events, and use the query result as a chart, shared dashboard widget, or alert input.