Record and query rag_retrieval_evaluated
Document what each rag_retrieval_evaluated row represents and which service sends it. Send test events to check the fields, then verify that the query answers your question.
- 1
Outcome becomes final
RAG evaluation pipeline emits only after the retrieval and answer evaluators finish or the candidate becomes ineligible.
- 2
Choose the fields
10 required fields preserve the declared grain: One evaluator result per eligible query, retrieval version, evaluator, and candidate run.
- 3
Check the test event
Check types, UTC time, alternate outcomes, idempotency, and every pseudonymous or review-classified field.
- 4
Query the result
Which retrieval and prompt versions produce relevant context and grounded accepted answers?
Grain
One evaluator result per eligible query, retrieval version, evaluator, and candidate run.
Owner
RAG evaluation pipeline
Emit when
After the retrieval and answer evaluators finish or the candidate becomes ineligible.
Field contract
Field types and data to exclude
Keep field names and types stable once production queries depend on them. Document optional fields and add them only when they answer a specific question.
| Field | Type | Required | Privacy | Meaning |
|---|---|---|---|---|
| timestamp_utc | timestamp | yes | non-sensitive | UTC time when the operation finishes. |
| event_id | string | yes | non-sensitive | Stable unique identifier used for deduplication. |
| release | string | yes | non-sensitive | Application or service version that emitted the event. |
| account_id | string | yes | pseudonymous | Stable internal account identifier, never an email or name. |
| evaluation_id | string | yes | pseudonymous | Stable identifier for the evaluation case. |
| retrieval_version | string | yes | non-sensitive | Version of the index, retriever, and ranking contract. |
| evaluator_version | string | yes | non-sensitive | Versioned evaluator and rubric identifier. |
| retrieval_score | number | yes | non-sensitive | Normalized retrieval relevance score. |
| grounded | boolean | yes | non-sensitive | Whether the reviewed answer met the grounding rule. |
| accepted | boolean | no | non-sensitive | Reviewed downstream acceptance when available. |
| duration_ms | number | yes | non-sensitive | End-to-end evaluated workflow duration. |
Synthetic JSON event
{
"timestamp_utc": "2026-07-29T10:12:43Z",
"event_id": "evt_rag_eval_01",
"account_id": "acct_8f31",
"release": "2026.07.3",
"evaluation_id": "eval_72ab",
"retrieval_version": "hybrid_v4",
"evaluator_version": "grounding_rubric_v2",
"retrieval_score": 0.86,
"grounded": true,
"accepted": true,
"duration_ms": 1280
}Privacy review
Review identifiers before ingestion
This example uses synthetic identifiers. Pseudonymous values can still be personal data, and review fields can expose business or provider context. Apply your own consent, retention, access, residency, and deletion requirements.
account_id: pseudonymousevaluation_id: pseudonymous
Validation checklist
Test the schema before building a dashboard
- Send one known rag_retrieval_evaluated fixture after the documented outcome boundary.
- Verify all 10 required fields arrive with the documented types.
- Retry the same event identifier and confirm the chosen deduplication behavior.
- Send a controlled failure or alternate outcome when the workflow supports one.
- Run the related SQL over a fixed window and reconcile the result to the fixture.
Common mistakes
Record one result per row
- Emitting rag_retrieval_evaluated before rag evaluation pipeline knows the final outcome.
- Mixing different kinds of results in one table, which makes counts and rates ambiguous.
- Replacing controlled categories with raw URLs, payloads, prompts, or error text.
- Changing a field type in place after saved queries and dashboards depend on it.
- Adding identifiers without a documented investigation, access, and retention need.
Use the contract
Query the event and set up monitoring
Related contracts
Send a test event before production traffic
Create a free API key, send the synthetic event, and inspect the inferred table before connecting a live workflow.