What are we measuring?
RAG evaluation asks two separate questions: did retrieval find the necessary evidence, and did the answer convey the necessary information? Fluent answers can hide missing evidence; complete evidence does not guarantee a complete answer.
The example was prepared and run locally on Node.js v24.19.0. It uses three synthetic documents, two questions and keyword matching, without embeddings, a vector database or an LLM. It is a small evaluation fixture, not a full RAG system or a production benchmark.
Dataset and expected evidence
The release question needs two documents: build covers test/build and secrets covers the secret check. The media question needs only images. Expected document IDs and answer facts are stored in the evaluation fixtures, not read by the retrieval function.
Recall is relevant evidence retrieved divided by relevant evidence expected. Precision is relevant evidence retrieved divided by all evidence retrieved. factCoverage checks only whether expected strings appear in the answer; it cannot establish semantic correctness or hallucination rates. It can miss valid paraphrases or accept a phrase inside a negation.
Run the example
Save rag-evaluation.mjs from the source links and run node rag-evaluation.mjs using Node.js 24. No package, account, API key or network access is needed. A failed assertion terminates the command with an error.
// Synthetic documents; no LLM, API key, embedding model or external request.
import assert from 'node:assert/strict';
const corpus = [
{ id: 'build', text: 'Yayın öncesi test ve build çalıştırılır.' },
{ id: 'secrets', text: 'Yayın öncesi gizli anahtar kontrolü yapılır.' },
{ id: 'images', text: 'Blog görselleri WebP biçiminde saklanır.' },
];
const cases = [
{ id: 'release', query: 'yayın test build gizli anahtar', expected: ['build', 'secrets'], facts: ['test ve build', 'gizli anahtar'] },
{ id: 'media', query: 'blog görselleri', expected: ['images'], facts: ['WebP'] },
];
const tokens = text => new Set(text.toLocaleLowerCase('tr-TR').match(/[\p{L}\p{N}]+/gu) || []);
function retrieve(query, k) {
const terms = tokens(query);
return corpus.map(doc => ({ ...doc, score: [...tokens(doc.text)].filter(t => terms.has(t)).length }))
.filter(doc => doc.score > 0).sort((a, b) => b.score - a.score || a.id.localeCompare(b.id)).slice(0, k);
}
function evaluate(item, docs, answer) {
const hits = docs.filter(doc => item.expected.includes(doc.id)).length;
return { recall: hits / item.expected.length, precision: docs.length ? hits / docs.length : 0,
factCoverage: item.facts.filter(fact => answer.includes(fact)).length / item.facts.length };
}
for (const item of cases) {
for (const k of [1, 2]) {
const docs = retrieve(item.query, k);
// This is evidence concatenation, NOT an LLM-generated answer.
const result = evaluate(item, docs, docs.map(doc => doc.text).join(' '));
if (item.id === 'release') assert.equal(result.recall, k === 1 ? 0.5 : 1);
else assert.equal(result.recall, 1);
assert.equal(result.precision, 1);
console.log(JSON.stringify({ case: item.id, k, retrieved: docs.map(d => d.id), ...result }));
}
}
const release = cases[0];
const fullEvidence = retrieve(release.query, 2);
const incompleteAnswer = fullEvidence[0].text;
const incomplete = evaluate(release, fullEvidence, incompleteAnswer);
assert.equal(incomplete.recall, 1);
assert.equal(incomplete.factCoverage, 0.5);
console.log(JSON.stringify({ case: 'answer-omission', ...incomplete }));
assert.equal(retrieve('pasaport yenileme', 2).length, 0);
console.log('PASS: retrieval, answer omission and out-of-corpus checks');Observed results
For the release question, k=1 retrieved only build: recall 0.5, precision 1 and factCoverage 0.5. k=2 retrieved build and secrets, giving 1 for all three ratios. Both k values returned only images for the media question. These are fixture results, not overall accuracy claims.
The answer-omission case intentionally uses only the first retrieved document despite complete evidence. Recall remains 1 while factCoverage drops to 0.5, reproducing the distinction between retrieval and answer completeness. An out-of-corpus query returns no evidence; a real product still needs a separate abstention policy.
{"case":"release","k":1,"retrieved":["build"],"recall":0.5,"precision":1,"factCoverage":0.5}
{"case":"release","k":2,"retrieved":["build","secrets"],"recall":1,"precision":1,"factCoverage":1}
{"case":"answer-omission","recall":1,"precision":1,"factCoverage":0.5}Moving to a real system
Two questions cannot establish that larger k is always better. Real evaluations need varied questions, irrelevant documents, missing evidence and contradictions, alongside cost, latency and context-noise measurements. Review evidence labels and avoid tuning on the same cases used to claim success.
Do not accept real LLM answers using substring matching alone. Review grounding, abstention and domain risk separately. Traces can retain document IDs and versions, but private text and credentials must not enter public logs.
Your first evaluation record
Record question ID, corpus version, expected and retrieved evidence IDs, ranks, k, answer and separate evaluation results together. Identify the failing layer, change that layer and rerun the same cases. The linked debugging guide complements this with one hypothesis per change.
Source: Runnable evaluation file
Source: RAG research paper — conceptual background
Source: Node.js assertion documentation
