Skip to content
r/adithya.
All projects
Search engineering · Evaluation framework2 min read

RAG-Bench — A Reproducible Retrieval Evaluation Framework

A framework for benchmarking RAG retrieval pipelines: keyword, dense, hybrid, and reranked search compared on the same questions, with statistical significance and latency.

FiQA · Search quality+21% relative
Hybrid search35.3%
With reranking42.7%

Relevant documents found in the first five results, averaged over 648 questions.

Recall@5: 0.353 → 0.427. Reordering the search results improves retrieval, at the cost of additional query time.

Inspect the benchmark

The same 648 FiQA questions, documents, and search settings. The second run adds a BGE reranker over the top 50 candidates. Each result file includes the resolved configuration and per-question scores.

Hybrid search 0.353 Recall@5
Config YAML Results JSON
With reranking 0.427 Recall@5
Config YAML Results JSON

Shared settings: 512-token chunks, 64-token overlap, RRF k=60, seed 42, and 10,000 bootstrap resamples. Base configuration · Full results report

FiQA p95 query time: 102 ms → 14.06 s with reranking, on an 8 GB Apple Silicon laptop.

Overview

RAG-Bench is a Python framework for answering a practical question: which retrieval pipeline actually finds the right documents? It runs keyword, dense, hybrid, and reranked search on the same labeled questions, then reports quality, latency, and statistical confidence side by side.

Features

  • Config-driven experiments: YAML files choose the dataset, embedding model, chunker, fusion method, and reranker.
  • Full retrieval stack: BM25 keyword search, BGE-M3 dense retrieval on Qdrant, Reciprocal Rank Fusion and weighted fusion, and a BGE cross-encoder reranker.
  • Standard IR metrics: Recall@5/10/20, nDCG@10, and MRR@10, validated against pytrec_eval.
  • Statistical rigor: 95% bootstrap confidence intervals with 10,000 resamples, and paired significance tests.
  • Reproducible runs: every run saves its full configuration, hashes, Git commit, and per-query results.
  • Reports: JSON results, CSV summaries, Markdown tables, plots, and an analysis notebook.

How it works

  1. A YAML config defines the pipeline under test.
  2. Documents are chunked with one of four strategies and embedded, with embeddings cached in SQLite.
  3. Each query runs through retrieval, optional fusion, and optional reranking of the top 50 candidates.
  4. Chunk scores are max-pooled back to documents and scored against ground-truth relevance labels.
  5. Results are compared across configurations with confidence intervals and significance tests.

Tech stack

  • Language: Python
  • Retrieval: BM25, BGE-M3 embeddings, Qdrant, BGE cross-encoder reranker
  • Evaluation: pytrec_eval, SciPy (bootstrap and Wilcoxon tests)
  • Storage: SQLite embedding cache
  • Testing: pytest

Results

  • 26 benchmark runs across 4 datasets (SciFact, NFCorpus, FiQA, and an English MLDR subset), covering 2,071 queries and 67,654 documents.
  • On FiQA, reranking raised Recall@5 from 0.353 to 0.427, a 21% relative improvement across 648 queries (Wilcoxon p ≈ 1.6 × 10⁻¹⁰).
  • On SciFact, the best pipeline reached 0.803 Recall@5 and 0.757 nDCG@10.

Inspect the configurations and per-query results behind the FiQA comparison.

Engineering highlights

Diagnose, don't just score. Candidate-pool Recall@50 separates "the document was never retrieved" from "the reranker ranked it poorly", pointing to which stage to fix.

Fair document-level scoring. Chunk results are deduplicated and max-pooled, so a document split into many chunks can't inflate the metrics.

Fast and resumable. Embeddings are cached by model and text hash, and results are written incrementally, so interrupted runs pick up where they stopped.

Built withPythonBM25BGE-M3QdrantSQLiteSciPypytest

Write-up updated .

Vote

Your vote stays in this browser.

Let's talk about it

Ask u/adithya-bot about this post. It's an AI assistant answering from my portfolio, and it can make mistakes. This conversation is private to your visit.

0/500

Keep exploring