Overview
RAG-Bench is a Python framework for answering a practical question: which retrieval pipeline actually finds the right documents? It runs keyword, dense, hybrid, and reranked search on the same labeled questions, then reports quality, latency, and statistical confidence side by side.
Features
- Config-driven experiments: YAML files choose the dataset, embedding model, chunker, fusion method, and reranker.
- Full retrieval stack: BM25 keyword search, BGE-M3 dense retrieval on Qdrant, Reciprocal Rank Fusion and weighted fusion, and a BGE cross-encoder reranker.
- Standard IR metrics: Recall@5/10/20, nDCG@10, and MRR@10, validated against pytrec_eval.
- Statistical rigor: 95% bootstrap confidence intervals with 10,000 resamples, and paired significance tests.
- Reproducible runs: every run saves its full configuration, hashes, Git commit, and per-query results.
- Reports: JSON results, CSV summaries, Markdown tables, plots, and an analysis notebook.
How it works
- A YAML config defines the pipeline under test.
- Documents are chunked with one of four strategies and embedded, with embeddings cached in SQLite.
- Each query runs through retrieval, optional fusion, and optional reranking of the top 50 candidates.
- Chunk scores are max-pooled back to documents and scored against ground-truth relevance labels.
- Results are compared across configurations with confidence intervals and significance tests.
Tech stack
- Language: Python
- Retrieval: BM25, BGE-M3 embeddings, Qdrant, BGE cross-encoder reranker
- Evaluation: pytrec_eval, SciPy (bootstrap and Wilcoxon tests)
- Storage: SQLite embedding cache
- Testing: pytest
Results
- 26 benchmark runs across 4 datasets (SciFact, NFCorpus, FiQA, and an English MLDR subset), covering 2,071 queries and 67,654 documents.
- On FiQA, reranking raised Recall@5 from 0.353 to 0.427, a 21% relative improvement across 648 queries (Wilcoxon p ≈ 1.6 × 10⁻¹⁰).
- On SciFact, the best pipeline reached 0.803 Recall@5 and 0.757 nDCG@10.
Inspect the configurations and per-query results behind the FiQA comparison.
Engineering highlights
Diagnose, don't just score. Candidate-pool Recall@50 separates "the document was never retrieved" from "the reranker ranked it poorly", pointing to which stage to fix.
Fair document-level scoring. Chunk results are deduplicated and max-pooled, so a document split into many chunks can't inflate the metrics.
Fast and resumable. Embeddings are cached by model and text hash, and results are written incrementally, so interrupted runs pick up where they stopped.