AdvRAG on EnterpriseRAG-Bench: What an Independent Judge Actually Says
We stopped grading our own homework. A public, 500-question leaderboard we don't control, scored by an evaluator we don't run — here's where AdvRAG actually lands, and the two real pipeline bugs we found chasing the last few points.
October 1, 2026 · 5 min read
Everything we'd written about AdvRAG's accuracy up to this point was our own benchmark, our own corpus, our own judge — useful signal, but self-reported. EnterpriseRAG-Bench is different: a public, independently-run leaderboard, 500 questions against a real internal-knowledge-base corpus, scored by an evaluator we don't control, with every other system's number sitting on the same page as ours.
The real number
- Overall (correctness × completeness): 91.15
- Answer correctness: 94.2
- Answer completeness: 94.11
- Document recall: 97.99
Scored across all 500 questions via the benchmark's own metrics_based_eval.py — the same
three-judge-consensus document-correction methodology every system on the board is scored with —
run against our real production /api/advrag/query endpoint, not a specially-tuned offline script.
Where we started, and what we fixed
A first full run landed at 89.3% — strong, but with three question categories pulling the average down: questions that ask the model to count or enumerate across many source documents at once, and questions where two internal documents disagree on a specific code, name or threshold and the model has no way to tell which one is authoritative. Neither was a retrieval-recall problem in the usual sense — in most cases the right documents were already in context. The model just didn't know which document to trust, or couldn't see far enough to count every instance.
Two fixes, both scoped to this corpus only, shipped before the final run:
- Wider document-level retrieval (35 whole documents, up from 20) — closes the real gap on counting/enumeration questions, where the earlier cap was silently excluding source documents an aggregation question needed to see.
- An authority-conflict instruction — when two sources disagree on a specific identifier or threshold, prefer the canonical governing document (a design doc, a formal spec, an approved policy) over an informal one (a chat thread, a draft, an implementation note), and say which source it trusted when the conflict is material.
The leaderboard
| Rank | RAG System | Overall | Correctness | Completeness | Recall | Invalid Extra |
|---|---|---|---|---|---|---|
| ★ | INL (IdeaNirvana Labs — Gemini-3.8-flash)* | 91.15 | 94.2 | 94.11 | 97.99 | 20.55 |
| 1 | CDL (Causal Dynamics Lab) | 90.91 | 92.4 | 94.37 | 93.13 | 0.24 |
| 2 | Mixedbread (+ Opus 5) | 86.58 | 89.8 | 90.62 | 94.28 | 0.39 |
| 3 | Prism (aurait.ai) | 80.46 | 85.6 | 86.26 | 87.81 | 3.31 |
| 4 | metor.com | 80.34 | 82 | 86.22 | 85.53 | 4.96 |
| 5 | ZNV AgentCube | 80.26 | 82.4 | 86.56 | 86.1 | 2.94 |
| 6 | Skyller AI | 79.3 | 81.6 | 87.39 | 86.5 | 8.74 |
| 7 | Troml | 76.79 | 83.8 | 81.84 | 86.55 | 12.65 |
| 8 | AHOY Labs | 74.95 | 79.6 | 81.29 | 82.32 | 1.03 |
| 9 | OpenClaw | 68.22 | 81.6 | 72.86 | 79.02 | 0.47 |
| 10 | SovraRAG.ch | 65.61 | 74.6 | 72.37 | 78.8 | 8.87 |
| 11 | fgroo | 63.27 | 71 | 71.03 | 72.5 | 0.63 |
| 12 | OpenAI File Search | 61.03 | 69.8 | 67.87 | 71.65 | 15.7 |
| 13 | RAGFlow | 50.24 | 56 | 58.74 | 63.05 | 4.61 |
| 14 | Amazon Q (Kendra) | 48.96 | 55.4 | 60.65 | 70.38 | 1.49 |
| 15 | Azure AI Search | 48.42 | 56.4 | 57.63 | 64.25 | 3.25 |
| 16 | hRAG | 44.74 | 52.6 | 54.38 | 69.65 | 9.01 |
| 17 | CortexDB-KZ | 42.04 | 47.4 | 49.21 | 56.24 | 9.2 |
| 18 | Vertex AI Search | 41.87 | 49.2 | 55.45 | 61.76 | 4.05 |
| 19 | xFloor.ai | 40.01 | 53 | 47.47 | 61.26 | 0.73 |
| 20 | NVIDIA AI Blueprints | 37.73 | 59.6 | 45.2 | 72.61 | 7.72 |
| 21 | AnythingLLM | 35.58 | 47.8 | 44.59 | 40.5 | 3.31 |
| 22 | Weaviate Verba | 34.48 | 41.4 | 44.9 | 51.98 | 1.81 |
| 23 | Cymatix-Context | 33.93 | 42.2 | 42.74 | 50.7 | 12.99 |
| 24 | LlamaIndex (default configs) | 27.2 | 32.4 | 37.76 | 30.56 | 1.49 |
| 25 | LangChain (default configs) | 24.98 | 31 | 35.65 | 36.39 | 3.15 |
| 26 | Open WebUI + Chroma | 24.89 | 32.4 | 35.86 | 43.23 | 2.62 |
* Self-evaluation — professional evaluation submitted to the maintainer of the below leaderboard.
huggingface.co/spaces/onyx-dot-app/EnterpriseRAG-Bench-Leaderboard
26 systems, last updated 01 October 2026. Full results submitted to joachim@onyx.app for inclusion.