Ask your technical documents
a question. Get an honest answer.
Hybrid retrieval, cross-encoder reranking, and source citationson every answer - refuses to answer rather than fabricate one.
Explainable retrieval
BM25, dense similarity, fusion rank, and rerank score for every candidate - not a black box.
Never fabricates
Below-threshold confidence returns an explicit failure state, not a confident-sounding guess.
Entity graph
Co-occurring technical terms extracted per section, browsable as an interactive graph.
Under the hood
Built for evidence, not vibes
A node tree, a hybrid retrieval pipeline, and a judge model
that checks the answer against what it was actually given.

Structure-aware, not chunk-first

Every document becomes a KnowledgeNode tree - headings, tables, warnings kept intact, not flattened.

Grounded Q&A

Every reply is traceable to a source file and page - ask a follow-up, get the same rigor.

Every format, one pipeline

Markdown, PDF, DOCX, XLSX, and CSV all normalize into the same tree - no format-specific retrieval logic.

Retrieval quality, measured

Evaluated against a naive baseline on a labeled set - Precision@5, Recall@5, and MRR, not vibes.

Our pipeline

Four stages, each explainable end to end - no black box between your documents and the answer.

  • 01

    Structure-aware parsing

    Every document becomes a hierarchical KnowledgeNode tree - headings, tables, warnings, and figures kept intact instead of flattened into a wall of text.

  • 02

    Hybrid retrieval

    BM25 keyword search and dense embeddings run in parallel and get fused by reciprocal rank, so exact terminology and paraphrased questions both surface the right passages.

  • 03

    Rerank & grade

    A cross-encoder reranks the fused candidates, then a grading pass drops weak or duplicate matches - every score stays visible in the explainability matrix.

  • 04

    Grounded generation

    The answer is generated only from what survived grading, traced back to a source file and page. If the evidence is too weak, the pipeline says so instead of guessing.

Frequently Asked Questions
Explainable retrieval, honest failure, and citationsyou can actually check.
EvidenceRAG is a document question-answering system built for technical, industrial, and engineering documents - the kind where a sentence taken out of context is worse than no answer at all. It's for teams who need to trust the answer, not just get one.
Markdown, PDF (including scanned/image-only pages via Gemini vision), Word, Excel, and CSV. Every parser produces the same hierarchical KnowledgeNode tree, so retrieval logic doesn't depend on the source format.
Hybrid BM25 + dense retrieval with reciprocal rank fusion, cross-encoder reranking, relevance grading, full parent-section reconstruction, and one bounded follow-up retrieval hop when confidence is low - all before generation even starts.
It says so. There are no hardcoded fallback answers - when retrieval confidence is too low or the evidence doesn't support a claim, the pipeline returns an explicit failure state and shows you the retrieved references instead of guessing.
Every answer ships with source file, page number, and section heading path. An explainability matrix also exposes BM25 score, dense score, fusion rank, and rerank score for every candidate chunk that was considered.
A second, cheap LLM call acts as a judge: given the generated answer and the exact chunks it was allowed to use, it returns a 0-1 support score and flags any unsupported claims. This catches hallucination that retrieval metrics alone can't.
Ask your documents a question
Point it at your technical manuals and get answers with page
numbers, not confident-sounding guesses.
EvidenceRAG
Structure-aware hybrid RAG for technical document intelligence
Product
Chat
Entity Graph
Documents
Source Citations
Explainability
Pipeline
Hybrid Retrieval
Cross-Encoder Reranking
Parent Reconstruction
Faithfulness Scoring
Telemetry
Resources
README
Eval Report
API Docs
Sample Corpus
GitHub