AI · 2026
VeriFaith
Status: ShippedCatches hallucinations in RAG answers, without an LLM judge.
- Role
- Co-built
- Timeline
- Feb 2026 – Apr 2026
- Status
- Status: Shipped
- Stack
- Python, FastAPI, Pydantic +5
Problem
Retrieval-augmented generation systems still produce claims their sources don't support. Most evaluation relies on an LLM-as-judge, which carries the same biases as the model it grades. In benchmarking, a RAGAS-style LLM-as-judge baseline overestimated faithfulness (Recall = 1.0).
Solution
Remove the LLM from the judging step. Each response is decomposed into atomic claims, the most relevant evidence is retrieved with FAISS semantic search, and every claim is classified with a DeBERTa-v3 NLI model as entailed, neutral, or contradicted.
How it works
- Step 01Atomic claim decomposition
- Step 02FAISS semantic retrieval of evidence
- Step 03NLI entailment classification (DeBERTa-v3)
- Step 04Structured JSON / Markdown faithfulness report
Product
VeriFaith: an NLI-based RAG faithfulness evaluation system exposed as a FastAPI REST API with Pydantic validation and structured JSON/Markdown reports, ready to drop into LangChain, LlamaIndex, or custom RAG pipelines.
Key features
- Deterministic 3-stage pipeline, no LLM-as-judge bias
- Claim-level verdicts with supporting evidence
- FastAPI REST API with Pydantic validation
- Integrates with LangChain, LlamaIndex and custom pipelines
Tech used
Results
- 100% contradiction detection on benchmark datasets
- Showed a RAGAS-style LLM-as-judge baseline overestimated faithfulness
- Invention disclosure submitted; patent filing under review at LPU School of CSE (Apr 2026)