Skip to content
All builds

AI · 2026

VeriFaith

Status: Shipped

Catches hallucinations in RAG answers, without an LLM judge.

Role
Co-built
Timeline
Feb 2026 – Apr 2026
Status
Status: Shipped
Stack
Python, FastAPI, Pydantic +5

01Problem

Retrieval-augmented generation systems still produce claims their sources don't support. Most evaluation relies on an LLM-as-judge, which carries the same biases as the model it grades. In benchmarking, a RAGAS-style LLM-as-judge baseline overestimated faithfulness (Recall = 1.0).

02Solution

Remove the LLM from the judging step. Each response is decomposed into atomic claims, the most relevant evidence is retrieved with FAISS semantic search, and every claim is classified with a DeBERTa-v3 NLI model as entailed, neutral, or contradicted.

How it works

  1. Step 01Atomic claim decomposition
  2. Step 02FAISS semantic retrieval of evidence
  3. Step 03NLI entailment classification (DeBERTa-v3)
  4. Step 04Structured JSON / Markdown faithfulness report

03Product

VeriFaith: an NLI-based RAG faithfulness evaluation system exposed as a FastAPI REST API with Pydantic validation and structured JSON/Markdown reports, ready to drop into LangChain, LlamaIndex, or custom RAG pipelines.

Key features

  • Deterministic 3-stage pipeline, no LLM-as-judge bias
  • Claim-level verdicts with supporting evidence
  • FastAPI REST API with Pydantic validation
  • Integrates with LangChain, LlamaIndex and custom pipelines

Tech used

  • Python
  • FastAPI
  • Pydantic
  • LLaMA 3.3 70B (Groq)
  • DeBERTa-v3
  • FAISS
  • sentence-transformers
  • NLTK

04Results

  • 100% contradiction detection on benchmark datasets
  • Showed a RAGAS-style LLM-as-judge baseline overestimated faithfulness
  • Invention disclosure submitted; patent filing under review at LPU School of CSE (Apr 2026)

© 2026 Suharshit Singh

Built with Next.js, deployed on Vercel

Last updated

Back to top