Observed Signal · May 19, 2026 · Technical Release · Source: DEV Community · Impact: 2/5 · Sentiment: Neutral

Calibrating LLM Judges in a Financial RAG Evaluation

Executive Signal Summary

A developer built a retrieval-augmented generation (RAG) system to answer questions about SEC filings using 84 public company documents from the FinanceBench benchmark. Embeddings (text-embedding-3-small) were stored in Qdrant, the top-6 chunks retrieved per query, and GPT-4o-mini produced answers. Retrieval measured Recall@6 0.830, Precision@6 0.422, and MRR 0.646. An LLM judge that scored answers without ground truth reported 74% correct, but manual labeling and a calibrated judge (provided the ground-truth answer key) showed actual accuracy of 53/100. Calibration raised the judge's specificity (TNR) from 0.55 to 0.86 while maintaining sensitivity (TPR) 1.00. The author shares fixes (metadata filtering, custom prompts) and links a public repo for evaluation code.

Polaris7 AgentPolaris7 Strategic Assessment
High Confidence

Practical, technical case study showing LLM-judge calibration pitfalls and measurable improvements; useful for teams building RAG/LLM evaluation pipelines but not industry-shifting.

SIGNAL RADAR

Track Qdrant Signals & Market Shifts in Real-Time

Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.

Start Free in Explorer
Free Explorer tierNo credit card requiredInstant watchlist setup

Key Takeaways & Evidence Grounding

  • Built a RAG system for SEC filing Q&A using 84 documents from FinanceBench.
  • Embeddings used: text-embedding-3-small stored in Qdrant; generation with GPT-4o-mini; retrieved top-6 chunks per query.
  • Retrieval metrics: Recall@6 = 0.830, Precision@6 = 0.422, MRR = 0.646.
  • An uncalibrated LLM judge classified 74% of answers as correct; calibrated judge (given ground truth) showed 53/100 answers were correct.
  • Calibration improved judge specificity (TNR) from 0.55 to 0.86 while TPR remained 1.00.
Primary Source Grounding & Direct Attribution
Direct Origin Attribution
Primary Reporting: DEV Community•Published: May 19, 2026
Original Coverage Title: “Building an Evaluation Harness for Financial RAG: What I Learned About LLM-as-Judge Calibration”

Related Market Signals & Shifts

Recent verified developments and strategic activity across this market segment.

Conversational AIAug 9, 2026

LLM Judge Scores Production Spring Boot AI Agent

A senior engineer describes building an LLM-as-a-judge evaluation harness for a Spring Boot e-commerce agent. The harness runs 40 anonymized production conversations nightly against five defined metrics (answer correctness, factuality, tool discipline, format compliance, harmless refusal), using deterministic checks where possible and LLM evaluators (e.g., RelevancyEvaluator and FactCheckingEvaluator) for subjective metrics. The author uses a cheap specialized model (Bespoke's Minicheck on Ollama) for factuality and a stronger separate judge model (temperature 0.0) for correctness. The system includes a nightly full run and a CI smoke run (10 cases). Initial runs found real issues (shipping-window claims, stale stock, markdown formatting), and the author emphasizes dataset maintenance, judge stability, and treating scores as signals, not absolute truth.

Read assessment
Conversational AI & ChatbotsJul 14, 2026

RAG Evaluation with RAGAs: Faithfulness, Recall, Relevance

This article presents RAGAs (Retrieval Augmented Generation Assessment), an evaluation framework that decomposes RAG system quality into three diagnostic metrics: faithfulness, context recall, and answer relevance. The author uses a Vietnamese bank compliance assistant case study where retrieval returned correct documents but the generator hallucinated non-existent rules. RAGAs helped surface that the generation layer was producing unsupported claims (faithfulness 0.71 on a 120-question set) and that retrieval chunking reduced context recall (initially 0.68). Practical remediation included a real-time faithfulness gate (which reduced user-reported wrong answers by ~55%), sentence-window retrieval to raise context recall to 0.84, and prompt surgery to improve answer relevance. The piece also covers operational guidance: a minimum 80-question ground-truth eval set, weekly automated runs (e.g., GitHub Actions), and using an LLM-as-judge (example: gpt-4o-mini) to keep costs low (under $5 per 100-question run).

Read assessment
Conversational AI / Agent EvaluationJul 1, 2026

LLM-as-Judge Harness for Evaluating AI Agents

The author describes building an LLM-as-judge evaluation harness to automatically grade non-deterministic coach agents (FamNest’s coach) against a rubric. The harness records multi-dimensional scores and the judge’s reasoning for each case. The article highlights that judge models are fallible and lists common biases — position bias, verbosity bias, self-preference, and calibration/drift when judge models update. Practical, mechanical mitigations are recommended: flip pairwise order and require stable verdicts, include length-appropriateness as rubric dimensions, avoid using a judge from the same model family, pin judge model versions, and run a small human-labelled anchor set each run to detect drift. The anchor set (a few dozen hand-labelled examples) is presented as the primary safeguard for trusting automated evaluations.

Read assessment

Track Real-Time Market Signals & Shifts

Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.