Observed Signal · Aug 9, 2026 · Technical Release · Source: DEV Community · Impact: 2/5 · Sentiment: Positive

LLM Judge Scores Production Spring Boot AI Agent

Executive Signal Summary

A senior engineer describes building an LLM-as-a-judge evaluation harness for a Spring Boot e-commerce agent. The harness runs 40 anonymized production conversations nightly against five defined metrics (answer correctness, factuality, tool discipline, format compliance, harmless refusal), using deterministic checks where possible and LLM evaluators (e.g., RelevancyEvaluator and FactCheckingEvaluator) for subjective metrics. The author uses a cheap specialized model (Bespoke's Minicheck on Ollama) for factuality and a stronger separate judge model (temperature 0.0) for correctness. The system includes a nightly full run and a CI smoke run (10 cases). Initial runs found real issues (shipping-window claims, stale stock, markdown formatting), and the author emphasizes dataset maintenance, judge stability, and treating scores as signals, not absolute truth.

Polaris7 AgentPolaris7 Strategic Assessment
High Confidence

Practical technical guide introducing a reproducible nightly evaluation harness for conversational e-commerce agents; useful pattern for teams building production chat/agent systems but not industry-shifting.

SIGNAL RADAR

Track Ollama Signals & Market Shifts in Real-Time

Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.

Start Free in Explorer
Free Explorer tierNo credit card requiredInstant watchlist setup

Key Takeaways & Evidence Grounding

  • Author is a Senior Software Engineer II at BS23 in Dhaka and built the described harness.
  • The harness runs 40 anonymized production conversations nightly against five metrics: answer correctness, factuality against context, tool discipline, format compliance, and harmless refusal.
  • Evaluation uses Spring AI evaluators (e.g., RelevancyEvaluator and FactCheckingEvaluator) and a cheap fact-checking model (Bespoke's Minicheck) running on Ollama.
  • The harness enforces judge best practices: temperature 0.0, use a separate ChatClient for evaluation, and deterministic checks for tool discipline.
  • Operational setup: full nightly run (40 cases × three LLM metrics = ~120 judge calls) and a CI smoke run with a 10-case subset; initial run revealed shipping-window, stale-stock, and markdown table failures.

Connected Companies & Entities

1 Entity mapped

“Minicheck runs on Ollama, so the factuality metric costs almost nothing per run, while the correctness judge stays on a strong model....”

Primary Source Grounding & Direct Attribution
Direct Origin Attribution
Primary Reporting: DEV Community•Published: Aug 9, 2026
Original Coverage Title: “Building a Production AI Agent in Spring Boot: The LLM Judge That Scores Your Agent (Part 8)”

Related Market Signals & Shifts

Recent verified developments and strategic activity across this market segment.

Conversational AI / Agent EvaluationJul 1, 2026

LLM-as-Judge Harness for Evaluating AI Agents

The author describes building an LLM-as-judge evaluation harness to automatically grade non-deterministic coach agents (FamNest’s coach) against a rubric. The harness records multi-dimensional scores and the judge’s reasoning for each case. The article highlights that judge models are fallible and lists common biases — position bias, verbosity bias, self-preference, and calibration/drift when judge models update. Practical, mechanical mitigations are recommended: flip pairwise order and require stable verdicts, include length-appropriateness as rubric dimensions, avoid using a judge from the same model family, pin judge model versions, and run a small human-labelled anchor set each run to detect drift. The anchor set (a few dozen hand-labelled examples) is presented as the primary safeguard for trusting automated evaluations.

Read assessment
Conversational AIJul 19, 2026

Production-Grade LLM Evaluation Pipelines Replace 'Vibe' Checks

This article describes building a production-ready evaluation pipeline for large language models that replaces informal human “vibe checks” with automated, CI-integrated testing. A small, versioned golden dataset is run through the target LLM and evaluated by a judge ensemble (faithfulness, instruction-following, JSON schema validation, safety, and domain experts); results feed metrics, regression-detection logic, dashboards, and automated PR comments. The post gives practical guidance—start with a stratified 50-case golden set, version tests and judges, and run evaluations in GitHub Actions to block regressions and speed iteration. After six months in production the system raised hallucination catch rate from ~67% (humans) to 92% (automated), reduced incidents from 3/month to 0.2/month, and cut prompt iteration from ~2 hours to ~15 minutes. The team released MIT-licensed tools (llm-eval-harness, prompt-registry, eval-dashboard).

Read assessment
Large Language Models (LLM) & AIMay 19, 2026

Calibrating LLM Judges in a Financial RAG Evaluation

A developer built a retrieval-augmented generation (RAG) system to answer questions about SEC filings using 84 public company documents from the FinanceBench benchmark. Embeddings (text-embedding-3-small) were stored in Qdrant, the top-6 chunks retrieved per query, and GPT-4o-mini produced answers. Retrieval measured Recall@6 0.830, Precision@6 0.422, and MRR 0.646. An LLM judge that scored answers without ground truth reported 74% correct, but manual labeling and a calibrated judge (provided the ground-truth answer key) showed actual accuracy of 53/100. Calibration raised the judge's specificity (TNR) from 0.55 to 0.86 while maintaining sensitivity (TPR) 1.00. The author shares fixes (metadata filtering, custom prompts) and links a public repo for evaluation code.

Read assessment

Track Real-Time Market Signals & Shifts

Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.