Observed Signal · Apr 11, 2026 · Technical Release · Source: DEV Community · Impact: 2/5 · Sentiment: Positive
Eval Stack for LangGraph Agent: LangFuse vs AgentCore
The article describes a practical two-week evaluation sprint to build an LLM-agent evaluation stack for a LangGraph-based agent. The team implemented a layered eval pipeline (conversation, orchestration, retrieval) using LangFuse for tracing, Ragas for RAG-specific metrics, DeepEval for custom metrics and test running, and a FastMCP server for tool calls. They formalized test fixtures in a .eval.yaml format separating deterministic checks from LLM-judge metrics. The team also evaluated AWS Bedrock AgentCore’s native tracing and built-in metrics, found semantic differences (e.g., Ragas’ RAG-grounding faithfulness vs AgentCore’s Builtin.Faithfulness), and adopted a small PoC decision framework (two weeks, ~10 fixtures, 15% divergence threshold) to decide on migration, hybrid use, or custom metrics. The post includes local Docker/Ollama examples and several operational lessons about metric definitions, judge-model bias, canary tests, and keeping fixtures tool-agnostic.
Practical guide for building and comparing LLM-agent evaluation stacks; highlights semantic metric differences between native cloud tooling and specialized RAG metrics, reducing operational risk for teams evaluating migration.
Track Langfuse Signals & Market Shifts in Real-Time
Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.
Key Takeaways & Evidence Grounding
- The authors built a multi-layer eval stack for a LangGraph agent using LangFuse (tracing), Ragas (RAG metrics), DeepEval (custom metrics and test runner) and a FastMCP server for tool calls.
- They defined a .eval.yaml fixture format with test_input and success_criteria to capture realistic scenarios and separate deterministic checks from LLM-judge metrics.
- They compared Ragas metrics with AWS Bedrock AgentCore built-in metrics and found semantic mismatches (e.g., Ragas faithfulness = grounding to retrieved context; AgentCore Builtin.Faithfulness = consistency with conversation history).
- They specified a three-outcome PoC: Adopt, Swap (hybrid), or Build custom, using a two-week PoC with ~10 fixtures and a 15% divergence threshold to compare stacks in parallel.
- The article provides a fully local eval example using Ollama as an LLM judge and Docker Compose for LangFuse and Postgres.
Connected Companies & Entities
2 Entities mappedOntology Mapping & Concepts
Related Market Signals & Shifts
Recent verified developments and strategic activity across this market segment.
Local RAG Evolved into Agentic AI with LangGraph
A developer describes converting a locally hosted RAG assistant (built with Ollama, ChromaDB, LangChain, Docker) into an agentic AI architecture using LangGraph. The author introduces a shared AgentState contract and implements three single-purpose agents — a RAG agent for documentation lookup, a Diagnostic agent with a fast known-error lookup and LLM fallback, and an Escalation agent that generates structured tickets when human intervention is required. An orchestrator uses a classifier to route queries conditionally through a state graph. The article discusses design lessons (classifier fragility, embedding initialization overhead, hardcoded escalation thresholds) and recommends starting with RAG and adding agents where needed.
Making LLM Agents Useful in Production
This curated newsletter edition surveys recent work showing how to move language-model agents from demos to production. Highlights include a Galileo field engineer who built a Claude Code-based system that queries 15 repositories to answer customer questions; OpenAI’s Codex team dogfooding their tooling; and Databricks’ analysis of orchestration and choreography patterns needed as agents scale. The edition also summarizes multiple technical papers and benchmarks (Video-MME-v2, Claw‑Eval, DataFlex), argues for focusing on the agent harness and runtime (Sebastian Raschka), and presents tools addressing agent memory and runtimes (mem0, goose in Rust). The newsletter covers architectural patterns (multi-source MCP), an approach called Recursive Language Models (RLMs) that reduces RAG reliance, and practical walkthroughs for using Claude Code as a personal operating system.
Layered Stack for Reliable LLM Tool Selection
A developer guide describes a production architecture to avoid tool-selection hallucinations in LLM-driven agents. Instead of loading hundreds of tools into context or using pure semantic search, the author recommends a five-step layered filtering stack: intent classification, deterministic metadata filtering, semantic search within the filtered subset, confidence scoring, and a final LLM pick among top candidates. The post cites using lightweight local models—gemma4:e4b via Ollama for intent routing and nomic-embed-text via Ollama for embeddings—reports end-to-end latency under 2 seconds, improved tool-selection accuracy versus pure RAG, and fully local/private model infrastructure. The article also emphasizes writing user-facing tool descriptions and notes concurrent-scaling is the next challenge.
Track Real-Time Market Signals & Shifts
Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.
