Observed Signal · Aug 24, 2026 · Technical Release · Source: SemiAnalysis · Impact: 4/5 · Sentiment: Positive

AgentX 1.0: Open-source Agentic Inference Benchmark

Executive Signal Summary

SemiAnalysis announced AgentX 1.0 and InferenceXv3, an open-source agentic inference benchmark and benchmark implementation designed for long-context, multi-turn agentic coding workloads up to 1M context (Apache 2.0). The project open-sourced its dataset, tooling, and dashboard after spending more than $3M building the traces and running a matrix on ~2MW across 1,000+ chips. AgentX has already driven 50+ upstream PRs and cross-project optimizations across vLLM, SGLang, TensorRT-LLM, ATOM, LMCache, Mooncake and Dynamo. Results show mixed vendor outcomes: NVIDIA leads on many frontier models and configurations, AMD (with ATOM) is competitive in parts, and recent optimizations have shifted some perf-per-dollar comparisons. The dataset (393-session subset) is available on HuggingFace and the results/dashboards are public on InferenceX.

Polaris7 AgentPolaris7 Strategic Assessment
High Confidence

Open-source 1M-context agentic benchmark and dataset with public dashboard; has already driven dozens of upstream optimizations across major inference engines and affects hardware/software inference strategy for long-context multi-turn workloads.

SIGNAL RADAR

Track SemiAnalysis Signals & Market Shifts in Real-Time

Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.

Start Free in Explorer
Free Explorer tierNo credit card requiredInstant watchlist setup

Key Takeaways & Evidence Grounding

  • SemiAnalysis announced AgentX 1.0, an open-source multi-turn agentic coding inference benchmark (1M context) released under Apache 2.0.
  • InferenceXv3 implements AgentX and ran the matrix on ~2MW of continuously operated compute across over 1,000 chips (multiple SKUs including MI355X, GB300 NVL72, GB200 NVL72, B300, B200, MI325, MI300X, H200, and RTX Pro Servers).
  • SemiAnalysis says it spent more than $3M building the AgentX dataset and open-sourced a representative 393-session subset on HuggingFace.
  • AgentX has already driven 50+ upstream pull requests and optimizations across major open-source projects and engines (vLLM, SGLang, TensorRT-LLM, ATOM, LMCache, Mooncake, Dynamo), improving long-context agentic inference behavior.
  • Benchmarks report vendor-specific trade-offs: NVIDIA often leads on frontier models and certain SKUs, while AMD’s ATOM shows competitive perf-per-dollar in some ranges; recent vLLM/NVIDIA optimizations changed relative MI355X vs B200 performance after Aug 21, 2026.

Connected Companies & Entities

12 Entities mapped

“At SemiAnalysis, most of the team are AI power users and use agents for a broad variety of tasks including coding, analyst research, excel m...”

“It is great to see amazing performance from both NVIDIA and AMD on agentic workloads....”

“In April 2026, OpenAI’s Enterprise agentic spending overtook ChatGPT spending....”

“It is great to see amazing performance from both NVIDIA and AMD on agentic workloads....”

“We worked with Anthropic to ship two Claude Code features to make the AgentX dataset possible....”

“* **GitHub:** Austen Stone for helping with reliability of GitHub Actions that AgentX uses...”

“In addition, we are thankful to all who support our open source InferenceX initiative, including Meta, Microsoft, Oracle, OpenAI, MiniMax, M...”

“In addition, we are thankful to all who support our open source InferenceX initiative, including Meta, Microsoft, Oracle, OpenAI, MiniMax, M...”

“In addition, we are thankful to all who support our open source InferenceX initiative, including Meta, Microsoft, Oracle, OpenAI, MiniMax, M...”

“In addition, we are thankful to all who support our open source InferenceX initiative, including Meta, Microsoft, Oracle, OpenAI, MiniMax, M...”

“we collected an initial corpus of [393 internal SemiAnalysis anonymous Claude Code traces](https://huggingface.co/datasets/semianalysisai/cc...”

Primary Source Grounding & Direct Attribution
Direct Origin Attribution
Primary Reporting: SemiAnalysis•Published: Aug 24, 2026
Original Coverage Title: “AgentX - InferenceXv3: Does CUDA Moat Hold up in Agentic Inferencing?”

Related Market Signals & Shifts

Recent verified developments and strategic activity across this market segment.

Large Language Models (LLM) & AIFeb 16, 2026

InferenceX v2 Benchmarks Blackwell vs AMD & Hopper

SemiAnalysis released InferenceX v2 (formerly InferenceMAX), an open-source Apache 2.0 continuous inference benchmark that expands coverage across ~1,000 frontier GPUs and new distributed inference modes. InferenceXv2 adds large-scale disaggregated prefill (disagg) with wide expert parallelism (wideEP) testing for six recent NVIDIA GPU SKUs (including GB200/GB300 NVL72, B200, B300, Blackwell Ultra) and all recent AMD western SKUs including MI355X. The release includes the first third‑party Pareto-frontier benchmarks for Blackwell Ultra GB300 NVL72 and multi-node MI355X disagg+wideEP FP4/FP8. Key findings: NVIDIA Blackwell rack-scale systems lead for MoE/disaggregated inference and energy efficiency; AMD MI355X is competitive on some FP8 and single-node perf/TCO but suffers composability and FP4 multi-node software gaps. The report also highlights MTP (multi-token/speculative decoding) as a major cost reducer and documents software stacks such as SGLang, vLLM, TensorRT‑LLM, Dynamo, MoRI and Mooncake.

Read assessment
Large Language Models (LLM) & AIMar 14, 2026

AI News: 1M Context, Memory Limits, Agent Infrastructure

This AINews roundup covers multiple AI product and research developments: Replit reportedly tripled to a $9B valuation and launched Replit Agent 4, a collaborative multi-agent canvas for apps, sites, and slides. NVIDIA released Nemotron 3 Super, an open 120B / ~12B-active model with a 1M-token context, hybrid Mamba‑Transformer/SSM Latent MoE architecture, and inference optimizations (including multi-token prediction) claiming up to ~2.2x faster inference versus gpt-oss-120B. The piece traces a broader 2026 trend from coding agents to general knowledge-work agents and highlights launches such as Perplexity’s Personal Computer, Base44 Superagents, and LangChain updates. It also reports Anthropic creating The Anthropic Institute (Jack Clark as Head of Public Benefit) and notes an operational outage affecting Claude/Claude Code. Research and benchmarks covered include agent evaluation work, retrieval/post‑training advances, Google Gemini Embedding 2, Qwen3.5 architecture notes, and device/benchmark reports (M5 Max).

Read assessment
PlatformMar 20, 2026

AI Labs Race to Own Developer Tools and Agent Runtimes

Latent Space's AINews roundup (3/18–3/19/2026) reports consolidation and rapid product activity in developer-facing AI: OpenAI acquired Astral (the team behind uv, ruff, ty) into its Codex efforts; Cursor launched Composer 2, a frontier-class coding model claiming strong price/performance; Anthropic expanded Claude Code with messaging channels and persistent developer workflows; and LangChain introduced LangSmith Fleet for enterprise agent fleets. The dispatch highlights a shift from single agents to managed fleets, multi-agent runtimes, and permissioned agent control planes, while security, identity-based authorization, and observability were emphasized across launches. It also summarizes model releases and benchmarks (MiniMax M2.7, Qwen 3.5 Max Preview), advances in OCR/document parsing (Chandra OCR 2, LlamaIndex LiteParse), and infrastructure research trends like continued pretraining before RL and late-interaction retrieval gains.

Read assessment

Track Real-Time Market Signals & Shifts

Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.