Artificial Analysis
Independent AI model benchmarking and selection platform.
Available information varies by company and source.
Profile record updated:
Company facts
- Official name
- Artificial Analysis, Inc.
- Entity type
- COMPANY
- Headquarters
- United States
- Company size
- 10–49
- Market role
- B2B SaaS Provider
- Official website
- artificialanalysis.ai
What Artificial Analysis does
The company aggregates, standardises and produces benchmark data on AI models, then packages that data into software interfaces, APIs, leaderboards and enterprise decision-support tools. It creates value by reducing the time and uncertainty involved in AI model evaluation and procurement. The commercial engine combines public distribution for top-of-funnel adoption with paid access to richer datasets, advanced analytics and enterprise workflows.
Category differentiation
Artificial Analysis is not a foundation model developer or a general MLOps platform. It is an independent benchmarking and analytics layer used to compare and select AI models.
Strategic context
AI-supported assessment from the existing company research; distinguish interpretation from sourced facts.
Artificial Analysis, Inc. is a private US-based B2B SaaS provider focused on independent AI benchmarking and analysis. The company publishes model benchmarks, leaderboards, evaluation frameworks and recommendation tools that compare AI models on quality, speed, latency and cost. Its products help enterprises, developers, researchers and procurement teams assess competing AI models using standardised and vendor-neutral data. The company generates revenue through a freemium commercial model. Public leaderboards and a free API drive awareness and adoption, while premium subscriptions, commercial API access and enterprise-grade data access monetise deeper usage. Its direct customers are business and technical teams making model selection, procurement and validation decisions rather than end-consumers.
Company news briefing
Briefing updated:
Artificial Analysis maintains its role as a key independent benchmarking authority, with its intelligence indexes and comparative performance metrics extensively cited across recent model releases. Its evaluation frameworks track major industry deployments including Anthropic's Claude Opus 5.5, OpenAI's GPT-6 Astra and GPT-5.6 families, DeepSeek-V4.1-Flash, Z.ai's GLM-5.3-Flash, and Xiaomi's MiMo-V2.6-Pro. These independent metrics provide objective performance and cost data across the rapidly evolving proprietary and open-weight large language model landscape.
Business model & monetisation
Artificial Analysis uses a freemium SaaS and data subscription model. Free access to public benchmarks, leaderboards and a free API builds market visibility and user adoption. Revenue comes from premium subscriptions, commercial API access, expanded benchmark datasets, advanced analytics, downloadable reporting and enterprise plans with deeper access and additional seats.
- Premium subscriptions
- Software Subscription
- Commercial API access
- Pay-per-Use
- Enterprise plans with deeper data access and seats
- Software Subscription
- Free API and public leaderboards
Products & capabilities
No products with linked sources are available in this view.
Products & market categories
Technology
Recent recorded signals
Dates refer to the source publication. Older entries are historical context, not evidence of a new event.
Anthropic Launches Claude Opus 5.5 Despite Slowdown Call
AI Model Launch · Recorded impact score: 4/5
Anthropic has released Claude Opus 5.5, its flagship AI model and the first in the 5.5 family, achieving the highest score ever recorded on the Artificial Analysis Intelligence Index (58) and leading six out of ten benchmarks, including outperforming OpenAI's GPT-6 Astra on Terminal-Bench 4.0. Priced 20% lower for input/output tokens and 60% lower for cache reads, it claims a 40% cost reduction for typical workloads, though higher output token usage may offset savings; it's also 30% faster and now the default in Claude Code and the Claude app. OpenAI responded by releasing GPT-6 Sol and Luna, priced ~50% lower than predecessors and offering up to 90% discounts on cached input. Both cite efficiency gains. Claude Opus 5.5 is available on Claude apps, Claude Platform, AWS, Google Cloud, and Azure, with smaller models to follow.
- Claude Opus 5.5 scores 58 on the Artificial Analysis Intelligence Index, the highest recorded, and leads six out of ten benchmarks.
- Pricing is reduced: $4 per million input tokens, $20 per million output tokens, and cache reads at $0.20 per million tokens, promising a 40% cost reduction per task (though higher output token usage may offset savings).
Xiaomi MiMo-V2.6-Pro tops open weights, trained for $3M
AI / LLM · Recorded impact score: 4/5
Xiaomi released MiMo-V2.6-Pro, a natively omnimodal open-weights model with 1.02T total / 42B active parameters, trained for $3M (about 130 hours and 75B tokens). It debuts as the top open-weights model on Artificial Analysis' Intelligence Index (46) with cost efficiency at $0.435/M input and $0.87/M output tokens, under an MIT license. Xiaomi also open-sourced the RL training environment code and recipes, but not the full 7k+ task datasets, signaling an emphasis on transparency in RL training.
- Xiaomi released MiMo-V2.6-Pro and MiMo-V2.6-Flash, natively omnimodal open-weights models.
- MiMo-V2.6-Pro has 1.02T total / 42B active parameters and tops the Artificial Analysis Intelligence Index at 46.
DeepSeek Launches V4.1 Flash with Novel Encoder-Decoder Architecture
AI Model Launch · Recorded impact score: 5/5
DeepSeek released DeepSeek-V4.1-Flash, a 763B-parameter mixture-of-experts model employing a novel causal encoder-decoder architecture with 8B active parameters for prefill and 16B for decode. It features native vision understanding, 1M token context, an MIT license, and extreme inference efficiency, claiming up to 1/8 KV cache footprint versus V4 Flash. Independent evals (Artificial Analysis Index 40, Vals Index #1 open-weight) show it surpasses V4 Pro at lower cost. API pricing is $0.30/1M input and $1.20/1M output tokens. DeepSeek has soft-retired V4 Pro, routing traffic to V4.1 Flash. The model supports SSD offload and local deployment, with Ollama and Baseten offering day-0 support. Technical discussions highlight the architecture's novelty and potential impact on long-context agents.
- DeepSeek launched V4.1-Flash with a causal encoder-decoder architecture, 763B total params (8B prefill/16B decode active).
- Artificial Analysis Index scores V4.1-Flash at 40, above V4 Pro and below GLM-5.3-Flash.
ModelBest Releases MiniCPM5-2B Edge AI Model
AI · Recorded impact score: 4/5
Chinese AI startup ModelBest, in collaboration with the OpenBMB open-source community, has released MiniCPM5-2B, a 2-billion-parameter open-source language model designed for edge devices. The model natively supports tool calling, deep search, code generation, and multi-step reasoning, enabling general-purpose agentic capabilities on resource-constrained hardware. It ranks #1 on the Intelligence Index among open-source models under 4 billion parameters, according to Artificial Analysis, and scores 20 on the Agentic Index. ModelBest has open-sourced the full-stack technical suite, including datasets, training recipes, and reinforcement learning infrastructure, to foster reproducibility. The model aims to shift advanced AI from centralized clouds to edge devices, enhancing privacy, reducing latency, and cutting cloud API costs. Downloads across the MiniCPM family have surpassed 50 million.
- ModelBest released MiniCPM5-2B, a 2-billion-parameter open-source language model for edge devices.
- MiniCPM5-2B ranks #1 on the Intelligence Index among open-source models under 4 billion parameters.
New Articles: Benchmarking GPT-6 Astra, Intelligence Index v4.3, and more
Recorded impact score: 4/5
The articles page now shows 105 articles (up from 94), with new entries including 'Benchmarking GPT-6 Astra' (Sep 9, 2026), 'Announcing the Artificial Analysis Intelligence Index v4.3' (Sep 7, 2026), 'OpenBMB releases MiniCPM5-2B' (Sep 7, 2026), 'Announcing Artificial Analysis Intelligence Index v4.2' (Sep 4, 2026), 'Muse Spark 1.3: Meta reaches the frontier' (Sep 2, 2026), 'Google has released Gemini 3.8 Flash' (Sep 2, 2026), 'Claude Fable 5.1 tops the Artificial Analysis Intelligence Index' (Sep 1, 2026), 'Agnes AI releases Agnes 2.5 Pro Beta' (Aug 27, 2026), 'Intelligence at pocket scale' (Aug 24, 2026), 'Announcing the Speech Agent Arena' (Aug 24, 2026), and 'Announcing the Artificial Analysis Search Index' (Aug 18, 2026).
Explore company relationships
Questions about Artificial Analysis
What is Artificial Analysis?
Artificial Analysis is a B2B AI benchmarking and analysis company that compares models using performance, cost, speed and evaluation data.
Who uses Artificial Analysis?
AI engineers, researchers, procurement teams and enterprise decision-makers use it to evaluate and select AI models.
How does Artificial Analysis make money?
It monetises through premium subscriptions, commercial API access and enterprise plans built on top of its benchmark data and analytics tools.
Sources & coverage
This profile uses public, official and technically observable information. Missing information does not prove that a product or relationship does not exist. The list below does not imply that every profile statement has been verified.
11 publicly documented primary sources and citations linked across the market graph.
Continue your research on Artificial Analysis
Explorer includes additional company details, a Watchlist for up to 25 companies and your personal Strategic Intelligence Agent. It monitors your market daily and delivers tailored briefings with clear strategic context whenever relevant news occurs.
Free, with no time limit.
