Observed Signal · May 26, 2026 · Technical Release · Source: DEV Community · Impact: 2/5 · Sentiment: Neutral
13.5k GPU Dataset Shows 20 Years of FLOPS Growth
An open GPU specification dataset and analysis covering 13,566 GPU models (1999–2025) tracks theoretical peak FP32 performance, TDP, and efficiency trends across generations. The author reports ~400× growth in flagship FP32 peak from 2006 to 2025, a ~100× improvement in FP32-per-watt, and a sharp rise in datacenter TDPs (up to 1,400 W for MI355X in 2025) driven by rack cooling and AI accelerators. The timeline shows alternating FP32 leadership between NVIDIA and AMD across different eras, and the write-up explains measurement methodology, caveats (theoretical vs. real-world FLOPS, vendor sparsity claims for tensor metrics), and publishes the cleaned dataset (CSV/SQLite) under CC BY 4.0 for public use.
An openly published, large GPU-spec dataset and accompanying analysis provide useful empirical data on compute, power, and efficiency trends relevant to AI infrastructure planning, but it is not an industry‑shifting policy or platform announcement.
Track NVIDIA Signals & Market Shifts in Real-Time
Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.
Key Takeaways & Evidence Grounding
- GPU Ark published a cleaned open dataset of 13,566 GPUs with specs and 993 third-party benchmark results (CSV, SQLite) under CC BY 4.0.
- Peak flagship FP32 grew from ~0.3 TFLOPS (GeForce 8800 GTX, 2006) to 126 TFLOPS (Blackwell, 2025) — roughly a 400× increase (~37% CAGR).
- TDP for datacenter accelerators rose sharply after 2020: examples include H100 (700 W), MI325X/B200 (1000 W, 2024) and MI355X (1400 W, 2025).
- Performance per watt (FP32/W) improved approximately 100× from 2006 to peak efficiency around 2022; process-node and architecture were primary drivers.
- Raw FP32 leadership between NVIDIA and AMD shifted over time; raw FP16/tensor metrics are not directly comparable across vendors due to sparsity accounting differences.
Connected Companies & Entities
2 Entities mappedOntology Mapping & Concepts
Related Market Signals & Shifts
Recent verified developments and strategic activity across this market segment.
AI Accelerators Break 50-Year Compute Trend
Exponential View published a research report, The State of the AI Economy, highlighting a break in a 50-year trend of global compute growth. The author says the global stock of compute grew at roughly 66% compounded annually until about 2023, and that the arrival of AI accelerators since 2020 has created a new surge in floating-point compute (FLOPs) — described as bringing more "FLOP-factories" online. The newsletter contextualises prior inflection points (mid-1990s consumer PC/Internet wave and the mid-2000s end of Dennard scaling), and argues the current AI-driven pace may persist for years before reverting toward the long-term trend. The piece also notes mainstream coverage (Bloomberg, WSJ) and opens with a remembrance of journalist Om Malik.
2026 GPU Comparison: NVIDIA, AMD, Intel for AI
This article evaluates workstation and prosumer GPUs for local LLM inference and AI workloads in mid-2026, comparing NVIDIA's Blackwell (RTX 50-series), AMD's Radeon AI Pro R9700, and Intel's Arc Pro B70. It argues that VRAM capacity, memory bandwidth, and software ecosystem maturity matter more than peak theoretical compute (AI TOPS) for real-world transformer inference. The piece provides recommended VRAM ranges for common model sizes, a complete spec and price table for relevant consumer and professional cards, and practical guidance on power, thermal behavior, form factor, PCIe bandwidth, and multi-GPU considerations. Conclusions highlight NVIDIA's Blackwell family as the inference benchmark due to bandwidth and CUDA/TensorRT maturity, AMD's R9700 as a value workstation option with ROCm support, and Intel's B70 as an affordable 32 GB workstation GPU with a maturing oneAPI ecosystem.
InferenceX v2 Benchmarks Blackwell vs AMD & Hopper
SemiAnalysis released InferenceX v2 (formerly InferenceMAX), an open-source Apache 2.0 continuous inference benchmark that expands coverage across ~1,000 frontier GPUs and new distributed inference modes. InferenceXv2 adds large-scale disaggregated prefill (disagg) with wide expert parallelism (wideEP) testing for six recent NVIDIA GPU SKUs (including GB200/GB300 NVL72, B200, B300, Blackwell Ultra) and all recent AMD western SKUs including MI355X. The release includes the first third‑party Pareto-frontier benchmarks for Blackwell Ultra GB300 NVL72 and multi-node MI355X disagg+wideEP FP4/FP8. Key findings: NVIDIA Blackwell rack-scale systems lead for MoE/disaggregated inference and energy efficiency; AMD MI355X is competitive on some FP8 and single-node perf/TCO but suffers composability and FP4 multi-node software gaps. The report also highlights MTP (multi-token/speculative decoding) as a major cost reducer and documents software stacks such as SGLang, vLLM, TensorRT‑LLM, Dynamo, MoRI and Mooncake.
Track Real-Time Market Signals & Shifts
Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.
