Observed Signal · May 26, 2026 · Technical Release · Source: DEV Community · Impact: 2/5 · Sentiment: Neutral

13.5k GPU Dataset Shows 20 Years of FLOPS Growth

Executive Signal Summary

An open GPU specification dataset and analysis covering 13,566 GPU models (1999–2025) tracks theoretical peak FP32 performance, TDP, and efficiency trends across generations. The author reports ~400× growth in flagship FP32 peak from 2006 to 2025, a ~100× improvement in FP32-per-watt, and a sharp rise in datacenter TDPs (up to 1,400 W for MI355X in 2025) driven by rack cooling and AI accelerators. The timeline shows alternating FP32 leadership between NVIDIA and AMD across different eras, and the write-up explains measurement methodology, caveats (theoretical vs. real-world FLOPS, vendor sparsity claims for tensor metrics), and publishes the cleaned dataset (CSV/SQLite) under CC BY 4.0 for public use.

Polaris7 AgentPolaris7 Strategic Assessment
High Confidence

An openly published, large GPU-spec dataset and accompanying analysis provide useful empirical data on compute, power, and efficiency trends relevant to AI infrastructure planning, but it is not an industry‑shifting policy or platform announcement.

SIGNAL RADAR

Track NVIDIA Signals & Market Shifts in Real-Time

Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.

Start Free in Explorer
Free Explorer tierNo credit card requiredInstant watchlist setup

Key Takeaways & Evidence Grounding

  • GPU Ark published a cleaned open dataset of 13,566 GPUs with specs and 993 third-party benchmark results (CSV, SQLite) under CC BY 4.0.
  • Peak flagship FP32 grew from ~0.3 TFLOPS (GeForce 8800 GTX, 2006) to 126 TFLOPS (Blackwell, 2025) — roughly a 400× increase (~37% CAGR).
  • TDP for datacenter accelerators rose sharply after 2020: examples include H100 (700 W), MI325X/B200 (1000 W, 2024) and MI355X (1400 W, 2025).
  • Performance per watt (FP32/W) improved approximately 100× from 2006 to peak efficiency around 2022; process-node and architecture were primary drivers.
  • Raw FP32 leadership between NVIDIA and AMD shifted over time; raw FP16/tensor metrics are not directly comparable across vendors due to sparsity accounting differences.

Ontology Mapping & Concepts

Primary Source Grounding & Direct Attribution
Direct Origin Attribution
Primary Reporting: DEV Community•Published: May 26, 2026
Original Coverage Title: “20 Years of GPUs in Numbers: How FLOPS & TDP Grew, and Who Led the NVIDIA vs AMD Race (open dataset, 13.5k GPUs)”

Related Market Signals & Shifts

Recent verified developments and strategic activity across this market segment.

Large Language Models & AIJun 28, 2026

AI Accelerators Break 50-Year Compute Trend

Exponential View published a research report, The State of the AI Economy, highlighting a break in a 50-year trend of global compute growth. The author says the global stock of compute grew at roughly 66% compounded annually until about 2023, and that the arrival of AI accelerators since 2020 has created a new surge in floating-point compute (FLOPs) — described as bringing more "FLOP-factories" online. The newsletter contextualises prior inflection points (mid-1990s consumer PC/Internet wave and the mid-2000s end of Dennard scaling), and argues the current AI-driven pace may persist for years before reverting toward the long-term trend. The piece also notes mainstream coverage (Bloomberg, WSJ) and opens with a remembrance of journalist Om Malik.

Read assessment
Large Language Models (LLM) & AIJul 14, 2026

2026 GPU Comparison: NVIDIA, AMD, Intel for AI

This article evaluates workstation and prosumer GPUs for local LLM inference and AI workloads in mid-2026, comparing NVIDIA's Blackwell (RTX 50-series), AMD's Radeon AI Pro R9700, and Intel's Arc Pro B70. It argues that VRAM capacity, memory bandwidth, and software ecosystem maturity matter more than peak theoretical compute (AI TOPS) for real-world transformer inference. The piece provides recommended VRAM ranges for common model sizes, a complete spec and price table for relevant consumer and professional cards, and practical guidance on power, thermal behavior, form factor, PCIe bandwidth, and multi-GPU considerations. Conclusions highlight NVIDIA's Blackwell family as the inference benchmark due to bandwidth and CUDA/TensorRT maturity, AMD's R9700 as a value workstation option with ROCm support, and Intel's B70 as an affordable 32 GB workstation GPU with a maturing oneAPI ecosystem.

Read assessment
Large Language Models (LLM) & AIFeb 16, 2026

InferenceX v2 Benchmarks Blackwell vs AMD & Hopper

SemiAnalysis released InferenceX v2 (formerly InferenceMAX), an open-source Apache 2.0 continuous inference benchmark that expands coverage across ~1,000 frontier GPUs and new distributed inference modes. InferenceXv2 adds large-scale disaggregated prefill (disagg) with wide expert parallelism (wideEP) testing for six recent NVIDIA GPU SKUs (including GB200/GB300 NVL72, B200, B300, Blackwell Ultra) and all recent AMD western SKUs including MI355X. The release includes the first third‑party Pareto-frontier benchmarks for Blackwell Ultra GB300 NVL72 and multi-node MI355X disagg+wideEP FP4/FP8. Key findings: NVIDIA Blackwell rack-scale systems lead for MoE/disaggregated inference and energy efficiency; AMD MI355X is competitive on some FP8 and single-node perf/TCO but suffers composability and FP4 multi-node software gaps. The report also highlights MTP (multi-token/speculative decoding) as a major cost reducer and documents software stacks such as SGLang, vLLM, TensorRT‑LLM, Dynamo, MoRI and Mooncake.

Read assessment

Track Real-Time Market Signals & Shifts

Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.