Observed Signal · Jul 28, 2026 · Technical Report · Source: DEV Community · Impact: 2/5 · Sentiment: Neutral
GPU ledger only bills finished processes
A technical analysis demonstrating limits of NVIDIA's per-process GPU accounting: the driver keeps a post‑mortem ledger of departed processes (via nvidia-smi accounting.mode) but the entries are written or finalized at process exit, contain only PIDs (not preserved command names), and include a sampled gpu_utilization column that often reads 0%. Accounting mode is driver state and can reset on driver unload, producing an empty ledger indistinguishable from an idle GPU unless the mode is checked. Accurate attribution therefore requires a live sampler that maps PIDs while processes are running; otherwise work by long‑lived resident daemons or short jobs that start and exit between sampler ticks will be unattributed. The author argues for exposing the size of the unattributed bucket rather than silently redistributing it.
A practical, technical finding about GPU accounting and attribution that matters for engineering teams measuring local inference costs and attribution; relevant to infrastructure/observability but not industry-shifting policy or platform change.
Track NVIDIA Signals & Market Shifts in Real-Time
Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.
Key Takeaways & Evidence Grounding
- NVIDIA's driver supports per-process accounting via nvidia-smi accounting.mode, and the driver's accounted-apps list persists after a process exits.
- The driver's accounting buffer size is 4000 entries (accounting.buffer_size = 4000).
- The ledger exposes a sampled gpu_utilization column and an accumulated time column; in the author's sample, 97.6% of GPU-milliseconds belonged to rows reporting 0% utilization.
- Accounting mode is driver state (not persistent); if the driver unloads (persistence_mode disabled) the accounting mode resets and the accounted list can be empty even after heavy use.
- The accounted entries store PIDs but not command names; reliable attribution requires a separate live sampler that maps PIDs to service names while processes are running.
Connected Companies & Entities
2 Entities mapped“It turns out NVIDIA's driver has kept per-process accounting for years, and it survives the process....”
“The `13149789` attributed to `ollama` is not `ollama serve`; it is the `llama-server` worker processes it spawns and reaps on model swaps....”
Ontology Mapping & Concepts
Related Market Signals & Shifts
Recent verified developments and strategic activity across this market segment.
AI GPU Clusters Often Misprovisioned, Idle 95%
The article argues that reported GPU utilization metrics often conflate memory residency (models loaded into VRAM) with actual compute activity, leading teams to provision and pay for far more GPU capacity than they use. It defines three idle modes—Batch Idle, Inference Idle, and Provisioning Idle—each tracing back to poor demand-curve forecasting, incorrect concurrency assumptions, and treating loaded memory as active compute. The author gives a cost example (an 8× A100 cluster at ~$38,000/month) to show how sustained low utilization compounds into six‑figure annual waste, and concludes that the root fix is better demand modeling at design time rather than scheduler tuning alone.
Hidden Costs of Cloud GPU Training: Egress, Idle, Lock-In
This analysis (published 2026-05-28) argues that the advertised GPU hourly rate understates real training costs by omitting three major drivers: idle GPU time, data egress fees, and vendor lock‑in. Citing 2026 industry studies, the piece notes average GPU utilization can be as low as ~5% in some Kubernetes deployments, making idle time a dominant cost. It lists typical 2026 egress rates (AWS ~$0.09/GB, Google Cloud ~$0.12/GB) and explains how recurring dataset and checkpoint transfers amplify bills and create data gravity that raises exit costs. Recommended mitigations include idle detection (monitoring nvidia-smi), right‑sizing hardware, co‑locating compute and storage, compressing transfers, and modeling exit costs upfront. The article highlights a shift toward specialized and regional GPU providers that compete on transparent pricing and low or zero egress.
Detect GPU Waste in Kubernetes Clusters
This technical guide explains how GPU capacity in Kubernetes clusters can be wasted despite healthy-looking pod-level metrics, and it describes practical methods to surface and quantify that waste. The article defines common waste modes — idle allocations, tier misplacement, CPU-bound stalls, KV cache pressure, and orphaned workloads — and argues standard kubectl/node metrics miss these signals. It recommends deploying NVIDIA DCGM via dcgm-exporter to export per-GPU telemetry to Prometheus, lists specific DCGM metrics and threshold heuristics for waste alerts, and provides a Prometheus query to find idle GPU allocations. The post also describes tooling: the open-source scanner piqc for quick scans and Paralleliq Introspect for model-aware tier misplacement analysis, and gives a simple cost formula to convert waste into daily dollar figures.
Track Real-Time Market Signals & Shifts
Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.
