Observed Signal · Aug 18, 2026 · Technical Release · Source: DEV Community · Impact: 2/5 · Sentiment: Neutral

What actually fits in 8GB of VRAM

Executive Signal Summary

A hands-on technical analysis of what "fits" when running LLMs locally on an 8GB VRAM laptop (RTX 5060). The author shows four commonly shown memory figures (dxdiag display memory, dedicated memory, shared memory, and runtime/tool fit outputs) each report a real but different quantity; only dedicated VRAM determines whether a model will load without catastrophic slowdown. The post documents that Windows reports a large shared-memory pool backed by system RAM, that typical desktop use can consume 1–3.5GB of GPU-resident memory before model loading, and that allocation shape (context size, KV cache, attention buffers) often determines failures regardless of free VRAM. Two practical fit rules are given: dense models must fit entirely in GPU VRAM (≈7GB or smaller at 4-bit quant), while MoE models must fit in system RAM with experts offloaded to CPU.

Polaris7 AgentPolaris7 Strategic Assessment
High Confidence

Practical, hands-on guidance for local LLM inference and memory-fit estimation; useful to engineers evaluating model/runtime choices but not industry-shifting.

SIGNAL RADAR

Track NVIDIA Signals & Market Shifts in Real-Time

Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.

Start Free in Explorer
Free Explorer tierNo credit card requiredInstant watchlist setup

Key Takeaways & Evidence Grounding

  • The test machine is described as a laptop with an RTX 5060 and 8GB of dedicated VRAM.
  • dxdiag reported Display Memory: 24,144 MB (Dedicated Memory: 7,899 MB; Shared Memory: 16,245 MB).
  • Shared memory reported by Windows is actually system RAM allocated as overflow and is much slower than dedicated VRAM.
  • Typical desktop usage on the machine consumes between ~1GB and ~3.5GB of the dedicated GPU memory before any model loads.
  • Two fit rules: dense models must fit entirely in GPU VRAM (~7GB or smaller at 4-bit quantisation); mixture-of-experts (MoE) models must fit in system RAM with experts offloaded to CPU while VRAM holds attention layers and the key-value (KV) cache.

Connected Companies & Entities

4 Entities mapped

“The breakdown is four lines down the Display tab on every Windows machine there is....”

“The shared pool isn't the graphics card's at all, and every display adapter on the machine claims it: the integrated Intel graphics, the RTX...”

“Ollama, automatic 80/20 CPU/GPU split | Default behaviour...”

Primary Source Grounding & Direct Attribution
Direct Origin Attribution
Primary Reporting: DEV Community•Published: Aug 18, 2026
Original Coverage Title: “What really fits in 8GB VRAM”

Related Market Signals & Shifts

Recent verified developments and strategic activity across this market segment.

Large Language Models (LLM) & AIAug 8, 2026

Why 16 GB RAM Isn't Really 16 GB

A technical explanation showing why available memory on a machine is often less than its marketed RAM and why large AI models cannot simply be loaded onto consumer hardware. The article demonstrates that runtime heaps (e.g., Java's) are a policy-limited portion of system RAM (default ~25%), and different runtime flags (-Xmx, MaxRAMPercentage) produce different allocation ceilings. It explains model memory arithmetic (e.g., 7 billion parameters × 4 bytes = 28 GB), the performance impact of on-card VRAM bandwidth versus the host link, and why quantization (reducing bytes per value) enables smaller models to fit on consumer GPUs. The piece concludes that very large models (e.g., 70B parameters) require rack-scale specialized cards and are therefore typically offered via APIs.

Read assessment
Large Language Models (LLM) & AIApr 3, 2026

Gemma 4 VRAM Hardware Guide

A developer-published practical guide outlines real-world VRAM requirements for running Gemma 4 models locally. It maps model tiers to recommended memory: E2B/E4B for validation on 8GB laptops, the 26B A4B variant as a sweet spot for 16–24GB GPUs, and the 31B model for users with 24GB+ GPUs. The post includes an Ollama setup guide (Gemma4Guide) and specific optimization tips for Apple Silicon unified memory (M1–M4). The author invites community discussion about hardware setups and runtime experiences.

Read assessment
Large Language Models (LLM) & AIJun 11, 2026

Alibaba’s Qwen 3.6 35B-A3B MoE Model and Local 24GB VRAM Guide

The article reviews Alibaba’s Qwen 3.6 35B‑A3B, a Mixture‑of‑Experts (MoE) LLM released April 16, 2026 under Apache 2.0, and explains why the model requires all 35B parameters to be resident in memory (creating a practical 24GB VRAM minimum). Benchmarks and quantization guidance show that on consumer 24GB GPUs the model can achieve high token throughput (e.g., ~120 tok/s on an RTX 4090 with Q4_K_M and tuned llama.cpp settings). The piece compares the MoE 35B-A3B to the dense Qwen 3.6 27B (which fits in ~16GB and scores higher on SWE‑bench), details VRAM usage by quantization and KV cache, and provides hardware and backend recommendations for local deployment (Ollama, llama.cpp, vLLM, Unsloth quant). Published on Dev.to (source runaihome.com republished) on 2026-06-11.

Read assessment

Track Real-Time Market Signals & Shifts

Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.