Observed Signal · Apr 29, 2026 · Technical Release · Source: DEV Community · Impact: 2/5 · Sentiment: Positive

Running Qwen 3.6-35B-A3B on a 30GB Mixed GPU Rig

Executive Signal Summary

A developer describes upgrading an autonomous Minecraft AI (“Kiwi‑chan”) running on a heterogeneous 30GB consumer-GPU rig (RTX 3060, RTX 3050, GTX 1660 Ti, GTX 1660 Super) to use the MoE model Qwen 3.6-35B-A3B-Instruct. The author found Qwen’s Active 3 Billion design (≈3.5B parameters active per token) outperforms dense 30B+ models on mixed PCIe rigs by reducing compute and bandwidth bottlenecks. Combined with ik_llama.cpp’s Split Mode Graph tensor parallelism and 8-bit quantization of the KV cache (-ctk q8_0 -ctv q8_0), the setup yields higher throughput (~70–90 tokens/s vs ~20–35 tokens/s for a dense model) and a larger usable context window (32K–64K) while running on Ubuntu Server 24.04 LTS with CUDA 13.0. The post includes deployment flags and manual tensor-split recommendations for maximizing performance on mismatched GPUs.

Polaris7 AgentPolaris7 Strategic Assessment
High Confidence

Provides practical, reproducible optimizations for self-hosting large MoE LLMs on inexpensive, heterogeneous consumer GPU rigs—useful to practitioners deploying local inference but of limited broad industry impact.

SIGNAL RADAR

Track NVIDIA Signals & Market Shifts in Real-Time

Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.

Start Free in Explorer
Free Explorer tierNo credit card requiredInstant watchlist setup

Key Takeaways & Evidence Grounding

  • Author runs a heterogeneous GPU rig composed of NVIDIA RTX 3060 (12GB), RTX 3050 (6GB), GTX 1660 Ti (6GB) and GTX 1660 Super (6GB), totaling 30GB VRAM.
  • The chosen model is Qwen 3.6-35B-A3B-Instruct; its Active 3 Billion approach activates roughly 3.5B parameters per token instead of computing all 35B.
  • ik_llama.cpp (a fork of llama.cpp) implements Split Mode Graph tensor parallelism, reportedly improving throughput 3x–4x over serial layer-split execution on mixed GPUs.
  • Applying 8-bit KV cache quantization (-ctk q8_0 -ctv q8_0) doubled the usable context to roughly 32K–64K with minimal quality degradation.
  • Reported inference throughput for Qwen 3.6-35B-A3B on this rig is ~70–90 tokens/s versus ~20–35 tokens/s for a dense 31B model; environment used Ubuntu Server 24.04 LTS with CUDA 13.0.
Primary Source Grounding & Direct Attribution
Direct Origin Attribution
Primary Reporting: DEV Community•Published: Apr 29, 2026
Original Coverage Title: “Upgrading Kiwi-chan’s Brain: Pushing a 30GB "Frankenstein" GPU Rig to the Limit with Qwen 3.6-35B-A3B”

Related Market Signals & Shifts

Recent verified developments and strategic activity across this market segment.

Large Language Models (LLM) & AIJun 11, 2026

Alibaba’s Qwen 3.6 35B-A3B MoE Model and Local 24GB VRAM Guide

The article reviews Alibaba’s Qwen 3.6 35B‑A3B, a Mixture‑of‑Experts (MoE) LLM released April 16, 2026 under Apache 2.0, and explains why the model requires all 35B parameters to be resident in memory (creating a practical 24GB VRAM minimum). Benchmarks and quantization guidance show that on consumer 24GB GPUs the model can achieve high token throughput (e.g., ~120 tok/s on an RTX 4090 with Q4_K_M and tuned llama.cpp settings). The piece compares the MoE 35B-A3B to the dense Qwen 3.6 27B (which fits in ~16GB and scores higher on SWE‑bench), details VRAM usage by quantization and KV cache, and provides hardware and backend recommendations for local deployment (Ollama, llama.cpp, vLLM, Unsloth quant). Published on Dev.to (source runaihome.com republished) on 2026-06-11.

Read assessment
Large Language Models (LLM) & AIJul 4, 2026

Local LLMs Reach Practical Usability

A developer revisits running large language models locally and reports that the landscape has shifted: newer Qwen models (dense and MoE variants) now run acceptably on consumer-class hardware with two RX6800 GPUs and 64 GB RAM. The author highlights Qwen3.6-27B (dense) for accuracy, Qwen3.6-35B-A3B (MoE) for speed, and Qwen-Coder-Next-80B (MoE) for coding tasks. Infrastructure improvements include llama.cpp's experimental router mode, ongoing work to persist attention checkpoints and context slots, and a personal fork that adds slot save/restore. The piece also compares harnesses (Hermes, Pi) and argues that capable local inference enables offline, libre-software experimentation without relying on commercial inference providers.

Read assessment
Large Language Models (LLM) & AIAug 17, 2026

Run Qwen 3.8‑27B Locally with Unsloth & DeepSeek

Technical how‑to by Jacques Gariépy describing step‑by‑step instructions to run the Qwen 3.8‑27B model locally on an NVIDIA RTX 3090 (24 GB) using Unsloth (llama.cpp CUDA 13) as the local inference engine and DeepSeek Harness as the agent orchestration runtime. The guide covers obtaining Unsloth and Hugging Face tokens, a Windows-specific SSLKEYLOGFILE installation fix, compiling DeepSeek Harness, recommended llama-server.exe startup flags (FlashAttention‑2, Q8 KV cache, 32k context), .env and Cordis configuration for automatic local provider selection, common Windows troubleshooting, and measured benchmarks (~125 prompt tokens/s, ~38 predicted tokens/s) with ~23.5 GB VRAM usage.

Read assessment

Track Real-Time Market Signals & Shifts

Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.