Observed Signal · May 10, 2026 · Technical Release · Source: DEV Community · Impact: 3/5 · Sentiment: Positive
Qwen 3.5 Wins Local Benchmark Using llama.cpp
An independent Round 3 benchmark replaced Ollama with a direct llama.cpp server to measure local LLM performance more precisely. The author built a 12-task automated suite across five categories (coding, multi-file agentic coding, reasoning, tool use, and speed) and tested five models: Qwen 3.5, Gemma 4, Devstral, Codestral, and DeepSeek R1. Running on an NVIDIA RTX 5090 system, Qwen 3.5 swept the leaderboard—best coding, best agentic performance, and best single-model weighted score—reaching ~206.7 tokens/sec and a weighted overall score of 85.3. The migration reclaimed ~44 GB of disk from Ollama, enabled fine-grained inference flags (e.g., --reasoning-budget, --chat-template chatml), and highlighted Mixture-of-Experts (MoE) models’ throughput advantage for local deployment.
Demonstrates that direct llama.cpp deployment and Mixture-of-Experts models can deliver significantly higher local throughput and operational control, informing local inference and agent deployment choices for teams evaluating on-prem or hybrid LLM setups.
Track Ollama Signals & Market Shifts in Real-Time
Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.
Key Takeaways & Evidence Grounding
- The author migrated from Ollama to a direct llama.cpp server to avoid benchmarking Ollama’s runtime choices instead of models.
- Round 3 ran five local models across 12 tasks spanning coding, reasoning, tool use, agentic multi-file coding, and speed microbenchmarks.
- Qwen 3.5 achieved peak throughput of 206.7 tokens/sec and a weighted overall score of 85.3, finishing first on the leaderboard.
- Hardware used: NVIDIA RTX 5090 (32 GB), AMD Ryzen 9 9950X3D (16 cores), 64 GB RAM, Samsung 9100 Pro 2 TB NVMe, Ubuntu 24.04.
- Removing Ollama reclaimed ~44 GB of disk; the llama-server was launched with flags like --reasoning-budget 8192 and --chat-template chatml.
Connected Companies & Entities
5 Entities mappedOntology Mapping & Concepts
Related Market Signals & Shifts
Recent verified developments and strategic activity across this market segment.
Local LLMs Reach Practical Usability
A developer revisits running large language models locally and reports that the landscape has shifted: newer Qwen models (dense and MoE variants) now run acceptably on consumer-class hardware with two RX6800 GPUs and 64 GB RAM. The author highlights Qwen3.6-27B (dense) for accuracy, Qwen3.6-35B-A3B (MoE) for speed, and Qwen-Coder-Next-80B (MoE) for coding tasks. Infrastructure improvements include llama.cpp's experimental router mode, ongoing work to persist attention checkpoints and context slots, and a personal fork that adds slot save/restore. The piece also compares harnesses (Hermes, Pi) and argues that capable local inference enables offline, libre-software experimentation without relying on commercial inference providers.
Nine local LLM interfaces tested on one GPU
A hands-on survey evaluated nine local-model interfaces on the same GPU over roughly two weeks, comparing reliability, offload behaviour, file I/O honesty, and context handling rather than just tokens/sec. Results showed large variance driven by the runtime/harness rather than model weights: Ollama was the default reliable harness (32.9 tokens/sec on a 30B MoE with 443 tokens/sec prefill); llama.cpp was faster when carefully tuned; LM Studio reliably extracted structured data to files; several tools exhibited silent failures or context-window bugs (Unsloth capped at 4096 tokens on Windows); and Claude Code could not connect reliably because local models did not parse its system-prompt format. The author concludes benchmarks must target specific real-world use cases because tool behavior, not model choice alone, determines practical outcomes.
Running Qwen 3.6-35B-A3B on a 30GB Mixed GPU Rig
A developer describes upgrading an autonomous Minecraft AI (“Kiwi‑chan”) running on a heterogeneous 30GB consumer-GPU rig (RTX 3060, RTX 3050, GTX 1660 Ti, GTX 1660 Super) to use the MoE model Qwen 3.6-35B-A3B-Instruct. The author found Qwen’s Active 3 Billion design (≈3.5B parameters active per token) outperforms dense 30B+ models on mixed PCIe rigs by reducing compute and bandwidth bottlenecks. Combined with ik_llama.cpp’s Split Mode Graph tensor parallelism and 8-bit quantization of the KV cache (-ctk q8_0 -ctv q8_0), the setup yields higher throughput (~70–90 tokens/s vs ~20–35 tokens/s for a dense model) and a larger usable context window (32K–64K) while running on Ubuntu Server 24.04 LTS with CUDA 13.0. The post includes deployment flags and manual tensor-split recommendations for maximizing performance on mismatched GPUs.
Track Real-Time Market Signals & Shifts
Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.
