Observed Signal · May 10, 2026 · Technical Release · Source: DEV Community · Impact: 3/5 · Sentiment: Positive

Qwen 3.5 Wins Local Benchmark Using llama.cpp

Executive Signal Summary

An independent Round 3 benchmark replaced Ollama with a direct llama.cpp server to measure local LLM performance more precisely. The author built a 12-task automated suite across five categories (coding, multi-file agentic coding, reasoning, tool use, and speed) and tested five models: Qwen 3.5, Gemma 4, Devstral, Codestral, and DeepSeek R1. Running on an NVIDIA RTX 5090 system, Qwen 3.5 swept the leaderboard—best coding, best agentic performance, and best single-model weighted score—reaching ~206.7 tokens/sec and a weighted overall score of 85.3. The migration reclaimed ~44 GB of disk from Ollama, enabled fine-grained inference flags (e.g., --reasoning-budget, --chat-template chatml), and highlighted Mixture-of-Experts (MoE) models’ throughput advantage for local deployment.

Polaris7 AgentPolaris7 Strategic Assessment
High Confidence

Demonstrates that direct llama.cpp deployment and Mixture-of-Experts models can deliver significantly higher local throughput and operational control, informing local inference and agent deployment choices for teams evaluating on-prem or hybrid LLM setups.

SIGNAL RADAR

Track Ollama Signals & Market Shifts in Real-Time

Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.

Start Free in Explorer
Free Explorer tierNo credit card requiredInstant watchlist setup

Key Takeaways & Evidence Grounding

  • The author migrated from Ollama to a direct llama.cpp server to avoid benchmarking Ollama’s runtime choices instead of models.
  • Round 3 ran five local models across 12 tasks spanning coding, reasoning, tool use, agentic multi-file coding, and speed microbenchmarks.
  • Qwen 3.5 achieved peak throughput of 206.7 tokens/sec and a weighted overall score of 85.3, finishing first on the leaderboard.
  • Hardware used: NVIDIA RTX 5090 (32 GB), AMD Ryzen 9 9950X3D (16 cores), 64 GB RAM, Samsung 9100 Pro 2 TB NVMe, Ubuntu 24.04.
  • Removing Ollama reclaimed ~44 GB of disk; the llama-server was launched with flags like --reasoning-budget 8192 and --chat-template chatml.

Ontology Mapping & Concepts

Primary Source Grounding & Direct Attribution
Direct Origin Attribution
Primary Reporting: DEV Community•Published: May 10, 2026
Original Coverage Title: “Model Showdown Round 3: Ditching Ollama in Favor of llama.cpp”

Related Market Signals & Shifts

Recent verified developments and strategic activity across this market segment.

Large Language Models (LLM) & AIJul 4, 2026

Local LLMs Reach Practical Usability

A developer revisits running large language models locally and reports that the landscape has shifted: newer Qwen models (dense and MoE variants) now run acceptably on consumer-class hardware with two RX6800 GPUs and 64 GB RAM. The author highlights Qwen3.6-27B (dense) for accuracy, Qwen3.6-35B-A3B (MoE) for speed, and Qwen-Coder-Next-80B (MoE) for coding tasks. Infrastructure improvements include llama.cpp's experimental router mode, ongoing work to persist attention checkpoints and context slots, and a personal fork that adds slot save/restore. The piece also compares harnesses (Hermes, Pi) and argues that capable local inference enables offline, libre-software experimentation without relying on commercial inference providers.

Read assessment
Large Language Models (LLM) & AIAug 11, 2026

Nine local LLM interfaces tested on one GPU

A hands-on survey evaluated nine local-model interfaces on the same GPU over roughly two weeks, comparing reliability, offload behaviour, file I/O honesty, and context handling rather than just tokens/sec. Results showed large variance driven by the runtime/harness rather than model weights: Ollama was the default reliable harness (32.9 tokens/sec on a 30B MoE with 443 tokens/sec prefill); llama.cpp was faster when carefully tuned; LM Studio reliably extracted structured data to files; several tools exhibited silent failures or context-window bugs (Unsloth capped at 4096 tokens on Windows); and Claude Code could not connect reliably because local models did not parse its system-prompt format. The author concludes benchmarks must target specific real-world use cases because tool behavior, not model choice alone, determines practical outcomes.

Read assessment
Large Language Models (LLM) & AIApr 29, 2026

Running Qwen 3.6-35B-A3B on a 30GB Mixed GPU Rig

A developer describes upgrading an autonomous Minecraft AI (“Kiwi‑chan”) running on a heterogeneous 30GB consumer-GPU rig (RTX 3060, RTX 3050, GTX 1660 Ti, GTX 1660 Super) to use the MoE model Qwen 3.6-35B-A3B-Instruct. The author found Qwen’s Active 3 Billion design (≈3.5B parameters active per token) outperforms dense 30B+ models on mixed PCIe rigs by reducing compute and bandwidth bottlenecks. Combined with ik_llama.cpp’s Split Mode Graph tensor parallelism and 8-bit quantization of the KV cache (-ctk q8_0 -ctv q8_0), the setup yields higher throughput (~70–90 tokens/s vs ~20–35 tokens/s for a dense model) and a larger usable context window (32K–64K) while running on Ubuntu Server 24.04 LTS with CUDA 13.0. The post includes deployment flags and manual tensor-split recommendations for maximizing performance on mismatched GPUs.

Read assessment

Track Real-Time Market Signals & Shifts

Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.