Observed Signal · Aug 17, 2026 · Technical Release · Source: DEV Community · Impact: 2/5 · Sentiment: Positive

Run Qwen 3.8‑27B Locally with Unsloth & DeepSeek

Executive Signal Summary

Technical how‑to by Jacques Gariépy describing step‑by‑step instructions to run the Qwen 3.8‑27B model locally on an NVIDIA RTX 3090 (24 GB) using Unsloth (llama.cpp CUDA 13) as the local inference engine and DeepSeek Harness as the agent orchestration runtime. The guide covers obtaining Unsloth and Hugging Face tokens, a Windows-specific SSLKEYLOGFILE installation fix, compiling DeepSeek Harness, recommended llama-server.exe startup flags (FlashAttention‑2, Q8 KV cache, 32k context), .env and Cordis configuration for automatic local provider selection, common Windows troubleshooting, and measured benchmarks (~125 prompt tokens/s, ~38 predicted tokens/s) with ~23.5 GB VRAM usage.

Polaris7 AgentPolaris7 Strategic Assessment
High Confidence

Practical, reproducible guide demonstrating self‑hosting a 27B LLM on consumer GPU hardware; relevant to AI/LLM infrastructure but niche and not directly industry‑wide for AdTech.

SIGNAL RADAR

Track DeepSeek Signals & Market Shifts in Real-Time

Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.

Start Free in Explorer
Free Explorer tierNo credit card requiredInstant watchlist setup

Key Takeaways & Evidence Grounding

  • Guide shows how to run Qwen 3.8‑27B quantized as UD-Q4_K_XL locally on an NVIDIA RTX 3090 (24 GB VRAM).
  • Uses Unsloth's llama.cpp CUDA 13 inference server (llama-server.exe) with FlashAttention‑2 and Q8 KV cache compression.
  • DeepSeek Harness (deepseek-ai) is used as the Node.js/TypeScript agent orchestration runtime with Cordis plugins.
  • Reported benchmarks: prompt_tokens_per_second = 125.37 and predicted_tokens_per_second = 38.19; VRAM used ≈ 23.5 GB.
  • Article documents a Windows installation bug caused by SSLKEYLOGFILE and provides a step‑by‑step patch and PowerShell mitigation.

Connected Companies & Entities

3 Entities mapped

“Clone and build the DeepSeek Harness repository (https://github.com/deepseek-ai/deepseek-harness) to use the DeepSeek Harness runtime (Node....”

“If downloading GGUF models from the Hugging Face Hub, create an access token (hf_...) at huggingface.co/settings/tokens and configure HF_TOK...”

“The NVIDIA RTX 3090 (24 GB VRAM) is recommended for self-hosting 27–32B parameter models and is used for the documented setup and benchmarks...”

Primary Source Grounding & Direct Attribution
Direct Origin Attribution
Primary Reporting: DEV Community•Published: Aug 17, 2026
Original Coverage Title: “Faire tourner Qwen 3.8–27B en local avec Unsloth et DeepSeek Harness sur une RTX 3090 (24 Go) sous Windows 11.”

Related Market Signals & Shifts

Recent verified developments and strategic activity across this market segment.

Large Language Models (LLM) & AIApr 29, 2026

Running Qwen 3.6-35B-A3B on a 30GB Mixed GPU Rig

A developer describes upgrading an autonomous Minecraft AI (“Kiwi‑chan”) running on a heterogeneous 30GB consumer-GPU rig (RTX 3060, RTX 3050, GTX 1660 Ti, GTX 1660 Super) to use the MoE model Qwen 3.6-35B-A3B-Instruct. The author found Qwen’s Active 3 Billion design (≈3.5B parameters active per token) outperforms dense 30B+ models on mixed PCIe rigs by reducing compute and bandwidth bottlenecks. Combined with ik_llama.cpp’s Split Mode Graph tensor parallelism and 8-bit quantization of the KV cache (-ctk q8_0 -ctv q8_0), the setup yields higher throughput (~70–90 tokens/s vs ~20–35 tokens/s for a dense model) and a larger usable context window (32K–64K) while running on Ubuntu Server 24.04 LTS with CUDA 13.0. The post includes deployment flags and manual tensor-split recommendations for maximizing performance on mismatched GPUs.

Read assessment
Large Language Models (LLM) & AIJul 4, 2026

Local LLMs Reach Practical Usability

A developer revisits running large language models locally and reports that the landscape has shifted: newer Qwen models (dense and MoE variants) now run acceptably on consumer-class hardware with two RX6800 GPUs and 64 GB RAM. The author highlights Qwen3.6-27B (dense) for accuracy, Qwen3.6-35B-A3B (MoE) for speed, and Qwen-Coder-Next-80B (MoE) for coding tasks. Infrastructure improvements include llama.cpp's experimental router mode, ongoing work to persist attention checkpoints and context slots, and a personal fork that adds slot save/restore. The piece also compares harnesses (Hermes, Pi) and argues that capable local inference enables offline, libre-software experimentation without relying on commercial inference providers.

Read assessment
Large Language Models (LLM) & AIMay 10, 2026

Qwen 3.5 Wins Local Benchmark Using llama.cpp

An independent Round 3 benchmark replaced Ollama with a direct llama.cpp server to measure local LLM performance more precisely. The author built a 12-task automated suite across five categories (coding, multi-file agentic coding, reasoning, tool use, and speed) and tested five models: Qwen 3.5, Gemma 4, Devstral, Codestral, and DeepSeek R1. Running on an NVIDIA RTX 5090 system, Qwen 3.5 swept the leaderboard—best coding, best agentic performance, and best single-model weighted score—reaching ~206.7 tokens/sec and a weighted overall score of 85.3. The migration reclaimed ~44 GB of disk from Ollama, enabled fine-grained inference flags (e.g., --reasoning-budget, --chat-template chatml), and highlighted Mixture-of-Experts (MoE) models’ throughput advantage for local deployment.

Read assessment

Track Real-Time Market Signals & Shifts

Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.