Observed Signal · Aug 28, 2026 · Technical Release · Source: DEV Community · Impact: 2/5 · Sentiment: Positive
Large Language Models & Local Agents Market: Flash Onyx 2.2: Local Model for Law and Game Feel
Flash Onyx 2.2 is a development update for a local-first agent shell (Flash Onyx) built around gemma4 configured as an engineering agent that runs fully on a 16 GB laptop GPU at single-digit tokens per second. Version 2.2 expands the system-prompt domains beyond code into legal and game-design guidance, with strict rules (e.g., refuse invented citations, jurisdiction-first for law; game feel heuristics for games). The author compressed prompt text to reduce per-conversation prefill cost, tuned sampling (temperature moved from 0.7 to 0.6 and min_p from 0 to 0.05), and increased context window to 65536 after probes showed 32k/64k/128k used ~8.1–8.2 GB GPU memory. Build tooling now injects the repo license into models; Onyx 2.2 is not yet published but is planned as Natuworkguy/flash-onyx-2.2 in 12b and 31b variants.
Practical technical guidance on prompt design, sampling, and context-window tuning for local LLMs is useful to developers building private/local agents but is a niche developer-level update, not industry-shifting.
Wichtigste Kernpunkte & Evidenz
- Flash Onyx 2 is based on gemma4 configured as an engineering agent.
- The model runs locally on a 16 GB laptop GPU at single-digit tokens per second.
- Version 2.2 adds focused system-prompt domains for LAW and GAMES with explicit rules (e.g., do not invent legal citations; prioritize feel in games).
- Sampling was tuned: temperature lowered from 0.7 to 0.6 and min_p changed from 0 to 0.05.
- Context window was set to 65536 after probes showed 32k/64k/128k windows used ~8.1–8.2 GB GPU memory; prefill rate is ~95 tokens/second.
Verknüpfte Unternehmen
3 verknüpfte UnternehmenApple
Globaler Technologiekonzern mit integriertem Betriebssystem-Ökosystem, proprietärer Consumer-Hardware und einem hochskalierbaren First-Party-Advertising- und Services-Netzwerk.
“At 0.5 it also stopped, but reached for du --max-depth=1, a GNU flag this Mac does not have, breaking a different rule in the same prompt....”
Ollama
Lokale und cloudbasierte Infrastruktur für die effiziente Entwicklung und Bereitstellung von Open-Source-KI-Modellen.
“The build tooling grew up a little too. models/build.py fills the model name into the prompt from the header, and injects the repo license i...”
GitHub
GitHub ist die führende cloudbasierte Entwicklungsplattform für kollaborative Softwareentwicklung, CI/CD-Automatisierung und KI-gestützte Codierung.
“Flash Onyx is the model line behind FLASH (https://github.com/natuworkguy), the local-first agent shell I work on....”
Ontology Mapping & Concepts
Track Real-Time Market Signals & Shifts
Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.
