Observed Signal · Aug 28, 2026 · Technical Release · Source: DEV Community · Impact: 2/5 · Sentiment: Positive
Large Language Models & Local Agents Market: Flash Onyx 2.2: Local Model for Law and Game Feel
Flash Onyx 2.2 is a development update for a local-first agent shell (Flash Onyx) built around gemma4 configured as an engineering agent that runs fully on a 16 GB laptop GPU at single-digit tokens per second. Version 2.2 expands the system-prompt domains beyond code into legal and game-design guidance, with strict rules (e.g., refuse invented citations, jurisdiction-first for law; game feel heuristics for games). The author compressed prompt text to reduce per-conversation prefill cost, tuned sampling (temperature moved from 0.7 to 0.6 and min_p from 0 to 0.05), and increased context window to 65536 after probes showed 32k/64k/128k used ~8.1–8.2 GB GPU memory. Build tooling now injects the repo license into models; Onyx 2.2 is not yet published but is planned as Natuworkguy/flash-onyx-2.2 in 12b and 31b variants.
Practical technical guidance on prompt design, sampling, and context-window tuning for local LLMs is useful to developers building private/local agents but is a niche developer-level update, not industry-shifting.
Key Takeaways & Evidence Grounding
- Flash Onyx 2 is based on gemma4 configured as an engineering agent.
- The model runs locally on a 16 GB laptop GPU at single-digit tokens per second.
- Version 2.2 adds focused system-prompt domains for LAW and GAMES with explicit rules (e.g., do not invent legal citations; prioritize feel in games).
- Sampling was tuned: temperature lowered from 0.7 to 0.6 and min_p changed from 0 to 0.05.
- Context window was set to 65536 after probes showed 32k/64k/128k windows used ~8.1–8.2 GB GPU memory; prefill rate is ~95 tokens/second.
Connected Companies & Entities
3 Entities mappedApple
Consumer electronics giant with integrated software, services and advertising platforms.
“At 0.5 it also stopped, but reached for du --max-depth=1, a GNU flag this Mac does not have, breaking a different rule in the same prompt....”
Ollama
Local and cloud infrastructure for open-model AI development.
“The build tooling grew up a little too. models/build.py fills the model name into the prompt from the header, and injects the repo license i...”
GitHub
Developer platform for code collaboration, automation and AI coding.
“Flash Onyx is the model line behind FLASH (https://github.com/natuworkguy), the local-first agent shell I work on....”
Ontology Mapping & Concepts
Track Real-Time Market Signals & Shifts
Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.
