Observed Signal · Jun 14, 2026 · Technical Release · Source: DEV Community · Impact: 1/5 · Sentiment: Neutral
Engineer Announces Genesis 2 CPU-Only Cascade MoE AI
Aleksandr Larionov, an industrial automation engineer with 20+ years of experience, describes his career path from programming on a ZX Spectrum to building Genesis 2 — an AI system based on a patented Cascade Mixture-of-Experts architecture. He claims Genesis 2 contains 10,800 experts, achieves 100% accuracy on its domains, delivers 18ms inference, runs entirely on CPU (no GPU), and can execute tasks such as running code and managing servers. The article is a personal narrative explaining his motivation to build AI from scratch and links to the project's GitHub repository.
Personal project/technical announcement about a proprietary AI architecture with limited, early-stage impact on the broader AdTech/MarTech industry.
Track GitHub Signals & Market Shifts in Real-Time
Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.
Key Takeaways & Evidence Grounding
- Author Aleksandr Larionov is an industrial automation engineer with over 20 years of experience.
- Larionov worked at Rosneft for 12 years and rose to acting head of the production automation group, managing 105 people across 14 oil fields.
- He is building Genesis 2, an AI system based on a novel Cascade Mixture-of-Experts architecture that he says is patented.
- Genesis 2 is described as having 10,800 experts, 100% accuracy on its domains, 18ms inference latency, and running entirely on CPU without GPUs.
- The article links to the Genesis 2 code/project hosted on GitHub.
Connected Companies & Entities
1 Entity mappedOntology Mapping & Concepts
Related Market Signals & Shifts
Recent verified developments and strategic activity across this market segment.
Developer Builds Private Self‑Hosted AI Brain Locally
A developer published a detailed walkthrough of building a private, self‑hosted AI “brain” called NEXUS on a consumer Windows laptop (Intel i7, 16GB RAM, no GPU). The system ingests files and web feeds, stores semantic memory as vector embeddings, and answers questions from the author's personal data. The stack is entirely open source and runs locally: Ollama (models Llama 3.2 3B and Mistral 7B), Open WebUI, Qdrant (vector store), n8n for automation, SearXNG for private search, PostgreSQL, Redis, MinIO, Neo4j, and Docker/WSL2. The author reports zero software/API costs (only electricity) and documents the full build publicly, including automation (watched folder, web scraping every two hours) and mobile notifications (Telegram). Publication date: 2026-06-14.
Author Builds Private Local AI 'NEXUS' on Laptop
After cancelling a $240/year ChatGPT Plus subscription, the author built a fully private AI assistant called NEXUS that runs entirely on a 2018 Intel i7 laptop with no GPU. Using Ollama to host local LLMs (llama3.2:3b and mistral:7b), a 274 MB nomic-embed-text model to produce 768-dimensional embeddings, and Qdrant as a local vector database in Docker containers, the author implemented a four-step pipeline (parse, chunk, embed, store) enabling persistent semantic memory and retrieval-augmented generation. The system includes autonomous agents (LangGraph), a watcher for ingestion, and safety design choices (local-only embeddings, timeouts, human review). The project emphasizes data ownership, privacy, and the practical feasibility of local RAG workflows on commodity hardware.
Delivery Rider Builds Multi‑Expert AI Agent MVP
A self-taught developer and night‑shift delivery rider published a detailed account of building an MVP AI system that runs parallel "expert" agents and aggregates their responses. Beginning formal coding on May 24, 2026, the author implemented multi-expert parallel execution (medical, legal, strategy, general fallback) using LLM APIs (Zhipu, Aliyun, OpenRouter), a persistent memory module (last 20 turns), an input/output safety filter with violation logs, and a "director brain" that aggregates expert outputs. The system supports multi-round debate where each expert sees the full discussion history; the author notes higher token consumption, slow response speed, and basic concatenation aggregation as current limitations. Source code and additional design notes are available on GitHub. The post frames the project as a work‑in‑progress and invites feedback on engineering and learning cadence.
Track Real-Time Market Signals & Shifts
Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.
