Observed Signal · Jun 24, 2026 · Technical Release · Source: DEV Community · Impact: 2/5 · Sentiment: Positive

Neonmem 0.9.7 Released with Offline Vector Recall

Executive Signal Summary

Neonmem 0.9.7 is a technical release that introduces a two-level importer separating a deduplicated, searchable facts pool from typed memories extracted from agent chats. The release replaces the previous embedder with IBM Granite-30M run as a fused fp16 ONNX graph via ONNX Runtime, enabling FAISS-based retrieval on CPU without GPUs, cloud services, or third‑party LLMs. New features include tag canonicalization, deduplicated chat capture, a single compressed "cartridge" containing the full source corpus, and opt-in AES-256-GCM encryption at rest. The build is distributed as Windows (signed installer + portable) and Linux (AppImage) binaries, with macOS support forthcoming. All core components are open-source / permissively licensed.

Polaris7 AgentPolaris7 Strategic Assessment
High Confidence

Technical release improves local, CPU-based embeddings and vector retrieval with clear privacy and reproducibility features, but it is a niche developer tool rather than a platform-changing announcement for the advertising industry.

SIGNAL RADAR

Track IBM Signals & Market Shifts in Real-Time

Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.

Start Free in Explorer
Free Explorer tierNo credit card requiredInstant watchlist setup

Key Takeaways & Evidence Grounding

  • Neonmem released version 0.9.7.
  • Introduces a two-level importer separating a deduplicated 'facts pool' and typed 'memories' extracted from agent chats.
  • Replaced previous embedder with IBM Granite-30M executed as a fused fp16 ONNX graph through ONNX Runtime for CPU-only retrieval.
  • Uses FAISS for vector search; claims no GPU, no PyTorch, no API key, and no cloud dependencies.
  • Supports an opt-in AES-256-GCM encryption-at-rest option and ships Windows (signed installer + portable) and Linux (AppImage) builds; macOS is planned.

Connected Companies & Entities

2 Entities mapped

“0.9.7 replaces the old embedder with IBM Granite-30M, run as a fused fp16 ONNX graph through ONNX Runtime:...”

“Point Neonmem at a Claude (or other agent) transcript and it pulls out only what's worth keeping — the decisions, dead-ends and rules — as c...”

Primary Source Grounding & Direct Attribution
Direct Origin Attribution
Primary Reporting: DEV Community•Published: Jun 24, 2026
Original Coverage Title: “Neonmem 0.9.7 is out.”

Related Market Signals & Shifts

Recent verified developments and strategic activity across this market segment.

Large language model memory & retrievalMar 31, 2026

nan-forget: Brain-Inspired Memory for LLMs

nan-forget is an open-source, long-term memory system for LLM-powered coding tools that applies three neuroscience ideas: forgetting (time-based decay), spreading activation (multi-stage retrieval), and sleep-like consolidation. It scores memories by combining vector similarity with a decay_weight (30-day half-life) and a frequency_boost, and uses a three-stage retrieval pipeline (Recognition → Recall → Spreading Activation) to surface related context. A consolidation engine runs after 10 saves or 24 hours to cluster, summarize and archive originals; garbage collection deduplicates (cosine > 0.95) and expires stale entries. Implementation uses a single SQLite database with sqlite-vec for vector KNN (replacing a prior Qdrant setup), structured JSON memory records, four automatic capture hooks, cross-LLM support (MCP server, REST API, CLI), and is published under an MIT license on GitHub (NaNMesh/nan-forget).

Read assessment
Large Language Models (LLM) & AIMar 14, 2026

AI News: 1M Context, Memory Limits, Agent Infrastructure

This AINews roundup covers multiple AI product and research developments: Replit reportedly tripled to a $9B valuation and launched Replit Agent 4, a collaborative multi-agent canvas for apps, sites, and slides. NVIDIA released Nemotron 3 Super, an open 120B / ~12B-active model with a 1M-token context, hybrid Mamba‑Transformer/SSM Latent MoE architecture, and inference optimizations (including multi-token prediction) claiming up to ~2.2x faster inference versus gpt-oss-120B. The piece traces a broader 2026 trend from coding agents to general knowledge-work agents and highlights launches such as Perplexity’s Personal Computer, Base44 Superagents, and LangChain updates. It also reports Anthropic creating The Anthropic Institute (Jack Clark as Head of Public Benefit) and notes an operational outage affecting Claude/Claude Code. Research and benchmarks covered include agent evaluation work, retrieval/post‑training advances, Google Gemini Embedding 2, Qwen3.5 architecture notes, and device/benchmark reports (M5 Max).

Read assessment
Large Language Models (LLM) & AIJul 5, 2026

cliMEM adds persistent memory to CLI coding agents

Authors describe cliMEM, a local proxy that gives command-line coding agents persistent, per-project memory by intercepting agent requests, extracting distilled facts from chat logs, and storing them in Cognee (a graph + vector memory engine). Built by Team AIALCHEMISTS at a WeMakeDevs hackathon, cliMEM injects relevant remembered facts and a live file tree into new sessions so agents retain decisions, conventions, and open threads. The post recounts major implementation challenges (missing DB migrations, embedding provider API mismatches with NVIDIA NIM, tokenizer mapping issues with Jina) and pragmatic fixes: running Cognee migrations at startup, switching to local embeddings (fastembed) during the hackathon, and contributing an EMBEDDING_INPUT_TYPE config and provider detection patch for Cognee to support NVIDIA NIM. The team plans further hardening and to land the Cognee PR.

Read assessment

Track Real-Time Market Signals & Shifts

Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.