Observed Signal · Jun 25, 2026 · Technical Release · Source: DEV Community · Impact: 1/5 · Sentiment: Positive
Jarvis implements semantic memory with pgvector
A technical walkthrough of how the open-source Jarvis AI Platform implemented semantic memory retrieval in Java using embeddings and PostgreSQL's pgvector extension. The article describes the end-to-end memory pipeline: generating 768-dimensional embeddings locally with Ollama (nomic-embed-text), storing vectors in a vector(768) column, performing cosine-similarity search via pgvector, and assembling prompt context in a reactive Spring Boot application. It covers engineering decisions (JDBC for vector ops because R2DBC lacks vector support), performance numbers (embedding ~200ms, search <20ms), defenses against prompt-injection, HNSW indexing for document chunks, challenges building pgvector on Alpine Linux, and contributor guidance. The project is open source under Apache 2.0 (GitHub: sujankim/jarvis-ai-platform).
Practical, in-depth engineering guide for building semantic memory with embeddings and pgvector; useful to engineers building AI applications but not a major industry-shifting announcement for AdTech/MarTech.
Track Ollama Signals & Market Shifts in Real-Time
Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.
Key Takeaways & Evidence Grounding
- Jarvis AI Platform implemented semantic memory retrieval using embeddings and PostgreSQL's pgvector extension with a vector(768) column.
- The project uses Ollama's nomic-embed-text model (768-dimensional embeddings) running locally; embedding generation is ~200ms per text.
- JDBC is used for vector read/write and Flyway migrations because R2DBC does not natively support PostgreSQL's vector type.
- Document chunk search uses an HNSW index (idx_chunks_embedding_hnsw) for fast approximate nearest-neighbor search on embeddings.
- The memory system is open source under Apache 2.0 and available at https://github.com/sujankim/jarvis-ai-platform.
Connected Companies & Entities
3 Entities mapped“We use Ollama's nomic-embed-text model....”
“Whisper transcription is running via Groq API....”
“Session history (Redis ~1ms)...”
Ontology Mapping & Concepts
Related Market Signals & Shifts
Recent verified developments and strategic activity across this market segment.
PostgreSQL Semantic Search with pgvector
This technical guide explains how to implement semantic search directly inside PostgreSQL using the open-source pgvector extension. It covers the end-to-end flow: choosing an embedding model, storing embeddings alongside relational data, chunking long documents, generating embeddings (example using OpenAI), indexing options (HNSW and IVFFlat), distance operators (cosine, L2, inner product, etc.), and integrating with .NET via Npgsql and Pgvector. The author argues pgvector is a pragmatic choice for many applications when PostgreSQL is already the primary datastore, while recommending dedicated vector stores once scale, latency, or multi-tenant isolation requirements exceed Postgres’s operational fit. The piece emphasizes embedding-model compatibility, index tuning, and treating model changes as data migrations.
AI Agents Need Vector Databases for Memory
This technical blog post explains why retrieval-backed long-term memory for AI agents is best implemented with vector databases. It defines three memory types (working, long-term, episodic), outlines the memory stack (embedding model, vector store, chunking, metadata), recommends practical tooling (pgvector, Qdrant, Chroma) and embedding-dimension trade-offs, and provides a minimal Python example using pgvector and OpenAI embeddings. The author lists common production failure modes (stale memory, poor chunking, blind cosine similarity, context overflow, cost, privacy, and silent quality rot) and a practitioner's checklist for safe, private, and maintainable memory-enabled agents.
Service Layer for Production Vector Search
Part 4 of a technical series demonstrating a production-ready semantic search API built with Java, Spring Boot, PostgreSQL + pgvector, and the OpenAI embeddings API. The article explains the service layer's role in orchestrating document lifecycle and search pipelines: saving documents as PENDING, calling the embedding service, and updating status to READY or FAILED while recording errors. It describes a save-first/embed-second failure pattern, embedding and re-embedding on updates, a search flow that embeds queries and runs a two-layer SQL subquery to compute cosine distance and apply score thresholds, and why JPA alone is insufficient for dynamic vector search SQL. The post also covers a QueryBuilder helper, metadata filter validation to avoid injection, consistent global error responses, and performance benefits from lifecycle-driven indexing. The full reference implementation and tests are available on GitHub.
Track Real-Time Market Signals & Shifts
Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.
