Observed Signal · Jul 19, 2026 · Technical Release · Source: DEV Community · Impact: 4/5 · Sentiment: Positive

Building Semantic Search with GPT-5 and Microsoft Foundry

Executive Signal Summary

This technical tutorial demonstrates building a production-grade semantic search pipeline using GPT-5 for query planning and Microsoft Foundry's Foundry IQ (an agentic retrieval layer on Azure AI Search) for retrieval. The guide walks through creating a knowledge source (backed by blob storage and automatically indexed by Azure AI Search), configuring a knowledge base that uses an LLM deployment (example: gpt-5-mini) for query planning and answer synthesis, wiring the knowledge base into a Foundry agent via an MCP endpoint, and tuning retrieval_reasoning_effort to balance cost, latency, and answer quality. The article emphasizes multi-hop queries, iterative planning, and the trade-offs between extractive and synthesized answers.

Polaris7 AgentPolaris7 Strategic Assessment
High Confidence

Microsoft's Foundry IQ and GPT-5 integration represents a major platform-level capability for agentic, multi-hop semantic retrieval on Azure; this can materially change enterprise search, knowledge-driven automation, and how MarTech/AdTech systems perform document-level reasoning and synthesis, with implications for cost/latency and deployment patterns.

SIGNAL RADAR

Track Microsoft Signals & Market Shifts in Real-Time

Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.

Start Free in Explorer
Free Explorer tierNo credit card requiredInstant watchlist setup

Key Takeaways & Evidence Grounding

  • Microsoft Foundry's agentic retrieval layer is called Foundry IQ and is built on Azure AI Search.
  • The tutorial uses GPT-5 (example deployment 'gpt-5-mini') for query planning and answer synthesis within a Foundry knowledge base.
  • Creating a knowledge source (e.g., pointing at a Blob Storage container) triggers Azure AI Search to generate the index, skillset, and indexer to chunk and vectorize content automatically.
  • Each knowledge base exposes a standalone MCP endpoint that MCP-compatible clients (including Foundry Agent Service and GitHub Copilot) can call using the knowledge_base_retrieve tool.
  • The knowledge base setting retrieval_reasoning_effort (Minimal, Medium, High) controls LLM-driven planning, affecting latency, token consumption, and the ability to perform multi-hop queries and iterative retrieval.

Connected Companies & Entities

3 Entities mapped

“Microsoft Foundry's answer to this is Foundry IQ: an agentic retrieval layer built on Azure AI Search that treats retrieval as a reasoning t...”

“Any MCP-compatible client — including Foundry Agent Service, but also GitHub Copilot or other MCP clients — can call its `knowledge_base_ret...”

“KnowledgeBaseAzureOpenAIModel(azure_open_ai_parameters=AzureOpenAIVectorizerParameters(resource_url="https://<your-foundry-resource>.openai....”

Primary Source Grounding & Direct Attribution
Direct Origin Attribution
Primary Reporting: DEV Community•Published: Jul 19, 2026
Original Coverage Title: “Building Production-Grade Semantic Search with GPT-5 and Microsoft Foundry, From Scratch”

Related Market Signals & Shifts

Recent verified developments and strategic activity across this market segment.

PlatformAug 11, 2026

Build Meeting-Audio RAG on Microsoft Foundry

This technical walkthrough shows how to build an end-to-end retrieval-augmented generation (RAG) pipeline for recorded meetings using Microsoft Foundry. It combines Fast Transcription (synchronous diarized speech-to-text), chunking speaker turns with contextual summaries, indexing those JSONL transcripts into Foundry IQ (an Azure AI Search–based knowledge layer), and exposing a Foundry agent that answers questions with grounded citations and timecodes. The article covers project setup, authentication (Entra ID / managed identities), API endpoints and versions (fast transcription API v2025-10-15; agentic retrieval v2026-04-01 / v2026-05-01-preview), evaluation strategies (WER and judge models), production considerations (permissions, reindexing, tracing), and migration guidance from legacy Azure OpenAI “Add your data” flows (deprecated, retires October 14, 2026).

Read assessment
AI SearchMay 12, 2026

How Modern AI Search Engines Work

This technical article outlines the architecture and key components of modern AI-native search engines. It describes a multi-stage pipeline—query understanding, hybrid semantic retrieval (sparse + dense), contextual extraction and semantic chunking, reranking, model routing/orchestration, grounded response generation, streaming output, and caching/feedback loops—often implemented as Retrieval-Augmented Generation (RAG). The piece explains why hybrid retrieval (BM25/SPLADE plus dense embeddings) and rank fusion (e.g., RRF) are used, names common vector database and tooling options (FAISS, Pinecone, Milvus, Weaviate), and highlights reranking approaches (cross-encoder rerankers, open-source BGE rerankers, Cohere Rerank). It emphasizes semantic chunking and precision-focused reranking as methods to improve relevance, reduce token costs, and ground generated responses.

Read assessment
Site Search / Vector Search ImplementationAug 7, 2026

How to Build a Semantic Site Search Engine

Technical how-to describing a practical, efficient architecture for building semantic site search using embeddings and incremental indexing. The author recommends splitting the pipeline into four jobs (crawl, extract, index, serve), keeping raw HTML, hashing chunks to avoid re-embedding unchanged content, and serving queries with cached query embeddings plus a hybrid keyword+embedding merge. The guide covers content extraction heuristics, an example incremental reindex algorithm, latency budgeting for search boxes, and operational recommendations for running and migrating indexes and embedding models.

Read assessment

Track Real-Time Market Signals & Shifts

Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.