Ask any team that shipped a "chat with your docs" bot in 2024 what happened six months later, and you'll hear a familiar story: the chunking strategy needed retuning, the reranker was hand-rolled and brittle, permissions leaked because the vector index didn't respect SharePoint ACLs, and every new agent needed its own copy-pasted retrieval pipeline. Retrieval-Augmented Generation (RAG) was never really the hard part β productionizing RAG as a shared, governed, multi-tenant capability was.
Microsoft Foundry's answer to that problem is Foundry IQ, a managed knowledge layer released as the productized wrapper around Azure AI Search's agentic retrieval engine. It is one of the more quietly significant additions to the Foundry ecosystem this year, because it changes the unit of reuse in enterprise AI from "a RAG pipeline I built" to "a knowledge base I connect to N agents," with permission enforcement baked into the query path instead of bolted on after retrieval.
This article is a deep, implementation-level walkthrough of Foundry IQ: what it actually is, how the agentic retrieval pipeline works internally, how it plugs into Foundry Agent Service over MCP, what the security model really enforces (and doesn't), and where it breaks down at scale.
A Foundry model β even the largest ones you can deploy β has a knowledge cutoff and zero awareness of your tenant's SharePoint sites, blob containers, Fabric lakehouses, or internal wikis. The standard fix is RAG: chunk your documents, embed them, index them in a vector store, retrieve the top-k chunks at query time, and stuff them into the prompt.
The problem isn't the concept β it's everything around it:
Foundry IQ addresses this by promoting retrieval from "a pipeline you write" to "a first-class, shareable resource" β the knowledge base β that sits on top of Azure AI Search's agentic retrieval engine and is consumable by any number of Foundry agents (or Microsoft Agent Framework apps, or Copilot Studio agents) via a standard protocol.
Three objects matter here, and it's worth being precise about the layering because the docs use "Foundry IQ" and "agentic retrieval" almost interchangeably, which causes confusion.
A knowledge source is a top-level Azure AI Search resource describing where content comes from and how it's queried. Knowledge sources are either:
This indexed-vs-remote split matters architecturally: indexed sources trade freshness for query speed and semantic reranking quality; remote sources trade some query-time latency and reduced reranking control for zero duplication of source-of-truth data and native enforcement of the origin system's permission model.
A knowledge base is the orchestration object. It references one or more knowledge sources and holds the parameters that control retrieval behavior β most importantly the retrieval reasoning effort (minimal, low, or medium), which determines whether an LLM is used to plan/decompose the query before execution. Multiple agents can point at the same knowledge base. This is the reusable unit: build it once, govern it once, connect N agents to it.
Agentic retrieval is the actual multi-query pipeline that a knowledge base executes when called. It is a genuinely distinct pattern from naive single-query vector search:
minimal effort): an LLM β an Azure OpenAI deployment you configure on the knowledge base β takes the user's query plus conversation history and decomposes it into a set of focused subqueries. This is where "compare parental leave in the US and Germany" becomes two or three separate, well-formed sub-questions instead of one blurry embedding.
The important design decision here: agentic retrieval returns grounding data, not necessarily a final answer. Whether you consume it as raw extractive passages (GA path) or ask it to synthesize a natural-language answer (preview path) is your choice, made per knowledge base configuration.
At a component level:
| Component | Owning service | Role |
|---|---|---|
| Knowledge base | Azure AI Search | Orchestrates the pipeline; owns query parameters and reasoning effort |
| Knowledge source(s) | Azure AI Search | Define what content is queried and how |
| Search index | Azure AI Search | Backing store for indexed sources; holds text + vectors + semantic config |
| Semantic ranker | Azure AI Search | L2 reranking of subquery results |
| LLM | Azure OpenAI (via Foundry Models) | Powers query planning, web-result summarization, and answer synthesis |
| Foundry Agent Service | Microsoft Foundry | Consumes the knowledge base as an MCP tool from a PromptAgentDefinition |
Note what's not in this list: there is no separate "Foundry IQ service" runtime. Foundry IQ is the productized, governed front door β the naming and portal experience layer β over Azure AI Search's agentic retrieval, surfaced inside the Microsoft Foundry portal and consumable through Foundry Agent Service. This matters operationally: your quotas, region availability, and REST API versioning all live under Azure AI Search, not under a separate Foundry billing meter.
As of the 2026-04-01 GA REST API, Azure AI Search supports agentic retrieval for GA knowledge source types with minimal reasoning effort only (i.e., no LLM-based query planning, extractive results only). The 2026-08-01-preview REST API version unlocks preview knowledge source types (SharePoint, Fabric, MCP-as-source, Web), non-minimal reasoning effort (LLM query planning), answer synthesis, and multi-turn message arrays.
The Microsoft Foundry portal and Azure portal currently only expose the preview surface β meaning anything you wire up through the portal UI may need a deliberate migration pass before it's a supportable GA production configuration. If you're building for production today, decide explicitly which REST API version you're targeting rather than letting the portal default you into preview schemas you didn't intend to depend on.
This is the part developers most need to internalize: Foundry Agent Service talks to a Foundry IQ knowledge base exclusively through the Model Context Protocol. The knowledge base itself exposes an MCP endpoint:
{search_service_endpoint}/knowledgebases/{knowledge_base_name}/mcp?api-version=2026-08-01-preview
That endpoint exposes exactly one MCP tool today: knowledge_base_retrieve. Your agent's PromptAgentDefinition gets an MCPTool pointed at that endpoint via a project connection β a RemoteTool connection category with ProjectManagedIdentity auth, which is specific to Foundry project connections and lets the project's system-assigned managed identity authenticate to Azure AI Search without you juggling API keys in agent config.
This design has a consequence worth calling out explicitly: the knowledge base is not a Foundry-native resource β it's a remote tool the agent calls over network protocol. That means:
mcp_approval_request / tool-call spans), which is good for observability but means retrieval latency shows up as tool latency in your traces, not model latency β budget your P95 SLAs accordingly.
Here's the real, end-to-end path: create the connection, then create the agent, then call it. This mirrors the officially supported pattern (Python SDK β₯ 2.0.0, REST API 2026-08-01-preview for the knowledge base MCP endpoint, 2025-10-01-preview for the ARM connection).
import requests
from azure.identity import DefaultAzureCredential, get_bearer_token_provider
credential = DefaultAzureCredential()
project_resource_id = (
"/subscriptions/<sub-id>/resourceGroups/<rg>/providers/"
"Microsoft.CognitiveServices/accounts/<account>/projects/<project>"
)
project_connection_name = "hr-kb-mcp-connection"
mcp_endpoint = (
"https://hr-search-svc.search.windows.net/knowledgebases/hr-policy-kb/mcp"
"?api-version=2026-08-01-preview"
)
bearer_token_provider = get_bearer_token_provider(
credential, "https://management.azure.com/.default"
)
headers = {"Authorization": f"Bearer {bearer_token_provider()}"}
response = requests.put(
f"https://management.azure.com{project_resource_id}/connections/{project_connection_name}"
"?api-version=2025-10-01-preview",
headers=headers,
json={
"name": project_connection_name,
"type": "Microsoft.MachineLearningServices/workspaces/connections",
"properties": {
"authType": "ProjectManagedIdentity",
"category": "RemoteTool",
"target": mcp_endpoint,
"isSharedToAll": True,
"audience": "https://search.azure.com/",
"metadata": {"ApiType": "Azure"},
},
},
)
response.raise_for_status()
print(f"Connection '{project_connection_name}' created or updated successfully.")
Before this call succeeds in a real tenant, three RBAC assignments have to be in place:
That third one trips people up: it's Azure AI Search's identity, not the agent's, that needs the model-calling permission, because Search is the thing invoking the LLM mid-pipeline for query planning β the agent never sees that intermediate call.
from azure.ai.projects import AIProjectClient
from azure.ai.projects.models import PromptAgentDefinition, MCPTool
from azure.identity import DefaultAzureCredential
credential = DefaultAzureCredential()
project_endpoint = "https://hr-foundry.services.ai.azure.com/api/projects/hr-project"
mcp_endpoint = (
"https://hr-search-svc.search.windows.net/knowledgebases/hr-policy-kb/mcp"
"?api-version=2026-08-01-preview"
)
project_connection_name = "hr-kb-mcp-connection" # from step 5.1
project_client = AIProjectClient(endpoint=project_endpoint, credential=credential)
instructions = """
You are a helpful HR assistant that must use the knowledge base to answer all
questions from the user. You must never answer from your own knowledge under
any circumstances.
Every answer must include citations for the knowledge base sources you used,
rendered as: [message_idx:search_idx | source_name]
If the knowledge base does not contain the answer, respond with "I don't know".
"""
mcp_kb_tool = MCPTool(
server_label="knowledge-base",
server_url=mcp_endpoint,
require_approval="never", # skip human-in-the-loop approval per call
allowed_tools=["knowledge_base_retrieve"], # explicit allow-list β the only tool exposed anyway
project_connection_id=project_connection_name,
)
agent = project_client.agents.create_version(
agent_name="hr-policy-assistant",
definition=PromptAgentDefinition(
model="gpt-4.1-mini",
instructions=instructions,
tools=[mcp_kb_tool],
),
)
print(f"Agent '{agent.name}' version {agent.version} created.")
A few implementation details worth flagging:
require_approval="never" is a real security decision, not a formality. Because the knowledge base's only exposed tool is a read-only retrieval call, blanket auto-approval is defensible here β but if you later add a knowledge source or MCP-server-as-source configuration that can trigger side effects (e.g., a Fabric Data Agent that can kick off a query job with cost implications), revisit this.allowed_tools=["knowledge_base_retrieve"] looks redundant since that's the only tool the endpoint exposes today, but it's cheap insurance against a future server-side addition silently expanding your agent's capability surface.
from azure.ai.projects.models import ResponsesHostServer # illustrative import path
response = project_client.agents.responses.create(
agent_name="hr-policy-assistant",
input="How does parental leave differ between our US and Germany offices, "
"and which of our team leads have direct reports in Germany?",
)
print(response.output_text)
Under the hood, this single call triggers: query decomposition into "US parental leave policy" and "Germany parental leave policy" and "team leads with direct reports in Germany" sub-questions (assuming low/ medium reasoning effort), three parallel hybrid searches possibly across two different knowledge sources (an HR policy blob index and an org-chart SQL-backed index), semantic reranking of each, and a synthesized, citation-bearing answer handed back to the agent's model to finalize.
This is the section worth reading twice before you put Foundry IQ in front of sensitive data.
The permission story has two genuinely different enforcement mechanisms depending on knowledge source type, and conflating them is a common and dangerous mistake:
For indexed sources (Blob, SQL, OneLake, Indexed SharePoint): permission enforcement requires you to have synchronized access control list (ACL) metadata fields into your search index yourself, and to pass the calling user's identity via the x-ms-query-source-authorization header at query time so Search can filter results per-caller. This is query-time RBAC/ACL enforcement, and as of this writing it is itself a preview capability layered on top of the base agentic retrieval feature. If you don't populate those permission fields and don't forward that header, your index has zero awareness of who's asking β it will happily surface a document to anyone whose query is a good semantic match, regardless of whether that person should be able to read it.
For remote SharePoint sources specifically: content isn't indexed into Azure AI Search at all. Instead, the same authorization header is forwarded to SharePoint's own Copilot Retrieval API, and SharePoint enforces its native permission model directly at query time. No duplicated data, no duplicated ACL sync job, no drift between the source system's permissions and a stale index copy.
The practical takeaway: "Foundry IQ enforces permissions" is true only if you built the plumbing for it. The platform gives you the mechanism β header propagation, ACL metadata schema support, Purview sensitivity label honoring for supported sources β but it does not retroactively secure an index you built without permission metadata. Teams migrating an existing, permission-naive Azure AI Search index into a Foundry IQ knowledge source need to treat ACL backfill as a hard blocking prerequisite, not a nice-to-have.
There's a second, more subtle risk: who runs the query planning LLM call, and against what? Query decomposition sends the user's raw query and conversation history to the configured Azure OpenAI model. If your conversation history contains sensitive context from a prior turn (say, a previous answer that quoted a restricted document), that content is now part of the payload sent to the planning model, which may sit in a different resource/network boundary than the eventual retrieval target. Model input/output logging and data residency policy on that Azure OpenAI deployment therefore becomes part of your knowledge base's overall data-handling boundary β not an implementation detail you can ignore because "it's just doing retrieval."
Consider a mid-size enterprise with HR policy PDFs in SharePoint, a structured headcount/org-chart table in Azure SQL, and a benefits FAQ maintained as a Fabric lakehouse table. Before Foundry IQ, three separate retrieval pipelines, three separate embedding refresh jobs, and three separate places for permissions to silently diverge from the source systems.
With Foundry IQ:
hr-policy-kb, references all of them, with reasoning effort set to low so multi-part questions get decomposed.
This is the actual value proposition in concrete terms: not "better RAG," but organizational reuse of a governed retrieval capability across otherwise-independent agent teams.
2026-04-01 GA (stability, but 2026-08-01-preview (LLM query planning, answer synthesis, preview sources) and document the migration path before you have production traffic depending on preview-only behavior.medium) based on your actual SLA, not just answer quality.
Foundry IQ has no independent billing meter β costs roll up through the underlying services it composes:
The practical implication: a single user question against a medium-reasoning-effort knowledge base with three knowledge sources can trigger one planning LLM call, three-plus parallel search queries, three reranking passes, and one synthesis LLM call β all before your agent's own model ever generates a token. Budget and monitor this as a distinct cost center from your agent's model spend, and prefer minimal or low reasoning effort for high-volume, latency- and cost-sensitive endpoints where query complexity doesn't warrant full decomposition.
knowledge_base_retrieve, especially on questions the model "thinks" it already knows the answer to β exactly the class of question where a stale or wrong parametric answer is most dangerous.
| Approach | When it makes sense | Trade-off vs. Foundry IQ |
|---|---|---|
| Hand-rolled RAG pipeline (your own chunker, embedder, vector store, reranker) | Highly specialized domains (e.g., genomics, legal citations) where generic chunking/embedding underperforms, or when you need a non-Azure vector store | Full control, but you own query decomposition, reranking, ACL enforcement, and multi-source orchestration yourself β and you rebuild it per agent unless you invest separately in your own shared-service layer |
| Foundry Agent Optimizer prompt/tool tuning alone (no retrieval layer) | Small, static knowledge bases that fit comfortably in a system prompt or a single small index | Doesn't scale past a few documents; no dynamic multi-source retrieval, no citation infrastructure |
| Direct Azure AI Search agentic retrieval without Foundry IQ framing | Non-Foundry applications, or when you need the GA 2026-04-01 REST API surface without any Foundry-specific portal/MCP conventions |
Same underlying engine, but you lose the Foundry-portal knowledge-base authoring UX and the standardized MCP tool contract for Foundry agents specifically |
| Vendor RAG-as-a-service platforms outside Azure | Multi-cloud strategies or existing investment in another vector database ecosystem | Loses native ACL/Purview integration and native MCP wiring into Foundry Agent Service; you're back to building your own bridge |
The honest framing: Foundry IQ doesn't introduce a fundamentally new retrieval algorithm β multi-query decomposition, parallel hybrid search, and semantic reranking are all patterns you could implement yourself against Azure AI Search directly, or against any vector database. What it buys you is standardization and reuse: one governed object, one permission model, one MCP contract, consumable by every agent in your tenant instead of reinvented per team.
Foundry IQ is best understood not as a new retrieval technology but as Microsoft formalizing agentic retrieval β Azure AI Search's multi-query, LLM-assisted retrieval pipeline β into a reusable, governed, MCP-addressable resource inside the Foundry ecosystem. The technical substance (query planning, parallel hybrid search, semantic reranking, optional answer synthesis) has existed in Azure AI Search independently; what Foundry IQ adds is the organizational contract: one knowledge base, many agents, one place to reason about permissions, cost, and quality.
The permission story is genuinely good when built correctly β ACL sync plus header propagation plus Purview label honoring is a real, defensible security model β but it is opt-in machinery, not a default you get for free by pointing an agent at an index. Teams that skip the ACL and header plumbing get a bot that looks secure and isn't. Teams that respect the preview/GA API boundary, budget for the extra LLM calls query planning and synthesis introduce, and treat instructions as a tested contract rather than a one-off prompt will get real leverage: a knowledge layer that multiple agent teams can build on without each reinventing retrieval from scratch.
If you're currently maintaining more than one hand-rolled RAG pipeline against overlapping enterprise content inside the same tenant, that's the strongest signal that Foundry IQ's reuse model is worth the migration effort.
(Note: Some figures and API version references above reflect documentation current as of late September 2026 and reference preview features that may change before general availability β verify against current Microsoft Learn documentation before building production systems.)