{"slug": "agentic-retrieval-with-langchain-and-amazon-bedrock-knowledge-bases", "title": "Agentic retrieval with LangChain and Amazon Bedrock Knowledge Bases", "summary": "Amazon Bedrock Managed Knowledge Bases now offers agentic retrieval through the AgenticRetrieveStream API, which plans a retrieval loop by breaking a question into sub-queries, running them, judging whether it has enough evidence, and searching again if it doesn't, according to an AWS machine-learning blog post. The post contrasts this with the single-shot Retrieve API, which runs one hybrid search and returns scored chunks, and notes that the langchain-aws package exposes both paths so developers can choose either from a LangChain application. The walkthrough uses the US East (N. Virginia) Region (us-east-1) and covers what each retrieval path costs and when the cheaper standard path is the right choice.", "body_md": "## [Artificial Intelligence](https://aws.amazon.com/blogs/machine-learning/)\n\n# Agentic retrieval with LangChain and Amazon Bedrock Knowledge Bases\n\nWhen a user asks the support assistant, a Retrieval Augmented Generation (RAG) application built with LangChain to compare two products across three dimensions, they’re effectively posing six questions simultaneously. Similarity search uses a single query vector to encapsulate all the intents. The retriever then generates the best approximation of the average of those intents. The resulting answer comes back concise. The search executes without errors. The relevance scores look reasonable. Yet the retrieved chunks, while topically relevant, only cover a fraction of what the question actually asked.\n\nIn this post, we showcase a RAG application on Amazon Bedrock Managed Knowledge Base with [LangChain](https://github.com/langchain-ai/langchain-aws). We run the same multi-part question through standard and agentic retrieval, and read the trace events to see the plan the model produced. We also cover what the two retrieval paths cost and when the cheaper one is the right choice.\n\nAgentic retrieval is available on Amazon Bedrock Managed Knowledge Base. Instead of one search, Amazon Bedrock Managed Knowledge Base plans the retrieval. It breaks the question into sub-queries, runs them, judges whether it has enough evidence, and searches again if it doesn’t. The `langchain-aws` package exposes both agentic and standard retrieval, so you can use either from a LangChain application.\n\n## Solution overview\n\n[Amazon Bedrock Managed Knowledge Base](https://aws.amazon.com/blogs/aws/introducing-amazon-bedrock-managed-knowledge-base-for-faster-more-accurate-enterprise-ai-applications/), the fully managed RAG capability in Amazon Bedrock, removes the self-managed vector store, embeddings, and re-ranking models from the RAG architecture. You configure a data source, and Amazon Bedrock Managed Knowledge Bases handles chunking, embedding, storage, and retrieval. This walkthrough uses [Amazon Simple Storage Service (Amazon S3)](https://aws.amazon.com/pm/serv-s3/?trk=50b671a1-06f5-4224-9505-fa45ee881c08&sc_channel=ps&ef_id=EAIaIQobChMI1uiUo4eXlgMVy1J_AB2pfydUEAAYASAAEgJ6PPD_BwE&gads_camp=23522747487&gads_ag=196433733807&gads_ad=795876995201&gads_kw=amazon%20s3&gads_matchtype=e&gads_network=g&gads_device=c&gads_geo=9022830&gad_campaignid=23522747487&gbraid=0AAAAADjHtp-DJgH1kF6NUslUNqoqqj14q&gclid=EAIaIQobChMI1uiUo4eXlgMVy1J_AB2pfydUEAAYASAAEgJ6PPD_BwE).\n\nAmazon Bedrock Managed Knowledge Bases provides two APIs. We briefly discuss those differences in this post. The [Retrieve](https://docs.aws.amazon.com/bedrock/latest/APIReference/API_agent-runtime_Retrieve.html) API runs one hybrid search and returns scored chunks. The [AgenticRetrieveStream](https://docs.aws.amazon.com/bedrock/latest/APIReference/API_agent-runtime_AgenticRetrieveStream.html) API runs a planning loop and streams the steps back to you as trace events. In the `langchain-aws` package, the first is a standard LangChain retriever you can drop into a chain. The second is a function retrieval directly from a knowledge base.\n\nThe following diagram shows the solution architecture. The application queries Amazon Bedrock Knowledge Bases using either the Retrieve API (standard, single-shot) or the AgenticRetrieveStream API (multi-step planning loop). Both paths return document chunks from the knowledge base, which the application then uses to generate a grounded response.\n\n## Implementation walkthrough\n\nThe following sections walk you through creating a knowledge base, querying it with both retrieval methods, and reading the trace events the agentic planner produces.\n\n## Prerequisites\n\nTo follow along you need:\n\n- An AWS account with access to Amazon Bedrock in a Region where Amazon Bedrock Managed Knowledge Bases and agentic retrieval are available. This walkthrough uses the US East (N. Virginia) Region (`us-east-1` ), and the code assumes it throughout. Check the AWS[documentation](https://docs.aws.amazon.com/bedrock/latest/userguide/models-region-compatibility.html) for other Regional availability and support.\n- Two AWS Identity and Access Management (IAM) identities, described in the next section: a service role the knowledge base assumes, and permissions on the identity you call the APIs from.\n- [Python 3.12](https://www.python.org/downloads/) or later.\n- An S3 bucket holding the sample documents. The corpus needs several documents that cover overlapping topics so that a comparative question has somewhere to go. A single flat document cannot demonstrate query planning.\n\nInstall the packages. The [Boto3](https://boto3.amazonaws.com/v1/documentation/api/latest/index.html) version matters: `agentic_retrieve_stream` did not exist before 1.43.32.\n\n## Permissions\n\nTwo identities are involved and separating them is worth doing deliberately. The knowledge base assumes a service role to read your documents and call the embedding model. Your application uses an AWS Security Token Service (AWS STS) caller identity to query. Neither needs the other’s permissions.\n\nAmazon Bedrock creates the service role for you if you let it. To supply your own, give it a trust policy that lets Amazon Bedrock assume it. Scope it with `aws:SourceAccount` and `aws:SourceArn` so that another account can’t use it as a confused deputy:\n\nThe service role also needs `s3:ListBucket` on your bucket and `s3:GetObject` on its contents, both conditioned on `aws:ResourceAccount`. Scope the `knowledge-base/*` wildcard character down to specific knowledge base IDs after you have created them.\n\nThe AWS STS caller identity needs a different set. `bedrock:AgenticRetrieveStream` and `bedrock:InvokeModelWithResponseStream` can’t be scoped to a knowledge base Amazon Resource Name (ARN). `bedrock:Retrieve` and `bedrock:GetDocumentContent` can:\n\n`bedrock:GetDocumentContent` is often overlooked. Agentic retrieval calls it when a `FullDocumentExpansion` step decides a passage lacks the context to answer. A policy with only `bedrock:Retrieve` works until the planner reaches for a whole document and then fails partway through a query.\n\nTo create and manage the knowledge base itself, the calling role additionally needs `bedrock:CreateKnowledgeBase` on `*`, and the `GetKnowledgeBase`, `UpdateKnowledgeBase`, `DeleteKnowledgeBase`, `StartIngestionJob`, `GetIngestionJob`, and `ListIngestionJobs` actions on `knowledge-base/*`. If you’re using guardrails, add `bedrock:GetGuardrail` and `bedrock:ApplyGuardrail`.\n\nRunning this walkthrough might incur costs for document storage and ingestion in the knowledge base, retrieval calls, and foundation model (FM) inference.\n\nFor more information about pricing, see the Knowledge Bases section of [Amazon Bedrock pricing](https://aws.amazon.com/bedrock/pricing/).\n\nDelete the resources when you complete this experiment.\n\n## Creating and populating the knowledge base\n\nCreate the knowledge base with a `managedKnowledgeBaseConfiguration`. Setting `embeddingModelType` to `MANAGED` uses the service-managed embedding model.\n\nThere’s no `storageConfiguration` in that request. For a self-managed knowledge base you would pass one describing your vector store. Amazon Bedrock Managed Knowledge Base does not take one, which is the clearest signal in the API that Amazon Bedrock owns the storage layer.\n\nAttach the S3 bucket as a data source, then start an ingestion job. Ingestion is asynchronous, so poll until the job reaches a terminal state rather than sleeping for a fixed interval and hoping.\n\nThe full data source configuration and error handling are in the [sample repository](https://github.com/aws-samples/sample-rag-bedrock-langchain-blog).\n\n## Querying with the LangChain retriever\n\n[AmazonKnowledgeBasesRetriever](https://reference.langchain.com/python/langchain-aws/retrievers/bedrock/AmazonKnowledgeBasesRetriever) wraps the Retrieve API and behaves like any other LangChain retriever. For Amazon Bedrock Managed Knowledge Bases, pass `managedSearchConfiguration`. This is the part that trips people up: `vectorSearchConfiguration` is the previous path for knowledge bases where you run your own vector store. It is what most existing examples show.\n\nEach result comes back as a LangChain `Document`. The relevance score is in `metadata[\"score\"]`, and the source document’s own metadata is under `metadata[\"source_metadata\"]`, renamed so it does not collide. If you want to drop low-confidence results, set `min_score_confidence` on the retriever instead of filtering afterward.\n\nFor a question with one clear intent, this is the right tool. It is one call. The latency is the lowest of the two options, and you keep full control of how the answer gets generated. Most of the queries a production assistant sees are this shape, and reaching for a planning loop to answer them wastes money and time.\n\n## Where single-shot retrieval runs out\n\nNow give the same retriever a question with several parts:\n\nFive chunks come back, ranked by hybrid score against one embedding of that whole question.\n\nThat question contains six intents: two services across three dimensions. Scoring the retrieved text for evidence of each one gives a concrete measure of what a single embedding recovers.\n\n| **numberOfResults** | **Chunks** | **Share of corpus** | **Sub-intents covered** | **Missing** | \n| 5 | 5 | 10% | 4 of 6 | checkout on-call, inventory restore | \n| 10 | 10 | 19% | 6 of 6 | none | \n\nAt five results, one embedding standing for six intents misses two of them. At ten it covers all six, with visible waste: two sub-intents are covered twice and one chunk carries none.\n\nThe retriever did its job. The limitation is structural: one vector cannot represent six intents, and there is no step in the process that asks whether the returned evidence is enough to answer the question.\n\n## Running agentic retrieval\n\nAgentic retrieval is not a LangChain retriever, but a feature of Amazon Bedrock Managed Knowledge Bases. The `langchain-aws` package exposes it as a standalone function, [agentic_retrieve](https://python.langchain.com/api_reference/aws/retrievers/langchain_aws.retrievers.bedrock.html), because the underlying API streams its results and doesn’t fit the synchronous `BaseRetriever` interface. There’s no flag on `AmazonKnowledgeBasesRetriever` that switches it on.\n\nWith `generate_response=True`, the service returns a grounded answer and citations alongside the retrieved chunks, so you get an answer without wiring up a separate model call. The function works only against Amazon Bedrock Managed Knowledge Base.\n\nInternally, the service plans, retrieves, evaluates whether the evidence is sufficient, and iterates if it isn’t. The helper hides all of that and hands back the final chunks, which is convenient and means you can’t see the plan.\n\n## Reading the trace events\n\nTo watch the model decompose the question, call [agentic_retrieve_stream](https://docs.aws.amazon.com/boto3/latest/reference/services/bedrock-agent-runtime/client/agentic_retrieve_stream.html) on the [bedrock-agent-runtime](https://docs.aws.amazon.com/boto3/latest/reference/services/bedrock-agent-runtime.html) client directly. This is the one place in this walkthrough where we step around `langchain-aws`, because the helper discards trace events and doesn’t expose `maxAgentIteration` or a custom planner model.\n\nSet `generateResponse` to `False` when you only want the retrieval behavior. The API generates a grounded answer by default, which costs an extra model call you may not need while you are inspecting the plan.\n\nA trace event’s `step` tells you where the planner is. `SpeculativeRetrieval` runs before the first plan to cut latency and doesn’t count against your iteration budget. `Planning` is where the model reads the question and prior results and emits sub-queries. `Retrieval` fires once per sub-query. `FullDocumentExpansion` appears when the model decides a passage lacks the context to answer and pulls the whole document instead. Each carries a status of `IN_PROGRESS`, `SUCCEEDED`, or `FAILED`, plus a human-readable message.\n\nThe final chunks arrive separately. The `result` is its own event type rather than a fifth step, and it holds the deduplicated chunks from every iteration along with the grounded answer when response generation is on. Branch on the event key, as the preceding loop demonstrates, rather than expect a terminal step value.\n\nThe sub-query text is the part worth logging. It sits in `attributes.actions[].retrieve.inputQuery.text`, not in the top-level trace fields, so a handler that reads only `step` and `status` shows you that planning happened without showing you what it decided.\n\nThe following diagram shows the agentic retrieval planning loop, including the speculative retrieval, planning, sub-query retrieval, evaluation, and optional re-planning steps.\n\nTwo details are worth knowing before you build on this. Deduplication applies only to the `result` event, so a chunk retrieved by three sub-queries appears once at the end but three times across the traces.\n\nThe second is about scores. A Retrieve response gives each chunk a typed `score` field holding its relevance to the query. Agentic retrieval results carry `content`, `metadata`, and `sourceRetriever`, with no equivalent typed field. Code that reads `result[\"score\"]` after switching APIs gets nothing. If you rank or filter relevance, plan for that difference.\n\nIn production, use Amazon Bedrock Guardrails to enforce content policies and grounding checks on generated responses. Both retrieval paths support guardrails. Agentic retrieval supports guardrails through `policyConfiguration.bedrockGuardrailConfiguration` rather than the `guardrail_config` argument the LangChain retriever takes, and supports `BLOCK` mode only. If you rely on `MASK` mode, that is a reason to stay on the Retrieve API.\n\n`maxAgentIteration` accepts two through ten and defaults to five. Leave it at the default. At two or three the planner runs one cycle, emits no sub-queries, and returns what the speculative retrieval step already found. This is single-shot behavior at the agentic price. Decomposition begins at four. The planner often stops early when it judges the evidence sufficient, so the ceiling is a bound rather than a target.\n\n## Comparing the two retrieval paths\n\nFor context on how this behaves at scale, AWS evaluated agentic retrieval on MuSiQue, a public multi-hop benchmark. The evaluation showed improved recall over single-shot retrieval, with the largest gains on the hardest questions. Single-hop questions saw gains under five points. That last figure matches the shape of the trade-off: decomposition helps when there is something to decompose.\n\n## Building the RAG chain\n\nFor the standard retriever, the usual [LangChain Expression Language (LCEL)](https://python.langchain.com/docs/concepts/lcel/) composition works directly:\n\n`format_docs` matters more than it looks. Passing `Document` objects straight into a prompt renders their `repr`, and the model gets metadata noise mixed into the context.\n\nTo put agentic retrieval in the same position, wrap it in a `RunnableLambda`, since it is a function rather than a retriever:\n\nNote that `generate_response` is off here. The service can generate the answer itself, but inside a chain you usually want your own prompt and model, so you take the chunks and generate downstream. Use the service generation when you want one call and less code, and the wrapped version when the prompt is yours to control.\n\n## Choosing between standard and agentic retrieval\n\nUse Retrieve for short, well-scoped questions. It is cheaper, faster, works against self-managed knowledge bases, and returns scores in the results. Most production traffic looks like this.\n\nUse AgenticRetrieveStream when questions are multi-part, comparative, or exploratory, or when the evidence spans more than one knowledge base. It registers up to five knowledge bases in one request and routes sub-queries using a natural-language description you attach to each. The other API cannot do this at all. It costs more per call, makes several model invocations, and has the higher latency of the two.\n\nRouting on query shape rather than picking one for everything is the pattern we recommend. A classifier or a heuristic on the question can send most traffic down the cheap path and reserve the planner for questions that need it.\n\n## Clean up resources\n\nDelete the knowledge base, its data source, the S3 objects and bucket, and the IAM role that you created. A knowledge base with documents in it continues to incur storage charges.\n\nThe [repository](https://github.com/aws-samples/sample-rag-bedrock-langchain-blog) includes a cleanup script that also empties the bucket and removes the role.\n\n## Conclusion\n\nWe showed how to build a RAG application on Amazon Bedrock Knowledge Bases with LangChain, and how agentic retrieval handles multi-part questions that single-shot retrieval answers poorly. We also showed the friction in the current integration. Agentic retrieval is a function rather than a LangChain retriever, so it needs a `RunnableLambda` to sit in a chain. The trace events that show the query plan require a direct `boto3` call.\n\nAgentic retrieval trades higher per-call cost for improved recall on multi-hop questions, using a built-in model for query planning. The next useful step is measuring your own query mix before you route everything through a planner.\n\nTo get started, see the Amazon Bedrock Knowledge Bases [documentation](https://docs.aws.amazon.com/bedrock/latest/userguide/kb-build-managed.html) and the accompanying sample code. For help applying this to your own workload, contact your AWS account team.", "url": "https://wpnews.pro/news/agentic-retrieval-with-langchain-and-amazon-bedrock-knowledge-bases", "canonical_source": "https://aws.amazon.com/blogs/machine-learning/agentic-retrieval-with-langchain-and-amazon-bedrock-knowledge-bases/", "published_at": "2026-10-05 15:53:56+00:00", "updated_at": "2026-10-05 16:20:11.132561+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-agents", "ai-infrastructure", "developer-tools", "large-language-models"], "entities": ["Amazon Bedrock Managed Knowledge Bases", "LangChain", "langchain-aws", "Amazon Web Services", "Amazon S3", "AgenticRetrieveStream API", "Retrieve API", "us-east-1"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/agentic-retrieval-with-langchain-and-amazon-bedrock-knowledge-bases", "markdown": "https://wpnews.pro/news/agentic-retrieval-with-langchain-and-amazon-bedrock-knowledge-bases.md", "text": "https://wpnews.pro/news/agentic-retrieval-with-langchain-and-amazon-bedrock-knowledge-bases.txt", "jsonld": "https://wpnews.pro/news/agentic-retrieval-with-langchain-and-amazon-bedrock-knowledge-bases.jsonld"}}