cd /news/artificial-intelligence/fine-tune-a-search-agent-with-multi-… · home › topics › artificial-intelligence › article
[ARTICLE · art-143984] src=aws.amazon.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

Fine-tune a search agent with multi-turn RL on Amazon SageMaker AI

Amazon Web Services detailed Amazon SageMaker AI multi-turn reinforcement learning (MTRL), a fine-tuning capability that trains LLM search agents across full multi-turn trajectories rather than single responses, using multi-turn rollouts and policy gradient algorithms. SageMaker AI MTRL offers a modular agent-environment interface, serverless per-token execution, asynchronous rollout with bounded off-policy staleness, and a native algorithm library including Proximal Policy Optimization (PPO), Clipped Importance Sampling Policy Optimization (CISPO), and importance-sampling losses paired with group-based advantage estimators such as GRPO, GRPO pass@k, and RLOO. AWS said the approach targets the gap left by supervised fine-tuning, which needs costly expert demonstrations, and single-turn RLVR, which scores one response at a time and misses interdependent decisions.

by read11 min views1 publishedOct 2, 2026
Fine-tune a search agent with multi-turn RL on Amazon SageMaker AI
Image: AWS ML Blog

Artificial Intelligence #

Search agents powered by large language models (LLMs) are transforming how enterprises retrieve information. Rather than requiring users to craft the perfect query, a search agent autonomously decides what to search for, which retrieval strategy to use, and when to stop searching. It does this across multiple rounds of interaction, refining its approach based on what it has already retrieved.

However, getting this multi-step behavior to work well is hard. No base model arrives knowing your tools or your environment. Prompt a small model and you rarely get dependable multi-turn behavior. Prompt a frontier model and it often works, but you pay for that capability in latency and cost. Fine-tuning offers a third path: you teach a small model your tools and environment directly. The result is a small model’s speed and cost with the reliability that would otherwise require a frontier model.

Even though fine-tuning is the natural next step, the traditional approaches each fall short. Supervised fine-tuning (SFT) depends on expert demonstrations of ideal multi-turn trajectories, which are costly to collect and usually don’t exist for your setup. Single-turn reinforcement learning (RL), such as RL with verifiable rewards (RLVR), scores one response at a time. But a search agent makes interdependent decisions across many turns, each one building on the context before it. Optimizing a single step in isolation misses those dependencies entirely.

What you need is a training approach that optimizes the agent across the full multi-turn trajectory. The reward signal only needs to reflect whether the final outcome was good. That is exactly what multi-turn reinforcement learning (MTRL) provides. It trains the agent to make good decisions across a full sequence of steps, bakes in your environment-specific behavior, and gives you control over output quality. Because you run a smaller, specialized model, you also get faster, cheaper inference.

In this post, we describe how we fine-tuned a search agent using Amazon SageMaker AI multi-turn reinforcement learning (MTRL) and share the results we observed in retrieval quality and reliability. We begin by explaining what Amazon SageMaker AI MTRL is and how it works.

What is Amazon SageMaker AI MTRL? #

With Amazon SageMaker AI MTRL, you can fine-tune LLMs using reinforcement learning in multi-turn interaction settings. It frames an agentic task as a sequence of decisions, uses multi-turn rollouts to generate training data, and optimizes the model with policy gradient algorithms.

Amazon SageMaker AI MTRL offers:

  • Modular agent-environment interface : You keep integration low-code. You define custom rewards, custom tool loops, and multi-turn conversation shapes.
  • Serverless execution: You get production-scale agentic RL at per-token pricing without provisioning or managing GPU clusters.
  • Asynchronous rollout and trajectory collection : You run generation and gradient updates in parallel with bounded off-policy staleness, so training stays fast without drifting too far from the current policy.
  • A native algorithm library: You choose from Proximal Policy Optimization (PPO), Clipped Importance Sampling Policy Optimization (CISPO), and importance-sampling (IS) losses, paired with group-based advantage estimators (GRPO, GRPO pass@k, RLOO, and more).
  • Resumable training : You can split long training runs across multiple jobs to work beyond single-job time limits.
  • Trajectory and reward observability : You can inspect what your agent did turn by turn and across training steps in MLflow managed by Amazon SageMaker AI.
  • Evaluation jobs : You can report reward, pass@k, and trajectory metrics before deploying to an Amazon SageMaker AI endpoint or Amazon Bedrock.

This makes MTRL a natural fit for search agents: you have a clear reward signal (retrieval quality), a multi-turn interaction loop (the agent issuing queries and receiving results), and a well-defined environment (the search tools).

Environment and setup #

Our setup relied on the following:

  • An AWS account with access to Amazon SageMaker AI in the US West (Oregon) AWS Region (us-west-2).
  • Training and validation datasets uploaded to Amazon Simple Storage Service (Amazon S3) in the required format (see the MTRL documentation for details).
  • A deployed agent endpoint that exposes BM25 and vector search tools for the MTRL environment to call during rollouts.
  • Familiarity with Amazon SageMaker AI and Python.

Solution overview #

In this post, we use Amazon SageMaker AI MTRL to fine-tune a Qwen3.6-27B model (supported in the US West (Oregon) Region (us-west-2)) for a search agent. The search agent is an LLM-powered system that autonomously uses search tools to find, gather, and synthesize information to answer a question or complete a task.

We focus on an enterprise search setting where the agent has two tools available:

  • Lexical search (BM25): Finds exact keyword matches by counting word frequencies. Best for queries with specific terms or identifiers.
  • Vector search: Converts queries and documents into embedding vectors and computes similarities. This is recommended for semantic or conceptual queries.

We also limit the number of turns (one turn is one round of user-assistant interaction), preventing the agent from generating excessively long responses and encouraging efficient search behavior.

Training setup #

This section walks through the three key components of our training setup: the datasets, the reward function, and the MTRL job configuration.

Datasets

We use the following datasets for training and testing:

| Dataset | Train/Test | Public Link | Description | | FRAMES | Train | HuggingFace | Multi-hop factoid QA requiring synthesis across multiple Wikipedia articles. | | BRIGHT | Train | HuggingFace | Reasoning-intensive retrieval across 12 domains, where relevance needs reasoning, not keyword overlap. | | Enterprise RAG | Train | GitHub | A benchmark of over 500,000 synthetic enterprise documents and 500 questions for training and evaluating RAG systems on internal company knowledge | | ESCI | Train | GitHub | A multilingual product-search dataset labeling query–product pairs as Exact, Substitute, Complement, or Irrelevant | | Musique | Train | GitHub | A question-answering dataset built by composing single-hop questions into challenging multi-hop reasoning problems. | | MLQA | Train | HuggingFace | A parallel extractive question-answering benchmark covering seven languages for evaluating cross-lingual comprehension. | | FreshStack | Test | HuggingFace | Recent developer-tech Q&A over Stack Overflow + docs for five topics. | | WixQA | Test | HuggingFace | Help-center support QA over the Wix knowledge base. | | BrowseComp-Plus | Test | GitHub | Hard deep-research queries over ~100K human-verified open-web documents. | | Wands | Test | HuggingFace | A Wayfair product-search dataset containing over 42,000 query–product relevance judgments labeled as Exact, Partial, or Irrelevant. |

We preprocess the datasets into the format recognized by the MTRL service according to the documentation. We further reserve 5 percent of the training instances within each dataset as validation instances.

Reward function

Our main metric is nDCG@10 (Normalized Discounted Cumulative Gain at rank 10). nDCG@10 is a standard information retrieval metric that measures how well the top 10 retrieved documents match the ideal ranking. It rewards systems that place highly relevant documents near the top of the list, with a score of 1.0 meaning perfect ranking and 0.0 meaning no relevant documents retrieved. We use nDCG@10 directly as the reward function in MTRL. This is a trajectory-level reward where the agent completes its full multi-turn search, and the reward reflects how well the final retrieved documents match the ground truth.

When the agent reaches the maximum number of turns or the maximum sampling tokens in a single turn, we assign a reward of -1 to explicitly teach the model to avoid these failure modes. This kind of penalty-based reward design is effective at steering the model away from undesirable behavior without needing to hand-craft complex intermediate rewards.

MTRL job configuration

One of the appeals of Amazon SageMaker AI MTRL is how little you need to configure to get going. We change three hyperparameters and leave everything else at the default values. The following code snippet shows how to configure and launch the training job using the MultiTurnRLTrainer SDK:

The key hyperparameters we change are:

  • max_epochs: 1 controls how many full passes over the training data.
  • global_batch_size: 128 sets the number of prompts per training step.
  • rollout_max_concurrency: 32 controls how many rollouts run in parallel during trajectory collection.

That’s the whole setup. The choices that usually demand RL expertise, the algorithm, the advantage estimator, the off-policy staleness bounds, all run on defaults. Getting from the base model to the results in the next section required nothing more than the preceding configuration.

Results #

MTRL fine-tuning improved the search agent on three of four held-out benchmarks and made it substantially more reliable across all of them. The largest gains came on BrowseComp-Plus (+23.7 percent nDCG@10) and WixQA (+18.4 percent), with a smaller gain on Wands and a slight regression on FreshStack. The reliability improvement was the more striking result. On BrowseComp-Plus, the failure rate fell from 22.89 percent to 0.68 percent. This means the agent learned not only to search better but also to finish the task within its turn and token budget. The sections that follow walk through the training curves and the per-benchmark numbers.

Training progress

The following figure shows the reward (nDCG@10) on both training and validation instances over the course of training. The x-axis represents training steps and the y-axis shows the nDCG@10 score. Two lines track training and validation reward, both rising steadily before plateauing. MTRL training can span multiple days and the MTRL service has a default time limit of 24 hours, which is adjustable through the CreateJob JSON schema. If a job is stopped or has failed because of timeout or infrastructure error, you can continue training from a previous checkpoint with resume support.

The figure plots nDCG@10 against training steps for both training and validation sets. Both curves increase steadily as fine-tuning progresses and saturate by the end of training, suggesting further training would not be beneficial.

Test performance

The following table compares the fine-tuned model against the original Qwen3.6-27B on four held-out test datasets. All metrics in the table are computed from our evaluation runs on the datasets listed in the preceding section.

Column definitions:

  • Size is the number of questions in the dataset.
  • nDCG@10 is the standardNormalized Discounted Cumulative Gain . The failed tasks are assigned nDCG@10 of 0.
  • Failure rate is the percentage of questions where the agent hits an error (for example, exceeding the turn limit or token budget).
  • Turns (avg) is the average number of turns per question.

| benchmark | size | model | ndcg@10 | fail_rate | turns | | WixQA | 400 | Qwen3.6-27B | 0.5725 | 0.67% | 4.3 | | WixQA | 400 | Qwen3.6-27B-finetune | 0.6781 | 0.17% | 4.5 | | Wands | 147 | Qwen3.6-27B | 0.5762 | 0.00% | 2.2 | | Wands | 147 | Qwen3.6-27B-finetune | 0.6112 | 0.00% | 2.9 | | FreshStack | 672 | Qwen3.6-27B | 0.4112 | 0.20% | 3.1 | | FreshStack | 672 | Qwen3.6-27B-finetune | 0.4089 | 0.05% | 2.8 | | BrowseComp-Plus | 830 | Qwen3.6-27B | 0.5136 | 22.89% | 7.0 | | BrowseComp-Plus | 830 | Qwen3.6-27B-finetune | 0.6354 | 0.68% | 6.3 |

The results show that RL fine-tuning made a meaningful difference:

  • Better search quality: As shown in the preceding table, nDCG@10 improved on WixQA (from 0.5725 to 0.6781, a +18.4 percent gain), on Wands (from 0.5762 to 0.6112, a +6 percent), and on BrowseComp-Plus (from 0.5136 to 0.6354, a +23.7 percent gain). The score decreased slightly on FreshStack.
  • Fewer failures: Assigning a reward of -1 for failure cases was highly effective. On BrowseComp-Plus, the failure rate dropped from 22.89 percent to only 0.68 percent.

Clean up #

If you do explore this approach, remember to clean up the resources afterward to avoid ongoing charges:

  • Stop or delete the MTRL training job in the Amazon SageMaker AI console if it’s still running.
  • Delete the model artifacts in Amazon S3 if they’re no longer needed.
  • Delete any deployed endpoints that were created for evaluation.

Conclusion #

In this post, we described our experience using Amazon SageMaker AI MTRL to fine-tune an LLM-powered search agent with reinforcement learning. With minimal configuration, the fine-tuned model showed substantial improvements in retrieval quality, failure rate, and turn efficiency (fewer turns per question).

Key takeaways from our experience with Amazon SageMaker AI MTRL:

  • Low barrier to entry: Default configurations work well, and no deep RL expertise is required to get started.
  • Serverless infrastructure: No need to provision GPUs or manage distributed training clusters. You pay per-token.
  • Resumable training: Long training runs can be split across multiple jobs.
  • Direct optimization: You train against your actual task metric rather than proxy losses.
  • Full observability: You can inspect trajectories turn by turn in MLflow managed by Amazon SageMaker AI to understand what the agent learned.

Next steps #

If you want to apply this approach to your own agents, here are some useful starting points:

1. Consult the [Amazon SageMaker AI MTRL documentation](https://docs.aws.amazon.com/sagemaker/latest/dg/model-customize-mtrl.html) .
2. Prepare your dataset in the [required format](https://docs.aws.amazon.com/sagemaker/latest/dg/model-customize-mtrl-assets.html#model-customize-mtrl-assets-prompt-dataset-format) .
  1. Prepare your agent and define your reward function based on your task-specific metric.
  2. Start with default configurations, which worked well in our experience.
  3. Evaluate and iterate.
── more in #artificial-intelligence 4 stories · sorted by recency
── more on @amazon web services 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/fine-tune-a-search-a…] indexed:0 read:11min 2026-10-02 · —