LLM optimization integration for Amazon SageMaker Python SDK Amazon SageMaker Python SDK v3.17.0 now integrates generative AI inference recommendations, allowing users to benchmark endpoints, generate deployment recommendations, and deploy optimized configurations directly from notebooks. The new sagemaker.serve.ai_inference_recommender package automates benchmarking, recommendation generation, and deployment, reducing manual trial and error for optimal instance and framework selection. Artificial Intelligence /blogs/machine-learning/ LLM optimization integration for Amazon SageMaker Python SDK Optimizing generative AI inference deployments requires benchmarking endpoints, evaluating instance configurations, and iterating on deployment settings. The Amazon SageMaker Python SDK v3 now exposes generative AI inference recommendations in Amazon SageMaker AI /blogs/machine-learning/amazon-sagemaker-ai-now-supports-optimized-generative-ai-inference-recommendations/ directly in your notebook workflow. These recommendations are also accessible through the Amazon SageMaker AI UI /blogs/machine-learning/launching-ui-for-generative-ai-inference-recommendations-in-amazon-sagemaker-ai/ and Boto3 APIs https://docs.aws.amazon.com/sagemaker/latest/dg/instance-recommendation-create.html . With this release, you can benchmark an endpoint, generate data-driven deployment recommendations, and deploy the recommended configuration directly from a notebook https://github.com/aws-samples/sagemaker-genai-hosting-examples/blob/main/03-features/gen-ai-inference-recommendations/pysdk-ai-inference-recommender-demo.ipynb using the Amazon SageMaker Python SDK v3. In this post, we demonstrate how to use the new SDK interface for the end-to-end workflow to optimize generative AI inference deployments. Benefits of generative AI inference recommendations in Amazon SageMaker AI Generative AI inference recommendations in Amazon SageMaker AI automate inference optimization by: - Benchmarking a live Amazon SageMaker endpoint against a synthetic or real-traffic workload, measuring throughput, time-to-first-token TTFT , end-to-end latency, and more. - Generating deployment recommendations ranked by cost-performance tradeoff using your actual usage patterns. - Deploying the top-ranked configuration directly to an Amazon SageMaker real-time endpoint. Previously, these capabilities required using Amazon SageMaker Studio or constructing AWS SDK for Python Boto3 API calls. With this launch, they become Python SDK operations, fitting naturally into existing notebook and pipeline workflows. New SDK interfaces The new functionality is available under the sagemaker.serve.ai inference recommender https://sagemaker.readthedocs.io/en/3.17.0/api/generated/sagemaker.serve.ai inference recommender.html package starting with version 3.17.0 and exposes the following primary operations: Entry point | What it does | | ModelBuilder.from jumpstart config … | Builds a ModelBuilder from a JumpStart model ID and compute config | | start benchmark endpoint, … | Runs a load test against a deployed endpoint with a configurable synthetic workload | | mb.generate deployment recommendations … | Explores instance/framework configs against your workload and returns ranked recommendations | | mb.deploy … | Deploys the top recommendation to a real-time endpoint | | ModelBuilder.from recommendation job job name | Hydrates a ModelBuilder from a completed recommendation job — deploy in a different process or session | Prerequisites Verify you have the latest version of the Amazon SageMaker Python SDK installed: You will also need: - An AWS account with an AWS Identity and Access Management IAM role with Amazon SageMaker execution permissions https://docs.aws.amazon.com/sagemaker/latest/dg/security-iam-awsmanpol.html . - A deployed Amazon SageMaker real-time endpoint or a JumpStart model to deploy; see the following section . Solution overview Consider a common scenario: you have a generative AI model ready for production and need to determine the optimal instance type, framework configuration, and serving parameters. Traditionally, this involves manual trial and error across multiple instance types, container versions, and concurrency settings. With the Amazon SageMaker Python SDK integration, you can automate this entire workflow in a single notebook. The following walkthrough guides you through the end-to-end journey using this notebook https://github.com/aws-samples/sagemaker-genai-hosting-examples/blob/main/03-features/gen-ai-inference-recommendations/pysdk-ai-inference-recommender-demo.ipynb : Generate deployment recommendations: Let the service explore instance and framework configurations against your workload profile and return ranked options. Interpret and select: Review the ranked results, understand the tradeoffs, and pick the best fit. Deploy: Push the winning configuration to a live Amazon SageMaker endpoint. Benchmark: Validate the deployed endpoint under realistic load conditions. Compare frameworks: Optionally run LMI and vLLM head-to-head to find the best serving stack. Generate recommendations from real traffic data Your first step is to find the best deployment configuration for your model and workload. Rather than manually deploying across multiple instance types, call mb.generate deployment recommendations … to let the service explore instance types and framework configurations against your workload profile. The service deploys your model on each candidate, runs a load test matching your traffic pattern, and returns a ranked list of configurations optimized for your chosen performance target. Recommendation results can also be represented as a Python data frame. How to interpret recommendation results The recommendations table shows two candidate configurations rank 0 and rank 1 , both on ml.g6.2xlarge but with different LMI container versions. Here’s how to read the key metrics and choose between them: Key metrics to compare: - RequestThroughput avg : Requests the endpoint can serve per second. Higher is better. - OutputTokenThroughput avg : Total tokens generated per second across all concurrent requests. Higher is better. - RequestLatency p50/p90/p99 : End-to-end time from request to full response. Lower is better. - TimeToFirstToken p50/p90 : How quickly the user sees the first streamed token. Lower is better. - InterTokenLatency p50/p90 : Delay between successive tokens during streaming. Lower is better. Choosing between the two configurations in this example: Rank 0 lmi-26-0-0 delivers 112.8 req/s throughput and 3,609 tokens/s, with p90 TTFT of 983 ms and p90 latency of 1,000 ms. Rank 1 lmi-27-0-0 delivers 96.9 req/s throughput and 3,099 tokens/s, with p90 TTFT of 1,088 ms and p90 latency of 1,122 ms. Rank 0 wins on every dimension: approximately 16% higher throughput and approximately 10 percent lower latency. The service ranks it first because the job was configured with performance target=PerformanceTarget.TTFT MS, meaning the optimizer prioritized configurations that minimize time-to-first-token. General decision framework Latency-sensitive applications chatbots, interactive UIs : Prioritize low TTFT p90/p99 so users perceive fast responses. Throughput-sensitive workloads batch summarization, offline processing : Prioritize high RequestThroughput and OutputTokenThroughput to maximize tokens per dollar. If two configurations are close on your primary metric, use the secondary metrics as tiebreakers, then factor in cost a smaller instance at similar performance saves money . In this example, the top-ranked configuration lmi-26-0-0 is the clear choice because it dominates across all metrics at the same concurrency level 64 . Deploy from previously run recommendation job In production workflows, you often generate recommendations in one session and deploy in another. For example, a data scientist might run the recommendation job during experimentation, while an MLOps pipeline deploys the result during a release cycle. Use ModelBuilder.from recommendation job job name to hydrate a ModelBuilder from a completed job: Deploy a JumpStart model and benchmark it After you have deployed your recommended configuration, the next step is to validate its performance under controlled conditions. Benchmarking confirms that the endpoint meets your latency and throughput requirements before serving production traffic. The SDK makes this straightforward: deploy a JumpStart model and run a synthetic load test in only a few lines of code. Reading benchmark results After the benchmark completes, you need to understand whether your endpoint meets your service-level objectives. The benchmark returns a typed result object with a metrics accessor that gives you programmatic access to throughput, latency percentiles, and token-level timing. IDE autocomplete works on all fields. Benchmark results can similarly be represented as a Python data frame. Benchmark metrics reference This table breaks down the metrics reported by the benchmarking service. Metric | Description | Available statistics | request throughput | Completed requests per second | avg, p50, p90, p99 | time to first token | Time from request to first streamed token | avg, p50, p90, p99 | request latency | End-to-end request latency | avg, p50, p90, p99 | output token throughput | Output tokens per second system-wide aggregate | avg only | Compare LMI and vLLM, then deploy the preferred option Not sure which inference framework suits your model? Run two recommendation jobs in parallel, one for LMI and one for vLLM, compare the top results, and deploy from whichever wins. Sample notebook End-to-end runnable notebooks are available on GitHub https://github.com/aws-samples/sagemaker-genai-hosting-examples/tree/main/03-features/gen-ai-inference-recommendations : Notebook | What it covers | | Clean up To avoid ongoing charges, delete the endpoints created during this walkthrough. Run the following commands in your notebook: For pricing details on Amazon SageMaker real-time inference instances, see Amazon SageMaker AI Pricing /sagemaker/ai/pricing/ . Conclusion In this post, you walked through the complete inference optimization journey. You generated deployment recommendations that explore instance and framework configurations, interpreted the ranked results, and deployed the optimal configuration. You also validated performance with benchmarks under realistic load and compared serving frameworks head-to-head. With the Amazon SageMaker Python SDK integration for generative AI inference recommendations, this entire workflow lives in a single notebook, which removes the context-switching between console UIs and API calls. To get started, upgrade to Amazon SageMaker Python SDK v3 and explore the sample notebook included with this post https://github.com/aws-samples/sagemaker-genai-hosting-examples/tree/main/03-features/gen-ai-inference-recommendations . Additional resources Generative AI inference recommendations in Amazon SageMaker AI https://docs.aws.amazon.com/sagemaker/latest/dg/inference-recommender.html . Generative AI inference benchmarks in Amazon SageMaker AI https://docs.aws.amazon.com/sagemaker/latest/dg/generative-ai-inference-recommendations-benchmark.html . SageMaker Python SDK documentation https://sagemaker.readthedocs.io/ . Getting started with SageMaker JumpStart https://docs.aws.amazon.com/sagemaker/latest/dg/studio-jumpstart.html . Sample notebooks on GitHub https://github.com/aws-samples/sagemaker-genai-hosting-examples/tree/main/03-features/gen-ai-inference-recommendations . Have feedback or questions? Let us know in the Amazon SageMaker discussion forums https://repost.aws/tags/TAeaKkE5oeSuGz7XE3DRmLiA/amazon-sage-maker .