{"slug": "aim-unified-user-experience-from-profile-discovery-to-deployment", "title": "AIM: Unified User Experience from Profile Discovery to Deployment", "summary": "AMD detailed its AMD Inference Microservices (AIMs), standardized Docker-based inference microservices that automatically select a runtime profile from inputs including model, precision, engine, latency or throughput target, and detected accelerator count and model, then expose an OpenAI-compatible API. AMD validated the workflow separately on an AMD Instinct MI300X cluster, an AMD EPYC 9965 cluster, and an AMD Radeon PRO R9700 cluster, with the walkthrough covering profile selection, dry-run and list-profiles exploration, and Docker deployment. The post follows AMD's earlier multi-accelerator support coverage spanning Instinct GPUs, Radeon PRO GPUs, and EPYC CPUs.", "body_md": "# AIM: Unified User Experience from Profile Discovery to Deployment[#](#aim-unified-user-experience-from-profile-discovery-to-deployment)\n\nThis post is a continuation of the [Multi-Accelerator Support for AIMs and AMD Solution Blueprints](https://rocm.blogs.amd.com/software-tools-optimization/eai-hw-support/README.html) blog, which introduced (i) multi-accelerator coverage across AMD Instinct™ GPUs, Radeon™ PRO GPUs, and EPYC™ CPUs, and (ii) walked through the Document Summarization Solution Blueprint. **Here we focus on [AIMs](https://enterprise-ai.docs.amd.com/en/latest/aims/overview.html) specifically** and show how they deliver a unified user experience from Radeon to Instinct, i.e., similar profile discovery commands, container workflow, and OpenAI-compatible API regardless of the accelerator.\n\nWe will demonstrate this by running AIMs on a Radeon GPU, EPYC CPUs, a single Instinct GPU, and an 8-GPU Instinct node. This walkthrough covers three steps:\n\n1. **AIM containers** : How AIM profile selection works\n2. **AIM exploration** :`dry-run` and`list-profiles`\n3. **Docker deployment** : Start the server and call the OpenAI-compatible API\n\n## AMD Inference Microservices - AIMs[#](#amd-inference-microservices-aims)\n\n**AIMs** are standardized inference microservices for serving AI models on AMD hardware. They are distributed as Docker images, which makes them easy to deploy and manage. Each AIM ships with predefined [**profiles**](https://github.com/amd-enterprise-ai/aim-build/blob/main/docs/aim_architecture.md#35-profile-example): inference engine configurations for specific accelerators, precisions, tensor-parallel layouts, and latency or throughput targets.\n\nAs shown in Figure 1, AIMs abstract away the complexities involved in configuring and serving AI models by providing a [mechanism](https://github.com/amd-enterprise-ai/aim-build/blob/main/docs/aim_architecture.md#4-aim-runtime--command-execution) that automatically selects runtime parameters based on the user’s input, hardware, and model specifications.\n\n*Figure 1: AIM automated deployment sequence.*\n\nIn practice, the high-level steps look like this:\n\n1. **Initialize runtime** and**detect accelerators**\n  - Load and validate configurations.\n  - Example: AIM detects one AMD Instinct MI300X.\n2. **Select a profile** and**generate the launch command**\n  - Profiles are predefined configurations for specific models and hardware.\n  - Selection is automatic, based on inputs such as: \n    - Model (for example `meta-llama/Llama-3.1-8B-Instruct` )\n    - Precision (for example `auto` ,`fp16` , …)\n    - Engine (for example `vllm` )\n    - Metric ( `latency` or`throughput` )\n    - Detected accelerator count (for example `1` ,`2` ,`4` ,`8` GPUs)\n    - Detected accelerator model (for example `MI300X` ,`R9700` , …)\n  - It is possible to bypass automatic selection and specify a particular profile.\n  - See the [selection algorithm](https://github.com/amd-enterprise-ai/aim-build/blob/main/docs/aim_architecture.md#34-selection-algorithm-overview) for more information.\n3. **Launch the inference engine**\n  - AIM produces the command and environment variables using the selected profile to start the server.\n4. **Load model weights**\n  - AIM supports e.g. flexible [model caching strategies](https://github.com/amd-enterprise-ai/aim-build/blob/main/docs/aim_architecture.md#6-model-caching-support) .\n5. **Deployment completed**\n  - AIM exposes an [OpenAI-compatible API](https://platform.openai.com/docs/api-reference/introduction) for LLMs.\n\nTo illustrate this sequence, the rest of this post explores the AIM container with Docker. While we explore the container, be aware that there are multiple ways to deploy an AIM depending on your use case. See the following resources for more information:\n\n## Prerequisites[#](#prerequisites)\n\nThis post was validated separately on an AMD Instinct MI300X cluster, an AMD EPYC 9965 cluster, and an AMD Radeon PRO R9700 cluster. Also ensure that your system meets the following requirements:\n\n- AMD GPU with ROCm support, required for GPU deployment. See the [accelerator support](https://enterprise-ai.docs.amd.com/en/latest/aims/accelerator_support.html) page for the host ROCm version for your accelerator.\n- Docker installed, with device access on GPU hosts ( `/dev/kfd` and`/dev/dri` ).\n\n## Choosing the AIM[#](#choosing-the-aim)\n\nThis guide uses the following models:\n\n| Hardware | Model | Container image | \n|---|---|---|\n| Radeon | `Qwen/Qwen3.5-9B` | `amdenterpriseai/aim-radeon-qwen-qwen3-5-9b:0.12.0-preview` | \n| EPYC | `Qwen/Qwen3.5-9B` | `amdenterpriseai/aim-epyc-qwen-qwen3-5-9b:0.13.0` | \n| Instinct | `openai/gpt-oss-120b` | `amdenterpriseai/aim-openai-gpt-oss-120b:0.11.1` | \n\n*Version tags (for example `0.11.1`) change with new AIM releases.*\n\nFor a high-level view of which AIMs are available for which hardware, see the [accelerator support](https://enterprise-ai.docs.amd.com/en/latest/aims/accelerator_support.html) table.\n\nTo use a different model:\n\n- Open the [AIM catalog](https://enterprise-ai.docs.amd.com/en/latest/aims/catalog/models.html) .\n- Locate the model for your hardware.\n- Open the Docker Hub link ( [example](https://hub.docker.com/r/amdenterpriseai/aim-openai-gpt-oss-120b/tags) ) and copy the image name.\n- Open the **Technical specification** link in the catalog. That page lists available profiles for the specific model ([example](https://enterprise-ai.docs.amd.com/en/latest/aims/docs-aim/instinct/openai/gpt-oss-120b/README.html#model-specific-aim) ).**To follow this guide** , confirm the model has an “optimized” or “preview” profile for your hardware (see Figure 2 for an example). If it does not, pick another model from the catalog.\n\n*Figure 2: Subset of AIM profiles for gpt-oss-120b.*\n\n## Docker Profile Discovery[#](#docker-profile-discovery)\n\nTo illustrate the deployment sequence in action, we will use two AIM inspection commands:\n\n- [`dry-run`](https://github.com/amd-enterprise-ai/aim-build/blob/main/docs/cli.md#dry-run-dry-run) : shows which profile AIM would select and the exact engine command it would run\n- [`list-profiles`](https://github.com/amd-enterprise-ai/aim-build/blob/main/docs/cli.md#list-profiles-list-profiles) : lists and categorizes all available profiles by their compatibility with the current configuration.\n\nThis helps you understand which profiles are available and why certain profiles may or may not be selected. **Neither starts a server**, so both are safe to run before you commit to a deployment.\n\n### Explore the Profile Selection[#](#explore-the-profile-selection)\n\nLet’s begin with the `dry-run` command. The command shape is similar on every platform. What changes is the container image (and, for AMD Instinct 1 vs 8 GPUs, how many accelerators AIM detects). Pay attention to the detected hardware, accelerator count (e.g. detected GPUs), engine arguments, and environment variables — AIM fills those in for you.\n\nCommand:\n\n```\ndocker run --rm \\\n  --device=/dev/kfd --device=/dev/dri \\\n  amdenterpriseai/aim-radeon-qwen-qwen3-5-9b:0.12.0-preview \\\n  dry-run\n```\n\nTruncated output:\n\n```\nprofile:\n  aim_id: Qwen/Qwen3.5-9B\n  metadata:\n    accelerator_count: 1\n    accelerator_model: R9700\n    accelerator_type: gpu\n    metric: latency\n    precision: bf16\n    ...\n  engine_args:\n    tensor-parallel-size: 1\n    ...\n  env_vars:\n    FLASH_ATTENTION_TRITON_AMD_ENABLE: 'TRUE'\n    ...\n\n...\n\nexec python -m vllm.entrypoints.openai.api_server --model Qwen/Qwen3.5-9B ....\ndocker run --rm \\\n  --device=/dev/kfd --device=/dev/dri \\\n  amdenterpriseai/aim-openai-gpt-oss-120b:0.11.1 \\\n  dry-run\n```\n\nTruncated output:\n\n```\nprofile:\n  aim_id: openai/gpt-oss-120b\n  metadata:\n    accelerator_count: 1\n    accelerator_model: MI300X\n    accelerator_type: gpu\n    metric: latency\n    precision: fp4\n    ...\n  engine_args:\n    tensor-parallel-size: 1\n    ...\n  env_vars:\n    VLLM_ROCM_USE_AITER: '1'\n    ...\n\n...\n\nexec python -m vllm.entrypoints.openai.api_server --model openai/gpt-oss-120b ....\ndocker run --rm \\\n  --device=/dev/kfd --device=/dev/dri \\\n  amdenterpriseai/aim-openai-gpt-oss-120b:0.11.1 \\\n  dry-run\n```\n\nTruncated output:\n\n```\nprofile:\n  aim_id: openai/gpt-oss-120b\n  metadata:\n    accelerator_count: 8\n    accelerator_model: MI300X\n    accelerator_type: gpu\n    metric: latency\n    precision: fp4\n    ...\n  engine_args:\n    tensor-parallel-size: 8\n    ...\n  env_vars:\n    VLLM_ROCM_USE_AITER: '1'\n    ...\n\n...\n\nexec python -m vllm.entrypoints.openai.api_server --model openai/gpt-oss-120b ....\ndocker run --rm \\\n  amdenterpriseai/aim-epyc-qwen-qwen3-5-9b:0.13.0 \\\n  dry-run\n```\n\nTruncated output:\n\n```\nprofile:\n  aim_id: Qwen/Qwen3.5-9B\n  metadata:\n    accelerator_count: 188\n    accelerator_model: EPYC_9965\n    accelerator_type: cpu\n    metric: latency\n    precision: bf16\n    ...\n  engine_args:\n    enable-chunked-prefill: true\n    ...\n  env_vars:\n    VLLM_CPU_OMP_THREADS_BIND: auto\n    ...\n\n...\n\nexec python -m vllm.entrypoints.openai.api_server --model Qwen/Qwen3.5-9B ....\n```\n\nAs you can see, the profile configuration changes depending on the model and the hardware. Latency is picked as the performance metric by default. For a technical guide to the **selection algorithm**, see [algorithm overview](https://github.com/amd-enterprise-ai/aim-build/blob/main/docs/aim_architecture.md#34-selection-algorithm-overview).\n\nAIM also supports customizing the deployment with [environment variables](https://enterprise-ai.docs.amd.com/en/latest/aims/docker_deployment.html#customizing-deployment-with-environment-variables) and [custom profile configurations](https://enterprise-ai.docs.amd.com/en/latest/aims/custom_profiles.html) that extend beyond the built-in predefined profiles.\n\n### List Profiles[#](#list-profiles)\n\nUse `list-profiles` to see every profile available for the model. Here is an example for a single AMD Instinct MI300X GPU. Truncated output is shown in Figure 3:\n\n```\ndocker run --rm \\\n  --device=/dev/kfd --device=/dev/dri \\\n  amdenterpriseai/aim-openai-gpt-oss-120b:0.11.1 \\\n  list-profiles\n```\n\n*Figure 3: Subset of AIM profiles for gpt-oss-120b.*\n\nThe output lists and categorizes all available profiles by their compatibility with the current configuration. Each row shows the target accelerator, precision, inference engine, tensor parallelism (TP), performance metric, and profile type (`optimized`, `preview`, …). The **Compatibility** column shows whether a profile matches your current configuration; mismatches (for example `accelerator_mismatch` or `metric_mismatch`) explain why others are excluded.\n\n## Deployment[#](#deployment)\n\nThe following is a small example. See [Docker deployment guide](https://enterprise-ai.docs.amd.com/en/latest/aims/docker_deployment.html#docker-deployment) for more detail, including customization.\n\nNote: If you are running larger models in multi-GPU environments, read the following [guide](https://enterprise-ai.docs.amd.com/en/latest/aims/docker_deployment.html#running-larger-models-in-multi-gpu-environments) first.\n\nOnce you have validated the profile selection, start the inference server. The following command is an example for one AMD Instinct MI300X GPU. It runs the container with the profile that the container automatically selects on the hardware it detects:\n\n```\ndocker run \\\n  --device=/dev/kfd --device=/dev/dri \\\n  -p 8000:8000 \\\n  amdenterpriseai/aim-openai-gpt-oss-120b:0.11.1\n```\n\nWatch the container logs for profile selection, then for the inference server to start. Truncated output:\n\n```\n(APIServer pid=1) INFO:     Started server process [1]\n(APIServer pid=1) INFO:     Waiting for application startup.\n(APIServer pid=1) INFO:     Application startup complete.\n```\n\nOn first start:\n\n1. AIM detects your GPU and selects a profile automatically.\n2. Model weights are downloaded from Hugging Face.\n3. The inference engine starts and listens on port `8000` .\n\nSend a completion request (if you use another model, change the `model` field):\n\n```\ncurl http://localhost:8000/v1/completions \\\n  -H \"Content-Type: application/json\" \\\n  -d '{\n    \"model\": \"openai/gpt-oss-120b\",\n    \"prompt\": \"Once upon a time,\",\n    \"max_tokens\": 50,\n    \"temperature\": 0.7\n  }'\n```\n\nExample output (truncated):\n\n```\n{\n  \"text\": \" in a small village near a dense forest, there lived a wise ...\"\n}\n```\n\n## Summary[#](#summary)\n\nIn this blog, we demonstrated the flexibility and unified user experience of AIMs. You saw profile discovery across platforms, and a Docker deployment example on Instinct. AIM detects your accelerator, selects a profile automatically, and adjusts engine arguments and environment variables for each platform without manual tuning. To explore further, browse the [AIM catalog](https://enterprise-ai.docs.amd.com/en/latest/aims/catalog/models.html) for models and technical specifications to find the model for your use case. Read the [Docker deployment guide](https://enterprise-ai.docs.amd.com/en/latest/aims/docker_deployment.html) for customization options beyond the defaults shown here.\n\n## Disclaimers[#](#disclaimers)\n\nThe information presented in this document is for informational purposes only and may contain technical inaccuracies, omissions, and typographical errors. The information contained herein is subject to change and may be rendered inaccurate for many reasons, including but not limited to product and roadmap changes, component and motherboard version changes, new model and/or product releases, product differences between differing manufacturers, software changes, BIOS flashes, firmware upgrades, or the like. Any computer system has risks of security vulnerabilities that cannot be completely prevented or mitigated. AMD assumes no obligation to update or otherwise correct or revise this information. However, AMD reserves the right to revise this information and to make changes from time to time to the content hereof without obligation of AMD to notify any person of such revisions or changes. THIS INFORMATION IS PROVIDED ‘AS IS.” AMD MAKES NO REPRESENTATIONS OR WARRANTIES WITH RESPECT TO THE CONTENTS HEREOF AND ASSUMES NO RESPONSIBILITY FOR ANY INACCURACIES, ERRORS, OR OMISSIONS THAT MAY APPEAR IN THIS INFORMATION. AMD SPECIFICALLY DISCLAIMS ANY IMPLIED WARRANTIES OF NON-INFRINGEMENT, MERCHANTABILITY, OR FITNESS FOR ANY PARTICULAR PURPOSE. IN NO EVENT WILL AMD BE LIABLE TO ANY PERSON FOR ANY RELIANCE, DIRECT, INDIRECT, SPECIAL, OR OTHER CONSEQUENTIAL DAMAGES ARISING FROM THE USE OF ANY INFORMATION CONTAINED HEREIN, EVEN IF AMD IS EXPRESSLY ADVISED OF THE POSSIBILITY OF SUCH DAMAGES. AMD, the AMD Arrow logo, AMD Instinct, AMD Radeon, AMD EPYC, ROCm and combinations thereof are trademarks of Advanced Micro Devices, Inc. Other product names used in this publication are for identification purposes only and may be trademarks of their respective companies. © 2026 Advanced Micro Devices, Inc. All rights reserved\n\nThird-party content is licensed to you directly by the third party that owns the content and is not licensed to you by AMD. ALL LINKED THIRD-PARTY CONTENT IS PROVIDED “AS IS” WITHOUT A WARRANTY OF ANY KIND. USE OF SUCH THIRD-PARTY CONTENT IS DONE AT YOUR SOLE DISCRETION AND UNDER NO CIRCUMSTANCES WILL AMD BE LIABLE TO YOU FOR ANY THIRD-PARTY CONTENT. YOU ASSUME ALL RISK AND ARE SOLELY RESPONSIBLE FOR ANY DAMAGES THAT MAY ARISE FROM YOUR USE OF THIRD-PARTY CONTENT.", "url": "https://wpnews.pro/news/aim-unified-user-experience-from-profile-discovery-to-deployment", "canonical_source": "https://rocm.blogs.amd.com/artificial-intelligence/aims-deployment-exp/README.html", "published_at": "2026-09-30 00:00:00+00:00", "updated_at": "2026-09-30 16:21:15.235246+00:00", "lang": "en", "topics": ["ai-infrastructure", "ai-products", "ai-tools", "large-language-models"], "entities": ["AMD", "AMD Inference Microservices", "AMD Instinct MI300X", "AMD EPYC 9965", "AMD Radeon PRO R9700", "Docker", "vLLM", "meta-llama/Llama-3.1-8B-Instruct"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/aim-unified-user-experience-from-profile-discovery-to-deployment", "markdown": "https://wpnews.pro/news/aim-unified-user-experience-from-profile-discovery-to-deployment.md", "text": "https://wpnews.pro/news/aim-unified-user-experience-from-profile-discovery-to-deployment.txt", "jsonld": "https://wpnews.pro/news/aim-unified-user-experience-from-profile-discovery-to-deployment.jsonld"}}