cd /news/ai-infrastructure/aim-unified-user-experience-from-pro… · home › topics › ai-infrastructure › article
[ARTICLE · art-142634] src=rocm.blogs.amd.com ↗ pub= topic=ai-infrastructure verified=true sentiment=↑ positive

AIM: Unified User Experience from Profile Discovery to Deployment

AMD detailed its AMD Inference Microservices (AIMs), standardized Docker-based inference microservices that automatically select a runtime profile from inputs including model, precision, engine, latency or throughput target, and detected accelerator count and model, then expose an OpenAI-compatible API. AMD validated the workflow separately on an AMD Instinct MI300X cluster, an AMD EPYC 9965 cluster, and an AMD Radeon PRO R9700 cluster, with the walkthrough covering profile selection, dry-run and list-profiles exploration, and Docker deployment. The post follows AMD's earlier multi-accelerator support coverage spanning Instinct GPUs, Radeon PRO GPUs, and EPYC CPUs.

by read9 min views1 publishedSep 30, 2026
AIM: Unified User Experience from Profile Discovery to Deployment
Image: Rocm (auto-discovered)

This post is a continuation of the Multi-Accelerator Support for AIMs and AMD Solution Blueprints blog, which introduced (i) multi-accelerator coverage across AMD Instinct™ GPUs, Radeon™ PRO GPUs, and EPYC™ CPUs, and (ii) walked through the Document Summarization Solution Blueprint. Here we focus on AIMs specifically and show how they deliver a unified user experience from Radeon to Instinct, i.e., similar profile discovery commands, container workflow, and OpenAI-compatible API regardless of the accelerator.

We will demonstrate this by running AIMs on a Radeon GPU, EPYC CPUs, a single Instinct GPU, and an 8-GPU Instinct node. This walkthrough covers three steps:

  1. AIM containers : How AIM profile selection works
  2. AIM exploration :dry-run andlist-profiles
  3. Docker deployment : Start the server and call the OpenAI-compatible API

AMD Inference Microservices - AIMs# #

AIMs are standardized inference microservices for serving AI models on AMD hardware. They are distributed as Docker images, which makes them easy to deploy and manage. Each AIM ships with predefined profiles: inference engine configurations for specific accelerators, precisions, tensor-parallel layouts, and latency or throughput targets.

As shown in Figure 1, AIMs abstract away the complexities involved in configuring and serving AI models by providing a mechanism that automatically selects runtime parameters based on the user’s input, hardware, and model specifications.

Figure 1: AIM automated deployment sequence.

In practice, the high-level steps look like this:

  1. Initialize runtime anddetect accelerators
  • Load and validate configurations.
  • Example: AIM detects one AMD Instinct MI300X.
  1. Select a profile andgenerate the launch command
  • Profiles are predefined configurations for specific models and hardware.
  • Selection is automatic, based on inputs such as:
    • Model (for example meta-llama/Llama-3.1-8B-Instruct )
    • Precision (for example auto ,fp16 , …)
    • Engine (for example vllm )
    • Metric ( latency orthroughput )
    • Detected accelerator count (for example 1 ,2 ,4 ,8 GPUs)
    • Detected accelerator model (for example MI300X ,R9700 , …)
  • It is possible to bypass automatic selection and specify a particular profile.
  • See the selection algorithm for more information.
  1. Launch the inference engine
  • AIM produces the command and environment variables using the selected profile to start the server.
  1. Load model weights
  1. Deployment completed

To illustrate this sequence, the rest of this post explores the AIM container with Docker. While we explore the container, be aware that there are multiple ways to deploy an AIM depending on your use case. See the following resources for more information:

Prerequisites# #

This post was validated separately on an AMD Instinct MI300X cluster, an AMD EPYC 9965 cluster, and an AMD Radeon PRO R9700 cluster. Also ensure that your system meets the following requirements:

  • AMD GPU with ROCm support, required for GPU deployment. See the accelerator support page for the host ROCm version for your accelerator.
  • Docker installed, with device access on GPU hosts ( /dev/kfd and/dev/dri ).

Choosing the AIM# #

This guide uses the following models:

Hardware Model Container image
Radeon Qwen/Qwen3.5-9B amdenterpriseai/aim-radeon-qwen-qwen3-5-9b:0.12.0-preview
EPYC Qwen/Qwen3.5-9B amdenterpriseai/aim-epyc-qwen-qwen3-5-9b:0.13.0
Instinct openai/gpt-oss-120b amdenterpriseai/aim-openai-gpt-oss-120b:0.11.1

Version tags (for example 0.11.1) change with new AIM releases.

For a high-level view of which AIMs are available for which hardware, see the accelerator support table.

To use a different model:

  • Open the AIM catalog .
  • Locate the model for your hardware.
  • Open the Docker Hub link ( example ) and copy the image name.
  • Open the Technical specification link in the catalog. That page lists available profiles for the specific model (example ).To follow this guide , confirm the model has an “optimized” or “preview” profile for your hardware (see Figure 2 for an example). If it does not, pick another model from the catalog.

Figure 2: Subset of AIM profiles for gpt-oss-120b.

Docker Profile Discovery# #

To illustrate the deployment sequence in action, we will use two AIM inspection commands:

  • dry-run : shows which profile AIM would select and the exact engine command it would run
  • list-profiles : lists and categorizes all available profiles by their compatibility with the current configuration.

This helps you understand which profiles are available and why certain profiles may or may not be selected. Neither starts a server, so both are safe to run before you commit to a deployment.

Explore the Profile Selection#

Let’s begin with the dry-run command. The command shape is similar on every platform. What changes is the container image (and, for AMD Instinct 1 vs 8 GPUs, how many accelerators AIM detects). Pay attention to the detected hardware, accelerator count (e.g. detected GPUs), engine arguments, and environment variables — AIM fills those in for you.

Command:

docker run --rm \
  --device=/dev/kfd --device=/dev/dri \
  amdenterpriseai/aim-radeon-qwen-qwen3-5-9b:0.12.0-preview \
  dry-run

Truncated output:

profile:
  aim_id: Qwen/Qwen3.5-9B
  metadata:
    accelerator_count: 1
    accelerator_model: R9700
    accelerator_type: gpu
    metric: latency
    precision: bf16
    ...
  engine_args:
    tensor-parallel-size: 1
    ...
  env_vars:
    FLASH_ATTENTION_TRITON_AMD_ENABLE: 'TRUE'
    ...

...

exec python -m vllm.entrypoints.openai.api_server --model Qwen/Qwen3.5-9B ....
docker run --rm \
  --device=/dev/kfd --device=/dev/dri \
  amdenterpriseai/aim-openai-gpt-oss-120b:0.11.1 \
  dry-run

Truncated output:

profile:
  aim_id: openai/gpt-oss-120b
  metadata:
    accelerator_count: 1
    accelerator_model: MI300X
    accelerator_type: gpu
    metric: latency
    precision: fp4
    ...
  engine_args:
    tensor-parallel-size: 1
    ...
  env_vars:
    VLLM_ROCM_USE_AITER: '1'
    ...

...

exec python -m vllm.entrypoints.openai.api_server --model openai/gpt-oss-120b ....
docker run --rm \
  --device=/dev/kfd --device=/dev/dri \
  amdenterpriseai/aim-openai-gpt-oss-120b:0.11.1 \
  dry-run

Truncated output:

profile:
  aim_id: openai/gpt-oss-120b
  metadata:
    accelerator_count: 8
    accelerator_model: MI300X
    accelerator_type: gpu
    metric: latency
    precision: fp4
    ...
  engine_args:
    tensor-parallel-size: 8
    ...
  env_vars:
    VLLM_ROCM_USE_AITER: '1'
    ...

...

exec python -m vllm.entrypoints.openai.api_server --model openai/gpt-oss-120b ....
docker run --rm \
  amdenterpriseai/aim-epyc-qwen-qwen3-5-9b:0.13.0 \
  dry-run

Truncated output:

profile:
  aim_id: Qwen/Qwen3.5-9B
  metadata:
    accelerator_count: 188
    accelerator_model: EPYC_9965
    accelerator_type: cpu
    metric: latency
    precision: bf16
    ...
  engine_args:
    enable-chunked-prefill: true
    ...
  env_vars:
    VLLM_CPU_OMP_THREADS_BIND: auto
    ...

...

exec python -m vllm.entrypoints.openai.api_server --model Qwen/Qwen3.5-9B ....

As you can see, the profile configuration changes depending on the model and the hardware. Latency is picked as the performance metric by default. For a technical guide to the selection algorithm, see algorithm overview.

AIM also supports customizing the deployment with environment variables and custom profile configurations that extend beyond the built-in predefined profiles.

List Profiles#

Use list-profiles to see every profile available for the model. Here is an example for a single AMD Instinct MI300X GPU. Truncated output is shown in Figure 3:

docker run --rm \
  --device=/dev/kfd --device=/dev/dri \
  amdenterpriseai/aim-openai-gpt-oss-120b:0.11.1 \
  list-profiles

Figure 3: Subset of AIM profiles for gpt-oss-120b.

The output lists and categorizes all available profiles by their compatibility with the current configuration. Each row shows the target accelerator, precision, inference engine, tensor parallelism (TP), performance metric, and profile type (optimized, preview, …). The Compatibility column shows whether a profile matches your current configuration; mismatches (for example accelerator_mismatch or metric_mismatch) explain why others are excluded.

Deployment# #

The following is a small example. See Docker deployment guide for more detail, including customization.

Note: If you are running larger models in multi-GPU environments, read the following guide first.

Once you have validated the profile selection, start the inference server. The following command is an example for one AMD Instinct MI300X GPU. It runs the container with the profile that the container automatically selects on the hardware it detects:

docker run \
  --device=/dev/kfd --device=/dev/dri \
  -p 8000:8000 \
  amdenterpriseai/aim-openai-gpt-oss-120b:0.11.1

Watch the container logs for profile selection, then for the inference server to start. Truncated output:

(APIServer pid=1) INFO:     Started server process [1]
(APIServer pid=1) INFO:     Waiting for application startup.
(APIServer pid=1) INFO:     Application startup complete.

On first start:

  1. AIM detects your GPU and selects a profile automatically.
  2. Model weights are downloaded from Hugging Face.
  3. The inference engine starts and listens on port 8000 .

Send a completion request (if you use another model, change the model field):

curl http://localhost:8000/v1/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "openai/gpt-oss-120b",
    "prompt": "Once upon a time,",
    "max_tokens": 50,
    "temperature": 0.7
  }'

Example output (truncated):

{
  "text": " in a small village near a dense forest, there lived a wise ..."
}

Summary# #

In this blog, we demonstrated the flexibility and unified user experience of AIMs. You saw profile discovery across platforms, and a Docker deployment example on Instinct. AIM detects your accelerator, selects a profile automatically, and adjusts engine arguments and environment variables for each platform without manual tuning. To explore further, browse the AIM catalog for models and technical specifications to find the model for your use case. Read the Docker deployment guide for customization options beyond the defaults shown here.

Disclaimers# #

The information presented in this document is for informational purposes only and may contain technical inaccuracies, omissions, and typographical errors. The information contained herein is subject to change and may be rendered inaccurate for many reasons, including but not limited to product and roadmap changes, component and motherboard version changes, new model and/or product releases, product differences between differing manufacturers, software changes, BIOS flashes, firmware upgrades, or the like. Any computer system has risks of security vulnerabilities that cannot be completely prevented or mitigated. AMD assumes no obligation to update or otherwise correct or revise this information. However, AMD reserves the right to revise this information and to make changes from time to time to the content hereof without obligation of AMD to notify any person of such revisions or changes. THIS INFORMATION IS PROVIDED ‘AS IS.” AMD MAKES NO REPRESENTATIONS OR WARRANTIES WITH RESPECT TO THE CONTENTS HEREOF AND ASSUMES NO RESPONSIBILITY FOR ANY INACCURACIES, ERRORS, OR OMISSIONS THAT MAY APPEAR IN THIS INFORMATION. AMD SPECIFICALLY DISCLAIMS ANY IMPLIED WARRANTIES OF NON-INFRINGEMENT, MERCHANTABILITY, OR FITNESS FOR ANY PARTICULAR PURPOSE. IN NO EVENT WILL AMD BE LIABLE TO ANY PERSON FOR ANY RELIANCE, DIRECT, INDIRECT, SPECIAL, OR OTHER CONSEQUENTIAL DAMAGES ARISING FROM THE USE OF ANY INFORMATION CONTAINED HEREIN, EVEN IF AMD IS EXPRESSLY ADVISED OF THE POSSIBILITY OF SUCH DAMAGES. AMD, the AMD Arrow logo, AMD Instinct, AMD Radeon, AMD EPYC, ROCm and combinations thereof are trademarks of Advanced Micro Devices, Inc. Other product names used in this publication are for identification purposes only and may be trademarks of their respective companies. © 2026 Advanced Micro Devices, Inc. All rights reserved

Third-party content is licensed to you directly by the third party that owns the content and is not licensed to you by AMD. ALL LINKED THIRD-PARTY CONTENT IS PROVIDED “AS IS” WITHOUT A WARRANTY OF ANY KIND. USE OF SUCH THIRD-PARTY CONTENT IS DONE AT YOUR SOLE DISCRETION AND UNDER NO CIRCUMSTANCES WILL AMD BE LIABLE TO YOU FOR ANY THIRD-PARTY CONTENT. YOU ASSUME ALL RISK AND ARE SOLELY RESPONSIBLE FOR ANY DAMAGES THAT MAY ARISE FROM YOUR USE OF THIRD-PARTY CONTENT.

── more in #ai-infrastructure 4 stories · sorted by recency
── more on @amd 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/aim-unified-user-exp…] indexed:0 read:9min 2026-09-30 · —