Load & Serve Models on Amazon EKS
Load & Serve Models on Amazon EKS
[추출 작업 중]
출처: 문서
본문
Help improve this page
To contribute to this user guide, choose the Edit this page on GitHub link that is located in the right pane of every page.
Load & Serve Models on Amazon EKS
Tip
Register for upcoming Amazon EKS AI/ML workshops.
The steps in this section deploy a large language model (LLM) on Amazon EKS, serve it with vLLM, and interact with the inference endpoint.
The walkthrough uses the following tools:
vLLM — A high-throughput inference engine optimized for LLM serving and GPU memory management.
Run:ai Model Streamer — Streams model weights directly from Amazon S3 to GPU memory, reducing load time from minutes to seconds.
Open WebUI — A self-hosted chat frontend that connects to vLLM’s OpenAI-compatible API.
This section uses the Ministral-3-8B-Instruct-2512 model, but you can deploy any AI model that vLLM supports. For a list of supported models, see Supported models in the vLLM documentation.
Important
Use the cluster you created in the Set up Amazon EKS cluster for AI/ML workloads section. The instructions in this walkthrough work for both EKS Auto Mode and self-managed Karpenter.
The architecture diagram shows the end-to-end flow:
Model weights are downloaded from Hugging Face to Amazon S3.
vLLM streams the model directly from S3 to GPU memory using Run:ai Model Streamer.
Users send inference requests to the vLLM endpoint.
When you complete these steps, you have a vLLM inference endpoint that you can use to interact with a Ministral model through a chat frontend application. For additional information on optimizing model loading time on Amazon EKS, see Accelerate model loading on Amazon EKS.
Prerequisites
Complete the steps in the Cluster setup section.
If you opened a new terminal, set the cluster name and region you used in the Cluster Setup via CLI section:
export CLUSTER_NAME=ai-eks-docs
export AWS_REGION=us-east-2
Look up the model weights bucket you created in the Model weights S3 bucket step:
MODEL_BUCKET=$(aws s3api list-buckets \
--query "Buckets[?starts_with(Name, '${CLUSTER_NAME}-models-')].Name | [0]" \
--output text)
echo "Model bucket: ${MODEL_BUCKET}"
Step 1: Download the model from Hugging Face
In this step, you deploy a Kubernetes Job that downloads the model from Hugging Face and uploads it to the S3 bucket that you created in the prerequisites section.
To download the model, apply the following Job manifest:
Example Model download Job manifest
cat AWS Deep Learning Containers (DLCs), which are Docker images preinstalled with deep learning frameworks and optimized for performance on AWS infrastructure. DLCs include security patches, validated framework versions, and optimized GPU driver configurations.
This deployment uses the following AWS DLC for vLLM 0.21.0 with SOCI support: `public.ecr.aws/deep-learning-containers/vllm:0.21.0-gpu-py312-cu130-ubuntu22.04-ec2-v1.0-soci`.
The image tag indicates vLLM 0.21.0 with GPU support, Python 3.12, CUDA 13.0, Ubuntu 22.04, optimized for EC2-based workloads, and SOCI-enabled for faster container startup.
This manifest creates a Deployment that runs vLLM on a GPU node and streams the model directly from S3 into GPU memory using Run:ai Model Streamer. The manifest also creates a ClusterIP Service that exposes the vLLM endpoint on port 8000 for in-cluster access.
For additional information on optimizing model loading time on Amazon EKS, see Accelerate model loading on Amazon EKS. The following example uses `--enforce-eager` to accelerate load times in a getting started scenario. We recommend using other techniques to accelerate model load times as detailed in The --enforce-eager trade-off.
Apply the manifest:
Example vLLM Deployment and Service YAML
cat Cluster setup steps and view them on a pre-provisioned Grafana dashboard.
Important
You must complete the Monitoring subsection of the Cluster Setup via CLI section before continuing. This step depends on the kube-prometheus-stack being installed and the vLLM Grafana dashboard already provisioned in the values file.
Apply the vLLM ServiceMonitor
A ServiceMonitor tells Prometheus where to scrape vLLM metrics.
cat
To populate the dashboard with metrics, generate inference traffic against the vLLM endpoint you already exposed via port-forward in the validation step.
Discover the served model name:
MODEL_NAME=$(curl -s http://localhost:8000/v1/models | jq -r '.data[0].id') echo "Using model: $MODEL_NAME"
Send 50 chat completion requests in parallel:
for i in $(seq 1 50); do
curl -s -X POST http://localhost:8000/v1/chat/completions
-H "Content-Type: application/json"
-d "{"model": "$MODEL_NAME", "messages": [{"role": "user", "content": "Write a short poem about Kubernetes."}], "max_tokens": 128}"
> /dev/null &
done
wait
While traffic is flowing (or immediately after), check token-throughput metrics directly from the vLLM `/metrics` endpoint:
curl -s http://localhost:8000/metrics | grep -E '^vllm:(prompt_tokens_total|generation_tokens_total|avg_generation_throughput_toks_per_s|avg_prompt_throughput_toks_per_s)' | head
The `vllm:prompt_tokens_total` and `vllm:generation_tokens_total` metrics are monotonically increasing counters of input and output tokens served. The `vllm:avg_prompt_throughput_toks_per_s` and `vllm:avg_generation_throughput_toks_per_s` metrics are rolling-average throughput gauges. These same metrics power the Grafana dashboard you open in the following subsection.
View the vLLM Grafana dashboard
The kube-prometheus-stack values file from the Monitoring section already provisions the community vLLM dashboard (gnetId 25263) under the GPU Monitoring folder, so no extra import is needed.
Access Grafana through the load balancer you set up in the Access Grafana section. Print its URL:
echo "http://$(kubectl get ingress kube-prometheus-stack-grafana -n monitoring -o jsonpath='{.status.loadBalancer.ingress[0].hostname}')"
Open the URL in your browser and log in with username `admin` and the password from the following command:
kubectl --namespace monitoring get secrets kube-prometheus-stack-grafana -o jsonpath="{.data.admin-password}" | base64 -d ; echo
Navigate to Dashboards > GPU Monitoring > vLLM Metrics.
vLLM Grafana dashboard
The dashboard displays request rate, prompt and generation token throughput, latency percentiles, and GPU KV cache utilization for the vLLM inference endpoint.
Step 5: Deploy chat application
In this step, you deploy Open WebUI as a chat frontend to interact with the model. Open WebUI is an open source, self-hosted AI interface that supports OpenAI-compatible APIs and provides a chat interface with conversation history and markdown rendering. Because vLLM exposes an OpenAI-compatible API, Open WebUI connects to it directly as a backend.
To deploy the Open WebUI application, apply the following manifest:
Example Open WebUI Deployment and Service YAML
cat Set up load balancing.
Open WebUI is publicly accessible with no authentication
Open WebUI runs with authentication disabled (WEBUI_AUTH: "False"), so anyone who reaches the load balancer gets an unauthenticated chat interface backed by your GPU inference endpoint and can consume GPU capacity. Automated scanners find public load balancers within minutes. You must restrict access with the alb.ingress.kubernetes.io/inbound-cidrs annotation and treat source-IP allowlisting as a minimum safeguard rather than a complete one. For a stronger posture, use an internal scheme, add a TLS certificate, and enable Open WebUI authentication.
Find your public IP address and store it as a /32 CIDR:
export MY_CIDR="$(curl -s https://checkip.amazonaws.com)/32"
echo $MY_CIDR
The result looks like 203.0.113.4/32. If your network assigns addresses dynamically, your IP address can change, in which case you might need a broader range such as 203.0.113.0/24.
Apply the following Ingress to create the ALB:
cat Health check path
The `healthcheck-path` annotation points the load balancer health checks at the Open WebUI `/health` endpoint, because the ALB default health check matcher expects a 200 response.
The load balancer is created asynchronously and takes a minute or two. Print the URL:
echo "http://$(kubectl get ingress open-webui -o jsonpath='{.status.loadBalancer.ingress[0].hostname}')"
Open the URL in your browser. The chat interface appears and you can interact with the Ministral model.
Port-forwarding as an alternative
Port-forwarding (`kubectl port-forward svc/open-webui 8080:80`) remains an option for local testing without provisioning a load balancer.
Clean up
To remove the workload resources that you created in this section, delete the Open WebUI application, the vLLM inference server, and the model-download Job:
kubectl delete ingress open-webui kubectl delete deployment open-webui kubectl delete service open-webui kubectl delete deployment vllm-inference-app kubectl delete service vllm-inference-svc kubectl delete servicemonitor vllm-inference-app kubectl delete job model-download
For instructions on removing infrastructure resources such as the cluster, NodePool, and S3 bucket, see Cluster Setup Cleanup.
Javascript is disabled or is unavailable in your browser.
To use the Amazon Web Services Documentation, Javascript must be enabled. Please refer to your browser's Help pages for instructions.
Document Conventions
Inference
Autoscaling
Did this page help you? - Yes
Thanks for letting us know we're doing a good job!
If you've got a moment, please tell us what we did right so we can do more of it.
Did this page help you? - No
Thanks for letting us know this page needs work. We're sorry we let you down.
If you've got a moment, please tell us how we can make the documentation better.
## 더 알아보기 (Learn more)
- [공식 문서](https://docs.aws.amazon.com/eks/latest/userguide/ml-inference-load-serve-model.html)