CPU - Intel® Xeon®

CPU - Intel® Xeon®

이 페이지는 Intel Xeon에서 vLLM으로 테스트·검증된 하드웨어와 권장 모델 목록을 보여줘요. 자체 하드웨어에서 어떤 모델을 돌려야 할지 고민될 때, 이 목록이 좋은 출발점이 돼요.

출처: vLLM 공식 문서 — models-hardware_supported_models-cpu

검증된 하드웨어 (Validated Hardware)

Hardware
Intel® Xeon® 6 Processors
Intel® Xeon® 5 Processors

텍스트 전용 언어 모델 (Text-only Language Models)

Model Architecture Supported
unsloth/gpt-oss-20b GptOssForCausalLM
meta-llama/Llama-3.1-8B-Instruct LlamaForCausalLM
meta-llama/Llama-3.2-1B LlamaForCausalLM
meta-llama/Llama-3.2-3B-Instruct LlamaForCausalLM
meta-llama/Llama-3.3-70B-Instruct LlamaForCausalLM
RedHatAI/Meta-Llama-3.1-8B-quantized.w8a8 LlamaForCausalLM
RedHatAI/Meta-Llama-3.1-8B-Instruct-quantized.w8a8 LlamaForCausalLM
RedHatAI/Llama-3.2-1B-Instruct-quantized.w8a8 LlamaForCausalLM
RedHatAI/Llama-3.2-3B-Instruct-quantized.w8a8 LlamaForCausalLM
RedHatAI/DeepSeek-R1-Distill-Llama-70B-quantized.w8a8 LlamaForCausalLM
hugging-quants/Meta-Llama-3.1-8B-Instruct-AWQ-INT4 LlamaForCausalLM
AMead10/Llama-3.2-1B-Instruct-AWQ LlamaForCausalLM
AMead10/Llama-3.2-3B-Instruct-AWQ LlamaForCausalLM
TheBloke/TinyLlama-1.1B-Chat-v1.0-AWQ LlamaForCausalLM
TheBloke/TinyLlama-1.1B-Chat-v1.0-GPTQ LlamaForCausalLM
ibm-granite/granite-3.2-2b-instruct GraniteForCausalLM
Qwen/Qwen3-1.7B Qwen3ForCausalLM
Qwen/Qwen3-4B Qwen3ForCausalLM
Qwen/Qwen3-8B Qwen3ForCausalLM
Qwen/Qwen3-14B Qwen3ForCausalLM
Qwen/Qwen3-14B-AWQ Qwen3ForCausalLM
Qwen/Qwen3-30B-A3B Qwen3MoeForCausalLM
Qwen/QwQ-32B-AWQ Qwen2ForCausalLM
Qwen/Qwen1.5-0.5B-Chat-GPTQ-Int4 Qwen2ForCausalLM
RedHatAI/QwQ-32B-quantized.w8a8 Qwen2ForCausalLM
zai-org/glm-4-9b-hf GLMForCausalLM
google/gemma-7b GemmaForCausalLM
microsoft/Phi-4-reasoning Phi3ForCausalLM
TheBloke/Mistral-7B-Instruct-v0.2-AWQ MistralForCausalLM

멀티모달 언어 모델 (Multimodal Language Models)

Model Architecture Supported
meta-llama/Llama-4-Scout-17B-16E-Instruct Llama4ForConditionalGeneration
google/gemma-3-4b-it Gemma3ForConditionalGeneration
google/gemma-3-12b-it Gemma3ForConditionalGeneration
google/gemma-4-E4B-it Gemma4ForConditionalGeneration
google/gemma-4-E2B-it Gemma4ForConditionalGeneration
google/gemma-4-26B-A4B-it Gemma4ForConditionalGeneration
microsoft/Phi-4-multimodal-instruct Phi4MMForCausalLM
Qwen/Qwen2.5-VL-7B-Instruct Qwen2VLForConditionalGeneration
openai/whisper-large-v3 WhisperForConditionalGeneration

✅ 안내: "Runs and optimized" — 실행되고 최적화되어 있다는 의미예요.

더 알아보기 (Learn more)