컨텍스트 확장
컨텍스트 확장 (Context Extension)
이 예제는 YARN 방식(rope_parameters)으로 Qwen 모델의 컨텍스트 길이를 확장하고 간단한 채팅을 실행하는 방법을 보여줍니다. rope_theta와 original_max_position_embeddings, factor를 조정해 모델의 최대 위치 임베딩을 늘립니다.
출처: 문서
본문
소스: https://github.com/vllm-project/vllm/tree/main/examples/features/context_extension
오프라인 컨텍스트 확장 (Context Extension Offline)
# SPDX-License-Identifier: Apache-2.0
# SPDX-FileCopyrightText: Copyright contributors to the vLLM project
"""This script demonstrates how to extend the context length
of a Qwen model using the YARN method (rope_parameters)
and run a simple chat example.
Usage:
python examples/features/context_extension/context_extension_offline.py
"""
from vllm import LLM, RequestOutput, SamplingParams
def create_llm():
rope_theta = 1000000
original_max_position_embeddings = 32768
factor = 4.0
# Use yarn to extend context
hf_overrides = {
"rope_parameters": {
"rope_theta": rope_theta,
"rope_type": "yarn",
"factor": factor,
"original_max_position_embeddings": original_max_position_embeddings,
},
"max_model_len": int(original_max_position_embeddings * factor),
}
llm = LLM(model="Qwen/Qwen3-0.6B", hf_overrides=hf_overrides)
return llm
def run_llm_chat(llm):
sampling_params = SamplingParams(
temperature=0.8,
top_p=0.95,
max_tokens=128,
)
conversation = [
{"role": "system", "content": "You are a helpful assistant"},
{"role": "user", "content": "Hello"},
{"role": "assistant", "content": "Hello! How can I assist you today?"},
]
outputs = llm.chat(conversation, sampling_params, use_tqdm=False)
return outputs, [
conversation,
]
def print_outputs(outputs: list[RequestOutput], conversations: list):
print("\nGenerated Outputs:\n" + "-" * 80)
for i, output in enumerate(outputs):
prompt = conversations[i]
generated_text = output.outputs[0].text
print(f"Prompt: {prompt!r}\n")
print(f"Generated text: {generated_text!r}")
print("-" * 80)
def main():
llm = create_llm()
outputs, conversations = run_llm_chat(llm)
print_outputs(outputs, conversations)
if __name__ == "__main__":
main()
동작 요약:
rope_theta = 1000000,original_max_position_embeddings = 32768,factor = 4.0로 YARN(rope_type: "yarn") RoPE 설정을 적용합니다.hf_overrides를 통해rope_parameters와max_model_len = original_max_position_embeddings * factor(즉 131072)를 모델에 주입합니다.- 확장된 모델로
llm.chat()을 호출해 간단한 대화 예제를 실행하고 결과를 출력합니다. (llm.chatAPI는 채팅 템플릿을 자동 적용합니다.)
더 알아보기 (Learn more)
hf_overrides엔진 인자 — 모델 config 오버라이드- YARN / RoPE
rope_type설정 문서