컨텍스트 확장

컨텍스트 확장 (Context Extension)

이 예제는 YARN 방식(rope_parameters)으로 Qwen 모델의 컨텍스트 길이를 확장하고 간단한 채팅을 실행하는 방법을 보여줍니다. rope_thetaoriginal_max_position_embeddings, factor를 조정해 모델의 최대 위치 임베딩을 늘립니다.

출처: 문서

본문

소스: https://github.com/vllm-project/vllm/tree/main/examples/features/context_extension

오프라인 컨텍스트 확장 (Context Extension Offline)

# SPDX-License-Identifier: Apache-2.0
# SPDX-FileCopyrightText: Copyright contributors to the vLLM project
"""This script demonstrates how to extend the context length
of a Qwen model using the YARN method (rope_parameters)
and run a simple chat example.

Usage:
    python examples/features/context_extension/context_extension_offline.py
"""

from vllm import LLM, RequestOutput, SamplingParams

def create_llm():
    rope_theta = 1000000
    original_max_position_embeddings = 32768
    factor = 4.0

    # Use yarn to extend context
    hf_overrides = {
        "rope_parameters": {
            "rope_theta": rope_theta,
            "rope_type": "yarn",
            "factor": factor,
            "original_max_position_embeddings": original_max_position_embeddings,
        },
        "max_model_len": int(original_max_position_embeddings * factor),
    }

    llm = LLM(model="Qwen/Qwen3-0.6B", hf_overrides=hf_overrides)
    return llm

def run_llm_chat(llm):
    sampling_params = SamplingParams(
        temperature=0.8,
        top_p=0.95,
        max_tokens=128,
    )

    conversation = [
        {"role": "system", "content": "You are a helpful assistant"},
        {"role": "user", "content": "Hello"},
        {"role": "assistant", "content": "Hello! How can I assist you today?"},
    ]
    outputs = llm.chat(conversation, sampling_params, use_tqdm=False)
    return outputs, [
        conversation,
    ]

def print_outputs(outputs: list[RequestOutput], conversations: list):
    print("\nGenerated Outputs:\n" + "-" * 80)
    for i, output in enumerate(outputs):
        prompt = conversations[i]
        generated_text = output.outputs[0].text
        print(f"Prompt: {prompt!r}\n")
        print(f"Generated text: {generated_text!r}")
        print("-" * 80)

def main():
    llm = create_llm()
    outputs, conversations = run_llm_chat(llm)
    print_outputs(outputs, conversations)

if __name__ == "__main__":
    main()

동작 요약:

  • rope_theta = 1000000, original_max_position_embeddings = 32768, factor = 4.0로 YARN(rope_type: "yarn") RoPE 설정을 적용합니다.
  • hf_overrides를 통해 rope_parametersmax_model_len = original_max_position_embeddings * factor(즉 131072)를 모델에 주입합니다.
  • 확장된 모델로 llm.chat()을 호출해 간단한 대화 예제를 실행하고 결과를 출력합니다. (llm.chat API는 채팅 템플릿을 자동 적용합니다.)

더 알아보기 (Learn more)

  • hf_overrides 엔진 인자 — 모델 config 오버라이드
  • YARN / RoPE rope_type 설정 문서