샘플링

샘플링 (Sampling)

서버가 클라이언트를 통해 언어 모델에 LLM 샘플링("completions" 또는 "generations")을 요청하는 표준화된 방법을 설명하는 페이지예요. 클라이언트가 모델 접근·선택·권한을 통제하면서 서버가 AI 기능을 활용할 수 있게 해 주며, 서버 API 키가 필요 없어요. 이 기능은 **폐기(Deprecated)**되었어요.

출처: 문서

본문

경고 (Warning): Deprecated: Sampling 기능은 프로토콜 버전 2026-07-28 기준으로 폐기되었어요 (SEP-2577). 기능 수명주기 정책 아래에서 이 개정이 릴리스된 후 최소 12개월 동안 사양에 남아 있다가 제거 자격을 얻어요. 새 구현은 SHOULD NOT 이를 채택하고, 기존 구현은 LLM 제공자 API에 직접 통합하는 방식으로 SHOULD 이전해야 해요. 폐기 기능 레지스트리를 참조하세요.

Model Context Protocol (MCP)은 서버가 클라이언트를 통해 언어 모델에 LLM 샘플링("completions" 또는 "generations")을 요청하는 표준화된 방법을 제공해요. 이 흐름은 클라이언트가 모델 접근·선택·권한에 대한 통제를 유지하면서 서버가 AI 기능을 활용할 수 있게 해 주고, 서버 API 키가 필요 없어요. 서버는 텍스트·오디오·이미지 기반 상호작용을 요청하고, 선택적으로 MCP 서버의 컨텍스트를 프롬프트에 포함할 수 있어요.

사용자 상호작용 모델 (User Interaction Model)

MCP의 Sampling은 LLM 호출이 다른 MCP 서버 기능 안에 *중첩(nested)*되어 일어나도록 해서, 서버가 에이전트적(agentic) 동작을 구현할 수 있게 해요.

구현체는 자신의 필요에 맞는 어떤 인터페이스 패턴으로든 sampling을 노출할 자유가 있어요. 프로토콜 자체는 특정 사용자 상호작용 모델을 강제하지 않아요.

경고 (Warning): 신뢰·안전과 보안을 위해, sampling 요청을 거부할 수 있는 능력을 가진 사람(human)이 항상 루프 안에 있어야 SHOULD 해요.

애플리케이션은 SHOULD:

  • sampling 요청을 쉽고 직관적으로 검토할 수 있는 UI를 제공한다
  • 사용자가 보내기 전에 프롬프트를 보고 편집할 수 있게 한다
  • 전달 전에 생성된 응답을 검토용으로 제시한다

샘플링의 도구 (Tools in Sampling)

서버는 sampling 요청에 tools 배열과 선택적 toolChoice 설정을 제공해 클라이언트의 LLM이 sampling 중 도구를 사용하도록 요청할 수 있어요. tools 배열의 도구 정의는 sampling 요청에 한정된 것이고, 등록된 도구와 일치할 필요가 없어요. 이를 통해 서버는 LLM이 특별히 지정된 도구를 호출하고, 결과를 받고, 대화를 계속하는 에이전트적 동작을 단일 sampling 요청 흐름 안에서 구현할 수 있어요.

클라이언트는 도구 지원 sampling 요청을 받으려면 sampling.tools 기능을 통해 도구 사용 지원을 MUST 선언해야 해요. 서버는 sampling.tools 기능으로 도구 사용 지원을 선언하지 않은 클라이언트에게 도구 활성 sampling 요청을 보내서는 MUST NOT 안 돼요.

기능 (Capabilities)

sampling을 지원하는 클라이언트는 MUST 각 요청의 _meta.io.modelcontextprotocol/clientCapabilities에 sampling 기능을 선언해야 해요:

기본 샘플링 (Basic sampling):

{
  "_meta": {
    "io.modelcontextprotocol/clientCapabilities": {
      "sampling": {}
    }
  }
}

도구 사용 지원 포함 (With tool use support):

{
  "_meta": {
    "io.modelcontextprotocol/clientCapabilities": {
      "sampling": {
        "tools": {}
      }
    }
  }
}

컨텍스트 포함 지원 (폐기됨) (With context inclusion support (deprecated)):

{
  "_meta": {
    "io.modelcontextprotocol/clientCapabilities": {
      "sampling": {
        "context": {}
      }
    }
  }
}

참고 (Note): includeContext 파라미터 값 "thisServer"와 "allServers"는 기능 수명주기 정책 아래에서 폐기되었어요 (SEP-2596). 이 값들은 Sampling 기능 자체보다 늦어도 함께 제거될 거예요. 서버는 SHOULD 이 값 사용을 피하고(예: includeContext는 기본값이 "none"이므로 그냥 생략 가능), 클라이언트가 sampling.context 기능을 선언하지 않았다면 SHOULD NOT 사용해야 해요. 폐기 기능 레지스트리를 참조하세요.

프로토콜 메시지 (Protocol Messages)

메시지 만들기 (Creating Messages)

클라이언트 요청 처리 중에 언어 모델 생성을 요청하려면, 서버는 sampling/createMessage 요청을 담은 InputRequiredResult를 보내요:

입력 요청 (InputRequiredResult.inputRequests 안에 전달됨):

{
  "method": "sampling/createMessage",
  "params": {
    "messages": [
      {
        "role": "user",
        "content": {
          "type": "text",
          "text": "What is the capital of France?"
        }
      }
    ],
    "modelPreferences": {
      "hints": [
        {
          "name": "claude-3-sonnet"
        }
      ],
      "costPriority": 0.3,
      "intelligencePriority": 0.8,
      "speedPriority": 0.5
    },
    "temperature": 0.1,
    "systemPrompt": "You are a helpful assistant.",
    "includeContext": "thisServer",
    "maxTokens": 100
  }
}

클라이언트 결과 (재시도된 요청의 inputResponses 안에 반환됨):

{
  "role": "assistant",
  "content": {
    "type": "text",
    "text": "The capital of France is Paris."
  },
  "model": "claude-3-sonnet-20240307",
  "stopReason": "endTurn"
}

도구를 사용한 샘플링 (Sampling with Tools)

다음 다이어그램은 다중 턴 도구 루프를 포함한 도구 사용 샘플링의 전체 흐름을 보여줘요:

sequenceDiagram
    participant Server
    participant Client
    participant User
    participant LLM

    Client->>Server: tools/call(id:1)
    note right of Server: Server needs more info
    Server->>Client: InputRequiredResult(<br/>sampling/createMessage<br/>(messages + tools))

    Note over Client,User: Human-in-the-loop review
    Client->>User: Present request for approval
    User-->>Client: Approve/modify

    Client->>LLM: Forward request with tools
    LLM-->>Client: Response with tool_use<br/>(stopReason: "toolUse")

    Client->>User: Present tool calls for review
    User-->>Client: Approve tool calls
    Client-->>Server: tools/call(id:2, Return tool_use response)

    Note over Server: Execute tool(s)
    Server->>Server: Run get_weather("Paris")<br/>Run get_weather("London")

    Note over Server,Client: Continue with tool results
    Server->>Client: InputRequiredResult(<br/>sampling/createMessage<br/>(history + tool_results + tools))

    Client->>User: Present continuation
    User-->>Client: Approve

    Client->>LLM: Forward with tool results
    LLM-->>Client: Final text response<br/>(stopReason: "endTurn")

    Client->>User: Present response
    User-->>Client: Approve
    Client-->>Server: tools/call(id:3, Return final response)

    Note over Server: Server processes result<br/>(may continue conversation...)

도구 사용 기능으로 LLM 생성을 요청하려면, 서버는 요청에 tools와 선택적으로 toolChoice를 포함해요:

입력 요청 (Server -> Client, InputRequiredResult.inputRequests 안에 전달됨):

{
  "method": "sampling/createMessage",
  "params": {
    "messages": [
      {
        "role": "user",
        "content": {
          "type": "text",
          "text": "What's the weather like in Paris and London?"
        }
      }
    ],
    "tools": [
      {
        "name": "get_weather",
        "description": "Get current weather for a city",
        "inputSchema": {
          "type": "object",
          "properties": {
            "city": {
              "type": "string",
              "description": "City name"
            }
          },
          "required": ["city"]
        }
      }
    ],
    "toolChoice": {
      "mode": "auto"
    },
    "maxTokens": 1000
  }
}

클라이언트 결과 (Client -> Server, 재시도된 요청의 inputResponses 안에 반환됨):

{
  "role": "assistant",
  "content": [
    {
      "type": "tool_use",
      "id": "call_abc123",
      "name": "get_weather",
      "input": {
        "city": "Paris"
      }
    },
    {
      "type": "tool_use",
      "id": "call_def456",
      "name": "get_weather",
      "input": {
        "city": "London"
      }
    }
  ],
  "model": "claude-3-sonnet-20240307",
  "stopReason": "toolUse"
}

다중 턴 도구 루프 (Multi-turn Tool Loop)

LLM에서 도구 사용 요청을 받은 후, 서버는 일반적으로:

  1. 요청된 도구 사용을 실행한다.
  2. 도구 결과를 덧붙여 새 sampling 요청을 보낸다.
  3. LLM의 응답(새 도구 사용을 포함할 수 있음)을 받는다.
  4. 필요에 따라 반복한다(서버는 최대 반복 횟수를 제한할 수 있고, 마지막 반복에서 toolChoice: {mode: "none"}를 전달해 최종 결과를 강제할 수 있음).

도구 결과가 담긴 후속 입력 요청 (Server -> Client, InputRequiredResult.inputRequests 안에 전달됨):

{
  "method": "sampling/createMessage",
  "params": {
    "messages": [
      {
        "role": "user",
        "content": {
          "type": "text",
          "text": "What's the weather like in Paris and London?"
        }
      },
      {
        "role": "assistant",
        "content": [
          {
            "type": "tool_use",
            "id": "call_abc123",
            "name": "get_weather",
            "input": { "city": "Paris" }
          },
          {
            "type": "tool_use",
            "id": "call_def456",
            "name": "get_weather",
            "input": { "city": "London" }
          }
        ]
      },
      {
        "role": "user",
        "content": [
          {
            "type": "tool_result",
            "toolUseId": "call_abc123",
            "content": [
              {
                "type": "text",
                "text": "Weather in Paris: 18°C, partly cloudy"
              }
            ]
          },
          {
            "type": "tool_result",
            "toolUseId": "call_def456",
            "content": [
              {
                "type": "text",
                "text": "Weather in London: 15°C, rainy"
              }
            ]
          }
        ]
      }
    ],
    "tools": [
      {
        "name": "get_weather",
        "description": "Get current weather for a city",
        "inputSchema": {
          "type": "object",
          "properties": {
            "city": { "type": "string" }
          },
          "required": ["city"]
        }
      }
    ],
    "maxTokens": 1000
  }
}

최종 클라이언트 결과 (Client -> Server, 재시도된 요청의 inputResponses 안에 반환됨):

{
  "role": "assistant",
  "content": {
    "type": "text",
    "text": "Based on the current weather data:\n\n- **Paris**: 18°C and partly cloudy - quite pleasant!\n- **London**: 15°C and rainy - you'll want an umbrella.\n\nParis has slightly warmer and drier conditions today."
  },
  "model": "claude-3-sonnet-20240307",
  "stopReason": "endTurn"
}

메시지 콘텐츠 제약 (Message Content Constraints)

도구 결과 메시지 (Tool Result Messages)

사용자 메시지에 도구 결과(type: "tool_result")가 포함되면, MUST 오직 도구 결과만 포함해야 해요. 같은 메시지에 도구 결과와 다른 콘텐츠 유형(텍스트, 이미지, 오디오)을 섞는 것은 허용되지 않아요.

이 제약은 도구 결과에 전용 역할을 사용하는 제공자 API(예: OpenAI의 "tool" 역할, Gemini의 "function" 역할)와의 호환성을 보장해요.

유효함 - 단일 도구 결과:

{
  "role": "user",
  "content": {
    "type": "tool_result",
    "toolUseId": "call_123",
    "content": [{ "type": "text", "text": "Result data" }]
  }
}

유효함 - 여러 도구 결과:

{
  "role": "user",
  "content": [
    {
      "type": "tool_result",
      "toolUseId": "call_123",
      "content": [{ "type": "text", "text": "Result 1" }]
    },
    {
      "type": "tool_result",
      "toolUseId": "call_456",
      "content": [{ "type": "text", "text": "Result 2" }]
    }
  ]
}

유효하지 않음 - 혼합 콘텐츠:

{
  "role": "user",
  "content": [
    {
      "type": "text",
      "text": "Here are the results:"
    },
    {
      "type": "tool_result",
      "toolUseId": "call_123",
      "content": [{ "type": "text", "text": "Result data" }]
    }
  ]
}

도구 사용과 결과 균형 (Tool Use and Result Balance)

sampling에서 도구 사용을 쓸 때, ToolUseContent 블록을 담은 모든 어시스턴트 메시지는 다른 어떤 메시지보다 먼저, 완전히 ToolResultContent 블록으로만 구성된 사용자 메시지가 뒤따라야 MUST 해요. 각 도구 사용(예: id: $id)은 대응하는 도구 결과(toolUseId: $id)와 일치해야 해요.

이 요구사항은 다음을 보장해요:

  • 도구 사용이 대화가 계속되기 전에 항상 해결된다
  • 제공자 API가 여러 도구 사용을 동시에 처리하고 결과를 병렬로 가져올 수 있다
  • 대화가 일관된 요청-응답 패턴을 유지한다

유효한 시퀀스 예시:

  1. 사용자 메시지: "What's the weather like in Paris and London?"
  2. 어시스턴트 메시지: ToolUseContent (id: "call_abc123", name: "get_weather", input: {city: "Paris"}) + ToolUseContent (id: "call_def456", name: "get_weather", input: {city: "London"})
  3. 사용자 메시지: ToolResultContent (toolUseId: "call_abc123", content: "18°C, partly cloudy") + ToolResultContent (toolUseId: "call_def456", content: "15°C, rainy")
  4. 어시스턴트 메시지: 두 도시의 날씨를 비교하는 텍스트 응답

유효하지 않은 시퀀스 - 도구 결과 누락:

  1. 사용자 메시지: "What's the weather like in Paris and London?"
  2. 어시스턴트 메시지: ToolUseContent (id: "call_abc123", name: "get_weather", input: {city: "Paris"}) + ToolUseContent (id: "call_def456", name: "get_weather", input: {city: "London"})
  3. 사용자 메시지: ToolResultContent (toolUseId: "call_abc123", content: "18°C, partly cloudy") ← call_def456에 대한 결과 누락
  4. 어시스턴트 메시지: 텍스트 응답 (유효하지 않음 - 모든 도구 사용이 해결되지 않음)

크로스 API 호환성 (Cross-API Compatibility)

샘플링 사양은 여러 LLM 제공자 API(Claude, OpenAI, Gemini 등)에서 작동하도록 설계됐어요. 호환성을 위한 핵심 설계 결정:

메시지 역할 (Message Roles)

MCP는 두 가지 역할("user"와 "assistant")을 사용해요.

도구 사용 요청은 "assistant" 역할로 CreateMessageResult에 보내져요. 도구 결과는 "user" 역할의 메시지로 돌아와요. 도구 결과가 담긴 메시지는 다른 종류의 콘텐츠를 담을 수 없어요.

도구 선택 모드 (Tool Choice Modes)

CreateMessageRequest.params.toolChoice는 모델의 도구 사용 능력을 제어해요:

  • {mode: "auto"}: 모델이 도구 사용 여부를 결정 (기본값)
  • {mode: "required"}: 모델은 완료 전에 최소 하나의 도구를 사용해야 MUST 함
  • {mode: "none"}: 모델은 어떤 도구도 사용해서는 MUST NOT 안 됨

병렬 도구 사용 (Parallel Tool Use)

MCP는 모델이 여러 도구 사용 요청을 병렬로 만들 수 있게 허용해요 (ToolUseContent 배열 반환). 주요 제공자 API는 모두 이를 지원해요:

  • Claude: 병렬 도구 사용을 기본 지원
  • OpenAI: 병렬 도구 호출 지원 (parallel_tool_calls: false로 비활성화 가능)
  • Gemini: 병렬 함수 호출을 기본 지원

병렬 도구 사용 비활성화를 지원하는 제공자를 감싸는 구현체는 MAY 이를 확장으로 노출할 수 있지만, 핵심 MCP 사양의 일부는 아니에요.

메시지 흐름 (Message Flow)

sequenceDiagram
    participant Server
    participant Client
    participant User
    participant LLM

    Client->>Server: tools/call(id:1)
    note right of Server: Server needs more info
    Server->>Client: InputRequiredResult(<br/>sampling/createMessage<br/>(messages + tools))

    Note over Client,User: Human-in-the-loop review
    Client->>User: Present request for approval
    User-->>Client: Review and approve/modify

    Note over Client,LLM: Model interaction
    Client->>LLM: Forward approved request
    LLM-->>Client: Return generation

    Note over Client,User: Response review
    Client->>User: Present response for approval
    User-->>Client: Review and approve/modify

    Note over Server,Client: Replay Request with approved response
    Client-->>Server: tools/call(id:3, Return approved response)

데이터 타입 (Data Types)

메시지 (Messages)

sampling 메시지는 MUST "user" 또는 "assistant"의 role 필드와 메시지 데이터를 나타내는 content 필드를 담아야 해요.

sampling 요청의 메시지 목록은 SHOULD NOT 별도 요청들 사이에 보존되지 말아야 해요.

content 필드는 다음을 담을 수 있어요:

텍스트 콘텐츠 (Text Content)

{
  "type": "text",
  "text": "The message content"
}

이미지 콘텐츠 (Image Content)

{
  "type": "image",
  "data": "base64-encoded-image-data",
  "mimeType": "image/jpeg"
}

오디오 콘텐츠 (Audio Content)

{
  "type": "audio",
  "data": "base64-encoded-audio-data",
  "mimeType": "audio/wav"
}

모델 선호 (Model Preferences)

MCP의 모델 선택은 서버와 클라이언트가 서로 다른 모델을 제공하는 다른 AI 제공자를 사용할 수 있으므로 신중한 추상화가 필요해요. 서버는 클라이언트가 그 정확한 모델에 접근할 수 없거나 다른 제공자의 동등한 모델을 선호할 수 있으므로, 특정 모델을 이름으로 단순히 요청할 수 없어요.

이를 해결하기 위해 MCP는 추상적 기능 우선순위와 선택적 모델 힌트를 결합한 선호 시스템을 구현해요:

기능 우선순위 (Capability Priorities)

서버는 정규화된 세 가지 우선순위 값(0-1)으로 필요를 표현해요:

  • costPriority: 비용 최소화가 얼마나 중요한가? 값이 높을수록 더 저렴한 모델을 선호해요.
  • speedPriority: 낮은 레이턴시가 얼마나 중요한가? 값이 높을수록 더 빠른 모델을 선호해요.
  • intelligencePriority: 고급 기능이 얼마나 중요한가? 값이 높을수록 더 유능한 모델을 선호해요.

모델 힌트 (Model Hints)

우선순위가 특성에 따라 모델을 선택하는 데 도움을 주는 반면, hints는 서버가 특정 모델이나 모델 계열을 제안하게 해요:

  • 힌트는 모델 이름을 유연하게 매칭할 수 있는 부분 문자열로 취급됨
  • 여러 힌트는 선호 순서대로 평가됨
  • 클라이언트는 MAY 힌트를 다른 제공자의 동등한 모델로 매핑할 수 있음
  • 힌트는 자문적(advisory)이며, 최종 모델 선택은 클라이언트가 함

예를 들어:

{
  "hints": [
    { "name": "claude-3-sonnet" }, // Prefer Sonnet-class models
    { "name": "claude" } // Fall back to any Claude model
  ],
  "costPriority": 0.3, // Cost is less important
  "speedPriority": 0.8, // Speed is very important
  "intelligencePriority": 0.5 // Moderate capability needs
}

클라이언트는 이 선호를 처리해 사용 가능한 옵션에서 적절한 모델을 선택해요. 예를 들어, 클라이언트가 Claude 모델에 접근할 수 없지만 Gemini가 있다면, 유사한 기능에 기반해 sonnet 힌트를 gemini-1.5-pro로 매핑할 수 있어요.

시스템 프롬프트 (System Prompt)

선택적 systemPrompt 필드는 서버가 특정 시스템 프롬프트를 요청할 수 있게 해요. 클라이언트는 MAY 이 필드를 서버에 알리지 않고 수정하거나 무시할 수 있어요.

컨텍스트 포함 (Context Inclusion)

includeContext 파라미터는 클라이언트가 응답에 포함해야 하는 컨텍스트 정보를 지정해요:

  • "none": 추가 컨텍스트 없음.
  • "thisServer": 요청한 서버의 컨텍스트 포함.
  • "allServers": 연결된 모든 MCP 서버의 컨텍스트 포함.

"thisServer"와 "allServers" 값은 폐기되었어요. 기능 (Capabilities)를 참조하세요.

클라이언트는 MAY 이 필드를 서버에 알리지 않고 수정하거나 무시할 수 있어요. 예를 들어, 특정 요청에서 이 필드를 존중하면 민감한 정보를 서버와 공유해야 하게 된다고 판단하면, 클라이언트는 그에 따라 응답을 제한할 수 있어요.

샘플링 파라미터 (Sampling Parameters)

LLM 샘플링은 다음 파라미터로 미세 조정할 수 있어요:

  • temperature: 모델 응답의 무작위성을 제어. 값이 높을수록 무작위성이 높고, 낮을수록 더 안정적인 출력. 유효 범위는 모델 제공자에 따라 다름.
  • maxTokens: 생성할 최대 토큰. 필수.
  • stopSequences: 생성을 중지하는 시퀀스의 배열.
  • metadata: 추가 제공자별 파라미터.

클라이언트는 MUST maxTokens 파라미터를 존중해야 해요.

클라이언트는 MAY temperature, stopSequences, metadata를 수정하거나 무시할 수 있어요. 예를 들어, 클라이언트가 이 파라미터 중 하나 이상을 지원하지 않는 모델을 사용할 수 있고, 따라서 이를 활용하지 못할 수 있어요.

결과 필드 (Result Fields)

sampling 결과는 다음 필드를 담아요:

  • role: 메시지 역할. 메시지 (Messages)를 참조.
  • content: 메시지 콘텐츠. 다음 중 하나일 수 있음:
    • 응답에 단일 콘텐츠 블록만 있을 때(단일 텍스트 응답 같은) 단일 콘텐츠 블록.
    • 응답에 하나 이상의 콘텐츠 블록이 있을 때(여러 도구 사용이나 혼합 콘텐츠 같은) 콘텐츠 블록 배열.
    • 콘텐츠 블록 유형은 메시지 (Messages)를 참조.
  • model: 메시지를 생성한 모델의 이름.
  • stopReason: 알고 있다면 sampling이 멈춘 이유. 사양은 다음(비완전) 중지 이유를 정의하며, 구현체는 MAY 자신만의 임의 값을 제공할 수 있음:
    • "endTurn": 참여자가 대화를 상대방에게 넘기고 있음.
    • "stopSequence": 메시지 생성이 요청된 stopSequences 중 하나를 만남.
    • "maxTokens": 토큰 한도에 도달함.
    • "toolUse": 모델이 하나 이상의 도구를 사용하려 함.

오류 처리 (Error Handling)

오류가 발생하거나 사용자가 sampling 요청을 거부하면, InputRequiredResult 패턴에서는 서버가 오류 메시지가 담긴 응답을 기다리지 않으므로 클라이언트는 오류 메시지를 넣어 초기 호출을 다시 재생할 필요가 없어요.

보안 고려 사항 (Security Considerations)

  1. 클라이언트는 SHOULD 사용자 승인 통제를 구현한다
  2. 양측 모두 SHOULD 메시지 콘텐츠를 검증한다
  3. 클라이언트는 SHOULD 모델 선호 힌트를 존중한다
  4. 클라이언트는 SHOULD 비율 제한(rate limiting)을 구현한다
  5. 양측 모두 MUST 민감한 데이터를 적절히 처리한다

sampling에 도구를 사용하면 추가 보안 고려 사항이 적용돼요:

  1. 서버는 MUST stopReason: "toolUse"에 응답할 때 각 ToolUseContent 항목에 일치하는 toolUseId를 가진 ToolResultContent 항목으로 응답하고, 사용자 메시지가 도구 결과만(다른 콘텐츠 유형 없음) 포함하도록 보장해야 한다
  2. 양측 모두 SHOULD 도구 루프에 대한 반복 한도를 구현한다

더 알아보기 (Learn more)