OpenAI Responses 호환 엔드포인트 연결하기
OpenAI Responses 호환 엔드포인트 연결하기
OpenAI나 Responses API 스키마를 말하는 어떤 엔드포인트든 AI Connection으로 연결하고, 응답에서 실제 출력과 툴 호출을 모두 뽑아낼 수 있어요. 스트리밍 방식과 비스트리밍 방식 모두 준비해 볼게요.
출처: 문서
본문
개요
이 가이드는 AI Connection을 OpenAI Responses API(POST https://api.openai.com/v1/responses)나, 같은 스키마로 만든 여러분의 엔드포인트로 연결해요. 그리고 돌아오는 응답에서 실제 출력과 툴 호출을 모두 파싱합니다. 래퍼 서비스도, 별도 코드도 필요 없어요.
같은 엔드포인트에 대해 두 개의 Connection을 만들 거예요:
- 비스트리밍(Non-streaming). Response mode
HTTP Response. 응답 객체 전체가 한 번에 도착해요. - 스트리밍(Streaming). Response mode
SSE Streaming에 body에"stream": true를 넣어요. 응답은 Server-Sent Events로 도착하고, 툴 호출은 마지막response.completed프레임에 담깁니다.
둘 다 같은 두 값을 뽑아내요: 실제 출력과 호출된 툴.
여기 있는 어떤 것도 OpenAI의 서버에만 국한되진 않아요. 이 가이드의 모든 내용은 요청·응답 스키마를 기준으로 하기 때문에, 응답 shaped 요청 본문을 받고 응답 shaped
output배열로 답하는 엔드포인트라면 여러분 엔드포인트에도 그대로 적용할 수 있어요. 여기에는 Azure OpenAI의/openai/v1/responses, OpenAI 앞에 둔 게이트웨이나 프록시, 같은 계약으로 만든 셀프호스팅 모델 서버, 그리고 그 스키마로 감싼 여러분의 에이전트까지 전부 해당합니다. 엔드포인트 필드에 여러분 URL을 넣고 auth 헤더만 바꾸면, 페이로드와 이벤트 이름, 두 변환기(transformer)는 그대로 유지돼요.
왜 툴 호출에 Transformer가 필요한가요
Responses API는 툴 호출을 ToolCall이 기대하는 형태로 돌려주지 않아요. 그래서 key path로는 거기에 닿을 수 없습니다. 세 가지가 걸림돌이에요:
- 툴 호출이
output안에 다른 것들과 함께 들어 있어요.output은reasoning,message,function_call항목이 섞인 배열이고, 순서는 모델이 만든 대로에요. key path가 가리킬 고정 인덱스가 없습니다. - 인자가 JSON 문자열로 도착해요.
"{\"city\":\"Hong Kong\"}"처럼 객체가 아니라 문자열이요.inputParameters는 진짜 객체여야 해요. - 필드 이름이 달라요. API는
arguments라고 부르고,ToolCall은inputParameters라고 불러요.
Transformer가 이 세 가지 간극을 모두 메워요. Project Settings → Transformers에서 두 개를 한 번만 만들어 두면, 두 Connection이 함께 씁니다.
Transformer는 Team 플랜 이상에서 사용할 수 있어요.
만들어 보기
툴 호출 transformer 작성하기
Project Settings → Transformers → New Transformer로 가서 이름을 openai_responses_tools로 짓고, 아래를 붙여 넣어요.
from typing import Any
import json
def transformer(data: Any):
# Both connections share this transformer, so normalize whatever arrives
# into the object that holds the `output` array:
# non-streaming: the full response object
# streaming: a `response.completed` frame, under "response"
# ping preview: a list of every frame received
frames = data if isinstance(data, list) else [data]
output = []
for frame in frames:
if not isinstance(frame, dict):
continue
response = frame.get("response") if isinstance(frame.get("response"), dict) else frame
if isinstance(response.get("output"), list):
output = response["output"]
tools_called = []
for item in output:
if not isinstance(item, dict) or item.get("type") != "function_call":
continue
arguments = item.get("arguments")
if isinstance(arguments, str):
try:
arguments = json.loads(arguments)
except json.JSONDecodeError:
arguments = {"raw_arguments": arguments}
tools_called.append(
{
"name": item.get("name"),
"inputParameters": arguments if isinstance(arguments, dict) else {},
}
)
return tools_called
저장하기 전에 내장 디버거로 확인해 보세요. 실제 Responses 페이로드를 Input 패널에 붙여 넣고 Test를 누르면, {"name": ..., "inputParameters": {...}} 객체들의 리스트가 돌아와야 해요.
isinstance(data, list)분기를 꼭 유지하세요. Ping에서는 스트리밍 Connection이 받은 프레임 전체 리스트를 transformer에 넘기고, 평가 실행에서는 매칭된 프레임 하나만 넘겨요. 단일 프레임을 가정한 transformer는 실행 중엔 잘 동작해도 ping에서 Invalid Tools Called Transformation으로 실패하기 쉽습니다.
실제 출력 transformer 작성하기
실제 출력은 문자열이어야 하는데, Responses 호출이 툴을 쓰기로 하면 보통 assistant 텍스트를 전혀 돌려주지 않아요. 그대로 두면 실제 출력이 비어서 스트리밍 Connection의 ping이 "Empty streaming response"로 실패하고, 메트릭이 채점할 값도 없어집니다.
텍스트가 없을 때 툴 호출로 폴백하는 두 번째 transformer openai_responses_output을 만들어요:
from typing import Any
def transformer(data: Any):
frames = data if isinstance(data, list) else [data]
output = []
for frame in frames:
if not isinstance(frame, dict):
continue
response = frame.get("response") if isinstance(frame.get("response"), dict) else frame
if isinstance(response.get("output"), list):
output = response["output"]
texts = []
tool_calls = []
for item in output:
if not isinstance(item, dict):
continue
if item.get("type") == "message":
for part in item.get("content") or []:
if isinstance(part, dict) and isinstance(part.get("text"), str):
texts.append(part["text"])
elif item.get("type") == "function_call":
tool_calls.append(f"{item.get('name')}({item.get('arguments')})")
if texts:
return "".join(texts)
# No assistant message, so surface the calls instead of returning ""
return "\n".join(tool_calls)
비스트리밍 Connection 만들기
Project Settings → AI Connections → New AI Connection으로 가서 이름을 OpenAI Responses (non-streaming)으로 짓고, 아래를 채워요.
General → AI App Endpoint
| Field | Value |
|---|---|
| Endpoint | https://api.openai.com/v1/responses |
| Response Mode | HTTP Response |
Responses 호환 엔드포인트를 직접 운영 중인가요? 그 URL을 대신 넣으면 돼요. 이 가이드의 다른 모든 단계는 그대로예요.
Headers
| Key | Value |
|---|---|
Authorization |
Bearer sk-... |
Content-Type |
application/json |
Body는 JSON payload 모드로 넣어요. golden.input은 따옴표 없이 입력하면, 저장할 때 에디터가 인코딩하고 요청 시점에 각 golden의 input으로 치환돼요.
{
"model": "gpt-5.4-mini",
"input": golden.input,
"tools": [
{
"type": "function",
"name": "get_weather",
"description": "Look up the current weather for a city",
"parameters": {
"type": "object",
"properties": {
"city": {
"type": "string",
"description": "City name, for example San Francisco"
}
},
"required": ["city"],
"additionalProperties": false
},
"strict": true
}
]
}
Responses API의 툴은 평평해요. type, name, description, parameters가 Chat Completions처럼 function 키 아래에 중첩되지 않고 최상위에 있어요. "strict": true를 쓰면 properties의 모든 키가 required에도 들어가야 하고, additionalProperties는 false여야 합니다.
Output parsing
| Parser | Setting |
|---|---|
| Actual output | Transformer → openai_responses_output |
| Tools called | Transformer → openai_responses_tools |
retrieval context와 state는 비워 두세요. 각 파서를 JSON Key Path에서 Transformer로 바꾸고 드롭다운에서 고른 뒤 저장해요. 각 파서는 따로 저장되기 때문에, 선택만 하고 저장하지 않은 transformer가 툴 호출이 비어 오는 가장 흔한 원인이에요.
Ping Endpoint를 누르면, 패널에 원시 응답 옆으로 파싱된 실제 출력과 툴 호출이 표시돼요.
JSON payload 모드에서는
state,prompts,hyperparameters,testCaseId,turnId같은 단어가 body에서 따옴표 없이 등장하는 곳(툴의description안까지) 어디든 payload 변수로 읽혀요."The state to look up"처럼 설명에 들어 있으면 저장할 때 잘못된 JSON으로 다시 쓰여 버립니다. 다른 문구로 바꾸거나, body를 Code 모드로 구성하세요.
스트리밍용으로 복제하기
방금 만든 Connection의 점 세 개 메뉴를 열고 Duplicate를 선택한 뒤, 복사본 이름을 OpenAI Responses (streaming)으로 바꿔요. 세 가지만 바꾸면 돼요.
General → AI App Endpoint: Response Mode를 SSE Streaming으로 설정해요.
Body: "stream": true를 추가해요.
{
"model": "gpt-5.4-mini",
"input": golden.input,
"stream": true,
"tools": [ ... ]
}
Output parsing: SSE Connection은 각 값을 어느 프레임이 담는지 알아야 해요. Responses API는 모든 프레임에 event:를 붙여서 response.created, response.output_text.delta, response.output_item.done을 지나 마지막에 response.completed까지 흘러가요. 이 마지막 프레임에 대고 두 파서를 모두 지정하세요.
| Parser | SSE Event Name | Accumulate Events | Extraction |
|---|---|---|---|
| Actual output | response.completed |
Off | Transformer → openai_responses_output |
| Tools called | response.completed |
n/a | Transformer → openai_responses_tools |
Ping을 눌러 보세요. 비스트리밍 Connection과 같은 파싱 값을, 이번에는 스트림에서 조립해서 볼 수 있어요.
모든 프레임은 payload에도 이벤트 이름을
type필드로 담고 있어서, SSE Payload Typeresponse.completed가 같은 프레임과 매칭돼요. 여러분과 OpenAI 사이의 프록시가event:라벨을 지워 버리는 경우 이걸 쓰면 됩니다. 둘 다 채우면 프레임은 두 조건을 모두 만족해야 해요.
response.completed만 전체output배열을 담는 프레임이에요. 개별response.output_item.done프레임은 각각 항목 하나씩만 담고, 매칭된 프레임은 마지막 것만 유지되기 때문에, tools called를response.output_item.done에 맞추면 마지막 툴 호출을 제외한 나머지를 조용히 놓치게 됩니다.
텍스트를 토큰 단위로 스트리밍하고 싶으면 실제 출력 SSE 이벤트를
response.output_text.delta로, key path를["delta"]로 설정하고 Accumulate Events를 켜면 돼요. tools called는response.completed에 남겨 두세요. 툴만 호출하는 응답은output_textdelta를 내보내지 않아서, 그 경우 실제 출력은 비어 돌아옵니다.
스트리밍 Connection에서 실제 출력 SSE 이벤트를 비워 두지 마세요.
response.function_call_arguments.delta와response.reasoning_summary_text.delta프레임도 최상위delta필드를 가져서, 이름 없는 누적 파서가 원시 툴 인자와 reasoning 텍스트를 실제 출력에 섞어 넣어요.
평가 실행하기
두 Connection을 전체 데이터셋에 지정하기 전에 ping을 해 볼 만한 입력이 세 개 있어요:
| Input | What it checks |
|---|---|
What's the weather in San Francisco right now? |
One tool call is extracted, with arguments parsed into an inputParameters object. |
Compare the weather in San Francisco and Tokyo right now. |
Parallel tool calls. You should get two entries. Getting one means the tools parser is on response.output_item.done. |
Say pong three times. |
The text path still works. No tool call, and actual output comes back as the assistant's message. |
모델이 툴을 호출하는 대신 문장으로 답하면 body에 "tool_choice": "required"를 추가해요.
이제 두 Connection 모두 모든 테스트 케이스에서 toolsCalled 리스트를 만들어 내고, 이게 툴 메트릭이 채점하는 값이에요. Tool Correctness는 toolsCalled를 golden의 expectedTools와 비교하니까, 각 golden에 이렇게 설정해 두세요:
[{ "name": "get_weather", "inputParameters": { "city": "San Francisco" } }]
메트릭을 metric collection에 추가하고, 각 Connection을 출력 생성 방식으로 선택한 single-turn evaluation을 돌려요. 같은 모델, 같은 툴, 하나는 스트리밍 하나는 비스트리밍이므로 두 테스트 런은 직접 비교할 수 있어요.
문제 해결
| Ping error | Cause |
|---|---|
Tools called comes back null |
The parser is still set to JSON Key Path, or the transformer was selected but not saved. Each parser saves on its own. |
Invalid Tools Called Transformation |
The transformer raised. Usually the list-of-frames shape on a streaming ping, so keep the isinstance(data, list) branch. |
Invalid Tools Called Return Type |
The return value isn't a list of ToolCall. Every entry needs a string name, and inputParameters has to be an object. |
Invalid Actual Output Return Type |
The actual output transformer returned None or something that isn't a string. |
Empty streaming response |
No frame matched the actual output parser, or the model returned tool calls only. Use the fallback in openai_responses_output. |
No data received from streaming response |
"stream": true is missing from the body, so OpenAI replied with a single JSON response instead of an event stream. |
401 or 403 |
The Authorization header is missing or malformed, or the key can't reach that model. |
다음 단계
AI Connections
엔드포인트, 페이로드, 출력 파싱, 헤더를 전체적으로 다뤄요.
Streaming
SSE 이벤트 이름, key path, accumulate 모드가 어떻게 함께 동작하는지 설명해요.
Transformers
응답을 변형하는 Python 함수를 작성·테스트·관리하는 방법이에요.
Authorization
정적 헤더 대신 secrets manager에서 API 키를 가져오는 방법이에요.