구조화된 출력(Structured output)
구조화된 출력(Structured output)
이 항목에서는 JSON Schema를 사용해 에이전트 워크플로에서 스키마 검증된 JSON을 반환하는 방법을 설명해요. 에이전트는 태스크를 완료하는 데 필요한 모든 도구를 사용할 수 있으며, 검증 성공 시 결과에는 스키마와 일치하는 구조화된 데이터가 포함돼요.
본문
필요한 구조에 대한 JSON Schema를 정의하면 SDK가 모델의 최종 출력을 그 스키마에 대해 검증해요.
왜 구조화된 출력인가?
에이전트는 기본적으로 자유 형식 텍스트를 반환해요. 이는 대화형 사용 사례에는 적합하지만 출력을 프로그래밍 방식으로 사용해야 할 때는 적합하지 않아요. 구조화된 출력은 애플리케이션 로직, 데이터베이스, UI 구성 요소에 직접 전달할 수 있는 타입화된 데이터를 제공해요.
코드베이스를 분석하는 에이전트를 생각해 보세요. 구조화된 출력 없이 자유 형식 텍스트를 받아 직접 파싱해야 해요. 구조화된 출력으로 원하는 형태를 정의하고 직접 사용할 수 있는 타입화된 데이터를 얻어요.
| 구조화된 출력 없음 | 구조화된 출력 있음 |
|---|---|
| This codebase uses Python and TypeScript. It has 42 files and the main entry point is... | { "languages": ["Python", "TypeScript"], "file_count": 42, "entry_point": "src/main.ts" } |
빠른 시작
outputFormat(TypeScript) 또는 output_format(Python) 옵션에 JSON Schema를 전달해요. 검증이 성공하면 결과 메시지에 스키마와 일치하는 데이터가 있는 structured_output 필드가 포함돼요. 에이전트가 재시도 후에도 스키마를 충족할 수 없으면 SDK는 대신 오류 결과를 반환해요.
import { query } from "cortex-code-agent-sdk";
const schema = {
type: "object",
properties: {
company_name: { type: "string" },
founded_year: { type: "number" },
headquarters: { type: "string" },
},
required: ["company_name"],
};
for await (const message of query({
prompt: "Research Snowflake and provide key company information",
options: {
cwd: process.cwd(),
outputFormat: { type: "json_schema", schema },
},
})) {
if (message.type === "result" && message.structured_output) {
console.log(message.structured_output);
// { company_name: "Snowflake", founded_year: 2012, headquarters: "Bozeman, MT" }
}
}
import asyncio
from cortex_code_agent_sdk import query, CortexCodeAgentOptions, ResultMessage
schema = {
"type": "object",
"properties": {
"company_name": {"type": "string"},
"founded_year": {"type": "number"},
"headquarters": {"type": "string"},
},
"required": ["company_name"],
}
async def main():
async for message in query(
prompt="Research Snowflake and provide key company information",
options=CortexCodeAgentOptions(
cwd=".",
output_format={"type": "json_schema", "schema": schema},
),
):
if isinstance(message, ResultMessage) and message.structured_output:
print(message.structured_output)
asyncio.run(main())
Zod 및 Pydantic으로 타입 안전 스키마
JSON Schema를 직접 작성하는 대신 Zod(TypeScript) 또는 Pydantic(Python)을 사용해 스키마를 정의할 수 있어요. 이 라이브러리들은 JSON Schema를 생성하고 응답을 자동 완성과 타입 검사가 있는 완전히 타입화된 객체로 파싱하게 해줘요.
import { z } from "zod";
import { query } from "cortex-code-agent-sdk";
const FeaturePlan = z.object({
feature_name: z.string(),
summary: z.string(),
steps: z.array(
z.object({
step_number: z.number(),
description: z.string(),
estimated_complexity: z.enum(["low", "medium", "high"]),
})
),
risks: z.array(z.string()),
});
type FeaturePlan = z.infer<typeof FeaturePlan>;
const schema = z.toJSONSchema(FeaturePlan);
for await (const message of query({
prompt: "Plan how to add dark mode support to a React app.",
options: {
cwd: process.cwd(),
outputFormat: { type: "json_schema", schema },
},
})) {
if (message.type === "result" && message.structured_output) {
const parsed = FeaturePlan.safeParse(message.structured_output);
if (parsed.success) {
const plan: FeaturePlan = parsed.data;
console.log(`Feature: ${plan.feature_name}`);
plan.steps.forEach((step) => {
console.log(`${step.step_number}. [${step.estimated_complexity}] ${step.description}`);
});
}
}
}
import asyncio
from pydantic import BaseModel
from cortex_code_agent_sdk import query, CortexCodeAgentOptions, ResultMessage
class Step(BaseModel):
step_number: int
description: str
estimated_complexity: str # 'low', 'medium', 'high'
class FeaturePlan(BaseModel):
feature_name: str
summary: str
steps: list[Step]
risks: list[str]
async def main():
async for message in query(
prompt="Plan how to add dark mode support to a React app.",
options=CortexCodeAgentOptions(
cwd=".",
output_format={
"type": "json_schema",
"schema": FeaturePlan.model_json_schema(),
},
),
):
if isinstance(message, ResultMessage) and message.structured_output:
plan = FeaturePlan.model_validate(message.structured_output)
print(f"Feature: {plan.feature_name}")
for step in plan.steps:
print(f"{step.step_number}. [{step.estimated_complexity}] {step.description}")
asyncio.run(main())
예시: TODO 추적 에이전트
이 예시는 다단계 도구 사용과 함께 구조화된 출력을 보여줘요. 에이전트는 기본 제공 도구(Grep, Bash)를 사용해 코드베이스에서 TODO 주석을 찾은 뒤 결과를 구조화된 데이터로 반환해요. author 같은 선택 필드는 git blame 정보를 사용할 수 없을 때를 처리해요.
import { query } from "cortex-code-agent-sdk";
const todoSchema = {
type: "object",
properties: {
todos: {
type: "array",
items: {
type: "object",
properties: {
text: { type: "string" },
file: { type: "string" },
line: { type: "number" },
author: { type: "string" },
date: { type: "string" },
},
required: ["text", "file", "line"],
},
},
total_count: { type: "number" },
},
required: ["todos", "total_count"],
};
for await (const message of query({
prompt: "Find all TODO comments in this codebase and identify who added them",
options: {
cwd: process.cwd(),
outputFormat: { type: "json_schema", schema: todoSchema },
},
})) {
if (message.type === "result" && message.structured_output) {
const data = message.structured_output;
console.log(`Found ${data.total_count} TODOs`);
data.todos.forEach((todo) => {
console.log(`${todo.file}:${todo.line} - ${todo.text}`);
if (todo.author) {
console.log(` Added by ${todo.author} on ${todo.date}`);
}
});
}
}
import asyncio
from cortex_code_agent_sdk import query, CortexCodeAgentOptions, ResultMessage
todo_schema = {
"type": "object",
"properties": {
"todos": {
"type": "array",
"items": {
"type": "object",
"properties": {
"text": {"type": "string"},
"file": {"type": "string"},
"line": {"type": "number"},
"author": {"type": "string"},
"date": {"type": "string"},
},
"required": ["text", "file", "line"],
},
},
"total_count": {"type": "number"},
},
"required": ["todos", "total_count"],
}
async def main():
async for message in query(
prompt="Find all TODO comments in this codebase and identify who added them",
options=CortexCodeAgentOptions(
cwd=".",
output_format={"type": "json_schema", "schema": todo_schema},
),
):
if isinstance(message, ResultMessage) and message.structured_output:
data = message.structured_output
print(f"Found {data['total_count']} TODOs")
for todo in data["todos"]:
print(f"{todo['file']}:{todo['line']} - {todo['text']}")
if "author" in todo:
print(f" Added by {todo['author']} on {todo['date']}")
asyncio.run(main())
예시: SQL 쿼리 결과
Cortex Code는 기본 제공 Snowflake SQL 도구를 갖고 있어요. 이를 구조화된 출력과 결합해 타입화된 쿼리 결과를 얻을 수 있어요.
import { query } from "cortex-code-agent-sdk";
const schema = {
type: "object",
properties: {
top_customers: {
type: "array",
items: {
type: "object",
properties: {
name: { type: "string" },
total_revenue: { type: "number" },
order_count: { type: "number" },
},
required: ["name", "total_revenue", "order_count"],
},
},
query_used: { type: "string" },
},
required: ["top_customers", "query_used"],
};
for await (const message of query({
prompt: "Find the top 5 customers by revenue from the ORDERS table",
options: {
cwd: process.cwd(),
connection: "my-connection",
outputFormat: { type: "json_schema", schema },
},
})) {
if (message.type === "result" && message.structured_output) {
const { top_customers, query_used } = message.structured_output;
console.log(`Query: ${query_used}`);
top_customers.forEach((c) => {
console.log(`${c.name}: $${c.total_revenue} (${c.order_count} orders)`);
});
}
}
import asyncio
from cortex_code_agent_sdk import query, CortexCodeAgentOptions, ResultMessage
schema = {
"type": "object",
"properties": {
"top_customers": {
"type": "array",
"items": {
"type": "object",
"properties": {
"name": {"type": "string"},
"total_revenue": {"type": "number"},
"order_count": {"type": "number"},
},
"required": ["name", "total_revenue", "order_count"],
},
},
"query_used": {"type": "string"},
},
"required": ["top_customers", "query_used"],
}
async def main():
async for message in query(
prompt="Find the top 5 customers by revenue from the ORDERS table",
options=CortexCodeAgentOptions(
cwd=".",
connection="my-connection",
output_format={"type": "json_schema", "schema": schema},
),
):
if isinstance(message, ResultMessage) and message.structured_output:
data = message.structured_output
print(f"Query: {data['query_used']}")
for c in data["top_customers"]:
print(f"{c['name']}: ${c['total_revenue']} ({c['order_count']} orders)")
asyncio.run(main())