Voice Quickstart
Voice Quickstart (음성 퀵스타트)
출처: 문서
본문
사전 요구사항
Agents SDK의 기본 퀵스타트 지침을 따르고 가상 환경을 설정했는지 확인하세요. 그런 다음 SDK에서 선택적 음성 의존성을 설치해요.
pip install 'openai-agents[voice]'
아래 데모 코드는 마이크·스피커 I/O에 sounddevice도 사용하는데, 이건 voice extra에 포함되지 않아요.
pip install sounddevice
개념
알아야 할 핵심 개념은 VoicePipeline이에요. 이건 3단계 프로세스예요.
- 음성-텍스트(speech-to-text) 모델을 실행해 오디오를 텍스트로 바꿔요.
- 보통 에이전틱 워크플로인 당신의 코드를 실행해 결과를 만들어요.
- 텍스트-음성(text-to-speech) 모델을 실행해 결과 텍스트를 다시 오디오로 바꿔요.
graph LR
%% Input
A["🎤 Audio Input"]
%% Voice Pipeline
subgraph Voice_Pipeline [Voice Pipeline]
direction TB
B["Transcribe (speech-to-text)"]
C["Your Code"]:::highlight
D["Text-to-speech"]
B --> C --> D
end
%% Output
E["🎧 Audio Output"]
%% Flow
A --> Voice_Pipeline
Voice_Pipeline --> E
%% Custom styling
classDef highlight fill:#ffcc66,stroke:#333,stroke-width:1px,font-weight:700;
에이전트
먼저 몇몇 에이전트를 설정해 볼게요. 이 SDK로 에이전트를 만들어 봤다면 익숙할 거예요. 두 에이전트, 구성된 handoff, 하나의 도구를 만들 거예요.
import random
from agents import Agent
from agents.decorators import tool
from agents.extensions.handoff_prompt import prompt_with_handoff_instructions
@tool
def get_weather(city: str) -> str:
"""Get the weather for a given city."""
print(f"[debug] get_weather called with city: {city}")
choices = ["sunny", "cloudy", "rainy", "snowy"]
return f"The weather in {city} is {random.choice(choices)}."
spanish_agent = Agent(
name="Spanish",
handoff_description="A Spanish-speaking agent.",
instructions=prompt_with_handoff_instructions(
"You're speaking to a human, so be polite and concise. Speak in Spanish.",
),
model="gpt-5.6-sol",
)
agent = Agent(
name="Assistant",
instructions=prompt_with_handoff_instructions(
"You're speaking to a human, so be polite and concise. If the user speaks in Spanish, hand off to the Spanish agent.",
),
model="gpt-5.6-sol",
handoffs=[spanish_agent],
tools=[get_weather],
)
음성 파이프라인
SingleAgentVoiceWorkflow를 워크플로로 써서 간단한 음성 파이프라인을 설정해 볼게요.
from agents.voice import SingleAgentVoiceWorkflow, VoicePipeline
pipeline = VoicePipeline(workflow=SingleAgentVoiceWorkflow(agent))
파이프라인 실행
import numpy as np
import sounddevice as sd
from agents.voice import AudioInput
# For simplicity, we'll just create 3 seconds of silence
# In reality, you'd get microphone data
buffer = np.zeros(24000 * 3, dtype=np.int16)
audio_input = AudioInput(buffer=buffer)
result = await pipeline.run(audio_input)
# Create an audio player using `sounddevice`
player = sd.OutputStream(samplerate=24000, channels=1, dtype=np.int16)
player.start()
# Play the audio stream as it comes in
async for event in result.stream():
if event.type == "voice_stream_event_audio":
player.write(event.data)
전부 합치기
import asyncio
import random
import numpy as np
import sounddevice as sd
from agents import Agent
from agents.decorators import tool
from agents.voice import (
AudioInput,
SingleAgentVoiceWorkflow,
VoicePipeline,
)
from agents.extensions.handoff_prompt import prompt_with_handoff_instructions
@tool
def get_weather(city: str) -> str:
"""Get the weather for a given city."""
print(f"[debug] get_weather called with city: {city}")
choices = ["sunny", "cloudy", "rainy", "snowy"]
return f"The weather in {city} is {random.choice(choices)}."
spanish_agent = Agent(
name="Spanish",
handoff_description="A Spanish-speaking agent.",
instructions=prompt_with_handoff_instructions(
"You're speaking to a human, so be polite and concise. Speak in Spanish.",
),
model="gpt-5.6-sol",
)
agent = Agent(
name="Assistant",
instructions=prompt_with_handoff_instructions(
"You're speaking to a human, so be polite and concise. If the user speaks in Spanish, hand off to the Spanish agent.",
),
model="gpt-5.6-sol",
handoffs=[spanish_agent],
tools=[get_weather],
)
async def main():
pipeline = VoicePipeline(workflow=SingleAgentVoiceWorkflow(agent))
buffer = np.zeros(24000 * 3, dtype=np.int16)
audio_input = AudioInput(buffer=buffer)
result = await pipeline.run(audio_input)
# Create an audio player using `sounddevice`
player = sd.OutputStream(samplerate=24000, channels=1, dtype=np.int16)
player.start()
# Play the audio stream as it comes in
async for event in result.stream():
if event.type == "voice_stream_event_audio":
player.write(event.data)
if __name__ == "__main__":
asyncio.run(main())
이 예제를 실행하면 에이전트가 들을 수 있는 음성 오디오를 만들어 내요! 직접 에이전트에게 말해 볼 수 있는 데모는 examples/voice/static의 예제를 확인하세요.
더 알아보기 (Learn more)
- OpenAI Agents SDK 문서에서 더 많은 가이드를 확인하세요.