트레이스에서 테스트 런 만들기
트레이스에서 테스트 런 만들기
테스트 런을 열고, 앱의 트레이스를 테스트 케이스로 스트리밍하고, Confident AI가 각각을 트레이스 수준과 컴포넌트 수준에서 평가하게 해요.
출처: 문서
본문
개요
변경을 배포하기 전에 앱이 여전히 기준을 통과하는지, 프로덕션으로 흘러들어가는 회귀가 없는지 알고 싶을 거예요. 그게 테스트 런의 용도죠: 여러분이 신경 쓰는 메트릭으로 각각 평가되는 테스트 케이스 배치로, 끝에 통과/실패 한 가지 판독이 있어요. 이 가이드는 앱이 모든 요청에서 이미 만드는 것 — 트레이스 — 으로 그것을 만듭니다.
핵심 아이디어는 한 문장으로 요약돼요: 열린 테스트 런에 트레이스를 대면, 그것이 평가된 테스트 케이스가 돼요. Confident AI가 메트릭으로 채점하고, 런 아래에 정리하고, 마지막 케이스가 도착하면 런을 닫아요.
트레이스는 두 가지 방법으로 Confident AI에 닿아요 — API를 통하거나 OpenTelemetry를 통해서요 — 그래서 가장 편한 스택으로 런을 만들 수 있어요.
graph LR
T["Your app's traces"] --> R["An open<br/>test run"]
R --> TC["Each trace becomes<br/>an evaluated test case"]
TC --> Res["Pass / fail read<br/>before you ship"]
style R fill:#eef2ff,stroke:#6366f1
style TC fill:#eef2ff,stroke:#6366f1
style Res fill:#eef2ff,stroke:#6366f1
이 가이드에서 배울 것:
- 테스트 런을 열고 트레이스를 이 런 안의 테스트 케이스로 캡처해요.
- 런 id와 metric collection으로 트레이스를 보내 클라우드에서 평가하고 자동으로 테스트 케이스로 변환해요.
- 각 트레이스 안의 단계를 평가해요 — retriever, LLM 호출 — 컴포넌트 수준 평가로 최종 답변뿐 아니라 봐요.
- OpenTelemetry로 테스트 케이스를 보내 스팬 속성으로, 이미 인스트루먼테이션이 있는 사용자를 위해요.
끝나면 매 릴리스 전에 품질을 벤치마킹하는 반복 가능한 방법을, 전적으로 앱의 트레이스에서 만들어 갖게 돼요.
트레이스 기반 테스트 런은 single-turn 이에요 — 테스트 케이스당 입력 하나, 출력 하나. metric collection을 single-turn 메트릭으로 만드세요.
사전 준비
첫 테스트 케이스가 도착하기 전에 두 가지가 있어야 해요:
- Project API Key —
CONFIDENT_API_KEY(예:confident_us_proj_...). 여기서 가져오세요. - 프로젝트에 single-turn metric collection. 이것이 테스트 케이스의 "좋음"을 정의해요 — 통과율, 품질 임계값, 릴리스를 게이트할 메트릭. 이름으로 참조하므로 정확한 이름을 적어 두세요.
API로 트레이스 보내기
각 단계는 Confident API를 호출해요 — 모든 스니펫이 레퍼런스로 바로 연결되고, 거기서 라이브로 호출을 시도할 수 있어요. OpenTelemetry로 인스트루먼트하는 쪽을 선호하나요? OpenTelemetry 섹션이 같은 흐름을 다룹니다.
테스트 런 열기
빈 런을 여는 것부터 시작해요. Confident AI가 id를 돌려주고, 런은 보내는 동안 계속 in progress — 테스트 케이스를 받을 준비가 된 상태 — 를 유지해요.
Request (POST /v1/test-runs) — API reference
curl -X POST "https://api.confident-ai.com/v1/test-runs" \
-H "CONFIDENT_API_KEY: <PROJECT-API-KEY>" \
-H "Content-Type: application/json" \
-d '{
"identifier": "my-test-run",
"metricCollection": "Agent Quality"
}'
import requests
response = requests.post(
"https://api.confident-ai.com/v1/test-runs",
headers={
"CONFIDENT_API_KEY": "<PROJECT-API-KEY>",
},
json={
"identifier": "my-test-run",
"metricCollection": "Agent Quality"
},
)
print(response.json())
const response = await fetch("https://api.confident-ai.com/v1/test-runs", {
method: "POST",
headers: {
"CONFIDENT_API_KEY": "<PROJECT-API-KEY>",
"Content-Type": "application/json",
},
body: JSON.stringify({
"identifier": "my-test-run",
"metricCollection": "Agent Quality"
}),
});
const data = await response.json();
console.log(data);
package main
import (
"fmt"
"io"
"net/http"
"strings"
)
func main() {
body := `{
"identifier": "my-test-run",
"metricCollection": "Agent Quality"
}`
req, err := http.NewRequest("POST", "https://api.confident-ai.com/v1/test-runs", strings.NewReader(body))
if err != nil {
panic(err)
}
req.Header.Set("CONFIDENT_API_KEY", "<PROJECT-API-KEY>")
req.Header.Set("Content-Type", "application/json")
res, err := http.DefaultClient.Do(req)
if err != nil {
panic(err)
}
defer res.Body.Close()
out, err := io.ReadAll(res.Body)
if err != nil {
panic(err)
}
fmt.Println(string(out))
}
import java.net.URI;
import java.net.http.HttpClient;
import java.net.http.HttpRequest;
import java.net.http.HttpResponse;
public class Example {
public static void main(String[] args) throws Exception {
String body = """
{
"identifier": "my-test-run",
"metricCollection": "Agent Quality"
}""";
HttpRequest request = HttpRequest.newBuilder()
.uri(URI.create("https://api.confident-ai.com/v1/test-runs"))
.header("CONFIDENT_API_KEY", "<PROJECT-API-KEY>")
.header("Content-Type", "application/json")
.POST(HttpRequest.BodyPublishers.ofString(body))
.build();
HttpResponse<String> response = HttpClient.newHttpClient()
.send(request, HttpResponse.BodyHandlers.ofString());
System.out.println(response.body());
}
}
use serde_json::json;
#[tokio::main]
async fn main() -> Result<(), Box<dyn std::error::Error>> {
let response = reqwest::Client::new()
.post("https://api.confident-ai.com/v1/test-runs")
.header("CONFIDENT_API_KEY", "<PROJECT-API-KEY>")
.json(&json!({
"identifier": "my-test-run",
"metricCollection": "Agent Quality"
}))
.send()
.await?;
println!("{}", response.text().await?);
Ok(())
}
두 필드 모두 선택 사양이에요. identifier는 나중에 런을 다시 찾을 라벨이고, metricCollection은 자기 이름을 지정하지 않은 모든 케이스의 기본 collection을 설정해요. 런의 id와 플랫폼의 link를 돌려받아요:
Response (POST /v1/test-runs) — API reference
{
"success": true,
"data": {
"id": "<TEST-RUN-ID>"
},
"link": "https://app.confident-ai.com/project/<PROJECT-ID>/test-runs/<TEST-RUN-ID>",
"deprecated": false
}
그 id를 붙잡아 두세요 — 테스트 케이스로 보내는 모든 트레이스가 그것을 담아요.
metricCollection은 프로젝트에 이미 존재하고 single-turn이어야 해요. 이름이 일치하지 않으면404를 반환하고, multi-turn collection은400을 반환해요.
트레이스를 테스트 케이스로 보내기
이제 런의 id를 testRunId로 담아 트레이스를 보내요. 그 한 필드가 평범한 트레이스를 테스트 케이스로 만드는 것이에요: Confident AI가 런으로 끌어올리고, 평가하고, 결과를 정리해요. 평가하는 metric collection은 아래처럼 트레이스에 실리거나, 아니면 런의 기본값으로 폴백해요.
Request (POST /v1/traces) — API reference
curl -X POST "https://api.confident-ai.com/v1/traces" \
-H "CONFIDENT_API_KEY: <PROJECT-API-KEY>" \
-H "Content-Type: application/json" \
-d '{
"uuid": "<TRACE-UUID>",
"input": "What is the capital of France?",
"output": "The capital of France is Paris.",
"startTime": "2025-01-15T10:30:00Z",
"endTime": "2025-01-15T10:30:05Z",
"testRunId": "<TEST-RUN-ID>",
"metricCollection": "Collection Name",
"baseSpans": [
{
"uuid": "<SPAN-UUID>",
"name": "Agent",
"input": "What is the capital of France?",
"output": "Let me look that up for you.",
"startTime": "2025-01-15T10:30:00Z",
"endTime": "2025-01-15T10:30:02Z"
}
]
}'
import requests
response = requests.post(
"https://api.confident-ai.com/v1/traces",
headers={
"CONFIDENT_API_KEY": "<PROJECT-API-KEY>",
},
json={
"uuid": "<TRACE-UUID>",
"input": "What is the capital of France?",
"output": "The capital of France is Paris.",
"startTime": "2025-01-15T10:30:00Z",
"endTime": "2025-01-15T10:30:05Z",
"testRunId": "<TEST-RUN-ID>",
"metricCollection": "Collection Name",
"baseSpans": [
{
"uuid": "<SPAN-UUID>",
"name": "Agent",
"input": "What is the capital of France?",
"output": "Let me look that up for you.",
"startTime": "2025-01-15T10:30:00Z",
"endTime": "2025-01-15T10:30:02Z"
}
]
},
)
print(response.json())
const response = await fetch("https://api.confident-ai.com/v1/traces", {
method: "POST",
headers: {
"CONFIDENT_API_KEY": "<PROJECT-API-KEY>",
"Content-Type": "application/json",
},
body: JSON.stringify({
"uuid": "<TRACE-UUID>",
"input": "What is the capital of France?",
"output": "The capital of France is Paris.",
"startTime": "2025-01-15T10:30:00Z",
"endTime": "2025-01-15T10:30:05Z",
"testRunId": "<TEST-RUN-ID>",
"metricCollection": "Collection Name",
"baseSpans": [
{
"uuid": "<SPAN-UUID>",
"name": "Agent",
"input": "What is the capital of France?",
"output": "Let me look that up for you.",
"startTime": "2025-01-15T10:30:00Z",
"endTime": "2025-01-15T10:30:02Z"
}
]
}),
});
const data = await response.json();
console.log(data);
package main
import (
"fmt"
"io"
"net/http"
"strings"
)
func main() {
body := `{
"uuid": "<TRACE-UUID>",
"input": "What is the capital of France?",
"output": "The capital of France is Paris.",
"startTime": "2025-01-15T10:30:00Z",
"endTime": "2025-01-15T10:30:05Z",
"testRunId": "<TEST-RUN-ID>",
"metricCollection": "Collection Name",
"baseSpans": [
{
"uuid": "<SPAN-UUID>",
"name": "Agent",
"input": "What is the capital of France?",
"output": "Let me look that up for you.",
"startTime": "2025-01-15T10:30:00Z",
"endTime": "2025-01-15T10:30:02Z"
}
]
}`
req, err := http.NewRequest("POST", "https://api.confident-ai.com/v1/traces", strings.NewReader(body))
if err != nil {
panic(err)
}
req.Header.Set("CONFIDENT_API_KEY", "<PROJECT-API-KEY>")
req.Header.Set("Content-Type", "application/json")
res, err := http.DefaultClient.Do(req)
if err != nil {
panic(err)
}
defer res.Body.Close()
out, err := io.ReadAll(res.Body)
if err != nil {
panic(err)
}
fmt.Println(string(out))
}
import java.net.URI;
import java.net.http.HttpClient;
import java.net.http.HttpRequest;
import java.net.http.HttpResponse;
public class Example {
public static void main(String[] args) throws Exception {
String body = """
{
"uuid": "<TRACE-UUID>",
"input": "What is the capital of France?",
"output": "The capital of France is Paris.",
"startTime": "2025-01-15T10:30:00Z",
"endTime": "2025-01-15T10:30:05Z",
"testRunId": "<TEST-RUN-ID>",
"metricCollection": "Collection Name",
"baseSpans": [
{
"uuid": "<SPAN-UUID>",
"name": "Agent",
"input": "What is the capital of France?",
"output": "Let me look that up for you.",
"startTime": "2025-01-15T10:30:00Z",
"endTime": "2025-01-15T10:30:02Z"
}
]
}""";
HttpRequest request = HttpRequest.newBuilder()
.uri(URI.create("https://api.confident-ai.com/v1/traces"))
.header("CONFIDENT_API_KEY", "<PROJECT-API-KEY>")
.header("Content-Type", "application/json")
.POST(HttpRequest.BodyPublishers.ofString(body))
.build();
HttpResponse<String> response = HttpClient.newHttpClient()
.send(request, HttpResponse.BodyHandlers.ofString());
System.out.println(response.body());
}
}
use serde_json::json;
#[tokio::main]
async fn main() -> Result<(), Box<dyn std::error::Error>> {
let response = reqwest::Client::new()
.post("https://api.confident-ai.com/v1/traces")
.header("CONFIDENT_API_KEY", "<PROJECT-API-KEY>")
.json(&json!({
"uuid": "<TRACE-UUID>",
"input": "What is the capital of France?",
"output": "The capital of France is Paris.",
"startTime": "2025-01-15T10:30:00Z",
"endTime": "2025-01-15T10:30:05Z",
"testRunId": "<TEST-RUN-ID>",
"metricCollection": "Collection Name",
"baseSpans": [
{
"uuid": "<SPAN-UUID>",
"name": "Agent",
"input": "What is the capital of France?",
"output": "Let me look that up for you.",
"startTime": "2025-01-15T10:30:00Z",
"endTime": "2025-01-15T10:30:02Z"
}
]
}))
.send()
.await?;
println!("{}", response.text().await?);
Ok(())
}
input과 output은 트레이스 수준 메트릭이 평가하는 것 — 앱이 무엇을 물었고, 어떻게 답했는지 — 이고, uuid, name, startTime, endTime(ISO-8601)이 그것을 완성해요. 그게 완전한 테스트 케이스예요.
모든 테스트 케이스는 평가할 metric collection이 필요해요 — 이 케이스를 특별히 평가하려면 트레이스에 설정하거나, 런을 열 때 기본값을 설정해서 모든 트레이스가 상속하게 해요. 둘 다 있으면 트레이스의 collection이 우선해요.
트레이스 안의 단계 평가하기
트레이스 수준 점수는 최종 답변이 좋았음을 말해 줘요. 왜인지는 알려주지 않아요 — 답이 틀렸을 때 어느 단계가 실패했는지도요. 잘못된 컨텍스트를 끌어온 retriever였나요, 올바른 것을 무시한 모델이었나요?
그걸 답하려면 각 스팬에 자체 metricCollection을 줘서 스팬을 평가할 수 있어요. 각 스팬은 컴포넌트이고, 독자적으로 평가되며 Observatory나 테스트 케이스의 트레이스 뷰에서 볼 수 있어요.
Request (POST /v1/traces) — API reference
curl -X POST "https://api.confident-ai.com/v1/traces" \
-H "CONFIDENT_API_KEY: <PROJECT-API-KEY>" \
-H "Content-Type: application/json" \
-d '{
"uuid": "<TRACE-UUID>",
"input": "What is the capital of France?",
"output": "The capital of France is Paris.",
"startTime": "2025-01-15T10:30:00Z",
"endTime": "2025-01-15T10:30:05Z",
"testRunId": "<TEST-RUN-ID>",
"metricCollection": "Collection Name",
"retrieverSpans": [
{
"uuid": "<RETRIEVER-SPAN-UUID>",
"name": "retrieve_context",
"embedder": "text-embedding-3-small",
"input": "capital of France",
"retrievalContext": [
"Paris is the capital and most populous city of France."
],
"startTime": "2025-01-15T10:30:00Z",
"endTime": "2025-01-15T10:30:01Z",
"metricCollection": "Retriever Collection Name"
}
],
"llmSpans": [
{
"uuid": "<LLM-SPAN-UUID>",
"parentUuid": "<RETRIEVER-SPAN-UUID>",
"name": "generate_answer",
"model": "gpt-4o",
"input": "Answer using the retrieved context.",
"output": "The capital of France is Paris.",
"startTime": "2025-01-15T10:30:01Z",
"endTime": "2025-01-15T10:30:05Z",
"metricCollection": "LLM Collection Name"
}
]
}'
import requests
response = requests.post(
"https://api.confident-ai.com/v1/traces",
headers={
"CONFIDENT_API_KEY": "<PROJECT-API-KEY>",
},
json={
"uuid": "<TRACE-UUID>",
"input": "What is the capital of France?",
"output": "The capital of France is Paris.",
"startTime": "2025-01-15T10:30:00Z",
"endTime": "2025-01-15T10:30:05Z",
"testRunId": "<TEST-RUN-ID>",
"metricCollection": "Collection Name",
"retrieverSpans": [
{
"uuid": "<RETRIEVER-SPAN-UUID>",
"name": "retrieve_context",
"embedder": "text-embedding-3-small",
"input": "capital of France",
"retrievalContext": [
"Paris is the capital and most populous city of France."
],
"startTime": "2025-01-15T10:30:00Z",
"endTime": "2025-01-15T10:30:01Z",
"metricCollection": "Retriever Collection Name"
}
],
"llmSpans": [
{
"uuid": "<LLM-SPAN-UUID>",
"parentUuid": "<RETRIEVER-SPAN-UUID>",
"name": "generate_answer",
"model": "gpt-4o",
"input": "Answer using the retrieved context.",
"output": "The capital of France is Paris.",
"startTime": "2025-01-15T10:30:01Z",
"endTime": "2025-01-15T10:30:05Z",
"metricCollection": "LLM Collection Name"
}
]
},
)
print(response.json())
const response = await fetch("https://api.confident-ai.com/v1/traces", {
method: "POST",
headers: {
"CONFIDENT_API_KEY": "<PROJECT-API-KEY>",
"Content-Type": "application/json",
},
body: JSON.stringify({
"uuid": "<TRACE-UUID>",
"input": "What is the capital of France?",
"output": "The capital of France is Paris.",
"startTime": "2025-01-15T10:30:00Z",
"endTime": "2025-01-15T10:30:05Z",
"testRunId": "<TEST-RUN-ID>",
"metricCollection": "Collection Name",
"retrieverSpans": [
{
"uuid": "<RETRIEVER-SPAN-UUID>",
"name": "retrieve_context",
"embedder": "text-embedding-3-small",
"input": "capital of France",
"retrievalContext": [
"Paris is the capital and most populous city of France."
],
"startTime": "2025-01-15T10:30:00Z",
"endTime": "2025-01-15T10:30:01Z",
"metricCollection": "Retriever Collection Name"
}
],
"llmSpans": [
{
"uuid": "<LLM-SPAN-UUID>",
"parentUuid": "<RETRIEVER-SPAN-UUID>",
"name": "generate_answer",
"model": "gpt-4o",
"input": "Answer using the retrieved context.",
"output": "The capital of France is Paris.",
"startTime": "2025-01-15T10:30:01Z",
"endTime": "2025-01-15T10:30:05Z",
"metricCollection": "LLM Collection Name"
}
]
}),
});
const data = await response.json();
console.log(data);
package main
import (
"fmt"
"io"
"net/http"
"strings"
)
func main() {
body := `{
"uuid": "<TRACE-UUID>",
"input": "What is the capital of France?",
"output": "The capital of France is Paris.",
"startTime": "2025-01-15T10:30:00Z",
"endTime": "2025-01-15T10:30:05Z",
"testRunId": "<TEST-RUN-ID>",
"metricCollection": "Collection Name",
"retrieverSpans": [
{
"uuid": "<RETRIEVER-SPAN-UUID>",
"name": "retrieve_context",
"embedder": "text-embedding-3-small",
"input": "capital of France",
"retrievalContext": [
"Paris is the capital and most populous city of France."
],
"startTime": "2025-01-15T10:30:00Z",
"endTime": "2025-01-15T10:30:01Z",
"metricCollection": "Retriever Collection Name"
}
],
"llmSpans": [
{
"uuid": "<LLM-SPAN-UUID>",
"parentUuid": "<RETRIEVER-SPAN-UUID>",
"name": "generate_answer",
"model": "gpt-4o",
"input": "Answer using the retrieved context.",
"output": "The capital of France is Paris.",
"startTime": "2025-01-15T10:30:01Z",
"endTime": "2025-01-15T10:30:05Z",
"metricCollection": "LLM Collection Name"
}
]
}`
req, err := http.NewRequest("POST", "https://api.confident-ai.com/v1/traces", strings.NewReader(body))
if err != nil {
panic(err)
}
req.Header.Set("CONFIDENT_API_KEY", "<PROJECT-API-KEY>")
req.Header.Set("Content-Type", "application/json")
res, err := http.DefaultClient.Do(req)
if err != nil {
panic(err)
}
defer res.Body.Close()
out, err := io.ReadAll(res.Body)
if err != nil {
panic(err)
}
fmt.Println(string(out))
}
import java.net.URI;
import java.net.http.HttpClient;
import java.net.http.HttpRequest;
import java.net.http.HttpResponse;
public class Example {
public static void main(String[] args) throws Exception {
String body = """
{
"uuid": "<TRACE-UUID>",
"input": "What is the capital of France?",
"output": "The capital of France is Paris.",
"startTime": "2025-01-15T10:30:00Z",
"endTime": "2025-01-15T10:30:05Z",
"testRunId": "<TEST-RUN-ID>",
"metricCollection": "Collection Name",
"retrieverSpans": [
{
"uuid": "<RETRIEVER-SPAN-UUID>",
"name": "retrieve_context",
"embedder": "text-embedding-3-small",
"input": "capital of France",
"retrievalContext": [
"Paris is the capital and most populous city of France."
],
"startTime": "2025-01-15T10:30:00Z",
"endTime": "2025-01-15T10:30:01Z",
"metricCollection": "Retriever Collection Name"
}
],
"llmSpans": [
{
"uuid": "<LLM-SPAN-UUID>",
"parentUuid": "<RETRIEVER-SPAN-UUID>",
"name": "generate_answer",
"model": "gpt-4o",
"input": "Answer using the retrieved context.",
"output": "The capital of France is Paris.",
"startTime": "2025-01-15T10:30:01Z",
"endTime": "2025-01-15T10:30:05Z",
"metricCollection": "LLM Collection Name"
}
]
}""";
HttpRequest request = HttpRequest.newBuilder()
.uri(URI.create("https://api.confident-ai.com/v1/traces"))
.header("CONFIDENT_API_KEY", "<PROJECT-API-KEY>")
.header("Content-Type", "application/json")
.POST(HttpRequest.BodyPublishers.ofString(body))
.build();
HttpResponse<String> response = HttpClient.newHttpClient()
.send(request, HttpResponse.BodyHandlers.ofString());
System.out.println(response.body());
}
}
use serde_json::json;
#[tokio::main]
async fn main() -> Result<(), Box<dyn std::error::Error>> {
let response = reqwest::Client::new()
.post("https://api.confident-ai.com/v1/traces")
.header("CONFIDENT_API_KEY", "<PROJECT-API-KEY>")
.json(&json!({
"uuid": "<TRACE-UUID>",
"input": "What is the capital of France?",
"output": "The capital of France is Paris.",
"startTime": "2025-01-15T10:30:00Z",
"endTime": "2025-01-15T10:30:05Z",
"testRunId": "<TEST-RUN-ID>",
"metricCollection": "Collection Name",
"retrieverSpans": [
{
"uuid": "<RETRIEVER-SPAN-UUID>",
"name": "retrieve_context",
"embedder": "text-embedding-3-small",
"input": "capital of France",
"retrievalContext": [
"Paris is the capital and most populous city of France."
],
"startTime": "2025-01-15T10:30:00Z",
"endTime": "2025-01-15T10:30:01Z",
"metricCollection": "Retriever Collection Name"
}
],
"llmSpans": [
{
"uuid": "<LLM-SPAN-UUID>",
"parentUuid": "<RETRIEVER-SPAN-UUID>",
"name": "generate_answer",
"model": "gpt-4o",
"input": "Answer using the retrieved context.",
"output": "The capital of France is Paris.",
"startTime": "2025-01-15T10:30:01Z",
"endTime": "2025-01-15T10:30:05Z",
"metricCollection": "LLM Collection Name"
}
]
}))
.send()
.await?;
println!("{}", response.text().await?);
Ok(())
}
여기서 트레이스는 여전히 자체 metricCollection으로 엔드 투 엔드 점수를 받고, 각 스팬은 여러분이 붙인 collection — 하나는 retriever, 하나는 모델 — 으로 평가돼요. 각 스팬의 type — retriever, llm, tool, agent — 을 설정해서 올바른 종류의 메트릭으로 평가되게 해요.
컴포넌트에 그 역무에 맞는 메트릭을 주세요: retriever에는 검색 관련성, 모델에는 faithfulness나 답변 품질, 툴에는 인자 정확성. 그래야 실패한 테스트 케이스가 원인 단계를 정확히 가리켜요.
결과 읽기
Confident AI가 자동으로 런을 닫고, 통과·실패 수를 집계하고, 완료로 표시해요 — 자세한 내용은 런이 끝나면을 보세요.
1단계의 link를 열어 플랫폼에서 런을 읽어요 — 모든 테스트 케이스, 그 트레이스, 두 수준의 점수까지. 대신 결과를 자체 파이프라인으로 가져오려면 런을 가져와요:
Request (GET /v1/test-runs/{testRunId}) — API reference
curl -X GET "https://api.confident-ai.com/v1/test-runs/{testRunId}" \
-H "CONFIDENT_API_KEY: <PROJECT-API-KEY>"
import requests
response = requests.get(
"https://api.confident-ai.com/v1/test-runs/{testRunId}",
headers={
"CONFIDENT_API_KEY": "<PROJECT-API-KEY>",
},
)
print(response.json())
const response = await fetch("https://api.confident-ai.com/v1/test-runs/{testRunId}", {
method: "GET",
headers: {
"CONFIDENT_API_KEY": "<PROJECT-API-KEY>",
},
});
const data = await response.json();
console.log(data);
package main
import (
"fmt"
"io"
"net/http"
)
func main() {
req, err := http.NewRequest("GET", "https://api.confident-ai.com/v1/test-runs/{testRunId}", nil)
if err != nil {
panic(err)
}
req.Header.Set("CONFIDENT_API_KEY", "<PROJECT-API-KEY>")
res, err := http.DefaultClient.Do(req)
if err != nil {
panic(err)
}
defer res.Body.Close()
out, err := io.ReadAll(res.Body)
if err != nil {
panic(err)
}
fmt.Println(string(out))
}
import java.net.URI;
import java.net.http.HttpClient;
import java.net.http.HttpRequest;
import java.net.http.HttpResponse;
public class Example {
public static void main(String[] args) throws Exception {
HttpRequest request = HttpRequest.newBuilder()
.uri(URI.create("https://api.confident-ai.com/v1/test-runs/{testRunId}"))
.header("CONFIDENT_API_KEY", "<PROJECT-API-KEY>")
.GET()
.build();
HttpResponse<String> response = HttpClient.newHttpClient()
.send(request, HttpResponse.BodyHandlers.ofString());
System.out.println(response.body());
}
}
#[tokio::main]
async fn main() -> Result<(), Box<dyn std::error::Error>> {
let response = reqwest::Client::new()
.get("https://api.confident-ai.com/v1/test-runs/{testRunId}")
.header("CONFIDENT_API_KEY", "<PROJECT-API-KEY>")
.send()
.await?;
println!("{}", response.text().await?);
Ok(())
}
완료 ✅. 응답이 런의 전체 metricsScores와 분석할 testCases 배열을 줘요.
OpenTelemetry로 트레이스 보내기
앱이 이미 OpenTelemetry 스팬을 내보낸다면, 같은 필드를 스팬 속성으로 쉽게 실어 Confident AI의 OTel 엔드포인트로 스팬을 내보낼 수 있어요.
이동은 위 워크스루와 동일해요. 여전히 API로 런을 열어 testRunId를 얻고, 거기서부터 모든 것이 이미 내보내는 스팬에 실려요. OTLP/HTTP exporter를 x-confident-api-key 헤더로 인증된 Confident AI의 OpenTelemetry 엔드포인트로 지정해요:
https://otel.confident-ai.com/v1/traces
그다음 두 그룹의 속성을 설정해요:
- 루트 스팬 — 테스트 케이스 자체:
confident.trace.test_run_id,confident.trace.metric_collection,confident.trace.input,confident.trace.output,confident.trace.name. - 각 자식 스팬 — 컴포넌트:
confident.span.type,confident.span.name,confident.span.input,confident.span.output,confident.span.metric_collection.
속성 값은 문자열이므로, 구조화된 것 — 예컨대 검색된 청크 리스트 — 은 JSON으로 인코딩해요.
OpenTelemetry는 언어에 구애받지 않으므로, 앱에서 네이티브하게 트레이스를 보내도 API나 일반 SDK 사용자와 같은 테스트 런을 얻을 수 있어요. 다양한 언어의 예시 몇 가지예요:
Rust
루트 스팬이 테스트 케이스로 만드는 confident.trace.* 속성을 담고, 자식 스팬이 컴포넌트를 평가하는 confident.span.*을 담아요.
use opentelemetry::{global, trace::{Tracer, TraceContextExt}, KeyValue};
use opentelemetry_otlp::WithExportConfig;
use opentelemetry_sdk::runtime;
use std::collections::HashMap;
fn init_tracer() {
let mut headers = HashMap::new();
headers.insert(
"x-confident-api-key".to_string(),
std::env::var("CONFIDENT_API_KEY").expect("CONFIDENT_API_KEY not set"),
);
opentelemetry_otlp::new_pipeline()
.tracing()
.with_exporter(
opentelemetry_otlp::new_exporter()
.http()
.with_endpoint("https://otel.confident-ai.com/v1/traces")
.with_headers(headers),
)
.install_batch(runtime::Tokio)
.expect("failed to install tracer");
}
#[tokio::main]
async fn main() {
init_tracer();
let tracer = global::tracer("confident-test-run");
let test_run_id = "your-test-run-id"; // the id from step 1
let input = "Can I get a refund on my annual plan after two months?";
let output = "Annual plans are refundable on a prorated basis within the first 30 days...";
// Root span → the test case, evaluated end to end.
tracer.in_span("Refund policy question", |cx| {
let root = cx.span();
root.set_attribute(KeyValue::new("confident.trace.test_run_id", test_run_id));
root.set_attribute(KeyValue::new("confident.trace.metric_collection", "Agent Quality"));
root.set_attribute(KeyValue::new("confident.trace.name", "Refund policy question"));
root.set_attribute(KeyValue::new("confident.trace.input", input));
root.set_attribute(KeyValue::new("confident.trace.output", output));
// Child span → a component, evaluated on its own.
tracer.in_span("generate_answer", |cx| {
let span = cx.span();
span.set_attribute(KeyValue::new("confident.span.type", "llm"));
span.set_attribute(KeyValue::new("confident.span.name", "generate_answer"));
span.set_attribute(KeyValue::new("confident.span.input", "Answer using the policy context..."));
span.set_attribute(KeyValue::new("confident.span.output", output));
span.set_attribute(KeyValue::new("confident.span.metric_collection", "Answer Quality"));
});
});
global::shutdown_tracer_provider(); // flush before the process exits
}
Clojure
인터롭을 통해 OpenTelemetry Java SDK를 사용해요. exporter가 x-confident-api-key 헤더로 Confident AI를 가리키고, 루트와 자식 스팬이 같은 속성을 담아요.
(ns traces
(:import
[io.opentelemetry.api.common Attributes]
[io.opentelemetry.exporter.otlp.http.trace OtlpHttpSpanExporter]
[io.opentelemetry.sdk OpenTelemetrySdk]
[io.opentelemetry.sdk.trace SdkTracerProvider]
[io.opentelemetry.sdk.trace.export BatchSpanProcessor]
[java.util.concurrent TimeUnit]))
(defn build-sdk []
(let [exporter (-> (OtlpHttpSpanExporter/builder)
(.setEndpoint "https://otel.confident-ai.com/v1/traces")
(.addHeader "x-confident-api-key" (System/getenv "CONFIDENT_API_KEY"))
(.build))
provider (-> (SdkTracerProvider/builder)
(.addSpanProcessor (-> (BatchSpanProcessor/builder exporter) (.build)))
(.build))]
(-> (OpenTelemetrySdk/builder)
(.setTracerProvider provider)
(.build))))
(defn -main []
(let [sdk (build-sdk)
tracer (.getTracer sdk "confident-test-run")
test-run-id "your-test-run-id" ; the id from step 1
input "Can I get a refund on my annual plan after two months?"
output "Annual plans are refundable on a prorated basis within the first 30 days..."
;; Root span → the test case, evaluated end to end.
root (-> (.spanBuilder tracer "Refund policy question")
(.setAllAttributes
(-> (Attributes/builder)
(.put "confident.trace.test_run_id" test-run-id)
(.put "confident.trace.metric_collection" "Agent Quality")
(.put "confident.trace.name" "Refund policy question")
(.put "confident.trace.input" input)
(.put "confident.trace.output" output)
(.build)))
(.startSpan))]
(with-open [_ (.makeCurrent root)]
;; Child span → a component, evaluated on its own.
(let [child (-> (.spanBuilder tracer "generate_answer")
(.setAllAttributes
(-> (Attributes/builder)
(.put "confident.span.type" "llm")
(.put "confident.span.name" "generate_answer")
(.put "confident.span.input" "Answer using the policy context...")
(.put "confident.span.output" output)
(.put "confident.span.metric_collection" "Answer Quality")
(.build)))
(.startSpan))]
(.end child)))
(.end root)
;; Flush before the process exits.
(.. sdk getSdkTracerProvider (shutdown) (join 10 TimeUnit/SECONDS))))
짧은 수명 스크립트는 스팬을 보내기 전에 종료될 수 있고, 테스트 케이스가 도착하지 않아요. 프로세스가 끝나기 전에 트레이서를 종료하세요 — Rust의
shutdown_tracer_provider(), Clojure의shutdown().join(...)— 그래야 배치가 먼저 플러시돼요.
스팬이 도착하면 같은 런에 합류해 API로 보낸 케이스처럼 마무리돼요. 결과를 읽는 것도 같은 방식이에요.
필드 & 속성 레퍼런스
API 본문과 OpenTelemetry 속성은 같은 테스트 케이스를 담아요 — 하나는 JSON에서 필드를, 다른 하나는 스팬에서 이름 지어요. 둘 사이를 오갈 때 이걸 써요:
| What it is | API field | OpenTelemetry attribute |
|---|---|---|
| The run to file the case under | testRunId |
confident.trace.test_run_id |
| Metrics that evaluate the test case (unless the run sets a default) | metricCollection |
confident.trace.metric_collection |
| What the app was asked | input |
confident.trace.input |
| What the app answered | output |
confident.trace.output |
| A name for the case | name |
confident.trace.name |
| A component's kind | spans[].type |
confident.span.type |
| A component's name | spans[].name |
confident.span.name |
| A component's input | spans[].input |
confident.span.input |
| A component's output | spans[].output |
confident.span.output |
| Metrics that evaluate a component | spans[].metricCollection |
confident.span.metric_collection |
개념
워크스루만으로 벤치마크를 실행하기에 충분해요. 이 섹션은 그 아래에서 무슨 일이 일어나는지 — 트레이스가 어떻게 테스트 케이스가 되는지, 런이 언제 끝났는지 아는 방법 — 을 다뤄요.
트레이스가 테스트 케이스가 되는 법
트레이스는 그 자체로 앱이 프로덕션에서 한 일의 기록이에요. 두 가지가 그것을 테스트 케이스로 승격해요: 열린 런을 가리키는 testRunId와 평가할 metricCollection이요.
sequenceDiagram
participant App as Your app
participant CA as Confident AI
participant Run as Test run
App->>CA: Open a run
CA-->>App: Run id
loop One per test case
App->>CA: Trace + run id + metrics
CA->>Run: File as a test case
CA->>CA: Evaluate against the metrics
end
Note over Run: Closes on its own once<br/>every case is scored
둘 다 있으면 Confident AI가 트레이스의 input·output에서 테스트 케이스를 도출하고, 이름 붙인 collection으로 평가하며, 런 아래에 결과를 정리해요. 트레이스의 나머지 — 스팬, 타이밍, 메타데이터 — 는 보존되므로, 테스트 케이스를 열면 점수 뒤의 정확한 트레이스가 보여요.
전체 vs 부분 평가
두 수준은 서로 다른 질문에 답하고, 같은 케이스에서 함께 돌아요:
- 전체 — 트레이스의
metricCollection— 이렇게 묻는 셈이에요: 이 입력이 주어졌을 때 최종 답변이 좋았는가? - 부분 — 스팬의
metricCollection— 이렇게 묻는 셈이에요: 이 한 단계가 제 역무를 다했나? 검색된 컨텍스트가 관련 있나, 툴이 올바른 인자를 받았나, 모델이 정책을 지켰나?
트레이스 수준 점수는 케이스가 실패했는지, 컴포넌트 수준 점수는 어디서 실패했는지 알려줘요. 함께면 빨간 테스트 케이스가 진단이 돼요.
런이 끝나면
"done" 호출은 없어요. 런은 계속 열려 있고 테스트 케이스가 계속 도착하는 동안 받아들여요 — 그래서 테스트 스위트 수명 동안 케이스를 스트리밍할 수 있어요.
Confident AI는 4시간 동안 유휴 — 새 테스트 케이스나 업데이트 없이 4시간 — 상태가 되면 런을 자동으로 닫아요. 그 시점에 런의 통과·실패 합계를 계산하고 완료로 표시해요 — 따라서 새 테스트 케이스를 더는 받지 않아요. 보내는 각 케이스가 유휴 시계를 리셋하므로, 활성 런은 스위트 중간에 닫히지 않아요. 스위트가 멈추면 런은 마지막 케이스가 도착한 후 약 4시간쯤 스스로 정착해요.
모범 사례
앱과 테스트 스위트가 커져도 트레이스 기반 런을 신뢰할 수 있게 유지하는 몇 가지 습관이에요:
- 수준별로 metric collection 하나를 유지하고 재사용하세요. 런의 모든 테스트 케이스를 같은 트레이스 수준 collection으로 평가해서 점수가 비교 가능하게 하고, 신경 쓰는 컴포넌트용 전용 collection은 아껴 두세요.
- 트레이스
uuid를 직접 소유하세요. 직접 생성해서 테스트 케이스를 자체 로그와 연결하고 필요할 때 결정적으로 다시 보낼 수 있게 해요. - 구조화된 OpenTelemetry 값은 JSON으로 인코딩하세요. 속성은 문자열이에요 — 리스트와 객체(스팬의 검색된 컨텍스트 등)를 JSON으로 인코딩해서 온전히 통과하게 해요.
- 종료 전에 항상 플러시하세요. 배치 exporter는 프로세스가 먼저 죽으면 스팬을 버려요. 테스트 스크립트 끝에 트레이서를 종료해서 모든 케이스가 보내지게 해요.
- 벤치마크를 프로덕션 데이터에 섞지 마세요. 배포 전 벤치마크를 전용 프로젝트나 환경에서 돌려 테스트 케이스가 프로덕션 관측성에 흐려지지 않게 해요.
FAQ
내 트레이스가 테스트 케이스가 아니라 일반 트레이스로 나타나요.
트레이스는 프로젝트의 열린 런을 가리키는 testRunId와 평가할 metric collection(트레이스에 설정하거나 런에서 상속)을 담을 때만 테스트 케이스가 돼요. testRunId가 잘못되었거나, 이미 닫힌 런을 가리키거나, 다른 프로젝트에 속하면 트레이스는 평범한 텔레메트리로 유지돼요.
API와 OpenTelemetry 둘 다로 트레이스를 보내야 하나요?
아니요 — 앱이 이미 쓰는 것을 고르세요. 둘 다 같은 런에서 같은 테스트 케이스를 만들므로, 섞을 수도 있어요: 일부 케이스를 API로, 다른 것을 OpenTelemetry로 같은 테스트 런에 보내도 돼요.
이런 식으로 대화를 벤치마킹할 수 있나요?
아직은요 — 트레이스 기반 런은 single-turn, 케이스당 입력·출력 하나예요. multi-turn 대화 평가는 metric collections과 multi-turn 평가 옵션을 보세요.
런이 끝났는지 어떻게 알죠?
GET /v1/test-runs/{id}를 가져와 상태를 확인하거나, 플랫폼에서 지켜보세요. 4시간 동안 유휴 — 새 테스트 케이스나 업데이트 없이 4시간 — 상태가 되면 저절로 완료로 바뀌어요. 런이 끝나면을 보세요.
다음 단계
이제 매 릴리스 전에 품질을 벤치마킹할 수 있어요, 전적으로 앱의 트레이스에서 만들었죠. 더 나아가려면:
Metric Collections
"좋음"이 무엇인지 정의해요 — 테스트 케이스를 엔드 투 엔드로, 단계 단계로 평가하는 collection이에요.
OpenTelemetry
Confident AI가 스팬을 어떻게 수집하는지, 어떤 confident.* 속성을 읽는지 보세요.
Trace Broadcasting
Go, Java, Ruby, C# 등에서 OpenTelemetry 트레이스를 Confident AI로 보내요.
LLM 평가
시간에 따른 런을 비교하고, 그것으로 릴리스를 게이트하고, 앱이 진화하면서 품질을 추적해요.