이미지 이해
이미지 이해 (Image understanding)
Gemini 모델은 처음부터 멀티모달로 설계되어 특수 ML 모델을 훈련하지 않고도 이미지 캡션 생성, 분류, 시각 질의응답을 비롯한 다양한 이미지 처리·컴퓨터 비전 작업을 처리할 수 있어요.
일반적인 멀티모달 능력에 더해, Gemini 모델은 추가 훈련을 통해 객체 감지(object detection)와 분할(segmentation) 같은 특정 용례에 대해 향상된 정확도를 제공해요.
출처: 문서
본문
Gemini에 이미지 전달하기
Gemini에 이미지를 입력으로 제공하는 방법은 여러 가지가 있어요.
- URL로 이미지 전달: 공개적으로 접근 가능한 이미지에 적합해요.
- 인라인 이미지 데이터 전달: base64-encoded 이미지 데이터용이에요.
- File API로 이미지 업로드: 큰 파일이거나 여러 요청에서 이미지를 재사용할 때 권장돼요.
URL로 이미지 전달
Files API로 이미지를 업로드하고 요청에 전달할 수 있어요.
Python
from google import genai
client = genai.Client()
uploaded_file = client.files.upload(file="path/to/organ.jpg")
interaction = client.interactions.create(
model="gemini-3.8-flash",
input=[
{"type": "text", "text": "Caption this image."},
{
"type": "image",
"uri": uploaded_file.uri,
"mime_type": uploaded_file.mime_type
}
]
)
print(interaction.output_text)
JavaScript
import { GoogleGenAI } from "@google/genai";
const client = new GoogleGenAI({});
const uploadedFile = await client.files.upload({
file: "path/to/organ.jpg",
config: { mimeType: "image/jpeg" }
});
const interaction = await client.interactions.create({
model: "gemini-3.8-flash",
input: [
{type: "text", text: "Caption this image."},
{
type: "image",
uri: uploadedFile.uri,
mime_type: uploadedFile.mimeType
}
]
});
console.log(interaction.output_text);
Java
import com.google.genai.Client;
import com.google.genai.gaos.models.interactions.Content;
import com.google.genai.gaos.models.interactions.CreateModelInteraction;
import com.google.genai.gaos.models.interactions.ImageContent;
import com.google.genai.gaos.models.interactions.ImageContentMimeType;
import com.google.genai.gaos.models.interactions.Interaction;
import com.google.genai.gaos.models.interactions.InteractionsInput;
import com.google.genai.gaos.models.interactions.Model;
import com.google.genai.gaos.models.interactions.TextContent;
import com.google.genai.gaos.models.operations.CreateInteractionRequestBody;
import com.google.genai.types.File;
import com.google.genai.types.UploadFileConfig;
import java.util.Arrays;
import java.util.List;
Client client = new Client();
File uploadedFile =
client.files.upload(
new java.io.File("path/to/organ.jpg"),
UploadFileConfig.builder().mimeType("image/jpeg").build());
Content textContent = TextContent.builder().text("Caption this image.").build();
Content imageContent =
ImageContent.builder()
.uri(uploadedFile.uri().orElse(""))
.mimeType(ImageContentMimeType.of(uploadedFile.mimeType().orElse("image/jpeg")))
.build();
List<Content> contents = Arrays.asList(textContent, imageContent);
CreateModelInteraction params =
CreateModelInteraction.builder()
.model(Model.of("gemini-3.8-flash"))
.input(InteractionsInput.ofContent(contents))
.build();
Interaction interaction =
client.interactions.create(CreateInteractionRequestBody.of(params)).interaction().get();
System.out.println(interaction.outputText().orElse(""));
Go
package main
import (
"context"
"fmt"
"log"
"google.golang.org/genai"
"google.golang.org/genai/interactions/models/interactions"
"google.golang.org/genai/interactions/models/operations"
)
func main() {
ctx := context.Background()
client, err := genai.NewClient(ctx, nil)
if err != nil {
log.Fatal(err)
}
uploadedFile, err := client.Files.UploadFromPath(ctx, "path/to/organ.jpg", &genai.UploadFileConfig{
MIMEType: "image/jpeg",
})
if err != nil {
log.Fatal(err)
}
res, err := client.Interactions.Create(ctx, operations.CreateInteractionRequest{
Body: operations.NewCreateInteractionRequestBody(interactions.CreateModelInteraction{
Model: interactions.Model("gemini-3.8-flash"),
Input: interactions.NewInteractionsInput([]interactions.Content{
interactions.NewContent(interactions.TextContent{
Text: "Caption this image.",
}),
interactions.NewContent(interactions.ImageContent{
URI: genai.Ptr(uploadedFile.URI),
MimeType: interactions.ImageContentMimeType(uploadedFile.MIMEType).ToPointer(),
}),
}),
}),
})
if err != nil {
log.Fatal(err)
}
if res.Interaction.OutputText != nil {
fmt.Println(*res.Interaction.OutputText)
}
}
REST
# First upload the file using the Files API, then use the URI:
curl -X POST "https://generativelanguage.googleapis.com/v1beta/interactions" \
-H "x-goog-api-key: $GEMINI_API_KEY" \
-H 'Content-Type: application/json' \
-d '{
"model": "gemini-3.8-flash",
"input": [
{"type": "text", "text": "Caption this image."},
{
"type": "image",
"uri": "YOUR_FILE_URI",
"mime_type": "image/jpeg"
}
]
}'
인라인 이미지 데이터 전달
이미지 데이터를 base64-encoded 문자열로 제공할 수 있어요.
Python
import base64
from google import genai
with open('path/to/small-sample.jpg', 'rb') as f:
image_bytes = f.read()
client = genai.Client()
interaction = client.interactions.create(
model="gemini-3.8-flash",
input=[
{"type": "text", "text": "Caption this image."},
{
"type": "image",
"data": base64.b64encode(image_bytes).decode('utf-8'),
"mime_type": "image/jpeg"
}
]
)
print(interaction.output_text)
JavaScript
import { GoogleGenAI } from "@google/genai";
import * as fs from "node:fs";
const client = new GoogleGenAI({});
const base64ImageFile = fs.readFileSync("path/to/small-sample.jpg", {
encoding: "base64",
});
const interaction = await client.interactions.create({
model: "gemini-3.8-flash",
input: [
{type: "text", text: "Caption this image."},
{
type: "image",
data: base64ImageFile,
mime_type: "image/jpeg"
}
]
});
console.log(interaction.output_text);
Java
import com.google.genai.Client;
import com.google.genai.gaos.models.interactions.Content;
import com.google.genai.gaos.models.interactions.CreateModelInteraction;
import com.google.genai.gaos.models.interactions.ImageContent;
import com.google.genai.gaos.models.interactions.ImageContentMimeType;
import com.google.genai.gaos.models.interactions.Interaction;
import com.google.genai.gaos.models.interactions.InteractionsInput;
import com.google.genai.gaos.models.interactions.Model;
import com.google.genai.gaos.models.interactions.TextContent;
import com.google.genai.gaos.models.operations.CreateInteractionRequestBody;
import java.nio.file.Files;
import java.nio.file.Paths;
import java.util.Arrays;
import java.util.Base64;
import java.util.List;
byte[] imageBytes = Files.readAllBytes(Paths.get("path/to/small-sample.jpg"));
String base64Image = Base64.getEncoder().encodeToString(imageBytes);
Client client = new Client();
Content textContent = TextContent.builder().text("Caption this image.").build();
Content imageContent =
ImageContent.builder()
.data(base64Image)
.mimeType(ImageContentMimeType.IMAGE_JPEG)
.build();
List<Content> contents = Arrays.asList(textContent, imageContent);
CreateModelInteraction params =
CreateModelInteraction.builder()
.model(Model.of("gemini-3.8-flash"))
.input(InteractionsInput.ofContent(contents))
.build();
Interaction interaction =
client.interactions.create(CreateInteractionRequestBody.of(params)).interaction().get();
System.out.println(interaction.outputText().orElse(""));
Go
package main
import (
"context"
"encoding/base64"
"fmt"
"log"
"os"
"google.golang.org/genai"
"google.golang.org/genai/interactions/models/interactions"
"google.golang.org/genai/interactions/models/operations"
)
func main() {
ctx := context.Background()
client, err := genai.NewClient(ctx, nil)
if err != nil {
log.Fatal(err)
}
imageBytes, err := os.ReadFile("path/to/small-sample.jpg")
if err != nil {
log.Fatal(err)
}
base64Image := base64.StdEncoding.EncodeToString(imageBytes)
res, err := client.Interactions.Create(ctx, operations.CreateInteractionRequest{
Body: operations.NewCreateInteractionRequestBody(interactions.CreateModelInteraction{
Model: interactions.Model("gemini-3.8-flash"),
Input: interactions.NewInteractionsInput([]interactions.Content{
interactions.NewContent(interactions.TextContent{
Text: "Caption this image.",
}),
interactions.NewContent(interactions.ImageContent{
Data: genai.Ptr(base64Image),
MimeType: interactions.ImageContentMimeTypeImageJpeg.ToPointer(),
}),
}),
}),
})
if err != nil {
log.Fatal(err)
}
if res.Interaction.OutputText != nil {
fmt.Println(*res.Interaction.OutputText)
}
}
REST
IMG_PATH="/path/to/your/image1.jpg"
if [[ "$(base64 --version 2>&1)" = *"FreeBSD"* ]]; then
B64FLAGS="--input"
else
B64FLAGS="-w0"
fi
curl -X POST "https://generativelanguage.googleapis.com/v1beta/interactions" \
-H "x-goog-api-key: $GEMINI_API_KEY" \
-H 'Content-Type: application/json' \
-d '{
"model": "gemini-3.8-flash",
"input": [
{"type": "text", "text": "Caption this image."},
{
"type": "image",
"data": "'"$(base64 $B64FLAGS $IMG_PATH)"'",
"mime_type": "image/jpeg"
}
]
}'
참고: 인라인 이미지 데이터는 전체 요청 크기(텍스트 프롬프트, 시스템 지침, 인라인 바이트)를 20MB로 제한해요. 더 큰 요청은 File API로 이미지 파일을 업로드하세요.
File API로 이미지 업로드
큰 파일이거나 같은 이미지 파일을 반복해서 사용해야 할 때는 Files API를 사용하세요. Files API 가이드를 참고하세요.
Python
from google import genai
client = genai.Client()
my_file = client.files.upload(file="path/to/sample.jpg")
interaction = client.interactions.create(
model="gemini-3.8-flash",
input=[
{"type": "text", "text": "Caption this image."},
{
"type": "image",
"uri": my_file.uri,
"mime_type": my_file.mime_type
}
]
)
print(interaction.output_text)
JavaScript
import { GoogleGenAI } from "@google/genai";
const client = new GoogleGenAI({});
const myfile = await client.files.upload({
file: "path/to/sample.jpg",
config: { mimeType: "image/jpeg" },
});
const interaction = await client.interactions.create({
model: "gemini-3.8-flash",
input: [
{type: "text", text: "Caption this image."},
{
type: "image",
uri: myfile.uri,
mime_type: myfile.mimeType
}
]
});
console.log(interaction.output_text);
Java
import com.google.genai.Client;
import com.google.genai.gaos.models.interactions.Content;
import com.google.genai.gaos.models.interactions.CreateModelInteraction;
import com.google.genai.gaos.models.interactions.ImageContent;
import com.google.genai.gaos.models.interactions.ImageContentMimeType;
import com.google.genai.gaos.models.interactions.Interaction;
import com.google.genai.gaos.models.interactions.InteractionsInput;
import com.google.genai.gaos.models.interactions.Model;
import com.google.genai.gaos.models.interactions.TextContent;
import com.google.genai.gaos.models.operations.CreateInteractionRequestBody;
import com.google.genai.types.File;
import com.google.genai.types.UploadFileConfig;
import java.util.Arrays;
import java.util.List;
Client client = new Client();
File myFile =
client.files.upload(
new java.io.File("path/to/sample.jpg"),
UploadFileConfig.builder().mimeType("image/jpeg").build());
Content textContent = TextContent.builder().text("Caption this image.").build();
Content imageContent =
ImageContent.builder()
.uri(myFile.uri().orElse(""))
.mimeType(ImageContentMimeType.of(myFile.mimeType().orElse("image/jpeg")))
.build();
List<Content> contents = Arrays.asList(textContent, imageContent);
CreateModelInteraction params =
CreateModelInteraction.builder()
.model(Model.of("gemini-3.8-flash"))
.input(InteractionsInput.ofContent(contents))
.build();
Interaction interaction =
client.interactions.create(CreateInteractionRequestBody.of(params)).interaction().get();
System.out.println(interaction.outputText().orElse(""));
Go
package main
import (
"context"
"fmt"
"log"
"google.golang.org/genai"
"google.golang.org/genai/interactions/models/interactions"
"google.golang.org/genai/interactions/models/operations"
)
func main() {
ctx := context.Background()
client, err := genai.NewClient(ctx, nil)
if err != nil {
log.Fatal(err)
}
myFile, err := client.Files.UploadFromPath(ctx, "path/to/sample.jpg", &genai.UploadFileConfig{
MIMEType: "image/jpeg",
})
if err != nil {
log.Fatal(err)
}
res, err := client.Interactions.Create(ctx, operations.CreateInteractionRequest{
Body: operations.NewCreateInteractionRequestBody(interactions.CreateModelInteraction{
Model: interactions.Model("gemini-3.8-flash"),
Input: interactions.NewInteractionsInput([]interactions.Content{
interactions.NewContent(interactions.TextContent{
Text: "Caption this image.",
}),
interactions.NewContent(interactions.ImageContent{
URI: genai.Ptr(myFile.URI),
MimeType: interactions.ImageContentMimeType(myFile.MIMEType).ToPointer(),
}),
}),
}),
})
if err != nil {
log.Fatal(err)
}
if res.Interaction.OutputText != nil {
fmt.Println(*res.Interaction.OutputText)
}
}
REST
# First upload the file (see Files API guide for details)
# Then use the file URI in the request:
curl -X POST "https://generativelanguage.googleapis.com/v1beta/interactions" \
-H "x-goog-api-key: $GEMINI_API_KEY" \
-H 'Content-Type: application/json' \
-d '{
"model": "gemini-3.8-flash",
"input": [
{"type": "text", "text": "Caption this image."},
{
"type": "image",
"uri": "YOUR_FILE_URI",
"mime_type": "image/jpeg"
}
]
}'
여러 이미지로 프롬프팅
input 배열에 여러 이미지 객체를 포함하면 단일 프롬프트에서 여러 이미지를 제공할 수 있어요.
Python
from google import genai
client = genai.Client()
interaction = client.interactions.create(
model="gemini-3.8-flash",
input=[
{"type": "text", "text": "What is different between these two images?"},
{
"type": "image",
"uri": "https://example.com/image1.jpg",
"mime_type": "image/jpeg"
},
{
"type": "image",
"uri": "https://example.com/image2.jpg",
"mime_type": "image/jpeg"
}
]
)
print(interaction.output_text)
JavaScript
import { GoogleGenAI } from "@google/genai";
const client = new GoogleGenAI({});
const interaction = await client.interactions.create({
model: "gemini-3.8-flash",
input: [
{type: "text", text: "What is different between these two images?"},
{
type: "image",
uri: "https://example.com/image1.jpg",
mime_type: "image/jpeg"
},
{
type: "image",
uri: "https://example.com/image2.jpg",
mime_type: "image/jpeg"
}
]
});
console.log(interaction.output_text);
Java
import com.google.genai.Client;
import com.google.genai.gaos.models.interactions.Content;
import com.google.genai.gaos.models.interactions.CreateModelInteraction;
import com.google.genai.gaos.models.interactions.ImageContent;
import com.google.genai.gaos.models.interactions.ImageContentMimeType;
import com.google.genai.gaos.models.interactions.Interaction;
import com.google.genai.gaos.models.interactions.InteractionsInput;
import com.google.genai.gaos.models.interactions.Model;
import com.google.genai.gaos.models.interactions.TextContent;
import com.google.genai.gaos.models.operations.CreateInteractionRequestBody;
import java.util.Arrays;
import java.util.List;
Client client = new Client();
Content textContent =
TextContent.builder().text("What is different between these two images?").build();
Content image1 =
ImageContent.builder()
.uri("https://example.com/image1.jpg")
.mimeType(ImageContentMimeType.IMAGE_JPEG)
.build();
Content image2 =
ImageContent.builder()
.uri("https://example.com/image2.jpg")
.mimeType(ImageContentMimeType.IMAGE_JPEG)
.build();
List<Content> contents = Arrays.asList(textContent, image1, image2);
CreateModelInteraction params =
CreateModelInteraction.builder()
.model(Model.of("gemini-3.8-flash"))
.input(InteractionsInput.ofContent(contents))
.build();
Interaction interaction =
client.interactions.create(CreateInteractionRequestBody.of(params)).interaction().get();
System.out.println(interaction.outputText().orElse(""));
Go
package main
import (
"context"
"fmt"
"log"
"google.golang.org/genai"
"google.golang.org/genai/interactions/models/interactions"
"google.golang.org/genai/interactions/models/operations"
)
func main() {
ctx := context.Background()
client, err := genai.NewClient(ctx, nil)
if err != nil {
log.Fatal(err)
}
res, err := client.Interactions.Create(ctx, operations.CreateInteractionRequest{
Body: operations.NewCreateInteractionRequestBody(interactions.CreateModelInteraction{
Model: interactions.Model("gemini-3.8-flash"),
Input: interactions.NewInteractionsInput([]interactions.Content{
interactions.NewContent(interactions.TextContent{
Text: "What is different between these two images?",
}),
interactions.NewContent(interactions.ImageContent{
URI: genai.Ptr("https://example.com/image1.jpg"),
MimeType: interactions.ImageContentMimeTypeImageJpeg.ToPointer(),
}),
interactions.NewContent(interactions.ImageContent{
URI: genai.Ptr("https://example.com/image2.jpg"),
MimeType: interactions.ImageContentMimeTypeImageJpeg.ToPointer(),
}),
}),
}),
})
if err != nil {
log.Fatal(err)
}
if res.Interaction.OutputText != nil {
fmt.Println(*res.Interaction.OutputText)
}
}
REST
curl -X POST "https://generativelanguage.googleapis.com/v1beta/interactions" \
-H "x-goog-api-key: $GEMINI_API_KEY" \
-H 'Content-Type: application/json' \
-d '{
"model": "gemini-3.8-flash",
"input": [
{"type": "text", "text": "What is different between these two images?"},
{
"type": "image",
"uri": "https://example.com/image1.jpg",
"mime_type": "image/jpeg"
},
{
"type": "image",
"uri": "https://example.com/image2.jpg",
"mime_type": "image/jpeg"
}
]
}'
객체 감지 (Object detection)
모델은 이미지에서 객체를 감지하고 경계 상자(bounding box) 좌표를 얻도록 훈련돼요. 이미지 크기에 상대적인 좌표는 [0, 1000] 범위로 크기 조정돼요. 원본 이미지 크기에 맞춰 이 좌표를 다시 스케일(descale)해야 해요.
참고: TypeScript/JavaScript에서 구조화된 출력 스키마 검증에 Zod를 사용하려면 zod 패키지를 설치해야 해요(npm install zod).
Python
from google import genai
from pydantic import BaseModel, Field
from typing import List
import json
client = genai.Client()
prompt = "Detect the all of the prominent items in the image. The box_2d should be [ymin, xmin, ymax, xmax] normalized to 0-1000."
class BoundingBox(BaseModel):
box_2d: List[int] = Field(description="The 2D bounding box of the item as [ymin, xmin, ymax, xmax] normalized to 0-1000.")
mask: List[List[int]] = Field(description="The segmentation mask of the item as a polygon of [x,y] coordinates, normalized to 0-1000.")
label: str = Field(description="A descriptive label for the item.")
class BoundingBoxes(BaseModel):
boxes: List[BoundingBox]
interaction = client.interactions.create(
model="gemini-3.8-flash",
input=[
{"type": "text", "text": prompt},
{
"type": "image",
"uri": "https://example.com/image.png",
"mime_type": "image/png"
}
],
response_format={
"type": "text",
"mime_type": "application/json",
"schema": BoundingBoxes.model_json_schema()
}
)
bounding_boxes = BoundingBoxes.model_validate_json(interaction.output_text)
print(bounding_boxes)
JavaScript
import { GoogleGenAI } from "@google/genai";
import * as z from "zod";
const client = new GoogleGenAI({});
const prompt = "Detect the all of the prominent items in the image. The box_2d should be [ymin, xmin, ymax, xmax] normalized to 0-1000.";
const boundingBoxesSchema = z.object({
boxes: z.array(z.object({
box_2d: z.array(z.number()),
mask: z.array(z.array(z.number())),
label: z.string()
}))
});
const interaction = await client.interactions.create({
model: "gemini-3.8-flash",
input: [
{ type: "text", text: prompt },
{
type: "image",
uri: "https://example.com/image.png",
mime_type: "image/png"
}
],
response_format: {
type: 'text',
mime_type: 'application/json',
schema: z.toJSONSchema(boundingBoxesSchema)
},
});
const result = boundingBoxesSchema.parse(JSON.parse(interaction.output_text));
console.log(result);
Java
import com.google.genai.Client;
import com.google.genai.gaos.models.interactions.Content;
import com.google.genai.gaos.models.interactions.CreateModelInteraction;
import com.google.genai.gaos.models.interactions.CreateModelInteractionResponseFormat;
import com.google.genai.gaos.models.interactions.ImageContent;
import com.google.genai.gaos.models.interactions.ImageContentMimeType;
import com.google.genai.gaos.models.interactions.Interaction;
import com.google.genai.gaos.models.interactions.InteractionsInput;
import com.google.genai.gaos.models.interactions.Model;
import com.google.genai.gaos.models.interactions.ResponseFormat;
import com.google.genai.gaos.models.interactions.TextContent;
import com.google.genai.gaos.models.interactions.TextResponseFormat;
import com.google.genai.gaos.models.interactions.TextResponseFormatMimeType;
import com.google.genai.gaos.models.operations.CreateInteractionRequestBody;
import java.util.Arrays;
import java.util.List;
import java.util.Map;
Client client = new Client();
String prompt =
"Detect the all of the prominent items in the image. The box_2d should be [ymin, xmin, ymax, xmax] normalized to 0-1000.";
Map<String, Object> boundingBoxSchema =
Map.of(
"type", "object",
"properties",
Map.of(
"box_2d",
Map.of(
"type", "array",
"items", Map.of("type", "integer"),
"description",
"The 2D bounding box of the item as [ymin, xmin, ymax, xmax] normalized to 0-1000."),
"mask",
Map.of(
"type", "array",
"items", Map.of("type", "array", "items", Map.of("type", "integer")),
"description",
"The segmentation mask of the item as a polygon of [x,y] coordinates, normalized to 0-1000."),
"label",
Map.of("type", "string", "description", "A descriptive label for the item.")),
"required", List.of("box_2d", "mask", "label"));
Map<String, Object> boundingBoxesSchema =
Map.of(
"type", "object",
"properties", Map.of("boxes", Map.of("type", "array", "items", boundingBoxSchema)),
"required", List.of("boxes"));
CreateModelInteractionResponseFormat format =
CreateModelInteractionResponseFormat.of(
ResponseFormat.of(
TextResponseFormat.builder()
.mimeType(TextResponseFormatMimeType.APPLICATION_JSON)
.schema(boundingBoxesSchema)
.build()));
Content textContent = TextContent.builder().text(prompt).build();
Content imageContent =
ImageContent.builder()
.uri("https://example.com/image.png")
.mimeType(ImageContentMimeType.IMAGE_PNG)
.build();
CreateModelInteraction params =
CreateModelInteraction.builder()
.model(Model.of("gemini-3.8-flash"))
.input(InteractionsInput.ofContent(Arrays.asList(textContent, imageContent)))
.responseFormat(format)
.build();
Interaction interaction =
client.interactions.create(CreateInteractionRequestBody.of(params)).interaction().get();
System.out.println(interaction.outputText().orElse(""));
Go
package main
import (
"context"
"fmt"
"log"
"google.golang.org/genai"
"google.golang.org/genai/interactions/models/interactions"
"google.golang.org/genai/interactions/models/operations"
)
func main() {
ctx := context.Background()
client, err := genai.NewClient(ctx, nil)
if err != nil {
log.Fatal(err)
}
prompt := "Detect the all of the prominent items in the image. The box_2d should be [ymin, xmin, ymax, xmax] normalized to 0-1000."
boundingBoxSchema := map[string]any{
"type": "object",
"properties": map[string]any{
"box_2d": map[string]any{
"type": "array",
"items": map[string]any{"type": "integer"},
"description": "The 2D bounding box of the item as [ymin, xmin, ymax, xmax] normalized to 0-1000.",
},
"mask": map[string]any{
"type": "array",
"items": map[string]any{"type": "array", "items": map[string]any{"type": "integer"}},
"description": "The segmentation mask of the item as a polygon of [x,y] coordinates, normalized to 0-1000.",
},
"label": map[string]any{
"type": "string",
"description": "A descriptive label for the item.",
},
},
"required": []string{"box_2d", "mask", "label"},
}
boundingBoxesSchema := map[string]any{
"type": "object",
"properties": map[string]any{
"boxes": map[string]any{
"type": "array",
"items": boundingBoxSchema,
},
},
"required": []string{"boxes"},
}
format := interactions.NewCreateModelInteractionResponseFormat(
interactions.NewResponseFormat(interactions.TextResponseFormat{
MimeType: interactions.TextResponseFormatMimeTypeApplicationJSON.ToPointer(),
Schema: boundingBoxesSchema,
}),
)
res, err := client.Interactions.Create(ctx, operations.CreateInteractionRequest{
Body: operations.NewCreateInteractionRequestBody(interactions.CreateModelInteraction{
Model: interactions.Model("gemini-3.8-flash"),
Input: interactions.NewInteractionsInput([]interactions.Content{
interactions.NewContent(interactions.TextContent{
Text: prompt,
}),
interactions.NewContent(interactions.ImageContent{
URI: genai.Ptr("https://example.com/image.png"),
MimeType: interactions.ImageContentMimeTypeImagePng.ToPointer(),
}),
}),
ResponseFormat: genai.Ptr(format),
}),
})
if err != nil {
log.Fatal(err)
}
if res.Interaction.OutputText != nil {
fmt.Println(*res.Interaction.OutputText)
}
}
REST
curl -X POST "https://generativelanguage.googleapis.com/v1beta/interactions" \
-H "x-goog-api-key: $GEMINI_API_KEY" \
-H 'Content-Type: application/json' \
-d '{
"model": "gemini-3.8-flash",
"input": [
{"type": "text", "text": "Detect the all of the prominent items in the image. The box_2d should be [ymin, xmin, ymax, xmax] normalized to 0-1000."},
{
"type": "image",
"uri": "https://example.com/image.png",
"mime_type": "image/png"
}
],
"response_format": {
"type": "text",
"mime_type": "application/json",
"schema": {
"type": "object",
"properties": {
"boxes": {
"type": "array",
"items": {
"type": "object",
"properties": {
"box_2d": { "type": "array", "items": { "type": "integer" } },
"mask": { "type": "array", "items": { "type": "array", "items": { "type": "integer" } } },
"label": { "type": "string" }
},
"required": ["box_2d", "mask", "label"]
}
}
},
"required": ["boxes"]
}
}
}'
참고: 모델은 "이 이미지에서 모든 초록색 객체의 경계 상자를 보여줘" 같은 커스텀 지침에 따라 경계 상자를 생성하는 것도 지원해요.
더 많은 예시는 Gemini Cookbook에서 볼 수 있어요.
분할 (Segmentation)
Gemini 모델은 항목을 감지할 뿐만 아니라 분할하고 윤곽 마스크(contour mask)도 제공해요.
모델은 JSON 리스트를 예측하는데, 각 항목이 분할 마스크를 나타내요. 각 항목은 [ymin, xmin, ymax, xmax] 형식의 경계 상자("box_2d")를 가지며 0~1000 사이의 정규화된 좌표를 사용하고, 객체를 식별하는 라벨("label"), 그리고 경계 상자 안의 분할 마스크를 0-1000으로 정규화된 [x, y] 좌표의 폴리곤으로 제공해요.
참고: 더 나은 결과를 위해 thinking 레벨을 "minimal"로 설정해 thinking을 비활성화하세요.
Python
from google import genai
from pydantic import BaseModel, Field
from typing import List
import json
client = genai.Client()
prompt = """
Give the segmentation masks for the wooden and glass items.
Output a JSON list of segmentation masks where each entry contains the 2D
bounding box in the key "box_2d", the segmentation mask in key "mask", and
the text label in the key "label". Use descriptive labels.
"""
class BoundingBox(BaseModel):
box_2d: List[int] = Field(description="The 2D bounding box of the item as [ymin, xmin, ymax, xmax] normalized to 0-1000.")
mask: List[List[int]] = Field(description="The segmentation mask of the item as a polygon of [x,y] coordinates, normalized to 0-1000.")
label: str = Field(description="A descriptive label for the item.")
class BoundingBoxes(BaseModel):
boxes: List[BoundingBox]
interaction = client.interactions.create(
model="gemini-3.8-flash",
input=[
{"type": "text", "text": prompt},
{
"type": "image",
"uri": "https://example.com/image.png",
"mime_type": "image/png"
}
],
response_format={
"type": "text",
"mime_type": "application/json",
"schema": BoundingBoxes.model_json_schema()
},
generation_config={
"thinking_level": "minimal"
}
)
items = BoundingBoxes.model_validate_json(interaction.output_text)
print("Segmentation results:", items)
JavaScript
import { GoogleGenAI } from "@google/genai";
import * as z from "zod";
const client = new GoogleGenAI({});
const prompt = `
Give the segmentation masks for the wooden and glass items.
Output a JSON list of segmentation masks where each entry contains the 2D
bounding box in the key "box_2d", the segmentation mask in key "mask", and
the text label in the key "label". Use descriptive labels.
`;
const boundingBoxesSchema = z.object({
boxes: z.array(z.object({
box_2d: z.array(z.number()),
mask: z.array(z.array(z.number())),
label: z.string()
}))
});
const interaction = await client.interactions.create({
model: "gemini-3.8-flash",
input: [
{ type: "text", text: prompt },
{
type: "image",
uri: "https://example.com/image.png",
mime_type: "image/png"
}
],
response_format: {
type: 'text',
mime_type: 'application/json',
schema: z.toJSONSchema(boundingBoxesSchema)
},
generation_config: {
thinking_level: "minimal"
}
});
const result = boundingBoxesSchema.parse(JSON.parse(interaction.output_text));
console.log(result);
Java
import com.google.genai.Client;
import com.google.genai.gaos.models.interactions.Content;
import com.google.genai.gaos.models.interactions.CreateModelInteraction;
import com.google.genai.gaos.models.interactions.CreateModelInteractionResponseFormat;
import com.google.genai.gaos.models.interactions.GenerationConfig;
import com.google.genai.gaos.models.interactions.ImageContent;
import com.google.genai.gaos.models.interactions.ImageContentMimeType;
import com.google.genai.gaos.models.interactions.Interaction;
import com.google.genai.gaos.models.interactions.InteractionsInput;
import com.google.genai.gaos.models.interactions.Model;
import com.google.genai.gaos.models.interactions.ResponseFormat;
import com.google.genai.gaos.models.interactions.TextContent;
import com.google.genai.gaos.models.interactions.TextResponseFormat;
import com.google.genai.gaos.models.interactions.TextResponseFormatMimeType;
import com.google.genai.gaos.models.interactions.ThinkingLevel;
import com.google.genai.gaos.models.operations.CreateInteractionRequestBody;
import java.util.Arrays;
import java.util.List;
import java.util.Map;
Client client = new Client();
String prompt =
"Give the segmentation masks for the wooden and glass items.\n"
+ "Output a JSON list of segmentation masks where each entry contains the 2D\n"
+ "bounding box in the key \"box_2d\", the segmentation mask in key \"mask\", and\n"
+ "the text label in the key \"label\". Use descriptive labels.";
Map<String, Object> boundingBoxSchema =
Map.of(
"type", "object",
"properties",
Map.of(
"box_2d",
Map.of(
"type", "array",
"items", Map.of("type", "integer"),
"description",
"The 2D bounding box of the item as [ymin, xmin, ymax, xmax] normalized to 0-1000."),
"mask",
Map.of(
"type", "array",
"items", Map.of("type", "array", "items", Map.of("type", "integer")),
"description",
"The segmentation mask of the item as a polygon of [x,y] coordinates, normalized to 0-1000."),
"label",
Map.of("type", "string", "description", "A descriptive label for the item.")),
"required", List.of("box_2d", "mask", "label"));
Map<String, Object> boundingBoxesSchema =
Map.of(
"type", "object",
"properties", Map.of("boxes", Map.of("type", "array", "items", boundingBoxSchema)),
"required", List.of("boxes"));
CreateModelInteractionResponseFormat format =
CreateModelInteractionResponseFormat.of(
ResponseFormat.of(
TextResponseFormat.builder()
.mimeType(TextResponseFormatMimeType.APPLICATION_JSON)
.schema(boundingBoxesSchema)
.build()));
Content textContent = TextContent.builder().text(prompt).build();
Content imageContent =
ImageContent.builder()
.uri("https://example.com/image.png")
.mimeType(ImageContentMimeType.IMAGE_PNG)
.build();
CreateModelInteraction params =
CreateModelInteraction.builder()
.model(Model.of("gemini-3.8-flash"))
.input(InteractionsInput.ofContent(Arrays.asList(textContent, imageContent)))
.responseFormat(format)
.generationConfig(GenerationConfig.builder().thinkingLevel(ThinkingLevel.MINIMAL).build())
.build();
Interaction interaction =
client.interactions.create(CreateInteractionRequestBody.of(params)).interaction().get();
System.out.println("Segmentation results: " + interaction.outputText().orElse(""));
Go
package main
import (
"context"
"fmt"
"log"
"google.golang.org/genai"
"google.golang.org/genai/interactions/models/interactions"
"google.golang.org/genai/interactions/models/operations"
)
func main() {
ctx := context.Background()
client, err := genai.NewClient(ctx, nil)
if err != nil {
log.Fatal(err)
}
prompt := "Give the segmentation masks for the wooden and glass items.\n" +
"Output a JSON list of segmentation masks where each entry contains the 2D\n" +
"bounding box in the key \"box_2d\", the segmentation mask in key \"mask\", and\n" +
"the text label in the key \"label\". Use descriptive labels."
boundingBoxSchema := map[string]any{
"type": "object",
"properties": map[string]any{
"box_2d": map[string]any{
"type": "array",
"items": map[string]any{"type": "integer"},
"description": "The 2D bounding box of the item as [ymin, xmin, ymax, xmax] normalized to 0-1000.",
},
"mask": map[string]any{
"type": "array",
"items": map[string]any{"type": "array", "items": map[string]any{"type": "integer"}},
"description": "The segmentation mask of the item as a polygon of [x,y] coordinates, normalized to 0-1000.",
},
"label": map[string]any{
"type": "string",
"description": "A descriptive label for the item.",
},
},
"required": []string{"box_2d", "mask", "label"},
}
boundingBoxesSchema := map[string]any{
"type": "object",
"properties": map[string]any{
"boxes": map[string]any{
"type": "array",
"items": boundingBoxSchema,
},
},
"required": []string{"boxes"},
}
format := interactions.NewCreateModelInteractionResponseFormat(
interactions.NewResponseFormat(interactions.TextResponseFormat{
MimeType: interactions.TextResponseFormatMimeTypeApplicationJSON.ToPointer(),
Schema: boundingBoxesSchema,
}),
)
res, err := client.Interactions.Create(ctx, operations.CreateInteractionRequest{
Body: operations.NewCreateInteractionRequestBody(interactions.CreateModelInteraction{
Model: interactions.Model("gemini-3.8-flash"),
Input: interactions.NewInteractionsInput([]interactions.Content{
interactions.NewContent(interactions.TextContent{
Text: prompt,
}),
interactions.NewContent(interactions.ImageContent{
URI: genai.Ptr("https://example.com/image.png"),
MimeType: interactions.ImageContentMimeTypeImagePng.ToPointer(),
}),
}),
ResponseFormat: genai.Ptr(format),
GenerationConfig: &interactions.GenerationConfig{
ThinkingLevel: interactions.ThinkingLevelMinimal.ToPointer(),
},
}),
})
if err != nil {
log.Fatal(err)
}
if res.Interaction.OutputText != nil {
fmt.Println("Segmentation results: " + *res.Interaction.OutputText)
}
}
REST
curl -X POST "https://generativelanguage.googleapis.com/v1beta/interactions" \
-H "x-goog-api-key: $GEMINI_API_KEY" \
-H 'Content-Type: application/json' \
-d '{
"model": "gemini-3.8-flash",
"input": [
{"type": "text", "text": "Give the segmentation masks for the wooden and glass items.\nOutput a JSON list of segmentation masks where each entry contains the 2D\nbounding box in the key \"box_2d\", the segmentation mask in key \"mask\", and\nthe text label in the key \"label\". Use descriptive labels."},
{
"type": "image",
"uri": "https://example.com/image.png",
"mime_type": "image/png"
}
],
"response_format": {
"type": "text",
"mime_type": "application/json",
"schema": {
"type": "object",
"properties": {
"boxes": {
"type": "array",
"items": {
"type": "object",
"properties": {
"box_2d": { "type": "array", "items": { "type": "integer" } },
"mask": { "type": "array", "items": { "type": "array", "items": { "type": "integer" } } },
"label": { "type": "string" }
},
"required": ["box_2d", "mask", "label"]
}
}
},
"required": ["boxes"]
}
},
"generation_config": {
"thinking_level": "minimal"
}
}'
객체와 분할 마스크가 있는 분할 출력 예시
지원되는 이미지 형식
Gemini는 다음 이미지 형식 MIME 타입을 지원해요.
- PNG -
image/png - JPEG -
image/jpeg - WEBP -
image/webp - HEIC -
image/heic - HEIF -
image/heif
다른 파일 입력 방법에 대해 알아보려면 File input methods 가이드를 참고하세요.
기능 (Capabilities)
모든 Gemini 모델 버전은 멀티모달이며 이미지 캡션 생성, 시각 질의응답, 이미지 분류, 객체 감지, 분할을 포함한(물론 이에 국한되지 않는) 광범위한 이미지 처리·컴퓨터 비전 작업에 활용될 수 있어요. Gemini는 품질과 성능 요구 사항에 따라 특수 ML 모델을 사용할 필요를 줄여줄 수 있어요.
최신 모델 버전은 일반 능력에 더해 향상된 객체 감지와 분할 같은 특수 작업의 정확도를 개선하도록 특별히 훈련돼요.
제한 사항과 주요 기술 정보
파일 한도
Gemini 모델은 요청당 최대 3,600개의 이미지 파일을 지원해요.
토큰 계산
(원문 기준) 두 차원에 해당하는 경우 258 토큰.
다음 단계 (What's next)
이 가이드는 이미지 파일을 업로드하고 이미지 입력에서 텍스트 출력을 생성하는 방법을 보여줘요. 더 알아보려면 다음 리소스를 참고하세요.
- Files API: Gemini와 함께 사용할 파일 업로드·관리에 대해 더 알아보세요.
- System instructions: 시스템 지침으로 특정 요구 사항과 사용 사례에 맞게 모델의 동작을 조정할 수 있어요.
- File prompting strategies: Gemini API는 텍스트, 이미지, 오디오, 비디오 데이터로 프롬프팅하는 것을 지원해요. 이를 멀티모달 프롬프팅이라고도 해요.
- Safety guidance: 생성형 AI 모델은 때때로 부정확하거나 편향되거나 공격적인 출력 같은 예상치 못한 출력을 만들어내기도 해요. 이러한 출력으로 인한 피해 위험을 줄이기 위해 후처리와 사람의 평가가 필수적이에요.
더 알아보기 (Learn more)
- Files API 문서로 이미지 업로드 방법을 익혀 보세요.
- Image generation 가이드로 이미지 생성·편집을 배워 보세요.
- Gemini Cookbook에서 더 많은 예제를 살펴보세요.