Gemini thinking

Gemini thinking

Gemini 3 및 2.5 시리즈 모델은 내부 "thinking 프로세스"를 사용해 추론과 다단계 계획 능력을 크게 개선하며, 코딩, 고급 수학, 데이터 분석 같은 복잡한 작업에 매우 효과적으로 만듭니다. 이 가이드는 Gemini API를 사용해 Gemini의 thinking 기능을 다루는 방법을 보여줘요.

출처: 원문

본문

thinking으로 콘텐츠 생성

thinking 모델로 요청을 시작하는 것은 다른 콘텐츠 생성 요청과 비슷해요. 핵심 차이는 다음 텍스트 생성 예시에서처럼 model 필드에 thinking을 지원하는 모델 중 하나를 지정하는 거예요:

Python

from google import genai

client = genai.Client()
prompt = "Explain the concept of Occam's Razor and provide a simple, everyday example."
response = client.models.generate_content(
    model="gemini-3.8-flash",
    contents=prompt
)

print(response.text)

JavaScript

import { GoogleGenAI } from "@google/genai";

const ai = new GoogleGenAI({});

async function main() {
  const prompt = "Explain the concept of Occam's Razor and provide a simple, everyday example.";

  const response = await ai.models.generateContent({
    model: "gemini-3.8-flash",
    contents: prompt,
  });

  console.log(response.text);
}

main();

Go

package main

import (
  "context"
  "fmt"
  "log"
  "os"
  "google.golang.org/genai"
)

func main() {
  ctx := context.Background()
  client, err := genai.NewClient(ctx, nil)
  if err != nil {
      log.Fatal(err)
  }

  prompt := "Explain the concept of Occam's Razor and provide a simple, everyday example."
  model := "gemini-3.8-flash"

  resp, _ := client.Models.GenerateContent(ctx, model, genai.Text(prompt), nil)

  fmt.Println(resp.Text())
}

REST

curl "https://generativelanguage.googleapis.com/v1beta/models/gemini-3.8-flash:generateContent" \
 -H "x-goog-api-key: *** \
 -H 'Content-Type: application/json' \
 -X POST \
 -d '{
   "contents": [
     {
       "parts": [
         {
           "text": "Explain the concept of Occam''s Razor and provide a simple, everyday example."
         }
       ]
     }
   ]
 }'

생각 요약(Thought summaries)

생각 요약은 모델의 원시 생각을 요약한 버전이며 모델의 내부 추론 과정에 대한 통찰을 제공해요. thinking 수준과 예산은 생각 요약이 아니라 모델의 원시 생각에 적용된다는 점에 유의하세요.

요청 구성에서 includeThoughts를 true로 설정해 생각 요약을 활성화할 수 있어요. 그런 다음 response 매개변수 parts를 반복하고 thought 불리언을 확인해 요약에 접근할 수 있어요.

다음은 스트리밍 없이 생각 요약을 활성화하고 검색하는 예시예요. 응답과 함께 단일 최종 생각 요약을 반환해요:

Python

from google import genai
from google.genai import types

client = genai.Client()
prompt = "What is the sum of the first 50 prime numbers?"
response = client.models.generate_content(
  model="gemini-3.8-flash",
  contents=prompt,
  config=types.GenerateContentConfig(
    thinking_config=types.ThinkingConfig(
      include_thoughts=True
    )
  )
)

for part in response.candidates[0].content.parts:
  if not part.text:
    continue
  if part.thought:
    print("Thought summary:")
    print(part.text)
    print()
  else:
    print("Answer:")
    print(part.text)
    print()

JavaScript

import { GoogleGenAI } from "@google/genai";

const ai = new GoogleGenAI({});

async function main() {
  const response = await ai.models.generateContent({
    model: "gemini-3.8-flash",
    contents: "What is the sum of the first 50 prime numbers?",
    config: {
      thinkingConfig: {
        includeThoughts: true,
      },
    },
  });

  for (const part of response.candidates[0].content.parts) {
    if (!part.text) {
      continue;
    }
    else if (part.thought) {
      console.log("Thoughts summary:");
      console.log(part.text);
    }
    else {
      console.log("Answer:");
      console.log(part.text);
    }
  }
}

main();

Go

package main

import (
  "context"
  "fmt"
  "google.golang.org/genai"
  "os"
)

func main() {
  ctx := context.Background()
  client, err := genai.NewClient(ctx, nil)
  if err != nil {
      log.Fatal(err)
  }

  contents := genai.Text("What is the sum of the first 50 prime numbers?")
  model := "gemini-3.8-flash"
  resp, _ := client.Models.GenerateContent(ctx, model, contents, &genai.GenerateContentConfig{
    ThinkingConfig: &genai.ThinkingConfig{
      IncludeThoughts: true,
    },
  })

  for _, part := range resp.Candidates[0].Content.Parts {
    if part.Text != "" {
      if part.Thought {
        fmt.Println("Thoughts Summary:")
        fmt.Println(part.Text)
      } else {
        fmt.Println("Answer:")
        fmt.Println(part.Text)
      }
    }
  }
}

다음은 생성 중에 롤링되고 증분되는 요약을 반환하는 스트리밍과 함께 thinking을 사용하는 예시예요:

Python

from google import genai
from google.genai import types

client = genai.Client()

prompt = """
Alice, Bob, and Carol each live in a different house on the same street: red, green, and blue.
The person who lives in the red house owns a cat.
Bob does not live in the green house.
Carol owns a dog.
The green house is to the left of the red house.
Alice does not own a cat.
Who lives in each house, and what pet do they own?
"""

thoughts = ""
answer = ""

for chunk in client.models.generate_content_stream(
    model="gemini-3.8-flash",
    contents=prompt,
    config=types.GenerateContentConfig(
      thinking_config=types.ThinkingConfig(
        include_thoughts=True
      )
    )
):
  for part in chunk.candidates[0].content.parts:
    if not part.text:
      continue
    elif part.thought:
      if not thoughts:
        print("Thoughts summary:")
      print(part.text)
      thoughts += part.text
    else:
      if not answer:
        print("Answer:")
      print(part.text)
      answer += part.text

JavaScript

import { GoogleGenAI } from "@google/genai";

const ai = new GoogleGenAI({});

const prompt = `Alice, Bob, and Carol each live in a different house on the same
street: red, green, and blue. The person who lives in the red house owns a cat.
Bob does not live in the green house. Carol owns a dog. The green house is to
the left of the red house. Alice does not own a cat. Who lives in each house,
and what pet do they own?`;

let thoughts = "";
let answer = "";

async function main() {
  const response = await ai.models.generateContentStream({
    model: "gemini-3.8-flash",
    contents: prompt,
    config: {
      thinkingConfig: {
        includeThoughts: true,
      },
    },
  });

  for await (const chunk of response) {
    for (const part of chunk.candidates[0].content.parts) {
      if (!part.text) {
        continue;
      } else if (part.thought) {
        if (!thoughts) {
          console.log("Thoughts summary:");
        }
        console.log(part.text);
        thoughts = thoughts + part.text;
      } else {
        if (!answer) {
          console.log("Answer:");
        }
        console.log(part.text);
        answer = answer + part.text;
      }
    }
  }
}

await main();

Go

package main

import (
  "context"
  "fmt"
  "log"
  "os"
  "google.golang.org/genai"
)

const prompt = `
Alice, Bob, and Carol each live in a different house on the same street: red, green, and blue.
The person who lives in the red house owns a cat.
Bob does not live in the green house.
Carol owns a dog.
The green house is to the left of the red house.
Alice does not own a cat.
Who lives in each house, and what pet do they own?
`

func main() {
  ctx := context.Background()
  client, err := genai.NewClient(ctx, nil)
  if err != nil {
      log.Fatal(err)
  }

  contents := genai.Text(prompt)
  model := "gemini-3.8-flash"

  resp := client.Models.GenerateContentStream(ctx, model, contents, &genai.GenerateContentConfig{
    ThinkingConfig: &genai.ThinkingConfig{
      IncludeThoughts: true,
    },
  })

  for chunk := range resp {
    for _, part := range chunk.Candidates[0].Content.Parts {
      if len(part.Text) == 0 {
        continue
      }

      if part.Thought {
        fmt.Printf("Thought: %s\n", part.Text)
      } else {
        fmt.Printf("Answer: %s\n", part.Text)
      }
    }
  }
}

thinking 제어

Gemini 모델은 기본적으로 동적 thinking에 참여하며 사용자 요청의 복잡성에 따라 추론 노력을 자동 조정해요. 그러나 특정 지연 시간 제약이 있거나 모델이 평소보다 더 깊은 추론에 참여하도록 요구한다면 선택적으로 매개변수를 사용해 thinking 동작을 제어할 수 있어요.

thinking 수준(Gemini 3)

thinkingLevel 매개변수는 Gemini 3 모델 이상에 권장되며 추론 동작을 제어할 수 있게 해줘요.

다음 표는 각 모델 유형에 대한 thinkingLevel 설정을 자세히 설명해요:

Thinking 수준 Gemini 3.8 & 3.7 Flash Gemini 3.6 & 3.5 Flash Gemini 3.1 Pro Gemini 3.5 & 3.1 Flash-Lite Gemini 3.1 Flash-Lite Image Gemini 3 Flash Gemini Robotics ER 2 설명
minimal 미지원(오류) 지원 미지원 지원(기본값) 지원(기본값) 지원 지원 대부분의 쿼리에서 "no thinking" 설정과 일치해요. minimal은 thinking이 꺼진 것을 보장하지 않으며, 모델이 복잡한 작업에 대해 아주 최소한으로 추론할 수 있어요.
low 지원 지원 지원 지원 미지원 지원 지원 지연 시간과 비용을 최소화해요.
medium 지원(기본값) 지원(기본값) 지원 지원 미지원 지원 지원 대부분 작업에 균형 잡힌 thinking.
high 지원(동적) 지원(동적) 지원(기본값, 동적) 지원(동적) 지원(동적) 지원(기본값, 동적) 지원(기본값, 동적) 추론 깊이를 최대화해요. 모델이 첫(non-thinking) 출력 토큰에 도달하는 데 훨씬 오래 걸릴 수 있지만 출력은 더 신중하게 추론돼요.

다음 예시는 thinking 수준을 설정하는 방법을 보여줘요.

Python

from google import genai
from google.genai import types

client = genai.Client()

response = client.models.generate_content(
    model="gemini-3.8-flash",
    contents="Provide a list of 3 famous physicists and their key contributions",
    config=types.GenerateContentConfig(
        thinking_config=types.ThinkingConfig(thinking_level="low")
    ),
)

print(response.text)

JavaScript

import { GoogleGenAI, ThinkingLevel } from "@google/genai";

const ai = new GoogleGenAI({});

async function main() {
  const response = await ai.models.generateContent({
    model: "gemini-3.8-flash",
    contents: "Provide a list of 3 famous physicists and their key contributions",
    config: {
      thinkingConfig: {
        thinkingLevel: ThinkingLevel.LOW,
      },
    },
  });

  console.log(response.text);
}

main();

Go

package main

import (
  "context"
  "fmt"
  "google.golang.org/genai"
  "os"
)

func main() {
  ctx := context.Background()
  client, err := genai.NewClient(ctx, nil)
  if err != nil {
      log.Fatal(err)
  }

  thinkingLevelVal := "low"

  contents := genai.Text("Provide a list of 3 famous physicists and their key contributions")
  model := "gemini-3.8-flash"
  resp, _ := client.Models.GenerateContent(ctx, model, contents, &genai.GenerateContentConfig{
    ThinkingConfig: &genai.ThinkingConfig{
      ThinkingLevel: &thinkingLevelVal,
    },
  })

fmt.Println(resp.Text())
}

REST

curl "https://generativelanguage.googleapis.com/v1beta/models/gemini-3.8-flash:generateContent" \
-H "x-goog-api-key: *** \
-H 'Content-Type: application/json' \
-X POST \
-d '{
  "contents": [
    {
      "parts": [
        {
          "text": "Provide a list of 3 famous physicists and their key contributions"
        }
      ]
    }
  ],
  "generationConfig": {
    "thinkingConfig": {
          "thinkingLevel": "low"
    }
  }
}'

Gemini 3.1 Pro에서는 thinking을 비활성화할 수 없어요. Gemini 3 Flash와 Flash-Lite도 완전한 thinking-off를 지원하지 않아요. thinking 수준을 지정하지 않으면 Gemini는 Gemini 3 모델의 기본 thinking 수준을 사용해요(예: Gemini 3.1 Pro는 "high", Gemini 3.5 Flash는 "medium").

Gemini 2.5 시리즈 모델은 thinkingLevel을 지원하지 않아요. 대신 thinkingBudget을 사용하세요.

토큰 한도 및 max_output_tokens

max_output_tokens 생성 매개변수는 생성할 수 있는 최대 토큰 수(생각 토큰 포함)를 설정해요.

설정하면 이 매개변수는 모델이 thinking 예산(thinking_level)을 할당하는 방식을 바꾸지 않고 인프라가 강제하는 하드 컷오프로 작동해요.

모델이 추론 중 이 한도에 도달하면 finish_reason: MAX_TOKENS로 생성을 중단하고 잘리거나 빈 출력을 반환해요(생성된 thinking 토큰에 대해서는 여전히 청구). 응답을 자르지 않고 비용 또는 지연 시간을 줄이려면 작은 max_output_tokens를 설정하는 대신 thinking_level(low 또는 medium)을 낮추세요.

thinking 예산(Thinking budgets)

thinkingBudget 매개변수는 Gemini 2.5 시리즈와 함께 도입되었으며, 모델이 추론에 사용할 특정 thinking 토큰 수를 안내해요.

참고: Gemini 3 모델에는 thinkingLevel 매개변수를 사용하세요. thinkingBudget은 하위 호환성으로 허용되지만, Gemini 3 Pro에서 사용하면 예상치 못한 성능이 발생할 수 있어요.

다음은 각 모델 유형에 대한 thinkingBudget 구성 세부 사항이에요. thinkingBudget을 0으로 설정하면 thinking을 비활성화할 수 있어요. thinkingBudget을 -1로 설정하면 동적 thinking이 켜져 모델이 요청의 복잡성에 따라 예산을 조정해요.

| 모델 | 기본 설정 (Thinking budget 미설정) | 범위 | thinking 비활성화 | 동적 thinking 켜기 | |---|---|---|---|---| | 2.5 Pro | 동적 thinking | 128 ~ 32768 | N/A: thinking 비활성화 불가 | thinkingBudget = -1 (기본값) | | 2.5 Flash | 동적 thinking | 0 ~ 24576 | thinkingBudget = 0 | thinkingBudget = -1 (기본값) | | 2.5 Flash Preview | 동적 thinking | 0 ~ 24576 | thinkingBudget = 0 | thinkingBudget = -1 (기본값) | | 2.5 Flash Lite | 모델은 생각하지 않음 | 512 ~ 24576 | thinkingBudget = 0 | thinkingBudget = -1 | | 2.5 Flash Lite Preview | 모델은 생각하지 않음 | 512 ~ 24576 | thinkingBudget = 0 | thinkingBudget = -1 | | Robotics-ER 1.6 Preview | 동적 thinking | 0 ~ 24576 | thinkingBudget = 0 | thinkingBudget = -1 (기본값) | | 2.5 Flash Live Native Audio Preview (09-2025) | 동적 thinking | 0 ~ 24576 | thinkingBudget = 0 | thinkingBudget = -1 (기본값) |

Python

from google import genai
from google.genai import types

client = genai.Client()

response = client.models.generate_content(
    model="gemini-2.5-flash",
    contents="Provide a list of 3 famous physicists and their key contributions",
    config=types.GenerateContentConfig(
        thinking_config=types.ThinkingConfig(thinking_budget=1024)
        # Turn off thinking:
        # thinking_config=types.ThinkingConfig(thinking_budget=0)
        # Turn on dynamic thinking:
        # thinking_config=types.ThinkingConfig(thinking_budget=-1)
    ),
)

print(response.text)

JavaScript

import { GoogleGenAI } from "@google/genai";

const ai = new GoogleGenAI({});

async function main() {
  const response = await ai.models.generateContent({
    model: "gemini-2.5-flash",
    contents: "Provide a list of 3 famous physicists and their key contributions",
    config: {
      thinkingConfig: {
        thinkingBudget: 1024,
        // Turn off thinking:
        // thinkingBudget: 0
        // Turn on dynamic thinking:
        // thinkingBudget: -1
      },
    },
  });

  console.log(response.text);
}

main();

Go

package main

import (
  "context"
  "fmt"
  "google.golang.org/genai"
  "os"
)

func main() {
  ctx := context.Background()
  client, err := genai.NewClient(ctx, nil)
  if err != nil {
      log.Fatal(err)
  }

  thinkingBudgetVal := int32(1024)

  contents := genai.Text("Provide a list of 3 famous physicists and their key contributions")
  model := "gemini-2.5-flash"
  resp, _ := client.Models.GenerateContent(ctx, model, contents, &genai.GenerateContentConfig{
    ThinkingConfig: &genai.ThinkingConfig{
      ThinkingBudget: &thinkingBudgetVal,
      // Turn off thinking:
      // ThinkingBudget: int32(0),
      // Turn on dynamic thinking:
      // ThinkingBudget: int32(-1),
    },
  })

fmt.Println(resp.Text())
}

REST

curl "https://generativelanguage.googleapis.com/v1beta/models/gemini-2.5-flash:generateContent" \
-H "x-goog-api-key: *** \
-H 'Content-Type: application/json' \
-X POST \
-d '{
  "contents": [
    {
      "parts": [
        {
          "text": "Provide a list of 3 famous physicists and their key contributions"
        }
      ]
    }
  ],
  "generationConfig": {
    "thinkingConfig": {
          "thinkingBudget": 1024
    }
  }
}'

프롬프트에 따라 모델이 토큰 예산을 초과하거나 미만일 수 있어요.

생각 서명(Thought signatures)

중요: Google GenAI SDK는 생각 서명 반환을 자동으로 처리해줘요. 대화 기록을 수정하거나 REST API를 사용하는 경우에만 생각 서명을 수동으로 관리해야 해요.

Gemini API는 무상태이므로 모델은 모든 API 요청을 독립적으로 처리하며 다중 턴 상호작용에서 이전 턴의 생각 컨텍스트에 접근할 수 없어요.

다중 턴 상호작용에서 생각 컨텍스트를 유지할 수 있게 하기 위해 Gemini는 생각 서명(모델의 내부 생각 과정을 암호화한 표현)을 반환해요.

  • Gemini 2.5 모델은 thinking이 활성화되고 요청에 함수 호출, 특히 함수 선언이 포함될 때 생각 서명을 반환해요.
  • Gemini 3 모델은 모든 유형의 parts에 대해 생각 서명을 반환할 수 있어요. 모든 서명을 받은 대로 다시 전달하는 것을 권장하지만, 함수 호출 서명에는 필수예요. 자세한 내용은 Thought Signatures 페이지를 읽어 보세요.

함수 호출과 함께 고려할 다른 사용 제한 사항:

  • 서명은 응답의 다른 parts(예: 함수 호출 또는 텍스트 parts) 안에서 모델이 반환해요. 이후 턴에서 전체 응답을 모든 parts와 함께 모델에 다시 반환하세요.
  • 서명이 있는 parts를 서로 연결하지 마세요.
  • 서명이 있는 part를 서명이 없는 다른 part와 병합하지 마세요.

가격(Pricing)

참고: 요약은 API의 무료 및 유료 티어 모두에서 사용할 수 있어요. 생각 서명은 요청의 일부로 다시 보내면 청구되는 입력 토큰을 증가시켜요.

thinking이 켜지면 응답 가격은 출력 토큰과 thinking 토큰의 합이에요. 생성된 총 thinking 토큰 수는 thoughtsTokenCount 필드에서 얻을 수 있어요.

Python

# ...
print("Thoughts tokens:", response.usage_metadata.thoughts_token_count)
print("Output tokens:", response.usage_metadata.candidates_token_count)

JavaScript

// ...
console.log(`Thoughts tokens: ${response.usageMetadata.thoughtsTokenCount}`);
console.log(`Output tokens: ${response.usageMetadata.candidatesTokenCount}`);

Go

// ...
fmt.Println("Thoughts tokens:", response.UsageMetadata.ThoughtsTokenCount)
fmt.Println("Output tokens:", response.UsageMetadata.CandidatesTokenCount)

Thinking 모델은 최종 응답 품질을 개선하기 위해 전체 생각을 생성한 다음, 생각 과정에 대한 통찰을 제공하기 위해 요약을 출력해요. 따라서 가격은 API에서 요약만 출력되더라도 요약을 만들기 위해 모델이 생성해야 하는 전체 생각 토큰을 기준으로 해요.

참고: max_output_tokens는 thinking 토큰과 출력 토큰의 합계에 적용되므로 낮은 한도를 설정하면 응답이 잘릴 수 있어요. 자세한 내용은 토큰 한도 및 max_output_tokens를 참고하세요.

토큰에 대해 더 알아보려면 토큰 계산 가이드를 참고하세요.

모범 사례(Best practices)

이 섹션은 thinking 모델을 효율적으로 사용하기 위한 몇 가지 지침을 포함해요. 항상 그렇듯이 우리의 프롬프팅 지침과 모범 사례를 따르면 최상의 결과를 얻을 수 있어요.

디버깅과 스티어링

  • 추론 검토: 기대한 응답을 얻지 못할 때는 Gemini의 생각 요약을 신중히 분석하는 것이 도움이 돼요. 모델이 작업을 어떻게 분해하고 결론에 도달했는지 확인하고, 그 정보를 사용해 올바른 결과로 수정할 수 있어요.
  • 추론에서 지침 제공: 특히 긴 출력을 원한다면 프롬프트에서 모델이 사용하는 thinking량을 제약하는 지침을 제공하는 것이 좋아요. 이를 통해 응답에 더 많은 토큰 출력을 확보할 수 있어요.

작업 복잡성(Task complexity)

  • 쉬운 작업(thinking 꺼도 됨): 사실 검색이나 분류처럼 복잡한 추론이 필요 없는 단순한 요청에는 thinking이 필요 없어요. 예시:

"Where was DeepMind founded?" "Is this email asking for a meeting or just providing information?"

  • 중간 작업(기본값/약간의 thinking): 많은 일반 요청은 어느 정도 단계별 처리나 더 깊은 이해의 이점을 얻어요. Gemini는 다음과 같은 작업에 thinking 능력을 유연하게 사용할 수 있어요:

Analogize photosynthesis and growing up. Compare and contrast electric cars and hybrid cars.

  • 어려운 작업(최대 thinking 능력): 복잡한 수학 문제나 코딩 작업 같은 진정으로 복잡한 과제에는 높은 thinking 예산을 설정하는 것을 권장해요. 이러한 유형의 작업은 모델이 전체 추론 및 계획 능력을 사용해야 하며, 종종 답변 전에 많은 내부 단계가 필요해요. 예시:

Solve problem 1 in AIME 2025: Find the sum of all integer bases b > 9 for which 17b is a divisor of 97b. Write Python code for a web application that visualizes real-time stock market data, including user authentication. Make it as efficient as possible.

지원 모델, 도구, 기능

Thinking 기능은 모든 3 및 2.5 시리즈 모델에서 지원돼요. 모든 모델 기능은 모델 개요 페이지에서 확인할 수 있어요.

Thinking 모델은 Gemini의 모든 도구와 기능과 함께 작동해요. 이를 통해 모델이 외부 시스템과 상호작용하고, 코드를 실행하고, 실시간 정보에 접근하며, 그 결과를 추론과 최종 응답에 통합할 수 있어요.

Thinking cookbook에서 thinking 모델과 함께 도구를 사용하는 예시를 시도해 볼 수 있어요.

더 알아보기 (Learn more)