알림 메시지 모범 사례
알림 메시지 모범 사례 (Notification Message Best Practices)
모니터는 비즈니스와 시스템을 원활하게 유지하는 데 필수적이에요. 모니터가 경보를 발생시키면 주의가 필요하다는 신호예요. 하지만 문제를 감지하는 것은 빙산의 일각에 불과하고, 해결 시간에 크게 영향을 주는 것은 알림(notification)이에요.
알림 메시지는 모니터링 시스템과 문제 해결자 사이의 간극을 메워줘요. 불명확하거나 잘못 작성된 메시지는 혼란을 일으키고, 응답 시간을 늦추거나, 문제를 해결하지 못하게 할 수 있어요. 반면 명확하고 실행 가능한 메시지는 팀이 무엇이 잘못됐고 다음에 무엇을 해야 하는지 빠르게 이해하게 해줘요.
이 가이드를 사용해 알림 메시지를 개선하고 다음을 배워보세요:
- 효과적인 커뮤니케이션의 핵심 원칙
- 피해야 할 흔한 실수
- 결과를 내는 메시지를 만드는 팁
제품 관리자부터 개발자까지, 이 리소스는 알림이 시스템 신뢰성과 팀 효율성을 향상시키도록 보장해요.
출처: 문서
본문
알림 구성 (Notification Configuration)
첫 번째 단계는 필수 필드로 알림을 구성하는 거예요:
- Monitor Name(모니터 이름). 이는 알림 제목(Notification title)이기도 해요.
- Monitor Message(모니터 메시지). 이는 알림의 본문이에요.
{% image source="https://docs.dd-static.net/images/monitors/guide/notification_message_best_practices/monitor_notification_message.bf73a18a39d1708a45b2b5de4c966fc7.png?auto=format&fit=max&w=850 1x, https://docs.dd-static.net/images/monitors/guide/notification_message_best_practices/monitor_notification_message.bf73a18a39d1708a45b2b5de4c966fc7.png?auto=format&fit=max&w=850&dpr=2 2x" alt="Monitor notification message configuration" /%}
이름 (Name)
대응자가 경보 컨텍스트를 빠르게 이해할 수 있도록 핵심 정보를 포함해 Monitor Name을 작성해요. 모니터 제목은 다음을 포함해 신호에 대한 명확하고 간결한 설명을 줘야 해요:
- 실패 모드(failure mode) 또는 벗어난 메트릭
- 영향을 받는 리소스(예: 데이터센터, Kubernetes 클러스터, 호스트, 서비스)
| 수정 필요 (Needs Revision) | 개선된 제목 (Improved Title) |
|---|---|
| Memory usage | High memory usage on {{pod_name.name}} |
두 예시 모두 메모리 소비 모니터를 참조하지만, 개선된 제목은 집중 조사를 위한 필수 컨텍스트를 포함한 철저한 표현을 제공해요.
메시지 (Message)
온콜(on-call) 대응자는 알림 본문에 의존해 경보를 이해하고 조치를 취해요. 명확성을 위해 간결하고 정확하며 가독성 있는 메시지를 작성해요.
- 무엇이 실패하고 있는지 정확히 언급하고 주요 근본 원인을 나열해요
- 빠른 해결 지침을 위한 솔루션 런북(runbook)을 추가해요
- 명확한 다음 단계를 위해 관련 페이지 링크를 포함해요
- 이메일 직접 알림 또는 통합 핸들(integration handles)(예: Slack)을 통해 적절한 수신자에게 알림을 보내도록 해요
모니터 메시지를 더 향상시킬 수 있는 고급 기능을 살펴보려면 다음 섹션을 읽어보세요.
변수 (Variables)
모니터 메시지 변수는 실시간 컨텍스트 정보로 알림 메시지를 커스터마이즈할 수 있는 동적 플레이스홀더예요. 변수를 사용해 메시지 명확성을 높이고 상세한 컨텍스트를 제공해요. 변수에는 두 가지 유형이 있어요:
| 변수 유형 (Variable Type) | 설명 (Description) |
|---|---|
| 조건부 (Conditional) | 모니터 상태 같은 조건에 따라 메시지 컨텍스트를 조정하는 "if-else" 로직 사용 |
| 템플릿 (Template) | 컨텍스트 정보로 모니터 알림을 풍부하게 만듦 |
변수는 Multi-Alert 모니터에서 특히 중요해요. 트리거되면 어떤 그룹이 책임인지 알아야 하기 때문이에요. 예를 들어 컨테이너별 CPU 사용량을 호스트별로 그룹화해 모니터링한다고 해볼게요. 유용한 변수는 경보를 트리거한 호스트를 나타내는 {{host.name}}이에요.
{% image source="https://docs.dd-static.net/images/monitors/guide/notification_message_best_practices/query_parameters.bf371540ca6032b17fcf95c083ac2762.png?auto=format&fit=max&w=850 1x, https://docs.dd-static.net/images/monitors/guide/notification_message_best_practices/query_parameters.bf371540ca6032b17fcf95c083ac2762.png?auto=format&fit=max&w=850&dpr=2 2x" alt="Example monitor query of container.cpu.usage metric averaged by host" /%}
조건부 변수 (Conditional variables)
이 변수들은 필요와 사용 사례에 따라 분기 로직을 구현해 알림 메시지를 맞춤화할 수 있게 해줘요. 경보를 트리거한 그룹에 따라 다른 사람/그룹에게 알리려면 조건부 변수를 사용해요.
{{#is_exact_match "role.name" "network"}}
# The content displays if the host triggering the alert contains `network` in the role name, and only notifies @[email protected].
@[email protected]
{{/is_exact_match}}
경보를 트리거한 그룹에 특정 문자열이 포함되어 있으면 알림을 받을 수 있어요.
{{#is_match "datacenter.name" "us"}}
# The content displays if the region triggering the alert contains `us` (such as us1 or us3)
@[email protected]
{{/is_match}}
더 많은 정보와 예시는 조건부 변수(Conditional Variables) 문서를 참고해요.
템플릿 변수 (Template variables)
모니터가 경보를 발생시킨 원인이 된 메타데이터에 접근하려면 {{value}} 같은 모니터 템플릿 변수를 추가하고, 경보 컨텍스트와 관련된 정보도 추가해요.
예를 들어 호스트 이름, IP, 모니터 쿼리 값을 보고 싶다면:
The CPU for {{host.name}} (IP:{{host.ip}}) reached a critical value of {{value}}.
사용 가능한 템플릿 변수 목록은 문서를 참고해요.
또한 템플릿 변수를 사용해 알림을 자동으로 라우팅하는 동적 링크와 핸들을 만들 수도 있어요. 핸들 예시:
@slack-{{service.name}} There is an ongoing issue with {{service.name}}.
그룹 service:ad-server가 트리거될 때 다음과 같은 결과가 나와요:
@slack-ad-server There is an ongoing issue with ad-server.
링크 예시:
[https://app.datadoghq.com/dash/integration/system_overview?tpl_var_scope=host:{{host.name](https://app.datadoghq.com/dash/integration/system_overview?tpl_var_scope=host:{{host.name)}}
모범 사례를 따르는 알림 메시지 예시 (Example of a notification message following best practices)
## What's happening? The CPU usage on {{host.name}} has exceeded the defined threshold.
Current CPU Usage: {{value}} Threshold: {{threshold}} Time: {{last_triggered_at_epoch}}
## Impact 1. Customers are experiencing lag on the website. 2. Timeouts and Errors.
## Why? There can be several reasons as to why the CPU usage exceeded the threshold:
- Increase in traffic
- Hardware Issues
- External Attack
## How to troubleshoot/solve the issue? 1. Analyze workload to identify CPU-intensive processes. a. for OOM - increase pod limits if too low 2. Upscale {{host.name}} capacity by adding more replicas: a. directly: b. change configuration through add more replicas runbook 3. Check for any Kafka issues 4. Check for any other outages/incident (attempted connections)
## Related links * Troubleshooting Dashboard * App Dashboard * Logs * Infrastructure * Pipeline Overview * App Documentation * Failure Modes
더 알아보기 (Learn more)