본문 바로가기
WIKI 기술 지식 베이스

메트릭 분석

원문 보기 위키 갱신

메트릭 분석 (Metrics Analysis)

분석 과정의 일부로 Flagger는 가용성, 오류율, 평균 응답 시간, 그리고 앱별 메트릭에 기반한 기타 목표 같은 서비스 수준 목표(SLO, Service Level Objectives)를 검증할 수 있습니다. SLO 분석 중 성능 저하가 감지되면 최종 사용자에 대한 영향은 최소화하면서 릴리스가 자동으로 롤백됩니다.

출처: 문서

본문

내장 메트릭 (Builtin metrics)

Flagger는 두 가지 내장 메트릭 검사를 제공합니다: HTTP 요청 성공률과 지속 시간(duration)입니다.

  analysis:
    metrics:
    - name: request-success-rate
      interval: 1m
      # minimum req success rate (non 5xx responses)
      # percentage (0-100)
      thresholdRange:
        min: 99
    - name: request-duration
      interval: 1m
      # maximum req duration P99
      # milliseconds
      thresholdRange:
        max: 500

각 메트릭에 대해 thresholdRange로 허용되는 값의 범위를, interval로 창 크기 또는 시계열을 지정할 수 있습니다. 내장 검사는 모든 서비스 메시 / 인그레스 컨트롤러에서 사용할 수 있으며 Prometheus 쿼리로 구현됩니다.

커스텀 메트릭 (Custom metrics)

카나리아 분석은 커스텀 메트릭 검사로 확장할 수 있습니다. MetricTemplate 커스텀 리소스를 사용해 Flagger가 메트릭 프로바이더에 연결하고, float64 값을 반환하는 쿼리를 실행하도록 구성합니다. 쿼리 결과는 지정된 임계값 범위에 따라 canary를 검증하는 데 사용됩니다.

apiVersion: flagger.app/v1beta1
kind: MetricTemplate
metadata:
  name: my-metric
spec:
  provider:
    type: # can be prometheus, datadog, etc
    address: # API URL
    insecureSkipVerify: # if set to true, disables the TLS cert validation
    secretRef:
      name: # name of the secret containing the API credentials
  query: # metric query

쿼리 템플릿에서 사용할 수 있는 변수는 다음과 같습니다:

  • name (canary.metadata.name)
  • namespace (canary.metadata.namespace)
  • target (canary.spec.targetRef.name)
  • service (canary.spec.service.name)
  • ingress (canary.spec.ingresRef.name)
  • interval (canary.spec.analysis.metrics[]​.interval)
  • variables (canary.spec.analysis.metrics[]​.templateVariables)

canary 분석 메트릭은 templateRef로 템플릿을 참조할 수 있습니다:

  analysis:
    metrics:
      - name: "my metric"
        templateRef:
          name: my-metric
          # namespace is optional
          # when not specified, the canary namespace will be used
          namespace: flagger
        # accepted values
        thresholdRange:
          min: 10
          max: 1000
        # metric query time window
        interval: 1m

canary 분석 메트릭은 templateVariables로 커스텀 변수 집합을 참조할 수 있습니다. 이 변수들은 카나리아 분석 중에 참조되는 MetricTemplate 오브젝트에 정의된 쿼리로 주입됩니다:

  analysis:
    metrics:
      - name: "my metric"
        templateRef:
          name: my-metric
          namespace: flagger
        # accepted values
        thresholdRange:
          min: 10
          max: 1000
        # metric query time window
        interval: 1m
        # custom variables used within the referenced metric template
        templateVariables:
          direction: inbound
apiVersion: flagger.app/v1beta1
kind: MetricTemplate
metadata:
  name: my-metric
spec:
  provider:
    type: prometheus
    address: http://prometheus.linkerd-viz:9090
  query: |
    histogram_quantile(
      0.99,
      sum(
        rate(
          response_latency_ms_bucket{
            namespace="{{ namespace }}",
            deployment=~"{{ target }}",
            direction="{{ variables.direction }}"
          }[{{ interval }}]
        )
      ) by (le)
    )

Prometheus

프로바이더 타입을 prometheus로 설정하고 PromQL로 쿼리를 작성하면 Prometheus 서버를 대상으로 하는 커스텀 메트릭 검사를 만들 수 있습니다.

Prometheus 템플릿 예시:

apiVersion: flagger.app/v1beta1
kind: MetricTemplate
metadata:
  name: not-found-percentage
  namespace: istio-system
spec:
  provider:
    type: prometheus
    address: http://prometheus.istio-system:9090
  query: |
    100 - sum(
        rate(
            istio_requests_total{
              reporter="destination",
              destination_workload_namespace="{{ namespace }}",
              destination_workload="{{ target }}",
              response_code!="404"
            }[{{ interval }}]
        )
    )
    /
    sum(
        rate(
            istio_requests_total{
              reporter="destination",
              destination_workload_namespace="{{ namespace }}",
              destination_workload="{{ target }}"
            }[{{ interval }}]
        )
    ) * 100

카나리아 분석에서 템플릿을 참조합니다:

  analysis:
    metrics:
      - name: "404s percentage"
        templateRef:
          name: not-found-percentage
          namespace: istio-system
        thresholdRange:
          max: 5
        interval: 1m

위 구성은 HTTP 404 req/sec 비율이 전체 트래픽의 5% 미만인지 확인해 canary를 검증합니다. 404 비율이 5% 임계값에 도달하면 canary는 실패합니다.

Prometheus gRPC 오류율 예시:

apiVersion: flagger.app/v1beta1
kind: MetricTemplate
metadata:
  name: grpc-error-rate-percentage
  namespace: flagger
spec:
  provider:
    type: prometheus
    address: http://flagger-prometheus.flagger-system:9090
  query: |
    100 - sum(
        rate(
            grpc_server_handled_total{
              grpc_code!="OK",
              kubernetes_namespace="{{ namespace }}",
              kubernetes_pod_name=~"{{ target }}-[0-9a-zA-Z]+(-[0-9a-zA-Z]+)"
            }[{{ interval }}]
        )
    )
    /
    sum(
        rate(
            grpc_server_started_total{
              kubernetes_namespace="{{ namespace }}",
              kubernetes_pod_name=~"{{ target }}-[0-9a-zA-Z]+(-[0-9a-zA-Z]+)"
            }[{{ interval }}]
        )
    ) * 100

위 템플릿은 go-grpc-prometheus로 계측된 gRPC 서비스를 위한 것입니다.

Prometheus 인증

Prometheus API가 기본 인증(basic authentication)을 요구한다면, MetricTemplate과 같은 네임스페이스에 basic-auth 자격 증명으로 시크릿을 만들 수 있습니다:

apiVersion: v1
kind: Secret
metadata:
  name: prom-auth
  namespace: flagger
data:
  username: your-user
  password: your-password

또는 베어러 토큰 인증(SA 토큰 사용)이 필요하다면:

apiVersion: v1
kind: Secret
metadata:
  name: prom-auth
  namespace: flagger
data:
  token: ey1234...

그런 다음 MetricTemplate에서 시크릿을 참조합니다:

apiVersion: flagger.app/v1beta1
kind: MetricTemplate
metadata:
  name: my-metric
  namespace: flagger
spec:
  provider:
    type: prometheus
    address: http://prometheus.monitoring:9090
    secretRef:
      name: prom-auth

Datadog

Datadog 프로바이더를 사용해 커스텀 메트릭 검사를 만들 수 있습니다.

Datadog API 자격 증명으로 시크릿을 만듭니다:

apiVersion: v1
kind: Secret
metadata:
  name: datadog
  namespace: istio-system
data:
  datadog_api_key: your-datadog-api-key
  datadog_application_key: your-datadog-application-key

Datadog 템플릿 예시:

apiVersion: flagger.app/v1beta1
kind: MetricTemplate
metadata:
  name: not-found-percentage
  namespace: istio-system
spec:
  provider:
    type: datadog
    address: https://api.datadoghq.com
    secretRef:
      name: datadog
  query: |
    100 - (
      sum:istio.mesh.request.count{
        reporter:destination,
        destination_workload_namespace:{{ namespace }},
        destination_workload:{{ target }},
        !response_code:404
      }.as_count()
      / 
      sum:istio.mesh.request.count{
        reporter:destination,
        destination_workload_namespace:{{ namespace }},
        destination_workload:{{ target }}
      }.as_count()
    ) * 100

카나리아 분석에서 템플릿을 참조합니다:

  analysis:
    metrics:
      - name: "404s percentage"
        templateRef:
          name: not-found-percentage
          namespace: istio-system
        thresholdRange:
          max: 5
        interval: 1m

Amazon CloudWatch

CloudWatch 메트릭 프로바이더를 사용해 커스텀 메트릭 검사를 만들 수 있습니다.

CloudWatch 템플릿 예시:

apiVersion: flagger.app/v1alpha1
kind: MetricTemplate
metadata:
  name: cloudwatch-error-rate
spec:
  provider:
    type: cloudwatch
    region: ap-northeast-1 # specify the region of your metrics
  query: |
    [
        {
            "Id": "e1",
            "Expression": "m1 / m2",
            "Label": "ErrorRate"
        },
        {
            "Id": "m1",
            "MetricStat": {
                "Metric": {
                    "Namespace": "MyKubernetesCluster",
                    "MetricName": "ErrorCount",
                    "Dimensions": [
                        {
                            "Name": "appName",
                            "Value": "{{ name }}.{{ namespace }}"
                        }
                    ]
                },
                "Period": 60,
                "Stat": "Sum",
                "Unit": "Count"
            },
            "ReturnData": false
        },
        {
            "Id": "m2",
            "MetricStat": {
                "Metric": {
                    "Namespace": "MyKubernetesCluster",
                    "MetricName": "RequestCount",
                    "Dimensions": [
                        {
                            "Name": "appName",
                            "Value": "{{ name }}.{{ namespace }}"
                        }
                    ]
                },
                "Period": 60,
                "Stat": "Sum",
                "Unit": "Count"
            },
            "ReturnData": false
        }
    ]

쿼리 형식 문서는 여기에서 확인할 수 있습니다.

카나리아 분석에서 템플릿을 참조합니다:

  analysis:
    metrics:
      - name: "app error rate"
        templateRef:
          name: cloudwatch-error-rate
        thresholdRange:
          max: 0.1
        interval: 1m

참고 이 프로바이더를 사용하려면 Flagger에 cloudwatch:GetMetricData를 수행할 AWS IAM 권한이 필요합니다.

New Relic

New Relic 프로바이더를 사용해 커스텀 메트릭 검사를 만들 수 있습니다.

New Relic Insights 자격 증명으로 시크릿을 만듭니다:

apiVersion: v1
kind: Secret
metadata:
  name: newrelic
  namespace: istio-system
data:
  newrelic_account_id: your-account-id
  newrelic_query_key: your-insights-query-key

New Relic 템플릿 예시:

apiVersion: flagger.app/v1beta1
kind: MetricTemplate
metadata:
  name: newrelic-error-rate
  namespace: ingress-nginx
spec:
  provider:
    type: newrelic
    secretRef:
      name: newrelic
  query: |
    SELECT 
        filter(sum(nginx_ingress_controller_requests), WHERE status >= '500') / 
        sum(nginx_ingress_controller_requests) * 100
    FROM Metric 
    WHERE metricName = 'nginx_ingress_controller_requests' 
    AND ingress = '{{ ingress }}' AND  namespace = '{{ namespace }}'

카나리아 분석에서 템플릿을 참조합니다:

  analysis:
    metrics:
      - name: "error rate"
        templateRef:
          name: newrelic-error-rate
          namespace: ingress-nginx
        thresholdRange:
          max: 5
        interval: 1m

Graphite

Graphite 프로바이더를 사용해 커스텀 메트릭 검사를 만들 수 있습니다.

Graphite 템플릿 예시:

apiVersion: flagger.app/v1beta1
kind: MetricTemplate
metadata:
  name: graphite-request-success-rate
spec:
  provider:
    type: graphite
    address: http://graphite.monitoring
  query: |
    target=summarize(
      asPercent(
        sumSeries(
          stats.timers.httpServerRequests.app.{{target}}.exception.*.method.*.outcome.{CLIENT_ERROR,INFORMATIONAL,REDIRECTION,SUCCESS}.status.*.uri.*.count
        ),
        sumSeries(
          stats.timers.httpServerRequests.app.{{target}}.exception.*.method.*.outcome.*.status.*.uri.*.count
        )
      ),
      {{interval}},
      'avg'
    )

카나리아 분석에서 템플릿을 참조합니다:

  analysis:
    metrics:
      - name: "success rate"
        templateRef:
          name: graphite-request-success-rate
        thresholdRange:
          min: 90
        interval: 1min

Graphite 인증

Graphite API가 기본 인증을 요구한다면, MetricTemplate과 같은 네임스페이스에 basic-auth 자격 증명으로 시크릿을 만들 수 있습니다:

apiVersion: v1
kind: Secret
metadata:
  name: graphite-basic-auth
  namespace: flagger
data:
  username: your-user
  password: your-password

그런 다음 MetricTemplate에서 시크릿을 참조합니다:

apiVersion: flagger.app/v1beta1
kind: MetricTemplate
metadata:
  name: my-metric
  namespace: flagger
spec:
  provider:
    type: graphite
    address: http://graphite.monitoring
    secretRef:
      name: graphite-basic-auth

Google Cloud Monitoring (Stackdriver)

클러스터에서 Workload Identity를 활성화하고, Cloud Monitoring API에 대한 읽기 권한이 있는 서비스 계정 키를 만든 다음, GCP 서비스 계정과 Kubernetes의 Flagger 서비스 계정 사이에 IAM 정책 바인딩을 만드세요. 이 가이드를 참조할 수 있습니다.

flagger 서비스 계정에 애노테이션을 답니다:

kubectl annotate serviceaccount flagger \
    --namespace <namespace> \
    iam.gke.io/gcp-service-account=<gcp-serviceaccount-name>@<project-id>.iam.gserviceaccount.com

대안으로, json 키를 다운로드해 serviceAccountKey 키로 시크릿에 추가할 수 있습니다 (이 방법은 권장되지 않습니다).

project-id를 담은 시크릿을 만듭니다 (클러스터에서 Workload Identity가 활성화되지 않았다면 서비스 계정 json도 포함합니다):

 kubectl create secret generic gcloud-sa --from-literal=project=<project-id>

그런 다음 메트릭 템플릿에서 시크릿을 참조합니다. 참고: 여기 사용된 특정 MQL 쿼리는 GKE에 Istio가 설치되어 있을 때 동작합니다.

apiVersion: flagger.app/v1beta1
kind: MetricTemplate
metadata:
  name: bytes-sent
  namespace: test
spec:
  provider:
    type: stackdriver
    secretRef: 
      name: gcloud-sa
  query: |
    fetch k8s_container
    | metric 'istio.io/service/server/response_latencies'
    | filter
        (metric.destination_service_name == '{{ service }}-canary'
        && metric.destination_service_namespace == '{{ namespace }}')
    | align delta(1m)
    | every 1m
    | group_by [],
        [value_response_latencies_percentile:
          percentile(value.response_latencies, 99)]

쿼리 언어의 참조는 여기에서 찾을 수 있습니다.

InfluxDB

InfluxDB 프로바이더는 flux 쿼리 언어를 사용합니다.

InfluxDB UI에서 찾을 수 있는 인증 토큰을 담은 시크릿을 만듭니다:

 kubectl create secret generic influx-token --from-literal=token=<token>

그런 다음 메트릭 템플릿에서 시크릿을 참조합니다.

참고: 여기 사용된 특정 MQL 쿼리는 GKE에 Istio가 설치되어 있을 때 동작합니다.

apiVersion: flagger.app/v1beta1
kind: MetricTemplate
metadata:
  name: not-found
  namespace: test
spec:
  provider:
    type: influxdb
    secretRef:
      name: influx-token
  query: |
    from(bucket: "default")
    |> range(start: -2h)
    |> filter(fn: (r) => r["_measurement"] == "istio_requests_total")
    |> filter(fn: (r) => r[" destination_workload_namespace"] == "{{ namespace }}")
    |> filter(fn: (r) => r["destination_workload"] == "{{ target }}")
    |> filter(fn: (r) => r["response_code"] == "500")
    |> count()
    |> yield(name: "count")

Dynatrace

Dynatrace 프로바이더를 사용해 커스텀 메트릭 검사를 만들 수 있습니다.

Dynatrace 토큰으로 시크릿을 만듭니다:

apiVersion: v1
kind: Secret
metadata:
  name: dynatrace
  namespace: istio-system
data:
  dynatrace_token: ZHQwYz...

Dynatrace 메트릭 템플릿 예시:

apiVersion: flagger.app/v1beta1
kind: MetricTemplate
metadata:
  name: response-time-95pct
  namespace: istio-system
spec:
  provider:
    type: dynatrace
    address: https://xxxxxxxx.live.dynatrace.com
    secretRef:
      name: dynatrace
  query: |
    builtin:service.response.time:filter(eq(dt.entity.service,SERVICE-ABCDEFG0123456789)):percentile(95)

카나리아 분석에서 템플릿을 참조합니다:

  analysis:
    metrics:
      - name: "response-time-95pct"
        templateRef:
          name: response-time-95pct
          namespace: istio-system
        thresholdRange:
          max: 1000
        interval: 1m

Keptn

Keptn 프로바이더를 사용해 커스텀 메트릭 검사를 만들 수 있습니다. 이 프로바이더는 단일 메트릭의 값을 나타내는 KeptnMetric 하나의 값을 검증하거나, 서로 다른 데이터 소스에서 온 다양한 메트릭 값을 분석하고 우선순위를 정하는 유연한 등급 부여(grading) 로직을 제공하는 Keptn Analysis를 검증할 수 있습니다.

이 프로바이더는 클러스터에 Keptn이 설치되어 있어야 합니다.

Keptn 메트릭 템플릿 예시:

apiVersion: flagger.app/v1beta1
kind: MetricTemplate
metadata:
  name: response-time
  namespace: istio-system
spec:
  provider:
    type: keptn
  query: keptnmetric/my-namespace/response-time/2m/reporter=destination

이는 my-namespace 네임스페이스에서 response-time이라는 이름의 KeptnMetric을 참조하며, 다음과 같을 수 있습니다:

apiVersion: metrics.keptn.sh/v1beta1
kind: KeptnMetric
metadata:
  name: response-time
  namespace: my-namespace
spec:
  fetchIntervalSeconds: 10
  provider:
    name: my-prometheus-keptn-provider
  query: histogram_quantile(0.8, sum by(le) (rate(http_server_request_latency_seconds_bucket{status_code='200',
    job='simple-go-backend'}[5m[])))

query에는 다음 구성 요소가 포함되며 / 문자로 구분됩니다:

<type>/<namespace>/<resource-name>/<timeframe>/<arguments>
  • type (필수): keptnmetric 또는 analysis 중 하나여야 합니다.
  • namespace (필수): 참조되는 KeptnMetric/AnalysisDefinition의 네임스페이스입니다.
  • resource-name (필수): 참조되는 KeptnMetric/AnalysisDefinition의 이름입니다.
  • timeframe (선택): Analysis에 사용되는 시간 범위입니다. 보통은 Canary의 분석 간격과 같은 값으로 설정합니다. type이 analysis일 때만 관련이 있습니다.
  • arguments (선택): Analysis에 전달되는 인자입니다. 인자는 ; 문자로 구분된 키-값 쌍 목록으로 전달됩니다, 예: foo=bar;bar=foo. type이 analysis일 때만 관련이 있습니다.

analysis 타입의 경우, 프로바이더가 반환하는 값은 0(분석 실패) 또는 1(분석 통과)입니다.

Splunk

Splunk 프로바이더를 사용해 커스텀 메트릭 검사를 만들 수 있습니다.

Splunk o11y UI에서 찾을 수 있는 인증 토큰을 담은 시크릿을 만듭니다:

apiVersion: v1
kind: Secret
metadata:
  name: splunk
  namespace: istio-system
data:
  sf_token_key: your-access-token

Splunk 템플릿 예시:

apiVersion: flagger.app/v1beta1
kind: MetricTemplate
metadata:
  name: success-rate
  namespace: istio-system
spec:
  provider:
    type: splunk
    address: https://api.<REALM>.signalfx.com
    secretRef:
      name: splunk
  query: |
    total = data('traces.count', filter=filter('sf_service', '{{target}}')).sum().publish(enable=False)
    success = data('traces.count', filter=filter('sf_service', '{{target}}') and filter('sf_error', 'false')).sum().publish(enable=False)
    ((success/total) * 100).publish()

쿼리 형식 문서는 여기에서 확인할 수 있습니다.

카나리아 분석에서 템플릿을 참조합니다:

  analysis:
    metrics:
      - name: "success rate"
        templateRef:
          name: success-rate
          namespace: istio-system
        thresholdRange:
          max: 99
        interval: 1m

더 알아보기 (Learn more)