본문 바로가기
WIKI 기술 지식 베이스

Gloo 카나리아 배포

원문 보기 위키 갱신

Gloo 카나리아 배포 (Gloo Canary Deployments)

Gloo Edge ingress controller와 Flagger를 함께 사용해서 canary 릴리스와 A/B 테스트를 자동화하는 방법을 이 문서에서 알려드릴게요. Gloo RouteTable을 활용해 canary 트래픽을 전환하며 배포를 검증하는 구성을 살펴봅시다.

출처: 문서

본문

사전 준비 (Prerequisites)

Flagger는 Kubernetes 클러스터 v1.16 이상과 Gloo Edge ingress 1.6.0 이상이 필요합니다.

이 가이드는 Flagger 버전 1.6.0 이상을 위해 작성되었습니다. 이전 Flagger 버전은 canary를 처리하기 위해 Gloo UpstreamGroup을 사용했지만, 새 버전의 Flagger는 canary와 A/B 테스트를 처리하기 위해 Gloo RouteTable을 사용합니다.

Helm v3로 Gloo를 설치합니다:

helm repo add gloo https://storage.googleapis.com/solo-public-helm
kubectl create ns gloo-system
helm upgrade -i gloo gloo/gloo \
--namespace gloo-system

Gloo와 같은 네임스페이스에 Flagger와 Prometheus 애드온을 설치합니다:

helm repo add flagger https://flagger.app

helm upgrade -i flagger flagger/flagger \
--namespace gloo-system \
--set prometheus.install=true \
--set meshProvider=gloo

부트스트랩 (Bootstrap)

Flagger는 Kubernetes deployment와 선택적으로 horizontal pod autoscaler(HPA)를 받아서 일련의 오브젝트(Kubernetes deployments, ClusterIP services, Gloo route tables, upstreams)를 생성합니다. 이 오브젝트들은 앱을 클러스터 외부에 노출하며 canary 분석과 승격을 진행시킵니다.

테스트 네임스페이스를 생성합니다:

kubectl create ns test

deployment와 horizontal pod autoscaler를 생성합니다:

kubectl -n test apply -k https://github.com/fluxcd/flagger//kustomize/podinfo?ref=main

canary 분석 중 트래픽을 발생시킬 부하 테스트 서비스를 배포합니다:

kubectl -n test apply -k https://github.com/fluxcd/flagger//kustomize/tester?ref=main

Flagger가 생성할 route table을 참조하는 virtual service 정의를 생성합니다(app.example.com을 자신의 도메인으로 바꾸세요):

apiVersion: gateway.solo.io/v1
kind: VirtualService
metadata:
  name: podinfo
  namespace: test
spec:
  virtualHost:
    domains:
      - 'app.example.com'
    routes:
      - matchers:
         - prefix: /
        delegateAction:
          ref:
            name: podinfo
            namespace: test

위 리소스를 podinfo-virtualservice.yaml로 저장한 뒤 적용합니다:

kubectl apply -f ./podinfo-virtualservice.yaml

canary 커스텀 리소스를 생성합니다(app.example.com을 자신의 도메인으로 바꾸세요):

apiVersion: flagger.app/v1beta1
kind: Canary
metadata:
  name: podinfo
  namespace: test
spec:
  # upstreamRef (optional)
  # defines an upstream to copy the spec from when flagger generates new upstreams.
  # necessary to copy over TLS config, circuit breakers, etc. (anything nonstandard)
#  upstreamRef:
#    apiVersion: gloo.solo.io/v1
#    kind: Upstream
#    name: podinfo-upstream
#    namespace: gloo-system
  provider: gloo
  # deployment reference
  targetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: podinfo
  # HPA reference (optional)
  autoscalerRef:
    apiVersion: autoscaling/v2
    kind: HorizontalPodAutoscaler
    name: podinfo
  service:
    # ClusterIP port number
    port: 9898
    # container port number or name (optional)
    targetPort: 9898
  analysis:
    # schedule interval (default 60s)
    interval: 10s
    # max number of failed metric checks before rollback
    threshold: 5
    # max traffic percentage routed to canary
    # percentage (0-100)
    maxWeight: 50
    # canary increment step
    # percentage (0-100)
    stepWeight: 5
    # Gloo Prometheus checks
    metrics:
    - name: request-success-rate
      # minimum req success rate (non 5xx responses)
      # percentage (0-100)
      thresholdRange:
        min: 99
      interval: 1m
    - name: request-duration
      # maximum req duration P99
      # milliseconds
      thresholdRange:
        max: 500
      interval: 30s
    # testing (optional)
    webhooks:
      - name: acceptance-test
        type: pre-rollout
        url: http://flagger-loadtester.test/
        timeout: 10s
        metadata:
          type: bash
          cmd: "curl -sd 'test' http://podinfo-canary:9898/token | grep token"
      - name: load-test
        url: http://flagger-loadtester.test/
        timeout: 5s
        metadata:
          type: cmd
          cmd: "hey -z 2m -q 5 -c 2 -host app.example.com http://gateway-proxy.gloo-system"

참고: upstreamRef를 사용할 때 다음 필드가 원본 upstream에서 복사됩니다: Labels, SslConfig, CircuitBreakers, ConnectionConfig, UseHttp2, InitialStreamWindowSize

위 리소스를 podinfo-canary.yaml로 저장한 뒤 적용합니다:

kubectl apply -f ./podinfo-canary.yaml

몇 초 후 Flagger가 canary 오브젝트들을 생성합니다:

# applied 
deployment.apps/podinfo
horizontalpodautoscaler.autoscaling/podinfo
virtualservices.gateway.solo.io/podinfo
canary.flagger.app/podinfo

# generated 
deployment.apps/podinfo-primary
horizontalpodautoscaler.autoscaling/podinfo-primary
service/podinfo
service/podinfo-canary
service/podinfo-primary
routetables.gateway.solo.io/podinfo
upstreams.gloo.solo.io/test-podinfo-canaryupstream-9898
upstreams.gloo.solo.io/test-podinfo-primaryupstream-9898

부트스트랩이 끝나면 Flagger는 canary 상태를 initialized로 설정합니다:

kubectl -n test get canary podinfo

NAME      STATUS        WEIGHT   LASTTRANSITIONTIME
podinfo   Initialized   0        2019-05-17T08:09:51Z

자동화된 canary 승격 (Automated canary promotion)

Flagger는 HTTP 요청 성공률, 요청 평균 지속 시간, pod 상태 같은 주요 성능 지표(KPI)를 측정하면서 canary로 트래픽을 점진적으로 이동시키는 컨트롤 루프를 구현합니다. KPI 분석 결과에 따라 canary는 승격되거나 중단되며, 분석 결과는 Slack으로 게시됩니다.

컨테이너 이미지를 업데이트해 canary 배포를 트리거합니다:

kubectl -n test set image deployment/podinfo \
podinfod=ghcr.io/stefanprodan/podinfo:6.0.1

Flagger는 deployment 리비전이 바뀐 것을 감지하고 새로운 롤아웃을 시작합니다:

kubectl -n test describe canary/podinfo

Status:
  Canary Weight:         0
  Failed Checks:         0
  Phase:                 Succeeded
Events:
  Type     Reason  Age   From     Message
  ----     ------  ----  ----     -------
  Normal   Synced  3m    flagger  New revision detected podinfo.test
  Normal   Synced  3m    flagger  Scaling up podinfo.test
  Warning  Synced  3m    flagger  Waiting for podinfo.test rollout to finish: 0 of 1 updated replicas are available
  Normal   Synced  3m    flagger  Advance podinfo.test canary weight 5
  Normal   Synced  3m    flagger  Advance podinfo.test canary weight 10
  Normal   Synced  3m    flagger  Advance podinfo.test canary weight 15
  Normal   Synced  2m    flagger  Advance podinfo.test canary weight 20
  Normal   Synced  2m    flagger  Advance podinfo.test canary weight 25
  Normal   Synced  1m    flagger  Advance podinfo.test canary weight 30
  Normal   Synced  1m    flagger  Advance podinfo.test canary weight 35
  Normal   Synced  55s   flagger  Advance podinfo.test canary weight 40
  Normal   Synced  45s   flagger  Advance podinfo.test canary weight 45
  Normal   Synced  35s   flagger  Advance podinfo.test canary weight 50
  Normal   Synced  25s   flagger  Copying podinfo.test template spec to podinfo-primary.test
  Warning  Synced  15s   flagger  Waiting for podinfo-primary.test rollout to finish: 1 of 2 updated replicas are available
  Normal   Synced  5s    flagger  Promotion completed! Scaling down podinfo.test

canary 분석 중 deployment에 새 변경사항을 적용하면 Flagger가 분석을 다시 시작한다는 점을 참고하세요.

모든 canary를 다음 명령으로 모니터링할 수 있습니다:

watch kubectl get canaries --all-namespaces

NAMESPACE   NAME      STATUS        WEIGHT   LASTTRANSITIONTIME
test        podinfo   Progressing   15       2019-05-17T14:05:07Z
prod        frontend  Succeeded     0        2019-05-17T16:15:07Z
prod        backend   Failed        0        2019-05-17T17:05:07Z

자동화된 롤백 (Automated rollback)

canary 분석 중 HTTP 500 오류와 높은 지연을 발생시켜 Flagger가 결함 버전을 일시 중지하고 롤백하는지 테스트할 수 있습니다.

또 다른 canary 배포를 트리거합니다:

kubectl -n test set image deployment/podinfo \
podinfod=ghcr.io/stefanprodan/podinfo:6.0.2

HTTP 500 오류를 발생시킵니다:

watch curl -H 'Host: app.example.com' http://gateway-proxy.gloo-system/status/500

높은 지연을 발생시킵니다:

watch curl -H 'Host: app.example.com' http://gateway-proxy.gloo-system/delay/2

실패한 검사 횟수가 canary 분석 임계값에 도달하면 트래픽이 primary로 되돌아가고, canary는 0으로 스케일 다운되며 롤아웃은 실패로 표시됩니다.

kubectl -n test describe canary/podinfo

Status:
  Canary Weight:         0
  Failed Checks:         10
  Phase:                 Failed
Events:
  Type     Reason  Age   From     Message
  ----     ------  ----  ----     -------
  Normal   Synced  3m    flagger  Starting canary deployment for podinfo.test
  Normal   Synced  3m    flagger  Advance podinfo.test canary weight 5
  Normal   Synced  3m    flagger  Advance podinfo.test canary weight 10
  Normal   Synced  3m    flagger  Advance podinfo.test canary weight 15
  Normal   Synced  3m    flagger  Halt podinfo.test advancement success rate 69.17% < 99%
  Normal   Synced  2m    flagger  Halt podinfo.test advancement success rate 61.39% < 99%
  Normal   Synced  2m    flagger  Halt podinfo.test advancement success rate 55.06% < 99%
  Normal   Synced  2m    flagger  Halt podinfo.test advancement success rate 47.00% < 99%
  Normal   Synced  2m    flagger  (combined from similar events): Halt podinfo.test advancement success rate 38.08% < 99%
  Warning  Synced  1m    flagger  Rolling back podinfo.test failed checks threshold reached 10
  Warning  Synced  1m    flagger  Canary failed! Scaling down podinfo.test

커스텀 메트릭 (Custom metrics)

canary 분석은 Prometheus 쿼리로 확장할 수 있습니다.

데모 앱은 Prometheus로 계측되어 있으므로 HTTP 요청 지속 시간 히스토그램을 사용해 canary를 검증하는 커스텀 검사를 만들 수 있습니다.

메트릭 템플릿을 생성하고 클러스터에 적용합니다:

apiVersion: flagger.app/v1beta1
kind: MetricTemplate
metadata:
  name: not-found-percentage
  namespace: test
spec:
  provider:
    type: prometheus
    address: http://flagger-prometheus.gloo-system:9090
  query: |
    100 - sum(
        rate(
            http_request_duration_seconds_count{
              kubernetes_namespace="{{ namespace }}",
              kubernetes_pod_name=~"{{ target }}-[0-9a-zA-Z]+(-[0-9a-zA-Z]+)"
              status!="{{ interval }}"
            }[1m]
        )
    )
    /
    sum(
        rate(
            http_request_duration_seconds_count{
              kubernetes_namespace="{{ namespace }}",
              kubernetes_pod_name=~"{{ target }}-[0-9a-zA-Z]+(-[0-9a-zA-Z]+)"
            }[{{ interval }}]
        )
    ) * 100

canary 분석을 수정하고 다음 메트릭을 추가합니다:

  analysis:
    metrics:
      - name: "404s percentage"
        templateRef:
          name: not-found-percentage
        thresholdRange:
          max: 5
        interval: 1m

위 구성은 HTTP 404 req/sec 비율이 전체 트래픽의 5% 미만인지 확인해 canary를 검증합니다. 404 비율이 5% 임계값에 도달하면 canary는 실패합니다.

컨테이너 이미지를 업데이트해 canary 배포를 트리거합니다:

kubectl -n test set image deployment/podinfo \
podinfod=ghcr.io/stefanprodan/podinfo:6.0.3

404를 발생시킵니다:

watch curl -H 'Host: app.example.com' http://gateway-proxy.gloo-system/status/404

Flagger 로그를 확인합니다:

kubectl -n gloo-system logs deployment/flagger -f | jq .msg

Starting canary deployment for podinfo.test
Advance podinfo.test canary weight 5
Advance podinfo.test canary weight 10
Advance podinfo.test canary weight 15
Halt podinfo.test advancement 404s percentage 6.20 > 5
Halt podinfo.test advancement 404s percentage 6.45 > 5
Halt podinfo.test advancement 404s percentage 7.60 > 5
Halt podinfo.test advancement 404s percentage 8.69 > 5
Halt podinfo.test advancement 404s percentage 9.70 > 5
Rolling back podinfo.test failed checks threshold reached 5
Canary failed! Scaling down podinfo.test

알림이 구성되어 있다면 Flagger는 canary가 실패한 이유와 함께 알림을 보냅니다.

A/B 테스트 (A/B Testing)

가중치 기반 라우팅 외에도 Flagger는 HTTP 매치 조건을 기반으로 canary로 트래픽을 라우팅하도록 구성할 수 있습니다. A/B 테스트 시나리오에서는 HTTP 헤더나 쿠키를 사용해 특정 사용자 세그먼트를 타깃팅합니다. 이는 세션 어피니티가 필요한 프론트엔드 앱에 특히 유용합니다.

canary 분석을 수정하고 max/step 가중치를 제거하고 매치 조건과 iterations를 추가합니다:

analysis:
  interval: 1m
  threshold: 5
  iterations: 10
  match:
  - headers:
      x-canary:
        exact: "insider"
  webhooks:
  - name: load-test
    url: http://flagger-loadtester.test/
    metadata:
      cmd: "hey -z 1m -q 5 -c 5 -H 'X-Canary: insider' -host app.example.com http://gateway-proxy.gloo-system"

위 구성은 X-Canary: insider 헤더가 있는 사용자를 타깃으로 10분간 분석을 실행합니다.

컨테이너 이미지를 업데이트해 canary 배포를 트리거합니다:

kubectl -n test set image deployment/podinfo \
podinfod=ghcr.io/stefanprodan/podinfo:6.0.4

Flagger는 deployment 리비전이 바뀐 것을 감지하고 A/B 테스트를 시작합니다:

kubectl -n gloo-system logs deploy/flagger -f | jq .msg

New revision detected! Progressing canary analysis for podinfo.test
Advance podinfo.test canary iteration 1/10
Advance podinfo.test canary iteration 2/10
Advance podinfo.test canary iteration 3/10
Advance podinfo.test canary iteration 4/10
Advance podinfo.test canary iteration 5/10
Advance podinfo.test canary iteration 6/10
Advance podinfo.test canary iteration 7/10
Advance podinfo.test canary iteration 8/10
Advance podinfo.test canary iteration 9/10
Advance podinfo.test canary iteration 10/10
Copying podinfo.test template spec to podinfo-primary.test
Waiting for podinfo-primary.test rollout to finish: 1 of 2 updated replicas are available
Routing all traffic to primary
Promotion completed! Scaling down podinfo.test

웹 브라우저의 user agent 헤더는 기기나 OS에 따라 사용자 세그먼트를 나눌 수 있게 합니다.

예를 들어 모든 모바일 사용자를 canary 인스턴스로 라우팅하려면:

match:
- headers:
    user-agent:
      regex: ".*Mobile.*"

또는 Android 사용자만 타깃으로 하려면:

match:
- headers:
    user-agent:
      regex: ".*Android.*"

또는 특정 브라우저 버전:

match:
- headers:
    user-agent:
      regex: ".*Firefox.*"

분석 과정을 깊이 있게 보려면 사용 문서를 읽어보세요.

더 알아보기 (Learn more)