본문 바로가기
WIKI 기술 지식 베이스

Linkerd 카나리아 배포

원문 보기 위키 갱신

Linkerd 카나리아 배포 (Linkerd Canary Deployments)

Linkerd와 Flagger를 함께 사용해서 canary 배포를 자동화하는 방법을 이 문서에서 알려드릴게요. SMI TrafficSplit과 Gateway API HTTPRoute를 활용해 canary 트래픽을 점진적으로 전환하며 검증하는 구성을 살펴봅시다.

출처: 문서

본문

사전 준비 (Prerequisites)

Flagger는 Kubernetes 클러스터 v1.21 이상과 Linkerd 2.14 이상이 필요합니다.

Linkerd와 Prometheus(Linkerd Viz의 일부)를 설치합니다:

# The CRDs need to be installed beforehand
linkerd install --crds | kubectl apply -f -

linkerd install | kubectl apply -f -
linkerd viz install | kubectl apply -f -

# For linkerd versions 2.12 and later, the SMI extension needs to be install in
# order to enable TrafficSplits
curl -sL https://linkerd.github.io/linkerd-smi/install | sh
linkerd smi install | kubectl apply -f -

flagger-system 네임스페이스에 Flagger를 설치합니다:

kubectl apply -k github.com/fluxcd/flagger//kustomize/linkerd

Helm을 선호한다면 Linkerd, Linkerd Viz, Linkerd-SMI, Flagger를 설치하는 명령은 다음과 같습니다:

helm repo add linkerd https://helm.linkerd.io/stable
helm install linkerd-crds linkerd/linkerd-crds -n linkerd --create-namespace
# See https://linkerd.io/2/tasks/generate-certificates/ for how to generate the
# certs referred below
helm install linkerd-control-plane linkerd/linkerd-control-plane \
  -n linkerd \
  --set-file identityTrustAnchorsPEM=ca.crt \
  --set-file identity.issuer.tls.crtPEM=issuer.crt \
  --set-file identity.issuer.tls.keyPEM=issuer.key \

helm install linkerd-viz linkerd/linkerd-viz -n linkerd-viz --create-namespace

helm install flagger flagger/flagger \
  --n flagger-system \
  --set meshProvider=gatewayapi:v1beta1 \
  --set metricsServer=http://prometheus.linkerd-viz:9090 \
  --set linkerdAuthPolicy.create=true

부트스트랩 (Bootstrap)

Flagger는 Kubernetes deployment와 선택적으로 horizontal pod autoscaler(HPA)를 받아서 일련의 오브젝트(Kubernetes deployments, ClusterIP services, SMI traffic split)를 생성합니다. 이 오브젝트들은 메시 안에서 앱을 노출하며 canary 분석과 승격을 진행시킵니다.

테스트 네임스페이스를 생성하고 Linkerd 프록시 주입을 활성화합니다:

kubectl create ns test
kubectl annotate namespace test linkerd.io/inject=enabled

canary 분석 중 트래픽을 발생시킬 부하 테스트 서비스를 설치합니다:

kubectl apply -k https://github.com/fluxcd/flagger//kustomize/tester?ref=main

deployment와 horizontal pod autoscaler를 생성합니다:

kubectl apply -k https://github.com/fluxcd/flagger//kustomize/podinfo?ref=main

podinfo deployment용 메트릭 템플릿과 canary 커스텀 리소스를 생성합니다:

---
apiVersion: flagger.app/v1beta1
kind: MetricTemplate
metadata:
  name: success-rate
  namespace: test
spec:
  provider:
    type: prometheus
    address: http://prometheus.linkerd-viz:9090
  query: |
    sum(
      rate(
        response_total{
          namespace="{{ namespace }}",
          deployment=~"{{ target }}",
          classification!="failure",
          direction="{{ variables.direction }}"
        }[{{ interval }}]
      )
    ) 
    / 
    sum(
      rate(
        response_total{
          namespace="{{ namespace }}",
          deployment=~"{{ target }}",
          direction="{{ variables.direction }}"
        }[{{ interval }}]
      )
    ) 
    * 100
---
apiVersion: flagger.app/v1beta1
kind: MetricTemplate
metadata:
  name: latency
  namespace: test
spec:
  provider:
    type: prometheus
    address: http://prometheus.linkerd-viz:9090
  query: |
    histogram_quantile(
        0.99,
        sum(
            rate(
                response_latency_ms_bucket{
                    namespace="{{ namespace }}",
                    deployment=~"{{ target }}",
                    direction="{{ variables.direction }}"
                    }[{{ interval }}]
                )
            ) by (le)
        )
---
apiVersion: flagger.app/v1beta1
kind: Canary
metadata:
  name: podinfo
  namespace: test
spec:
  # deployment reference
  targetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: podinfo
  # HPA reference (optional)
  autoscalerRef:
    apiVersion: autoscaling/v2
    kind: HorizontalPodAutoscaler
    name: podinfo
  # the maximum time in seconds for the canary deployment
  # to make progress before it is rollback (default 600s)
  progressDeadlineSeconds: 60
  service:
    # ClusterIP port number
    port: 9898
    # container port number or name (optional)
    targetPort: 9898
    # Reference to the Service that the generated HTTPRoute would attach to.
    gatewayRefs:
      - name: podinfo
        namespace: test
        group: core
        kind: Service
        port: 9898
  analysis:
    # schedule interval (default 60s)
    interval: 30s
    # max number of failed metric checks before rollback
    threshold: 5
    # max traffic percentage routed to canary
    # percentage (0-100)
    maxWeight: 50
    # canary increment step
    # percentage (0-100)
    stepWeight: 5
    # Linkerd Prometheus checks
    metrics:
    - name: success-rate
      templateRef:
        name: success-rate
        namespace: test
      # minimum req success rate (non 5xx responses)
      # percentage (0-100)
      thresholdRange:
        min: 99
      interval: 1m
      templateVariables:
        direction: inbound
    - name: latency
      templateRef:
        name: latency
        namespace: test
      # maximum req duration P99
      # milliseconds
      thresholdRange:
        max: 500
      interval: 30s
      templateVariables:
        direction: inbound
    # testing (optional)
    webhooks:
      - name: acceptance-test
        type: pre-rollout
        url: http://flagger-loadtester.test/
        timeout: 30s
        metadata:
          type: bash
          cmd: "curl -sd 'test' http://podinfo-canary.test:9898/token | grep token"
      - name: load-test
        type: rollout
        url: http://flagger-loadtester.test/
        metadata:
          cmd: "hey -z 2m -q 10 -c 2 http://podinfo-canary.test:9898/"

위 리소스를 podinfo-canary.yaml로 저장한 뒤 적용합니다:

kubectl apply -f ./podinfo-canary.yaml

canary 분석이 시작되면 Flagger는 canary로 트래픽을 라우팅하기 전에 pre-rollout 웹훅을 호출합니다. canary 분석은 30초마다 HTTP 메트릭과 롤아웃 훅을 검증하면서 5분간 실행됩니다.

몇 초 후 Flagger가 canary 오브젝트들을 생성합니다:

# applied
deployment.apps/podinfo
horizontalpodautoscaler.autoscaling/podinfo
ingresses.extensions/podinfo
canary.flagger.app/podinfo

# generated
deployment.apps/podinfo-primary
horizontalpodautoscaler.autoscaling/podinfo-primary
service/podinfo
service/podinfo-canary
service/podinfo-primary
trafficsplits.split.smi-spec.io/podinfo

부트스트랩 후 podinfo deployment는 0으로 스케일 다운되고 podinfo.test로의 트래픽은 primary pod로 라우팅됩니다. canary 분석 중에는 podinfo-canary.test 주소로 canary pod를 직접 타깃할 수 있습니다.

자동화된 canary 승격 (Automated canary promotion)

Flagger는 HTTP 요청 성공률, 요청 평균 지속 시간, pod 상태 같은 주요 성능 지표(KPI)를 측정하면서 canary로 트래픽을 점진적으로 이동시키는 컨트롤 루프를 구현합니다. KPI 분석 결과에 따라 canary는 승격되거나 중단되며, 분석 결과는 Slack으로 게시됩니다.

컨테이너 이미지를 업데이트해 canary 배포를 트리거합니다:

kubectl -n test set image deployment/podinfo \
podinfod=ghcr.io/stefanprodan/podinfo:6.0.1

Flagger는 deployment 리비전이 바뀐 것을 감지하고 새로운 롤아웃을 시작합니다:

kubectl -n test describe canary/podinfo

Status:
  Canary Weight:         0
  Failed Checks:         0
  Phase:                 Succeeded
Events:
 New revision detected! Scaling up podinfo.test
 Waiting for podinfo.test rollout to finish: 0 of 1 updated replicas are available
 Pre-rollout check acceptance-test passed
 Advance podinfo.test canary weight 5
 Advance podinfo.test canary weight 10
 Advance podinfo.test canary weight 15
 Advance podinfo.test canary weight 20
 Advance podinfo.test canary weight 25
 Waiting for podinfo.test rollout to finish: 1 of 2 updated replicas are available
 Advance podinfo.test canary weight 30
 Advance podinfo.test canary weight 35
 Advance podinfo.test canary weight 40
 Advance podinfo.test canary weight 45
 Advance podinfo.test canary weight 50
 Copying podinfo.test template spec to podinfo-primary.test
 Waiting for podinfo-primary.test rollout to finish: 1 of 2 updated replicas are available
 Promotion completed! Scaling down podinfo.test

canary 분석 중 deployment에 새 변경사항을 적용하면 Flagger가 분석을 다시 시작한다는 점을 참고하세요.

canary 배포는 다음 오브젝트 중 하나가 변경되면 트리거됩니다:

  • Deployment PodSpec (컨테이너 이미지, command, ports, env, resources 등)
  • 볼륨으로 마운트되거나 환경 변수에 매핑된 ConfigMaps
  • 볼륨으로 마운트되거나 환경 변수에 매핑된 Secrets

모든 canary를 다음 명령으로 모니터링할 수 있습니다:

watch kubectl get canaries --all-namespaces

NAMESPACE   NAME      STATUS        WEIGHT   LASTTRANSITIONTIME
test        podinfo   Progressing   15       2019-06-30T14:05:07Z
prod        frontend  Succeeded     0        2019-06-30T16:15:07Z
prod        backend   Failed        0        2019-06-30T17:05:07Z

자동화된 롤백 (Automated rollback)

canary 분석 중 HTTP 500 오류와 높은 지연을 발생시켜 Flagger가 결함 버전을 일시 중지하고 롤백하는지 테스트할 수 있습니다.

또 다른 canary 배포를 트리거합니다:

kubectl -n test set image deployment/podinfo \
podinfod=ghcr.io/stefanprodan/podinfo:6.0.2

부하 테스터 pod에 접속합니다:

kubectl -n test exec -it flagger-loadtester-xx-xx sh

HTTP 500 오류를 발생시킵니다:

watch -n 1 curl http://podinfo-canary.test:9898/status/500

지연을 발생시킵니다:

watch -n 1 curl http://podinfo-canary.test:9898/delay/1

실패한 검사 횟수가 canary 분석 임계값에 도달하면 트래픽이 primary로 되돌아가고, canary는 0으로 스케일 다운되며 롤아웃은 실패로 표시됩니다.

kubectl -n test describe canary/podinfo

Status:
  Canary Weight:         0
  Failed Checks:         10
  Phase:                 Failed
Events:
 Starting canary analysis for podinfo.test
 Pre-rollout check acceptance-test passed
 Advance podinfo.test canary weight 5
 Advance podinfo.test canary weight 10
 Advance podinfo.test canary weight 15
 Halt podinfo.test advancement success rate 69.17% < 99%
 Halt podinfo.test advancement success rate 61.39% < 99%
 Halt podinfo.test advancement success rate 55.06% < 99%
 Halt podinfo.test advancement request duration 1.20s > 0.5s
 Halt podinfo.test advancement request duration 1.45s > 0.5s
 Rolling back podinfo.test failed checks threshold reached 5
 Canary failed! Scaling down podinfo.test

커스텀 메트릭 (Custom metrics)

canary 분석은 Prometheus 쿼리로 확장할 수 있습니다.

not found 오류에 대한 검사를 정의해 봅시다. canary 분석을 수정하고 다음 메트릭을 추가합니다:

  analysis:
    metrics:
    - name: "404s percentage"
      threshold: 3
      query: |
        100 - sum(
            rate(
                response_total{
                    namespace="test",
                    deployment="podinfo",
                    status_code!="404",
                    direction="inbound"
                }[1m]
            )
        )
        /
        sum(
            rate(
                response_total{
                    namespace="test",
                    deployment="podinfo",
                    direction="inbound"
                }[1m]
            )
        )
        * 100

위 구성은 HTTP 404 req/sec 비율이 전체 트래픽의 3% 미만인지 확인해 canary 버전을 검증합니다. 404 비율이 3% 임계값에 도달하면 분석이 중단되고 canary는 실패로 표시됩니다.

컨테이너 이미지를 업데이트해 canary 배포를 트리거합니다:

kubectl -n test set image deployment/podinfo \
podinfod=ghcr.io/stefanprodan/podinfo:6.0.3

404를 발생시킵니다:

watch -n 1 curl http://podinfo-canary:9898/status/404

Flagger 로그를 확인합니다:

kubectl -n flagger-system logs deployment/flagger -f | jq .msg

Starting canary deployment for podinfo.test
Pre-rollout check acceptance-test passed
Advance podinfo.test canary weight 5
Halt podinfo.test advancement 404s percentage 6.20 > 3
Halt podinfo.test advancement 404s percentage 6.45 > 3
Halt podinfo.test advancement 404s percentage 7.22 > 3
Halt podinfo.test advancement 404s percentage 6.50 > 3
Halt podinfo.test advancement 404s percentage 6.34 > 3
Rolling back podinfo.test failed checks threshold reached 5
Canary failed! Scaling down podinfo.test

Slack이 구성되어 있다면 Flagger는 canary가 실패한 이유와 함께 알림을 보냅니다.

Linkerd Ingress

Flagger와 Linkerd 모두와 호환되는 ingress 컨트롤러는 두 가지가 있습니다: NGINX와 Gloo.

NGINX 설치:

helm upgrade -i nginx-ingress stable/nginx-ingress \
--namespace ingress-nginx

podinfo용 ingress 정의를 생성합니다. 이것은 들어오는 헤더를 내부 서비스 이름으로 다시 씁니다(Linkerd에 필요):

apiVersion: extensions/v1beta1
kind: Ingress
metadata:
  name: podinfo
  namespace: test
  labels:
    app: podinfo
  annotations:
    kubernetes.io/ingress.class: "nginx"
    nginx.ingress.kubernetes.io/configuration-snippet: |
      proxy_set_header l5d-dst-override $service_name.$namespace.svc.cluster.local:9898;
      proxy_hide_header l5d-remote-ip;
      proxy_hide_header l5d-server-id;
spec:
  rules:
    - host: app.example.com
      http:
        paths:
          - backend:
              serviceName: podinfo
              servicePort: 9898

ingress 컨트롤러를 사용할 때 NGINX는 메시 밖에서 실행되므로 Linkerd 트래픽 분할은 들어오는 트래픽에 적용되지 않습니다. 프론트엔드 앱에 대한 canary 분석을 실행하기 위해 Flagger는 shadow ingress를 만들고 NGINX 특정 annotation을 설정합니다.

A/B 테스트 (A/B Testing)

가중치 기반 라우팅 외에도 Flagger는 HTTP 매치 조건을 기반으로 canary로 트래픽을 라우팅하도록 구성할 수 있습니다. A/B 테스트 시나리오에서는 HTTP 헤더나 쿠키를 사용해 특정 사용자 세그먼트를 타깃팅합니다. 이는 세션 어피니티가 필요한 프론트엔드 앱에 특히 유용합니다.

podinfo canary 분석을 수정하고 provider를 nginx로 설정하고 ingress 참조를 추가하고 max/step 가중치를 제거하고 매치 조건과 iterations를 추가합니다:

apiVersion: flagger.app/v1beta1
kind: Canary
metadata:
  name: podinfo
  namespace: test
spec:
  # ingress reference
  provider: nginx
  ingressRef:
    apiVersion: extensions/v1beta1
    kind: Ingress
    name: podinfo
  targetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: podinfo
  autoscalerRef:
    apiVersion: autoscaling/v2
    kind: HorizontalPodAutoscaler
    name: podinfo
  service:
    # container port
    port: 9898
  analysis:
    interval: 1m
    threshold: 10
    iterations: 10
    match:
      # curl -H 'X-Canary: always' http://app.example.com
      - headers:
          x-canary:
            exact: "always"
      # curl -b 'canary=always' http://app.example.com
      - headers:
          cookie:
            exact: "canary"
    # Linkerd Prometheus checks
    metrics:
    - name: request-success-rate
      thresholdRange:
        min: 99
      interval: 1m
    - name: request-duration
      thresholdRange:
        max: 500
      interval: 30s
    webhooks:
      - name: acceptance-test
        type: pre-rollout
        url: http://flagger-loadtester.test/
        timeout: 30s
        metadata:
          type: bash
          cmd: "curl -sd 'test' http://podinfo-canary:9898/token | grep token"
      - name: load-test
        type: rollout
        url: http://flagger-loadtester.test/
        metadata:
          cmd: "hey -z 2m -q 10 -c 2 -H 'Cookie: canary=always' http://app.example.com"

위 구성은 canary 쿠키가 always로 설정된 사용자 또는 X-Canary: always 헤더로 서비스를 호출하는 사용자를 타깃으로 10분간 분석을 실행합니다.

참고 이제 부하 테스트는 외부 주소를 타깃으로 하고 canary 쿠키를 사용합니다.

컨테이너 이미지를 업데이트해 canary 배포를 트리거합니다:

kubectl -n test set image deployment/podinfo \
podinfod=ghcr.io/stefanprodan/podinfo:6.0.4

Flagger는 deployment 리비전이 바뀐 것을 감지하고 A/B 테스트를 시작합니다:

kubectl -n test describe canary/podinfo

Events:
 Starting canary deployment for podinfo.test
 Pre-rollout check acceptance-test passed
 Advance podinfo.test canary iteration 1/10
 Advance podinfo.test canary iteration 2/10
 Advance podinfo.test canary iteration 3/10
 Advance podinfo.test canary iteration 4/10
 Advance podinfo.test canary iteration 5/10
 Advance podinfo.test canary iteration 6/10
 Advance podinfo.test canary iteration 7/10
 Advance podinfo.test canary iteration 8/10
 Advance podinfo.test canary iteration 9/10
 Advance podinfo.test canary iteration 10/10
 Copying podinfo.test template spec to podinfo-primary.test
 Waiting for podinfo-primary.test rollout to finish: 1 of 2 updated replicas are available
 Promotion completed! Scaling down podinfo.test

위 절차는 커스텀 메트릭 검사, 웹훅, 수동 승격 승인, Slack 또는 MS Teams 알림으로 확장할 수 있습니다.

더 알아보기 (Learn more)