본문 바로가기
WIKI 기술 지식 베이스

Apache APISIX 카나리아 배포

원문 보기 위키 갱신

Apache APISIX 카나리아 배포 (Apache APISIX Canary Deployments)

Apache APISIX와 Flagger를 함께 사용해서 canary 배포를 자동화하는 방법을 이 문서에서 알려드릴게요. APISIX의 ApisixRoute를 활용해 canary 트래픽을 점진적으로 전환하며 배포를 검증하는 구성을 살펴봅시다.

출처: 문서

본문

사전 준비 (Prerequisites)

Flagger는 Kubernetes 클러스터 v1.19 이상, Apache APISIX v2.15 이상, Apache APISIX Ingress Controller v1.5.0 이상이 필요합니다.

Helm v3로 Apache APISIX와 Apache APISIX Ingress Controller를 설치합니다:

helm repo add apisix https://charts.apiseven.com
kubectl create ns apisix

helm upgrade -i apisix apisix/apisix --version=0.11.3 \
--namespace apisix \
--set apisix.podAnnotations."prometheus\.io/scrape"=true \
--set apisix.podAnnotations."prometheus\.io/port"=9091 \
--set apisix.podAnnotations."prometheus\.io/path"=/apisix/prometheus/metrics \
--set pluginAttrs.prometheus.export_addr.ip=0.0.0.0 \
--set pluginAttrs.prometheus.export_addr.port=9091 \
--set pluginAttrs.prometheus.export_uri=/apisix/prometheus/metrics \
--set pluginAttrs.prometheus.metric_prefix=apisix_ \
--set ingress-controller.enabled=true \
--set ingress-controller.config.apisix.serviceNamespace=apisix

Apache APISIX와 같은 네임스페이스에 Flagger와 Prometheus 애드온을 설치합니다:

helm repo add flagger https://flagger.app

helm upgrade -i flagger flagger/flagger \
--namespace apisix \
--set prometheus.install=true \
--set meshProvider=apisix

부트스트랩 (Bootstrap)

Flagger는 Kubernetes deployment와 선택적으로 horizontal pod autoscaler(HPA)를 받아서 일련의 오브젝트(Kubernetes deployments, ClusterIP services, ApisixRoute)를 생성합니다. 이 오브젝트들은 앱을 클러스터 외부에 노출하며 canary 분석과 승격을 진행시킵니다.

테스트 네임스페이스를 생성합니다:

kubectl create ns test

deployment와 horizontal pod autoscaler를 생성합니다:

kubectl apply -k https://github.com/fluxcd/flagger//kustomize/podinfo?ref=main

canary 분석 중 트래픽을 발생시킬 부하 테스트 서비스를 배포합니다:

helm upgrade -i flagger-loadtester flagger/loadtester \
--namespace=test

Apache APISIX ApisixRoute를 생성합니다. Flagger는 이를 참조하여 canary ApisixRoute를 생성합니다(app.example.com을 자신의 도메인으로 바꾸세요):

apiVersion: apisix.apache.org/v2
kind: ApisixRoute
metadata:
  name: podinfo
  namespace: test
spec:
  http:
    - backends:
        - serviceName: podinfo
          servicePort: 80
      match:
        hosts:
          - app.example.com
        methods:
          - GET
        paths:
          - /*
      name: method
      plugins:
        - name: prometheus
          enable: true
          config:
            disable: false
            prefer_name: true

위 리소스를 podinfo-apisixroute.yaml로 저장한 뒤 적용합니다:

kubectl apply -f ./podinfo-apisixroute.yaml

canary 커스텀 리소스를 생성합니다(app.example.com을 자신의 도메인으로 바꾸세요):

apiVersion: flagger.app/v1beta1
kind: Canary
metadata:
  name: podinfo
  namespace: test
spec:
  provider: apisix
  targetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: podinfo
  # apisix route reference
  routeRef:
    apiVersion: apisix.apache.org/v2
    kind: ApisixRoute
    name: podinfo
  # the maximum time in seconds for the canary deployment
  # to make progress before it is rollback (default 600s)
  progressDeadlineSeconds: 60
  service:
    # ClusterIP port number
    port: 80
    # container port number or name
    targetPort: 9898
  analysis:
    # schedule interval (default 60s)
    interval: 10s
    # max number of failed metric checks before rollback
    threshold: 10
    # max traffic percentage routed to canary
    # percentage (0-100)
    maxWeight: 50
    # canary increment step
    # percentage (0-100)
    stepWeight: 10
    # APISIX Prometheus checks
    metrics:
      - name: request-success-rate
        # minimum req success rate (non 5xx responses)
        # percentage (0-100)
        thresholdRange:
          min: 99
        interval: 1m
      - name: request-duration
        # builtin Prometheus check
        # maximum req duration P99
        # milliseconds
        thresholdRange:
          max: 500
        interval: 30s
    webhooks:
      - name: load-test
        url: http://flagger-loadtester.test/
        timeout: 5s
        type: rollout
        metadata:
          cmd: |-
            hey -z 1m -q 10 -c 2 -h2 -host app.example.com http://apisix-gateway.apisix/api/info

위 리소스를 podinfo-canary.yaml로 저장한 뒤 적용합니다:

kubectl apply -f ./podinfo-canary.yaml

몇 초 후 Flagger가 canary 오브젝트들을 생성합니다:

# applied 
deployment.apps/podinfo
horizontalpodautoscaler.autoscaling/podinfo
apisixroute/podinfo
canary.flagger.app/podinfo

# generated 
deployment.apps/podinfo-primary
horizontalpodautoscaler.autoscaling/podinfo-primary
service/podinfo
service/podinfo-canary
service/podinfo-primary
apisixroute/podinfo-podinfo-canary

자동화된 canary 승격 (Automated canary promotion)

Flagger는 HTTP 요청 성공률, 요청 평균 지속 시간, pod 상태 같은 주요 성능 지표(KPI)를 측정하면서 canary로 트래픽을 점진적으로 이동시키는 컨트롤 루프를 구현합니다. KPI 분석 결과에 따라 canary는 승격되거나 중단되며, 분석 결과는 Slack이나 MS Teams로 게시됩니다.

컨테이너 이미지를 업데이트해 canary 배포를 트리거합니다:

kubectl -n test set image deployment/podinfo \
podinfod=stefanprodan/podinfo:6.0.1

Flagger는 deployment 리비전이 바뀐 것을 감지하고 새로운 롤아웃을 시작합니다:

kubectl -n test describe canary/podinfo

Status:
  Canary Weight:  0
  Conditions:
    Message:               Canary analysis completed successfully, promotion finished.
    Reason:                Succeeded
    Status:                True
    Type:                  Promoted
  Failed Checks:           1
  Iterations:              0
  Phase:                   Succeeded

Events:
  Type     Reason  Age                    From     Message
  ----     ------  ----                   ----     -------
  Warning  Synced  2m59s                  flagger  podinfo-primary.test not ready: waiting for rollout to finish: observed deployment generation less than desired generation
  Warning  Synced  2m50s                  flagger  podinfo-primary.test not ready: waiting for rollout to finish: 0 of 1 (readyThreshold 100%) updated replicas are available
  Normal   Synced  2m40s (x3 over 2m59s)  flagger  all the metrics providers are available!
  Normal   Synced  2m39s                  flagger  Initialization done! podinfo.test
  Normal   Synced  2m20s                  flagger  New revision detected! Scaling up podinfo.test
  Warning  Synced  2m (x2 over 2m10s)     flagger  canary deployment podinfo.test not ready: waiting for rollout to finish: 0 of 1 (readyThreshold 100%) updated replicas are available
  Normal   Synced  110s                   flagger  Starting canary analysis for podinfo.test
  Normal   Synced  109s                   flagger  Advance podinfo.test canary weight 10
  Warning  Synced  100s                   flagger  Halt advancement no values found for apisix metric request-success-rate probably podinfo.test is not receiving traffic: running query failed: no values found
  Normal   Synced  90s                    flagger  Advance podinfo.test canary weight 20
  Normal   Synced  80s                    flagger  Advance podinfo.test canary weight 30
  Normal   Synced  69s                    flagger  Advance podinfo.test canary weight 40
  Normal   Synced  59s                    flagger  Advance podinfo.test canary weight 50
  Warning  Synced  30s (x2 over 40s)      flagger  podinfo-primary.test not ready: waiting for rollout to finish: 1 old replicas are pending termination
  Normal   Synced  9s (x3 over 50s)       flagger  (combined from similar events): Promotion completed! Scaling down podinfo.test

canary 분석 중 deployment에 새 변경사항을 적용하면 Flagger가 분석을 다시 시작한다는 점을 참고하세요.

모든 canary를 다음 명령으로 모니터링할 수 있습니다:

watch kubectl get canaries --all-namespaces

NAMESPACE   NAME      STATUS      WEIGHT   LASTTRANSITIONTIME
test        podinfo-2   Progressing   10       2022-11-23T05:00:54Z
test        podinfo     Succeeded     0        2022-11-23T06:00:54Z

자동화된 롤백 (Automated rollback)

canary 분석 중 HTTP 500 오류를 발생시켜 Flagger가 결함 버전을 일시 중지하고 롤백하는지 테스트할 수 있습니다.

또 다른 canary 배포를 트리거합니다:

kubectl -n test set image deployment/podinfo \
podinfod=stefanprodan/podinfo:6.0.2

부하 테스터 pod에 접속합니다:

kubectl -n test exec -it deploy/flagger-loadtester bash

HTTP 500 오류를 발생시킵니다:

hey -z 1m -c 5 -q 5 -host app.example.com http://apisix-gateway.apisix/status/500

지연을 발생시킵니다:

watch -n 1 curl -H \"host: app.example.com\" http://apisix-gateway.apisix/delay/1

실패한 검사 횟수가 canary 분석 임계값에 도달하면 트래픽이 primary로 되돌아가고, canary는 0으로 스케일 다운되며 롤아웃은 실패로 표시됩니다.

kubectl -n apisix logs deploy/flagger -f | jq .msg

"New revision detected! Scaling up podinfo.test"
"canary deployment podinfo.test not ready: waiting for rollout to finish: 0 of 1 (readyThreshold 100%) updated replicas are available"
"Starting canary analysis for podinfo.test"
"Advance podinfo.test canary weight 10"
"Halt podinfo.test advancement success rate 0.00% < 99%"
"Halt podinfo.test advancement success rate 26.76% < 99%"
"Halt podinfo.test advancement success rate 34.19% < 99%"
"Halt podinfo.test advancement success rate 37.32% < 99%"
"Halt podinfo.test advancement success rate 39.04% < 99%"
"Halt podinfo.test advancement success rate 40.13% < 99%"
"Halt podinfo.test advancement success rate 48.28% < 99%"
"Halt podinfo.test advancement success rate 50.35% < 99%"
"Halt podinfo.test advancement success rate 56.92% < 99%"
"Halt podinfo.test advancement success rate 67.70% < 99%"
"Rolling back podinfo.test failed checks threshold reached 10"
"Canary failed! Scaling down podinfo.test"

커스텀 메트릭 (Custom metrics)

canary 분석은 Prometheus 쿼리로 확장할 수 있습니다.

메트릭 템플릿을 생성하고 클러스터에 적용합니다:

apiVersion: flagger.app/v1beta1
kind: MetricTemplate
metadata:
  name: not-found-percentage
  namespace: test
spec:
  provider:
    type: prometheus
    address: http://flagger-prometheus.apisix:9090
  query: |
    sum(
      rate(
        apisix_http_status{
          route=~"{{ namespace }}_{{ route }}-{{ target }}-canary_.+",
          code!~"4.."
        }[{{ interval }}]
      )
    )
    /
    sum(
      rate(
        apisix_http_status{
          route=~"{{ namespace }}_{{ route }}-{{ target }}-canary_.+"
        }[{{ interval }}]
      )
    ) * 100

canary 분석을 수정하고 404 오류율 검사를 추가합니다:

  analysis:
    metrics:
      - name: "404s percentage"
        templateRef:
          name: not-found-percentage
        thresholdRange:
          max: 5
        interval: 1m

위 구성은 HTTP 404 req/sec 비율이 전체 트래픽의 5% 미만인지 확인해 canary를 검증합니다. 404 비율이 5% 임계값에 도달하면 canary는 실패합니다.

위 절차는 더 많은 커스텀 메트릭 검사, 웹훅, 수동 승격 승인, Slack 또는 MS Teams 알림으로 확장할 수 있습니다.

더 알아보기 (Learn more)