본문 바로가기
WIKI 기술 지식 베이스

Istio 카나리아 배포

원문 보기 위키 갱신

Istio 카나리아 배포 (Istio Canary Deployments)

Istio와 Flagger를 함께 사용해서 canary 배포를 자동화하는 방법을 이 문서에서 알려드릴게요. Istio의 Virtual Service와 Destination Rule을 활용해 트래픽을 점진적으로 전환하며 배포를 검증하는 구성을 살펴봅시다.

출처: 문서

본문

사전 준비 (Prerequisites)

Flagger는 Kubernetes 클러스터 v1.16 이상과 Istio v1.5 이상이 필요합니다.

텔레메트리 지원과 Prometheus가 있는 Istio를 설치합니다:

istioctl manifest install --set profile=default

# Suggestion: Please change release-1.8 in below command, to your real istio version.
kubectl apply -f https://raw.githubusercontent.com/istio/istio/release-1.18/samples/addons/prometheus.yaml

istio-system 네임스페이스에 Flagger를 설치합니다:

kubectl apply -k github.com/fluxcd/flagger//kustomize/istio

데모 앱을 메시 밖으로 노출할 ingress gateway를 생성합니다:

apiVersion: networking.istio.io/v1alpha3
kind: Gateway
metadata:
  name: public-gateway
  namespace: istio-system
spec:
  selector:
    istio: ingressgateway
  servers:
    - port:
        number: 80
        name: http
        protocol: HTTP
      hosts:
        - "*"

부트스트랩 (Bootstrap)

Flagger는 Kubernetes deployment와 선택적으로 horizontal pod autoscaler(HPA)를 받아서 일련의 오브젝트(Kubernetes deployments, ClusterIP services, Istio destination rules, virtual services)를 생성합니다. 이 오브젝트들은 메시 안에서 앱을 노출하며 canary 분석과 승격을 진행시킵니다.

Istio 사이드카 주입이 활성화된 테스트 네임스페이스를 생성합니다:

kubectl create ns test
kubectl label namespace test istio-injection=enabled

deployment와 horizontal pod autoscaler를 생성합니다:

kubectl apply -k https://github.com/fluxcd/flagger//kustomize/podinfo?ref=main

canary 분석 중 트래픽을 발생시킬 부하 테스트 서비스를 배포합니다:

kubectl apply -k https://github.com/fluxcd/flagger//kustomize/tester?ref=main

canary 커스텀 리소스를 생성합니다(example.com을 자신의 도메인으로 바꾸세요):

apiVersion: flagger.app/v1beta1
kind: Canary
metadata:
  name: podinfo
  namespace: test
spec:
  # deployment reference
  targetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: podinfo
  # the maximum time in seconds for the canary deployment
  # to make progress before it is rollback (default 600s)
  progressDeadlineSeconds: 60
  # HPA reference (optional)
  autoscalerRef:
    apiVersion: autoscaling/v2
    kind: HorizontalPodAutoscaler
    name: podinfo
  service:
    # service port number
    port: 9898
    # container port number or name (optional)
    targetPort: 9898
    # Istio gateways (optional)
    gateways:
    - istio-system/public-gateway
    # Istio virtual service host names (optional)
    hosts:
    - app.example.com
    # Istio traffic policy (optional)
    trafficPolicy:
      tls:
        # use ISTIO_MUTUAL when mTLS is enabled
        mode: DISABLE
    # Istio retry policy (optional)
    retries:
      attempts: 3
      perTryTimeout: 1s
      retryOn: "gateway-error,connect-failure,refused-stream"
  analysis:
    # schedule interval (default 60s)
    interval: 1m
    # max number of failed metric checks before rollback
    threshold: 5
    # max traffic percentage routed to canary
    # percentage (0-100)
    maxWeight: 50
    # canary increment step
    # percentage (0-100)
    stepWeight: 10
    metrics:
    - name: request-success-rate
      # minimum req success rate (non 5xx responses)
      # percentage (0-100)
      thresholdRange:
        min: 99
      interval: 1m
    - name: request-duration
      # maximum req duration P99
      # milliseconds
      thresholdRange:
        max: 500
      interval: 30s
    # testing (optional)
    webhooks:
      - name: acceptance-test
        type: pre-rollout
        url: http://flagger-loadtester.test/
        timeout: 30s
        metadata:
          type: bash
          cmd: "curl -sd 'test' http://podinfo-canary:9898/token | grep token"
      - name: load-test
        url: http://flagger-loadtester.test/
        timeout: 5s
        metadata:
          cmd: "hey -z 1m -q 10 -c 2 http://podinfo-canary.test:9898/"

Istio 1.4를 사용할 때는 request-duration을 메트릭 템플릿으로 바꿔야 한다는 점을 참고하세요.

위 리소스를 podinfo-canary.yaml로 저장한 뒤 적용합니다:

kubectl apply -f ./podinfo-canary.yaml

canary 분석이 시작되면 Flagger는 canary로 트래픽을 라우팅하기 전에 pre-rollout 웹훅을 호출합니다. canary 분석은 매분 HTTP 메트릭과 롤아웃 훅을 검증하면서 5분간 실행됩니다.

몇 초 후 Flagger가 canary 오브젝트들을 생성합니다:

# applied 
deployment.apps/podinfo
horizontalpodautoscaler.autoscaling/podinfo
canary.flagger.app/podinfo

# generated 
deployment.apps/podinfo-primary
horizontalpodautoscaler.autoscaling/podinfo-primary
service/podinfo
service/podinfo-canary
service/podinfo-primary
destinationrule.networking.istio.io/podinfo-canary
destinationrule.networking.istio.io/podinfo-primary
virtualservice.networking.istio.io/podinfo

자동화된 canary 승격 (Automated canary promotion)

컨테이너 이미지를 업데이트해 canary 배포를 트리거합니다:

kubectl -n test set image deployment/podinfo \
podinfod=ghcr.io/stefanprodan/podinfo:6.0.1

Flagger는 deployment 리비전이 바뀐 것을 감지하고 새로운 롤아웃을 시작합니다:

kubectl -n test describe canary/podinfo

Status:
  Canary Weight:         0
  Failed Checks:         0
  Phase:                 Succeeded
Events:
  Type     Reason  Age   From     Message
  ----     ------  ----  ----     -------
  Normal   Synced  3m    flagger  New revision detected podinfo.test
  Normal   Synced  3m    flagger  Scaling up podinfo.test
  Warning  Synced  3m    flagger  Waiting for podinfo.test rollout to finish: 0 of 1 updated replicas are available
  Normal   Synced  3m    flagger  Advance podinfo.test canary weight 5
  Normal   Synced  3m    flagger  Advance podinfo.test canary weight 10
  Normal   Synced  3m    flagger  Advance podinfo.test canary weight 15
  Normal   Synced  2m    flagger  Advance podinfo.test canary weight 20
  Normal   Synced  2m    flagger  Advance podinfo.test canary weight 25
  Normal   Synced  1m    flagger  Advance podinfo.test canary weight 30
  Normal   Synced  1m    flagger  Advance podinfo.test canary weight 35
  Normal   Synced  55s   flagger  Advance podinfo.test canary weight 40
  Normal   Synced  45s   flagger  Advance podinfo.test canary weight 45
  Normal   Synced  35s   flagger  Advance podinfo.test canary weight 50
  Normal   Synced  25s   flagger  Copying podinfo.test template spec to podinfo-primary.test
  Warning  Synced  15s   flagger  Waiting for podinfo-primary.test rollout to finish: 1 of 2 updated replicas are available
  Normal   Synced  5s    flagger  Promotion completed! Scaling down podinfo.test

canary 분석 중 deployment에 새 변경사항을 적용하면 Flagger가 분석을 다시 시작한다는 점을 참고하세요.

canary 배포는 다음 오브젝트 중 하나가 변경되면 트리거됩니다:

  • Deployment PodSpec (컨테이너 이미지, command, ports, env, resources 등)
  • 볼륨으로 마운트되거나 환경 변수에 매핑된 ConfigMaps
  • 볼륨으로 마운트되거나 환경 변수에 매핑된 Secrets

모든 canary를 다음 명령으로 모니터링할 수 있습니다:

watch kubectl get canaries --all-namespaces

NAMESPACE   NAME      STATUS        WEIGHT   LASTTRANSITIONTIME
test        podinfo   Progressing   15       2019-01-16T14:05:07Z
prod        frontend  Succeeded     0        2019-01-15T16:15:07Z
prod        backend   Failed        0        2019-01-14T17:05:07Z

자동화된 롤백 (Automated rollback)

canary 분석 중 HTTP 500 오류와 높은 지연을 발생시켜 Flagger가 롤아웃을 일시 중지하는지 테스트할 수 있습니다.

또 다른 canary 배포를 트리거합니다:

kubectl -n test set image deployment/podinfo \
podinfod=ghcr.io/stefanprodan/podinfo:6.0.2

부하 테스터 pod에 접속합니다:

kubectl -n test exec -it flagger-loadtester-xx-xx sh

HTTP 500 오류를 발생시킵니다:

watch curl http://podinfo-canary:9898/status/500

지연을 발생시킵니다:

watch curl http://podinfo-canary:9898/delay/1

실패한 검사 횟수가 canary 분석 임계값에 도달하면 트래픽이 primary로 되돌아가고, canary는 0으로 스케일 다운되며 롤아웃은 실패로 표시됩니다.

kubectl -n test describe canary/podinfo

Status:
  Canary Weight:         0
  Failed Checks:         10
  Phase:                 Failed
Events:
  Type     Reason  Age   From     Message
  ----     ------  ----  ----     -------
  Normal   Synced  3m    flagger  Starting canary deployment for podinfo.test
  Normal   Synced  3m    flagger  Advance podinfo.test canary weight 5
  Normal   Synced  3m    flagger  Advance podinfo.test canary weight 10
  Normal   Synced  3m    flagger  Advance podinfo.test canary weight 15
  Normal   Synced  3m    flagger  Halt podinfo.test advancement success rate 69.17% < 99%
  Normal   Synced  2m    flagger  Halt podinfo.test advancement success rate 61.39% < 99%
  Normal   Synced  2m    flagger  Halt podinfo.test advancement success rate 55.06% < 99%
  Normal   Synced  2m    flagger  Halt podinfo.test advancement success rate 47.00% < 99%
  Normal   Synced  2m    flagger  (combined from similar events): Halt podinfo.test advancement success rate 38.08% < 99%
  Warning  Synced  1m    flagger  Rolling back podinfo.test failed checks threshold reached 10
  Warning  Synced  1m    flagger  Canary failed! Scaling down podinfo.test

세션 어피니티 (Session Affinity)

Flagger는 가중치 기반 라우팅과 A/B 테스트를 개별적으로 수행할 수 있지만, Istio와 함께라면 둘을 결합해 세션 어피니티가 있는 canary 릴리스를 만들 수 있습니다. 자세한 내용은 배포 전략 문서를 읽어보세요.

canary 커스텀 리소스를 생성합니다(app.example.com을 자신의 도메인으로 바꾸세요):

apiVersion: flagger.app/v1beta1
kind: Canary
metadata:
  name: podinfo
  namespace: test
spec:
  # deployment reference
  targetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: podinfo
  # the maximum time in seconds for the canary deployment
  # to make progress before it is rollback (default 600s)
  progressDeadlineSeconds: 60
  # HPA reference (optional)
  autoscalerRef:
    apiVersion: autoscaling/v2
    kind: HorizontalPodAutoscaler
    name: podinfo
  service:
    # service port number
    port: 9898
    # container port number or name (optional)
    targetPort: 9898
    # Istio gateways (optional)
    gateways:
    - istio-system/public-gateway
    # Istio virtual service host names (optional)
    hosts:
    - app.example.com
    # Istio traffic policy (optional)
    trafficPolicy:
      tls:
        # use ISTIO_MUTUAL when mTLS is enabled
        mode: DISABLE
    # Istio retry policy (optional)
    retries:
      attempts: 3
      perTryTimeout: 1s
      retryOn: "gateway-error,connect-failure,refused-stream"
  analysis:
    # schedule interval (default 60s)
    interval: 1m
    # max number of failed metric checks before rollback
    threshold: 5
    # max traffic percentage routed to canary
    # percentage (0-100)
    maxWeight: 50
    # canary increment step
    # percentage (0-100)
    stepWeight: 10
    # session affinity config
    sessionAffinity:
      # name of the cookie used
      cookieName: flagger-cookie
      # max age of the cookie (in seconds)
      # optional; defaults to 86400
      maxAge: 21600
    metrics:
    - name: request-success-rate
      # minimum req success rate (non 5xx responses)
      # percentage (0-100)
      thresholdRange:
        min: 99
      interval: 1m
    - name: request-duration
      # maximum req duration P99
      # milliseconds
      thresholdRange:
        max: 500
      interval: 30s
    # testing (optional)
    webhooks:
      - name: acceptance-test
        type: pre-rollout
        url: http://flagger-loadtester.test/
        timeout: 30s
        metadata:
          type: bash
          cmd: "curl -sd 'test' http://podinfo-canary:9898/token | grep token"
      - name: load-test
        url: http://flagger-loadtester.test/
        timeout: 5s
        metadata:
          cmd: "hey -z 1m -q 10 -c 2 http://podinfo-canary.test:9898/"

위 리소스를 podinfo-canary-session-affinity.yaml로 저장한 뒤 적용합니다:

kubectl apply -f ./podinfo-canary-session-affinity.yaml

컨테이너 이미지를 업데이트해 canary 배포를 트리거합니다:

kubectl -n test set image deployment/podinfo \
podinfod=ghcr.io/stefanprodan/podinfo:6.0.1

브라우저에서 app.example.com을 로드하고 podinfo:6.0.1이 요청을 처리하는 것을 볼 때까지 새로고침할 수 있습니다. 이후의 모든 요청은 Flagger가 Istio로 구성한 세션 어피니티 덕분에 podinfo:6.0.0이 아니라 podinfo:6.0.1이 처리합니다.

트래픽 미러링 (Traffic mirroring)

읽기 작업을 수행하는 앱의 경우 Flagger는 트래픽 미러링으로 canary 릴리스를 진행하도록 구성할 수 있습니다. Istio 트래픽 미러링은 들어오는 각 요청을 복사해 primary로 하나, canary 서비스로 하나를 보냅니다. primary의 응답은 사용자에게 돌아가고 canary의 응답은 버려집니다. 두 요청 모두에 대해 메트릭이 수집되므로 canary 메트릭이 임계값 안에 있을 때만 배포가 진행됩니다.

미러링은 멱등성(idempotent) 이 있거나 두 번 처리할 수 있는(primary 한 번, canary 한 번) 요청에 사용해야 한다는 점을 참고하세요.

stepWeight/maxWeight를 iterations로 바꾸고 analysis.mirror를 true로 설정하면 미러링을 활성화할 수 있습니다:

apiVersion: flagger.app/v1beta1
kind: Canary
metadata:
  name: podinfo
  namespace: test
spec:
  analysis:
    # schedule interval
    interval: 1m
    # max number of failed metric checks before rollback
    threshold: 5
    # total number of iterations
    iterations: 10
    # enable traffic shadowing 
    mirror: true
    # weight of the traffic mirrored to your canary (defaults to 100%)
    mirrorWeight: 100
    metrics:
    - name: request-success-rate
      thresholdRange:
        min: 99
      interval: 1m
    - name: request-duration
      thresholdRange:
        max: 500
      interval: 1m
    webhooks:
      - name: acceptance-test
        type: pre-rollout
        url: http://flagger-loadtester.test/
        timeout: 30s
        metadata:
          type: bash
          cmd: "curl -sd 'test' http://podinfo-canary:9898/token | grep token"
      - name: load-test
        url: http://flagger-loadtester.test/
        timeout: 5s
        metadata:
          cmd: "hey -z 1m -q 10 -c 2 http://podinfo.test:9898/"

위 구성으로 Flagger는 다음 단계로 canary 릴리스를 진행합니다:

  • 새 리비전 감지(deployment spec, secrets 또는 configmaps 변경)
  • canary 배포를 0에서 스케일 업
  • HPA가 canary 최소 레플리카를 설정할 때까지 대기
  • canary pod 상태 확인
  • 인수 테스트(acceptance tests) 실행
  • 테스트가 실패하면 canary 릴리스 중단
  • 부하 테스트 시작
  • primary에서 canary로 트래픽의 100% 미러링
  • 매분 요청 성공률과 요청 지속 시간 확인
  • 메트릭 검사 실패 임계값에 도달하면 canary 릴리스 중단
  • 반복 횟수에 도달하면 트래픽 미러링 중지
  • 실제 트래픽을 canary pod로 라우팅
  • canary 승격(primary secrets, configmaps, deployment spec 업데이트)
  • primary 배포 롤아웃이 끝날 때까지 대기
  • HPA가 primary 최소 레플리카를 설정할 때까지 대기
  • primary pod 상태 확인
  • 실제 트래픽을 primary로 다시 전환
  • canary를 0으로 스케일 다운
  • canary 분석 결과 알림 전송

위 절차는 커스텀 메트릭 검사, 웹훅, 수동 승격 승인, Slack 또는 MS Teams 알림으로 확장할 수 있습니다.

TCP 서비스용 Canary 배포 (Canary Deployments for TCP Services)

TCP(비 HTTP) 서비스에서 Canary 배포를 수행하는 것은 HTTP Canary와 거의 동일합니다. Gateway 문서를 TCP 라우팅을 지원하도록 업데이트하는 것 외에도, Canary 문서의 service 섹션 안에서 appProtocol 필드를 TCP로 설정해야 한다는 점만 다릅니다.

예시:

apiVersion: networking.istio.io/v1alpha3
kind: Gateway
metadata:
  name: public-gateway
  namespace: istio-system
spec:
  selector:
    istio: ingressgateway
  servers:
    - port:
        number: 7070
        name: tcp-service
        protocol: TCP # <== set the protocol to tcp here
      hosts:
        - "*"
apiVersion: flagger.app/v1beta1
kind: Canary
# omitted for brevity
spec:
  service:
    port: 7070
    appProtocol: TCP # <== set the appProtocol here
    targetPort: 7070
    portName: "tcp-service-port"

appProtocol이 TCP와 같으면 Flagger는 이를 TCP 서비스에 대한 Canary 배포로 취급합니다. VirtualService 문서를 만들 때 primary와 canary 서비스 사이에 요청을 라우팅하는 TCP 섹션을 추가합니다. 이 스펙에 대한 자세한 내용은 Istio 문서를 참고하세요.

결과 VirtualService는 아래와 유사한 tcp 섹션을 포함합니다:

tcp:
  - route:
    - destination:
        host: tcp-service-primary
        port:
          number: 7070
      weight: 100
    - destination:
        host: tcp-service-canary
        port:
          number: 7070
      weight: 0

Canary 분석이 시작되면 Flagger는 이 tcp 섹션 안의 가중치를 조정해 Canary 배포를 진행합니다. 오류가 나면(그리고 중단되거나) 분석 끝에 성공적으로 도달해 승격될 때까지 계속 진행합니다.

appProtocol을 TCP 외의 다른 값(예: HTTP)으로 설정하면 HTTP 서비스로 취급하여 Canary를 수행한다는 점도 중요합니다. appProtocol을 전혀 설정하지 않아도 마찬가지입니다. appProtocal이 TCP와 같은 경우에 만 Canary를 TCP 서비스로 취급합니다.

더 알아보기 (Learn more)