Knative 카나리아 배포
Knative 카나리아 배포 (Knative Canary Deployments)
Knative와 Flagger를 함께 사용해서 canary 배포를 자동화하는 방법을 이 문서에서 알려드릴게요. Knative Revision 사이에서 트래픽을 분산하며 canary 분석과 승격을 진행하는 구성을 살펴봅시다.
출처: 문서
본문
사전 준비 (Prerequisites)
Flagger는 Kubernetes 클러스터 v1.19 이상과 serving.knative.dev/v1를 API 버전으로 지원하는 리소스를 포함한 Knative Serving 설치가 필요합니다.
Knative v1.17.0을 설치합니다:
kubectl apply -f https://github.com/knative/serving/releases/download/knative-v1.17.0/serving-crds.yaml
kubectl apply -f https://github.com/knative/serving/releases/download/knative-v1.17.0/serving-core.yaml
kubectl apply -f https://github.com/knative/net-kourier/releases/download/knative-v1.17.0/kourier.yaml
kubectl patch configmap/config-network \
--namespace knative-serving \
--type merge \
--patch '{"data":{"ingress-class":"kourier.ingress.networking.knative.dev"}}'
flagger-system 네임스페이스에 Flagger를 설치합니다:
kubectl apply -k github.com/fluxcd/flagger//kustomize/knative
Knative Service용 네임스페이스를 생성합니다:
kubectl create namespace test
podinfo를 배포하는 Knative Service를 생성합니다:
apiVersion: serving.knative.dev/v1
kind: Service
metadata:
name: podinfo
namespace: test
spec:
template:
spec:
containers:
- image: ghcr.io/stefanprodan/podinfo:6.0.0
ports:
- containerPort: 9898
protocol: TCP
command:
- ./podinfo
- --port=9898
- --port-metrics=9797
- --grpc-port=9999
- --grpc-service-name=podinfo
- --level=info
- --random-delay=false
- --random-error=false
canary 분석 중 트래픽을 발생시킬 부하 테스트 서비스를 배포합니다:
kubectl apply -k https://github.com/fluxcd/flagger//kustomize/tester?ref=main
Canary 커스텀 리소스를 생성합니다:
apiVersion: flagger.app/v1beta1
kind: Canary
metadata:
name: podinfo
namespace: test
spec:
provider: knative
# knative service ref
targetRef:
apiVersion: serving.knative.dev/v1
kind: Service
name: podinfo
# the maximum time in seconds for the canary deployment
# to make progress before it is rollback (default 600s)
progressDeadlineSeconds: 60
analysis:
# schedule interval (default 60s)
interval: 15s
# max number of failed metric checks before rollback
threshold: 15
# max traffic percentage routed to canary
maxWeight: 50
# canary increment step
# percentage (0-100)
stepWeight: 10
metrics:
- name: request-success-rate
# min success rate (non-5xx responses)
# percentage (0-100)
thresholdRange:
min: 99
interval: 1m
- name: request-duration
# milliseconds
thresholdRange:
max: 500
interval: 1m
webhooks:
- name: load-test
url: http://flagger-loadtester.test/
timeout: 5s
metadata:
type: cmd
cmd: "hey -z 1m -q 5 -c 2 http://podinfo.test"
logCmdOutput: "true"
참고:
.spec.provider가knative로 설정된 Canary 리소스는.spec.targetRef.kind가Service이고.spec.targetRef.apiVersion이serving.knative.dev/v1인 경우에만 유효합니다.
위 리소스를 podinfo-canary.yaml로 저장한 뒤 적용합니다:
kubectl apply -f ./podinfo-canary.yaml
canary 분석이 시작되면 Flagger는 canary로 트래픽을 라우팅하기 전에 pre-rollout 웹훅을 호출합니다. canary 분석은 매분 HTTP 메트릭과 롤아웃 훅을 검증하면서 5분간 실행됩니다.
몇 초 후 Flagger는 Knative Service podinfo에 다음 변경을 가합니다:
flagger.app/primary-revision이름의 annotation을 오브젝트에 추가합니다.- primary와 canary Knative Revision 사이의 트래픽 분산을 조작할 수 있도록 오브젝트의
.spec.traffic섹션을 수정합니다.
자동화된 canary 승격 (Automated canary promotion)
컨테이너 이미지를 업데이트해 canary 배포를 트리거합니다:
kubectl -n test patch services.serving podinfo --type=json \
-p '[{"op": "replace", "path": "/spec/template/spec/containers/0/image", "value": "ghcr.io/stefanprodan/podinfo:6.0.1"}]'
Flagger는 deployment 리비전이 바뀐 것을 감지하고 새로운 롤아웃을 시작합니다:
kubectl -n test describe canary/podinfo
Status:
Canary Weight: 0
Failed Checks: 0
Phase: Succeeded
Events:
Type Reason Age From Message
---- ------ ---- ---- -------
Normal Synced 3m flagger New revision detected podinfo.test
Normal Synced 3m flagger Scaling up podinfo.test
Normal Synced 3m flagger Advance podinfo.test canary weight 5
Normal Synced 3m flagger Advance podinfo.test canary weight 10
Normal Synced 3m flagger Advance podinfo.test canary weight 15
Normal Synced 2m flagger Advance podinfo.test canary weight 20
Normal Synced 2m flagger Advance podinfo.test canary weight 25
Normal Synced 1m flagger Advance podinfo.test canary weight 30
Normal Synced 1m flagger Advance podinfo.test canary weight 35
Normal Synced 55s flagger Advance podinfo.test canary weight 40
Normal Synced 45s flagger Advance podinfo.test canary weight 45
Normal Synced 35s flagger Advance podinfo.test canary weight 50
Normal Synced 25s flagger Copying podinfo.test template spec to podinfo-primary.test
Normal Synced 5s flagger Promotion completed! Scaling down podinfo.test
새 Knative Revision이 생성될 때마다 canary 배포가 트리거됩니다.
canary 분석 중 Knative Service에 새 변경사항을 적용하면 Flagger가 분석을 다시 시작한다는 점을 참고하세요.
Flagger가 Knative Revision 사이 트래픽을 분산하도록 Knative Service 오브젝트를 점진적으로 변경하는 것을 모니터링할 수 있습니다:
watch kubectl get httproute -n test podinfo -o=jsonpath='{.spec.traffic}'
모든 canary를 다음 명령으로 모니터링할 수 있습니다:
watch kubectl get canaries --all-namespaces
NAMESPACE NAME STATUS WEIGHT LASTTRANSITIONTIME
test podinfo Progressing 15 2025-03-16T14:05:07Z
prod frontend Succeeded 0 2025-03-16T16:15:07Z
prod backend Failed 0 2025-03-16T17:05:07Z
자동화된 롤백 (Automated rollback)
canary 분석 중 HTTP 500 오류와 높은 지연을 발생시켜 Flagger가 롤아웃을 일시 중지하는지 테스트할 수 있습니다.
또 다른 canary 배포를 트리거합니다:
kubectl -n test patch services.serving podinfo --type=json \
-p '[{"op": "replace", "path": "/spec/template/spec/containers/0/image", "value": "ghcr.io/stefanprodan/podinfo:6.0.2"}]'
부하 테스터 pod에 접속합니다:
kubectl -n test exec -it flagger-loadtester-xx-xx sh
HTTP 500 오류를 발생시킵니다:
watch curl http://podinfo-canary:9898/status/500
지연을 발생시킵니다:
watch curl http://podinfo-canary:9898/delay/1
실패한 검사 횟수가 canary 분석 임계값에 도달하면 트래픽이 primary Knative Revision으로 되돌아가고 롤아웃은 실패로 표시됩니다.
kubectl -n test describe canary/podinfo
Status:
Canary Weight: 0
Failed Checks: 10
Phase: Failed
Events:
Type Reason Age From Message
---- ------ ---- ---- -------
Normal Synced 3m flagger Starting canary deployment for podinfo.test
Normal Synced 3m flagger Advance podinfo.test canary weight 5
Normal Synced 3m flagger Advance podinfo.test canary weight 10
Normal Synced 3m flagger Advance podinfo.test canary weight 15
Normal Synced 3m flagger Halt podinfo.test advancement error rate 69.17% > 1%
Normal Synced 2m flagger Halt podinfo.test advancement error rate 61.39% > 1%
Normal Synced 2m flagger Halt podinfo.test advancement error rate 55.06% > 1%
Normal Synced 2m flagger Halt podinfo.test advancement error rate 47.00% > 1%
Normal Synced 2m flagger (combined from similar events): Halt podinfo.test advancement error rate 38.08% > 1%
Warning Synced 1m flagger Rolling back podinfo.test failed checks threshold reached 10
Warning Synced 1m flagger Canary failed! Scaling down podinfo.test