노드에서 Amazon FSx for Lustre 성능 최적화하기

노드에서 Amazon FSx for Lustre 성능 최적화하기 (EFA) (Optimize Amazon FSx for Lustre performance on nodes (EFA))

이 주제는 Amazon EKS 및 Amazon FSx for Lustre와 함께 EFA(Elastic Fabric Adapter) 튜닝을 설정하는 방법을 설명해요.

참고

FSx for Lustre CSI 드라이버 생성 및 배포에 대한 정보는 FSx for Lustre 드라이버 배포하기를 참고하세요.

EFA가 없는 표준 노드 최적화는 노드에서 Amazon FSx for Lustre 성능 최적화하기(비-EFA)를 참고하세요.

출처: 문서

본문

1단계: EKS 클러스터 생성

제공된 구성 파일을 사용해 클러스터를 만들어요.

# Create cluster using efa-cluster.yaml
eksctl create cluster -f efa-cluster.yaml

예시 efa-cluster.yaml:

#efa-cluster.yaml

apiVersion: eksctl.io/v1alpha5
kind: ClusterConfig

metadata:
  name: csi-fsx
  region: us-east-1
  version: "1.30"

iam:
  withOIDC: true

availabilityZones: ["us-east-1a", "us-east-1d"]

managedNodeGroups:
  - name: my-efa-ng
    instanceType: c6gn.16xlarge
    minSize: 1
    desiredCapacity: 1
    maxSize: 1
    availabilityZones: ["us-east-1b"]
    volumeSize: 300
    privateNetworking: true
    amiFamily: Ubuntu2204
    efaEnabled: true
    preBootstrapCommands:
      - |
        #!/bin/bash
        eth_intf="$(/sbin/ip -br -4 a sh | grep $(hostname -i)/ | awk '{print $1}')"
        efa_version=$(modinfo efa | awk '/^version:/ {print $2}' | sed 's/[^0-9.]//g')
        min_efa_version="2.12.1"

        if [[ "$(printf '%s\n' "$min_efa_version" "$efa_version" | sort -V | head -n1)" != "$min_efa_version" ]]; then
            sudo curl -O https://efa-installer.amazonaws.com/aws-efa-installer-1.37.0.tar.gz
            tar -xf aws-efa-installer-1.37.0.tar.gz && cd aws-efa-installer
            echo "Installing EFA driver"
            sudo apt-get update && apt-get upgrade -y
            sudo apt install -y pciutils environment-modules libnl-3-dev libnl-route-3-200 libnl-route-3-dev dkms
            sudo ./efa_installer.sh -y
            modinfo efa
        else
            echo "Using EFA driver version $efa_version"
        fi

        echo "Installing Lustre client"
        sudo wget -O - https://fsx-lustre-client-repo-public-keys.s3.amazonaws.com/fsx-ubuntu-public-key.asc | gpg --dearmor | sudo tee /usr/share/keyrings/fsx-ubuntu-public-key.gpg > /dev/null
        sudo echo "deb [signed-by=/usr/share/keyrings/fsx-ubuntu-public-key.gpg] https://fsx-lustre-client-repo.s3.amazonaws.com/ubuntu jammy main" > /etc/apt/sources.list.d/fsxlustreclientrepo.list
        sudo apt update | tail
        sudo apt install -y lustre-client-modules-$(uname -r) amazon-ec2-utils | tail
        modinfo lustre

        echo "Loading Lustre/EFA modules..."
        sudo /sbin/modprobe lnet
        sudo /sbin/modprobe kefalnd ipif_name="$eth_intf"
        sudo /sbin/modprobe ksocklnd
        sudo lnetctl lnet configure

        echo "Configuring TCP interface..."
        sudo lnetctl net del --net tcp 2> /dev/null
        sudo lnetctl net add --net tcp --if $eth_intf

        # For P5 instance type which supports 32 network cards,
        # by default add 8 EFA interfaces selecting every 4th device (1 per PCI bus)
        echo "Configuring EFA interface(s)..."
        instance_type="$(ec2-metadata --instance-type | awk '{ print $2 }')"
        num_efa_devices="$(ls -1 /sys/class/infiniband | wc -l)"
        echo "Found $num_efa_devices available EFA device(s)"

        if [[ "$instance_type" == "p5.48xlarge" || "$instance_type" == "p5e.48xlarge" ]]; then
           for intf in $(ls -1 /sys/class/infiniband | awk 'NR % 4 == 1'); do
               sudo lnetctl net add --net efa --if $intf --peer-credits 32
          done
        else
        # Other instances: Configure 2 EFA interfaces by default if the instance supports multiple network cards,
        # or 1 interface for single network card instances
        # Can be modified to add more interfaces if instance type supports it
            sudo lnetctl net add --net efa --if $(ls -1 /sys/class/infiniband | head -n1) --peer-credits 32
            if [[ $num_efa_devices -gt 1 ]]; then
               sudo lnetctl net add --net efa --if $(ls -1 /sys/class/infiniband | tail -n1) --peer-credits 32
            fi
        fi

        echo "Setting discovery and UDSP rule"
        sudo lnetctl set discovery 1
        sudo lnetctl udsp add --src efa --priority 0
        sudo /sbin/modprobe lustre

        sudo lnetctl net show
        echo "Added $(sudo lnetctl net show | grep -c '@efa') EFA interface(s)"

2단계: 노드 그룹 생성

EFA 지원 노드 그룹을 만들어요.

# Create node group using efa-ng.yaml
eksctl create nodegroup -f efa-ng.yaml

중요

# 5. Mount FSx filesystem 섹션에서 환경에 맞게 다음 값을 조정하세요.

FSX_DNS="" # Needs to be adjusted.
MOUNT_NAME="" # Needs to be adjusted.
MOUNT_POINT="" # Needs to be adjusted.

예시 efa-ng.yaml:

apiVersion: eksctl.io/v1alpha5
kind: ClusterConfig

metadata:
  name: final-efa
  region: us-east-1

managedNodeGroups:
  - name: ng-1
    instanceType: c6gn.16xlarge
    minSize: 1
    desiredCapacity: 1
    maxSize: 1
    availabilityZones: ["us-east-1a"]
    volumeSize: 300
    privateNetworking: true
    amiFamily: Ubuntu2204
    efaEnabled: true
    preBootstrapCommands:
      - |
        #!/bin/bash
        exec 1> >(logger -s -t $(basename $0)) 2>&1
        # ... (Lustre/EFA 설치, 모듈 관리, 네트워크 튜닝, FSx 마운트, 튜닝 적용 스크립트)

efa-ng.yaml의 preBootstrapCommands 스크립트는 다음을 수행해요.

  1. Lustre 클라이언트 설치 – Ubuntu를 감지하고 Lustre 저장소를 추가한 후 lustre-client-modules-$(uname -r) 및 lustre-client 패키지를 설치해요.
  2. 네트워크 및 RPC 튜닝 적용 – modprobe.conf에 ptlrpc ptlrpcd_per_cpt_max=64 및 ksocklnd credits=2560 옵션을 설정해요.
  3. Lustre 모듈 관리 – 기존 Lustre 모듈을 확인하고, 마운트된 파일 시스템이 있으면 언마운트한 후 lustre_rmmod로 제거하고, modprobe lustre로 새롭게 로드해요.
  4. Lustre 네트워킹 초기화 – lctl network up으로 Lustre 네트워킹을 초기화해요.
  5. EFA 설정 및 구성 – EFA 드라이버를 설치/확인하고, LNet 모듈(lnet, kefalnd, ksocklnd)을 로드하며, TCP 인터페이스와 EFA 인터페이스를 구성하고 discovery 및 UDSP 규칙을 설정해요. P5 인스턴스 유형은 8개의 EFA 인터페이스를, 다른 인스턴스는 1개 또는 2개의 EFA 인터페이스를 추가해요.
  6. FSx 파일 시스템 마운트 – FSX_DNS, MOUNT_NAME, MOUNT_POINT 값을 사용해 FSx 파일 시스템을 마운트해요.
  7. Lustre 성능 튜닝 적용 – LRU 튜닝(lru_max_age, lru_size=100×CPU 수), 클라이언트 캐시(max_cached_mb=64), RPC 제어(OST max_rpcs_in_flight=32, MDC max_rpcs_in_flight=64, MDC max_mod_rpcs_in_flight=50)를 적용해요.
  8. 모든 튜닝 확인 – 각 튜닝 매개변수의 실제 값을 검증하고 성공 또는 경고를 보고해요.
  9. 영속 튜닝 설정 – lustre-functions.sh 및 apply_lustre_tunings.sh 스크립트를 만들고, systemd 서비스(lustre-tunings.service)를 생성해 재부팅 후에도 튜닝이 유지되도록 해요.

(선택 사항) 3단계: EFA 설정 확인

노드에 SSH로 접속해요.

# Get instance ID from EKS console or {aws} CLI
ssh -i /path/to/your-key.pem ec2-user@...

EFA 구성을 확인해요.

sudo lnetctl net show

설정 로그를 확인해요.

sudo cat /var/log/cloud-init-output.log

lnetctl net show의 예상 출력 예시:

net:
    - net type: tcp
      ...
    - net type: efa
      local NI(s):
        - nid: xxx.xxx.xxx.xxx@efa
          status: up

예시 배포

a. claim.yaml 생성

#claim.yaml

apiVersion: v1
kind: PersistentVolumeClaim
metadata:
  name: fsx-claim-efa
spec:
  accessModes:
    - ReadWriteMany
  storageClassName: ""
  resources:
    requests:
      storage: 4800Gi
  volumeName: fsx-pv

클레임을 적용해요.

kubectl apply -f claim.yaml

b. pv.yaml 생성

``을 업데이트해요.

#pv.yaml

apiVersion: v1
kind: PersistentVolume
metadata:
  name: fsx-pv
spec:
  capacity:
    storage: 4800Gi
  volumeMode: Filesystem
  accessModes:
    - ReadWriteMany
  mountOptions:
    - flock
  persistentVolumeReclaimPolicy: Recycle
  csi:
    driver: fsx.csi.aws.com
    volumeHandle: fs-...
    volumeAttributes:
      dnsname: fs-....fsx.us-east-1.amazonaws.com
      mountname: ...

영구 볼륨을 적용해요.

kubectl apply -f pv.yaml

c. pod.yaml 생성

#pod.yaml

apiVersion: v1
kind: Pod
metadata:
  name: fsx-efa-app
spec:
  containers:
  - name: app
    image: amazonlinux:2
    command: ["/bin/sh"]
    args: ["-c", "while true; do dd if=/dev/urandom bs=100M count=20 > data/test_file; sleep 10; done"]
    resources:
      requests:
        vpc.amazonaws.com/efa: 1
      limits:
        vpc.amazonaws.com/efa: 1
    volumeMounts:
    - name: persistent-storage
      mountPath: /data
  volumes:
  - name: persistent-storage
    persistentVolumeClaim:
      claimName: fsx-claim-efa

Pod를 적용해요.

kubectl apply -f pod.yaml

추가 확인 명령

Pod가 파일 시스템을 마운트하고 쓰는지 확인해요.

kubectl exec -ti fsx-efa-app -- df -h | grep data
# Expected output:
# @tcp:/  4.5T  1.2G  4.5T   1% /data

kubectl exec -ti fsx-efa-app -- ls /data
# Expected output:
# test_file

트래픽이 EFA를 통해 가는지 확인하기 위해 노드에 SSH로 접속해요.

sudo lnetctl net show -v

예상 출력은 트래픽 통계가 있는 EFA 인터페이스를 보여줘요.

관련 정보

  • FSx for Lustre 드라이버 배포하기
  • 노드에서 Amazon FSx for Lustre 성능 최적화하기(비-EFA)
  • Amazon FSx for Lustre 성능
  • Elastic Fabric Adapter

더 알아보기 (Learn more)