본문 바로가기
WIKI 기술 지식 베이스

lakeFS와 Vertex AI 함께 사용하기

원문 보기 위키 갱신

Vertex AI는 Google Cloud 사용자가 어떤 사용 사례든 완전 관리형 ML 도구로 머신러닝(ML) 모델을 더 빠르게 구축·배포·확장할 수 있게 해줘요. lakeFS는 GCS 버킷 위에 저장소를 만들고, Dataset API로 lakeFS 버전 위에 관리형 Dataset을 만들거나, Cloud Storage Mounts가 읽을 수 있는 형태로 lakeFS 오브젝트 버전을 자동 내보내는 방식으로 Vertex AI와 함께 동작해요.

출처: 문서

본문

Vertex AI를 쓰면 Google Cloud 사용자가 모든 사용 사례에 맞는 완전 관리형 ML 도구로 머신러닝(ML) 모델을 더 빠르게 구축하고, 배포하고, 확장할 수 있어요.

lakeFS는 사용자가 GCS 버킷 위에 저장소를 만들고, lakeFS 버전 위에 Dataset API로 관리형 Dataset을 만들거나, Cloud Storage Mounts가 읽을 수 있는 형태로 lakeFS 오브젝트 버전을 자동으로 내보내는 방식으로 Vertex AI와 함께 동작해요.

lakeFS를 Vertex Managed Datasets와 함께 사용하기

Vertex의 ImageDataset과 VideoDataset은 gcs에서 CSV 파일을 임포트해 데이터셋을 만들 수 있게 해줘요(gcs_source 참고).

이 CSV 파일에는 이미지 파일의 GCS 주소와 그에 대응하는 라벨이 담겨 있어요.

lakeFS API는 버전 관리되는 오브젝트의 실제 GCS 주소를 내보내는 기능을 지원하기 때문에, 데이터셋을 만들 때 이런 CSV 파일을 생성할 수 있어요:

#!/usr/bin/env python

# Requirements:
# google-cloud-aiplatform>=1.31.0
# lakefs>=1.0.0

import csv
from pathlib import PosixPath
from io import StringIO

import lakefs
from google.cloud import storage
from google.cloud import aiplatform

# Dataset configuration
lakefs_repo = 'my-repository'
lakefs_ref = 'main'
img_dataset = 'datasets/my-images/'

# Vertex configuration
import_bucket = 'underlying-gcs-bucket'

# produce import file for Vertex's SDK
buf = StringIO()
csv_writer = csv.writer(buf)
for obj in lakefs.repository(lakefs_repo).ref(lakefs_ref).objects(prefix=img_dataset):
    p = PosixPath(obj.path)
    csv_writer.writerow((obj.physical_address, p.parent.name))

# spit out CSV
print('Generated path and labels CSV')
buf.seek(0)

# Write it to storage
storage_client = storage.Client()
bucket = storage_client.bucket(import_bucket)
blob = bucket.blob(f'vertex/imports/{lakefs_repo}/{lakefs_ref}/labels.csv')
with blob.open('w') as out:
    out.write(buf.read())

print(f'Wrote CSV to gs://{import_bucket}/vertex/imports/{lakefs_repo}/{lakefs_ref}/labels.csv')

# import in Vertex, as dataset
print('Importing dataset...')
ds = aiplatform.ImageDataset.create(
    display_name=f'{lakefs_repo}_{lakefs_ref}_imgs',
    gcs_source=f'gs://{import_bucket}/vertex/imports/{lakefs_repo}/{lakefs_ref}/labels.csv',
    import_schema_uri=aiplatform.schema.dataset.ioformat.image.single_label_classification,
    sync=True
)
ds.wait()
print(f'Done! {ds.display_name} ({ds.resource_name})')

lakeFS를 Cloud Storage Fuse와 함께 사용하기

Vertex는 Google Cloud Storage를 Fuse Filesystem으로 마운트해 학습 잡의 커스텀 입력으로 쓰는 걸 허용해요.

소비하려는 버전마다 lakeFS 파일을 복사하는 대신, gcsfuse의 네이티브 symlink inode로 심볼릭 링크를 만들 수 있어요.

이 과정은 예제 gcsfuse_symlink_exporter.lua Lua 훅으로 완전 자동화할 수 있어요.

해야 할 일은 다음과 같아요:

  • 예제 .lua 파일을 lakeFS 저장소에 업로드하세요. 이 예제에서는 scripts/gcsfuse_symlink_exporter.lua 아래에 두겠어요.

  • 새 훅 정의 파일을 만들어 _lakefs_actions/export_images.yaml로 업로드하세요:

---
# Example hook declaration: (_lakefs_actions/export_images.yaml):
name: export_images

on:
  post-commit:
    branches: ["main"]
  post-merge:
    branches: ["main"]
  post-create-tag:

hooks:
- id: gcsfuse_export_images
  type: lua
  properties:
    script_path: scripts/export_gcs_fuse.lua  # Path to the script we uploaded in the previous step
    args:
      prefix: "datasets/images/"  # Path we want to export every commit
      destination: "gs://my-bucket/exports/my-repo/"  # Where should we create the symlinks?
      mount:
        from: "gs://my-bucket/repos/my-repo/"  # Symlinks are to a unix-mounted file
        to: "/gcs/my-bucket/repos/my-repo/"    #  This will ensure they point to a location that exists.

      # Should be the contents of a valid credentials.json file
      # See: https://developers.google.com/workspace/guides/create-credentials
      # Will be used to write the symlink files
      gcs_credentials_json_string: |
        {
          "client_id": "...",
          "client_secret": "...",
          "refresh_token": "...",
          "type": "..."
        }

끝! 다음에 태그가 만들어지거나 main 브랜치에 업데이트가 일어나면, datasets/images/의 lakeFS 버전이 마운트 가능한 위치로 자동 내보내져요.

심볼릭 링크가 걸린 파일을 소비하려면, 마운트에서 평소처럼 읽으면 돼요:

with open('/gcs/my-bucket/exports/my-repo/branches/main/datasets/images/001.jpg') as f:
    image_data = f.read()

과거에 내보냈던 커밋도 읽을 수 있어요:

commit_id = 'abcdef123deadbeef567'
with open(f'/gcs/my-bucket/exports/my-repo/commits/{commit_id}/datasets/images/001.jpg') as f:
    image_data = f.read()

Cloud Storage Fuse와 함께 lakeFS를 쓸 때 유의할 점

lakeFS 경로가 gcsfuse에서 읽히려면 마운트 옵션 --implicit-dirs를 반드시 지정해야 해요.

더 알아보기 (Learn more)

공식 문서의 자세한 내용은 https://docs.lakefs.io/integrations/vertex_ai/에서 확인하실 수 있어요.