레퍼런스, 커밋 & 태그 다루기
레퍼런스, 커밋, 태그는 lakeFS에서 버전을 이해하고 관리하는 데 핵심이에요. 이 가이드는 커밋 히스토리 탐색, 레퍼런스 다루기, 태그로 불변 스냅샷 만들기, 추적과 데이터 계보(lineage)를 위한 메타데이터 활용을 다룹니다.
본문
레퍼런스 이해하기
레퍼런스란?
레퍼런스는 lakeFS에서 커밋을 가리키는 어떤 포인터든 될 수 있어요:
-
브랜치(Branch): 변경 가능한(mutable) 레퍼런스예요 (새 커밋이 만들어질 때마다 함께 바뀌어요)
-
태그(Tag): 불변(immutable) 레퍼런스예요 (항상 같은 커밋을 가리켜요)
-
커밋 ID: 특정 커밋의 고유 식별자예요
-
Ref 표현식:
main~2처럼 고급 레퍼런스 문법이에요 (main보다 2커밋 이전)
레퍼런스 생성
유효한 어떤 레퍼런스든 레퍼런스 객체로 얻을 수 있어요:
import lakefs
repo = lakefs.repository("my-data-repo")
# Reference to branch head
main_ref = repo.ref("main")
print(f"Main reference: {main_ref.id}")
# Reference to specific commit
commit_ref = repo.ref("abc123def456")
print(f"Commit reference: {commit_ref.id}")
# Reference to tag
tag_ref = repo.ref("v1.0.0")
print(f"Tag reference: {tag_ref.id}")
# Advanced reference expressions
two_back = repo.ref("main~2") # Two commits before main
print(f"Two commits back: {two_back.id}")
레퍼런스에서 커밋 정보 가져오기
import lakefs
repo = lakefs.repository("my-data-repo")
ref = repo.ref("main")
# Get the underlying commit
commit = ref.get_commit()
print(f"Commit ID: {commit.id}")
print(f"Message: {commit.message}")
print(f"Committer: {commit.committer}")
print(f"Created: {commit.creation_date}")
print(f"Parents: {commit.parents}")
print(f"Metadata: {commit.metadata}")
커밋 이해하기
커밋이란?
커밋은 브랜치上的 변경 사항의 불변 스냅샷을 만들어요. 각 커밋은 고유 ID와 선택적인 메타데이터를 가집니다:
branch = lakefs.repository("my-repo").branch("main")
# Create a commit
ref = branch.commit(
message="Add new dataset",
metadata={"author": "data-team", "version": "1.0"}
)
print(f"Committed: {ref.id}")
커밋은 lakeFS 버전 컨트롤의 기본 구성 요소예요. 커밋으로 할 수 있는 일은:
-
고유 식별자로 변경 사항을 시간순으로 추적해요
-
감사(auditing)와 추적을 위해 메타데이터를 기록해요
-
데이터 계보를 위해 재현 가능한 스냅샷을 만들어요
-
누가 언제 변경했는지 파악해요
커밋 다루기
커밋 상세 정보 가져오기
특정 커밋에 대한 자세한 정보를 조회해요:
import lakefs
from datetime import datetime
repo = lakefs.repository("my-data-repo")
try:
# Get commit by ID
commit_ref = repo.commit("abc123def456xyz")
commit = commit_ref.get_commit()
print(f"Commit Details:")
print(f" ID: {commit.id}")
print(f" Message: {commit.message}")
print(f" Committer: {commit.committer}")
print(f" Timestamp: {datetime.fromtimestamp(commit.creation_date)}")
print(f" Parents: {', '.join(commit.parents) if commit.parents else 'None'}")
# Check for merge commit
if len(commit.parents) > 1:
print(f" Type: Merge commit (from {len(commit.parents)} parents)")
else:
print(f" Type: Regular commit")
except Exception as e:
print(f"Commit not found: {e}")
커밋 메타데이터 접근
커밋에 붙은 커스텀 메타데이터를 조회해요:
import lakefs
repo = lakefs.repository("my-data-repo")
branch = repo.branch("main")
# Get the latest commit
commit = branch.get_commit()
print(f"Commit: {commit.id[:8]}")
print(f"Message: {commit.message}")
if commit.metadata:
print("Metadata:")
for key, value in commit.metadata.items():
print(f" {key}: {value}")
else:
print("No metadata")
메타데이터와 함께 커밋 만들기
추적을 위해 커스텀 메타데이터를 붙여 커밋을 만들어요:
import lakefs
import json
from datetime import datetime
repo = lakefs.repository("my-data-repo")
branch = repo.branch("main")
# Upload data
branch.object("data/dataset.csv").upload(data=b"id,value\n1,100\n2,200")
# Commit with rich metadata
commit_ref = branch.commit(
message="Add customer dataset v2",
metadata={
"author": "data-team",
"version": "2.0",
"dataset-type": "raw",
"source": "database-export",
"record-count": "10000",
"timestamp": datetime.now().isoformat(),
"data-owner": "[email protected]"
}
)
print(f"Committed: {commit_ref.id}")
print(f"Metadata stored for tracking")
커밋 히스토리 탐색
커밋 목록 조회 (Log)
브랜치의 커밋 히스토리를 확인해요:
import lakefs
from datetime import datetime
repo = lakefs.repository("my-data-repo")
branch = repo.branch("main")
print("Recent commits:")
for i, commit in enumerate(branch.log(max_amount=10)):
timestamp = datetime.fromtimestamp(commit.creation_date)
print(f" {i+1}. {commit.id[:8]} - {commit.message[:40]} ({timestamp})")
메타데이터로 커밋 추적
커스텀 메타데이터를 기준으로 커밋을 찾아요:
import lakefs
def find_commits_by_metadata(repo_name, branch_name, key, value):
"""Find commits with specific metadata"""
repo = lakefs.repository(repo_name)
branch = repo.branch(branch_name)
matching_commits = []
for commit in branch.log(max_amount=1000):
if commit.metadata and commit.metadata.get(key) == value:
matching_commits.append(commit)
return matching_commits
# Usage:
commits = find_commits_by_metadata("analytics-repo", "main", "dataset-type", "clean")
print(f"Found {len(commits)} commits with dataset-type=clean")
for commit in commits[:5]:
print(f" {commit.id[:8]} - {commit.message}")
레퍼런스 비교하기 (Diff)
두 레퍼런스 사이의 Diff
임의의 두 레퍼런스 사이에 무엇이 바뀌었는지 확인해요:
import lakefs
repo = lakefs.repository("my-data-repo")
main = repo.ref("main")
dev = repo.ref("develop")
print("Changes from main to develop:")
for change in main.diff(other_ref=dev):
print(f" {change.type:10} {change.path} ({change.size_bytes} bytes)")
# Count changes
changes = list(main.diff(other_ref=dev))
print(f"\nTotal changes: {len(changes)}")
필터링이 있는 Diff
diff 결과를 경로나 변경 유형으로 필터링해요:
import lakefs
repo = lakefs.repository("my-data-repo")
tag_v1 = repo.ref("v1.0.0")
tag_v2 = repo.ref("v2.0.0")
# Get all changes
all_changes = list(tag_v1.diff(other_ref=tag_v2))
# Filter by change type
added = [c for c in all_changes if c.type == "added"]
removed = [c for c in all_changes if c.type == "removed"]
changed = [c for c in all_changes if c.type == "changed"]
print(f"Added: {len(added)}")
print(f"Removed: {len(removed)}")
print(f"Changed: {len(changed)}")
# Filter by path prefix
data_changes = [c for c in all_changes if c.path.startswith("data/")]
print(f"Changes in data/ folder: {len(data_changes)}")
크기 분석이 있는 상세 Diff
크기 정보와 함께 변경 내용을 분석해요:
import lakefs
repo = lakefs.repository("my-data-repo")
ref1 = repo.ref("commit1")
ref2 = repo.ref("commit2")
print("Detailed changes:")
for change in ref1.diff(other_ref=ref2):
size_info = f" ({change.size_bytes} bytes)" if change.size_bytes else ""
print(f" {change.type:10} {change.path}{size_info}")
태그 다루기
태그는 lakeFS에서 특정 커밋을 가리키는 불변 포인터예요. 릴리스, 데이터 버전, 중요한 스냅샷을 표시하는 데 딱 맞아요.
태그란?
태그는 특정 커밋을 중요한 지점(예: 릴리스)으로 표시해요:
import lakefs
tag = lakefs.repository("my-repo").tag("v1.0.0").create(
source_ref="main"
)
태그는 커밋을 가리키는 불변 포인터로, 다음을 할 수 있게 해줘요:
-
버저닝과 배포를 위해 릴리스를 표시해요
-
재현성과 아카이빙을 위해 스냅샷을 만들어요
-
데이터 히스토리의 중요한 지점을 참조해요
-
버전 간 데이터 계보를 추적해요
브랜치와 달리 태그는 한 번 만들면 절대 바뀌지 않아서, 안정적인 참조 지점으로 활용하기 좋아요.
태그 생성
간단한 태그 만들기
브랜치의 현재 헤드를 가리키는 태그를 만들어요:
import lakefs
repo = lakefs.repository("my-data-repo")
# Create a tag from the main branch's head
tag = repo.tag("v1.0.0").create(source_ref="main")
print(f"Created tag: v1.0.0")
print(f"Points to commit: {tag.get_commit().id}")
특정 커밋에서 태그 만들기
임의의 커밋을 가리키는 태그를 만들어요:
import lakefs
repo = lakefs.repository("my-data-repo")
main = repo.branch("main")
# Get a specific commit from history
commits = list(main.log(max_amount=10))
if commits:
# Tag an older commit
commit_to_tag = commits[0] # Most recent
tag = repo.tag("v1.0.0-rc1").create(source_ref=commit_to_tag.id)
print(f"Tagged commit: {commit_to_tag.id[:8]}")
print(f"Tag name: v1.0.0-rc1")
다른 태그에서 태그 만들기
기존 태그를 기반으로 새 태그를 만들어요:
import lakefs
repo = lakefs.repository("my-data-repo")
try:
# Create a new tag from an existing tag
existing_tag = repo.tag("v1.0.0")
new_tag = repo.tag("stable").create(source_ref=existing_tag)
print(f"New tag 'stable' points to same commit as 'v1.0.0'")
except Exception as e:
print(f"Error: {e}")
조건부 태그 생성
이미 존재하지 않는 경우에만 태그를 만들어요:
import lakefs
from lakefs.exceptions import ConflictException
repo = lakefs.repository("my-data-repo")
tag_name = "v2.0.0"
try:
# Create tag with exist_ok=False (will fail if exists)
tag = repo.tag(tag_name).create(source_ref="main", exist_ok=False)
print(f"Created new tag: {tag_name}")
except ConflictException:
print(f"Tag already exists: {tag_name}")
tag = repo.tag(tag_name)
print(f"Using existing tag: {tag.get_commit().id}")
태그 목록 조회
저장소의 모든 태그를 나열해요:
import lakefs
repo = lakefs.repository("my-data-repo")
print("All tags in repository:")
for tag in repo.tags():
commit = tag.get_commit()
print(f" {tag.id:20} -> {commit.id[:8]}... ({commit.message})")
태그 정보 가져오기
특정 태그에 대한 자세한 정보를 가져와요:
import lakefs
repo = lakefs.repository("my-data-repo")
tag = repo.tag("v1.0.0")
try:
commit = tag.get_commit()
print(f"Tag: {tag.id}")
print(f"Commit ID: {commit.id}")
print(f"Message: {commit.message}")
print(f"Committer: {commit.committer}")
print(f"Created: {commit.creation_date}")
print(f"Metadata: {commit.metadata}")
except Exception as e:
print(f"Tag not found: {e}")
태그에서 데이터 접근하기
태그된 버전의 객체 목록 조회
특정 태그 버전의 모든 객체를 나열해요:
import lakefs
repo = lakefs.repository("my-data-repo")
tag_ref = repo.ref("v1.0.0") # Use ref() for tag access
# List all objects in this tag
print(f"Objects in v1.0.0:")
for obj in tag_ref.objects():
print(f" {obj.path}")
# List specific prefix
print(f"\nModels in v1.0.0:")
for obj in tag_ref.objects(prefix="models/"):
if hasattr(obj, 'path'): # It's a file, not a folder
print(f" {obj.path} ({obj.size_bytes} bytes)")
태그된 버전에서 데이터 읽기
특정 태그에서 객체 내용을 읽어요:
import lakefs
import csv
import io
repo = lakefs.repository("my-data-repo")
tag_ref = repo.ref("v1.0.0")
# Read a CSV file from the tag
try:
obj = tag_ref.object("data/dataset.csv")
with obj.reader(mode='r') as f:
reader = csv.reader(f)
headers = next(reader)
print(f"Headers: {headers}")
for row in reader:
print(f" {row}")
except Exception as e:
print(f"Error reading file: {e}")
태그된 버전 간 데이터 비교
두 태그 버전 사이에 무엇이 바뀌었는지 비교해요:
import lakefs
repo = lakefs.repository("my-data-repo")
tag_v1 = repo.ref("v1.0.0")
tag_v2 = repo.ref("v2.0.0")
# See what changed
print("Changes from v1.0.0 to v2.0.0:")
for change in tag_v1.diff(other_ref=tag_v2):
print(f" {change.type:10} {change.path}")
# Count change types
changes = list(tag_v1.diff(other_ref=tag_v2))
added = len([c for c in changes if c.type == "added"])
removed = len([c for c in changes if c.type == "removed"])
changed = len([c for c in changes if c.type == "changed"])
print(f"\nSummary: +{added} -{removed} ~{changed}")
태그 삭제
단일 태그 삭제
더 이상 필요 없는 태그를 제거해요:
import lakefs
repo = lakefs.repository("my-data-repo")
try:
tag = repo.tag("old-release")
tag.delete()
print("Tag deleted: old-release")
except Exception as e:
print(f"Delete failed: {e}")
커밋 관계
머지 커밋 식별하기
머지 커밋을 찾고 분석해요:
import lakefs
repo = lakefs.repository("my-data-repo")
branch = repo.branch("main")
print("Merge commits:")
for i, commit in enumerate(branch.log(max_amount=50)):
if len(commit.parents) > 1:
print(f" {commit.id[:8]} - Merged {len(commit.parents)} branches")
print(f" Message: {commit.message}")
print(f" Parents: {', '.join([p[:8] for p in commit.parents])}")
커밋 계보 추적하기
부모를 따라가며 커밋을 거슬러 올라가요. 커밋을 더 잘 이해하기 위한 방법이고, 변경 사항을 추적할 때는 log 연산을 사용하는 쪽을 권장해요:
import lakefs
def trace_ancestry(repo_name, commit_id, depth=5):
"""Trace commit ancestry up to specified depth"""
repo = lakefs.repository(repo_name)
ancestry = []
current_id = commit_id
for level in range(depth):
try:
commit_ref = repo.commit(current_id)
commit = commit_ref.get_commit()
ancestry.append({
"level": level,
"commit_id": commit.id[:8],
"message": commit.message,
"parents": commit.parents
})
# Move to first parent
if commit.parents:
current_id = commit.parents[0]
else:
break
except Exception as e:
print(f"Error at level {level}: {e}")
break
return ancestry
# Usage:
ancestry = trace_ancestry("my-repo", "abc123def456", depth=5)
print("Commit Ancestry:")
for entry in ancestry:
indent = " " * entry["level"]
print(f"{indent}└─ {entry['commit_id']} - {entry['message']}")
실전 워크플로
ML 모델 릴리스 워크플로
학습된 모델을 버저닝해서 릴리스해요:
import lakefs
import json
def release_ml_model(repo_name, model_version, model_metrics):
"""
Create a versioned release of an ML model
"""
repo = lakefs.repository(repo_name)
try:
# Create release tag
tag_name = f"model-v{model_version}"
tag = repo.tag(tag_name).create(source_ref="main")
commit = tag.get_commit()
print(f"ML Model Released: {tag_name}")
print(f" Commit: {commit.id[:8]}")
# Read model metadata from tagged version
tag_ref = repo.ref(tag_name)
try:
with tag_ref.object("models/metadata.json").reader() as f:
metadata = json.load(f)
print(f" Model: {metadata.get('name')}")
print(f" Framework: {metadata.get('framework')}")
print(f" Version: {metadata.get('version')}")
except:
print(" (No metadata file)")
# Store release info
release_info = {
"version": model_version,
"commit": commit.id,
"metrics": model_metrics,
"tag": tag_name
}
return release_info
except Exception as e:
print(f"Model release failed: {e}")
return None
# Usage:
metrics = {
"accuracy": 0.945,
"precision": 0.92,
"recall": 0.96,
"f1": 0.939
}
release_info = release_ml_model("ml-repo", "3.2.0", metrics)
if release_info:
print(f"\nModel released and can be retrieved from tag: {release_info['tag']}")
프로덕션 배포 워크플로
프로덕션 데이터 버전을 관리해요:
import lakefs
def promote_to_production(repo_name, from_tag, environment):
"""
Promote a tagged version to production by creating an environment tag
"""
repo = lakefs.repository(repo_name)
try:
# Create environment-specific tag
env_tag_name = f"prod-{environment}"
# Delete old environment tag if it exists
try:
old_tag = repo.tag(env_tag_name)
old_tag.delete()
print(f"Removed old {env_tag_name} tag")
except:
pass # Tag didn't exist
# Create new environment tag pointing to the same commit as version tag
source_tag = repo.tag(from_tag)
env_tag = repo.tag(env_tag_name).create(source_ref=source_tag)
print(f"Promoted to {environment}")
print(f" Source: {from_tag}")
print(f" Target: {env_tag_name}")
print(f" Commit: {env_tag.get_commit().id[:8]}")
return env_tag_name
except Exception as e:
print(f"Promotion failed: {e}")
return None
# Usage:
env_tag = promote_to_production("prod-repo", "v2.1.0", "us-west-1")
if env_tag:
print(f"Production data updated to use {env_tag}")
에러 처리
자주 발생하는 레퍼런스 에러 다루기
import lakefs
from lakefs.exceptions import NotFoundException
repo = lakefs.repository("my-data-repo")
# Reference doesn't exist
try:
ref = repo.ref("non-existent-ref")
commit = ref.get_commit()
except NotFoundException:
print("Reference not found")
# Commit doesn't exist
try:
commit_ref = repo.commit("nonexistent123")
commit = commit_ref.get_commit()
except NotFoundException:
print("Commit not found")
# Invalid reference expression
try:
ref = repo.ref("main~1000") # Try to get 1000 commits back
commit = ref.get_commit()
except NotFoundException:
print("Reference expression invalid or out of range")
태그 에러 다루기
import lakefs
from lakefs.exceptions import ConflictException, NotFoundException
repo = lakefs.repository("my-data-repo")
# Tag already exists
try:
tag = repo.tag("v1.0.0").create(source_ref="main", exist_ok=False)
except ConflictException:
print("Tag already exists")
tag = repo.tag("v1.0.0")
# Tag doesn't exist
try:
tag = repo.tag("non-existent-tag")
commit = tag.get_commit()
except NotFoundException:
print("Tag not found")
더 알아보기 (Learn more)
공식 문서의 원문은 https://docs.lakefs.io/reference/python/refs/ 에서 확인할 수 있어요.