본문 바로가기
WIKI 기술 지식 베이스

레퍼런스, 커밋 & 태그 다루기

원문 보기 위키 갱신

레퍼런스, 커밋, 태그는 lakeFS에서 버전을 이해하고 관리하는 데 핵심이에요. 이 가이드는 커밋 히스토리 탐색, 레퍼런스 다루기, 태그로 불변 스냅샷 만들기, 추적과 데이터 계보(lineage)를 위한 메타데이터 활용을 다룹니다.

출처: 레퍼런스, 커밋 & 태그 다루기

본문

레퍼런스 이해하기

레퍼런스란?

레퍼런스는 lakeFS에서 커밋을 가리키는 어떤 포인터든 될 수 있어요:

  • 브랜치(Branch): 변경 가능한(mutable) 레퍼런스예요 (새 커밋이 만들어질 때마다 함께 바뀌어요)

  • 태그(Tag): 불변(immutable) 레퍼런스예요 (항상 같은 커밋을 가리켜요)

  • 커밋 ID: 특정 커밋의 고유 식별자예요

  • Ref 표현식: main~2처럼 고급 레퍼런스 문법이에요 (main보다 2커밋 이전)

레퍼런스 생성

유효한 어떤 레퍼런스든 레퍼런스 객체로 얻을 수 있어요:

import lakefs

repo = lakefs.repository("my-data-repo")

# Reference to branch head
main_ref = repo.ref("main")
print(f"Main reference: {main_ref.id}")

# Reference to specific commit
commit_ref = repo.ref("abc123def456")
print(f"Commit reference: {commit_ref.id}")

# Reference to tag
tag_ref = repo.ref("v1.0.0")
print(f"Tag reference: {tag_ref.id}")

# Advanced reference expressions
two_back = repo.ref("main~2")  # Two commits before main
print(f"Two commits back: {two_back.id}")

레퍼런스에서 커밋 정보 가져오기

import lakefs

repo = lakefs.repository("my-data-repo")
ref = repo.ref("main")

# Get the underlying commit
commit = ref.get_commit()

print(f"Commit ID: {commit.id}")
print(f"Message: {commit.message}")
print(f"Committer: {commit.committer}")
print(f"Created: {commit.creation_date}")
print(f"Parents: {commit.parents}")
print(f"Metadata: {commit.metadata}")

커밋 이해하기

커밋이란?

커밋은 브랜치上的 변경 사항의 불변 스냅샷을 만들어요. 각 커밋은 고유 ID와 선택적인 메타데이터를 가집니다:

branch = lakefs.repository("my-repo").branch("main")

# Create a commit
ref = branch.commit(
    message="Add new dataset",
    metadata={"author": "data-team", "version": "1.0"}
)
print(f"Committed: {ref.id}")

커밋은 lakeFS 버전 컨트롤의 기본 구성 요소예요. 커밋으로 할 수 있는 일은:

  • 고유 식별자로 변경 사항을 시간순으로 추적해요

  • 감사(auditing)와 추적을 위해 메타데이터를 기록해요

  • 데이터 계보를 위해 재현 가능한 스냅샷을 만들어요

  • 누가 언제 변경했는지 파악해요

커밋 다루기

커밋 상세 정보 가져오기

특정 커밋에 대한 자세한 정보를 조회해요:

import lakefs
from datetime import datetime

repo = lakefs.repository("my-data-repo")

try:
    # Get commit by ID
    commit_ref = repo.commit("abc123def456xyz")
    commit = commit_ref.get_commit()

    print(f"Commit Details:")
    print(f"  ID: {commit.id}")
    print(f"  Message: {commit.message}")
    print(f"  Committer: {commit.committer}")
    print(f"  Timestamp: {datetime.fromtimestamp(commit.creation_date)}")
    print(f"  Parents: {', '.join(commit.parents) if commit.parents else 'None'}")

    # Check for merge commit
    if len(commit.parents) > 1:
        print(f"  Type: Merge commit (from {len(commit.parents)} parents)")
    else:
        print(f"  Type: Regular commit")

except Exception as e:
    print(f"Commit not found: {e}")

커밋 메타데이터 접근

커밋에 붙은 커스텀 메타데이터를 조회해요:

import lakefs

repo = lakefs.repository("my-data-repo")
branch = repo.branch("main")

# Get the latest commit
commit = branch.get_commit()

print(f"Commit: {commit.id[:8]}")
print(f"Message: {commit.message}")

if commit.metadata:
    print("Metadata:")
    for key, value in commit.metadata.items():
        print(f"  {key}: {value}")
else:
    print("No metadata")

메타데이터와 함께 커밋 만들기

추적을 위해 커스텀 메타데이터를 붙여 커밋을 만들어요:

import lakefs
import json
from datetime import datetime

repo = lakefs.repository("my-data-repo")
branch = repo.branch("main")

# Upload data
branch.object("data/dataset.csv").upload(data=b"id,value\n1,100\n2,200")

# Commit with rich metadata
commit_ref = branch.commit(
    message="Add customer dataset v2",
    metadata={
        "author": "data-team",
        "version": "2.0",
        "dataset-type": "raw",
        "source": "database-export",
        "record-count": "10000",
        "timestamp": datetime.now().isoformat(),
        "data-owner": "[email protected]"
    }
)

print(f"Committed: {commit_ref.id}")
print(f"Metadata stored for tracking")

커밋 히스토리 탐색

커밋 목록 조회 (Log)

브랜치의 커밋 히스토리를 확인해요:

import lakefs
from datetime import datetime

repo = lakefs.repository("my-data-repo")
branch = repo.branch("main")

print("Recent commits:")
for i, commit in enumerate(branch.log(max_amount=10)):
    timestamp = datetime.fromtimestamp(commit.creation_date)
    print(f"  {i+1}. {commit.id[:8]} - {commit.message[:40]} ({timestamp})")

메타데이터로 커밋 추적

커스텀 메타데이터를 기준으로 커밋을 찾아요:

import lakefs

def find_commits_by_metadata(repo_name, branch_name, key, value):
    """Find commits with specific metadata"""
    repo = lakefs.repository(repo_name)
    branch = repo.branch(branch_name)

    matching_commits = []

    for commit in branch.log(max_amount=1000):
        if commit.metadata and commit.metadata.get(key) == value:
            matching_commits.append(commit)

    return matching_commits

# Usage:
commits = find_commits_by_metadata("analytics-repo", "main", "dataset-type", "clean")
print(f"Found {len(commits)} commits with dataset-type=clean")

for commit in commits[:5]:
    print(f"  {commit.id[:8]} - {commit.message}")

레퍼런스 비교하기 (Diff)

두 레퍼런스 사이의 Diff

임의의 두 레퍼런스 사이에 무엇이 바뀌었는지 확인해요:

import lakefs

repo = lakefs.repository("my-data-repo")
main = repo.ref("main")
dev = repo.ref("develop")

print("Changes from main to develop:")
for change in main.diff(other_ref=dev):
    print(f"  {change.type:10} {change.path} ({change.size_bytes} bytes)")

# Count changes
changes = list(main.diff(other_ref=dev))
print(f"\nTotal changes: {len(changes)}")

필터링이 있는 Diff

diff 결과를 경로나 변경 유형으로 필터링해요:

import lakefs

repo = lakefs.repository("my-data-repo")
tag_v1 = repo.ref("v1.0.0")
tag_v2 = repo.ref("v2.0.0")

# Get all changes
all_changes = list(tag_v1.diff(other_ref=tag_v2))

# Filter by change type
added = [c for c in all_changes if c.type == "added"]
removed = [c for c in all_changes if c.type == "removed"]
changed = [c for c in all_changes if c.type == "changed"]

print(f"Added: {len(added)}")
print(f"Removed: {len(removed)}")
print(f"Changed: {len(changed)}")

# Filter by path prefix
data_changes = [c for c in all_changes if c.path.startswith("data/")]
print(f"Changes in data/ folder: {len(data_changes)}")

크기 분석이 있는 상세 Diff

크기 정보와 함께 변경 내용을 분석해요:

import lakefs

repo = lakefs.repository("my-data-repo")
ref1 = repo.ref("commit1")
ref2 = repo.ref("commit2")

print("Detailed changes:")
for change in ref1.diff(other_ref=ref2):
    size_info = f" ({change.size_bytes} bytes)" if change.size_bytes else ""
    print(f"  {change.type:10} {change.path}{size_info}")

태그 다루기

태그는 lakeFS에서 특정 커밋을 가리키는 불변 포인터예요. 릴리스, 데이터 버전, 중요한 스냅샷을 표시하는 데 딱 맞아요.

태그란?

태그는 특정 커밋을 중요한 지점(예: 릴리스)으로 표시해요:

import lakefs

tag = lakefs.repository("my-repo").tag("v1.0.0").create(
    source_ref="main"
)

태그는 커밋을 가리키는 불변 포인터로, 다음을 할 수 있게 해줘요:

  • 버저닝과 배포를 위해 릴리스를 표시해요

  • 재현성과 아카이빙을 위해 스냅샷을 만들어요

  • 데이터 히스토리의 중요한 지점을 참조해요

  • 버전 간 데이터 계보를 추적해요

브랜치와 달리 태그는 한 번 만들면 절대 바뀌지 않아서, 안정적인 참조 지점으로 활용하기 좋아요.

태그 생성

간단한 태그 만들기

브랜치의 현재 헤드를 가리키는 태그를 만들어요:

import lakefs

repo = lakefs.repository("my-data-repo")

# Create a tag from the main branch's head
tag = repo.tag("v1.0.0").create(source_ref="main")

print(f"Created tag: v1.0.0")
print(f"Points to commit: {tag.get_commit().id}")
특정 커밋에서 태그 만들기

임의의 커밋을 가리키는 태그를 만들어요:

import lakefs

repo = lakefs.repository("my-data-repo")
main = repo.branch("main")

# Get a specific commit from history
commits = list(main.log(max_amount=10))

if commits:
    # Tag an older commit
    commit_to_tag = commits[0]  # Most recent
    tag = repo.tag("v1.0.0-rc1").create(source_ref=commit_to_tag.id)

    print(f"Tagged commit: {commit_to_tag.id[:8]}")
    print(f"Tag name: v1.0.0-rc1")
다른 태그에서 태그 만들기

기존 태그를 기반으로 새 태그를 만들어요:

import lakefs

repo = lakefs.repository("my-data-repo")

try:
    # Create a new tag from an existing tag
    existing_tag = repo.tag("v1.0.0")
    new_tag = repo.tag("stable").create(source_ref=existing_tag)

    print(f"New tag 'stable' points to same commit as 'v1.0.0'")

except Exception as e:
    print(f"Error: {e}")
조건부 태그 생성

이미 존재하지 않는 경우에만 태그를 만들어요:

import lakefs
from lakefs.exceptions import ConflictException

repo = lakefs.repository("my-data-repo")
tag_name = "v2.0.0"

try:
    # Create tag with exist_ok=False (will fail if exists)
    tag = repo.tag(tag_name).create(source_ref="main", exist_ok=False)
    print(f"Created new tag: {tag_name}")

except ConflictException:
    print(f"Tag already exists: {tag_name}")
    tag = repo.tag(tag_name)
    print(f"Using existing tag: {tag.get_commit().id}")

태그 목록 조회

저장소의 모든 태그를 나열해요:

import lakefs

repo = lakefs.repository("my-data-repo")

print("All tags in repository:")
for tag in repo.tags():
    commit = tag.get_commit()
    print(f"  {tag.id:20} -> {commit.id[:8]}... ({commit.message})")

태그 정보 가져오기

특정 태그에 대한 자세한 정보를 가져와요:

import lakefs

repo = lakefs.repository("my-data-repo")
tag = repo.tag("v1.0.0")

try:
    commit = tag.get_commit()

    print(f"Tag: {tag.id}")
    print(f"Commit ID: {commit.id}")
    print(f"Message: {commit.message}")
    print(f"Committer: {commit.committer}")
    print(f"Created: {commit.creation_date}")
    print(f"Metadata: {commit.metadata}")

except Exception as e:
    print(f"Tag not found: {e}")

태그에서 데이터 접근하기

태그된 버전의 객체 목록 조회

특정 태그 버전의 모든 객체를 나열해요:

import lakefs

repo = lakefs.repository("my-data-repo")
tag_ref = repo.ref("v1.0.0")  # Use ref() for tag access

# List all objects in this tag
print(f"Objects in v1.0.0:")
for obj in tag_ref.objects():
    print(f"  {obj.path}")

# List specific prefix
print(f"\nModels in v1.0.0:")
for obj in tag_ref.objects(prefix="models/"):
    if hasattr(obj, 'path'):  # It's a file, not a folder
        print(f"  {obj.path} ({obj.size_bytes} bytes)")

태그된 버전에서 데이터 읽기

특정 태그에서 객체 내용을 읽어요:

import lakefs
import csv
import io

repo = lakefs.repository("my-data-repo")
tag_ref = repo.ref("v1.0.0")

# Read a CSV file from the tag
try:
    obj = tag_ref.object("data/dataset.csv")

    with obj.reader(mode='r') as f:
        reader = csv.reader(f)
        headers = next(reader)
        print(f"Headers: {headers}")

        for row in reader:
            print(f"  {row}")

except Exception as e:
    print(f"Error reading file: {e}")

태그된 버전 간 데이터 비교

두 태그 버전 사이에 무엇이 바뀌었는지 비교해요:

import lakefs

repo = lakefs.repository("my-data-repo")
tag_v1 = repo.ref("v1.0.0")
tag_v2 = repo.ref("v2.0.0")

# See what changed
print("Changes from v1.0.0 to v2.0.0:")
for change in tag_v1.diff(other_ref=tag_v2):
    print(f"  {change.type:10} {change.path}")

# Count change types
changes = list(tag_v1.diff(other_ref=tag_v2))
added = len([c for c in changes if c.type == "added"])
removed = len([c for c in changes if c.type == "removed"])
changed = len([c for c in changes if c.type == "changed"])

print(f"\nSummary: +{added} -{removed} ~{changed}")

태그 삭제

단일 태그 삭제

더 이상 필요 없는 태그를 제거해요:

import lakefs

repo = lakefs.repository("my-data-repo")

try:
    tag = repo.tag("old-release")
    tag.delete()
    print("Tag deleted: old-release")

except Exception as e:
    print(f"Delete failed: {e}")

커밋 관계

머지 커밋 식별하기

머지 커밋을 찾고 분석해요:

import lakefs

repo = lakefs.repository("my-data-repo")
branch = repo.branch("main")

print("Merge commits:")
for i, commit in enumerate(branch.log(max_amount=50)):
    if len(commit.parents) > 1:
        print(f"  {commit.id[:8]} - Merged {len(commit.parents)} branches")
        print(f"    Message: {commit.message}")
        print(f"    Parents: {', '.join([p[:8] for p in commit.parents])}")

커밋 계보 추적하기

부모를 따라가며 커밋을 거슬러 올라가요. 커밋을 더 잘 이해하기 위한 방법이고, 변경 사항을 추적할 때는 log 연산을 사용하는 쪽을 권장해요:

import lakefs

def trace_ancestry(repo_name, commit_id, depth=5):
    """Trace commit ancestry up to specified depth"""
    repo = lakefs.repository(repo_name)
    ancestry = []

    current_id = commit_id

    for level in range(depth):
        try:
            commit_ref = repo.commit(current_id)
            commit = commit_ref.get_commit()

            ancestry.append({
                "level": level,
                "commit_id": commit.id[:8],
                "message": commit.message,
                "parents": commit.parents
            })

            # Move to first parent
            if commit.parents:
                current_id = commit.parents[0]
            else:
                break

        except Exception as e:
            print(f"Error at level {level}: {e}")
            break

    return ancestry

# Usage:
ancestry = trace_ancestry("my-repo", "abc123def456", depth=5)
print("Commit Ancestry:")
for entry in ancestry:
    indent = "  " * entry["level"]
    print(f"{indent}└─ {entry['commit_id']} - {entry['message']}")

실전 워크플로

ML 모델 릴리스 워크플로

학습된 모델을 버저닝해서 릴리스해요:

import lakefs
import json

def release_ml_model(repo_name, model_version, model_metrics):
    """
    Create a versioned release of an ML model
    """
    repo = lakefs.repository(repo_name)

    try:
        # Create release tag
        tag_name = f"model-v{model_version}"
        tag = repo.tag(tag_name).create(source_ref="main")

        commit = tag.get_commit()

        print(f"ML Model Released: {tag_name}")
        print(f"  Commit: {commit.id[:8]}")

        # Read model metadata from tagged version
        tag_ref = repo.ref(tag_name)

        try:
            with tag_ref.object("models/metadata.json").reader() as f:
                metadata = json.load(f)
                print(f"  Model: {metadata.get('name')}")
                print(f"  Framework: {metadata.get('framework')}")
                print(f"  Version: {metadata.get('version')}")
        except:
            print("  (No metadata file)")

        # Store release info
        release_info = {
            "version": model_version,
            "commit": commit.id,
            "metrics": model_metrics,
            "tag": tag_name
        }

        return release_info

    except Exception as e:
        print(f"Model release failed: {e}")
        return None

# Usage:
metrics = {
    "accuracy": 0.945,
    "precision": 0.92,
    "recall": 0.96,
    "f1": 0.939
}

release_info = release_ml_model("ml-repo", "3.2.0", metrics)
if release_info:
    print(f"\nModel released and can be retrieved from tag: {release_info['tag']}")

프로덕션 배포 워크플로

프로덕션 데이터 버전을 관리해요:

import lakefs

def promote_to_production(repo_name, from_tag, environment):
    """
    Promote a tagged version to production by creating an environment tag
    """
    repo = lakefs.repository(repo_name)

    try:
        # Create environment-specific tag
        env_tag_name = f"prod-{environment}"

        # Delete old environment tag if it exists
        try:
            old_tag = repo.tag(env_tag_name)
            old_tag.delete()
            print(f"Removed old {env_tag_name} tag")
        except:
            pass  # Tag didn't exist

        # Create new environment tag pointing to the same commit as version tag
        source_tag = repo.tag(from_tag)
        env_tag = repo.tag(env_tag_name).create(source_ref=source_tag)

        print(f"Promoted to {environment}")
        print(f"  Source: {from_tag}")
        print(f"  Target: {env_tag_name}")
        print(f"  Commit: {env_tag.get_commit().id[:8]}")

        return env_tag_name

    except Exception as e:
        print(f"Promotion failed: {e}")
        return None

# Usage:
env_tag = promote_to_production("prod-repo", "v2.1.0", "us-west-1")
if env_tag:
    print(f"Production data updated to use {env_tag}")

에러 처리

자주 발생하는 레퍼런스 에러 다루기

import lakefs
from lakefs.exceptions import NotFoundException

repo = lakefs.repository("my-data-repo")

# Reference doesn't exist
try:
    ref = repo.ref("non-existent-ref")
    commit = ref.get_commit()
except NotFoundException:
    print("Reference not found")

# Commit doesn't exist
try:
    commit_ref = repo.commit("nonexistent123")
    commit = commit_ref.get_commit()
except NotFoundException:
    print("Commit not found")

# Invalid reference expression
try:
    ref = repo.ref("main~1000")  # Try to get 1000 commits back
    commit = ref.get_commit()
except NotFoundException:
    print("Reference expression invalid or out of range")

태그 에러 다루기

import lakefs
from lakefs.exceptions import ConflictException, NotFoundException

repo = lakefs.repository("my-data-repo")

# Tag already exists
try:
    tag = repo.tag("v1.0.0").create(source_ref="main", exist_ok=False)
except ConflictException:
    print("Tag already exists")
    tag = repo.tag("v1.0.0")

# Tag doesn't exist
try:
    tag = repo.tag("non-existent-tag")
    commit = tag.get_commit()
except NotFoundException:
    print("Tag not found")

더 알아보기 (Learn more)

공식 문서의 원문은 https://docs.lakefs.io/reference/python/refs/ 에서 확인할 수 있어요.