Amazon S3 Tables Catalog for Apache Iceberg로 테이블 접근하기

Amazon S3 Tables Catalog for Apache Iceberg로 테이블 접근하기 (Accessing tables with the client catalog)

Amazon S3 Tables Catalog for Apache Iceberg 클라이언트 카탈로그를 사용해 Apache Spark 같은 오픈 소스 쿼리 엔진에서 S3 테이블에 접근할 수 있어요. Amazon S3 Tables Catalog for Apache Iceberg는 AWS Labs가 호스팅하는 오픈 소스 라이브러리예요. 쿼리 엔진의 Apache Iceberg 작업(테이블 발견, 메타데이터 업데이트, 테이블 추가·제거 등)을 S3 Tables API 작업으로 변환해 동작해요.

출처: 문서

본문

Amazon S3 Tables Catalog for Apache Iceberg는 s3-tables-catalog-for-iceberg.jar라는 Maven JAR로 배포돼요. 클라이언트 카탈로그 JAR는 AWS Labs GitHub 저장소에서 빌드하거나 Maven에서 다운로드할 수 있어요. 테이블에 연결할 때는 Apache Iceberg용 Spark 세션을 초기화하면서 클라이언트 카탈로그 JAR를 종속성으로 사용해요.

Apache Spark에서 Amazon S3 Tables Catalog for Apache Iceberg 사용하기

Spark 세션을 초기화할 때 Amazon S3 Tables Catalog for Apache Iceberg 클라이언트 카탈로그를 사용해 오픈 소스 애플리케이션에서 테이블에 연결할 수 있어요. 세션 구성에서 Iceberg와 Amazon S3 종속성을 지정하고, 테이블 버킷을 메타데이터 웨어하우스로 사용하는 커스텀 카탈로그를 만들어요.

사전 조건

  • 테이블 버킷과 S3 Tables 작업에 접근할 수 있는 IAM 자격 증명. 자세한 내용은 "Access management for S3 Tables" 문서를 참고하세요.

Amazon S3 Tables Catalog for Apache Iceberg로 Spark 세션을 초기화하려면:

다음 명령으로 Spark를 초기화해요. 이 명령을 사용할 때는 Amazon S3 Tables Catalog for Apache Iceberg 버전 번호를 AWS Labs GitHub 저장소의 최신 버전으로, 테이블 버킷 ARN을 자신의 테이블 버킷 ARN으로 바꾸세요.

spark-shell \
--packages org.apache.iceberg:iceberg-spark-runtime-3.5_2.12:1.6.1,software.amazon.s3tables:s3-tables-catalog-for-iceberg-runtime:0.1.4 \
--conf spark.sql.catalog.s3tablesbucket=org.apache.iceberg.spark.SparkCatalog \
--conf spark.sql.catalog.s3tablesbucket.catalog-impl=software.amazon.s3tables.iceberg.S3TablesCatalog \
--conf spark.sql.catalog.s3tablesbucket.warehouse=arn:aws:s3tables:us-east-1:111122223333:bucket/amzn-s3-demo-table-bucket \
--conf spark.sql.extensions=org.apache.iceberg.spark.extensions.IcebergSparkSessionExtensions

Spark SQL로 S3 테이블 쿼리하기

Spark를 사용해 S3 테이블에서 DQL, DML, DDL 작업을 실행할 수 있어요. 테이블을 쿼리할 때는 세션 카탈로그 이름을 포함한 정규화된 테이블 이름을 다음 패턴으로 사용해요.

CatalogName.NamespaceName.TableName

다음 예시 쿼리들은 S3 테이블과 상호작용할 수 있는 몇 가지 방법을 보여줘요. 이 예시 쿼리를 쿼리 엔진에서 사용하려면 사용자 입력 자리표시자 값을 자신의 값으로 바꾸세요.

Spark로 테이블을 쿼리하려면:

  1. 네임스페이스 만들기
    spark.sql("CREATE NAMESPACE IF NOT EXISTS s3tablesbucket.my_namespace")
    
  2. 테이블 만들기
    spark.sql("CREATE TABLE IF NOT EXISTS s3tablesbucket.my_namespace.`my_table` (id INT, name STRING, value INT) USING iceberg")
    
  3. 테이블 쿼리하기
    spark.sql("SELECT * FROM s3tablesbucket.my_namespace.`my_table`").show()
    
  4. 테이블에 데이터 삽입하기
    spark.sql("""
        INSERT INTO s3tablesbucket.my_namespace.my_table VALUES
            (1, 'ABC', 100),
            (2, 'XYZ', 200)
    """)
    
  5. 기존 데이터 파일을 테이블로 로드하기
    • 데이터를 Spark로 읽어요.
      val data_file_location = "Path such as S3 URI to data file"
      val data_file = spark.read.parquet(data_file_location)
      
    • 데이터를 Iceberg 테이블로 작성해요.
      data_file.writeTo("s3tablesbucket.my_namespace.my_table").using("Iceberg").tableProperty("format-version", "2").createOrReplace()
      

더 알아보기 (Learn more)

  • Amazon S3 Tables Iceberg REST 엔드포인트로 테이블 접근하기 (Accessing tables using the Amazon S3 Tables Iceberg REST endpoint)
  • Athena로 S3 테이블 쿼리하기 (Amazon Athena)