디렉터리 버킷에 멀티파트 업로드 사용하기

디렉터리 버킷에 멀티파트 업로드 사용하기 (Using multipart uploads with directory buckets)

멀티파트 업로드(multipart upload) 프로세스로 단일 객체를 파트(part) 집합으로 업로드할 수 있어요. 각 파트는 객체 데이터의 연속된 일부예요. 이러한 객체 파트는 독립적으로, 그리고 어떤 순서로든 업로드할 수 있어요. 어느 파트의 전송이 실패해도 다른 파트에 영향을 주지 않고 해당 파트만 다시 전송할 수 있어요. 객체의 모든 파트가 업로드된 뒤 Amazon S3가 이 파트들을 조립해 객체를 만들어요. 일반적으로 객체 크기가 100MB에 도달하면 단일 작업으로 업로드하는 대신 멀티파트 업로드를 사용하는 것을 고려하세요.

출처: 문서

본문

멀티파트 업로드를 사용하면 다음과 같은 장점이 있어요.

  • 향상된 처리량(Throughput) – 파트를 병렬로 업로드해 처리량을 높일 수 있어요.
  • 네트워크 문제로부터의 빠른 복구 – 파트 크기가 작을수록 네트워크 오류로 실패한 업로드를 다시 시작할 때의 영향이 줄어들어요.
  • 업로드 일시 중지·재개 – 시간을 두고 객체 파트를 업로드할 수 있어요. 멀티파트 업로드를 시작한 뒤에는 만료 날짜가 없어요. 멀티파트 업로드를 명시적으로 완료(complete)하거나 중단(abort)해야 해요.
  • 최종 객체 크기를 알기 전에 업로드 시작 – 객체를 만드는 동안에 업로드를 시작할 수 있어요.

다음과 같은 경우 멀티파트 업로드를 사용하는 것을 권장해요.

  • 안정적인 고대역폭 네트워크에서 대용량 객체를 업로드한다면, 객체 파트를 병렬로 업로드해 멀티스레드 성능으로 사용 가능한 대역폭을 최대한 활용하세요.
  • 불안정한 네트워크에서 업로드한다면, 업로드 재시작을 피해 네트워크 오류에 대한 복원력을 높이려고 멀티파트 업로드를 사용하세요. 멀티파트 업로드에서는 업로드 중 중단된 파트만 다시 업로드하면 돼요. 객체 업로드를 처음부터 다시 시작할 필요가 없어요.

디렉터리 버킷의 S3 Express One Zone 스토리지 클래스로 객체를 업로드하기 위해 멀티파트 업로드를 사용할 때, 그 프로세스는 일반 용도 버킷에 객체를 업로드하는 멀티파트 업로드와 비슷해요. 다만 몇 가지 눈에 띄는 차이점이 있어요.

S3 Express One Zone으로 객체를 업로드하기 위해 멀티파트 업로드를 사용하는 방법에 대한 자세한 내용은 아래 주제들을 참고하세요.

멀티파트 업로드 프로세스

멀티파트 업로드는 세 단계 프로세스예요.

  • 업로드를 시작합니다.
  • 객체 파트를 업로드합니다.
  • 모든 파트를 업로드한 뒤 멀티파트 업로드를 완료합니다.

멀티파트 업로드 완료 요청을 받으면 Amazon S3는 업로드된 파트에서 객체를 구성하며, 그 뒤에는 버킷의 다른 객체처럼 그 객체에 접근할 수 있어요.

멀티파트 업로드 시작

멀티파트 업로드 시작 요청을 보내면 Amazon S3는 업로드 ID(upload ID)가 포함된 응답을 반환해요. 업로드 ID는 멀티파트 업로드의 고유 식별자예요. 파트를 업로드하거나, 파트를 나열하거나, 업로드를 완료하거나, 업로드를 중단할 때마다 이 업로드 ID를 포함해야 해요.

파트 업로드

파트를 업로드할 때는 업로드 ID 외에도 파트 번호(part number)를 지정해야 해요. S3 Express One Zone에서 멀티파트 업로드를 사용할 때 멀티파트 파트 번호는 연속된 파트 번호여야 해요. 연속되지 않은 파트 번호로 멀티파트 업로드 완료 요청을 시도하면 HTTP 400 Bad Request(Invalid Part Order) 오류가 발생해요.

파트 번호는 파트와 업로드 중인 객체에서의 위치를 고유하게 식별해요. 이전에 업로드한 파트와 같은 파트 번호로 새 파트를 업로드하면 이전 파트가 덮어써져요.

파트를 업로드할 때마다 Amazon S3는 응답에서 엔티티 태그(ETag) 헤더를 반환해요. 각 파트 업로드에 대해 파트 번호와 ETag 값을 기록해야 해요. 모든 객체 파트 업로드의 ETag 값은 동일하게 유지되지만, 각 파트에는 서로 다른 파트 번호가 할당돼요. 멀티파트 업로드 완료를 위한 후속 요청에 이 값들을 포함해야 해요.

Amazon S3는 S3 버킷에 업로드되는 모든 새 객체를 자동으로 암호화해요. 멀티파트 업로드를 할 때 요청에 암호화 정보를 지정하지 않으면 업로드된 파트의 암호화 설정은 대상 버킷의 기본 암호화 구성으로 설정돼요. Amazon S3 버킷의 기본 암호화 구성은 항상 활성화되어 있으며, 최소한 Amazon S3 관리 키를 사용한 서버 측 암호화(SSE-S3)로 설정돼요. 디렉터리 버킷에서는 SSE-S3와 AWS KMS 키를 사용한 서버 측 암호화(SSE-KMS)가 지원돼요. 자세한 내용은 데이터 보호 및 암호화 를 참고하세요.

멀티파트 업로드 완료

멀티파트 업로드를 완료하면 Amazon S3는 파트 번호를 기준으로 파트를 오름차순으로 연결해 객체를 만들어요. 완료 요청이 성공한 뒤에는 파트가 더 이상 존재하지 않아요.

완료된 멀티파트 업로드 요청에는 업로드 ID와 파트 번호 및 해당 ETag 값 목록이 모두 포함되어야 해요. Amazon S3 응답에는 결합된 객체 데이터를 고유하게 식별하는 ETag가 포함돼요. 이 ETag는 객체 데이터의 MD5 해시가 아니에요.

멀티파트 업로드 나열

특정 멀티파트 업로드의 파트 또는 진행 중인 모든 멀티파트 업로드를 나열할 수 있어요. 파트 나열 작업(list parts)은 특정 멀티파트 업로드에 대해 업로드한 파트 정보를 반환해요. 각 파트 나열 요청에 대해 Amazon S3는 지정된 멀티파트 업로드의 파트 정보를 최대 1,000개까지 반환해요. 멀티파트 업로드에 1,000개가 넘는 파트가 있으면 페이징(pagination)으로 모든 파트를 검색해야 해요.

반환된 파트 목록에는 완료되지 않은 파트가 포함되지 않아요. 멀티파트 업로드 나열 작업(list multipart uploads)으로 진행 중인 멀티파트 업로드 목록을 얻을 수 있어요.

진행 중인 멀티파트 업로드란 시작했지만 아직 완료하거나 중단하지 않은 업로드예요. 각 요청은 최대 1,000개의 멀티파트 업로드를 반환해요. 진행 중인 멀티파트 업로드가 1,000개가 넘으면 나머지를 검색하려면 추가 요청을 보내야 해요. 반환된 목록은 검증용으로만 사용하세요. 멀티파트 업로드 완료 요청을 보낼 때는 이 목록의 결과를 사용하지 마세요. 대신 파트를 업로드할 때 지정한 파트 번호와 Amazon S3가 반환하는 해당 ETag 값의 목록을 직접 유지하세요.

멀티파트 업로드 나열에 대한 자세한 내용은 Amazon Simple Storage Service API Reference의 ListParts 를 참고하세요.

중요

멀티파트 업로드 완료 요청이 성공적으로 전송되지 않으면 객체 파트가 조립되지 않고 객체가 생성되지 않아요. 업로드된 파트와 관련된 모든 스토리지는 과금돼요. 객체를 만들기 위해 멀티파트 업로드를 완료하거나, 업로드된 파트를 제거하기 위해 멀티파트 업로드를 중단하는 것이 중요해요.

디렉터리 버킷을 삭제하려면 먼저 진행 중인 모든 멀티파트 업로드를 완료하거나 중단해야 해요. 디렉터리 버킷은 S3 Lifecycle 구성을 지원하지 않아요. 필요한 경우 활성 멀티파트 업로드를 나열한 뒤 업로드를 중단하고, 그다음 버킷을 삭제할 수 있어요.

멀티파트 업로드 작업과 체크섬

객체를 업로드할 때 객체 무결성을 확인할 체크섬 알고리즘을 지정할 수 있어요. 디렉터리 버킷에서는 MD5가 지원되지 않아요. 다음 Secure Hash Algorithms(SHA) 또는 Cyclic Redundancy Check(CRC) 데이터 무결성 검사 알고리즘 중 하나를 지정할 수 있어요.

  • CRC32
  • CRC32C
  • SHA-1
  • SHA-256

Amazon S3 REST API 또는 AWS SDK를 사용해 GetObject나 HeadObject로 개별 파트의 체크섬 값을 검색할 수 있어요. 아직 진행 중인 멀티파트 업로드의 개별 파트 체크섬 값을 검색하려면 ListParts를 사용할 수 있어요.

중요

위 체크섬 알고리즘을 사용할 때 멀티파트 파트 번호는 연속된 파트 번호여야 해요. 연속되지 않은 파트 번호로 멀티파트 업로드 완료 요청을 시도하면 Amazon S3가 HTTP 400 Bad Request(Invalid Part Order) 오류를 생성해요.

체크섬이 멀티파트 업로드 객체에서 어떻게 동작하는지에 대한 자세한 내용은 Amazon S3의 객체 무결성 확인 을 참고하세요.

동시 멀티파트 업로드 작업

분산 개발 환경에서 애플리케이션은 같은 객체에 대해 동시에 여러 업데이트를 시작할 수 있어요. 예를 들어 애플리케이션이 같은 객체 키로 여러 멀티파트 업로드를 시작할 수 있어요. 각 업로드에 대해 애플리케이션은 파트를 업로드한 뒤 Amazon S3에 완료 업로드 요청을 보내 객체를 만들 수 있어요. S3 Express One Zone의 경우 객체 생성 시간은 멀티파트 업로드의 완료 날짜예요.

중요

디렉터리 버킷에 저장된 객체에는 버전 관리(Versioning)가 지원되지 않아요.

멀티파트 업로드와 요금

멀티파트 업로드를 시작한 뒤에는 업로드를 완료하거나 중단할 때까지 Amazon S3가 모든 파트를 보관해요. 수명 기간 동안 이 멀티파트 업로드와 관련 파트에 대한 모든 스토리지, 대역폭, 요청이 과금돼요. 멀티파트 업로드를 중단하면 Amazon S3는 업로드 아티팩트와 업로드한 모든 파트를 삭제하고, 더 이상 과금되지 않아요. 지정된 스토리지 클래스와 관계없이 불완전한 멀티파트 업로드를 삭제하는 데는 조기 삭제 요금이 없어요. 요금에 대한 자세한 내용은 Amazon S3 요금 을 참고하세요.

멀티파트 업로드 API 작업과 권한

디렉터리 버킷의 객체 관리 API 작업에 대한 접근을 허용하려면 버킷 정책이나 AWS Identity and Access Management(IAM) ID 기반 정책에서 s3express:CreateSession 권한을 부여해요.

멀티파트 업로드 작업을 사용하려면 필요한 권한이 있어야 해요. 이러한 작업을 수행할 권한을 IAM 보안 주체에 부여하려면 버킷 정책이나 IAM ID 기반 정책을 사용할 수 있어요. 다음 표는 다양한 멀티파트 업로드 작업에 필요한 권한을 나열해요.

Initiator 요소를 통해 멀티파트 업로드의 시작자(initiator)를 식별할 수 있어요. 시작자가 AWS 계정이면 이 요소는 Owner 요소와 동일한 정보를 제공해요. 시작자가 IAM 사용자이면 이 요소는 사용자 ARN과 표시 이름을 제공해요.

| 작업 | 필요한 권한 | | 멀티파트 업로드 만들기 | 멀티파트 업로드를 만들려면 디렉터리 버킷에서 s3express:CreateSession 작업을 수행하도록 허용되어야 해요. | | 멀티파트 업로드 시작 | 멀티파트 업로드를 시작하려면 디렉터리 버킷에서 s3express:CreateSession 작업을 수행하도록 허용되어야 해요. | | 파트 업로드 | 파트를 업로드하려면 디렉터리 버킷에서 s3express:CreateSession 작업을 수행하도록 허용되어야 해요. 시작자가 파트를 업로드하려면 버킷 소유자가 시작자가 디렉터리 버킷에서 s3express:CreateSession 작업을 수행하도록 허용해야 해요. | | 파트 업로드(복사) | 파트를 업로드하려면 디렉터리 버킷에서 s3express:CreateSession 작업을 수행하도록 허용되어야 해요. 시작자가 객체에 대한 파트를 업로드하려면 버킷 소유자가 시작자가 객체에서 s3express:CreateSession 작업을 수행하도록 허용해야 해요. | | 멀티파트 업로드 완료 | 멀티파트 업로드를 완료하려면 디렉터리 버킷에서 s3express:CreateSession 작업을 수행하도록 허용되어야 해요. 시작자가 멀티파트 업로드를 완료하려면 버킷 소유자가 시작자가 객체에서 s3express:CreateSession 작업을 수행하도록 허용해야 해요. | | 멀티파트 업로드 중단 | 멀티파트 업로드를 중단하려면 s3express:CreateSession 작업을 수행하도록 허용되어야 해요. 시작자가 멀티파트 업로드를 중단하려면 s3express:CreateSession 작업을 수행하도록 명시적 허용 접근이 부여되어야 해요. | | 파트 나열 | 멀티파트 업로드의 파트를 나열하려면 디렉터리 버킷에서 s3express:CreateSession 작업을 수행하도록 허용되어야 해요. | | 진행 중인 멀티파트 업로드 나열 | 버킷의 진행 중인 멀티파트 업로드를 나열하려면 해당 버킷에서 s3:ListBucketMultipartUploads 작업을 수행하도록 허용되어야 해요. |

멀티파트 업로드에 대한 API 작업 지원

Amazon Simple Storage Service API Reference의 다음 섹션에서 멀티파트 업로드용 Amazon S3 REST API 작업을 설명해요.

  • CreateMultipartUpload
  • UploadPart
  • UploadPartCopy
  • CompleteMultipartUpload
  • AbortMultipartUpload
  • ListParts
  • ListMultipartUploads

예제

멀티파트 업로드로 디렉터리 버킷의 S3 Express One Zone에 객체를 업로드하려면 다음 예제를 참고하세요.

  • 멀티파트 업로드 만들기
  • 멀티파트 업로드의 파트 업로드하기
  • 멀티파트 업로드 완료하기
  • 멀티파트 업로드 중단하기
  • 멀티파트 업로드 복사 작업 만들기
  • 진행 중인 멀티파트 업로드 나열하기
  • 멀티파트 업로드의 파트 나열하기

멀티파트 업로드 만들기

참고

디렉터리 버킷에서 CreateMultipartUpload 작업과 UploadPartCopy 작업을 수행할 때 버킷의 기본 암호화가 원하는 암호화 구성을 사용해야 하며, CreateMultipartUpload 요청에 제공하는 요청 헤더는 대상 버킷의 기본 암호화 구성과 일치해야 해요.

다음 예제들은 멀티파트 업로드를 만드는 방법을 보여줘요.

Java 예제 (AWS SDK for Java 2.x)

/**
 * This method creates a multipart upload request that generates a unique upload ID that is used to track
 * all the upload parts
 *
 * @param s3
 * @param bucketName - for example, 'doc-example-bucket--use1-az4--x-s3'
 * @param key
 * @return
 */
 private static String createMultipartUpload(S3Client s3, String bucketName, String key) {
 
     CreateMultipartUploadRequest createMultipartUploadRequest = CreateMultipartUploadRequest.builder() 
             .bucket(bucketName)
             .key(key)
             .build();
             
     String uploadId = null;
     
     try {
         CreateMultipartUploadResponse response = s3.createMultipartUpload(createMultipartUploadRequest);
         uploadId = response.uploadId();
     }
     catch (S3Exception e) {
         System.err.println(e.awsErrorDetails().errorMessage());
         System.exit(1);
     }
     return uploadId;

Python 예제 (Boto3)

def create_multipart_upload(s3_client, bucket_name, key_name):
    '''
   Create a multipart upload to a directory bucket
   
   :param s3_client: boto3 S3 client
   :param bucket_name: The destination bucket for the multipart upload
   :param key_name: The key name for the object to be uploaded
   :return: The UploadId for the multipart upload if created successfully, else None
   '''
   
   try:
        mpu = s3_client.create_multipart_upload(Bucket = bucket_name, Key = key_name)
        return mpu['UploadId'] 
    except ClientError as e:
        logging.error(e)
        return None

AWS CLI 예제

다음 예제는 AWS CLI로 디렉터리 버킷에 멀티파트 업로드를 만드는 방법을 보여줘요. 이 명령은 bucket-base-name--zone-id--x-s3 디렉터리 버킷의 KEY_NAME 객체에 대해 멀티파트 업로드를 시작해요. 이 명령을 사용하려면 사용자 입력 자리표시자를 직접 채워 넣어야 해요.

aws s3api create-multipart-upload --bucket bucket-base-name--zone-id--x-s3 --key KEY_NAME

자세한 내용은 AWS Command Line Interface의 create-multipart-upload를 참고하세요.

멀티파트 업로드의 파트 업로드하기

다음 예제들은 멀티파트 업로드의 파트를 업로드하는 방법을 보여줘요. (이 섹션의 나머지 예제들은 AWS SDK로 파트를 업로드·복사하는 예제를 포함하지만, 여기서는 AWS CLI로 파트 업로드 예제 하나를 보여줘요.)

파트를 업로드하려면 AWS CLI upload-part 명령을 사용할 수 있어요.

aws s3api upload-part --bucket bucket-base-name--zone-id--x-s3 --key KEY_NAME --part-number 1 --upload-id "UPLOAD_ID" --body localfile

자세한 내용은 AWS Command Line Interface의 upload-part를 참고하세요.

멀티파트 업로드 복사 작업 만들기

다음 예제들은 AWS SDK로 멀티파트 업로드를 사용해 한 버킷에서 다른 버킷으로 객체를 프로그래밍 방식으로 복사하는 방법을 보여줘요.

Java 예제 (AWS SDK for Java 2.x)

다음 예제는 단일 객체를 파트로 나눈 뒤 SDK for Java 2.x로 디렉터리 버킷에 그 파트들을 업로드하는 방법을 보여줘요.

/**
 * This method creates part requests and uploads individual parts to S3 and then returns all the completed parts
 *
 * @param s3
 * @param bucketName
 * @param key
 * @param uploadId
 * @throws IOException
 */
 private static ListCompletedPartmultipartUpload(S3Client s3, String bucketName, String key, String uploadId, String filePath) throws IOException {

        int partNumber = 1;
        ListCompletedPart completedParts = new ArrayList<>();
        ByteBuffer bb = ByteBuffer.allocate(1024 * 1024 * 5); // 5 MB byte buffer

        // read the local file, breakdown into chunks and process
        try (RandomAccessFile file = new RandomAccessFile(filePath, "r")) {
            long fileSize = file.length();
            int position = 0;
            while (position ();
        while (bytePosition < objectSize) {
            // The last part might be smaller than partSize, so check to make sure
            // that lastByte isn't beyond the end of the object.
            long lastByte = Math.min(bytePosition + partSize - 1, objectSize - 1);

            System.out.println("part no: " + partNum + ", bytePosition: " + bytePosition + ", lastByte: " + lastByte);

            // Copy this part.
            UploadPartCopyRequest req = UploadPartCopyRequest.builder()
                    .uploadId(uploadId)
                    .sourceBucket(sourceBucket)
                    .sourceKey(sourceKey)
                    .destinationBucket(destnBucket)
                    .destinationKey(destnKey)
                    .copySourceRange("bytes="+bytePosition+"-"+lastByte)
                    .partNumber(partNum)
                    .build();
            UploadPartCopyResponse res = s3.uploadPartCopy(req);
            CompletedPart part = CompletedPart.builder()
                    .partNumber(partNum)
                    .eTag(res.copyPartResult().eTag())
                    .build();
            completedParts.add(part);
            partNum++;
            bytePosition += partSize;
        }
        return completedParts;
    }

    public static void multipartCopyUploadTest(S3Client s3, String srcBucket, String srcKey, String destnBucket, String destnKey)  {
        System.out.println("Starting multipart copy for: " + srcKey);
        try {
            String uploadId = createMultipartUpload(s3, destnBucket, destnKey);
            System.out.println(uploadId);
            ListCompletedPart parts = multipartUploadCopy(s3, srcBucket, srcKey,destnBucket,  destnKey, uploadId);
            completeMultipartUpload(s3, destnBucket, destnKey, uploadId, parts);
            System.out.println("Multipart copy completed for: " + srcKey);
        } catch (Exception e) {
            System.err.println(e.getMessage());
            System.exit(1);
        }
    }

Python 예제 (Boto3)

다음 예제는 SDK for Python(Boto3)로 멀티파트 업로드를 사용해 한 버킷에서 다른 버킷으로 객체를 프로그래밍 방식으로 복사하는 방법을 보여줘요.

import logging
import boto3
from botocore.exceptions import ClientError

def head_object(s3_client, bucket_name, key_name):
    '''
    Returns metadata for an object in a directory bucket

    :param s3_client: boto3 S3 client
    :param bucket_name: Bucket that contains the object to query for metadata
    :param key_name: Key name to query for metadata
    :return: Metadata for the specified object if successful, else None
    '''

    try:
        response = s3_client.head_object(
            Bucket = bucket_name,
            Key = key_name
        )
        return response
    except ClientError as e:
        logging.error(e)
        return None
    
def create_multipart_upload(s3_client, bucket_name, key_name):
    '''
    Create a multipart upload to a directory bucket

    :param s3_client: boto3 S3 client
    :param bucket_name: Destination bucket for the multipart upload
    :param key_name: Key name of the object to be uploaded
    :return: UploadId for the multipart upload if created successfully, else None
    '''
    
    try:
        mpu = s3_client.create_multipart_upload(Bucket = bucket_name, Key = key_name)
        return mpu['UploadId'] 
    except ClientError as e:
        logging.error(e)
        return None

def multipart_copy_upload(s3_client, source_bucket_name, key_name, target_bucket_name, mpu_id, part_size):
    '''
    Copy an object in a directory bucket to another bucket in multiple parts of a specified size
    
    :param s3_client: boto3 S3 client
    :param source_bucket_name: Bucket where the source object exists
    :param key_name: Key name of the object to be copied
    :param target_bucket_name: Destination bucket for copied object
    :param mpu_id: The UploadId returned from the create_multipart_upload call
    :param part_size: The size parts that the object will be broken into, in bytes. 
                      Minimum 5 MiB, Maximum 5 GiB. There is no minimum size for the last part of your multipart upload.
    :return: part_list for the multipart copy if all parts are copied successfully, else None
    '''
    
    part_list = []
    copy_source = {
        'Bucket': source_bucket_name,
        'Key': key_name
    }
    try:
        part_counter = 1
        object_size = head_object(s3_client, source_bucket_name, key_name)
        if object_size is not None:
            object_size = object_size['ContentLength']
        while (part_counter - 1) * part_size <object_size:
            bytes_start = (part_counter - 1) * part_size
            bytes_end = (part_counter * part_size) - 1
            upload_copy_part = s3_client.upload_part_copy (
                Bucket = target_bucket_name,
                CopySource = copy_source,
                CopySourceRange = f'bytes={bytes_start}-{bytes_end}',
                Key = key_name,
                PartNumber = part_counter,
                UploadId = mpu_id
            )
            part_list.append({'PartNumber': part_counter, 'ETag': upload_copy_part['CopyPartResult']['ETag']})
            part_counter += 1
    except ClientError as e:
        logging.error(e)
        return None
    return part_list

def complete_multipart_upload(s3_client, bucket_name, key_name, mpu_id, part_list):
    '''
    Completes a multipart upload to a directory bucket

    :param s3_client: boto3 S3 client
    :param bucket_name: Destination bucket for the multipart upload
    :param key_name: Key name of the object to be uploaded
    :param mpu_id: The UploadId returned from the create_multipart_upload call
    :param part_list: List of uploaded part numbers with associated ETags 
    :return: True if the multipart upload was completed successfully, else False
    '''
    
    try:
        s3_client.complete_multipart_upload(
            Bucket = bucket_name,
            Key = key_name,
            UploadId = mpu_id,
            MultipartUpload = {
                'Parts': part_list
            }
        )
    except ClientError as e:
        logging.error(e)
        return False
    return True

if __name__ == '__main__':
    MB = 1024 ** 2
    region = 'us-west-2'
    source_bucket_name = 'SOURCE_BUCKET_NAME'
    target_bucket_name = 'TARGET_BUCKET_NAME'
    key_name = 'KEY_NAME'
    part_size = 10 * MB
    s3_client = boto3.client('s3', region_name = region)
    mpu_id = create_multipart_upload(s3_client, target_bucket_name, key_name)
    if mpu_id is not None:
        part_list = multipart_copy_upload(s3_client, source_bucket_name, key_name, target_bucket_name, mpu_id, part_size)
        if part_list is not None:
            if complete_multipart_upload(s3_client, target_bucket_name, key_name, mpu_id, part_list):
                print (f'{key_name} successfully copied through multipart copy from {source_bucket_name} to {target_bucket_name}')
            else:
                print (f'Could not copy {key_name} through multipart copy from {source_bucket_name} to {target_bucket_name}')

AWS CLI 예제

다음 예제는 AWS CLI로 멀티파트 업로드를 사용해 한 버킷에서 디렉터리 버킷으로 객체를 프로그래밍 방식으로 복사하는 방법을 보여줘요. 이 명령을 사용하려면 사용자 입력 자리표시자를 직접 채워 넣어야 해요.

aws s3api upload-part-copy --bucket bucket-base-name--zone-id--x-s3 --key TARGET_KEY_NAME --copy-source SOURCE_BUCKET_NAME/SOURCE_KEY_NAME --part-number 1 --upload-id "AS_mgt9RaQE9GEaifATue15dAAAAAAAAAAEMAAAAAAAAADQwNzI4MDU0MjUyMBYAAAAAAAAAAA0AAAAAAAAAAAH2AfYAAAAAAAAEBnJ4cxKMAQAAAABiNXpOFVZJ1tZcKWib9YKE1C565_hCkDJ_4AfCap2svg"

자세한 내용은 AWS Command Line Interface의 upload-part-copy를 참고하세요.

진행 중인 멀티파트 업로드 나열

디렉터리 버킷의 진행 중인 멀티파트 업로드를 나열하려면 AWS SDK 또는 AWS CLI를 사용할 수 있어요.

Java 예제 (AWS SDK for Java 2.x)

다음 예제들은 SDK for Java 2.x로 진행 중인(불완전한) 멀티파트 업로드를 나열하는 방법을 보여줘요.

 public static void listMultiPartUploads( S3Client s3, String bucketName) {
        try {
            ListMultipartUploadsRequest listMultipartUploadsRequest = ListMultipartUploadsRequest.builder()
                .bucket(bucketName)
                .build();
                
            ListMultipartUploadsResponse response = s3.listMultipartUploads(listMultipartUploadsRequest);
            List MultipartUpload uploads = response.uploads();
            for (MultipartUpload upload: uploads) {
                System.out.println("Upload in progress: Key = \"" + upload.key() + "\", id = " + upload.uploadId());
            }
      }
      catch (S3Exception e) {
            System.err.println(e.getMessage());
            System.exit(1);
      }
  }

Python 예제 (Boto3)

다음 예제들은 SDK for Python(Boto3)로 진행 중인(불완전한) 멀티파트 업로드를 나열하는 방법을 보여줘요.

import logging
import boto3
from botocore.exceptions import ClientError

def list_multipart_uploads(s3_client, bucket_name):
    '''
    List any incomplete multipart uploads in a directory bucket in e specified gion

    :param s3_client: boto3 S3 client
    :param bucket_name: Bucket to check for incomplete multipart uploads
    :return: List of incomplete multipart uploads if there are any, None if not
    '''
    
    try:
        response = s3_client.list_multipart_uploads(Bucket = bucket_name)
        if 'Uploads' in response.keys():
            return response['Uploads']
        else:
            return None 
    except ClientError as e:
        logging.error(e)

if __name__ == '__main__':
    bucket_name = 'BUCKET_NAME'
    region = 'us-west-2'
    s3_client = boto3.client('s3', region_name = region)
    multipart_uploads = list_multipart_uploads(s3_client, bucket_name)
    if multipart_uploads is not None:
        print (f'There are {len(multipart_uploads)} ncomplete multipart uploads for {bucket_name}')
    else:
        print (f'There are no incomplete multipart uploads for {bucket_name}')

AWS CLI 예제

다음 예제들은 AWS CLI로 진행 중인(불완전한) 멀티파트 업로드를 나열하는 방법을 보여줘요. 이 명령을 사용하려면 사용자 입력 자리표시자를 직접 채워 넣어야 해요.

aws s3api list-multipart-uploads --bucket bucket-base-name--zone-id--x-s3

자세한 내용은 AWS Command Line Interface의 list-multipart-uploads를 참고하세요.

멀티파트 업로드의 파트 나열

다음 예제들은 디렉터리 버킷의 멀티파트 업로드 파트를 나열하는 방법을 보여줘요.

Java 예제 (AWS SDK for Java 2.x)

다음 예제는 SDK for Java 2.x로 디렉터리 버킷의 멀티파트 업로드 파트를 나열하는 방법을 보여줘요.

public static void listMultiPartUploadsParts( S3Client s3, String bucketName, String objKey, String uploadID) {
         
         try {
             ListPartsRequest listPartsRequest = ListPartsRequest.builder()
                 .bucket(bucketName)
                 .uploadId(uploadID)
                 .key(objKey)
                 .build();

             ListPartsResponse response = s3.listParts(listPartsRequest);
             ListPart parts = response.parts();
             for (Part part: parts) {
                 System.out.println("Upload in progress: Part number = \"" + part.partNumber() + "\", etag = " + part.eTag());
             }

         } 
         
         catch (S3Exception e) {
             System.err.println(e.getMessage());
             System.exit(1);
         }
         
         
     }

Python 예제 (Boto3)

다음 예제는 SDK for Python(Boto3)로 디렉터리 버킷의 멀티파트 업로드 파트를 나열하는 방법을 보여줘요.

import logging
import boto3
from botocore.exceptions import ClientError

def list_parts(s3_client, bucket_name, key_name, upload_id):
    '''
    Lists the parts that have been uploaded for a specific multipart upload to a directory bucket.
    
    :param s3_client: boto3 S3 client
    :param bucket_name: Bucket that multipart uploads parts have been uploaded to
    :param key_name: Name of the object that has parts uploaded
    :param upload_id: Multipart upload ID that the parts are associated with
    :return: List of parts associated with the specified multipart upload, None if there are no parts
    '''
    parts_list = []
    next_part_marker = ''
    continuation_flag = True
    try:
        while continuation_flag:
            if next_part_marker == '':
                response = s3_client.list_parts(
                    Bucket = bucket_name,
                    Key = key_name,
                    UploadId = upload_id
                )
            else:
                response = s3_client.list_parts(
                    Bucket = bucket_name,
                    Key = key_name,
                    UploadId = upload_id,
                    NextPartMarker = next_part_marker
                )
            if 'Parts' in response:
                for part in response['Parts']:
                    parts_list.append(part)
                if response['IsTruncated']:
                    next_part_marker = response['NextPartNumberMarker']
                else:
                    continuation_flag = False
            else:
                continuation_flag = False
        return parts_list
    except ClientError as e:
        logging.error(e)
        return None

if __name__ == '__main__':
    region = 'us-west-2'
    bucket_name = 'BUCKET_NAME'
    key_name = 'KEY_NAME'
    upload_id = 'UPLOAD_ID'
    s3_client = boto3.client('s3', region_name = region)
    parts_list = list_parts(s3_client, bucket_name, key_name, upload_id)
    if parts_list is not None:
        print (f'{key_name} has {len(parts_list)} parts uploaded to {bucket_name}')
    else:
        print (f'There are no multipart uploads with that upload ID for {bucket_name} bucket')

AWS CLI 예제

다음 예제들은 AWS CLI로 디렉터리 버킷의 멀티파트 업로드 파트를 나열하는 방법을 보여줘요. 이 명령을 사용하려면 사용자 입력 자리표시자를 직접 채워 넣어야 해요.

aws s3api list-parts --bucket bucket-base-name--zone-id--x-s3 --key KEY_NAME --upload-id "AS_mgt9RaQE9GEaifATue15dAAAAAAAAAAEMAAAAAAAAADQwNzI4MDU0MjUyMBYAAAAAAAAAAA0AAAAAAAAAAAH2AfYAAAAAAAAEBSD0WBKMAQAAAABneY9yBVsK89iFkvWdQhRCcXohE8RbYtc9QvBOG8tNpA"

자세한 내용은 AWS Command Line Interface의 list-parts를 참고하세요.

더 알아보기 (Learn more)