수집 Job Spec
수집 Job Spec (Ingestion Job Specification)
Apache Pinot 배치 수집 job specification 레퍼런스 문서예요.
출처: 문서
본문
Pinot 배치 수집 작업은 YAML job spec 파일을 사용해 구성해요. 이 파일은 데이터 입력 위치, 세그먼트 생성 및 push 방식을 정의해요. 작업은 bin/pinot-admin.sh LaunchDataIngestionJob -jobSpecFile <path>로 실행해요.
최상위 필드
| 필드 | 설명 |
|---|---|
| executionFrameworkSpec | 실행 프레임워크 정의. |
| jobType | 수행할 수집 작업 유형. |
| inputDirURI | 입력 데이터의 루트 디렉터리 URI(예: s3://bucket/rawdata). |
| includeFileNamePattern / excludeFileNamePattern | 포함/제외할 파일 이름 패턴. |
| outputDirURI | 출력 세그먼트의 루트 디렉터리 URI. |
| overwriteOutput | 기존 출력 세그먼트 덮어쓰기 여부. |
| pinotFSSpecs | 모든 관련 Pinot 파일 시스템 정의. |
| recordReaderSpec | record reader(입력 형식) 정의. |
| tableSpec | 테이블 이름과 테이블 config/schema 위치 정의. |
| segmentNameGeneratorSpec | 세그먼트 이름 생성기 정의. |
| pinotClusterSpecs | Pinot 클러스터 접근 지점 정의. |
| pushJobSpec | 세그먼트 push 관련 구성 정의. |
executionFrameworkSpec
수집 작업을 실행할 분산 실행기(standalone, hadoop, spark 등)를 정의해요.
| 필드 | 설명 |
|---|---|
| name | 실행 프레임워크 이름: standalone, hadoop, spark, flink 등. |
| segmentGenerationJobRunnerClassName | SegmentGenerationJobRunner 인터페이스 구현 클래스. |
| segmentTarPushJobRunnerClassName | SegmentTarPushJobRunner 인터페이스 구현 클래스. |
| segmentUriPushJobRunnerClassName | SegmentUriPushJobRunner 인터페이스 구현 클래스. |
| segmentMetadataPushJobRunnerClassName | SegmentMetadataPushJobRunner 인터페이스 구현 클래스. |
| extraConfigs | 실행 프레임워크용 추가 config. Spark/Hadoop의 stagingDir(분산 파일시스템에서 모든 세그먼트를 호스팅한 뒤 출력 디렉터리로 이동) 포함. |
jobType
수집 작업 유형. 지원되는 유형:
SegmentCreationSegmentTarPushSegmentUriPushSegmentCreationAndTarPushSegmentCreationAndUriPushSegmentCreationAndMetadataPush
파일 이름 패턴 (File name patterns)
includeFileNamePattern / excludeFileNamePattern은 Java NIO PathMatcher 패턴(glob: 또는 regex:)을 지원해요.
glob:*.avro— inputDirURI 바로 아래의 avro 파일만 매칭(하위 디렉터리 제외).glob:**/*.avro— inputDirURI 아래의 모든 avro 파일을 재귀적으로 매칭.regex:.*[.]avro— 정규식 기반 매칭.
pinotFSSpecs
각 항목은 하나의 Pinot 파일시스템을 정의해요:
| 필드 | 설명 |
|---|---|
| scheme | PinotFS 식별 scheme(예: file, hdfs, s3, adl, gs). |
| className | PinotFS 인스턴스 생성에 사용되는 클래스 이름(예: org.apache.pinot.spi.filesystem.LocalPinotFS, org.apache.pinot.plugin.filesystem.S3PinotFS, org.apache.pinot.plugin.filesystem.HadoopPinotFS, org.apache.pinot.plugin.filesystem.AzurePinotFS). |
| configs | 파일시스템별 속성(예: region, accessKey). |
recordReaderSpec
입력 형식의 record reader를 정의해요:
| 필드 | 설명 |
|---|---|
| dataFormat | 레코드 데이터 형식: avro, parquet, orc, csv, json, thrift 등. |
| className | 해당 RecordReader 클래스 이름(예: org.apache.pinot.plugin.inputformat.avro.AvroRecordReader, org.apache.pinot.plugin.inputformat.csv.CSVRecordReader, org.apache.pinot.plugin.inputformat.parquet.ParquetRecordReader, org.apache.pinot.plugin.inputformat.json.JSONRecordReader). |
| configClassName | RecordReader config 클래스(선택). |
| configs | reader별 추가 config. |
tableSpec
| 필드 | 설명 |
|---|---|
| tableName | 테이블 이름. |
| schemaURI | 테이블 스키마를 읽을 위치. PinotFS(hdfs://path) 또는 HTTP(http://localhost:9000/tables/myTable/schema) 지원. |
| tableConfigURI | 테이블 config를 읽을 위치. 참고: controller에서 직접 읽는 API는 JSON wrapper를 포함하며 실제 config는 OFFLINE 필드 아래의 객체. |
segmentNameGeneratorSpec
| 필드 | 설명 |
|---|---|
| type | simple 또는 normalizedDate. |
| configs | 설정: segment.name.prefix, exclude.sequence.id 등. |
normalizedDate 유형은 입력 데이터의 날짜 폴더(예: 2014/01/01) 일부를 세그먼트 이름에 반영해요.
pinotClusterSpecs
| 필드 | 설명 |
|---|---|
| controllerURI | 테이블/스키마 정보와 데이터 push에 사용되는 controller URI(예: http://localhost:9000). |
pushJobSpec
| 필드 | 설명 |
|---|---|
| pushParallelism | 세그먼트 push 병렬도. |
| pushAttempts | push 시도 횟수. 기본 1(재시도 없음). |
| pushRetryIntervalMillis | 재시도 대기 시간(밀리초). 기본 1초. |
| segmentUriPrefix | URI/METADATA push에서 세그먼트 URI 접두사. Pinot < 0.6.0에서는 [scheme]://[bucket.name](예: s3://my.bucket), 0.6.0+에서는 빈 문자열 또는 무시. |
| segmentUriSuffix | 세그먼트 URI 접미사. |
| copyToDeepStoreForMetadataPush | URI/METADATA push 적용. true이고 세그먼트가 아직 딥 스토어에 없으면 딥 스토어로 이동. |
| preferMetadataTarGz | metadata tar gz 파일이 있으면 push에 사용 선호. |
샘플 job spec
executionFrameworkSpec:
name: 'standalone'
segmentGenerationJobRunnerClassName: 'org.apache.pinot.plugin.ingestion.batch.standalone.SegmentGenerationJobRunner'
segmentTarPushJobRunnerClassName: 'org.apache.pinot.plugin.ingestion.batch.standalone.SegmentTarPushJobRunner'
segmentUriPushJobRunnerClassName: 'org.apache.pinot.plugin.ingestion.batch.standalone.SegmentUriPushJobRunner'
segmentMetadataPushJobRunnerClassName: 'org.apache.pinot.plugin.ingestion.batch.standalone.SegmentMetadataPushJobRunner'
jobType: SegmentCreationAndTarPush
inputDirURI: 'examples/batch/airlineStats/rawdata'
includeFileNamePattern: 'glob:**/*.avro'
excludeFileNamePattern: 'glob:**/*.tmp'
outputDirURI: 'examples/batch/airlineStats/segments'
overwriteOutput: true
createMetadataTarGz: true
segmentCreationJobParallelism: 4
pinotFSSpecs:
- scheme: file
className: org.apache.pinot.spi.filesystem.LocalPinotFS
- scheme: s3
className: org.apache.pinot.plugin.filesystem.S3PinotFS
configs:
region: 'us-west-2'
recordReaderSpec:
dataFormat: 'avro'
className: 'org.apache.pinot.plugin.inputformat.avro.AvroRecordReader'
tableSpec:
tableName: 'airlineStats'
schemaURI: 'http://localhost:9000/tables/airlineStats/schema'
tableConfigURI: 'http://localhost:9000/tables/airlineStats'
segmentNameGeneratorSpec:
type: normalizedDate
configs:
segment.name.prefix: 'airlineStats_batch'
exclude.sequence.id: true
pinotClusterSpecs:
- controllerURI: 'http://localhost:9000'
pushJobSpec:
pushParallelism: 4
pushAttempts: 2
pushRetryIntervalMillis: 1000
copyToDeepStoreForMetadataPush: false
preferMetadataTarGz: true
작업은 다음 명령으로 실행해요:
bin/pinot-admin.sh LaunchDataIngestionJob -jobSpecFile examples/batch/airlineStats/ingestionJobSpec.yaml