수집 Job Spec

수집 Job Spec (Ingestion Job Specification)

Apache Pinot 배치 수집 job specification 레퍼런스 문서예요.

출처: 문서

본문

Pinot 배치 수집 작업은 YAML job spec 파일을 사용해 구성해요. 이 파일은 데이터 입력 위치, 세그먼트 생성 및 push 방식을 정의해요. 작업은 bin/pinot-admin.sh LaunchDataIngestionJob -jobSpecFile <path>로 실행해요.

최상위 필드

필드 설명
executionFrameworkSpec 실행 프레임워크 정의.
jobType 수행할 수집 작업 유형.
inputDirURI 입력 데이터의 루트 디렉터리 URI(예: s3://bucket/rawdata).
includeFileNamePattern / excludeFileNamePattern 포함/제외할 파일 이름 패턴.
outputDirURI 출력 세그먼트의 루트 디렉터리 URI.
overwriteOutput 기존 출력 세그먼트 덮어쓰기 여부.
pinotFSSpecs 모든 관련 Pinot 파일 시스템 정의.
recordReaderSpec record reader(입력 형식) 정의.
tableSpec 테이블 이름과 테이블 config/schema 위치 정의.
segmentNameGeneratorSpec 세그먼트 이름 생성기 정의.
pinotClusterSpecs Pinot 클러스터 접근 지점 정의.
pushJobSpec 세그먼트 push 관련 구성 정의.

executionFrameworkSpec

수집 작업을 실행할 분산 실행기(standalone, hadoop, spark 등)를 정의해요.

필드 설명
name 실행 프레임워크 이름: standalone, hadoop, spark, flink 등.
segmentGenerationJobRunnerClassName SegmentGenerationJobRunner 인터페이스 구현 클래스.
segmentTarPushJobRunnerClassName SegmentTarPushJobRunner 인터페이스 구현 클래스.
segmentUriPushJobRunnerClassName SegmentUriPushJobRunner 인터페이스 구현 클래스.
segmentMetadataPushJobRunnerClassName SegmentMetadataPushJobRunner 인터페이스 구현 클래스.
extraConfigs 실행 프레임워크용 추가 config. Spark/Hadoop의 stagingDir(분산 파일시스템에서 모든 세그먼트를 호스팅한 뒤 출력 디렉터리로 이동) 포함.

jobType

수집 작업 유형. 지원되는 유형:

  • SegmentCreation
  • SegmentTarPush
  • SegmentUriPush
  • SegmentCreationAndTarPush
  • SegmentCreationAndUriPush
  • SegmentCreationAndMetadataPush

파일 이름 패턴 (File name patterns)

includeFileNamePattern / excludeFileNamePattern은 Java NIO PathMatcher 패턴(glob: 또는 regex:)을 지원해요.

  • glob:*.avro — inputDirURI 바로 아래의 avro 파일만 매칭(하위 디렉터리 제외).
  • glob:**/*.avro — inputDirURI 아래의 모든 avro 파일을 재귀적으로 매칭.
  • regex:.*[.]avro — 정규식 기반 매칭.

pinotFSSpecs

각 항목은 하나의 Pinot 파일시스템을 정의해요:

필드 설명
scheme PinotFS 식별 scheme(예: file, hdfs, s3, adl, gs).
className PinotFS 인스턴스 생성에 사용되는 클래스 이름(예: org.apache.pinot.spi.filesystem.LocalPinotFS, org.apache.pinot.plugin.filesystem.S3PinotFS, org.apache.pinot.plugin.filesystem.HadoopPinotFS, org.apache.pinot.plugin.filesystem.AzurePinotFS).
configs 파일시스템별 속성(예: region, accessKey).

recordReaderSpec

입력 형식의 record reader를 정의해요:

필드 설명
dataFormat 레코드 데이터 형식: avro, parquet, orc, csv, json, thrift 등.
className 해당 RecordReader 클래스 이름(예: org.apache.pinot.plugin.inputformat.avro.AvroRecordReader, org.apache.pinot.plugin.inputformat.csv.CSVRecordReader, org.apache.pinot.plugin.inputformat.parquet.ParquetRecordReader, org.apache.pinot.plugin.inputformat.json.JSONRecordReader).
configClassName RecordReader config 클래스(선택).
configs reader별 추가 config.

tableSpec

필드 설명
tableName 테이블 이름.
schemaURI 테이블 스키마를 읽을 위치. PinotFS(hdfs://path) 또는 HTTP(http://localhost:9000/tables/myTable/schema) 지원.
tableConfigURI 테이블 config를 읽을 위치. 참고: controller에서 직접 읽는 API는 JSON wrapper를 포함하며 실제 config는 OFFLINE 필드 아래의 객체.

segmentNameGeneratorSpec

필드 설명
type simple 또는 normalizedDate.
configs 설정: segment.name.prefix, exclude.sequence.id 등.

normalizedDate 유형은 입력 데이터의 날짜 폴더(예: 2014/01/01) 일부를 세그먼트 이름에 반영해요.

pinotClusterSpecs

필드 설명
controllerURI 테이블/스키마 정보와 데이터 push에 사용되는 controller URI(예: http://localhost:9000).

pushJobSpec

필드 설명
pushParallelism 세그먼트 push 병렬도.
pushAttempts push 시도 횟수. 기본 1(재시도 없음).
pushRetryIntervalMillis 재시도 대기 시간(밀리초). 기본 1초.
segmentUriPrefix URI/METADATA push에서 세그먼트 URI 접두사. Pinot < 0.6.0에서는 [scheme]://[bucket.name](예: s3://my.bucket), 0.6.0+에서는 빈 문자열 또는 무시.
segmentUriSuffix 세그먼트 URI 접미사.
copyToDeepStoreForMetadataPush URI/METADATA push 적용. true이고 세그먼트가 아직 딥 스토어에 없으면 딥 스토어로 이동.
preferMetadataTarGz metadata tar gz 파일이 있으면 push에 사용 선호.

샘플 job spec

executionFrameworkSpec:
  name: 'standalone'
  segmentGenerationJobRunnerClassName: 'org.apache.pinot.plugin.ingestion.batch.standalone.SegmentGenerationJobRunner'
  segmentTarPushJobRunnerClassName: 'org.apache.pinot.plugin.ingestion.batch.standalone.SegmentTarPushJobRunner'
  segmentUriPushJobRunnerClassName: 'org.apache.pinot.plugin.ingestion.batch.standalone.SegmentUriPushJobRunner'
  segmentMetadataPushJobRunnerClassName: 'org.apache.pinot.plugin.ingestion.batch.standalone.SegmentMetadataPushJobRunner'

jobType: SegmentCreationAndTarPush

inputDirURI: 'examples/batch/airlineStats/rawdata'
includeFileNamePattern: 'glob:**/*.avro'
excludeFileNamePattern: 'glob:**/*.tmp'
outputDirURI: 'examples/batch/airlineStats/segments'
overwriteOutput: true
createMetadataTarGz: true
segmentCreationJobParallelism: 4

pinotFSSpecs:
  - scheme: file
    className: org.apache.pinot.spi.filesystem.LocalPinotFS
  - scheme: s3
    className: org.apache.pinot.plugin.filesystem.S3PinotFS
    configs:
      region: 'us-west-2'

recordReaderSpec:
  dataFormat: 'avro'
  className: 'org.apache.pinot.plugin.inputformat.avro.AvroRecordReader'

tableSpec:
  tableName: 'airlineStats'
  schemaURI: 'http://localhost:9000/tables/airlineStats/schema'
  tableConfigURI: 'http://localhost:9000/tables/airlineStats'

segmentNameGeneratorSpec:
  type: normalizedDate
  configs:
    segment.name.prefix: 'airlineStats_batch'
    exclude.sequence.id: true

pinotClusterSpecs:
  - controllerURI: 'http://localhost:9000'

pushJobSpec:
  pushParallelism: 4
  pushAttempts: 2
  pushRetryIntervalMillis: 1000
  copyToDeepStoreForMetadataPush: false
  preferMetadataTarGz: true

작업은 다음 명령으로 실행해요:

bin/pinot-admin.sh LaunchDataIngestionJob -jobSpecFile examples/batch/airlineStats/ingestionJobSpec.yaml

더 알아보기 (Learn more)