Pinot의 딥 스토리지로 OSS 사용하기

Pinot의 딥 스토리지로 OSS 사용하기

AliCloud Object Storage Service(OSS)를 Pinot 딥 스토리지로 구성하는 방법을 알려드려요.

출처: 문서

본문

OSS 파일시스템 플러그인을 구현하지 않고도 OSS를 Apache Pinot의 HDFS 딥 스토리지로 사용할 수 있어요. 아래 단계를 따르세요:

1. hdfs-site.xml과 core-site.xml 파일을 구성하세요. 그 후 이 구성을 임의 경로 아래에 두고, controller/server 구성의 pinot.<node>.storage.factory.oss.hadoop.conf 값에 이 경로를 설정하세요.

hdfs-site.xml의 경우 아무 구성도 제공할 필요가 없어요:

<?xml version="1.0" encoding="UTF-8"?>
<configuration>
</configuration>

core-site.xml의 경우 아래처럼 OSS access/secret과 bucket 구성을 제공해야 해요:

<?xml version="1.0"?>
<configuration>
    <property>
	      <name>fs.defaultFS</name>
	      <value>oss://your-bucket-name/</value>
	  </property>
    <property>
        <name>fs.oss.accessKeyId</name>
        <value>your-access-key-id</value>
    </property>
    <property>
        <name>fs.oss.accessKeySecret</name>
        <value>your-access-key-secret</value>
    </property>
    <property>
        <name>fs.oss.impl</name>
        <value>com.aliyun.emr.fs.oss.OssFileSystem</value>
    </property>
    <property>
        <name>fs.oss.endpoint</name>
        <value>your-oss-endpoint</value>
    </property>
</configuration>

2. OSS에 접근하기 위해 OSS와 관련된 HDFS jar들을 찾아 PINOT_DIR/lib 아래에 두세요. 아래 jar를 사용할 수 있지만 충돌을 피하기 위해 버전에 주의하세요.

  • smartdata-aliyun-oss
  • smartdata-hadoop-common
  • guava

3. controller.conf와 server.conf에 OSS 딥 스토리지 구성을 설정하세요:

Controller 구성

controller.data.dir=oss://your-bucket-name/path/to/segments
controller.local.temp.dir=/path/to/local/temp/directory
controller.enable.split.commit=true
pinot.controller.storage.factory.class.oss=org.apache.pinot.plugin.filesystem.HadoopPinotFS
pinot.controller.storage.factory.oss.hadoop.conf.path=path/to/conf/directory/
pinot.controller.segment.fetcher.protocols=file,http,oss
pinot.controller.segment.fetcher.oss.class=org.apache.pinot.common.utils.fetcher.PinotFSSegmentFetcher

Server 구성

pinot.server.instance.enable.split.commit=true
pinot.server.storage.factory.class.oss=org.apache.pinot.plugin.filesystem.HadoopPinotFS
pinot.server.storage.factory.oss.hadoop.conf.path=path/to/conf/directory/
pinot.server.segment.fetcher.protocols=file,http,oss
pinot.server.segment.fetcher.oss.class=org.apache.pinot.common.utils.fetcher.PinotFSSegmentFetcher

예시 작업 스펙

같은 HDFS 딥 스토리지 구성과 jar를 사용해 OSS에서 데이터를 읽고, 세그먼트를 만들어 다시 OSS로 push할 수 있어요. 예시 standalone 배치 수집 작업은 아래와 같을 수 있어요:

executionFrameworkSpec:
  name: 'standalone'
  segmentGenerationJobRunnerClassName: 'org.apache.pinot.plugin.ingestion.batch.standalone.SegmentGenerationJobRunner'
  segmentTarPushJobRunnerClassName: 'org.apache.pinot.plugin.ingestion.batch.standalone.SegmentTarPushJobRunner'
  segmentUriPushJobRunnerClassName: 'org.apache.pinot.plugin.ingestion.batch.standalone.SegmentUriPushJobRunner'
  segmentMetadataPushJobRunnerClassName: 'org.apache.pinot.plugin.ingestion.batch.standalone.SegmentMetadataPushJobRunner'
jobType: SegmentCreationAndMetadataPush
inputDirURI: 'oss://your-bucket-name/input'
includeFileNamePattern: 'glob:**/*.csv'
outputDirURI: 'oss://your-bucket-name/output'
overwriteOutput: true
pinotFSSpecs:
  - scheme: oss
    className: org.apache.pinot.plugin.filesystem.HadoopPinotFS
    configs:
      hadoop.conf.path: '/path/to/hadoop/conf'
recordReaderSpec:
  dataFormat: 'csv'
  className: 'org.apache.pinot.plugin.inputformat.csv.CSVRecordReader'
  configClassName: 'org.apache.pinot.plugin.inputformat.csv.CSVRecordReaderConfig'
tableSpec:
  tableName: 'transcript'
pinotClusterSpecs:
  - controllerURI: 'http://localhost:9000'

더 알아보기 (Learn more)