Pinot의 딥 스토리지로 OSS 사용하기
Pinot의 딥 스토리지로 OSS 사용하기
AliCloud Object Storage Service(OSS)를 Pinot 딥 스토리지로 구성하는 방법을 알려드려요.
출처: 문서
본문
OSS 파일시스템 플러그인을 구현하지 않고도 OSS를 Apache Pinot의 HDFS 딥 스토리지로 사용할 수 있어요. 아래 단계를 따르세요:
1. hdfs-site.xml과 core-site.xml 파일을 구성하세요. 그 후 이 구성을 임의 경로 아래에 두고, controller/server 구성의 pinot.<node>.storage.factory.oss.hadoop.conf 값에 이 경로를 설정하세요.
hdfs-site.xml의 경우 아무 구성도 제공할 필요가 없어요:
<?xml version="1.0" encoding="UTF-8"?>
<configuration>
</configuration>
core-site.xml의 경우 아래처럼 OSS access/secret과 bucket 구성을 제공해야 해요:
<?xml version="1.0"?>
<configuration>
<property>
<name>fs.defaultFS</name>
<value>oss://your-bucket-name/</value>
</property>
<property>
<name>fs.oss.accessKeyId</name>
<value>your-access-key-id</value>
</property>
<property>
<name>fs.oss.accessKeySecret</name>
<value>your-access-key-secret</value>
</property>
<property>
<name>fs.oss.impl</name>
<value>com.aliyun.emr.fs.oss.OssFileSystem</value>
</property>
<property>
<name>fs.oss.endpoint</name>
<value>your-oss-endpoint</value>
</property>
</configuration>
2. OSS에 접근하기 위해 OSS와 관련된 HDFS jar들을 찾아 PINOT_DIR/lib 아래에 두세요. 아래 jar를 사용할 수 있지만 충돌을 피하기 위해 버전에 주의하세요.
- smartdata-aliyun-oss
- smartdata-hadoop-common
- guava
3. controller.conf와 server.conf에 OSS 딥 스토리지 구성을 설정하세요:
Controller 구성
controller.data.dir=oss://your-bucket-name/path/to/segments
controller.local.temp.dir=/path/to/local/temp/directory
controller.enable.split.commit=true
pinot.controller.storage.factory.class.oss=org.apache.pinot.plugin.filesystem.HadoopPinotFS
pinot.controller.storage.factory.oss.hadoop.conf.path=path/to/conf/directory/
pinot.controller.segment.fetcher.protocols=file,http,oss
pinot.controller.segment.fetcher.oss.class=org.apache.pinot.common.utils.fetcher.PinotFSSegmentFetcher
Server 구성
pinot.server.instance.enable.split.commit=true
pinot.server.storage.factory.class.oss=org.apache.pinot.plugin.filesystem.HadoopPinotFS
pinot.server.storage.factory.oss.hadoop.conf.path=path/to/conf/directory/
pinot.server.segment.fetcher.protocols=file,http,oss
pinot.server.segment.fetcher.oss.class=org.apache.pinot.common.utils.fetcher.PinotFSSegmentFetcher
예시 작업 스펙
같은 HDFS 딥 스토리지 구성과 jar를 사용해 OSS에서 데이터를 읽고, 세그먼트를 만들어 다시 OSS로 push할 수 있어요. 예시 standalone 배치 수집 작업은 아래와 같을 수 있어요:
executionFrameworkSpec:
name: 'standalone'
segmentGenerationJobRunnerClassName: 'org.apache.pinot.plugin.ingestion.batch.standalone.SegmentGenerationJobRunner'
segmentTarPushJobRunnerClassName: 'org.apache.pinot.plugin.ingestion.batch.standalone.SegmentTarPushJobRunner'
segmentUriPushJobRunnerClassName: 'org.apache.pinot.plugin.ingestion.batch.standalone.SegmentUriPushJobRunner'
segmentMetadataPushJobRunnerClassName: 'org.apache.pinot.plugin.ingestion.batch.standalone.SegmentMetadataPushJobRunner'
jobType: SegmentCreationAndMetadataPush
inputDirURI: 'oss://your-bucket-name/input'
includeFileNamePattern: 'glob:**/*.csv'
outputDirURI: 'oss://your-bucket-name/output'
overwriteOutput: true
pinotFSSpecs:
- scheme: oss
className: org.apache.pinot.plugin.filesystem.HadoopPinotFS
configs:
hadoop.conf.path: '/path/to/hadoop/conf'
recordReaderSpec:
dataFormat: 'csv'
className: 'org.apache.pinot.plugin.inputformat.csv.CSVRecordReader'
configClassName: 'org.apache.pinot.plugin.inputformat.csv.CSVRecordReaderConfig'
tableSpec:
tableName: 'transcript'
pinotClusterSpecs:
- controllerURI: 'http://localhost:9000'