Modin 퀵스타트 — 60배 빨라지는 concat 예제

Modin 퀵스타트 — 60배 빨라지는 concat 예제

Modin이 실제로 얼마나 빨라지는지 한눈에 보려면 concatapply 를 큰 데이터셋으로 비교해 보는 게 가장 확실해요. 여기서는 NYC 택시 데이터(약 200MB)를 예로 들어요.

출처: https://modin.readthedocs.io/en/stable/getting_started/quickstart.html

설치와 import 준비부터 할게요.

pip install "modin[all]"
import modin.pandas as pd
import pandas
import ray
ray.init()

비교를 위해 데이터를 받아 두 개의 데이터프레임을 만들어요. 하나는 pandas, 하나는 Modin(pd)이에요.

25개를 이어 붙이는 concat 을 비교하면 Modin이 60배 이상 빨라져요. pandas는 1분 가까이 걸리는 일을 Modin은 1초 미만으로 끝내요.

big_pandas_df = pandas.concat([pandas_df for _ in range(25)])
big_modin_df   = pd.concat([modin_df for _ in range(25)])

단일 열에 apply 를 돌려 반올림하는 작업은 더 극적이에요. 1억 3천만 개가 넘는 행을 1초 만에 처리하면서 30배 이상 빨라져요.

rounded_trip_distance_pandas = big_pandas_df["trip_distance"].apply(round)
rounded_trip_distance_modin   = big_modin_df["trip_distance"].apply(round)

이 예제에서는 100MB 수준의 데이터에서 시작해서 20GB 수준까지 늘리면서도 코드를 전혀 바꾸지 않았어요. read_csv, concat, apply 외에도 Modin은 pandas API의 90% 이상을 지원해서 흔한 연산 대부분에서 속도 향상을 누릴 수 있어요.

참고: MODIN_ENGINE 환경변수로 Ray·Dask·Unidist 중 어떤 엔진을 쓸지 고를 수 있어요.

더 알아보기