5분 만에 의미 검색 엔진 구축하기
5분 만에 의미 검색 엔진 구축하기 (tutorials-basics-search-beginners-local)
| 시간: 5 - 15분 | 난이도: 초급 |
|---|
이 튜토리얼에는 두 가지 버전이 있어요:
- 이 페이지의 버전으로는 자신의 컴퓨터에서 Qdrant를 직접 실행해요. 이 경우 클러스터와 벡터 임베딩 인프라를 직접 관리해야 해요.
- 또는 Qdrant Cloud를 사용해 클러스터를 배포하고, Qdrant Cloud의 평생 무료(forever free) 티어로 벡터 임베딩을 생성할 수도 있어요 (신용카드 불필요). 이 옵션을 선호한다면 Qdrant Cloud 버전의 튜토리얼을 확인해 보세요.
출처: Qdrant 공식문서 - Build a Semantic Search Engine in 5 Minutes
개요 (Overview)
벡터 검색 엔진이 처음이라면 이 튜토리얼이 딱이에요. 5분 안에 공상과학(SF) 소설을 위한 의미 검색 엔진을 구축할 거예요. 설정을 끝낸 뒤에는 엔진에게 다가오는 외계인 위협에 대해 물어볼 거예요. 여러분이 만든 검색 엔진이 잠재적인 우주 공격에 대비할 책들을 추천해 줄 거예요.
시작하기 전에 최신 버전의 Python이 설치돼 있어야 해요. 가상 환경에서 이 코드를 실행하는 방법을 모른다면, Python 문서의 Creating Virtual Environments를 먼저 따라 해 보세요.
이 튜토리얼은 bash 셸을 사용한다고 가정해요. Python 문서를 참고해 가상 환경을 활성화하세요. 예를 들면:
source tutorial-env/bin/activate
1. 설치 (Installation)
검색 엔진이 데이터를 처리할 수 있도록 데이터를 가공해야 해요. Sentence Transformers 프레임워크를 사용하면 원시 데이터를 임베딩으로 바꿔 주는 흔한 대규모 언어 모델들에 접근할 수 있어요.
pip install -U sentence-transformers
인코딩된 데이터는 어딘가에 보관해야 해요. Qdrant를 사용하면 데이터를 임베딩으로 저장할 수 있어요. 이 데이터에 대해 검색 쿼리를 실행하는 데도 Qdrant를 사용할 수 있어요. 즉 엔진에게 키워드 매칭을 훨씬 넘어서는 관련성 있는 답변을 요청할 수 있다는 뜻이에요.
pip install -U qdrant-client
모델 가져오기 (Import the Models)
두 가지 주요 프레임워크가 정의되면, 이 엔진이 사용할 정확한 모델을 지정해야 해요.
from qdrant_client import models , QdrantClient from sentence_transformers import SentenceTransformer
Sentence Transformers 프레임워크에는 많은 임베딩 모델이 들어 있어요. 이 튜토리얼에서는 all-MiniLM-L6-v2를 사용할게요. 속도와 임베딩 품질이 이 튜토리얼에 맞게 잘 균형 잡혀 있거든요.
encoder = SentenceTransformer ( "all-MiniLM-L6-v2" )
2. 데이터셋 추가 (Add the Dataset)
all-MiniLM-L6-v2는 여러분이 제공한 데이터를 인코딩해요. 여기서는 도서관에 있는 모든 공상과학 소설을 나열할 거예요. 각 책에는 메타데이터인 이름, 저자, 출판 연도, 짧은 설명이 있어요.
documents = [ { "name" : "The Time Machine" , "description" : "A man travels through time and witnesses the evolution of humanity." , "author" : "H.G. Wells" , "year" : 1895 , }, { "name" : "Ender's Game" , "description" : "A young boy is trained to become a military leader in a war against an alien race." , "author" : "Orson Scott Card" , "year" : 1985 , }, { "name" : "Brave New World" , "description" : "A dystopian society where people are genetically engineered and conditioned to conform to a strict social hierarchy." , "author" : "Aldous Huxley" , "year" : 1932 , }, { "name" : "The Hitchhiker's Guide to the Galaxy" , "description" : "A comedic science fiction series following the misadventures of an unwitting human and his alien friend." , "author" : "Douglas Adams" , "year" : 1979 , }, { "name" : "Dune" , "description" : "A desert planet is the site of political intrigue and power struggles." , "author" : "Frank Herbert" , "year" : 1965 , }, { "name" : "Foundation" , "description" : "A mathematician develops a science to predict the future of humanity and works to save civilization from collapse." , "author" : "Isaac Asimov" , "year" : 1951 , }, { "name" : "Snow Crash" , "description" : "A futuristic world where the internet has evolved into a virtual reality metaverse." , "author" : "Neal Stephenson" , "year" : 1992 , }, { "name" : "Neuromancer" , "description" : "A hacker is hired to pull off a near-impossible hack and gets pulled into a web of intrigue." , "author" : "William Gibson" , "year" : 1984 , }, { "name" : "The War of the Worlds" , "description" : "A Martian invasion of Earth throws humanity into chaos." , "author" : "H.G. Wells" , "year" : 1898 , }, { "name" : "The Hunger Games" , "description" : "A dystopian society where teenagers are forced to fight to the death in a televised spectacle." , "author" : "Suzanne Collins" , "year" : 2008 , }, { "name" : "The Andromeda Strain" , "description" : "A deadly virus from outer space threatens to wipe out humanity." , "author" : "Michael Crichton" , "year" : 1969 , }, { "name" : "The Left Hand of Darkness" , "description" : "A human ambassador is sent to a planet where the inhabitants are genderless and can change gender at will." , "author" : "Ursula K. Le Guin" , "year" : 1969 , }, { "name" : "The Three-Body Problem" , "description" : "Humans encounter an alien civilization that lives in a dying system." , "author" : "Liu Cixin" , "year" : 2008 , }, ]
3. 저장 위치 정의 (Define Storage Location)
Qdrant에 임베딩을 어디에 저장할지 알려줘야 해요. 이건 기본 데모이므로 로컬 컴퓨터의 메모리를 임시 저장소로 사용할게요.
client = QdrantClient ( ":memory:" )
4. 컬렉션 생성 (Create a Collection)
Qdrant의 모든 데이터는 컬렉션으로 구성돼요. 이 경우 책을 저장하고 있으므로 컬렉션 이름을 my_books라고 부를게요.
client . create_collection ( collection_name = "my_books" , vectors_config = models . VectorParams ( size = encoder . get_sentence_embedding_dimension (), # Vector size is defined by used model distance = models . Distance . COSINE , ), )
vector_size매개변수는 특정 컬렉션의 벡터 크기를 정의해요. 크기가 다르다면 벡터 사이의 거리를 계산하는 것이 불가능해요. 384는 인코더의 출력 차원이에요. 사용 중인 모델의 차원을 얻으려면model.get_sentence_embedding_dimension()을 사용할 수도 있어요.distance매개변수는 두 점 사이의 거리를 측정하는 데 사용할 함수를 지정할 수 있게 해 줘요.
5. 데이터를 컬렉션에 업로드 (Upload Data to Collection)
데이터베이스에게 documents를 my_books 컬렉션에 업로드하라고 알려줘요. 이렇게 하면 각 레코드에 id와 페이로드가 부여돼요. 페이로드는 데이터셋의 메타데이터일 뿐이에요.
client . upload_points ( collection_name = "my_books" , points = [ models . PointStruct ( id = idx , vector = encoder . encode ( doc [ "description" ]) . tolist (), payload = doc ) for idx , doc in enumerate ( documents ) ], )
6. 엔진에게 질문하기 (Ask the Engine a Question)
이제 데이터가 Qdrant에 저장됐으니 질문을 던지고 의미적으로 관련된 결과를 받을 수 있어요.
hits = client . query_points ( collection_name = "my_books" , query = encoder . encode ( "alien invasion" ) . tolist (), limit = 3 , ) . points for hit in hits : print ( hit . payload , "score:" , hit . score )
결과 (Response):
검색 엔진은 외계인 침공과 관련된 가장 가능성 높은 세 가지 응답을 보여줘요. 각 응답에는 원래 질문과 얼마나 가까운지를 보여주는 점수가 부여돼요.
{'name': 'The War of the Worlds', 'description': 'A Martian invasion of Earth throws humanity into chaos.', 'author': 'H.G. Wells', 'year': 1898} score: 0.570093257022374 {'name': "The Hitchhiker's Guide to the Galaxy", 'description': 'A comedic science fiction series following the misadventures of an unwitting human and his alien friend.', 'author': 'Douglas Adams', 'year': 1979} score: 0.5040468703143637 {'name': 'The Three-Body Problem', 'description': 'Humans encounter an alien civilization that lives in a dying system.', 'author': 'Liu Cixin', 'year': 2008} score: 0.45902943411768216
쿼리 좁히기 (Narrow Down the Query)
2000년대 초반의 가장 최근 책은 어떨까요?
hits = client . query_points ( collection_name = "my_books" , query = encoder . encode ( "alien invasion" ) . tolist (), query_filter = models . Filter ( must = [ models . FieldCondition ( key = "year" , range = models . Range ( gte = 2000 ))] ), limit = 1 , ) . points for hit in hits : print ( hit . payload , "score:" , hit . score )
결과 (Response):
쿼리가 2008년 결과 하나로 좁혀졌어요.
{'name': 'The Three-Body Problem', 'description': 'Humans encounter an alien civilization that lives in a dying system.', 'author': 'Liu Cixin', 'year': 2008} score: 0.45902943411768216
다음 단계 (Next Steps)
축하해요, 여러분은 방금 최초의 검색 엔진을 만들었어요! 믿어 주세요, Qdrant의 나머지도 그렇게 복잡하지 않아요. 다음 튜토리얼로는 자체 하이브리드 검색 서비스 구축을 시도하거나, 무료 Qdrant Essentials 강좌를 수강해 보세요.