Selenium 스크래퍼

Selenium 스크래퍼 (SeleniumScrapingTool)

CrewAI의 SeleniumScrapingTool은 고효율 웹 스크래핑 태스크를 위해 설계된 도구예요. CSS 선택자를 사용해 웹 페이지의 특정 요소를 정밀하게 추출할 수 있으며, JavaScript로 렌더링되는 동적 콘텐츠가 있는 사이트를 스크래핑할 때 특히 유용하답니다.

참고: 이 도구는 현재 개발 중입니다. 기능을 다듬는 과정에서 예상치 못한 동작이 발생할 수 있어요. 여러분의 피드백이 개선에 큰 도움이 됩니다.

출처: 문서

본문

SeleniumScrapingTool은 고효율 웹 스크래핑 태스크를 위해 제작되었습니다. CSS 선택자로 특정 요소를 겨냥해 웹 페이지에서 콘텐츠를 정밀하게 추출할 수 있어요. 어떤 웹사이트 URL과도 작동하는 유연성을 제공하며 다양한 스크래핑 요구를 충족합니다.

설치 (Installation)

이 도구를 사용하려면 CrewAI tools 패키지와 Selenium을 설치해야 합니다.

pip install 'crewai[tools]'
uv add selenium webdriver-manager

이 도구는 브라우저 자동화에 Chrome WebDriver를 사용하므로, 시스템에 Chrome이 설치되어 있어야 합니다.

예시 (Example)

다음 예시는 CrewAI 에이전트와 함께 SeleniumScrapingTool을 사용하는 방법을 보여줍니다.

from crewai import Agent, Task, Crew, Process
from crewai_tools import SeleniumScrapingTool

# 도구 초기화
selenium_tool = SeleniumScrapingTool()

# 이 도구를 사용하는 에이전트 정의
web_scraper_agent = Agent(
    role="Web Scraper",
    goal="Extract information from websites using Selenium",
    backstory="An expert web scraper who can extract content from dynamic websites.",
    tools=[selenium_tool],
    verbose=True,
)

# 웹사이트에서 콘텐츠를 스크래핑하는 예시 태스크
scrape_task = Task(
    description="Extract the main content from the homepage of example.com. Use the CSS selector 'main' to target the main content area.",
    expected_output="The main content from example.com's homepage.",
    agent=web_scraper_agent,
)

# 크루 생성 및 실행
crew = Crew(
    agents=[web_scraper_agent],
    tasks=[scrape_task],
    verbose=True,
    process=Process.sequential,
)
result = crew.kickoff()

미리 정의된 파라미터로 도구를 초기화할 수도 있습니다.

# 미리 정의된 파라미터로 도구 초기화
selenium_tool = SeleniumScrapingTool(
    website_url='https://example.com',
    css_element='.main-content',
    wait_time=5
)

# 이 도구를 사용하는 에이전트 정의
web_scraper_agent = Agent(
    role="Web Scraper",
    goal="Extract information from websites using Selenium",
    backstory="An expert web scraper who can extract content from dynamic websites.",
    tools=[selenium_tool],
    verbose=True,
)

파라미터 (Parameters)

SeleniumScrapingTool은 초기화 시 다음 파라미터를 받습니다.

  • website_url: 선택. 스크래핑할 웹사이트의 URL. 초기화 시 제공하면 에이전트가 도구를 사용할 때 지정할 필요가 없습니다.
  • css_element: 선택. 추출할 요소의 CSS 선택자. 초기화 시 제공하면 에이전트가 도구를 사용할 때 지정할 필요가 없습니다.
  • cookie: 선택. 쿠키 정보를 담은 딕셔너리. 제한된 콘텐츠에 접근하기 위해 로그인 세션을 시뮬레이션할 때 유용.
  • wait_time: 선택. 스크래핑 전 지연(초)을 지정해 웹사이트와 동적 콘텐츠가 완전히 로드되도록 합니다. 기본값 3초.
  • return_html: 선택. 텍스트 대신 HTML 콘텐츠를 반환할지 여부. 기본값 False.

에이전트와 함께 도구를 사용할 때 에이전트는 다음 파라미터를 제공해야 합니다 (초기화 시 지정하지 않은 경우).

  • website_url: 필수. 스크래핑할 웹사이트의 URL.
  • css_element: 필수. 추출할 요소의 CSS 선택자.

에이전트 통합 예시 (Agent Integration Example)

SeleniumScrapingTool을 CrewAI 에이전트와 통합하는 더 자세한 예시는 다음과 같습니다.

from crewai import Agent, Task, Crew, Process
from crewai_tools import SeleniumScrapingTool

# 도구 초기화
selenium_tool = SeleniumScrapingTool()

# 이 도구를 사용하는 에이전트 정의
web_scraper_agent = Agent(
    role="Web Scraper",
    goal="Extract and analyze information from dynamic websites",
    backstory="""You are an expert web scraper who specializes in extracting
    content from dynamic websites that require browser automation. You have
    extensive knowledge of CSS selectors and can identify the right selectors
    to target specific content on any website.""",
    tools=[selenium_tool],
    verbose=True,
)

# 에이전트를 위한 태스크 생성
scrape_task = Task(
    description="""
    Extract the following information from the news website at {website_url}:

    1. The headlines of all featured articles (CSS selector: '.headline')
    2. The publication dates of these articles (CSS selector: '.pub-date')
    3. The author names where available (CSS selector: '.author')

    Compile this information into a structured format with each article's details grouped together.
    """,
    expected_output="A structured list of articles with their headlines, publication dates, and authors.",
    agent=web_scraper_agent,
)

# 태스크 실행
crew = Crew(
    agents=[web_scraper_agent],
    tasks=[scrape_task],
    verbose=True,
    process=Process.sequential,
)
result = crew.kickoff(inputs={"website_url": "https://news-example.com"})

구현 세부 사항 (Implementation Details)

SeleniumScrapingTool은 Selenium WebDriver를 사용해 브라우저 상호작용을 자동화합니다.

class SeleniumScrapingTool(BaseTool):
    name: str = "Read a website content"
    description: str = "A tool that can be used to read a website content."
    args_schema: Type[BaseModel] = SeleniumScrapingToolSchema

    def _run(self, **kwargs: Any) -> Any:
        website_url = kwargs.get("website_url", self.website_url)
        css_element = kwargs.get("css_element", self.css_element)
        return_html = kwargs.get("return_html", self.return_html)
        driver = self._create_driver(website_url, self.cookie, self.wait_time)

        content = self._get_content(driver, css_element, return_html)
        driver.close()

        return "\n".join(content)

이 도구는 다음 단계를 수행합니다.

  1. 헤드리스 Chrome 브라우저 인스턴스 생성
  2. 지정된 URL로 이동
  3. 페이지가 로드되도록 지정된 시간 동안 대기
  4. 제공된 쿠키가 있으면 추가
  5. CSS 선택자를 기반으로 콘텐츠 추출
  6. 추출된 콘텐츠를 텍스트 또는 HTML로 반환
  7. 브라우저 인스턴스 닫기

동적 콘텐츠 처리 (Handling Dynamic Content)

SeleniumScrapingTool은 JavaScript로 로드되는 동적 콘텐츠가 있는 웹사이트를 스크래핑할 때 특히 유용합니다. 실제 브라우저 인스턴스를 사용하면:

  1. 페이지에서 JavaScript 실행
  2. 동적 콘텐츠가 로드될 때까지 대기
  3. 필요하면 요소와 상호작용
  4. 단순 HTTP 요청으로는 얻을 수 없는 콘텐츠 추출

추출 전에 모든 동적 콘텐츠가 로드되도록 wait_time 파라미터를 조정할 수 있어요.

결론 (Conclusion)

SeleniumScrapingTool은 브라우저 자동화를 사용해 웹사이트에서 콘텐츠를 추출하는 강력한 방법을 제공합니다. 에이전트가 실제 사용자처럼 웹사이트와 상호작용할 수 있게 함으로써, 더 단순한 방법으로는 어렵거나 불가능한 동적 콘텐츠 스크래핑을 가능하게 해줘요. JavaScript로 렌더링되는 콘텐츠가 있는 현대 웹 애플리케이션을 다루는 리서치, 데이터 수집, 모니터링 태스크에 특히 유용합니다.

더 알아보기 (Learn more)