Skip to main content
This notebook provides a quick overview for getting started with FireCrawlLoader document loaders. For detailed documentation of all FireCrawlLoader features and configurations head to the API reference.

概述

集成详情

Loader features

FireCrawl crawls and convert any website into LLM-ready data. It crawls all accessible sub-pages and give you clean markdown and metadata for each. No sitemap required. FireCrawl handles complex tasks such as reverse proxies, caching, rate limits, and content blocked by JavaScript. Built by the mendable.ai team. This guide shows how to scrap and crawl entire websites and load them using the FireCrawlLoader in LangChain.

设置

要访问 FireCrawlLoader document loader,你需要install the @langchain/community integration, and the @mendable/firecrawl-js@0.0.36 package. Then create a FireCrawl account and get an API key.

凭证

Sign up and get your free FireCrawl API key to start. FireCrawl offers 300 free credits to get you started, and it’s open-source in case you want to self-host. 完成后设置 FIRECRAWL_API_KEY 环境变量:
如果你想要自动追踪模型调用,还可以设置你的 LangSmith API 密钥,取消注释以下内容:

安装

LangChain 的 FireCrawlLoader 集成位于 @langchain/community 包中:

实例化

Here’s an example of how to use the FireCrawlLoader to load web search results: Firecrawl offers 3 modes: scrape, crawl, and map. In scrape mode, Firecrawl will only scrape the page you provide. In crawl mode, Firecrawl will crawl the entire website. In map mode, Firecrawl will return semantic links related to the website. The formats (scrapeOptions.formats for crawl mode) parameter allows selection from "markdown", "html", or "rawHtml". However, the Loaded Document will return content in only one format, prioritizing as follows: markdown, then html, and finally rawHtml. Now we can instantiate our model object and load documents:

Load

Additional parameters

For params you can pass any of the params according to the Firecrawl documentation.

API 参考

有关所有 FireCrawlLoader 功能和配置的详细文档,请前往 API 参考