Firecrawl Integration Guide

Technology: firecrawl · Category: tooling · Last reviewed: 2026-08-23

Source: https://tech-stack.codeamanilabs.org/guide/firecrawl

Insight:

Firecrawl is the LLM-ready web data API — give it a URL (or a search query) and get back clean markdown / structured JSON, with proxies, anti-bot, and JavaScript rendering already handled. It's the managed alternative to hand-rolled Playwright/Puppeteer for RAG ingestion and agent web-browsing. Five primitives — scrape, crawl, map, search, extract — cover single-page, whole-site, URL-discovery, web-search, and schema-to-JSON. Bill per credit (1 credit = 1 page); map before crawl and onlyMainContent keep both cost and token counts low. Verified against docs.firecrawl.dev on 2026-08-23 — SDK v4.x (npm firecrawl 4.35.0 / pip firecrawl-py 4.38.0) on the v2 REST API. Docs install firecrawl; the legacy scoped alias @mendable/firecrawl-js still publishes the same 4.x releases in lockstep (it is not deprecated — only the v1 method names like crawlUrl are).

███████╗██╗██████╗ ███████╗ ██████╗██████╗  █████╗ ██╗    ██╗██╗
██╔════╝██║██╔══██╗██╔════╝██╔════╝██╔══██╗██╔══██╗██║    ██║██║
█████╗  ██║██████╔╝█████╗  ██║     ██████╔╝███████║██║ █╗ ██║██║
██╔══╝  ██║██╔══██╗██╔══╝  ██║     ██╔══██╗██╔══██║██║███╗██║██║
██║     ██║██║  ██║███████╗╚██████╗██║  ██║██║  ██║╚███╔███╔╝███████╗
╚═╝     ╚═╝╚═╝  ╚═╝╚══════╝ ╚═════╝╚═╝  ╚═╝╚═╝  ╚═╝ ╚══╝╚══╝ ╚══════╝

Firecrawl Integration Guide

Focus: Turning any URL — or a web search — into clean, LLM-ready markdown or schema-validated JSON, with proxies, anti-bot, and JS rendering handled for you. The managed alternative to hand-rolled Playwright/Puppeteer for RAG and agents.

Verification note (2026-08-23): Every endpoint path, SDK method name, credit cost, price, and rate-limit number below was live-verified against docs.firecrawl.dev. Firecrawl is on the v2 REST API, driven by the v4.x SDKs — firecrawl (npm, 4.35.0) / firecrawl-py (pip, 4.38.0). The docs install the unscoped firecrawl package; the older scoped name @mendable/firecrawl-js is not deprecated — it publishes the identical 4.x releases in lockstep (both point at github.com/firecrawl/firecrawl). What is deprecated is the v1 SDK method names (crawlUrl, scrapeUrl, asyncCrawlUrl) from SDK ≤1.x. Two facts shifted since the last review: prices rose (Standard $49.99→$83/mo) and enhanced/"stealth" proxy no longer costs +4 credits (now 1 credit, same as basic). Pricing dollar figures below reflect the annual-billing effective monthly rate; monthly-billed is higher — re-check firecrawl.dev/pricing before quoting a client.

Overview

Firecrawl is a web data API for AI. You hand it a URL and it returns the page as clean markdown, raw/processed HTML, a screenshot, a link list, or structured JSON — having already solved the parts that make DIY scraping painful: rotating proxies, anti-bot challenges, JavaScript/SPA rendering, PDFs, and dynamic content. It exposes five core primitives plus interactive browser control, all behind one API key.

For codeAmani, Firecrawl is the "web → context" lever: it's how you feed external pages into a RAG pipeline (pairs with the pinecone / pgvector guides), how an agent reads a page it was asked about, and how you pull structured facts (prices, listings, docs) off sites that have no API.

When Firecrawl wins — and when raw Playwright/Cheerio wins

flowchart TD
  A["Need data off a web page"] --> B{"Do you control the<br/>site / have an API?"}
  B -->|"yes"| C["Use the API / DB directly<br/>(don't scrape)"]
  B -->|"no"| D{"LLM-ready output?<br/>proxies + anti-bot?<br/>many sites?"}
  D -->|"yes — RAG / agents"| E["Firecrawl<br/>managed, per-credit"]
  D -->|"no — 1 static site,<br/>full DOM control, free"| F["Cheerio / Playwright<br/>self-hosted"]
  E --> G["scrape · crawl · map<br/>search · extract"]
  F --> H["You own proxies,<br/>retries, JS, parsing"]
Tool Reach for it when… Cost shape
Firecrawl RAG ingestion, agent web-browsing, scraping many sites, JS-heavy pages, you want markdown/JSON not HTML, you don't want to run proxy/anti-bot infra Per-credit (managed)
Cheerio One known static site, server-rendered HTML, you only need a few selectors, zero budget Free (your CPU)
Playwright / Puppeteer You need full programmatic browser control, custom auth flows, screenshots of your own app, and you're happy to operate proxies + anti-bot yourself Free (your infra)

Rule of thumb: Firecrawl converts the web into LLM input; Playwright/Cheerio give you a browser/parser you operate yourself. Firecrawl even offers a Browser Sandbox (managed Playwright-over-CDP) when you do need raw browser control without running the infra.

Primary use cases

Official Documentation

Resource URL
Introduction https://docs.firecrawl.dev/introduction
Node SDK https://docs.firecrawl.dev/sdks/node
Python SDK https://docs.firecrawl.dev/sdks/python
Crawl feature (async + webhooks) https://docs.firecrawl.dev/features/crawl
Parse files (PDF/DOCX/XLSX) https://docs.firecrawl.dev/features/parse
Enhanced ("stealth") proxy mode https://docs.firecrawl.dev/features/enhanced-mode
Rate limits & concurrency https://docs.firecrawl.dev/rate-limits
Webhooks & signature verification https://docs.firecrawl.dev/webhooks/overview
MCP server https://docs.firecrawl.dev/mcp-server
Open source / self-host https://docs.firecrawl.dev/contributing/open-source-or-cloud

Setup

Get an API key at firecrawl.dev/app/api-keys (keys are prefixed fc-). Both SDKs read FIRECRAWL_API_KEY from the environment automatically, or you can pass it explicitly.

Node / TypeScript

npm install firecrawl        # v4.x — the docs-recommended package name
# @mendable/firecrawl-js is the legacy scoped alias for the same 4.x releases
import { Firecrawl } from "firecrawl";

const firecrawl = new Firecrawl({ apiKey: process.env.FIRECRAWL_API_KEY });

// Single page → clean markdown
const doc = await firecrawl.scrape("https://firecrawl.dev", {
  formats: ["markdown"],
  onlyMainContent: true,
});
console.log(doc.markdown);

Python

pip install firecrawl-py
from firecrawl import Firecrawl  # AsyncFirecrawl is also exported

firecrawl = Firecrawl(api_key="fc-YOUR-API-KEY")  # or omit to read FIRECRAWL_API_KEY

doc = firecrawl.scrape("https://firecrawl.dev", formats=["markdown"], only_main_content=True)
print(doc.markdown)

Node uses camelCase option keys (onlyMainContent, scrapeOptions); Python uses snake_case (only_main_content, scrape_options). The REST API itself is camelCase.


Core Endpoints

All endpoints live under https://api.firecrawl.dev/v2/ and authenticate with Authorization: Bearer fc-....

const doc = await firecrawl.scrape("https://example.com/article", {
  formats: ["markdown", "html", "links", "screenshot"],
  onlyMainContent: true,          // strip nav/footer/ads — fewer tokens
  includeTags: ["article", "main"],
  excludeTags: ["nav", "footer", ".ad"],
  maxAge: 600000,                 // serve from cache if scraped < 10 min ago
});

Response (SDKs return the data object directly; cURL wraps it in { success, data }):

{
  "markdown": "# Article title\n\nClean body text…",
  "html": "<!DOCTYPE html>…",
  "links": ["https://example.com/next", "…"],
  "screenshot": "https://…/screenshot.png",
  "metadata": {
    "title": "Article title",
    "sourceURL": "https://example.com/article",
    "statusCode": 200,
    "scrapeId": "019eb884-…"
  }
}

Output formats: markdown, html, rawHtml, links, screenshot, summary, json, changeTracking, branding. Use onlyMainContent: true plus includeTags/excludeTags to cut boilerplate (and token count) before it ever reaches your LLM.

Actions (dynamic pages — click, scroll, wait, input)

Pass an actions array to drive a real browser before extraction. v2 action shape (verified):

const doc = await firecrawl.scrape("https://example.com", {
  actions: [
    { type: "wait", milliseconds: 1000 },
    { type: "click", selector: "#accept" },
    { type: "scroll", direction: "down" },
    { type: "click", selector: "#q" },
    { type: "write", text: "firecrawl" },   // text input
    { type: "press", key: "Enter" },
    { type: "wait", milliseconds: 2000 },
    { type: "screenshot" },
  ],
  formats: ["markdown"],
});

Action types: wait (milliseconds), click (selector), scroll (direction), write (text), press (key), screenshot. For a persistent interactive session, use interact(scrapeId, …) / stopInteraction(scrapeId) against a prior scrape's metadata.scrapeId, or the Browser Sandbox (firecrawl.browser(...) → CDP URL for full Playwright).

/map — fast sitemap / URL discovery

const res = await firecrawl.map("https://docs.firecrawl.dev", {
  search: "webhook",   // optional: rank URLs by relevance
  limit: 100,
});
console.log(res.links); // ["https://docs.firecrawl.dev/webhooks/overview", …]

map is the cheap reconnaissance step: discover the URL list first, decide what's worth scraping, then crawl/scrape only those — instead of crawling blind.

/crawl — recursive site crawl (async job + status polling)

Crawl-and-wait (handles the job + pagination for you — recommended):

const job = await firecrawl.crawl("https://docs.firecrawl.dev", {
  limit: 100,                       // default is 10,000 — always set a limit
  includePaths: ["^/features/.*"],  // regex on pathname
  excludePaths: ["^/blog/.*"],
  maxDiscoveryDepth: 3,
  sitemap: "include",               // "include" | "skip" | "only"
  scrapeOptions: { formats: ["markdown"], onlyMainContent: true },
});
console.log(job.status, job.data.length); // "completed", N docs

Start-and-poll (long crawls / custom polling):

const { id } = await firecrawl.startCrawl("https://docs.firecrawl.dev", { limit: 500 });
const status = await firecrawl.getCrawlStatus(id);
// status.status ∈ "scraping" | "completed" | "failed"; status.completed / status.total
// status.data = pages scraped so far; cancel with firecrawl.cancelCrawl(id)

Python mirrors this: firecrawl.crawl(url, limit=…, scrape_options=ScrapeOptions(...)), firecrawl.start_crawl(...), firecrawl.get_crawl_status(job.id).

Crawl gotchas (verified): default limit is 10,000 and the endpoint pre-checks that your credit balance covers it — set a real limit or you'll hit 402 Payment Required. By default crawl only follows children of the start path; use crawlEntireDomain, allowSubdomains, or allowExternalLinks to widen. Job results are retrievable via the API for 24 hours; after that, use the activity logs. data holds pages Firecrawl successfully scraped (even if the site returned 404) — fetch hard failures via the Get Crawl Errors endpoint (GET /crawl/{id}/errors).

/search — web search → scraped results

const results = await firecrawl.search("best dash cams 2026", {
  limit: 5,
  sources: ["web", "news", "images"],
  tbs: "qdr:m",                              // time filter: past month
  scrapeOptions: { formats: ["markdown"] },  // scrape each result inline
});
// results.web[] = { url, title, description, position, (markdown if scraped) }

One call searches the web and returns full page content for each hit — no separate scrape loop.

/extract — LLM structured extraction (schema → JSON)

const res = await firecrawl.extract({
  urls: ["https://example-forum.com/topic/123"],
  prompt: "Extract all user comments from this thread.",
  schema: {
    type: "object",
    properties: {
      comments: {
        type: "array",
        items: {
          type: "object",
          properties: { author: { type: "string" }, comment_text: { type: "string" } },
          required: ["author", "comment_text"],
        },
      },
    },
    required: ["comments"],
  },
});
console.log(res.data);
from pydantic import BaseModel

class Product(BaseModel):
    name: str
    price: str

data = firecrawl.extract(
    urls=["https://shop.example.com/item/42"],
    prompt="Extract the product name and price.",
    schema=Product,                 # a Pydantic model or a raw JSON Schema
    enable_web_search=True,         # optionally enrich from related pages
)
print(data.data)

Caching, proxies/stealth, PDFs & dynamic content


Developer Resources

Official SDKs

Framework integrations

MCP server (yes — first-class)

Firecrawl ships an official MCP server so Claude, Cursor, Windsurf, and VS Code can call scrape/search/crawl/etc. directly. Two ways to connect:

See docs.firecrawl.dev/mcp-server. Tools now span web (scrape/search/crawl/map/extract), page interaction, monitoring, and research-paper search. (This guide's research was done through that exact MCP server.)

Self-hosting / open source

Firecrawl is open source (github.com/firecrawl/firecrawl, AGPL-licensed) and self-hostable — run the API on your own infra for data-residency or cost control. The hosted cloud adds managed proxies, anti-bot, scale, and higher reliability; the self-host build asks you to bring your own proxy/anti-bot. Decision guide: docs.firecrawl.dev/contributing/open-source-or-cloud.

Rate limits & concurrency (verified 2026-08-23)

Two independent limits; exceeding either returns 429:

API rate limits (requests/min, current plans):

Plan /scrape /map /crawl /search
Free 10 10 2 10
Hobby 100 100 20 100
Standard 500 500 100 500
Growth 5,000 5,000 1,000 5,000
Scale 10,000 10,000 2,000 10,000

Concurrent browsers (parallel jobs ceiling): Free 2, Hobby 5, Standard 50, Growth 100, Scale/Enterprise 150+. Max queued jobs scale with the plan (Free/Hobby 50k, Standard 100k, Growth 200k, Scale 300k+). Jobs beyond the concurrency ceiling queue (and queue time counts against the request timeout). Check live headroom with the Queue Status endpoint. (Note: the pricing page also advertises a lower per-plan "concurrent requests" figure — Standard 25 / Growth 50 / Scale 100 — which is a distinct metric from the concurrent-browser ceilings above.)

Firecrawl's own guidance: rate limits exist mainly to prevent abuse — your real bottleneck is concurrent browsers, so size the plan by concurrency, not req/min.

Retry / backoff

The SDKs auto-retry and handle async polling. For your own loops, treat 429 and 5xx as retryable with exponential backoff + jitter; for 429, prefer reducing concurrency over hammering. 402 means out of credits (raise limit awareness or enable auto-recharge), not a transient error.

Webhooks for async crawl completion (verified)

Attach a webhook object to a crawl to get pushed events instead of polling:

// POST https://api.firecrawl.dev/v2/crawl
{
  "url": "https://docs.firecrawl.dev",
  "limit": 100,
  "webhook": {
    "url": "https://your-domain.com/api/webhooks/firecrawl",
    "metadata": { "tenant": "acme" },
    "events": ["started", "page", "completed"]
  }
}

Event types: crawl.started, crawl.page, crawl.completed, crawl.failed. Every request carries an X-Firecrawl-Signature header (sha256=…) — an HMAC-SHA256 of the raw body using your webhook secret (from the dashboard Advanced tab). Verify it with a timing-safe compare before processing — see codeAmani notes below.


Pricing (verify live before quoting)

These change — re-check firecrawl.dev/pricing. Dollar figures below are the annual-billing effective monthly rate shown on the pricing page on 2026-08-23; monthly-billed is higher. Verified live. Prices rose since the June review (Standard $49.99→$83, Growth $149.99→$333).

Plan Price (annual, eff. /mo) Credits / mo Concurrent browsers
Free $0 (no card) 1,000 2
Hobby $16 5,000 5
Standard (recommended) $83 100,000 50
Growth $333 500,000 100
Scale $599 1,000,000 150
Enterprise Custom Custom Custom (SSO, ZDR, SLA)

Credit-per-action model (verified)

Action Credit cost
Scrape 1 / page
Crawl 1 / page
Map 1 / page
Search 2 / 10 results
Interact 2 / browser-minute
Monitor 1 / page / check
Enhanced/"stealth" proxy 1 / page (no surcharge — changed 2026)
JSON mode (LLM structured extraction on a page) +4 / page
PII redaction · audio/video extraction · question/highlights format +4 / page (each)
Zero-Data-Retention (ZDR) +1 / page
PDF parsing 1 / PDF page

So a plain scrape/crawl page — even with the anti-bot enhanced proxy — is 1 credit. The LLM add-ons are what cost: turning on structured JSON (or PII redaction / A-V extraction / question/highlights) roughly 5×'s the per-page cost. Enhanced proxy is no longer one of those multipliers. Budget for the LLM formats, not the proxy.

Overage behavior

No pure pay-as-you-go. Auto-recharge can auto-purchase additional credit packs when you dip below a threshold (larger packs = better rate). Credits do not roll over month-to-month; credit packs have their own billing periods. Downgrades take effect at the next renewal.


Usage Monitoring


codeAmani Notes


Troubleshooting

Issue Fix
401 Unauthorized Check FIRECRAWL_API_KEY (must start with fc-); confirm it's read server-side
402 Payment Required on crawl Credit balance can't cover limit — lower limit or enable auto-recharge
429 Too Many Requests Hitting rate or concurrency limit — back off with jitter; reduce concurrency or upgrade plan
Empty / nav-only markdown JS-rendered SPA — add waitFor: 5000, or actions to trigger content; try map to find the real content URL
Old crawlUrl() / scrapeUrl() errors Those are v1 SDK method names (SDK ≤1.x) — upgrade to the v4.x SDK and use scrape, crawl, getCrawlStatus. Either package name works (firecrawl or the still-maintained @mendable/firecrawl-js); the package isn't the problem, the method name is
Crawl missing sibling/parent pages Crawl follows children by default — set crawlEntireDomain / allowSubdomains
Crawl results gone after a day API retains job results for 24h; pull from activity logs after that
Surprise high credit bill JSON mode (and PII redaction / A-V / question / highlights) add +4 credits/page; PDFs bill per page — audit metadata.creditsUsed. (Enhanced proxy is not a multiplier any more — it's 1 credit.)

Verification Status (2026-08-23)

Live-verified against docs.firecrawl.dev (via Firecrawl's own scrape API + Context7 + npm/PyPI registries):

Official docs: