Skip to content

About

Docker Sandboxes kit that gives any sandboxed agent live web search, scrape and crawl via Firecrawl

Topics

Resources

Contributing

Stars

1 star

Watchers

0 watching

Forks

Repository files navigation

Firecrawl kit for Docker Sandboxes

An agent running in a Docker Sandbox scraping the web through Firecrawl

A Docker Sandboxes kit (kind: mixin) that adds live web access to any sandbox agent via the Firecrawl Python SDK (firecrawl-py).

Layer it onto whatever agent you run and it can search the web, scrape a page to clean markdown, and crawl a site, instead of answering from training-cutoff knowledge.

The kit ships in both kit formats, with the same declarations in each:

Format Descriptor Published as Composes onto
kit spec v2 spec.yaml docker.io/firecrawl/firecrawl-docker-sandbox the built-in sbx agents (claude, codex, …), see section 2
Kit v3 firecrawl/firecrawl.yaml docker.io/firecrawl/sbx-kit-firecrawl v3 workloads such as docker/sbx-kit-shell, see section 2b

The two lines cannot be mixed: a v3 mixin only composes onto a v3 workload, and the v2 mixin only onto the built-in agents. Pick the row that matches what you run.

What the kit does

Four observable things, so each is independently verifiable (see Verify the kit):

  1. Installs firecrawl-py==4.44.0 as the agent user (1000), and fails the sandbox create if the install or the post-install import check fails.
  2. Declares a firecrawl credential. The key is swapped into the Authorization header by the sbx proxy on requests to api.firecrawl.dev, and is never baked into the image or written to the sandbox.
  3. Allows egress to api.firecrawl.dev (plus PyPI, at install time) via permissions.network.allow (network-policy@1, split by phase, in the v3 kit). Kit network rules add to your sandbox policy; they do not narrow it. Under Docker's default balanced policy, general websites (including most documentation sites) return a proxy 403, so Firecrawl is the agent's route to them. Under deny-all, api.firecrawl.dev is the only web host the agent can reach.
  4. Ships an agentInstructions note (agent-context@1 in the v3 kit) so the agent knows the capability exists, how to call it, and which sandbox constraints apply to it.

0. Prerequisites

  • The sbx CLI, v0.45 or later (v0.46 or later for the v3 kit). Docker Desktop is not required. On macOS, per Docker's install docs: brew install docker/tap/sbx.
  • Sign in once with sbx login.
  • A global network policy, set once with sbx policy init balanced (or deny-all). Until this runs, every sbx run fails with global network policy has not been initialized.

Kits are experimental and in Early Access. sbx supports both kit spec v2 (this repo's spec.yaml) and Kit v3 (firecrawl/firecrawl.yaml), and the v3 specification itself is marked experimental until its final release, targeted for Q4 2026. v3 workloads and mixins cannot be combined with v2 kits, which is why both forms are published.

1. Store the Firecrawl API key

Get a key from firecrawl.dev (it looks like fc-...). Store it once with sbx's secret manager. The key never enters the kit; the proxy injects it at runtime, which is why sbx run has no -e flag:

echo "$FIRECRAWL_API_KEY" | sbx secret set firecrawl

Service secrets are global (available to every sandbox) by default; sbx before v0.46 needed a -g flag for that. Running sbx secret set firecrawl with no piped value prompts for the key interactively instead. Confirm it stored:

sbx secret ls

On the first sbx run with this kit, sbx asks you to approve sending the firecrawl credential to api.firecrawl.dev and records a binding in ~/.config/sbx/credentials.yaml. Because the value already lives in the secret store, accept the defaults; no env-var or file source is needed. The kit declares only what it needs and where to inject it, so you stay in control of where the key comes from.

The credential is marked required. In an interactive sbx run, that means sbx prompts for the binding. In a non-interactive create (--detached, CI) there is no prompt: sbx creates the sandbox with the credential withheld and prints WARN: credential not sent: no binding authorizes this service, and scrapes then fail with a 401 because the literal proxy-managed placeholder reaches Firecrawl. For unattended use, write the binding to ~/.config/sbx/credentials.yaml first:

bindings:
  firecrawl:
    apiKey:
      domains: [api.firecrawl.dev]

If sbx policy ls shows Governance: Managed by <org>, one more step is needed before scrapes work; see Centrally governed hosts.

2. Launch the sandbox with the kit

From a local clone (the kit lives at the repo root):

git clone https://github.com/firecrawl/firecrawl-docker-sandbox.git
sbx run --kit ./firecrawl-docker-sandbox/ claude

Straight from git, pinned to a commit. Kit references must pin a 40-hex commit SHA; branches and tags are rejected, and :latest on an OCI reference is rejected too:

git ls-remote https://github.com/firecrawl/firecrawl-docker-sandbox.git main   # resolve the SHA
sbx run --kit "git+https://github.com/firecrawl/firecrawl-docker-sandbox.git#ref=<commit-sha>" claude

From the published image. Every merge to main publishes docker.io/firecrawl/firecrawl-docker-sandbox:latest via scripts/push-kit.sh; sbx rejects :latest, so resolve the tag to a digest first. Use the plain inspect output: the --format path parses the kit's YAML config blob as JSON and fails:

digest=$(docker buildx imagetools inspect docker.io/firecrawl/firecrawl-docker-sandbox:latest | awk '/^Digest:/ {print $2}')
sbx run --kit "oci://docker.io/firecrawl/firecrawl-docker-sandbox@$digest" claude

Choosing the agent

The trailing argument (claude above) is the coding agent that runs inside the sandbox, a separate axis from the kit. Any supported agent works, and sbx run --help lists them:

claude, codex, copilot, cursor, devin, docker-agent, droid, gemini, kiro, opencode, shell

So claude can be swapped for codex:

sbx run --kit ./firecrawl-docker-sandbox/ codex

Arguments meant for the agent itself go after a -- separator, e.g. sbx run --kit ./firecrawl-docker-sandbox/ codex -- --help.

The kit only assumes the base image ships python3 and python3-pip, which every docker/sandbox-templates image does. On a custom base image without them, the install hook stops with a message saying so.

2b. Launch with a Kit v3 workload

Kit v3 kits are ordinary OCI images: the descriptor rides in the manifest, so there is no digest dance and a version tag is fine. Compose the v3 mixin onto a v3 workload from Docker's docker organization, for example the plain shell:

sbx run docker/sbx-kit-shell:1.0.0 --kit docker.io/firecrawl/sbx-kit-firecrawl:1.0.0 .

Other v3 workloads and agent mixins are listed on Docker Hub (the docker/* kits are v3; sbx/* is the v2 line). For the Claude Code agent on that shell, add its mixin alongside ours:

sbx run docker/sbx-kit-shell:1.0.0 --kit docker/sbx-kit-claude-mixin:2.1.285 --kit docker.io/firecrawl/sbx-kit-firecrawl:1.0.0 .

From a local clone, point --kit at the firecrawl/ directory; sbx builds source-form kits on demand:

git clone https://github.com/firecrawl/firecrawl-docker-sandbox.git
sbx run docker/sbx-kit-shell:1.0.0 --kit ./firecrawl-docker-sandbox/firecrawl .

Two things differ from the v2 form, both by design of the v3 spec:

  • The firecrawl credential is required, and v3 enforces that at resolution: with no secret stored, the create is refused instead of creating a sandbox with the credential withheld. Store the key first (section 1).
  • The kit is self-describing inside the sandbox. cat /usr/share/sandbox/kit/firecrawl/kit.yaml prints the published descriptor, and the agent instructions live at /usr/share/sandbox/kit/firecrawl/firecrawl-context.md, indexed from the workload's AGENTS.md.

Everything in section 3 applies unchanged: same install hook, same proxy-managed FIRECRAWL_API_KEY, same egress.

3. Verify the kit

Inside the sandbox session, ! shell escapes prove the mixin is really there. Each check covers an independent layer, from a cheap import up to a full end-to-end scrape.

i. The package is installed, at the pinned version, in the user-site path:

!python3 -c "import firecrawl, importlib.metadata as m; print('firecrawl-py', m.version('firecrawl-py'), '->', firecrawl.__file__)"

Expect firecrawl-py 4.44.0 (the pin from this kit's spec.yaml and firecrawl/firecrawl.yaml) under /home/agent/.local/lib/.../site-packages/, the user-site location that matches the kit installing as user 1000 rather than as root.

ii. The credential is a proxy-managed sentinel, so the real key never enters the sandbox. credentials[].apiKey.proxyManaged makes the engine set FIRECRAWL_API_KEY to the literal proxy-managed in-container (the SDK refuses to send a request without some value), and the proxy swaps in the real key on outbound requests to api.firecrawl.dev:

!env | grep FIRECRAWL_API_KEY

Expect FIRECRAWL_API_KEY=proxy-managed. A real fc-… value here means the credential is coming from somewhere other than this kit's proxy injection. An unset variable means your sbx predates per-credential proxyManaged; update the sbx CLI.

iii. End-to-end proof, scraping a page through the cloud API. This transitively exercises the package, the credential, and egress to api.firecrawl.dev, so run this one if you run only one:

!python3 - <<'PY'
from firecrawl import Firecrawl
fc = Firecrawl()                                   # reads FIRECRAWL_API_KEY
doc = fc.scrape("https://docs.docker.com/ai/sandboxes/", formats=["markdown"])
md = getattr(doc, "markdown", None) or (doc.get("markdown") if isinstance(doc, dict) else "")
print(md[:300])
PY

Expect the first few hundred characters of the page's clean markdown.

Using Firecrawl from the agent

The agentInstructions note tells the agent the capability exists. The calls it points at:

from firecrawl import Firecrawl
fc = Firecrawl()

fc.scrape("https://example.com", formats=["markdown"])   # one page -> clean markdown
fc.search("docker sandboxes mixin kit", limit=5)         # search the web, get page content
fc.crawl("https://docs.example.com", limit=20)           # crawl a site/section
fc.search("flight prices", sources=["alexandria"])        # Alexandria (beta): discover providers/tools
fc.scrape_alexandria({"provider": "...", "capability": "...", "options": {}})  # Alexandria: execute

Alexandria calls need an API key that has been enabled for the beta; on other keys they return an authorization error rather than data.

See the Firecrawl Python SDK docs for the full API (formats, structured JSON extraction with a schema, crawl options).

Two sandbox-imposed limits are worth knowing, because they look like Firecrawl bugs and are not:

  • The kit adds one host; it does not open the web. Under balanced, general websites return a proxy 403 to curl or requests, so a URL found in a scrape result has to be scraped through Firecrawl too. Under deny-all, api.firecrawl.dev is the only web host.
  • Screenshot and other media formats return links on Firecrawl's storage host, which is outside the allow list, so the file cannot be downloaded from inside the sandbox. Widen permissions.network.allow in a fork if you need them.

Troubleshooting

mount policy denied: /Users/<you>: no applicable policies for op(...) at create time: sbx run mounts the current working directory into the sandbox, and mounting your whole home directory is blocked for safety. Run from any directory other than your home directory.

PaymentRequiredError: ... Insufficient credits (HTTP 402) on a scrape is the good failure: the request authenticated (a bad or missing key returns 401, not 402), the account is just out of credits. Top up at https://firecrawl.dev/pricing or lower the request limit.

Centrally governed hosts

A scrape that fails with a proxy-side 403 rather than a Firecrawl error:

WebsiteNotSupportedError: ... Blocked by network policy: domain api.firecrawl.dev:443 —
no matching allow rule — blocked by default deny policy

means the sandbox is under centralized governance and the managed policy is default-deny with no rule for api.firecrawl.dev. A managed policy overrides the kit's own permissions.network.allow, and local sbx policy allow rules are ignored for org-managed domains, so neither this kit nor you can widen egress: only the org can. Confirm with:

sbx policy ls <sandbox-name>        # look for "Managed by <org>" and whether api.firecrawl.dev is allowed

If you see network policy for "api.firecrawl.dev" is managed by your organization; local allow rules are not applied, ask whoever owns the governance profile to add an api.firecrawl.dev allow rule, mirroring the shape of any existing per-service rule. PyPI is usually already allowed for installs, so api.firecrawl.dev is the one host to add. Everything else in the kit (SDK install, proxy-injected credential) works either way; only the outbound scrape is gated. On an ungoverned host with a local balanced or open policy, no extra step is needed.

Developing

Validate the kit the way CI does, with no sbx or Docker daemon required. It runs the authoritative loader from docker/sbx-kits-contrib plus the engine-level and repo consistency checks that loader leaves to the runtime:

cd tools/kitcheck && go run . ../..

kitcheck also checks that firecrawl/firecrawl.yaml declares the same version, SDK pin, hosts, credential and agent instructions as spec.yaml. The v3 descriptor's grammar is validated by the kit frontend while it builds, and the image by Docker's conformance suite; with Docker Desktop and kit-tck:

PUSH=0 ./scripts/push-kit-v3.sh
kit-tck validate --layout /tmp/sbx-kit-firecrawl-layout 1.0.0

If you have the sbx CLI, also run the real thing and a smoke test of each form:

sbx kit validate .
sbx run --kit . shell                                 # v2 mixin on a built-in agent
sbx run docker/sbx-kit-shell:1.0.0 --kit ./firecrawl .       # v3 mixin on Docker's v3 shell workload

See CONTRIBUTING.md for how to bump the SDK pin and cut a release.

About

Docker Sandboxes kit that gives any sandboxed agent live web search, scrape and crawl via Firecrawl

Topics

Resources

Contributing

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages