How do I install extra Python packages on Scrapy Cloud?

My spider imports a package that is not in the Scrapy Cloud stack, and the job fails with an import error. How do I add it?

Asked and answered by the Zyte team, as a reference for the community.

List your dependencies in a requirements.txt in the project, and tell the deploy about it in scrapinghub.yml:

requirements:
  file: requirements.txt

Then deploy again. The packages are installed into the project’s environment as part of the deploy.

A few related points.

  • If a spider setting you need is not in the settings list in the dashboard, pick the first entry in the drop-down, Custom Name, then type the name and value. The Raw Settings tab lets you edit all settings as plain text.
  • The stack decides which Python version you run. See the stacks article in the Scrapy Cloud help center.
  • If you need more control over the environment, for example to run a browser next to your spider, a paying account can deploy a custom Docker image.

Docs: Deploying code to Scrapy Cloud projects - Zyte documentation and Frequently Asked Questions about Scrapy Cloud - Zyte documentation