Gerapy: the container rewrites ownership of whatever you mount into it
Distributed Crawler Management Framework Based on Scrapy, Scrapyd, Django and Vue.js
At a glance
- What is it?
- Gerapy is a crawler management framework that puts a web interface in front of Scrapy and Scrapyd, and its README is written in the voice of someone documenting a tool they ship rather than one they maintain. The docker image is where that shows: its startup script changes the ownership of the mounted workspace before doing anything else, and it creates an administrator account by a command the README never mentions.
- Who is it for?
- Gerapy fits teams that already write Scrapy spiders and want a shared place to build, deploy and watch them, since it replaces a per-developer scrapyd setup with a server and a client index. Three things to check before you run it anywhere shared.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 93 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Six commands before the interface answers anything
The source install is a short sequence, and the order matters because three of the steps depend on where you are.
Install the package, then check that the command exists on your path. Run the initialisation command, which creates a folder named after the tool unless you pass a name of your own. Change into that folder, because the next two steps operate on the workspace rather than on an installation:
gerapy init
cd gerapy
gerapy migrate
gerapy createsuperuser
gerapy runserverThat is the database migration, then an administrator account, then the server. Once it is up you have two addresses: the application on port 8000, and a separate admin path under the same port for the management backend.
There is one more form of the last command, for when you want the server reachable from outside the machine, and it is the same command with a host and port. The file describes that as running with a public host, which is to say binding to all interfaces, and it offers no guidance on what should be in front of it.
Documentation is offered in two places: a hosted documentation site and a separate documentation repository. For a project of this age, two documentation locations is one more than a reader needs.
The project configurator is called unstable, and hand-added projects sit outside it
The feature paragraph is where the tool admits its own weak point, and it does so without dressing it up.
The file says you can create a configurable project and then configure and generate Scrapy code automatically, and then immediately says that this module is unstable and is being refined.
The second route into a project is described next, and it carries its own limitation. You can drag a Scrapy project into a projects folder; refresh the browser and it appears on the project index page, but it comes to what the file calls un-configurable. You can still edit it through the web page.
Read carefully, that is a meaningful boundary: the visual generator only applies to projects the tool created itself. A spider you already wrote is visible and editable in the browser but is not driven by the configurator, so the feature you are evaluating does not apply to the code you already have.
The task list at the bottom of the file agrees with that reading and disagrees with itself at the same time. Four items are marked done, including a visual configuration of the spider with a previewing website. The remaining open item is the visual configuration of Scrapy itself, which is arguably the same feature named more broadly.
The deployment path is simpler: build the project, add a client on the client index page, then deploy by clicking a button, and watch the resulting job on the monitor page.
The docker path publishes a fixed administrator account
The container route is presented as the zero-configuration option, and it is the shortest section in the file.
One command brings up a compose stack that serves the interface on port 8000. The file then names the credentials you can use to log in: a temporary administrator account whose username and password are both the same four-letter word, followed by a request to change the password afterwards for safety.
That is better than publishing nothing, and it is still a fixed administrator credential in a README that is indexed and copied. The request to change it is a sentence in the documentation, not a mechanism.
There is a second layer to this. The container image does not wait for you to create an administrator. Its startup script includes an admin initialisation command, which is not the command the source instructions tell you to use. So on the container path an account is created for you at boot, without a prompt, and on the source path you are asked for credentials interactively.
That difference is easy to miss and easy to get wrong, because the two documented paths are described as interchangeable. If you switch between them, the account you made interactively does not follow you, and the one that was made for you is the one with the published password.
The documented mount path contradicts the command above it and the image
The manual docker invocation is given, and then the paragraph after it describes the same invocation with a different path.
The command mounts your workspace at a path inside the container and maps a host port to the container port. The prose immediately below says to mount the Gerapy workspace by the same volume flag with a different target directory, one under the application root rather than the home directory used in the command.
Only one of the two can be right, and the image settles it. The runtime stage defines its home directory with a default value, and that default is the path used in the command, not the one used in the prose. The compose file mounts the same home path too, and gives the volume a name.
So the sentence in the prose is wrong, and a reader who follows it exactly gets an empty directory inside the container. Nothing fails loudly: the server starts, it just cannot see the workspace, which for a crawler manager means an empty project list.
This is the kind of error that survives in a README because it is only visible when someone compares three places at once, and this one is visible in three places: the command, the prose, and the image.
The startup script chowns everything in the mount before it starts the server
The image builds a small shell script at build time, and its contents are the operational contract of the container.
The script is written to run with tracing on, and its first command is not the application. It runs a find over the home directory, selects every entry that is not already owned by the application user, and changes the ownership of those entries to the application user and group. Only after that does it switch to that user and run the application's own initialisation, migration, admin creation and server start.
Read that against the mounting advice in the README and the implication is clear. The documentation tells you to bind a directory on your host into that home directory. The container then rewrites the ownership of everything in it that does not already belong to the application user, which on a Linux host means your files end up owned by a numeric user id that the image chose.
That is a defensible design for a container that has to fix up a volume it did not create. It is also a destructive one for a host directory, and the two things a reader would want to know, that the ownership changes and that the id is configurable through an environment variable, are not in the README section at all.
The rest of the image is conventional and careful: two build stages on a slim Python base, a dedicated init process and a privilege-dropping helper installed as binaries, a cache purge, and cleanup of the package lists.
Four dependencies carry upper bounds, and the development dependency group is empty
The manifest is where the project's age shows, and it is worth reading before you pin anything.
Four of the runtime dependencies are capped rather than simply floored: a scheduler and a Django add-on are both limited to a fixed maximum version, another add-on is capped below its next minor, and the REST framework is held below a particular release. Capping a dependency is the right move when a transitive conflict forces it, and it is also a decision that eventually collides with whatever else in your environment needs a newer release of the same package.
The interpreter range starts at 3.7, the classifiers list 3.7, 3.8 and 3.9, and the container image is built on 3.10. So a 3.10 runtime is what you get in practice while the declared support stops at 3.9.
The dependency list otherwise describes the stack plainly: the crawler framework, its distributed scheduler, a Redis queue, a browser-automation library, a headless rendering client, a database driver pair for two engines, an egg-collection helper, an HTML parsing library, an API client for the scheduler, a websocket library, and a template engine.
And then there is an empty development dependency group, declared with a heading and nothing under it. So the tooling a contributor needs is not in the manifest, which means the install instructions for contributors live somewhere this file does not show.
Three years of commits with no release since July 2023
The version story is the last thing to check, and it is the one that decides whether you should pin.
The three most recent releases are all in the 0.9 series: one from the end of 2021, one from the end of 2022, and one from July 2023. There is no release after that, while the default branch was pushed in July 2026. So the branch has moved for three years without a tagged version.
The manifest still carries the version of the last release, which means the branch's own source tree identifies itself as that older release. Anything that reports its own version will tell you 0.9.13 whether or not it contains three years of commits.
The release list also skips a number between the 2022 and 2023 entries, so the series is not even complete between its own tags.
Put together with the capped dependencies and the interpreter classifiers that stop at 3.9, the picture is a project that is developed but not released, which is a normal state for an internal tool and an awkward one for anything you intend to depend on from a package index.
Editorial conclusion
Gerapy fits teams that already write Scrapy spiders and want a shared place to build, deploy and watch them, since it replaces a per-developer scrapyd setup with a server and a client index. Three things to check before you run it anywhere shared. The container's first action is a recursive ownership change across whatever directory you mount, which on a host directory means your files change owner, so mount a dedicated workspace. The published docker credentials are a fixed administrator pair, and the image creates an administrator account non-interactively at boot, so the first thing to do is change it. And the visual project configurator is described by its own documentation as unstable, with hand-added projects explicitly outside it, so plan on editing spiders as code.
Frequently asked questions
What is Gerapy?
A distributed crawler management framework built on Scrapy, Scrapyd, Scrapyd-Client, Scrapyd-API, Django and Vue.js, giving a web interface for creating projects, deploying them to clients and monitoring the resulting jobs.
How do I start a Gerapy server from source?
Install the package, run the init command to create a workspace, change into that folder, run the migrate command, create a superuser and then run the server. The interface is then on port 8000 with a separate admin path on the same port.
What are the default Gerapy Docker credentials?
The documentation gives a temporary administrator account whose username and password are both admin, and asks you to change it afterwards. The container image also runs an admin initialisation command at startup, which is not the command the source instructions use.
Can I use the Gerapy visual project configurator?
The file describes that module as unstable and under refinement. A Scrapy project you drop into the projects folder appears on the index page but is not configurable through the generator, only editable from the web page.
Which Python versions does Gerapy support?
The manifest declares Python 3.7 and later, with classifiers for 3.7, 3.8 and 3.9, while the container image is built on Python 3.10. The file also says Python 2 support may be added later.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/gerapy-gerapy)