Self-hosted service
owl234/ARL-Next avatar
owl234/ARL-Next

ARL-Next splits OSINT off the scan queue so long jobs stop blocking

🚀 自动化资产侦察与漏洞监控平台 (ARL-Next)。重构自经典 ARL,全面升级 Puppeteer + Nuclei 引擎,打通「天眼查/ICP ➔ 资产发现 ➔ 漏洞扫描 ➔ 威胁情报追踪」安全闭环。支持 Docker 极简部署,AI 二开友好。

448 stars71 forksPythonGPL-3.0

At a glance

What is it?
owl234/ARL-Next rebuilds the ARL asset reconnaissance platform as a Python 3.13 and MongoDB 7 stack with dedicated queues, a native MCP server and containerised workers. The engine that measures it is the queue split: OSINT collection is called directly so it cannot stall a scan.
Who is it for?
Fit for a security team running asset discovery at volume against their own estate, since the queue separation, the batched database writes and the worker lifecycle rotation are the three changes that address the failure modes of the original. A poor fit as a hosted multi-tenant service, since the deployment instructions assume a single server you control and the default credentials in the documentation are the reason that assumption matters.
Can I use it commercially?
Yes, with conditions. GPL-3.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
Is it still maintained?
Yes. The repository last received commits 7 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 6, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The bottleneck was one queue for everything

ARL-Next is a rebuild of the classic ARL asset reconnaissance platform, and the reason it exists is stated as three engineering problems in the original: task queue blocking under large-scale scanning, memory growth during long runs, and outdated dependencies.

The comparison table names the original architecture as a single shared Celery queue for all task execution, and says plainly that complex tasks at high concurrency prone to queue blocking and what amounts to a hang. That is the diagnosis, and it is specific enough to design against.

The fix is a service split. OSINT collection and screenshots became independent microservices, which allows heavy and light work to be separated and database writes batched. The stated outcome is that queue blocking and concurrency bottlenecks converge rather than accumulate.

The rest of the baseline moved forward at the same time: Python 3.6 became 3.13, Vue 2 became 3.5, MongoDB 3.x became 7.0, and the Nmap baseline moved from 7.70 to the 7.95 stable release. A rewrite that only fixed the queue would have left the dependency problem in place, so both moved.

OSINT is called directly, not queued

The architecture diagram is where the queue decision becomes concrete, and the single most important arrow is the one that does not go through the broker.

The backend calls the OSINT microservice directly, over a dedicated port, with a coroutine pool. Everything else that is a scan goes to RabbitMQ, which carries separate queues for light work, heavy work and threat intelligence. A production scan task is the thing that goes to the broker.

So the separation is not a naming convention. Company and ICP intelligence collection, which is the slow external-data part, never enters the queue that scan tasks compete in. The stated reason is to avoid long-running tasks blocking the scan queue.

That is the correct diagnosis of the original problem. If a pipeline blocks, the usual cause is a long job sharing a queue with short ones, and no amount of worker tuning fixes it. Moving the long job out is cheaper than making it faster.

The rest of the worker cluster is organised the same way, with the worker subprocess lifecycle rotation described as the mechanism that controls memory growth across long runs.

Database writes are batched and indexes carry uniqueness

The persistence layer change is the third leg, and it is the one that decides whether the throughput gains survive contact with real data.

Writes moved wholesale to the MongoDB bulk write interface. Core collections carry compound unique indexes. And the collection memory pool is set to 1GB as protection, with what the documentation calls an unreasonable hard memory limit removed so high-throughput scanning stays smooth.

Those three changes address three different failure modes. Bulk writes turn thousands of individual round trips into batches. Compound unique indexes make deduplication a database guarantee rather than an application convention, which matters for a pipeline that can retry. And the memory pool is the difference between the database being told how much to use and being forbidden from using what it needs.

The deduplication theme runs through the threat intelligence side too. The CVE and GitHub leak monitoring is described as using an atomic update lock based deduplication and strict timezone normalisation, which are the two things that stop the same leak being reported twice with different timestamps.

A native MCP server over a standard protocol

The original ARL had no standardised external scheduling interface, only a web console operated by hand. ARL-Next adds a built-in native MCP server, and the framing is about automation rather than convenience.

It is written on the Python standard library rather than pulled from a framework, and it exposes retrieval across a thirteen dimension full-panorama asset profile plus statistical dashboard data. The documentation is careful about the boundary: retrieval and search are available, while task dispatch and policy scheduling capability are still being connected.

That boundary matters when you point an agent at it. The diagram shows an AI agent reaching the MCP server over stdio, and the MCP server calling the backend with API token authentication. So the agent reads the estate; it does not yet get to launch scans through this surface.

Everything else in the access path is conventional. The browser talks HTTPS with basic auth to the frontend, the frontend proxies REST to the backend over a separate port, and both reach the backend as the single authenticated entry.

Containerised workers with a watchdog that restarts them

The screenshot and self-healing layer is the operational part of the architecture, and it has two mechanisms.

Puppeteer runs in containerised headless isolation with request-count based rotation, which is what keeps a long-lived browser process from accumulating memory across a scan run. Rotation by request count is a blunt instrument, and blunt is the point: it does not need to predict when the process degrades, it just replaces it on a schedule.

The autoheal daemon periodically probes the scanning workers and restarts abnormal containers automatically. So there are two layers of protection: rotation keeps a worker healthy while it works, and the watchdog replaces one that has already gone bad.

Sensitive information extraction moved to a pure Python engine with streaming extraction and no external binary dependency. The original used an external Go binary, which the comparison calls out as a cross-platform dependency and compile maintenance cost. Replacing a compiled helper with Python removes an entire class of deployment problem for a scanner that has to run on many machines.

The measurement engine is unified on the Nmap 7.95 stable release, and the plugin and fingerprint coverage is described as over 100 dedicated vulnerability proof-of-concept plugins and a web fingerprint library of 8,800 entries.

Two deployment routes, one aimed at network conditions

There are two production install paths, and the choice between them is really about where the server can reach GitHub.

The first route pulls the deployment orchestration from a domestic Chinese container registry mirror and installs from there, which the documentation frames as avoiding cross-border network timeouts and pull failures. The commands install base tooling, start Docker, create a working directory, pull the web image, create a temporary container, and copy the orchestration scripts out of it before starting.

The second route is a plain shallow clone and one script:

bash
git clone --depth 1 https://github.com/owl234/ARL-Next.git && cd ARL-Next
chmod +x start-prod.sh
bash start-prod.sh

The script carries its own environment probes and is described as installing Docker if missing and tuning kernel settings, so a bare machine is a supported starting point rather than a warning.

The estimates given are a couple of minutes on a machine that already has Docker and three to five from a bare machine, which is the cost difference between the two paths stated as a time budget.

Self-signed TLS and default credentials in the table

After deployment the stack is reached over HTTPS on a non-standard port, and the documentation tells you to ignore the browser warning about the certificate on first visit. That is a self-signed certificate generated by the deployment.

The default credential table is the item to read carefully. It gives a default account and default password per verification layer, and the point of publishing defaults is that they are known to anyone who reads the documentation, which is exactly why the deployment notes also describe SSL and a basic auth gateway in front of the core components with all core services on a private network.

So the security posture described is: terminate TLS at the edge, put basic auth in front, and keep the internal services off the public network. What is not described is what happens if you change the bind address, which is the same gap as in most self-hosted stacks.

The operational features around it are worth naming as well: container health probes, readiness polling, automatic restart of abnormal services, and an in-backend upgrade flow that syncs the orchestration scripts and the latest images from the management console.

A UI rewrite, and a repository with no package manifest at the root

Some of the changes are unglamorous and would not survive a feature list, which is why the comparison table lists them.

Tables and action bars now stick to the viewport when scrolling. Pagination state is persisted locally rather than reset. And the dictionary and policy configuration moved from nested submenus into a three column manager with online preview. These are the fixes for an interface where a long list had no fixed header and configuration was scattered.

The repository layout is worth a note because it explains the shape. There is a backend directory, a frontend directory, a separate MCP server directory with its own configuration guide, an updater directory, a production Dockerfile, a production compose file, an images directory, a changelog, and a version file. There is no root package manifest, which is consistent with a Python project where the frontend build is handled inside the container images rather than from a root script.

Releases are frequent and recent: v1.4.0 on 24 September 2026, v1.3.5 on 16 September, and v1.3.4 on 12 September, with the default branch last pushed on 30 September 2026. The project is GPL-3.0 licensed.

Editorial conclusion

Fit for a security team running asset discovery at volume against their own estate, since the queue separation, the batched database writes and the worker lifecycle rotation are the three changes that address the failure modes of the original. A poor fit as a hosted multi-tenant service, since the deployment instructions assume a single server you control and the default credentials in the documentation are the reason that assumption matters. Before deploying, read the credential table and the certificate note, because the stack terminates TLS itself and ships a self-signed certificate, and decide whether the MCP surface is one you want an agent pointed at before you enable it.

Frequently asked questions

What is ARL-Next?

ARL-Next is an asset reconnaissance and vulnerability monitoring platform rebuilt from the classic open source ARL project. It targets large-scale scanning, where the original shared a single Celery queue across all tasks and could block or hang under concurrency, and it splits OSINT and screenshots into independent microservices so long jobs cannot stall scan work.

Which technologies does ARL-Next use?

The stack baseline is Python 3.13 with Vue 3.5, MongoDB 7.0 and Nmap 7.95, moved up from Python 3.6, Vue 2, MongoDB 3.x and Nmap 7.70 in the original. Gunicorn serves the backend, RabbitMQ carries separate light, heavy and threat intelligence queues, Celery workers run the scan cluster, and Puppeteer does headless screenshots in containerised isolation with an autoheal daemon restarting abnormal workers.

Does ARL-Next have an MCP server?

Yes. It ships a native MCP server written on the Python standard library, reachable over stdio from an agent and authenticated to the backend with an API token. It supports retrieval across a thirteen dimension asset profile and statistical dashboard data, while task dispatch and policy scheduling are described as still being connected.

Official sources

  1. Issues
  2. License: GPL-3.0
  3. owl234/ARL-Next on GitHub
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/owl234-arl-next.svg)](https://hysenlabs.com/projects/owl234-arl-next)