auto-deep-researcher-24x7 submits training with sbatch --parsable and never leaves a process on the login node
🔥 An autonomous AI agent that runs your deep learning experiments 24/7 while you sleep. Zero-cost monitoring, Leader-Worker architecture, constant-size memory.
At a glance
- What is it?
- An unattended experiment loop for deep learning work, where the engineering detail worth reading is how it decides a job is finished: the controller stays local, Slurm accounting is the only authority on liveness, and a failed run is no longer filed as a completed one.
- Who is it for?
- This project fits someone with a GPU and a real experiment to iterate on who is willing to hand the loop to an agent overnight, especially on a cluster where process hygiene on the login node matters. It does not fit a machine with no GPU, since training is a hard requirement rather than an optional feature.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 122 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 2, 2026, and from our analysis. They are not legal advice.
Editorial analysis
One SSH call submits the job and exits, so the login node stays clean
The Slurm backend is designed around a single constraint: a controller on a cluster login node that stays resident is a nuisance to whoever administers it. The approach is to submit and leave. Training is submitted with sbatch in its parsable mode over one transient SSH call that exits immediately, and the controller itself never moves onto the cluster. File and repository reads reuse the same SSH path, which works because the login node shares the network filesystem with the compute nodes, so there is no need for a second transfer mechanism. An SSH execution backend was added earlier for the simpler case of a single remote host, where code edits, training, logs, process checks, and GPU queries all run remotely while the controller stays at home. The Slurm mode is opt-in and additive: the local and SSH backends behave exactly as before, and the release notes record twenty-one added unit tests that need no cluster to run.
sacct is the only liveness authority, and two bounds stop the monitor hanging
On Slurm, liveness is not decided by the controller at all. Job accounting is the sole authority, and the reason given is that Slurm itself enforces the time limit, so a job the scheduler has stopped is stopped regardless of what the login node can still see. GPU occupancy is read from the scheduler queue for the partition rather than from anything local. The harder case is the cluster becoming unreachable, and that is what the two bounds inside the liveness check exist for: a grace period for consecutive unknown results, and a wall-clock backstop derived from the submitted time limit. Together they guarantee the monitor loop terminates even when the cluster cannot be reached. The important detail is what those bounds refuse to do: they will not reap a job that accounting still reports as queued or running, so the guard against a hung loop cannot turn into a guard against killing live work.
A FAILED or TIMEOUT run is no longer written down as completed
This is the fix that changes what the agent reasons over. The monitor now asks the backend for the real terminal state of a finished job through a status call, so a run that ended in failure, a timeout, or a cancellation is no longer silently recorded as completed. That outcome flows into three places: the state file, the experiment ledger, and the context handed to the reflection phase, so the planning loop reasons about what actually happened rather than about an optimistic assumption. The backends differ in what they can report, and the difference is documented rather than papered over. On Slurm the state comes from accounting. The local and SSH backends identify jobs by process id only, so they report the terminal state as indeterminate and keep their previous behaviour. An unattended loop that treats a crashed run as a success will optimise towards a metric that was never achieved, which is the failure mode this addresses.
The experiment ledger is a JSONL file, so remembering costs no tokens
Memory here means a file rather than a context window. Every cycle's hypothesis, metrics, and outcome are appended to a ledger file in the workspace, and the file is described as crash-safe and as costing nothing in tokens, which is the property that makes it usable at scale. The ledger is fed back into planning, so the agent knows what it has already tried instead of repeating it. The stagnation signal is derived from that file: the planner is given the metric trajectory, and whether results are still improving or have stalled comes from the ledger rather than from a binary counter of repeated attempts. The metric to follow is named in configuration. A related pair of append-only journals records failed approaches that should not be retried and durable observations, and they are explicitly never compacted, instead being rotated to dated backups when they grow large, so history is never silently discarded.
The budget guard counts cycles per hour, not tokens or dollars
There is an optional cap on how many cycles the agent may start in an hour, and its stated purpose is protection rather than throughput. The description calls it proactive, and the reasoning is that an agent that is stuck in a loop will spend budget without producing new information. Counting cycles rather than tokens means the guard does not need to estimate what a cycle will cost, which is the harder and less reliable number to predict when the agent chooses its own prompts and its own experiments. Alongside it sits a zero-cost violation scanner and an advisory phase gate, both implemented as pure functions over the state and the ledger, so they can report whether the agent is stuck, whether a state is stale, and whether a baseline metric bar has been met without spending anything to find out. Every new configuration section in this release defaults to the previous behaviour, which is what makes the whole layer opt-in.
A Chinese vendor preset is four lines of alias over the OpenAI-compatible path
Running the agent against a domestic model endpoint is presented as a convenience rather than a new integration. You set the provider field to a single word: deepseek, qwen, which maps to the DashScope endpoint, kimi for Moonshot, or glm for Zhipu. The preset fills in the OpenAI-compatible base address and the default key environment variable, so the only other thing to set is the model, using that vendor's own model identifier. Both the base address and the key variable name stay overridable, which is the escape hatch for a self-hosted or proxied endpoint. The implementation is described as a thin alias over the path that already existed, with no new dependency, and the code change is a single module. The whole feature is two lines of configuration:
agent:
provider: "deepseek" # or qwen / kimi / glm
model: "deepseek-chat" # vendor's model idThe whole test suite runs with no GPU and no network
The test strategy is stated as a constraint rather than a boast: the suite is unit-tested without a GPU or a network connection, and the count went from sixty to ninety-nine tests with the autonomy layer added, plus twenty-one more for the Slurm work that need no cluster. That is what makes the safety scanners and the outcome-reporting logic testable at all, since both are pure functions over state. The runtime dependencies are correspondingly small: an Anthropic SDK, an OpenAI SDK, of which you install at least one, and a YAML parser. The tools added in the same release are equally plain: a regular-expression search across the workspace, a recursive and depth-limited directory listing, and line-range file reads so a large file is not truncated blindly. All three are symlink-safe in the sense that they never resolve outside the workspace, and they behave identically whether execution is local or over SSH.
One brief file is the control surface, and the skills are the interface
The shortest path to a running loop is three steps: create a project folder containing a single brief file, run the auto-experiment command with a project path and a GPU index, and then check progress with a status command or the optional note export. The brief is described as the main control file, and a project configuration file is optional, which is a deliberate inversion of the usual arrangement where configuration is mandatory. Progress reporting was added with optional sync to a note vault and a plain text fallback when no vault is configured, so the reporting path does not become a hard dependency on a second application. Skills are installed for both major coding agents, with an earlier release adding dual installation and ownership checks on the installer, and a separate guide exists so an assistant can walk a newcomer through the setup. The required list is short and worth reading in full: a recent Python, one NVIDIA GPU for training, an API key for either a compatible protocol, and the brief.
Editorial conclusion
This project fits someone with a GPU and a real experiment to iterate on who is willing to hand the loop to an agent overnight, especially on a cluster where process hygiene on the login node matters. It does not fit a machine with no GPU, since training is a hard requirement rather than an optional feature. Before starting, read the requirement list rather than the headline, because a GPU, an endpoint key, and a single brief file are all mandatory, and set a cycle cap before the first overnight run so a stuck loop cannot burn a budget.
Frequently asked questions
What does auto-deep-researcher-24x7 need to run?
Python 3.10 or newer, at least one NVIDIA GPU for training, an API key for an Anthropic-compatible or OpenAI-compatible endpoint, and a PROJECT_BRIEF.md in the project folder. A project config.yaml is optional, and at least one provider SDK must be installed.
How does the Slurm backend avoid leaving processes on the login node?
Training is submitted with sbatch in parsable mode over a single transient SSH call that exits immediately, while the controller stays local. File and repository reads reuse the SSH path because the login node shares the network filesystem with the compute nodes.
What happens when a training job fails in auto-deep-researcher-24x7?
The monitor asks the backend for the job's real terminal state, so a run that failed, timed out, or was cancelled is no longer recorded as completed. The outcome reaches the state file, the experiment ledger, and the reflection context. On Slurm it comes from job accounting; process-id-only backends report it as indeterminate.
How does auto-deep-researcher-24x7 remember what it already tried?
Every cycle's hypothesis, metrics, and outcome are appended to a JSONL ledger inside the workspace, which is crash-safe and costs no tokens. The planner reads the metric trajectory from that ledger to decide whether results are still improving or have stalled, instead of relying on a repeat counter.
Can auto-deep-researcher-24x7 use a Chinese LLM API?
Yes. Set the agent provider to a one-word preset such as deepseek, qwen, kimi, or glm, and set the model to that vendor's model identifier. The preset fills in the OpenAI-compatible base address and the default key environment variable, and both remain overridable for self-hosted or proxied endpoints.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/xiangyue-zhang-auto-deep-researcher-24x7)