ClearML Agent: a queue-driven execution agent for ML workloads
ClearML Agent - MLOps/LLMOps made easy. MLOps/LLMOps scheduler & orchestration solution
At a glance
- What is it?
- ClearML Agent polls ClearML job queues and turns Draft experiments into running processes on bare metal, cloud VMs or Kubernetes. It removes per-machine container plumbing at the cost of depending on a ClearML Server and its scheduler semantics.
- Who is it for?
- Adopt ClearML Agent if your team already runs a ClearML Server and you want queued experiments to execute on machines you add and remove without writing per-host container images. Do not adopt it if you have no ClearML Server, if you need a scheduler that works without a coordinating service, or if your jobs cannot tolerate an environment being rebuilt from the task's recorded requirements.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 14 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 22, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What ClearML Agent solves, and for whom
The README frames the agent as a job scheduler that listens on job queues, pulls jobs, sets the job environments, executes the job and monitors its progress. That single sentence describes the whole product boundary. The agent does not train models and does not track metrics itself; it takes a task definition that already exists in a ClearML Server and turns it into a running process on a machine you control.
The intended user is a team with more GPU machines than DevOps staff. The README is blunt about this: the agent was built so you can set up a dynamic cluster with "epsilon DevOps", and it explicitly says there is no need for yaml, json or template configuration of any kind. Machines can be added and removed from the cluster, and reused without dedicated containers or images. If your current workflow is a wiki page that tells researchers which GPU box is free and an SSH key, this is the gap the project targets.
The second audience is infrastructure teams who already run Kubernetes and do not want researchers touching it. The README states that users do not need direct Kubernetes access, and that the agent works in tandem with other tenants of the same cluster rather than claiming the whole thing. That combination, a UI for scheduling plus restricted cluster access for end users, is the actual selling point for larger organisations.
How the queue, the agent and the task fit together
The mechanism is a pull loop. A task in ClearML has a state. Any Draft experiment can be scheduled for execution by an agent, and the README describes two ways to get a previously run experiment back into Draft: the Reset action, which clears results and artifacts from the previous run, and the Clone action, which creates a new Draft experiment. Once the task is in a queue that an agent is listening on, the agent picks it up.
After pickup, the agent's job is environment construction. The README says execution environments are deployed either with virtualenv or fully docker containerized. So the agent reads what the task needs, builds that environment on the worker, then starts the process and monitors it. This is why the agent can reuse a machine without a dedicated image: the image is effectively derived per job rather than baked per host.
Two optional layers sit on top of the bare-metal case. In Kubernetes glue mode, the clearml-k8s glue pulls jobs from the ClearML job execution queue and prepares a Kubernetes job based on a provided yaml template; inside each pod a clearml-agent installs the experiment environment and monitors the process, visible in the ClearML UI. Slurm integration exists as well and the README points at the documentation rather than describing it inline. The architectural consequence is that the agent is the same component in all three cases, and only the thing that creates the worker changes.
Installing the agent and scheduling a first task
The README gives a five-step path. Step one is a ClearML Server, either self-hosted from the clearml-server repository or the free tier at app.clear.ml. Step two is installing the agent on any GPU machine. Step three is creating a job or adding ClearML to your code. Step four is changing parameters in the UI and scheduling for execution. Step five is left to the reader.
The install is a single package from PyPI:
pip install clearml-agentThe package name is clearml-agent and the distribution name in setup.py is clearml_agent. The runtime dependencies listed in requirements.txt are psutil, urllib3, virtualenv, requests and setuptools, plus pywin32 on Windows. Note the virtualenv pin: virtualenv>=16,<21. If your environment already pins virtualenv outside that range, pip will report a conflict rather than silently working.
Once installed, the agent needs to know which server to talk to and which queue to serve. The README does not spell out the daemon command in the excerpt available here, so check the documentation at clear.ml/docs for the exact flags before running anything in production. What the README does commit to is the workflow: after the agent is running against a queue, you take a Draft task in the UI, place it in that queue, and the agent picks it up, builds the environment and executes it. You should see the task move out of Draft and its console output appear in the ClearML UI.
For Kubernetes, the README points at the helm chart in the clearml-helm-charts repository and at examples/k8s_glue_example.py. The glue example is the file to read if you want to understand the mapping between a ClearML job and a Kubernetes job, because the README describes the behaviour but leaves the template details to the example.
Where the agent is the wrong tool
The dependency on a ClearML Server is the first constraint, and it is not optional. Every step in the README's five-step path assumes a server exists. If you want a scheduler that works from a git repository and a cron expression, the agent adds a service you now have to run, back up and upgrade. The self-hosted server is a separate repository with its own deployment story.
The second constraint is environment reconstruction. Because the agent builds the virtualenv or container per job, a task whose dependencies are not fully captured will fail on a worker that never had them. This is a real failure mode for research code that depends on something installed by hand on the original machine. The README's promise of reusing machines without dedicated images is the same property that makes under-specified tasks non-reproducible.
The third is the Kubernetes story. The README is explicit that Kubernetes is optional and "not a must", which is honest, but it also means the Kubernetes path is a glue layer over an existing cluster rather than a Kubernetes-native operator. Teams whose platform policy requires everything to be a CRD with a reconciliation loop will find that the ClearML job queue remains the source of truth, and the Kubernetes job is derived from it. That is a deliberate design choice, and it is the opposite of what a platform team used to declarative controllers would expect.
Finally, the README's own framing of "zero configuration" and "epsilon DevOps" should be read as marketing shorthand. The agent removes per-host container authoring, but you still operate a server, decide queue topology and manage credentials for whatever the tasks need to reach.
ClearML Agent compared with a general-purpose CI runner
The closest mental model for many engineers is a CI runner such as a self-hosted GitLab Runner or a Jenkins agent. Both poll a central service for work and execute it on a machine you registered. The difference is what the unit of work is and what the runner knows about it.
A CI runner executes a pipeline definition that lives in the repository, and the pipeline is written in terms of build steps. ClearML Agent executes a task that lives in the ClearML Server, and the task carries its own environment specification, parameters and artifacts. That is why the README can claim you change parameters in the UI and schedule for execution: the parameters are part of the task object, not part of a committed file. A CI runner has no equivalent concept, because changing a parameter means changing the repository.
The second difference is the resource model. CI runners are usually registered against a project or a tag. ClearML Agent listens on job queues, and the README describes the scheduler as flexible and controllable with priority support. Queues plus priorities is a scheduling abstraction that CI systems generally do not offer at this granularity, and it is the reason the agent can be described as a cluster manager rather than only a job executor. The trade-off is that the queue lives in the server, so the scheduling policy is not visible in your repository and cannot be reviewed like code.
Maintenance cost, release cadence and licence
The repository is not archived, and the last push was on 2026-09-10, which is recent relative to the release history. The most recent releases listed are v3.0.3 on 2026-06-02, v3.0.2 on 2026-05-25 and v3.0.1 on 2026-05-06. That is a patch-heavy stretch in the 3.0 line, and the spacing suggests the project ships fixes on a short cycle rather than sitting on a long-lived branch.
The practical upgrade cost is dominated by two things. First, the dependency pins in requirements.txt are narrow in places, notably virtualenv>=16,<21 and setuptools<82.0.0 for Python 3.9 and above. A shared environment where another tool needs a newer virtualenv will force you to isolate the agent. Second, because the agent reconstructs environments per job, an agent upgrade can change how existing tasks build even when the task definition is untouched. Pinning the agent version on workers is the obvious mitigation, and the version is read from clearml_agent/version.py at build time, so the installed version is discoverable.
The licence is Apache-2.0, as stated in setup.py and the repository's LICENSE file. That is a permissive licence, and the README also advertises enterprise features including RBAC, vault, multi-tenancy, quota management and fractional GPU support. Those are named as enterprise features, which means the open source package and the commercial offering are not the same surface. Anyone evaluating the project for an organisation should establish which of those capabilities exist in the Apache-2.0 code before assuming they are available. This is a description of what the repository states, not legal advice; get your own counsel for licence questions.
Editorial conclusion
Adopt ClearML Agent if your team already runs a ClearML Server and you want queued experiments to execute on machines you add and remove without writing per-host container images. Do not adopt it if you have no ClearML Server, if you need a scheduler that works without a coordinating service, or if your jobs cannot tolerate an environment being rebuilt from the task's recorded requirements. Before rolling it out, verify that your server version matches the agent's expectations, check what the agent does with an existing virtualenv on a shared machine, and confirm how the queue's priority ordering behaves for the job types you actually submit.
Frequently asked questions
Does ClearML Agent need a ClearML Server to work?
Yes. The README's five-step setup starts with a ClearML Server, either self-hosted from the clearml-server repository or the free tier at app.clear.ml, and the agent's job is to pull tasks from queues that live on that server.
How do I install ClearML Agent on a GPU machine?
The README gives pip install clearml-agent as the installation step on any GPU machine, whether on-premises or in the cloud. The runtime dependencies are listed in requirements.txt and include virtualenv and requests.
Can ClearML Agent run on Kubernetes instead of bare metal?
The README states you can run both bare metal and on top of Kubernetes in any combination. In Kubernetes glue mode, the clearml-k8s glue pulls jobs from the ClearML job execution queue and prepares a Kubernetes job from a yaml template, with a clearml-agent inside each pod installing and monitoring the experiment.
What happens to a previously run experiment when I schedule it again?
The README describes two routes back to Draft. Reset clears results and artifacts from the previous run, while Clone creates a new Draft experiment. Only Draft experiments can be scheduled for execution by an agent.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/clearml-clearml-agent)