IBM AssetOpsBench: an MCP-based benchmark for industrial maintenance agents
AssetOpsBench - Industry 4.0: A unified benchmark and framework for building, orchestrating, and evaluating domain-specific AI agents for Industry 4.0 asset operations and maintenance, with 460+ scenarios, 5 specialist agents (IoT, FMSR, TSFM, Work Order,...), and multi-agent orchestration blueprints (MetaAgent, AgentHive) over MCP.
At a glance
- What is it?
- IBM's AssetOpsBench ships 141+ scenarios, five domain MCP servers and two orchestration blueprints for asset operations and maintenance agents. It is a research harness, not a production maintenance platform, and the README leaves the runnable scenario path marked as not yet enabled.
- Who is it for?
- Adopt AssetOpsBench if you are building or evaluating LLM agents over industrial maintenance data and want a fixed scenario set plus MCP tool surfaces you do not have to invent. Do not adopt it as a production CMMS, an IoT historian or a Maximo replacement; the repository describes a benchmark and framework, and the README itself marks the single-scenario run command as 'to be enabled'.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The gap AssetOpsBench targets: agents that answer maintenance questions, not chatbots
Industrial maintenance questions are rarely answerable from one source. "Why is Chiller 6 drawing more power this week?" needs the sensor list for that asset, its recent readings, the failure modes recorded for that class of equipment, the open work orders, and possibly a forecast. A general-purpose LLM has none of these, and a hand-rolled retrieval layer over PDFs will not produce a numeric history either.
AssetOpsBench is IBM's attempt to make that class of question testable. The README describes it as a unified framework for developing, orchestrating and evaluating domain-specific AI agents in industrial asset operations and maintenance, with reproducible scenarios, agent tooling and evaluation pipelines over simulated industrial environments. The stated audience is maintenance engineers, reliability specialists, facility planners and Industry 4.0 researchers.
That framing matters when you decide whether to adopt it. This is not a product that monitors your plant. It is a benchmark plus a set of tool servers, and its value is that the tool surfaces and the scenarios are already defined, so you can compare two agent designs on the same tasks instead of on your own improvised demo.
Five MCP servers, one utility server, and what each one exposes
The architecture is a set of Model Context Protocol servers, each owning one slice of the maintenance domain, plus a shared utility server. The README lists them explicitly.
The IoT server covers asset and sensor lookup: sites, asset_ids, asset_detail, assets, find_assets_by_sensors, installed_sensors, measured_sensors, latest_reading, history and sensor_stats. The FMSR server handles failure modes with get_failure_modes, generate_failure_modes and add_failure_modes. The WO server splits into nine read tools (list_workorders, get_workorder, get_workorder_tasks, get_workorder_costs, get_workorder_actuals_vs_planned, get_workorder_kpis, get_schedule_calendar, get_my_assigned_workorders, get_failure_codes) and six write tools (generate_work_order, update_workorder, approve_workorder, assign_technician, close_workorder, cancel_workorder). Vibration adds get_vibration_data, list_vibration_sensors, compute_fft_spectrum, compute_envelope_spectrum, assess_vibration_severity, calculate_bearing_frequencies and diagnose_vibration. The utility server provides json_reader, get_sensor_catalog, get_asset_catalog, get_failure_mode_catalog, current_date_time and current_time_english.
The TSFM server is the largest surface: the README states it currently contains 41 tools split across tasks and evidence (list_tasks, profile_series, characterize_series, data_quality), a model catalog (list_models, search_models, find_models, resolve_model, model_template, register_model, register_finetuned, hf_stats), a feature catalog (list_features, search_features, extract_features, select_features) and a run and evaluation ledger (recipe_template, run_recipe, run_tabular_recipe, run_plan, evaluate, list_runs, list_results).
The design point worth noting is the write surface on the WO server. Six tools can mutate work orders. Any agent you plug in inherits that capability, and the README does not describe a permission or approval boundary around it.
Installing AssetOpsBench and running the plan-execute entry point
The README gives a clone-and-install sequence. It requires Python 3.12 or later according to pyproject.toml, and installs the package in editable mode.
git clone https://github.com/IBM/AssetOpsBench.git
cd AssetOpsBench
pip install -e .The README then shows a single-scenario invocation, but flags it as not yet active: the comment above the command reads "to be enabled". Treat the block below as the intended shape of the interface rather than something the README claims works today.
python -m assetopsbench.run --scenario "List all sensors of Chiller 6 in MAIN site"What actually installs is a set of console scripts declared in pyproject.toml rather than one monolithic runner. The distribution name is assetopsbench-mcp, and the entry points include plan-execute, iot-mcp-server, utilities-mcp-server, fmsr-mcp-server, tsfm-mcp-server, wo-mcp-server, vibration-mcp-server, and several agent CLIs: claude-agent, openai-agent, deep-agent, stirrup-agent, opencode-agent and direct-llm-agent, plus evaluate.
plan-execute --help
iot-mcp-serverThe second command starts the IoT MCP server in the foreground; you should see it block and listen rather than exit. The README points to INSTRUCTIONS.md for full setup, MCP server configuration and the plan-execute runner, so that file is the real starting point. Configuration is loaded from environment variables via python-dotenv, and the repository ships a .env.public file as the template.
The dependency list is heavy and worth reading before you commit: fastmcp and mcp[cli] for the servers, litellm pinned at 1.94.0, openai, claude-agent-sdk, openai-agents, deepagents, langchain-mcp-adapters, stirrup with docker, litellm and mcp extras, couchdb3, sktime, statsmodels, PyWavelets and granite-tsfm. The pyproject.toml comments state that sktime is a hard requirement of run_recipe and the model catalog, and that statsmodels backs ThetaForecaster, AutoETS and ExponentialSmoothing plus eight extractors.
Two orchestration blueprints, and why the choice is not cosmetic
The repository ships two orchestration frameworks, named in the README as MetaAgent and AgentHive, and the at-a-glance table counts them as a separate line item from the domain agents. It also counts five domain agents plus a utility server, and 141+ scenarios.
The distinction that matters for an adopter is where control flow lives. In a plan-and-execute design, a planner produces a step sequence and an executor walks it, calling MCP tools as it goes. The plan-execute console script in pyproject.toml points at agent.cli, which suggests that is the reference implementation of this pattern. A multi-agent design instead lets several agents hand work to each other, which is more flexible when the task decomposition is not known in advance and harder to reason about when a step fails.
The README does not publish a comparison of the two blueprints, so the honest position is that you have to read the code under src/agent to choose. The repository layout shows src/agent, src/evaluation, src/llm, src/observability, src/servers and src/mcphub as the shipped packages, which tells you observability and evaluation are treated as first-class concerns rather than afterthoughts. That is a reasonable signal for a benchmark project, though it says nothing about how mature either orchestrator is.
Where AssetOpsBench is the wrong tool
The clearest limitation is stated by the project itself. The Quick Start command that would run a single scenario carries the comment "to be enabled". A reader who clones the repository expecting a working end-to-end demo from the README alone will not get one; the README defers to INSTRUCTIONS.md for setup and the plan-execute runner.
The second limitation is branch sprawl. The README states that active development is on main, that the codebase used for various publication venues continues on separate branches, and it names two: IndustryAssetEQA for ACL 2026 and main-0.x for prior experimental work. The Colab link in the README points at a notebook on the main-0.x branch, not main. If you follow the README top to bottom you can end up running code from a different branch than the one you cloned, and the README does not tell you which branch is authoritative for any given feature.
The third is scope. Nine asset classes and 141+ scenarios is a benchmark, not a plant. There is no documented path for pointing the MCP servers at a live historian or a production CMMS, and the CouchDB dependency suggests the backing store is a specific database rather than a pluggable interface. If your goal is to put an agent in front of real technicians this week, this repository is not that. It is also the wrong tool if you need a time series model library rather than an agent harness: the TSFM server wraps sktime, statsmodels and granite-tsfm, and if forecasting is the whole job you would use those directly.
How it compares with ITBench and with building your own MCP servers
The related searches around this project include ITBench, which is the obvious comparison point: another IBM-published benchmark for IT operations agents. The difference is domain and tool surface. ITBench targets IT operations, while AssetOpsBench targets physical asset operations and maintenance, and the tool inventory reflects that: sensors, vibration spectra, bearing frequencies, failure modes and work order costs rather than service tickets and deployment pipelines. If your problem is a chiller or a pump, the AssetOpsBench tool set is closer to the data you actually have.
The other alternative is not a product but an approach: writing your own MCP servers over your own CMMS and historian, and skipping the benchmark entirely. That gives you a tool surface shaped to your data and no dependency on sktime, granite-tsfm or a pinned litellm version. What you lose is the scenario set and the evaluation pipeline. Without a fixed task list you cannot tell whether a change to your prompt or your orchestrator made the agent better or just different, and the repository's src/evaluation package exists precisely because that problem is real. A reasonable middle path is to copy the MCP server interfaces you need and keep AssetOpsBench as the scoring harness.
Maintenance status, licensing and the upgrade cost you should budget for
The repository is not archived, and the last push was on 2026-09-10, which is recent. That is the only maintenance signal available here; there are no retrieved releases, so there is no version history to reason about, and the package version in pyproject.toml is still 0.1.0.
The dependency set is the real upgrade cost. litellm is pinned to exactly 1.94.0 rather than a range, which means a security or compatibility fix in litellm requires a manual bump and a retest. The list also includes several fast-moving agent SDKs at low version floors (claude-agent-sdk, openai-agents, deepagents, langchain-mcp-adapters, stirrup), and sktime appears twice with different constraints. A renovate.json file at the repository root suggests automated dependency updates are configured, but automated pull requests still have to be merged and validated against the scenarios.
Licensing is Apache-2.0, which permits commercial use and modification and requires that you preserve the licence and notice files and state significant changes. That is a permissive licence, but the dependencies are a separate question: the project pulls in sktime, statsmodels, PyWavelets, granite-tsfm and several LLM SDKs, each under its own terms, and nothing here tells you how those interact with your distribution model. Check each dependency's licence rather than assuming the Apache-2.0 header covers the whole stack. This is not legal advice.
Editorial conclusion
Adopt AssetOpsBench if you are building or evaluating LLM agents over industrial maintenance data and want a fixed scenario set plus MCP tool surfaces you do not have to invent. Do not adopt it as a production CMMS, an IoT historian or a Maximo replacement; the repository describes a benchmark and framework, and the README itself marks the single-scenario run command as 'to be enabled'. Before committing, read INSTRUCTIONS.md, check which branch holds the code you need (main, main-0.x or IndustryAssetEQA), and confirm that the MCP servers you depend on can reach your own CouchDB and sensor endpoints, since the README does not document a deployment mode for that.
Frequently asked questions
Is IBM a leader in AI?
The repository does not make that claim and offers no evidence for or against it. What it does show is IBM publishing AssetOpsBench, an Apache-2.0 benchmark for Industry 4.0 maintenance agents, with a paper, a Hugging Face dataset and acceptance at several venues including KDD 2026.
What is IBM Maximo asset management?
The repository does not describe Maximo. AssetOpsBench is a separate IBM project: a benchmark and framework for building, orchestrating and evaluating AI agents over simulated industrial environments, with MCP servers for IoT data, failure modes, work orders, vibration and time series models.
What category is IBM in?
The repository does not answer this. It places AssetOpsBench in the Industry 4.0 asset operations and maintenance domain, tagged for condition-based maintenance, predictive maintenance, IoT and LLM agents, and licensed under Apache-2.0.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/ibm-assetopsbench)