MARVIS Agent: A Governed, Local Workbench for Credit-Risk Model Development
MARVIS-Agent: all-purpose credit risk agent for model development, validation, data processing, feature engineering, and strategy workflows.
At a glance
- What is it?
- MARVIS is a local-first credit-risk agent platform that converts natural-language risk goals into governed workflows covering model development, validation, strategy design, and vintage analysis. All computation runs on the analyst's own machine, and every deliverable carries structured audit evidence that points back to deterministic calculations.
- Who is it for?
- Credit-risk teams that need a self-contained, audit-ready platform covering the full model and strategy lifecycle will find MARVIS well-matched. Teams that require multi-user collaboration, cloud deployment, or explicit LLM provider configuration should check whether those capabilities are documented in the repository before committing to it.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 12 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
A Credit-Risk Agent Platform, Not a Script Wrapper
Most credit-risk teams maintain a fragmented set of notebooks, SQL scripts, and manually exported Excel reports. Each handoff between steps breaks audit lineage, and adding a new model type or strategy requires revisiting large parts of the chain.
MARVIS replaces that chain with a single governed agent platform. The README describes it explicitly as local-first and governed, and distinguishes it from a chatbot wrapped around a collection of scripts. The target users are credit-risk practitioners: data scientists building and retraining scorecards, model validators reviewing submitted PMML files, strategy analysts running impact studies, and risk officers who need audit-ready deliverables.
Version 2.5.0 was released on 2026-09-19. The platform supports the complete lifecycle from raw CSV and Excel data through to strategy code in Python, DuckDB SQL, and JSON, plus formatted Excel and Word reports. The agent and the Manual Workbench share the same validated tools and calculation kernels, so analysts can move between the two modes without changing the underlying methodology.
The Governed Workflow Mechanism: Request, Plan, Gate, Execute
MARVIS operates through a structured request-plan-execute loop. The analyst describes a risk goal in natural language. The agent identifies missing inputs and asks for them, including file paths, field definitions, observation windows, and assumption confirmations. It then builds a reviewable workflow plan and pauses at responsibility gates before running any computation.
The platform's design centres on deterministic calculation. KS, AUC, PSI, bad rate, approval rate, profit, and impact figures are computed by the platform's own code, not inferred by an LLM. Task state and data fingerprints carry ownership, provenance, and artifact references forward across steps. When the agent generates a report or produces strategy code, it links the deliverable to the structured evidence that produced it, rather than to a conversational summary.
High-impact actions pause for explicit confirmation before executing. This is the mechanism behind the platform's audit promise: an analyst can trace every number in a deliverable back to the governed step and the input data that produced it.
Installing MARVIS and Running a First Session
The package requires Python 3.11, 3.12, or 3.13. The pyproject.toml in the repository registers two shell entry points:
[project.scripts]
marvis = "marvis.__main__:main"
marvis-risk-agent = "marvis.__main__:main"After installation, either command starts the platform. The dependency list includes pandas, scikit-learn, XGBoost, LightGBM, CatBoost, DuckDB, FastAPI, sklearn2pmml, and pypmml, so the initial download is large. The pyproject.toml uses the `uv` lock format (`uv.lock` is present in the repository), so developers using uv can synchronise the environment from the lock file.
MARVIS runs a local FastAPI service that the agent communicates with. A first session starts by describing the business goal in natural language. The agent asks which CSV or Excel files to register, confirms column mappings and label definitions, then builds a plan the analyst reviews before any computation begins. High-impact actions pause for explicit confirmation at each responsibility gate.
Six Workflow Modules and Their Audit Deliverables
MARVIS organises its governed workflows into six areas.
Data processing handles CSV and Excel file registration, schema inference, data profiling, column alignment, join design with match-rate and fan-out diagnosis, deduplication, and governed transformations. Outputs include derived datasets, join evidence, and profiling summaries.
Labels, samples, and features covers DPD-based bad-label definition, observation and performance window design, cohort maturity checks, and development, validation, and OOT sample construction. Feature analysis includes IV, KS, AUC, PSI, Lift, and Coverage calculations, plus binning, correlation, collinearity, encoding, imputation, capping, and derivation.
Model development supports binary classification, regression, and multiclass recipes with governed reject inference, leakage-aware splits, multi-recipe training and comparison, model calibration, PMML export for supported recipes, dataset scoring, and monitoring handoff packages.
Model validation accepts Notebook, sample, PMML, and data dictionary materials for 1 to 10 models at once. It produces performance, stability, score consistency, binning, and stress evidence, plus individual Excel and Word reports and a batch summary workbook for multi-model submissions.
Strategy development builds approval, reject, limit, pricing, and segmentation rules using interactive trees, 2D Cross Matrix, 2D/3D cross-threshold search, scorecard cutoffs, and Voting/n-of-k combinations. Outputs include canonical strategies, backtests, ImpactCube evidence, and code in Python, DuckDB SQL, and JSON.
Vintage and risk analysis runs VTG-terminal and annualized bad-rate calculations plus Standard Vintage and roll-rate workflows, delivering audited Excel reports with structured evidence, charts, assumptions, conclusions, and red flags.
Where MARVIS Falls Short
MARVIS is built for single-workstation, local operation. The README does not document a multi-user collaboration model, a centralized model registry integration, or a cloud deployment path. Teams that need to share governed workflows across geographies must build that connectivity themselves.
PMML export is described as available for supported recipes. The README does not enumerate which scikit-learn estimators produce valid PMML via sklearn2pmml, so teams should test their specific pipeline configuration before committing to PMML as a production export path.
The pyproject.toml specifies `requires-python = ">=3.11,<3.14"`, which excludes Python 3.14 and later. Teams already on 3.14+ cannot install the current release without a compatible Python environment.
The agent depends on an external LLM for natural-language goal parsing. The README is silent on which LLM providers are supported and how to configure them. Teams operating in air-gapped environments or with strict LLM vendor policies should verify this before investing in the platform.
Comparison with Manual Pipeline Assembly and Commercial Tools
The most common alternative is a custom pipeline built from scikit-learn, XGBoost, and pandas, with reporting handled by a separate library such as evidently or great_expectations. That approach provides full control but produces no built-in governance layer, no audit lineage, and no strategy-design tooling. Analysts bear the full cost of maintaining provenance and reproducibility themselves.
Commercial risk platforms from vendors such as SAS or FICO cover credit-risk scoring conventions in depth and include model validation tooling, but they typically run in hosted or on-premises server environments with license-based pricing rather than a local Python install. MARVIS is MIT-licensed, which removes license-cost barriers for teams running it on internal workstations.
The key difference from both alternatives is that MARVIS encodes credit-risk domain conventions, including scorecard workflows, strategy pool design, and VTG analysis, directly into its governed workflows. A generic ML pipeline does not provide those conventions, and a commercial platform requires a vendor relationship.
Maintenance, Licensing, and Upgrade Cost
MARVIS is released under the MIT license, which permits commercial use, modification, and redistribution without royalties. The repository is not archived and the last push was on 2026-09-19. Version 2.5.0 introduced the complete seven-step strategy-development workflow, indicating ongoing feature investment.
The large dependency set, which includes major ML frameworks alongside DuckDB, FastAPI, and PMML tooling, means upstream package releases can require environment retesting. The version bounds in pyproject.toml limit this to known-good ranges, but teams should plan for periodic dependency review, particularly for sklearn2pmml and pypmml, which handle PMML export and whose API surfaces can change across minor versions.
Editorial conclusion
Credit-risk teams that need a self-contained, audit-ready platform covering the full model and strategy lifecycle will find MARVIS well-matched. Teams that require multi-user collaboration, cloud deployment, or explicit LLM provider configuration should check whether those capabilities are documented in the repository before committing to it. The first item to verify is PMML compatibility with the specific scikit-learn estimator the team intends to export.
Frequently asked questions
What is Marvis AI?
MARVIS-Agent is a local-first credit-risk agent platform for model development, validation, data processing, feature engineering, and strategy workflows. It accepts natural-language risk goals and executes governed, deterministic workflows that produce audit-ready deliverables without sending data to an external service.
What Python version does MARVIS require?
MARVIS requires Python 3.11, 3.12, or 3.13. The pyproject.toml specifies requires-python = ">=3.11,<3.14", so Python 3.14 and later are not yet supported.
Which machine learning frameworks does MARVIS install?
MARVIS installs scikit-learn, XGBoost, LightGBM, and CatBoost as direct dependencies. It also includes sklearn2pmml and pypmml for PMML export of supported scikit-learn-based recipes.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/eddyzzl-marvis-risk-agent)