# The risk metrics come from deterministic kernels, not from the agent

> A local-first credit risk platform that refuses to let a language model do arithmetic: KS, AUC, PSI and the rest are computed by platform code that the manual workbench and the agent mode share. The interesting engineering is in the handoffs, PMML export, and a Bandit baseline checked into the root.

**eddyzzl/marvis-risk-agent** — MARVIS-Agent: all-purpose credit risk agent for model development, validation, data processing, feature engineering, and strategy workflows.

- Repository: https://github.com/eddyzzl/marvis-risk-agent
- Stars: 476 · Forks: 8
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/eddyzzl-marvis-risk-agent

## KS, AUC and PSI are computed by code the agent does not control

The claim that separates this platform from a chatbot is stated as a bullet under why risk teams use it.

Trust the numbers. KS, AUC, PSI, bad rate, approval rate, profit and impact are described as calculated by deterministic platform code, not guessed by an LLM.

The mechanism behind that is one sentence later. Agent mode and the Manual Workbench share the same validated workflows, tools, schemas and deterministic calculation kernels.

That is the whole architecture in a line. There is no separate agent implementation of the statistics; there is one implementation, and the agent is a caller of it. A risk analyst who distrusts the agent can open the same workflow by hand and get the same number, which is the property that makes a result reviewable.

The rest of that bullet list is built on the same idea. Keep data close: files, task state, evidence and outputs stay in a controlled local workspace by default. Retain human responsibility: high-impact actions pause for confirmation, and key governed results carry lineage and audit evidence. Work in business language: start with the decision rather than with a chain of scripts.

And the README is explicit in its opening that this is not a chatbot wrapped around a collection of scripts, which is the failure mode those four properties exist to avoid.

## The strategy workflow is called seven-step and lists twelve things

The first of the three V2 deliverables is headed a complete seven-step strategy-development workflow, and the sentence under it then runs on.

Current and historical evidence, governed dual-population samples, univariate and model evidence, trees, Cross, scorecards, Voting, Strategy Pools, impact measurement, validation, code delivery, and four-format review reports.

That is twelve comma-separated elements under a heading that says seven steps. The README does not say which seven are the steps and which are the steps' outputs, so a reader planning a rollout has to guess at the shape of the workflow.

The vocabulary in that list is worth unpacking on its own, because it is a specific school of risk tooling. Dual-population samples means separate development and validation populations designed rather than split at random. Cross, Voting and n-of-k are cutpoint-combination methods over scorecards, and the strategy module entry confirms that a 2D Cross Matrix, 2D and 3D cross-threshold search, scorecard cutoffs and Voting combinations are all available. Strategy Pools is the mechanism for holding several candidate rules at once.

Impact measurement, stability checks and validation on independent partitions are the last three, and the module entry adds that adoption of local versions happens through human gates.

## Joins are proposed and confirmed, and deduplication is explicit

The data processing module is the least glamorous entry in the table and the one with the most telling verbs.

It registers CSV and Excel files, infers schemas, profiles data, aligns columns, proposes and confirms joins, diagnoses match rate, fans out and inflates rows, deduplicates explicitly, runs governed transformations, and exports safely.

Three of those describe behaviour rather than capability. Joins are proposed and then confirmed, which means the platform does not silently pick a key. Deduplication is explicit rather than automatic, which means removing a row is an act with a record behind it. And row inflation is something it diagnoses, which is the failure where a one-to-many join quietly multiplies your sample and every rate downstream becomes wrong.

The deliverables column is correspondingly concrete: derived datasets, join evidence, profiling summaries, and CSV or XLSX exports. Join evidence as a first-class output is what makes the confirm step auditable.

Schema inference and column alignment sit on top of that, and the dependency list explains how the alignment is done. `rapidfuzz` is a fuzzy string matching library, which is what you would reach for to match column names across files that spell the same field differently.

## Bad labels are built from DPD plus two windows, and cohort maturity is checked

The labels and samples module is where a credit risk platform either is rigorous or is not, and the README is specific about its inputs.

Bad labels are defined from DPD, which in this field means days past due, plus observation and performance windows. So the label is not a single field in the data; it is constructed from a delinquency threshold and two time horizons, and which combination you pick determines what your model is predicting.

From there the module checks cohort maturity, which is the step that catches a label definition applied to a period too recent to have matured. Then it designs development, validation and out-of-time samples, calculates IV, KS, AUC, PSI, Lift and Coverage, bins numeric and categorical features, analyses correlation and collinearity, and encodes, imputes, caps and derives.

Two phrases in there are worth pausing on. Out-of-time is not out-of-sample, and having it as a named sample type is a statement that the platform expects time to be part of your validation design. And capping is listed as an operation rather than a warning, which means winsorisation is something the workflow performs on the record.

The deliverables are feature evidence, governed sample definitions, selected feature sets and Excel reports. A governed sample definition as an artefact, rather than a line of code that generated a split, is what lets a validation team see what was built.

## Validation scores the submitted PMML, and one entry takes 1 to 10 models

The model validation module is the one with the most operational detail, and it is the clearest statement of what the platform is for.

It scans Notebook, sample, PMML and dictionary materials. It scores submitted PMML and calculates performance, stability, score consistency, binning and stress evidence, while retaining legacy Notebook model-score comparison.

Read that sequence carefully. The primary path is that a model arrives as PMML, and the platform scores that artefact rather than retraining anything. That is a real difference in posture: the validator is checking the thing that will actually run in production, and the legacy Notebook comparison is there for the older workflow.

There is a capacity limit, and it is stated. One entry accepts 1 to 10 models with individual evidence and editable report drafts, both manual and agent-assisted paths remain available, and two or more models also receive a batch summary workbook.

So the artefacts are structured validation evidence, individual Excel and Word reports, and a multi-model summary Excel.

That word individual is doing real work in a validation context. Each model gets its own evidence and its own draft, which means a reviewer can disagree with one model's conclusion without discarding the batch.

## PMML is the handoff format, and the export is qualified

Two independent mentions of PMML in the README, and they are the join between two modules.

The model development module ends its list of deliverables with PMML for supported recipes, alongside experiments, score evidence, model reports, scored data, model cards and handoff packages. The model validation module begins by accepting PMML as an input.

So the format is the contract between building a model and checking it, and it is the same contract the dependency list supports: `sklearn2pmml` converts from scikit-learn pipelines and `pypmml` scores PMML without running the original Python model.

The qualifier matters and should not be skipped. PMML export is available for supported recipes, not for all of them. The model development entry lists binary, regression and multiclass recipes among what it can build, and the library set underneath it is scikit-learn plus xgboost, lightgbm and catboost plus statsmodels. Which of those convert cleanly is not a property the README states.

A validation team that requires PMML has to establish coverage for their own model types before committing to the platform, and that check is not something the documentation answers.

The test configuration is telling about the project's own confidence in this path. The pytest markers include a `pmml_runtime` category for tests that execute the real PMML runtime, alongside slow, e2e and llm markers.

## A Bandit baseline is checked in, which means known findings are accepted

The repository root contains a file named `.bandit-baseline.json`.

Bandit is a Python security scanner, and a baseline file is a list of findings the project has looked at and chosen not to fix. Checking it in is an admission rather than a boast: these are the known issues, and a future scan that produces them is not a regression.

The rest of the root is agent-facing rather than user-facing. `CONTEXT.md` and `DESIGN.md` sit beside the README, which is the convention for giving an assistant the background of a repository before it edits it. There is a bilingual README in English and Chinese, a `marvis/` package, `docs/`, `packaging/`, `scripts/`, `tests/` and a `uv.lock`.

Two small packaging details are worth noting. The package registers two console scripts that point at the same entry point:

```toml
[project.scripts]
marvis = "marvis.__main__:main"
marvis-risk-agent = "marvis.__main__:main"
```

And the interpreter range is capped rather than open-ended:

```toml
requires-python = ">=3.11,<3.14"
```

A hard ceiling below 3.14 in a scientific stack is a deliberate choice, and it means an environment on a newer interpreter needs a resolution decision before it can install.

The test configuration adds a ten-minute per-test timeout enforced on threads, with markers separating real-training tests, service-starting tests, LLM evaluation tests and PMML runtime tests. That timeout is a guard against the hangs that real training runs cause, and its presence says the project has hit them.

## Conclusion

Adopt MARVIS if your team is regulated enough that every number needs provenance, since the design puts calculation in shared deterministic kernels and keeps lineage on each governed result, which is a rarer choice than adding an agent to an existing notebook chain. Do not adopt it as a chat interface over scripts, because that is the thing the project explicitly says it is not, and do not expect PMML export for every recipe since it is qualified as available for supported recipes. Two things to check first: whether your models are among those supported recipes, and whether Python is inside the declared range of 3.11 to below 3.14.

## FAQ

### What does MARVIS-Agent actually do?

It is a local-first, governed credit-risk agent platform covering data processing, labels and samples, model development, model validation, strategy development, and vintage and risk analysis. You describe the decision you need in natural language and it builds a reviewable plan, pauses at responsibility gates, and runs deterministic tools.

### Does MARVIS calculate risk metrics with an LLM?

No. The README states that KS, AUC, PSI, bad rate, approval rate, profit and impact are calculated by deterministic platform code rather than guessed by an LLM, and that Agent mode and the Manual Workbench share the same validated workflows, tools, schemas and calculation kernels.

### How many models can MARRIS validate in one entry?

Between 1 and 10, with individual evidence and editable report drafts for each. Two or more models also receive a batch summary workbook, and both manual and agent-assisted paths remain available.

### What Python versions does MARVIS-Agent support?

Python 3.11 or newer and below 3.14. The package builds with setuptools, is MIT licensed, and exposes its entry point under both the marvis and marvis-risk-agent console scripts.

## Sources

- [eddyzzl/marvis-risk-agent on GitHub](https://github.com/eddyzzl/marvis-risk-agent)
- [Issues](https://github.com/eddyzzl/marvis-risk-agent/issues)
- [License: MIT](https://github.com/eddyzzl/marvis-risk-agent/blob/main/LICENSE)
- [README](https://github.com/eddyzzl/marvis-risk-agent/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/eddyzzl-marvis-risk-agent
