Melusine: email threads cut into messages before anything is classified
đź“§ Melusine: Use python to automatize your email processing workflow
At a glance
- What is it?
- A Python library that starts by splitting an email conversation into individual messages and tagging signatures and footers, then qualifies each one with rules or deep learning models. The pipeline comes from a named config, the output is a dataframe, and the classifier still calls itself beta.
- Who is it for?
- Melusine fits a team already working in pandas that needs conversation segmentation and line tagging more than it needs a clever classifier, since the segmentation and tagging are the parts it actually ships. It does not fit anyone who needs a published accuracy figure, because none appears in the visible documentation.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 31 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.
Editorial analysis
A named config decides what the email becomes
The whole library in five lines:
from melusine.data import load_email_data
from melusine.pipeline import MelusinePipeline
# Load an email dataset
df = load_email_data()
# Load a pipeline
pipeline = MelusinePipeline.from_config("demo_pipeline")
# Run the pipeline
df = pipeline.transform(df)The design decision is in from_config. The qualification logic is not code you write in the transform call, it is a named configuration that the pipeline loads, which is why omegaconf is a runtime dependency and why the tutorial material points at a demo_pipeline rather than at a function signature.
Around that sit the parts described as boilerplate: debug mode, pipeline execution, and code parallelization. The claim is that you write the email qualification logic and the library handles running it.
The result is still a dataframe, so this composes with pandas rather than replacing it.
Segmentation comes before every model
The pre-packaged features are the unglamorous ones, and they are the reason the library exists. An email conversation is segmented into individual messages, message parts are tagged, and transferred emails are handled.
The README's segmentation example splits one email into two distinct messages at a transition pattern, then tags each message line by line. The stated payoff is that this segmentation can later improve the performance of machine learning models, which is the honest framing: the cut is the product, the classifier is downstream of it.
Tagging matters more than it sounds. If the body, the signature and the footer are all just text, a model learns to recognise boilerplate as content. Marking those parts means the qualification step sees the message body and can use the signature block as evidence about the sender instead. That is a cheap win available before any model is trained.
Rules and neural networks share one workflow
The claim is that Melusine integrates deep learning frameworks, naming HuggingFace, PyTorch and TensorFlow, alongside deterministic rules built on regex, keywords and heuristics, into a single email qualification workflow. The modular architecture is meant to let you keep the tools you already have.
Release 3.3 changed the shape of that arrangement. Its own summary line says it dropped sklearn inheritance, updated how debug mode is activated, and automated backend selection. For most users the first item is a non-event. For anyone who subclassed the scikit-learn estimators to add a custom feature, it is a breaking change in a minor release, and it is worth finding out before you upgrade rather than after.
Note that scikit-learn is still a hard runtime dependency at 1.0.0 or newer, so dropping the inheritance did not drop the library.
The newest tag is 3.4.0 from 2026-09-01, while the README header still reads Release 3.3, so the front page is one release behind the code.
The output is columns, and the demo shows two
The pipeline returns a qualified dataset, and the two columns named are the clearest statement of intent. messages holds the list of individual messages found in each email. emergency_result is a flag identifying urgent mail.
Those two map onto the four capabilities the overview advertises. Routing is the question of whether a message reaches its intended destination, prioritization is the urgent flag, snippet summaries extract the relevant part of a long message, and filtering is the removal of what you do not want. Each is a stage in a pipeline rather than a separate API call.
The shape worth noticing is that Melusine adds columns and rows of metadata to what you already had. An existing pandas workflow keeps working, and the decision about what to do with an urgent message stays in your code rather than inside the library.
What the visible documentation does not contain is any accuracy, precision or recall figure for the qualification steps, and no comparison against another email classifier.
Python 3.10 to 3.13 on five runtime dependencies
The packaging is conventional. The build backend is setuptools with setuptools 61 or newer plus wheel, requires-python is 3.10 or above, and the classifiers run from Python 3.10 through 3.13 with Typing :: Typed declared. The version is dynamic rather than written into the manifest.
Five packages are required at runtime: arrow for timestamps, pandas for the frames, scikit-learn at 1.0.0 or newer, tqdm for progress, and omegaconf for the pipeline configuration. That is a short list for a library that also talks to model frameworks, because the deep learning dependencies are deliberately not in it.
Installation is the single command pip install melusine. Development extras are tox, pre-commit, black, flake8, isort and mypy, so formatting, linting and type checking are all part of the intended workflow rather than optional polish.
The Makefile says Sphinx, the tree ships MkDocs
The development surface is well covered. The Makefile has help, which parses its own double-hash comments to print the target list, then lint running flake8 over melusine and tests, test running py.test, test-all running tox across every supported interpreter, pytest-coverage with term-missing, and a coverage target that runs coverage against the melusine package, writes an HTML report and opens it in a browser. There are clean targets for build, pyc and test artefacts.
The documentation target is described in the Makefile as generating Sphinx HTML including API docs. The repository, though, contains mkdocs.yml and a docs/ directory, and the published site is maif.github.io/melusine with a tutorials section starting at 00_GettingStarted. So there are two documentation toolchains visible in the tree and only one of them is named in the build.
Also at the root: tox.ini, a pre-commit configuration, AUTHORS.rst, CONTRIBUTING.md, MANIFEST.in and the test directory.
Called production ready, classified as beta
The README says Melusine is production ready and proven in the MAIF production environment, offering stability for real workloads. The project classifier in the manifest says Development Status :: 4 - Beta. Both are accurate about different things, and if you are making a procurement decision the classifier is the more conservative of the two.
Licensing is unambiguous: Apache Software License 2.0 in the README, a LICENSE file at the root and the same text declared in the manifest, with the OSI classifier to match. That is more consistently declared than most projects of this size manage.
History is steady rather than fast. 3.3.3 landed 2026-02-26, 3.3.4 on 2026-06-24 and 3.4.0 on 2026-09-01, and the last push to master carries that same date with the repository not archived. Three releases in six months on a library with no published benchmark is a project you should read the source of before trusting it with a filter.
Editorial conclusion
Melusine fits a team already working in pandas that needs conversation segmentation and line tagging more than it needs a clever classifier, since the segmentation and tagging are the parts it actually ships. It does not fit anyone who needs a published accuracy figure, because none appears in the visible documentation. Before building on 3.4, read what release 3.3 removed, namely scikit-learn inheritance, since subclassing code is what breaks.
Frequently asked questions
How do I install and run melusine on an email dataset?
Install it with pip install melusine, then load a dataset with load_email_data() and run MelusinePipeline.from_config("demo_pipeline").transform(df). The result is a dataframe with columns such as messages, the individual messages found in each email, and emergency_result, a flag for urgent mail.
What does melusine do before any model runs?
It segments an email conversation into individual messages, tags message parts such as the body, signatures and footers, and handles transferred emails. The example in the documentation splits one email into two distinct messages at a transition pattern and tags each message line by line.
Can melusine use rules instead of a neural network?
Yes. The library combines deterministic rules, meaning regex, keywords and heuristics, with deep learning frameworks such as HuggingFace, PyTorch and TensorFlow in one email qualification workflow, and its modular architecture is meant to let you keep the tools you already use.
Which Python versions and dependencies does melusine support?
Python 3.10 and newer, with classifiers through 3.13, and five runtime dependencies: arrow, pandas, scikit-learn, tqdm and omegaconf. The build uses setuptools with setuptools 61 or newer, and the development extras add tox, pre-commit, black, flake8, isort and mypy.
What changed in melusine release 3.3 that could break existing code?
The 3.3 line dropped scikit-learn inheritance, updated how debug mode is activated, and automated backend selection. Code that subclassed the scikit-learn classes is the part that breaks. The newest tag is 3.4.0 from 2026-09-01, while the README header still refers to release 3.3.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/maif-melusine)