Open-source project
Western-OC2-Lab/Intrusion-Detection-System-Using-Machine-Learning avatar
Western-OC2-Lab/Intrusion-Detection-System-Using-Machine-Learning

This IDS repository is three notebooks, three PDFs, and a claim of generality

Code for IDS-ML: intrusion detection system development using machine learning algorithms (Decision tree, random forest, extra trees, XGBoost, stacking, k-means, Bayesian optimization..)

598 stars163 forksJupyter NotebookMIT

At a glance

What is it?
Code for three intrusion detection systems published between 2019 and 2022, shipped as Jupyter notebooks beside the papers they came from. The README calls the models general-purpose while every dataset named is vehicular, and the last commit is dated 2026-04-01 with no release tags.
Who is it for?
Read it as the reference implementation for three specific papers rather than as a library, since there is no package, no dependency file and no test suite to install or run. The algorithms are ordinary and well documented in the abstracts, so the value is in seeing how a published vehicle-network detector is assembled.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Activity is slowing. The repository last received commits 6 months ago.
What is it written in?
Mainly Jupyter Notebook, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The repository is three notebooks and three PDFs

The whole artifact is visible in the top-level listing. Three Jupyter notebooks, each named for the system and the venue it was published at: a tree-based system for a 2019 conference, a multi-tiered hybrid system for an IoT journal, and a decision-based ensemble for a 2022 conference. Beside each notebook sits a PDF of the corresponding paper. Then a figures directory, a data directory, a licence and the readme.

There is no package manifest, no dependency file, no test directory and no visible continuous integration. The repository's primary language is recorded as a notebook, which is the honest answer for a codebase that is three notebooks.

That shape has consequences for anyone arriving with a software background. There is nothing to install and nothing to run with one command. Reproducing a result means opening a notebook and having the right versions of the gradient boosting libraries, and nothing in the repository pins them.

The PDFs are the authors' own preprints, each linked to an arXiv copy and a DOI, and the readme cites them formally with authors, venue, pages and DOI for all four publications.

The last commit is dated 2026-04-01, which puts the branch six months behind today with no release tags at all, so the state to pin is a commit rather than a version.

The code is called general, and every dataset named is vehicular

The opening sentence makes a claim that the rest of the repository does not test. It says the code and the proposed systems are general models that can be used in any intrusion detection and anomaly detection application.

The three papers are all about the Internet of Vehicles, and the only two datasets named anywhere in the visible text are vehicular: a CAN bus intrusion set representing intra-vehicle network traffic, and a public vehicular network dataset representing external traffic. So the architecture is argued on automotive network data and offered as generally applicable.

That is a defensible research position, since nothing about a gradient boosting classifier or a k-means cluster labeler is specific to a CAN bus. It is also a claim a reader should price in: the reported detection rates are rates on network data with a known structure and a known attack taxonomy, and there is no evidence here about other traffic.

The attack vocabulary in the 2019 abstract gives the shape of that taxonomy: denial of service, spoofing and sniffing. The system is described as applicable both to the CAN bus of a vehicle and to a general Internet of Vehicles setting.

So the honest framing is a set of general components demonstrated on one domain.

MTH-IDS stacks four tiers over two preprocessing stages

The multi-tiered hybrid system is the most fully specified of the three, and its abstract gives the whole structure.

Underneath the tiers sit two traditional stages, data pre-processing and feature engineering. Above them are four tiers of learning models.

The first tier is four tree-based supervised learners, a decision tree, a random forest, extra trees and extreme gradient boosting, used as multi-class classifiers for known attack detection. The second tier is a stacking ensemble together with a Bayesian optimisation method using a tree Parzen estimator, applied to tuning those supervised learners.

The third tier changes kind. It is a cluster-labelling k-means used as an unsupervised learner, and its stated job is zero-day attack detection. The fourth tier is two biased classifiers plus a second Bayesian optimisation method, this one using a Gaussian process, and its job is tuning the unsupervised learner.

The hybrid in the name comes from combining a signature-based detector with an anomaly-based one, so that known and unknown attacks are handled by different machinery rather than by one classifier stretched to cover both.

One detail in the results sentence is worth noting. The abstract says the system detects various types of known attacks on both named datasets. The zero-day tier is described architecturally, and the visible text does not attach a measured zero-day result to it.

LCCDE names the model it trusts and the decision it weighs

The third system expands its own acronym, which is helpful because the name is otherwise opaque: Leader Class and Confidence Decision Ensemble.

It is described as a decision-based ensemble framework, and the sentence constructing it says it is built by determining the best-performing model among three advanced machine learning algorithms. The list of those three is cut off partway through the name of the first one, so the visible text does not say which three.

What the acronym does tell you is the shape of the combination. A leader class implies the ensemble nominates one class or one model as authoritative, and a confidence decision implies the vote is weighted by how confident each contributor is rather than being a plain majority. That is a different combination rule from the stacking used in the second system, where the meta-learner is trained on out-of-fold predictions.

The motivation the abstract gives is the same as the other two papers: growing connectivity in vehicle networks increases the attack surface, and machine learning detectors are the proposed answer.

Across all three papers the recurring claim is a high detection rate at low computational cost, which is the pairing the 2019 abstract makes explicitly when it credits the ensemble and feature-selection approaches with achieving both at once.

The code has a paper of its own, and the lab splits variants across repositories

One of the four citations is not a research result but a description of this codebase. It is titled as an open-source code release for intrusion detection system development using machine learning, published in a software impacts journal with pages and a DOI of its own.

Publishing the artifact separately from the results is unusual practice and it is the citation to use if you want to reference the code rather than a finding. The readme places it under a heading of its own, above the three system papers.

The same readme also points at two sibling repositories from the group rather than folding them in here. One is a CNN and transfer-learning variant of intrusion detection development, and the other is a hyperparameter optimisation tutorial covering the tuning techniques the papers use. So the lab's work is split by technique: this repository for tree, unsupervised and ensemble methods, a separate one for deep learning, and a separate one for tuning.

The algorithm inventory in the summary is correspondingly wider than any single paper: tree-based methods including decision tree, random forest, gradient boosting, LightGBM and CatBoost, unsupervised k-means, ensemble methods including stacking and the proposed decision-based framework, and Bayesian optimisation as the tuning layer.

LightGBM and CatBoost do not appear in the 2019 abstract's list, which names four algorithms. The list grew as the three systems were published.

A maintenance state you can read off two dates

There is one date that matters here and one licence.

The repository publishes no releases. The last push is dated 2026-04-01, roughly six months before today, and the papers it accompanies run from a 2019 conference through a January 2022 journal issue to a 2022 conference, with the code-introduction paper also from 2022. So this is a repository whose visible work concluded some years ago and whose last commit is more recent than the papers but not recent.

That is not a criticism so much as a fact to plan around. For a notebook-based research artifact the practical consequence is modest: you pin a commit, you expect to fix library incompatibilities yourself, and you should not expect a notebook to be updated for a new release of a gradient boosting library.

The licence is MIT, which is the permissive choice and appropriate for code accompanying published research.

The data directory in the tree is worth a look before assuming the datasets are included, since the two named datasets are public and neither is bundled with most repositories of this kind. The readme does not describe what the directory contains.

Everything the project claims about itself is in the four abstracts, which are reproduced in full in the readme, so the claims can be read in the authors' own words rather than paraphrased.

Editorial conclusion

Read it as the reference implementation for three specific papers rather than as a library, since there is no package, no dependency file and no test suite to install or run. The algorithms are ordinary and well documented in the abstracts, so the value is in seeing how a published vehicle-network detector is assembled. Two caveats before you build on it. The stated generality is broader than the evidence: the README says the models apply to any intrusion detection or anomaly detection problem, while the only datasets named are a CAN bus set and a public vehicular network set. And there is no release, so pin the commit.

Frequently asked questions

What does the IDS-ML repository contain?

Three Jupyter notebooks, each paired with the PDF of the paper it came from, plus a figures directory, a data directory and an MIT licence. There is no package manifest, dependency file or test suite, so there is nothing to install.

Which machine learning algorithms does IDS-ML use?

Tree-based methods including decision tree, random forest, extra trees, XGBoost, LightGBM and CatBoost, unsupervised k-means for zero-day detection, ensemble methods including stacking and the proposed LCCDE framework, and Bayesian optimisation for tuning, with a tree Parzen estimator for supervised learners and a Gaussian process for unsupervised ones.

Which datasets do the intrusion detection systems use?

Two are named in the abstracts: a CAN intrusion dataset representing intra-vehicle network traffic, and CICIDS2017 representing external vehicular network traffic. The repository has a data directory, but the readme does not describe its contents.

What is MTH-IDS in this repository?

A multi-tiered hybrid intrusion detection system combining signature-based and anomaly-based detection. It has two stages of data pre-processing and feature engineering, then four tiers of learning models: four tree-based classifiers, a stacking ensemble with Bayesian optimisation, cluster-labelling k-means for zero-day detection, and two biased classifiers with a second Bayesian optimisation method.

What does LCCDE stand for?

Leader Class and Confidence Decision Ensemble, a decision-based ensemble framework for intrusion detection in Internet of Vehicles networks. The abstract says it is constructed by determining the best-performing model among three advanced machine learning algorithms, and the visible sentence cuts off before naming all three.

Is IDS-ML released as a versioned package?

No release tags are published. The repository is MIT licensed and its last commit is dated 2026-04-01, so a consumer should pin that commit and expect to resolve library version drift without upstream help.

Official sources

  1. Issues
  2. License: MIT
  3. README
  4. Western-OC2-Lab/Intrusion-Detection-System-Using-Machine-Learning on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/western-oc2-lab-intrusion-detection-system-using-machine-learning.svg)](https://hysenlabs.com/projects/western-oc2-lab-intrusion-detection-system-using-machine-learning)