3W: a commit-counted release tag, a badge pointed at a fork, and two licenses
Timely detections for more proactive and effective actions in offshore oil wells!
At a glance
- What is it?
- Petrobras 3W ships an offshore well fault dataset and a time-series toolkit from one Apache-licensed repository. Its configuration and version policy carry several quirks that affect how you should pin, cite, and reuse it.
- Who is it for?
- Treat 3W as a research resource rather than a pinned dependency: cite the dataset and toolkit versions separately from the repository tag, read LICENSE coverage for the file you actually want, and expect a heavy install that pulls torch and a notebook server whether or not you need them.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 4 days ago.
- What is it written in?
- Mainly Jupyter Notebook, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 9, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The coverage badge measures a fork on a branch that is never tagged
Both the coverage link and the coverage image at the top of the README point at a different owner and a different branch than the repository they sit in. The link target is `https://coveralls.io/github/rafaelpadilla/3W?branch=dev` and the image target is `https://coveralls.io/repos/github/rafaelpadilla/3W/badge.svg?branch=dev`. The project is `petrobras/3W`, and the versioning rules state that the project version is updated when there is a new commit in `main`.
So the coverage percentage on screen was produced from a branch named `dev` under an account that is not this organisation. It is the kind of badge that survives a fork or a rename, and it will keep rendering long after the numbers stop describing the code you cloned.
The neighbouring badges have the same shape of problem in milder form. The code style badge resolves its image from `raw.githubusercontent.com/astral-sh/ruff/main/assets/badge/v2.json`, which follows the ruff main branch rather than the version of ruff your environment installs. The project pins its own ruff to a range, `ruff>=0.14.13,<0.16`, but the badge cannot reflect a range. Read the badge as a pointer, not as a check.
Three version numbers move on three schedules, two of them five hours apart
This project tracks three separate versions by design, and the text is explicit that the toolkit and dataset versions are completely independent of each other.
name = "ThreeWToolkit"
version = "3.2.1"
description = "A modular and open-source AI toolkit for time-series processing, aimed at fault detection and classification in oil well operation"The toolkit version is 3.2.1 in `pyproject.toml`. The dataset version lives in `dataset/dataset.ini`. The project version is a git tag. Those three counters start from different numbers and increment on different events, so 3.2.1 tells you nothing about how far the dataset has moved.
The recent tags show what that produces. `v.1.90.0` was published at 18:32:45 on 2026-09-28 and `v.1.89.0` at 13:35:19 the same day, a gap of about five hours, with `v.1.88.0` four days earlier. Note the tag spelling as well, `v.1.90.0` with a dot after the letter, which is not the conventional SemVer prefix form. If you pin by tag, copy the tag exactly; if you parse tags, expect the extra character.
Every commit to main produces a release, so the tag count measures commits
The versioning section promises exclusive use of semantic versioning, and then defines the project version as updated whenever, and only when, there is a new commit in `main`, regardless of which resource changed. The parenthetical list is explicit: toolkit, dataset, project documentation, or an example of use.
Read together, those two statements pull in opposite directions. SemVer encodes meaning in the version number, and a patch bump is supposed to signal a backward-compatible fix. Here the trigger is a commit, so a commit that only fixes a typo in the documentation still advances the project version, and the project version advances toward a number like 1.90.0 while the toolkit sits at 3.2.1 and the dataset sits at whatever `dataset.ini` says.
The release history supports the reading. Three tags in five days, two of them on one day, on a repository whose primary language is Jupyter Notebook. Nothing in that sequence tells you whether the data changed, the code changed, or a sentence did. When you need to know what moved, read the tag notes rather than the version number, and pin the toolkit and dataset versions separately.
Two license badges sit above a repository that splits code from data
The badge block at the top of the README shows an Apache 2.0 shield and a Creative Commons Attribution 4.0 shield side by side. The split itself only appears further down, where code is stated to be Apache 2.0 and the dataset's Parquet files to be CC BY 4.0.
The grant for the data is worded narrowly: it covers Parquet files saved in subdirectories of the dataset directory. The repository root does contain a LICENSE.md, and the dataset directory is also where the dataset version file lives. So a reader who downloads a CSV, a metadata sidecar, or the ini file from that tree has to work out which grant applies, and the visible wording names only the Parquet files.
Neither shield tells you which is which, and neither names the directory. Two shields at the top of a repository that contains code, notebooks, papers, images and data read as two coequal licenses until you reach the Licenses section and find the split.
There is a second licence document in the tree as well, CONTRIBUTOR_LICENSE_AGREEMENT.md, which governs what happens to contributions rather than what you may do with what you received. Both documents matter for the same file, and only one of them is a licence you can act on as a user.
A contributor agreement and two contributing guides for one project
Before contributing, the text asks you to read and agree to three documents: CODE_OF_CONDUCT.md, CONTRIBUTOR_LICENSE_AGREEMENT.md, and CONTRIBUTING.md. The repository root holds a fourth, 3W_TOOLKIT_CONTRIBUTING.md, which the contributions section does not mention.
Two contributing guides in one root is a split worth resolving before you send a patch, because a reader who follows only the named file may submit against the wrong workflow. The root is crowded with process documents in general, including BACKLOG.md, CITATION.md, LISTS_OF_CITATIONS.md, and VERSIONING.md, which is the file the versioning section defers to for detailed rules.
The agreement is the part with teeth. It sits in front of a dataset whose outbound terms are attribution-only and permissive, and the visible text does not say how the two interact for a third party who contributes new dataset instances. If you are supplying labelled well data rather than code, that is the question to ask before you upload anything.
torch, timm and a live notebook server are unconditional dependencies
The toolkit declares thirty-one runtime dependencies, and the deep learning stack is not optional. `torch>=2.7.0`, `torchvision>=0.22.0`, `torchmetrics>=1.6.2`, and `timm>=1.0.15` sit in the main dependency list alongside the scientific set of numpy, scipy, pandas, scikit-learn, matplotlib, seaborn, plotly, imbalanced-learn, statsmodels, PyWavelets, scikit-image, and Pillow. The extras table holds only development tooling: mypy, a bounded ruff, and pytest.
So a user who only wants to read the Parquet files and fit a baseline classifier installs a GPU-oriented training stack. On top of that sit ipykernel, three notebook packages, tornado, nbconvert, and nbformat. That is a notebook server, pulled in by importing a time-series library.
The bounding policy is inconsistent inside the same list. Most entries carry a floor only, so numpy and torch float upward freely. Two carry a ceiling as well, `tornado>=6.5.5,<7` and `requests>=2.31,<3`. Three are pinned to an exact patch, `ipywidgets==8.1.7`, `jupyter-client==8.6.3`, and `jupyter-core==5.8.1`. And two have no constraint at all: nbconvert and nbformat.
Those last two are the pair worth worrying about, because they are the exporters that the pinned notebook frontend has to agree with, and they are the only dependencies with no floor and no ceiling.
Four of six package URLs point at the repository root, and the packaged README is a different file
The project URLs block sets Homepage, Repository, Documentation, and Source Code to the same address, the repository root. Only Bug Tracker and Changelog go anywhere else, pointing at the issues page and the releases page. The repository nevertheless carries a docs directory, a paper directory, a community directory, and an images directory.
So a package index will show the repository landing page where documentation is advertised, and the four identical entries give no hint that separate material exists in the tree.
The more consequential entry is the readme field, which points at `toolkit/ThreeWToolkit/README.md` rather than the root README.md. Everything a reader normally gets from this project, the table of contents, the motivation with its production and vessel cost figures, the governance history, the licence split, and the versioning rules, lives in the root file. The root README also identifies the repository as the first one Petrobras published on GitHub, a claim a package page cannot make for you.
Read the two files as different documents. The one on the index describes the toolkit package; the one at the root describes the project.
The README ends on an empty heading under a table of contents promising four more sections
The visible text stops at a bare `##` immediately after the sentence pointing at VERSIONING.md for detailed versioning rules. The heading above that break is Versioning. Nothing follows it.
The table of contents at the top promised considerably more: a 3W Dataset section with Structure and Overview, a 3W Toolkit section with Structure, Incorporated Problems, Examples of Use, and Reproducibility, and a 3W Community section. The questions entry in that same list resolves to the empty heading the document ends on. None of those bodies are present in what is visible here.
Some of that detail exists elsewhere in the tree, in 3W_DATASET_STRUCTURE.md and 3W_TOOLKIT_STRUCTURE.md, so the boundary is about where the prose lives rather than whether it was written. What is not recoverable from the visible text is the set of incorporated problems, any worked example of use, and the project's own statement of what reproducibility means for a dataset assembled from three different sources.
That last point matters most. The name comes from instances composed of three different sources containing undesirable events in oil wells, and the motivation section puts a figure on the stakes, losses reaching 5% of production in certain scenarios and a maritime probe costing more than US $500,000 per day. The governance section dates the launch to May 30, 2022 under the Flow Assurance department and CENPES, and says that from May 1st, 2024 the Well Integrity department joined the governance. Those facts are all in the visible text; the reproducibility claim is not.
Editorial conclusion
Treat 3W as a research resource rather than a pinned dependency: cite the dataset and toolkit versions separately from the repository tag, read LICENSE coverage for the file you actually want, and expect a heavy install that pulls torch and a notebook server whether or not you need them. Before building on it, check that the coverage badge matches your fork, decide whether the three exact notebook pins will collide with your own environment, and confirm which README you are quoting, since the package index shows a different file from the repository root.
Frequently asked questions
What is the 3W dataset from Petrobras and what does the name mean?
It is a public dataset of undesirable events in offshore oil wells, composed of instances from three different sources, which is where the name 3W comes from. It is published alongside the 3W Toolkit, a Python package for experimenting with those events.
Under what license is the 3W dataset published?
The data files are under Creative Commons Attribution 4.0, stated as covering the Parquet files saved in subdirectories of the dataset directory. The code is under Apache 2.0, and a contributor license agreement also sits in the repository for inbound contributions.
How is the 3W project versioned, and how do I pin it?
Three counters run independently: the toolkit version sits in pyproject.toml, the dataset version sits in dataset/dataset.ini, and the project version is a git tag updated on every commit to main. Recent tags are v.1.90.0, v.1.89.0, and v.1.88.0, so pin the toolkit and dataset versions separately from the repository tag.
Does the 3W toolkit require a GPU deep learning stack?
Installing it pulls torch, torchvision, torchmetrics, and timm as unconditional runtime dependencies, along with a Jupyter stack including ipykernel, ipywidgets, jupyter-client, jupyter-core, nbconvert, and nbformat. The optional extras contain only development tooling such as mypy, ruff, and pytest.
Who maintains 3W and what does the project say about governance?
The project was publicly launched on May 30, 2022, led by the Petrobras department responsible for Flow Assurance together with CENPES. From May 1st, 2024 its governance has included the department responsible for Well Integrity as well.
Why does Petrobras say 3W matters for offshore wells?
The stated motivation is that timely detection of undesirable events can prevent production losses, cut maintenance costs, and reduce environmental accidents. The text puts those losses at up to 5% of production in certain scenarios and a maritime probe at more than US $500,000 per day.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/petrobras-3w)