# Capital One DataProfiler: four install lines, a switched-off correlation matrix, and a format block that stops mid-key

> DataProfiler loads a file into a DataFrame and returns a profile dictionary of schema, statistics and detected entities. The install extras decide whether you get the entity recognizer at all, the correlation matrix is documented but currently switched off, and every example block in the profile format section ends before its last line does.

**capitalone/DataProfiler** — What's in your data? Extract schema, statistics and entities from datasets

- Repository: https://github.com/capitalone/DataProfiler
- Website: https://capitalone.github.io/DataProfiler
- Stars: 1,590 · Forks: 192
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/capitalone-dataprofiler

## The slim extra is the one that turns entity detection off

Four install lines appear under the install heading and they do not describe four flavours of the same thing. `pip install DataProfiler[full]` takes everything. `DataProfiler[ml]` takes the machine learning dependencies without generating reports. `DataProfiler[reports]` is the escape hatch for when the machine learning requirements are too strict, with tensorflow named as the thing you might not want, and it is the one that disables the default sensitive data detection and entity recognition. `pip install DataProfiler` on its own is the bare install, and it appears last under a second line that repeats the install from pypi wording. So the headline capability of the library, the pre-trained deep learning model used to identify PII and NPI, is exactly what the smallest sensible install does not give you. Reading the four lines in the order they are printed, the most complete package comes first and the most reduced one arrives as the fallback, which is the reverse of how most people scan an install section.

## correlation_matrix is in the profile shape but switched off

The structured profile shape carries a `correlation_matrix` field, described as a matrix of shape column_count by column_count holding the correlation coefficient between each pair of columns. The footnote under that block says the correlation matrix update is currently toggled off, that it will be reset in a later update, and that users can still use it as desired with the is_enable option set to True. Three things follow. Any consumer that reads the field out of an existing profile has to cope with it being stale or absent rather than assuming fresh numbers. The switch is named `is_enable`, a name that reads like a variable rather than a keyword argument the library accepts. And the promise of a later reset means the behaviour at the moment is a temporary state documented as temporary, so a pipeline built on top of it inherits a date as well as a schema.

## Every profile format block stops before its last key

Three profile shapes are given as format blocks and none of them closes. The structured block ends on a line holding a single double quote, immediately after the opening quote of the key that would have come next. The unstructured block ends mid-type name at `dict[strin`, and the graph block is the only one that reaches its own closing brace. The per-field list has the same problem in prose: the description of `order` stops at otherwise rando, in the middle of the word for unordered data. None of these are cosmetic when you are writing a consumer. The two level split is the part worth keeping in mind, since global_stats holds dataset level values such as row_count, column_count, encoding, file_type and times, while data_stats holds the per-column or per-label entries, each with its own column_name, data_type, data_label and categorical flag. What sits past the break in each block has to be read from the API documentation or from the keys your own run produces.

## A graph profile has no global_stats wrapper at all

The graph profile does not reuse the structure of the other two. Instead of a global_stats and data_stats pair it starts at the top level with num_nodes and num_edges, then categorical_attributes and continuous_attributes as two separate lists, then avg_node_degree and global_max_component_size, and it closes with times. The distribution blocks are shaped differently from each other too, which is easy to miss when code walks them generically. `continuous_distribution` maps an attribute name to an object carrying a name, a scale and a list of properties, while `categorical_distribution` maps an attribute name to an object carrying bin_counts and bin_edges. The examples show both of them partly empty, with one attribute populated and the other set to None. The structured shape also reports profile_schema, a description of the dataset format labelling each column and its index, and the unstructured shape reports empty_line_count and memory_size in MB, none of which have a counterpart in the graph shape.

## The unstructured profile counts entities at three levels

Inside the unstructured shape, data_stats is keyed by data_label rather than by column name, and each label carries two parallel structures. entity_counts maps a label to word_level, true_char_level and postprocess_char_level counts, and entity_percentages repeats the same three keys as floating point values. The presence of two character levels is the part that tells you something about how the recognizer works: one count is taken at character level before postprocessing and the other after it, so the two numbers are not a duplicate view of the same pass but the before and after of it. word_level sits alongside as a third granularity, counting by token. Because percentages are stored next to counts under the same three keys, a consumer that reads only one of the two structures will silently disagree with code that reads the other about the same text, and nothing in the shape forces you to look at both.

## The base requirements already carry a cloud SDK and two sketches

requirements.txt holds twenty-two entries, and several of them map directly onto a feature rather than a convenience. fastavro and python-snappy cover the AVRO and snappy paths behind the auto-detected load, pyarrow and h5py cover Parquet and HDF5, and chardet plus charset-normalizer sit under the encoding field that every profile reports. psutil is what makes memory_size in MB available in the unstructured shape, and networkx is the graph library behind the node and edge counts. The numeric stack is capped rather than open: numpy is held below 3.0.0, pandas below 3.0.0, pyarrow below 24.0.0 and chardet below 7.0.0, so a recent release of any of those is outside the range this version claims. Two cardinality sketching libraries, HLL and datasketches, are unconditional. And boto3 plus urllib3 and requests are unconditional too, which is why the plain install from pypi is not the small one people sometimes assume when the word slimmer appears further down the same section.

## make setup installs five requirement files and declares three

The Makefile has three targets and each one carries a fact worth knowing. The setup target declares three prerequisites, requirements.txt, requirements-dev.txt and requirements-test.txt, and then installs five files in its recipe: those three plus requirements-ml.txt and requirements-reports.txt. So the machine learning and report dependency sets arrive whether or not make decided they were needed, which makes the extras a choice you make later at install time rather than at setup time. The same recipe creates a venv, activates it, installs the project in editable mode with `pip3 install -e .`, then runs `pre-commit install` followed by `pre-commit run`, so the first thing a fresh checkout does after installing is reformat itself. The venv path it activates is venv/bin/activate, which is the POSIX layout, so this target is written for Linux and macOS. The test target pins the random seed:

```bash
DATAPROFILER_SEED=0 python3 -m unittest discover -p "test*.py"
```

A fixed seed means a failing test is meant to fail the same way twice, which is the right call for a library whose output includes sampled statistics.

## setup.py edits the README text to build its own description

The packaging step assembles several pieces by reading files rather than by declaring values. The long description is produced by reading README.md line by line, waiting for the line that contains `<p text-align="left">`, buffering everything from there until the matching `</p>`, and then deleting that span from the text with a string replace. What gets removed is the picture element that carries the logo, so the PyPI description ends up without the README's image markup and with the rest of the page intact. Install requirements are gathered the same way, by reading requirements.txt, requirements-ml.txt and requirements-reports.txt and splitting each into lines, which is why adding a dependency is a file edit rather than a setup.py edit. The version is not written down anywhere: it comes from versioneer.get_version(), with the matching command classes pulled from the same module, and a versioneer.py file sits at the repository root to do it. python_requires is set to >=3.10, and a resource_dir constant set to resources appears in the visible portion of the file.

## Conclusion

DataProfiler earns its place when a dataset arrives without a schema and somebody needs column types, null ratios and a first pass at sensitive values in one object, and the extras make that decision honest: the cheap install is cheap because it drops the entity recognizer that most people came for. Before choosing, check three things. Whether your platform is inside the numpy and pandas caps the base requirements set. Whether the slim extra is acceptable for your use, because with it the sensitive data detection that the project leads with is simply not there. And whether you need the correlation matrix, which is in the profile shape but off unless you set is_enable yourself. Read the profile keys against your own consumer code rather than against the examples, since the field lists on the page stop partway through and the version you install comes from versioneer rather than a pinned number.

## FAQ

### How do you install Capital One's Data Profiler without TensorFlow?

Use `pip install DataProfiler[reports]`, the slimmer package for when the machine learning requirements are too strict. It disables the default sensitive data detection and entity recognition, so PII and NPI detection is absent from that install.

### What is the difference between the DataProfiler full and ml extras?

`DataProfiler[full]` installs the full package from pypi, while `DataProfiler[ml]` installs the machine learning dependencies without generating reports.

### Does the capitalone/DataProfiler structured profile calculate a correlation matrix?

Not by default. The correlation matrix update is currently toggled off, and the page says users can still use it as desired by setting the is_enable option to True.

### What is the difference between global_stats and data_stats in a Data Profiler profile?

global_stats holds dataset level values such as row_count, column_count, file_type, encoding and times. data_stats holds the column or label level entries, each with its own column_name, data_type, data_label and categorical flag.

### Which Python versions does capitalone/DataProfiler support?

setup.py sets python_requires to >=3.10. The version number itself is not hardcoded in the packaging file, since it is read from versioneer.get_version() at build time.

### How are the Data Profiler tests run from the repository?

The Makefile test target runs DATAPROFILER_SEED=0 python3 -m unittest discover -p "test*.py", pinning the seed so a sampled statistic fails the same way on a repeat run.

## Sources

- [capitalone/DataProfiler on GitHub](https://github.com/capitalone/DataProfiler)
- [License: Apache-2.0](https://github.com/capitalone/DataProfiler/blob/main/LICENSE)
- [Project website](https://capitalone.github.io/DataProfiler)
- [README](https://github.com/capitalone/DataProfiler/blob/main/README.md)
- [Releases](https://github.com/capitalone/DataProfiler/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/capitalone-dataprofiler
