# Eyeballer: a CNN that labels pentest screenshots, and where its recall falls short

> Eyeballer is BishopFox's TensorFlow classifier that sorts web penetration test screenshots into five labels so a large scope can be triaged. It is a pure post-capture step with no scanning of its own, and the per-label recall figures decide how much trust a triage queue built on it can carry.

**BishopFox/eyeballer** — Convolutional neural network for analyzing pentest screenshots

- Repository: https://github.com/BishopFox/eyeballer
- Stars: 1,291 · Forks: 150
- Language: Python
- License: GPL-3.0
- Published: 2026-09-30 · Updated: 2026-09-30 · Language: en
- Canonical page: https://hysenlabs.com/projects/bishopfox-eyeballer

## Five fixed labels over rendered pixels, with no scanning of its own

Eyeballer is a convolutional neural network, written in Python on TensorFlow, that reads screenshots of web hosts and labels them. The scenario it is built for is a large-scope network penetration test, where a screenshotting tool has already produced images for a very large set of hosts and no human is going to look at all of them. The workflow the README describes is a two stage one. You run a screenshotting tool such as EyeWitness or GoWitness exactly as you normally would, then point Eyeballer at the resulting PNG files. It does not scan and it does not touch the hosts; it classifies images that already exist. The label set is fixed at five: Old-Looking Sites, Login Pages, Webapp, Custom 404 and Parked Domains. Predictions are multi-label rather than a single class, which follows from what the labels describe. An image can carry several at once, and the README's own accuracy definition confirms it, since an image counts as a failure if any single one of its labels is wrong. The evaluate mode confirms the same shape from the other end, reporting per-label metrics plus a none of the above pseudo-label for screenshots that match nothing.

## Why pixels instead of grepping for an input tag

The alternative Eyeballer argues against is a heuristic, and the README is unusually blunt about why it loses. For login pages it says you might think a simple heuristic would find them, but that in practice it is really hard, because modern sites do not just use a simple input tag you can grep for. That is a precise statement about the trade-off. Text matching reads markup the target controls and can be made to say almost anything, while Eyeballer reads the rendered pixels, which a server cannot falsify without changing what the user sees. The cost of the choice is equally concrete. The classifier can only report what it has seen before, so a login page painted in a style absent from the training set is a miss, and the output carries no confidence threshold the README describes, so a miss is indistinguishable from a confident negative. The second comparison is with the screenshotting tools the README names. EyeWitness and GoWitness do the capturing, Eyeballer consumes their output, and the instruction to use your favourite screenshotting tool as normal is what keeps the two separable.

## requirements.txt, a weights file you download yourself, and two output files

Setup is one requirements file and a weights path. The requirements are listed but not constrained: requirements.txt names Augmentor, click, matplotlib, numpy, pandas, pillow, scikit-learn, tensorflow, jinja2, progressbar2 and scipy with no version numbers, so a fresh install resolves whatever each distribution currently serves. Install it the way the README does, against the system interpreter:

```bash
sudo pip3 install -r requirements.txt
```

A second file, requirements-gpu.txt, covers a machine with a GPU, and the README states plainly that getting TensorFlow onto that GPU is outside its scope:

```bash
sudo pip3 install -r requirements-gpu.txt
```

The pretrained weights are not in the repository. They sit in the releases section as a file following the pattern bishop-fox-pretrained-vN.h5, and the path is passed explicitly on every run, so there is no default model. Start with a single image to see the output shape:

```bash
eyeballer.py --weights YOUR_WEIGHTS.h5 predict YOUR_FILE.png
```

Once one file behaves, hand it the whole directory:

```bash
eyeballer.py --weights YOUR_WEIGHTS.h5 predict PATH_TO/YOUR_FILES/
```

Two files come back, results.html so a person can browse the run and results.csv for whatever comes next. The top-level file list includes prediction_output_template.html and jinja2 is among the requirements, which is how the report gets rendered, but the README never documents the columns or their order in either output. A pipeline built on results.csv is therefore reading a format nobody pinned down.

## The 1.6x aspect ratio requirement fails silently

The first practical failure mode is not a crash. The README asks for screenshots captured at a native 1.6x aspect ratio and gives 1440x900 as the example, because Eyeballer scales an image down to the size the network expects and squishes it when the proportions are wrong. A screenshotting tool left on a different default, or an image already cropped or letterboxed upstream, produces a run that completes normally and still prints labels for all five classes. Nothing in the two output files marks a well proportioned input apart from a squished one, and no error is raised. Since the training screenshots are described as resized to 224x224, the correction belongs at capture time rather than in a preprocessing step you can add later. On a large engagement the images come from many hosts and many templates, so the capture configuration is the single thing worth standardising before the classifier sees anything.

## Where the triage queue actually breaks: 62.20% recall on Old Looking

The published numbers show where the operational risk sits. Overall Binary Accuracy is 93.52%, the chance that any one label on an image is correct. All-or-Nothing Accuracy is 76.09%, and for a work queue that second figure is the honest one, because an image counts as a failure if any of its labels is wrong, which means nearly a quarter of screenshots are mislabelled somewhere even though the per-label number looks excellent. The per-label table is more revealing still. Old Looking carries 91.70% precision against 62.20% recall, so the label the README treats as the most valuable target class misses a little over a third of the sites that belong in it, and a site never labelled is a site nobody looks at. Parked Domain is the weakest on both counts at 70.99% precision and 66.43% recall, and it is the label whose stated purpose is to shrink the scope. Custom 404 runs the other way, 91.01% recall against 80.20% precision, which means live pages get discarded as cosmetic. The evaluation set is 20% of the screenshots chosen at random and never used in training. The README does not document whether that split is stratified by label or grouped by host, and both answers would change how far these numbers travel to a new set of targets.

## Retraining from the Kaggle set: images, labels.csv and a GPU of your own

Training is one command against a fixed data contract. The training data is published on Kaggle under the link the README gives, spelled pentest-screensots, and two things are needed from it: an images folder holding the screenshots already resized to 224x224, and a labels.csv holding the labels. Both are copied into the root of the Eyeballer code tree rather than into a separate data directory, so the layout is rigid and refreshing the data means overwriting files at the top of the checkout. The command itself is:

```bash
eyeballer.py train
```

It writes a model file, weights.h5 by default, which you then pass to the other modes with the same --weights flag. The README says a machine with a good GPU is needed for this to finish in a reasonable time and leaves the hardware setup out of scope, so retraining is not a laptop task. Evaluating whatever you just produced looks like prediction with an explicit weights path:

```bash
eyeballer.py --weights YOUR_WEIGHTS.h5 evaluate
```

That mode reports precision and recall for each label including the none of the above pseudo-label, and that last figure is the one to watch after a retrain, since a model that starts routing a whole class into none of the above empties one column of the queue without any visible failure.

## A 2026 push, 2021 weights, unconstrained requirements and GPL-3.0

Two dates frame the operational risk, and they point in different directions. The last push to the repository was on 2026-03-08, while the two published releases are much older: 2.0 on 2021-03-01 and 3.0 on 2021-04-22. The code has therefore been touched more recently than the weights anyone actually downloads, and the README names no other channel for newer ones. Nothing in it promises a maintenance cadence, so those two dates are the whole signal. The unpinned requirements are the second concern. Because requirements.txt carries no version constraints, the TensorFlow an install resolves today is not the one the 2021 weights were written against, and a major step in that library is exactly the kind of change that turns a working load into an error. The repository does include a tests/ directory, which is where to look before attempting an upgrade. On licensing, the top-level LICENSE file is GPL-3.0 and the weights are distributed from the same repository's releases, so the same grant covers the model file and not only the Python. Whether that matters for your use is a question for your own counsel; the practical observation is that the licence is copyleft and it reaches the .h5 you drop into your pipeline.

## Conclusion

Eyeballer earns its place when a large-scope web test produces more screenshots than anyone will read, and only for that: it labels images, it does not scan, and it needs a screenshotting tool to do the capturing. The reason to look closely before adopting it is the distance between the headline 93.52% per-label accuracy and the 76.09% all-or-nothing figure, together with the 62.20% recall on Old Looking, because a triage step that silently drops a third of the oldest sites is worse than no triage step. Do not adopt it if your capture pipeline cannot hold the 1.6x proportion the README asks for, since squashed input still returns confident labels. Verify three things first: which weights file you actually downloaded, given that the newest release is 3.0 from 2021-04-22; whether the 20% holdout is grouped by host, which the README does not say; and the columns inside results.csv, because no output schema is documented, before anything in your process starts using it to cut scope.

## FAQ

### How do I run Eyeballer against a single screenshot?

Pass the weights path explicitly together with the predict mode and the file, in the form eyeballer.py --weights YOUR_WEIGHTS.h5 predict YOUR_FILE.png. The weights are not bundled with the code; they are published in the repository's releases as a file named bishop-fox-pretrained-vN.h5.

### What aspect ratio should the screenshots be?

A native 1.6x proportion, 1440x900 in the example the README gives. Eyeballer scales images down by itself, but the README warns that a wrong aspect ratio squishes the picture and that this affects prediction performance.

### Can I train my own Eyeballer model?

Yes. eyeballer.py train writes weights.h5 by default, and the README asks for a machine with a good GPU to finish in reasonable time. The training data is a Kaggle set from which you need an images folder resized to 224x224 and a labels.csv, both copied into the root of the code tree.

### What is the difference between Overall Binary Accuracy and All-or-Nothing Accuracy?

Overall Binary Accuracy, 93.52% in the README, is the chance that a single label is correct. All-or-Nothing Accuracy, 76.09%, counts an image as a failure if any one of its labels is wrong, so it measures whole-image correctness rather than per-label correctness.

## Sources

- [BishopFox/eyeballer on GitHub](https://github.com/BishopFox/eyeballer)
- [Issues](https://github.com/BishopFox/eyeballer/issues)
- [License: GPL-3.0](https://github.com/BishopFox/eyeballer/blob/master/LICENSE)
- [README](https://github.com/BishopFox/eyeballer/blob/master/README.md)
- [Releases](https://github.com/BishopFox/eyeballer/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/bishopfox-eyeballer
