Nextcloud Recognize: on-premises face, object and music tagging for your photo library
👁 👂 Smart media tagging for Nextcloud: recognizes faces, objects, landscapes, music genres
At a glance
- What is it?
- Recognize is a Nextcloud app that runs TensorFlow.js models on your own server to tag photos, videos and audio. It is built for self-hosters who want automatic tagging without shipping their library to a cloud API, and the trade-off is roughly 4GB of RAM and a model download measured in gigabytes.
- Who is it for?
- Adopt Recognize if you run a self-hosted Nextcloud on x86 64-bit hardware with a few gigabytes of RAM to spare and you want face, object and music genre tags without sending media to a third party. Skip it if your instance relies on end-to-end encrypted files, since the server cannot read them and recognition will not run, or if you are on Nextcloud AIO and expect native-speed inference.
- Can I use it commercially?
- Yes, with strict conditions. AGPL-3.0 is a network copyleft licence: if people use a modified version over a network, for example as a hosted service, you must offer them its source code under the same licence.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly PHP, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Recognize adds to a Nextcloud photo library
Nextcloud stores files. It does not know that a JPEG contains a face, a mountain or a plate of food. Recognize fills that gap by walking your media collection and attaching Nextcloud Collaborative Tags to each item. Those tags are not a private index inside the app: they are the same tags the rest of Nextcloud uses, which is why the README points to the Memories app for browsing photos grouped by face and to the audioplayer app for listening to music sorted by genre.
The scope is wider than photos. The app recognizes faces and groups photos by the faces that appear, detects animals, landscapes, food, vehicles and buildings, identifies landmarks and monuments, classifies audio into music genres, and recognizes human actions in video. The audience is narrow but clear: someone running their own Nextcloud instance who wants searchable, browsable media without handing the library to an external vision API. The README states plainly that no sensitive data is sent to cloud providers and that all image processing happens on the Nextcloud machine.
The model pipeline behind the tags
Recognize is a PHP app that shells out to a Node.js process. The front end is PHP and Vue; the recognition work happens in JavaScript under Node.js, with TensorFlow.js bundled inside the app. The README lists the specific models: a pre-trained EfficientNet v2 for ImageNet object detection, a model trained on the Landmarks v1 dataset for landmarks, face-api.js for extracting and comparing face features, a Musicnn network for music genres, and a pre-trained MoViNet model for video classification.
That split explains the install requirements. Because inference runs through TensorFlow.js in Node.js, the app needs the right native binaries for your architecture. The README describes two modes. Native speed requires an x86 64-bit processor with AVX support and a system with glibc. Sub-native speed uses WASM mode and works on x86 64-bit, arm64 and armv7l without AVX, and on glibc or musl systems. The practical consequence: Alpine Linux, FreeBSD and Nextcloud AIO are glibc-free, so they fall into the WASM path. The app also needs roughly 4GB of free RAM.
Model sizes are not small. Object recognition is 1GB, landmark recognition 300MB, video action recognition 50MB and music genre recognition 50MB. On a small VPS, that download and the resident memory are the real cost of adoption, not the PHP code.
Installing Recognize and tagging your first folder
The one-click route is the intended one. In your Nextcloud instance, go to Apps, search for "recognize" and click install. The README links a manual-install wiki page as the fallback when that fails. Two prerequisites apply before you start: PHP 8.0 or later, and the Collaborative Tags app must be enabled.
If you prefer to build from source, the README gives this sequence. It clones the repository into your apps directory and runs the project Makefile, which pulls npm and composer dependencies and builds the JavaScript.
cd /path/to/nextcloud/apps/
git clone https://github.com/marcelklehr/recognize.git
cd recognize
makeThe manual dependency list is make, git, Node.js v16.x with npm, PHP 8.0 or later, and composer. After the build finishes, enable the app in Nextcloud and open Settings, then Recognize. All configuration lives on that page. The first run will download the models, and the README's size table tells you what to expect on disk and in memory.
The only per-folder control is file-based. To exclude a directory from image recognition, place a .noimage file in it. For music, use .nomusic; for video, .novideo; to exclude a folder from every recognizer, use .nomedia. There is no documented per-tag or per-model toggle beyond these markers and the settings page.
If you run Nextcloud in Docker, the README notes that the app stages files in /tmp and suggests mounting a tmpfs there for speed. The docker compose snippet it gives is:
app:
image: nextcloud:26
volumes:
- type: tmpfs
target: /tmp:execThe README warns that your RAM must be large enough to hold big files, otherwise public uploads will fail. For docker run, the equivalent flag is --mount type=tmpfs,destination=/tmp:exec.
Where Recognize stops working
End-to-end encryption is a hard boundary. The README states that end-to-end encrypted files cannot be processed, because the server by design cannot read them. If your instance depends on that feature for the media you care about, Recognize has nothing to work with and no configuration will change that.
Hardware is the second boundary. The native backend needs x86 64-bit with AVX and glibc. That rules out common self-hosting choices: Alpine-based containers, FreeBSD, and Nextcloud AIO, which the README names explicitly. You can still run the app there, but through WASM, and the README calls that sub-native speed without quantifying the gap.
The maintenance note in the README is worth reading before you plan around the app. It says the app is currently maintained with limited effort, that the main functionality works for the majority of use cases, and that the maintainers will ensure it continues to work for future releases. That is a candid statement, not a guarantee of new model support. The repository shows a single maintainer, Marcel Klehr.
Memory is the last constraint. The 4GB figure is a floor, and the tmpfs advice for Docker cuts against it: the more of /tmp you keep in RAM, the less headroom you have for the models themselves.
Recognize compared with Immich's machine learning
The closest alternative for a self-hoster is Immich, which also performs face and object recognition locally. The difference is architectural. Immich is a standalone photo service with its own library, its own database and its own machine-learning container, so adopting it means moving your photos into it. Recognize is an app inside Nextcloud: it reads the files Nextcloud already manages and writes results back as Collaborative Tags, so nothing moves and other Nextcloud apps can consume the tags.
That difference decides the choice. If your library already lives in Nextcloud and you want tags visible to Memories and the audioplayer, Recognize is the lower-friction option. If you are starting fresh and want a dedicated photo application with its own interface, Immich is a different product with a different migration cost. Recognize's own README does not position it against Immich, and it does not claim to match a dedicated photo service on search or browsing.
Licence, packaging and upgrade cost
The repository is licensed AGPL-3.0, and the COPYING file is at the top level. The package.json declares MIT for the JavaScript package, which is not the same thing as the app licence; if you plan to redistribute or modify the app, read COPYING rather than the package manifest, and treat the discrepancy as something to confirm with the project rather than resolve yourself. This is not legal advice.
Upgrades are release-based. The recent tags are v13.1.0 and v12.0.2, both dated 2026-08-26, plus v12.0.1 from 2026-08-25. The Makefile carries version+=13.1.0 and a release target that strips dev dependencies and binaries before packaging, which means a packaged build re-downloads the TensorFlow binaries at install time. Plan for that download on every fresh deployment, not only the first.
The cost that matters over time is disk and RAM, not licensing. The models sit on the server, and the README's 1GB object model plus 300MB landmark model are the bulk of it. The last push to the repository was on 2026-09-08, so the project is not archived, but the README's own "limited effort" wording sets the expectation for how quickly new capabilities arrive.
Editorial conclusion
Adopt Recognize if you run a self-hosted Nextcloud on x86 64-bit hardware with a few gigabytes of RAM to spare and you want face, object and music genre tags without sending media to a third party. Skip it if your instance relies on end-to-end encrypted files, since the server cannot read them and recognition will not run, or if you are on Nextcloud AIO and expect native-speed inference. Before installing, confirm that the Collaborative Tags app is enabled and check whether your host has glibc, because that single detail decides whether you get the native backend or the slower WASM fallback.
Frequently asked questions
How do I use Recognize with Nextcloud?
Install it from Apps in your Nextcloud instance by searching for "recognize", or build it manually with git clone and make in the apps directory. Then open Settings, Recognize and let it process your media, using .noimage, .nomusic, .novideo or .nomedia files to exclude folders.
What is Nextcloud Recognize?
It is a Nextcloud app that goes through your media collection and adds fitting Collaborative Tags automatically, recognizing faces, objects, landmarks, music genres and human actions in video. All processing runs on your Nextcloud machine using TensorFlow.js in Node.js.
What does Nextcloud Recognize need to run?
PHP 8.0 or later, the Collaborative Tags app enabled, and about 4GB of free RAM. For native speed you need x86 64-bit with AVX and glibc; other systems such as Alpine Linux and Nextcloud AIO run in the slower WASM mode.
Can Nextcloud Recognize process end-to-end encrypted files?
No. The README states that end-to-end encrypted files cannot be processed by recognize, because the server by design cannot read them.
How large are the models Nextcloud Recognize downloads?
Object recognition is 1GB, landmark recognition 300MB, video action recognition 50MB and music genre recognition 50MB, according to the README's model size table.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/nextcloud-recognize)