Open-source project
chn-lee-yumi/MaterialSearch avatar
chn-lee-yumi/MaterialSearch

chn-lee-yumi/MaterialSearch: a GPL front end over a closed API and an obfuscated core

Semantic search. Search local photos and videos through natural language. AI语义搜索本地素材。以图搜图、查找本地素材、根据文字描述匹配画面、视频帧搜索、根据画面描述搜索视频。

1,972 stars215 forksHTMLGPL-3.0

At a glance

What is it?
MaterialSearch indexes local photos and video frames and lets you find them by description or by picture, which is a genuinely useful local search tool. What makes it worth reading carefully is its licensing shape: the repository now holds only the front end, parts of that front end are deliberately obfuscated, and the API implementation behind it is stated to be closed source.
Who is it for?
MaterialSearch is worth installing if you have a large local photo or video library and want to search it by description without uploading anything, and you are content to run a prebuilt bundle or a container image rather than build from source.
Can I use it commercially?
Yes, with conditions. GPL-3.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
Is it still maintained?
Yes. The repository last received commits 10 days ago.
What is it written in?
Mainly HTML, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What it does, and what the author thinks of one feature

The capability list is short and specific, which is unusual for a tool in this space. You can search images by typing a description. You can search images using another image as the query, the familiar reverse-image approach applied to your own library rather than to the internet. You can search video by a description, and what comes back is matching clips rather than whole files. You can search video using a screenshot, so a frame you saw somewhere becomes a query against your footage. And there is a similarity score between an image and a piece of text.

That last one comes with an annotation in the readme that the author describes as not very useful. It is worth pausing on, because it is the kind of honesty that tells you what to trust in the rest of the document. If the author is willing to label their own feature as not worth using, the other four claims are more likely to be descriptions of what actually works rather than aspirations.

The reason this is worth building at all is the privacy property. Everything runs against your own files. There is a hosted instance, given a domain on a Chinese academic network, but the deployment instructions that dominate the readme are for a Windows bundle and a container, and the configuration takes local filesystem paths. For anyone with tens of thousands of photographs, that is the difference between a tool you can use and a tool you cannot.

The underlying model is named, and it is a Chinese text and vision model rather than a general purpose one. That choice is consistent with the distribution: the Windows bundle is offered through two domestic file sharing services alongside the release page, and the container registry offered as an alternative is a regional one, recommended specifically for users in mainland China. The project is built for a particular audience and the packaging reflects it without pretending otherwise.

The licensing shape is the part to read twice

The repository is licensed under the GNU General Public License version 3, and the readme states that obfuscation of some source files does not restrict legitimate use under those terms. That assertion is the crux, and it deserves a direct reading rather than an assumed one.

Two things are stated in the same document. Parts of the source in this repository have been intentionally obfuscated, with the reason given as past incidents in which people removed or altered copyright and attribution information, and the author notes the obfuscation is intended to protect authorship and legal integrity. Separately, the API implementation is described as not open source and maintained by the author, with API-related requests directed to an issue tracker rather than to a pull request.

The copyleft licence is built on a specific condition: if you distribute the software, you must convey the source, in a form suitable for modification, under the same terms. Obfuscated source is not that. A project that takes the permissions of a copyleft licence while removing the ability to modify the covered code is in tension with the licence it invokes, and the readme's assurance that the arrangement does not restrict legitimate use addresses the permissions side of the bargain while leaving the obligations side unaddressed. The honest way to put it is that this is a disputed reading of the licence, and it is the author, not a user, who would have to live with that.

The motive is not hidden and it is not trivial. Being copied without attribution is a real grievance, and a copyleft licence is a reasonable tool for preventing it. The grievance is real; the remedy is the contested part. Obfuscation does not stop copying, it stops a copy from being maintained, and it does so for legitimate users as much as for the people it targets. The readme does invite contributions back, and asks that notices be kept, which is a coherent position for someone who has decided that attribution is worth more to them than modification. It is a position an organisation lawyer should read before the organisation depends on the code, and this article's job is to say so rather than to repeat the reassurance.

There is one more condition on contributions, and it is unusual enough to note: generated code contributions are not accepted, and submitters are asked to ensure they wrote and understand what they submitted. Combined with the split into two repositories, the effect is a project with a small, deliberate, human-reviewed surface and a closed middle.

The repository is now a front end, and the core is somewhere else

The structural change is stated at the top of the readme and reframes everything else. What was a single project is now two: this repository, which contains only front end code, and a separate repository holding the core as a standalone Python package that can be installed with the package manager. The stated reason is that the split makes version control and distribution easier, and that developers can use the core functions directly in their own projects.

That reason is a good one, and it is the ordinary reason people split a package from its interface. The consequences for a reader are less comfortable. The configuration file everyone is told to edit lives in the other repository, not this one, so the documented configuration surface of a repository that contains no backend is a file you cannot see from it. The Docker image is built by automated workflow from a source tree that is not this one. And the front end is talking to an API whose implementation is closed, so a front end developer cannot debug a failing request, only report it.

The contribution guidance is arranged around the same split, and it is unusually explicit. Front end features go to a pull request here. API related requests go to an issue here, with a parenthetical noting the implementation is not open source. Everything else can go to either repository. That is a coherent triage scheme for a project with a closed middle, and stating it plainly is better than leaving people to discover it.

What the repository does contain is small and says so. There is a configuration file for the graphical interface, a static asset directory, a compose file, a documentation directory, and two readmes in two languages. There is also an ignore file named for a specific editor's indexing, which is a small hint that the front end is developed with an AI-assisted editor. The primary language of the repository is recorded as markup rather than a programming language, which is what a repository that contains a front end and a compose file should look like in a language statistic.

Two distributions, and the offline switch that tells you what the image contains

Deployment is documented two ways, and the Windows route is written with the specificity of something the author has field-tested many times.

The minimum supported system is Windows 10, with an explicit note that anyone still on Windows 7 should upgrade or get a newer computer. The bundle must be extracted with one particular archiver, with a warning that other tools may fail, and there is an instruction file inside the archive to read afterwards. The bundle ships in two sizes: one without the model for people who already have it or want to choose their own, and one that includes the base model and is recommended for most users. The bundle picks between integrated and dedicated graphics hardware on its own.

The container route is the one a server operator would choose. It is built for one processor architecture only, it includes the base models, and it supports graphics acceleration. Two registries are offered, one international and one regional, with the regional one recommended for users who cannot reach the first. Before starting, you supply three things: where the database should live, which host directories to scan, and what those paths become inside the container.

The compose file is where the design decisions show. Two variables are set that are worth reading carefully. The host is bound to all interfaces rather than to loopback, which is correct in a container that publishes a port, and the assets path is a comma separated pair of container mount points with a skip path for anything that should not be traversed. The volume list maps the database directory separately from the two scanned trees, so the index survives a container rebuild while the media does not have to be copied in.

Then there is the environment variable that explains the whole image. The image sets the transformers library to offline mode, which means it will not contact the model hub to check whether a newer version of the model exists. The readme explains how to change that if you want a different model, and the fact that it is off by default tells you something about the author's priorities: an image that phones home on startup is an image that can start slowly, fail behind a firewall, or change behaviour without anyone deciding it should. Two operational warnings follow from the same place. Do not set a memory limit on the container, with a link to a specific issue describing what goes wrong, and expect a scan over a network filesystem to be slow, with remote shares explicitly discouraged as an assets path.

Configuration is four variables and a list of escape hatches

Every setting is an environment variable with a default in the configuration module, which is the pattern that makes a container image and a local install behave identically. The readme shows the idiom explicitly: read a variable with a fallback, and the fallback is what you get if you set nothing.

The example is two lines:

conf
ASSETS_PATH=C:/Users/Administrator/Pictures,C:/Users/Administrator/Videos
SKIP_PATH=C:/Users/Administrator/AppData

A comma separated list of directories to index, and a directory to skip. That is the whole required configuration, and the comma separated form is worth internalising because it is the difference between one scan root and several. A skip path exists so that the scanner can avoid descending into a directory that will never contain media but will cost a great deal of time to walk, which on a machine with a large application data directory is the difference between a scan that finishes and one that does not.

After those two, the readme documents a set of escape hatches, each attached to a symptom. If a format you know you have is not being found, the extension lists can be extended. If small images are being skipped, there are minimum width and height settings to lower. If the model hub is unreachable or you need a proxy, the standard proxy variables are honoured. Each of these is a diagnostic dressed as a feature, and the pattern suggests the author's support burden: most reported problems are almost certainly a file the scanner decided not to look at, and documenting the knob that changes that decision is cheaper than debugging each report.

One recommendation is repeated with a specific reason. Remote network shares are not advised as an assets path because they slow the scan. That is a statement about the architecture: an indexer that walks a tree and reads every file to compute embeddings is dominated by input throughput, and a network filesystem turns every read into a round trip. The same advice applies to any indexer of this shape, and it is the one piece of operational guidance in the readme that will save you the most time if you read it before starting rather than after.

The troubleshooting section closes by narrowing the author's responsibility, stating that issues relating to the project's own functionality, code and documentation are their concern while other classes of problem are not, and asking that any report include the operating system, the configuration, and what the application prints when it starts. The last of those three is the useful one, because the application prints its configuration at startup, which means every bug report can carry its own reproduction environment.

Editorial conclusion

MaterialSearch is worth installing if you have a large local photo or video library and want to search it by description without uploading anything, and you are content to run a prebuilt bundle or a container image rather than build from source. Read the licensing shape before you contribute to it, because a front end licensed under the GPL with obfuscated source and a closed API behind it is a different arrangement from what the licence headline suggests, and a company assessing it for use will want a straight answer on that rather than a reassurance that the terms still apply. Operationally, expect the first scan to be slow, expect the model to be downloaded with the image, and leave the container memory limit unset, which the project's own issue tracker says causes trouble.

Frequently asked questions

What does MaterialSearch do and where does it run?

It indexes local photos and video frames and searches them by text description or by using another image or screenshot as the query. The readme documents a Windows bundle and a Docker container, both running against local filesystem paths, with a hosted instance also available at a separate address.

Is MaterialSearch fully open source?

No. The repository states that the API implementation is not open source and is maintained by the author, and that some front end source has been intentionally obfuscated to protect authorship. Front end features are accepted as pull requests, while API requests are directed to the issue tracker. The project is licensed under the GPL version 3.

Which model does MaterialSearch use and can I change it?

The base model is a Chinese text and vision model, and both the Windows bundle and the container image include it. The container image defaults to an offline mode for the transformers library so it does not contact the model hub on startup, and setting that variable to zero is required before you can substitute a different model.

Why should I not set a memory limit on the MaterialSearch container?

The readme explicitly advises against it and links to a specific issue describing the problems that appear when a limit is set. The project also warns against pointing the assets path at a network share, because it slows the scan substantially.

What configuration does MaterialSearch need at minimum?

Two environment variables: a comma separated list of directories to index and a directory to skip. Everything else has a default, and the readme documents escape hatches for specific symptoms, including extending the recognised image and video extensions, lowering the minimum image dimensions, and setting a proxy.

Official sources

  1. chn-lee-yumi/MaterialSearch on GitHub
  2. License: GPL-3.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/chn-lee-yumi-materialsearch.svg)](https://hysenlabs.com/projects/chn-lee-yumi-materialsearch)