biniou: a self-hosted webui that bundles 30+ generative AI models behind one Gradio interface
a self-hosted webui for 30+ generative ai
At a glance
- What is it?
- biniou is a Python webui from Woolverine94 that wraps Stable Diffusion, Flux, Whisper, Bark, AudioCraft and llama.cpp chatbots in a single local install. It targets people who want one interface for many model families, and it pays for that breadth with pinned dependencies and a large disk footprint.
- Who is it for?
- Adopt biniou if you want one local interface across image, audio, video and chat models on a machine with at least 8GB RAM and enough disk for many checkpoints, and you accept that the dependency pins are tight. Do not adopt it if you need a supported API surface, a documented rollback path, or a single-purpose tool you can upgrade piecemeal.
- Can I use it commercially?
- Yes, with conditions. GPL-3.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What biniou actually solves, and for whom
Running several generative model families locally usually means several front ends. One folder for a Stable Diffusion webui, another for a Whisper transcription script, another for a llama.cpp chat loop, each with its own environment and its own way of downloading weights. biniou's premise is that this fragmentation is the problem worth solving. The README describes it as "a self-hosted webui for several kinds of GenAI" and states it can run "even without dedicated GPU and starting from 8GB RAM".
The intended user is someone with a single machine, often a desktop or a home server, who wants image generation, image restoration, speech transcription, music generation and a local chatbot reachable from one browser tab. The repository topics list the model families it touches: diffusers, flux, stable-diffusion, controlnet, ip-adapter, animatediff, photomaker, gfpgan, real-esrgan, insightface, whisper, bark, audiocraft, llama-cpp-python. That is a wide net. It is not aimed at someone who wants a minimal, scriptable image pipeline, because the interface is Gradio and the unit of work is a form in a browser.
The README also claims it "can work offline (once deployed and required models downloaded)". That is the second half of the pitch: after the initial setup, no network calls are needed to generate.
How the pieces fit: Gradio front end, Python modules, Hugging Face weights
The architecture is visible from the repository layout. webui.py is the entry point, and webui.sh launches it inside a virtual environment (the Dockerfile shows webui.sh being made executable and used as ENTRYPOINT). The lang/ directory holds interface translations, which matches the README's update notes mentioning French chatbot models. The ressources/ directory and the images/ directory hold static assets.
Each capability is a module. The weekly update entries reference an "Image Variation module", a "GFPGAN module" and a "Real ESRGAN module", so modules are the unit of feature work. Underneath, the heavy lifting comes from a pinned dependency set in requirements.txt: diffusers 0.34.0 for image pipelines, transformers 4.51.3, accelerate 1.9.0, peft 0.17.1 for LoRA adapters, audiocraft 1.2.0 and bark for audio, llama-cpp-python for GGUF chat models, plus gfpgan, insightface, controlnet-aux and mediapipe for the vision and face tooling.
Model weights are not vendored. They are pulled from Hugging Face, and the README's update log is essentially a list of new repository IDs added to the supported set, for example bartowski GGUF chat models and various Flux and SDXL LoRA repositories. The Dockerfile creates /home/biniou/.cache/huggingface, which confirms the cache location for those downloads.
The interface itself is gradio 3.50.2. That version pin matters more than it looks: Gradio 3.x and 4.x differ in component APIs, so anyone trying to patch the UI has to work against the older major version.
Installing biniou on Linux, Windows, macOS or Docker
The README links platform-specific installation instructions for OpenSUSE, RHEL-family distributions, Arch, Mandriva, Debian and Ubuntu, Windows 10 and 11, macOS Intel (marked experimental), and Docker. The repository root carries the scripts those instructions call: install.sh, install_win.cmd, and per-distribution helpers named oci-debian.sh, oci-rhel.sh, oci-opensuse.sh, oci-mandriva.sh and oci-arch.sh.
The Linux path is a clone followed by the installer script. Run these from the directory where you want the project to live:
git clone --branch main https://github.com/Woolverine94/biniou.git
cd biniou
./install.shinstall.sh creates the virtual environment and installs requirements.txt. Expect this step to be slow and disk-hungry, because the dependency list includes torch, xformers, jax, jaxlib and a git-sourced Real-ESRGAN. When it finishes, start the interface with the launcher script:
./webui.shThe Dockerfile gives the same result without touching your system Python. It builds on debian:bookworm-slim, creates a biniou user, clones the main branch, runs install.sh, and exposes port 7860:
FROM debian:bookworm-slim
ENV DEBIAN_FRONTEND=noninteractive
RUN apt-get update && apt-get -y upgrade
RUN apt-get install -y bash sudo apt-utils git pip python3 python3-venv python3-pkgconfig libavformat-dev libavdevice-dev gcc perl make ffmpeg openssl libtcmalloc-minimal4
RUN adduser --disabled-password --gecos '' biniou
USER biniou
RUN cd /home/biniou && git clone --branch main https://github.com/Woolverine94/biniou.git
WORKDIR /home/biniou/biniou
RUN ./install.sh
EXPOSE 7860/tcp
ENTRYPOINT ["/home/biniou/biniou/webui.sh"]Note what the Dockerfile does not do by default: the CUDA-enabled PyTorch replacement is present but commented out, so the image ships with the CPU-only build. The commented lines show the intended upgrade path, which uninstalls torch, torchvision, torchaudio and llama-cpp-python, then runs update_cuda.sh. For a first real use, start the webui, open port 7860 in a browser, pick an image module, and let it download its default checkpoint on the first generation. That first run is the slowest; subsequent runs hit the Hugging Face cache. Upgrades are handled by update.sh, with update_cuda.sh and update_rocm.sh for the GPU backends and update_win.cmd and update_win_cuda.cmd on Windows.
Where biniou gets in your way
The dependency pins are the first real constraint. requirements.txt pins numpy==1.25.2, gradio==3.50.2, diffusers==0.34.0, transformers==4.51.3, xformers==0.0.22.post7 and torchmetrics==1.3.2 among many others. A set that tight is fragile in a shared Python environment: if you already have a working diffusers or transformers install for another project, biniou will either conflict with it or force you to keep a separate virtual environment. That is why install.sh builds one.
Disk and download cost is the second. Nothing in the repository bundles weights, and the supported model list spans SDXL checkpoints, Flux LoRAs, GGUF chat models in the 2B to 30B range, Whisper variants and audio models. The README's weekly updates add several model repositories at a time. Storage planning is on you, and the README does not document a pruning or cache-management command.
The third issue is upgrade risk. The repository ships update.sh, update_cuda.sh, update_rocm.sh and Windows equivalents, but the README does not document rollback. If an update breaks a module, the documented recovery path is not stated. The README does mention an "Experimental bugfix for CUDA users who experiments missing CUDA libraries" and asks for feedback, which tells you the CUDA path has had rough edges.
Finally, biniou is the wrong tool if you need to call generation from your own code. There is no documented HTTP API in the README, only a Gradio webui. If your pipeline needs a stable programmatic contract, a library-level tool is a better fit.
biniou compared with a single-purpose diffusion webui
The obvious alternative is a dedicated Stable Diffusion webui, such as the AUTOMATIC1111 or ComfyUI style of tool. The difference is scope, and it cuts both ways. A single-purpose diffusion front end concentrates on image generation: its extension ecosystem, its sampling controls and its community documentation all point at that one job. biniou instead spreads its module count across image, audio, video and chat, and the README's update log reflects that spread, alternating between SDXL LoRAs, Flux LoRAs, Whisper models, audio restoration models and GGUF chatbot models in the same weekly cadence.
That means biniou's image controls are likely shallower than a dedicated tool's, and its chat and audio features are likely shallower than a dedicated Whisper or llama.cpp front end. What you get in exchange is one install, one port, one interface, and one Hugging Face cache. For a hobbyist machine that is a reasonable trade. For a production image pipeline where sampler behaviour and extension compatibility decide the outcome, the dedicated tool is the better choice.
A second comparison point is packaging. ComfyUI-style tools tend to let you update nodes independently. biniou updates as a unit through update.sh, with the whole requirements.txt moving together. That is simpler to reason about and harder to partially revert.
Licence and the cost of keeping biniou current
biniou is licensed GPL-3.0. For local, personal use that is unremarkable. It becomes a real consideration if you plan to redistribute a modified biniou, ship it inside a product, or combine it with code under an incompatible licence. GPL-3.0 also covers the network-use question differently from AGPL, so a self-hosted instance you do not distribute does not trigger the same obligations. This is a description of the licence identifier in the repository, not legal advice; if you intend to redistribute, read the LICENSE file and talk to someone qualified.
The dependencies carry their own licences, and they are not uniform. requirements.txt pulls from several sources: PyPI packages, a git URL for ai-forever/Real-ESRGAN, and a pinned commit of TencentARC/PhotoMaker. Those two git dependencies are not versioned releases, they are repository snapshots at a specific commit, which means the licence terms you inherit come from those repositories rather than from biniou. Model weights downloaded from Hugging Face have their own terms as well, and the README's update log shows models from many different publishers.
Maintenance cost is mostly update frequency. The README shows a weekly update cadence through August and September 2026, with the last push to the repository on 2026-09-09. Each update adds model repositories and occasionally changes defaults, such as replacing the default model for the Image Variation module. If you pin your own environment and skip updates, you keep stability but lose new model support. If you run update.sh regularly, you inherit whatever the current requirements.txt resolves to. The repository has a single pre-release, v0.0.1 from 2024-05-29, so there is no long-term support branch to fall back on.
Editorial conclusion
Adopt biniou if you want one local interface across image, audio, video and chat models on a machine with at least 8GB RAM and enough disk for many checkpoints, and you accept that the dependency pins are tight. Do not adopt it if you need a supported API surface, a documented rollback path, or a single-purpose tool you can upgrade piecemeal. Before installing, read install.sh and requirements.txt in full, confirm your Python version matches what the installer expects, and check that your GPU backend (CUDA or ROCm) has a matching update script in the repository root.
Frequently asked questions
Does biniou need a dedicated GPU to run?
No. The README states it can run even without a dedicated GPU and starting from 8GB RAM. The default Docker image also ships with the CPU-only PyTorch build, with CUDA replacement left commented out in the Dockerfile.
Which port does the biniou webui listen on?
The Dockerfile exposes port 7860/tcp and uses webui.sh as its entrypoint, so that is the port the interface is reached on in the container setup.
Can biniou run without an internet connection?
The README states it can work offline once deployed and the required models have been downloaded. Model weights come from Hugging Face, and the Dockerfile creates the cache directory at /home/biniou/.cache/huggingface.
What licence is biniou released under?
The repository is licensed GPL-3.0. Note that requirements.txt also pulls code from separate git repositories, including Real-ESRGAN and a pinned PhotoMaker commit, which carry their own terms.
How do you update a biniou installation?
The repository root contains update.sh for the base install, with update_cuda.sh and update_rocm.sh for GPU backends and update_win.cmd and update_win_cuda.cmd on Windows. The README does not document a rollback procedure if an update breaks a module.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/woolverine94-biniou)