MTools: a Flet desktop toolbox for media work, OCR and speech
MTools 是一个功能强大的多功能桌面应用程序,集成了音视频处理、图片编辑、文本操作和编码工具,内置AI增强功能。旨在简化您的工作流程,提升生产效率
At a glance
- What is it?
- A single Python and Flet application that bundles video conversion, image utilities, local speech recognition and an MCP server, shipped as prebuilt binaries per accelerator.
- Who is it for?
- MTools is worth a look if you want one window for the small media jobs that otherwise mean three terminal tabs, and if you would rather not assemble ffmpeg, an OCR engine and a speech model yourself. The shape is honest about what it is: a personal toolbox, MIT-licensed, published under `HG-ha`, with 1343 stars and 24 open issues, last pushed on 2026-08-21 alongside the v0.2.2 release.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 46 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 20, 2026, and from our analysis. They are not legal advice.
Editorial analysis
One window for the jobs that scatter across ffmpeg and PIL
MTools is a desktop application built with Flet, which is why the whole interface is Python rather than a compiled GUI framework. It collects media processing, video and audio editing, image editing, text operations and developer helpers into a single executable, with AI features layered on top. The repository description frames the intent plainly: simplify the workflow and raise productivity.
What sits inside is a mix of recognisable pieces and specific ones. Media work leans on `ffmpeg-python==0.2.0` and `opencv-python-headless==4.11.0.86`, image handling on `Pillow==12.0.0`, array work on `numpy==1.26.4` and encoding detection on `chardet==5.2.0`. On top of those sit features borrowed from other open source projects, which the README credits rather than hides: `HivisionIDPhotos` for AI identity photos, `video-subtitle-remover` for removing burned-in subtitles, and `ICP_Query` for the ICP filing lookup, the last one written by the same author.
The developer helper side is smaller and more concrete. Text operations that handle mixed encodings and batch conversions fit naturally here, and the fact that `chardet` is a pinned dependency suggests that a real chunk of the tool list is about taking text from one messy source and putting it somewhere clean.
The layout of the repository explains how the pieces are organised: `src/` holds the application, `extensions/` the feature plugins, `tools/` the helpers, `scripts/` the packaging work, `docs/` the guides including `build_guide.md` and `mcp.md`, and `skills/mtools/` the agent-facing instructions. Two build entry points at the root, `build.py` and `flet_build.py`, are what turn that into a standalone binary.
Running from source is four commands and one version constraint
The README offers a prebuilt binary path first and a source path second, and the source path is short enough to take on faith.
git clone https://github.com/HG-ha/MTools.git
cd MToolsThe remaining two steps use `uv`, which the README recommends as the package manager. `uv sync` creates the virtual environment from `uv.lock`, and `flet run` starts the app in development mode.
uv sync
uv run flet runThe one hard constraint is worth checking before anything else, because it is narrower than the badge in the README suggests. The badge says Python 3.11+, while `pyproject.toml` sets `requires-python = ">=3.11,<3.12"`. That upper bound means a machine with Python 3.12 or 3.13 as its default interpreter will fail to resolve the environment even though it satisfies the badge.
The Flet pin is equally specific: `flet==0.84.0`, with a comment in the packaging metadata explaining that the version is deliberately not hard-coded elsewhere to avoid Flet compilation problems. That is the sort of detail that saves an afternoon when a build breaks after a dependency bump.
If you want NVIDIA acceleration and the default runtime is not using it, the README gives an explicit swap. The default on Windows is DirectML, which covers Intel, AMD and NVIDIA, and the CUDA path replaces the ONNX Runtime packages outright.
uv remove onnxruntime-directml onnxruntime
uv add onnxruntime-gpu==1.24.4The comment notes that a variant carrying the full CUDA and cuDNN environment can be installed instead, at the cost of a much larger environment.
Six Windows and Linux builds, one macOS build, and what each costs
The release assets are the interesting engineering decision in the project. Rather than one installer, MTools publishes per-platform and per-accelerator binaries, and the README explains the trade for each.
On Windows 10 and 11 x64 there are three. `MTools_Windows_amd64` is the smallest and accelerates through nvidia, amd and intel DirectML, but cannot cap VRAM manually, so the README advises it for machines with under 8GB of VRAM or for anyone still on an older NVIDIA driver. `MTools_Windows_amd64_CUDA` is mid-sized and uses a CUDA runtime you install yourself, requiring CUDA 12.x plus cuDNN 9.x. `MTools_Windows_amd64_CUDA_FULL` is the largest and ships the complete CUDA and cuDNN runtime, so it runs out of the box at a cost of more than two gigabytes.
Linux mirrors that structure with `MTools_Linux_amd64`, `MTools_Linux_amd64_CUDA` and `MTools_Linux_amd64_CUDA_FULL`, with the smallest build carrying no GPU acceleration. macOS gets one asset, `MTools_Darwin_arm64`, limited to M-series chips with Core ML acceleration. The README marks Windows and macOS as the tested platforms and Linux as experimental.
The accelerator matrix underneath explains the asymmetry. Windows defaults to `onnxruntime-directml` for DirectML across Intel, AMD and NVIDIA hardware. Apple Silicon uses plain `onnxruntime==1.24.4` with hardware acceleration through Core ML, while Intel Macs fall back to CPU because there is no GPU path. Linux also defaults to CPU with `onnxruntime-gpu` available as an option.
One constraint carries across all of them: the DirectML build has no VRAM limit, and only the CUDA path can cap how much video memory is used. If you are running long transcription jobs on a small card, that is the line that decides your download.
Local speech recognition through sherpa-onnx, not a hosted API
The AI features are local by default, which is the more interesting architectural choice. Speech recognition runs through `sherpa-onnx` with the SenseVoice and Paraformer models, and the README is explicit that the `funasr` Python package is not required for this. The repository topics still list `funasr` and `whisper`, a leftover from an earlier approach that the sherpa-onnx path replaced.
OCR is the other local model, credited to `PPOCR_v5`, and both recognition paths run through ONNX Runtime, which is what ties the accelerator story to the feature list. Subtitles and identity photos follow the same pattern, borrowing a model-backed pipeline from the credited projects rather than shipping a new one.
There is a remote option too. The README documents an OpenAI-compatible endpoint for features such as subtitle repair and translation, with the provider configuration described in `docs/ai_providers.md`, and a partner section at the top of the README lists that API provider alongside a server sponsor. If you plan to keep everything on-device, the local paths cover recognition, which is the part that matters for privacy and latency.
The practical reading is that MTools treats local inference as the default and network access as an opt-in for the tasks where a hosted model is simply better.
A Streamable HTTP MCP server on port 8765
MTools exposes its own capabilities to agents. A Streamable HTTP MCP server ships inside the application and starts on demand, defaulting to `http://127.0.0.1:8765/mcp`. The README describes the sequence as opening Settings, entering the MCP service section and enabling it.
The agent-facing side lives in `skills/mtools/`, with configuration described in `skills/mtools/install.md` and the protocol details in `docs/mcp.md`. The README names OpenClaw, Hermes and Cursor as the callers this is aimed at, and the local bind address means an agent on the same machine can reach it without exposing anything to the network.
This is the feature that changes how you should think about the project. A toolbox of GUI utilities is normally something you click through. Exposing it over MCP turns the same utilities into tools an agent can invoke, which is why the `skills/` directory sits at the top level of the repository instead of under `docs/`. The cost is that the tool surface becomes part of your agent's attack surface, so the local address and the explicit enable toggle are the two details to keep in mind before you point anything at it.
The `skills/mtools/` directory being versioned in the repository, rather than generated at runtime, also means the integration is reviewable in a pull request. For a single-maintainer project that is a meaningful property.
A young release history that describes itself honestly
The version numbers tell the story of a project that found its footing in 2026. v0.1.0 published on 2026-05-10, described as a stability and compatibility release after a run of `0.0.12-beta` builds. That release note is the most informative one in the feed, because it names what changed: the Flet runtime, cross-platform packaging, the CUDA and ONNX Runtime dependencies, macOS and Windows verification, and new capabilities covering subtitles, text to speech, playback and log diagnosis. It also carries a troubleshooting note about `onnxruntime` errors, which tells you the packaging is still the hard part.
After that, v0.2.1 on 2026-07-18 and v0.2.2 on 2026-08-21, both with only a full changelog link and no body text. The last push to the repository matches v0.2.2 exactly, on 2026-08-21, and `pyproject.toml` carries the same 0.2.2 version.
The sub-1.0 number is honest. Breaking changes between minor versions are plausible, and two of the three releases will tell you nothing beyond a compare link. What you get instead is visible activity: three releases in a little over three months, a locked dependency file, a documented build guide and a `db.sh`-style discipline for repeatable builds in other projects of this shape.
For a personal toolbox that is a healthy signal. For anything you intend to wrap in your own tooling, pin a version and read `docs/build_guide.md` before you start, because the accelerator wiring is where the complexity lives and it is where the first release note says things went wrong.
Editorial conclusion
MTools is worth a look if you want one window for the small media jobs that otherwise mean three terminal tabs, and if you would rather not assemble ffmpeg, an OCR engine and a speech model yourself. The shape is honest about what it is: a personal toolbox, MIT-licensed, published under `HG-ha`, with 1343 stars and 24 open issues, last pushed on 2026-08-21 alongside the v0.2.2 release. Two things decide whether it belongs on your machine. The Python constraint is narrow, because `pyproject.toml` sets `requires-python = ">=3.11,<3.12"`, so a 3.12 or 3.13 interpreter will not resolve. And the release cadence is young rather than slow: v0.1.0 on 2026-05-10, v0.2.1 on 2026-07-18 and v0.2.2 on 2026-08-21, with the first release notes being the only ones that describe changes in detail. Start from the releases page and pick the binary that matches your accelerator, since the DirectML, CUDA and CUDA_FULL builds differ by more than two gigabytes of bundled runtime.
Frequently asked questions
What is mtools used for?
In this repository, MTools is an MIT-licensed Python and Flet desktop application that bundles media processing, video and audio conversion, image editing, text operations and developer helpers into one window. The AI features run locally through ONNX Runtime, using sherpa-onnx with SenseVoice and Paraformer for speech recognition and PPOCR for OCR, and the project also ships a Streamable HTTP MCP server so agents can call the same tools.
How do I install mtools?
The fastest route is a prebuilt binary from the releases page, with a separate asset per platform and accelerator: DirectML, CUDA and CUDA_FULL builds for Windows and Linux x64, and a Core ML build for Apple Silicon. To run from source, clone the repository, install `uv`, then use `uv sync` followed by `uv run flet run`. Note that `pyproject.toml` sets `requires-python = ">=3.11,<3.12"`, so the interpreter must be 3.11.
Which MTools build should I download for an NVIDIA GPU?
Use the CUDA build if you already have CUDA 12.x and cuDNN 9.x installed and want a smaller download, or the CUDA_FULL build to get the complete CUDA runtime bundled at more than two gigabytes extra. The plain Windows amd64 build uses DirectML, which covers NVIDIA, AMD and Intel but cannot cap VRAM usage, so it is a poor fit for long jobs on a card with less than 8GB of memory.
Does MTools send my media or audio to a server?
The core recognition paths are local. Speech recognition runs through sherpa-onnx with SenseVoice and Paraformer and OCR runs through PPOCR, both on ONNX Runtime, so audio and images do not leave the machine for those features. The README does document an optional OpenAI-compatible provider for tasks such as subtitle repair and translation, configured through `docs/ai_providers.md`, so remote calls are opt-in rather than baked into the default flow.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/hg-ha-mtools)