MyDumper: parallel logical backups for MySQL, MariaDB and TiDB
Official MyDumper Project
At a glance
- What is it?
- MyDumper splits a MySQL dump across threads and writes one file per table. It is a good fit when mysqldump is too slow or too hard to parse, and the wrong fit when you need non-transactional engines covered consistently.
- Who is it for?
- Adopt MyDumper when your dump is dominated by InnoDB tables and you want per-table files plus parallel export and import; skip it if you depend on non-transactional engines, since the README states consistent snapshots for those are not provided.
- Can I use it commercially?
- Yes, with conditions. GPL-3.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- Mainly C, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem MyDumper solves for MySQL backups
mysqldump writes a single stream. One connection reads, one file comes out, and restoring means replaying that file in order. That design is fine for small databases and painful for large ones: the wall clock is bounded by one thread, and the output is hard to inspect because every table is buried in the same file. MyDumper's README states the goals directly: parallelism for speed, output that is easier to manage because tables and dump metadata are separate files, consistency via a snapshot shared across all threads, and PCRE-based inclusion and exclusion of databases and tables. The audience is the operator who has to move a multi-hundred-gigabyte schema between servers on a schedule, or who needs to restore a single table without replaying everything. The README also makes a point of ownership: MyDumper is community-maintained and is not a Percona, MariaDB or MySQL product. That matters when you are deciding who to file a bug with.
How the consistent snapshot and threading actually work
The mechanism is described step by step in the README, and it follows MySQL practice rather than inventing anything. Slow running queries on the server either abort the dump or get killed, as a precaution. The main thread acquires a global read lock with FLUSH TABLES WITH READ LOCK, then reads metadata such as SHOW SLAVE STATUS and SHOW MASTER STATUS. Worker threads connect separately and each establishes a snapshot with START TRANSACTION WITH CONSISTENT SNAPSHOT. On servers before 4.1.8 the README says it creates a dummy InnoDB table and reads from it instead. Once every worker announces that its snapshot is established, the main thread runs UNLOCK TABLES and starts queueing jobs. Two consequences follow from that sequence. First, the lock is held only for the setup window, not for the whole dump, so writes resume while tables are still being exported. Second, the snapshot guarantee is transactional: the README states plainly that this does not yet provide consistent snapshots for non-transactional engines. The extra tools are worth noting too. myloader reads the directory mydumper produced and imports it with multiple threads, and --exec lets you pipe each finished file through an external command, single threaded, with the absolute path required and FILENAME substituted wherever you place it.
Installing MyDumper and running a first dump
The README no longer carries install instructions inline. It redirects to the official documentation for installing, and the same page covers compilation requirements for building from source. There is an official Docker image on Docker Hub, and the repository ships a Dockerfile under docker/ that the README says is aimed at development and building from source locally rather than distribution. The one-liner below builds that image from the master branch of the GitHub repository with ZSTD enabled, which is the example the README gives.
docker build --build-arg CMAKE_ARGS='-DWITH_ZSTD=ON' -t mydumper \
https://github.com/mydumper/mydumper.git#master:dockerAfter that, the tool is invoked as mydumper for export and myloader for import. A first real run is easier to manage through the defaults file, which the README says is becoming more important. The repository ships mydumper.cnf and myloader.cnf as complete samples. A minimal mydumper section looks like this.
[mydumper]
host = 127.0.0.1
user = root
password = p455w0rd
database = db
rows = 10000Point mydumper at that file and you should see a directory of per-table files plus dump metadata, rather than one monolithic stream. The rows key controls chunking and is the first knob to tune when a single large table dominates the runtime. Restoring is the mirror image: a [myloader] section with the destination host and database, for example the new_db and optimize-keys = AFTER_IMPORT_PER_TABLE shown in the README's sample, then myloader against the same directory.
Tuning string primary keys with the metadata planner
Chunking a table is easy when the primary key is an auto-increment integer and harder when it is a string, because the splitter has to guess where the data is dense. The README describes a bounded metadata-assisted planner that seeds prefix-based root chunks before falling back to the existing recursive splitter, and it says the defaults keep the current behavior as a safe fallback. The planner exposes three orthogonal bounds. --string-pk-planner-target-rows-per-prefix is the desired chunk size in rows; when it is 0, which the README gives as the default, the target is derived as table_rows divided by --string-pk-planner-max-prefixes, and a positive value is used directly. The planner deepens prefixes, using more leading characters, until each prefix's estimated row count is at or under that target. --max-char-size, default 2, caps how many leading characters a prefix may use and therefore the maximum planning depth; the README notes a larger value allows finer chunks on skewed key distributions at the cost of more EXPLAIN probes. The mode is selected with --string-pk-planner=auto, metadata or recursive. Read that as a trade-off rather than a free win: more planning probes mean more time spent before any rows move.
Where MyDumper is the wrong tool
The consistency model is the hard boundary. Because snapshots are established with START TRANSACTION WITH CONSISTENT SNAPSHOT, the guarantee applies to transactional engines, and the README says non-transactional engine support is not provided. If your schema still relies on MyISAM tables for anything you care about, a MyDumper backup can capture those tables in a state that never existed at a single point in time. That is a correctness problem, not a performance one. The second boundary is operational. The dump takes a global read lock at the start and may abort or kill slow queries, so running it against a busy primary has a visible cost; the README presents this as following best MySQL practices, but it is still a write stall for the duration of the setup window. Third, --exec is single threaded, so a pipeline built on it will not scale the way the export itself does. Fourth, the README points at documentation described as work in progress, and the install and usage sections have been moved out of the README entirely, so expect to read the generated docs site rather than the repository front page.
MyDumper compared with mysqldump and physical backups
The comparison people search for is mydumper vs mysqldump, and the difference is architectural. mysqldump emits one serialized stream over one connection; MyDumper opens multiple connections, each with its own consistent snapshot, and writes separate files per table plus metadata. That is why MyDumper can be faster on large schemas and why its output is easier to parse or partially restore. mysqldump remains the tool with no extra install step, since it ships with the server. Against physical backup tools the split is different again: MyDumper produces logical SQL output, which is portable across versions and engines and readable with a text editor, while a physical copy is tied to the storage format. Logical output also means restore time is dominated by replaying SQL, so a backup that exports quickly can still take a long time to load. Within the logical category, MySQL Shell's dump utilities and mysqlpump are the other names that come up; the README does not compare MyDumper to either, so treat any such comparison as something to verify against their own documentation rather than MyDumper's.
Maintenance, version drift and the GPL-3.0 licence
The repository is not archived and the last push was on 2026-09-23, with releases v1.0.5-1, v1.0.6-1 and v1.0.8-1 landing in the weeks before that. Version drift is the practical upgrade cost. The README documents a rename in the defaults file: prior to v0.14.0-1 the variables sections were [mydumper_variables] and [myloader_variables], and from v0.14.0-1 they became [mydumper_session_variables], [mydumper_global_variables], [myloader_session_variables] and [myloader_global_variables]. A config file written for an older release will not silently keep working the way you expect after an upgrade, so pin your version and re-read the sample mydumper.cnf on each bump. The same file shows a newer option, source-control-command = AWS with aws-session-command, for Aurora and MySQL 5.7 restores where SET SESSION SQL_LOG_BIN = 0 is rejected; the README notes that when --source-control-command=AWS is set, --enable-binlog is ignored. On licensing, the project is GPL-3.0. If you only run the binaries as a backup step, that is one thing; if you redistribute MyDumper inside a product, the copyleft terms apply and you should have someone qualified review the distribution, which is not something this article can decide for you.
Editorial conclusion
Adopt MyDumper when your dump is dominated by InnoDB tables and you want per-table files plus parallel export and import; skip it if you depend on non-transactional engines, since the README states consistent snapshots for those are not provided. Before rolling it out, check the defaults-file section names against your installed version, because the README shows mydumper_variables changing to mydumper_session_variables and mydumper_global_variables from v0.14.0-1, and old names will not behave as expected. The project is GPL-3.0, so verify how that interacts with your distribution plans.
Frequently asked questions
How do I install MyDumper?
The README no longer lists install steps inline and instead links to the official documentation's installing page, which also covers compilation requirements for building from source. An official Docker image is published on Docker Hub, and the repository's docker/Dockerfile can build from local or GitHub sources.
How do I use MyDumper?
Run mydumper to export and myloader to import, both multithreaded. The README points to the usage page in the official documentation for the full option list, and shows a defaults file with mydumper and myloader sections as the recommended way to configure a run.
MyDumper vs mysqldump: what is the difference?
MyDumper uses multiple threads with a consistent snapshot per worker and writes separate files for tables and dump metadata, while mysqldump produces a single stream. The README frames MyDumper's advantages as parallelism, easier-to-manage output, consistency across threads and PCRE-based filtering of databases and tables.
What is the best way to back up a MySQL database?
There is no single answer, because the right tool depends on your engines and restore needs. MyDumper suits transactional schemas where parallel export and per-table files help, while the README states its snapshot approach does not cover non-transactional engines, so a mixed-engine schema needs a different plan for those tables.
How can I automatically back up my MySQL database?
The README does not describe a scheduler or built-in automation. What it does provide is a defaults file with mydumper and myloader sections, so a recurring job can call mydumper with the same configuration each time, and --exec to pipe finished files through an external command such as gzip.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/mydumper-mydumper)