Model or dataset
xming521/WeClone avatar
xming521/WeClone

WeClone: fine-tune a Qwen model on your chat history and bind it to a Telegram bot

🚀 One-stop solution for creating your AI twin from chat history 💡 Fine-tune LLMs with your chat logs to capture your unique style, then bind to a chatbot to bring your digital self to life.

18,229 stars1,528 forksPythonAGPL-3.0

At a glance

What is it?
WeClone is a Python pipeline that exports chat logs, filters private data, fine-tunes an LLM with LoRA through LLaMA Factory, and serves the result as a Telegram or WeChat bot. It is a GPU project, not a hosted service.
Who is it for?
Adopt WeClone if you have a Linux box with at least 16GB of VRAM, a Telegram export you own, and a reason to run the whole pipeline locally rather than send logs to an API. Skip it if you need WhatsApp or Discord as the data source today, since both are marked as work in progress, or if you only want a prompt-engineered persona, because the project's own note says 7B output is average and 14B or larger is where results improve.
Can I use it commercially?
Yes, with strict conditions. AGPL-3.0 is a network copyleft licence: if people use a modified version over a network, for example as a hosted service, you must offer them its source code under the same licence.
Is it still maintained?
Yes. The repository last received commits 13 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 17, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem WeClone solves, and the people it is built for

Most chat assistants are generic. If you want a model that answers the way a specific person writes, you have two options: prompt an existing API with examples of that person's messages, or change the model's weights with their data. WeClone takes the second route and packages the whole path: exporting chat history, cleaning it, running supervised fine-tuning, and then attaching the resulting checkpoint to a chat platform so it can reply in real time. The README describes it as a "one-stop solution for creating your digital avatar from chat history", and the feature list names chat data export, preprocessing, model training and deployment as the four stages it covers.

The intended user is someone who owns the chat logs and has the hardware to train on them. The project's hardware table makes that explicit: the default recipe is Qwen2.5-VL-7B-Instruct fine-tuned with LoRA, which the table puts at 16GB of VRAM, while QLoRA at 4-bit precision is listed at 6GB for the same 7B model. Anyone without a GPU is not the audience. Neither is someone who wants a hosted signup flow, because there is no hosted version in the repository, only a project homepage and a documentation site.

There is also a privacy argument built into the design. The README lists "privacy information filtering with localized fine-tuning and deployment" as a core feature, and the dependency list includes presidio_analyzer and presidio_anonymizer, which are Microsoft's PII detection and anonymization libraries. The point is that the logs never have to leave your machine.

How the pipeline moves from a Telegram export to a trained checkpoint

The data flow has four visible stages. First, chat logs come in from a supported source. The platform table marks Telegram as fully supported for text and images, with stickers converted to emoji, while WhatsApp, Discord and Slack are all marked with the work-in-progress symbol for every column. Second, the logs are preprocessed and cleaned, and this is where the PII filtering runs. Third, the cleaned dataset goes to LLaMA Factory, which is pinned as llamafactory==0.9.4 in the main dependency group. Fourth, the fine-tuned model is loaded and bound to a bot.

The training side is not a custom trainer. WeClone delegates to LLaMA Factory and inherits its model and method coverage, which the README acknowledges by linking to LLaMA Factory's supported-models list. That is a sensible division: LLaMA Factory handles LoRA, QLoRA, GaLore, APOLLO, BAdam and full fine-tuning, and WeClone handles the parts specific to chat data. The default is supervised fine-tuning with LoRA on Qwen2.5-VL-7B-Instruct.

The deployment side is more uneven than the data side. The deployment table marks Telegram, Discord and Slack as supported, WeChat personal accounts as supported through a component named openclaw-weixin, and WhatsApp as still in progress. So the same project can read from Telegram and write to Discord, but it cannot yet read from Discord.

One configuration detail worth noting: the project keeps training and inference settings in a single file, settings.jsonc, and pyproject.toml carries a [tool.weclone] section with a config_version field and a config_changelog. That changelog records structural changes such as "[0.3.03] - 2025-11-01 - Add chat member relationship switch." If you upgrade across versions, your settings file may need edits, not just a reinstall.

Installing WeClone and running a first training pass

The README assumes a Linux environment with CUDA 12.6 or above already installed, and says the Windows environment has not been rigorously tested, suggesting WSL as the runtime. The recommended Python environment manager is uv. This block clones the repository, creates a Python 3.12 virtual environment, activates it, and installs the main dependency group in editable mode.

bash
git clone https://github.com/xming521/WeClone.git && cd WeClone
uv venv .venv --python=3.12
source .venv/bin/activate # windows .venv\Scripts\activate
uv pip install --group main -e .

Note that the project requires Python >=3.12,<3.13, so a 3.11 or 3.13 interpreter will not satisfy the package metadata. After the install finishes, the next step is the configuration file. The README instructs you to copy a template and rename it to settings.jsonc, and all later edits happen in that file.

bash
cp examples/tg.template.jsonc settings.jsonc

The repository ships two templates, examples/tg.template.jsonc and examples/mllm.template.jsonc, so the choice of template signals which pipeline you intend to run. The README then says to verify that the CUDA environment is recognized by PyTorch before going further, though the exact command is cut off in the README. What you should expect after a successful install is a working weclone-cli entry point, declared under [project.scripts] in pyproject.toml, plus a settings.jsonc that the CLI reads for both training and inference.

Where WeClone breaks down: data volume, model size and platform gaps

The most honest limitation is stated by the project itself. The README warns that fine-tuning effectiveness depends largely on model size and on the quantity and quality of chat data, that the 7B model's performance is "average", and that 14B or more parameters tend to deliver better results. That is a direct contradiction of the assumption many newcomers bring, which is that a small export and a small model will reproduce a person's voice. They usually will not.

The hardware table turns that into a budget problem. LoRA at 16-bit precision on 7B is listed at 16GB of VRAM, and 14B at 32GB. QLoRA at 4-bit brings 7B down to 6GB and 14B to 12GB, but quantization is a trade-off, not a free lunch, and the project does not publish quality comparisons between the two. If your GPU has 8GB, you are choosing between QLoRA on 7B and not training at all.

Platform support is the second gap. WhatsApp, Discord and Slack appear in the data source table but every cell is marked as work in progress, so Telegram is the only fully documented source for text and images. On the deployment side WeChat personal accounts are supported through openclaw-weixin, which the README does not explain further; anyone relying on a personal WeChat account should treat that path as less documented than the Telegram one. Finally, the project describes itself as "still in rapid iteration phase", and the config_version field in pyproject.toml confirms that settings structures change between releases. Pinning a working version is safer than tracking the default branch.

How WeClone differs from prompt-based persona tools and from LLaMA Factory alone

The obvious alternative is not another fine-tuning framework. It is skipping training entirely and putting a persona into a system prompt with a few dozen example messages, which works with any hosted model API and needs no GPU. The difference in approach is real: a prompt-based persona keeps the base model's knowledge and reasoning intact and only nudges style, while WeClone's supervised fine-tuning changes the weights themselves, which can capture phrasing and rhythm that a prompt struggles to hold, at the cost of a training run and a risk of degrading general ability if the dataset is narrow.

The closer comparison is LLaMA Factory by itself. WeClone depends on it, pins it at 0.9.4, and does not reimplement training. What WeClone adds is everything around the trainer: the chat export path, the preprocessing that turns conversation logs into training pairs, the presidio-based PII filtering, the settings.jsonc configuration layer, and the bot bindings for Telegram, Discord, Slack and WeChat. If you already have a clean instruction dataset and only need to fine-tune, LLaMA Factory alone is the shorter path. WeClone earns its place when the raw input is a messy chat export and the output needs to be a live bot.

A third option is a managed fine-tuning service from a model provider. That removes the GPU requirement but sends the chat logs to someone else, which is exactly the property WeClone's localized deployment is designed to avoid. Whether that matters depends on whose messages are in the export, and the export usually contains the other person's messages too.

Editorial conclusion

Adopt WeClone if you have a Linux box with at least 16GB of VRAM, a Telegram export you own, and a reason to run the whole pipeline locally rather than send logs to an API. Skip it if you need WhatsApp or Discord as the data source today, since both are marked as work in progress, or if you only want a prompt-engineered persona, because the project's own note says 7B output is average and 14B or larger is where results improve. Before committing, check three things: that your CUDA install is 12.6 or above, that the Qwen2.5-VL-7B-Instruct weights and the LLaMA Factory stack fit your disk and VRAM, and that you can live with AGPL-3.0 if the bot is ever exposed to other people over a network.

Frequently asked questions

What hardware does WeClone need to fine-tune a model?

The README's table lists LoRA at 16-bit precision as needing 16GB of VRAM for a 7B model and 32GB for 14B, while QLoRA at 4-bit drops those to 6GB and 12GB respectively. The default recipe is Qwen2.5-VL-7B-Instruct with LoRA.

Which chat platforms can WeClone read history from?

Telegram is the only fully supported data source, with text and images marked as available and stickers converted to emoji. WhatsApp, Discord and Slack are listed in the same table but every column is marked as work in progress.

Can WeClone deploy a bot to WeChat?

The deployment table marks WeChat personal accounts as supported through a component named openclaw-weixin. The README does not document that path in detail, so treat it as less explained than the Telegram deployment.

How does WeClone handle private information in chat logs?

The README lists privacy information filtering as a core feature, and the main dependency group includes presidio_analyzer and presidio_anonymizer. Combined with localized fine-tuning and deployment, the logs do not have to leave your machine.

Official sources

  1. License: AGPL-3.0
  2. Project website
  3. README
  4. Releases
  5. xming521/WeClone on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/xming521-weclone.svg)](https://hysenlabs.com/projects/xming521-weclone)