TaskMatrix: a quick start that clones a different repository, and pins from 2022
TaskMatrix
At a glance
- What is it?
- TaskMatrix is the implementation behind the Visual ChatGPT paper, routing chat messages to visual foundation models so you can send and receive images while talking. Its value is the template mechanism and a published GPU memory table; its cost is a dependency set pinned to torch 1.13.1, langchain 0.0.101 and Python 3.8, and a last push on 2024-01-06.
- Who is it for?
- Use TaskMatrix if you want to read how a chat model was wired to a dozen vision models as a research artefact, or if you are reproducing the Visual ChatGPT paper and need the GPU memory table. Do not start a new product on it: the dependency pins are from 2022, the quick start clones microsoft/TaskMatrix and changes into a visual-chatgpt directory that is not in the repository tree, and the last push was 2024-01-06 with no GitHub releases.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Probably not. The repository last received commits 33 months ago, on January 6, 2024.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The first two commands in the quick start do not match this repository
The quick start is a shell block, and the first two lines do not describe the repository you are reading. It says to clone the repo, and the command is `git clone https://github.com/microsoft/TaskMatrix.git`, then to go to directory, and the command is `cd visual-chatgpt`. Neither matches what is here: this repository is chenfei-wu/TaskMatrix, and its root holds visual_chatgpt.py as a file, not a directory of that name.
The rest of the block is a conventional Python setup: create a conda environment named visgpt on Python 3.8, activate it, install from requirements.txt, then install GroundingDINO and segment-anything straight from their git URLs. The last two lines set the OpenAI key, with `export OPENAI_API_KEY=...` for Linux and the Windows `set` form underneath it.
So the mechanical steps work, but the first two will send you to a different repository or fail outright, and the discrepancy is a useful signal about the state of the project. Read the block as a description of how the code was run when the paper was published, then fix the clone URL and the directory yourself before you start.
The environment half of the block, which is the part that still applies, is:
conda create -n visgpt python=3.8
conda activate visgpt
pip install -r requirements.txt
pip install git+https://github.com/IDEA-Research/GroundingDINO.git
pip install git+https://github.com/facebookresearch/segment-anything.gitNote the two spaces after pip install on the GroundingDINO and segment-anything lines. They are in the original, and they are the kind of detail that tells you the block was pasted rather than typed.
torch 1.13.1, langchain 0.0.101 and Python 3.8 are the pins
requirements.txt is where the age of the project shows. Three entries are pinned to exact versions: langchain at 0.0.101, torch at 1.13.1 and torchvision at 0.14.1, plus wget at 3.2. Everything else floats: accelerate, addict, albumentations, basicsr, controlnet-aux, diffusers, einops, gradio, imageio, kornia, numpy, omegaconf, open_clip_torch, openai, opencv-python, safetensors, streamlit, timm, torchmetrics, transformers, webdataset and yapf.
That combination belongs to a specific moment. A langchain release numbered 0.0.101 and a torch build from before the 2.0 line are both from the 2022 to 2023 window, and the conda environment is pinned to Python 3.8. Meanwhile the floating packages have moved a long way, so a fresh install resolves a modern diffusers and transformers against a two-year-old torch, which is not a combination anybody tested.
The last push to the repository was 2024-01-06 and there are no GitHub releases, so nothing upstream is going to relax those pins for you. If you want to run this, budget for resolving the environment yourself, and treat any code that imports langchain as code written against an API that no longer exists in the version you would install today.
The memory table is a budget, and the detectors are free
The README publishes a table of GPU memory in megabytes per foundation model, and it is the most immediately useful table in the repository. The pattern is what matters. Every detector in the Image2 family costs nothing: Image2Canny, Image2Line, Image2Hed, Image2Scribble, Image2Pose, Image2Depth and Image2Normal are all listed at 0 MB. Every generator in the matching Text2Image family sits around 3.5 GB: CannyText2Image and ScribbleText2Image and DepthText2Image at 3531, LineText2Image, HedText2Image, PoseText2Image, SegText2Image and NormalText2Image at 3529.
The larger single entries are ImageEditing at 3981, Text2Image at 3385 and InstructPix2Pix at 2827. The smaller ones are ImageCaptioning at 1209, VisualQuestionAnswering at 1495 and Image2Seg at 919.
Because the models are loaded on demand through a --load list, the arithmetic is yours. The editing command in the README loads four of them at once, Text2Box, Segmenting, Inpainting and ImageCaptioning, which on these figures is a little over 7 GB before anything else is resident. A single 24 GB card handles a couple of generators comfortably; a 16 GB card is where you start choosing, and a smaller one is where the detector plus one generator pairing is the realistic plan.
GroundingDINO, then segment-anything, then inpainting, in that order
The most recent feature work described is the addition of GroundingDINO and segment-anything, contributed by jordddan, and the pipeline is spelled out step by step. For the image editing case, GroundingDINO is first used to locate bounding boxes guided by the given text, then segment-anything generates the related mask, and finally stable diffusion inpainting edits the image based on that mask.
Two commands demonstrate it. Start the app with `python visual_chatgpt.py --load "Text2Box_cuda:0,Segmenting_cuda:0,Inpainting_cuda:0,ImageCaptioning_cuda:0"`, then say find xxx in the image or segment xxx in the image, where xxx is an object, and the app returns the detection or the segmentation result.
Notice what one sentence costs. Text in, boxes out is one detector. Text in, a mask out is that detector plus a segmentation model. Text in, an edited image is both plus an inpainting model, with the captioning model loaded alongside. The :0 suffixes put each on the first GPU, which is a reminder that the multi-GPU layout is opt-in per model rather than something the framework discovers for you.
A template is a class with the attribute template_model = True
The template idea is the contribution that is not just a wrapper, and its definition is short. A template is a pre-defined execution flow that assists ChatGPT in assembling complex tasks involving multiple foundation models. A template contains the experiential solution to complex tasks as determined by humans. And a template can invoke multiple foundation models, or even establish a new ChatGPT session.
The implementation is a convention rather than a framework. To define a template, you add a class with the attribute template_model = True. There is no base class to inherit and no registration call, which makes adding one cheap and makes the detection depend on that attribute being spelled correctly.
The worked example is InfinityOutPainting, contributed by ShengmingYin and thebestannie. Start it with `python visual_chatgpt.py --load "Inpainting_cuda:0,ImageCaptioning_cuda:0,VisualQuestionAnswering_cuda:0"` and then say extend the image to 2048x1024. The claim is that this single template extends an image to any size by combining the existing ImageCaptioning, Inpainting and VisualQuestionAnswering models, without the need for additional training. That is the idea in one sentence: the hard part is the choreography, and the choreography is the thing you write.
TaskMatrix.AI and LowCodeLLM sit beside the single entry script
The repository tree is small, and it says more about the project's direction than the README does. At the root there is one Python entry point, visual_chatgpt.py, alongside requirements.txt, LICENSE.txt, SECURITY.md, CONTRIBUTING.md, CODE_OF_CONDUCT.md, a README and an assets directory. Two directories carry names that suggest separate efforts rather than part of the demo: TaskMatrix.AI/ and LowCodeLLM/.
The tree also has a .DS_Store file committed, which is a harmless artefact of someone building on a Mac and a fair indication of how the repository is maintained by hand rather than by a release process.
Elsewhere the project points outward. There is a Hugging Face Space and a Google Colab notebook linked at the top of the README, so you can see the system without installing anything, and the paper is Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models, arXiv 2303.04671. The acknowledgements name the projects the implementation leans on: Hugging Face, LangChain, Stable Diffusion, ControlNet, InstructPix2Pix, CLIPSeg and BLIP. That list is effectively the dependency map, and it is the fastest way to see which part of the stack came from where.
The models are examples, and their licences are yours to comply with
The last two sections of the README are the ones to read before you cite or reuse any of this. The disclaimer states that the recommended models in the repository are just examples, used for scientific research exploring the concept of task automation and benchmarking against the published paper, and that users can replace the models according to their research needs. When you use those models you need to comply with the licences of those models respectively, Microsoft shall not be held liable for infringement of third-party rights resulting from your use of the repository, and users agree to defend, indemnify and hold Microsoft harmless against claims arising from it.
The trademark notice follows the same shape. The project may contain trademarks or logos of other projects, authorised use of Microsoft trademarks is subject to Microsoft's own guidelines, and use of those marks in a modified version must not imply Microsoft sponsorship.
Practical consequence: the licence file and the disclaimer are not the same thing, and neither of them grants you rights to the model weights. Those come from their own projects, one by one. If you are doing anything commercial, that per-model review is the part that cannot be skipped.
Editorial conclusion
Use TaskMatrix if you want to read how a chat model was wired to a dozen vision models as a research artefact, or if you are reproducing the Visual ChatGPT paper and need the GPU memory table. Do not start a new product on it: the dependency pins are from 2022, the quick start clones microsoft/TaskMatrix and changes into a visual-chatgpt directory that is not in the repository tree, and the last push was 2024-01-06 with no GitHub releases. Before you run anything: fix the clone URL and directory, decide which models you actually need, because ImageEditing alone is listed at 3981 MB and the *Text2Image family near 3.5 GB each, set OPENAI_API_KEY yourself, and read the disclaimer, which says the bundled models are examples and that you must comply with each model's own licence.
Frequently asked questions
What is TaskMatrix, the Visual ChatGPT code?
TaskMatrix connects ChatGPT and a series of visual foundation models to enable sending and receiving images during a conversation. It is the implementation behind the paper Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models, arXiv 2303.04671, and the project describes the design as pairing a general interface with domain experts.
How do I install TaskMatrix?
The quick start creates a conda environment with `conda create -n visgpt python=3.8`, activates it, runs `pip install -r requirements.txt`, then installs GroundingDINO and segment-anything from their git URLs, and sets OPENAI_API_KEY with the export form on Linux or the set form on Windows. Note that the clone command in that block points at github.com/microsoft/TaskMatrix and changes into a visual-chatgpt directory, neither of which matches this repository's tree.
How much GPU memory does TaskMatrix need?
The README publishes a table in megabytes. ImageEditing is listed at 3981, Text2Image at 3385, InstructPix2Pix at 2827, VisualQuestionAnswering at 1495, ImageCaptioning at 1209 and Image2Seg at 919, while the whole Image2 detector family is 0 and the matching Text2Image generators sit at 3529 or 3531 each. You choose which models to load with the --load list.
How do I use a TaskMatrix template?
You add a class with the attribute template_model = True, which marks it as a pre-defined execution flow that ChatGPT can assemble to handle a complex task across several models. The InfinityOutPainting example loads Inpainting, ImageCaptioning and VisualQuestionAnswering, and asking it to extend the image to 2048x1024 uses those existing models with no additional training.
Is TaskMatrix still being worked on?
The repository is not archived, but the last push was 2024-01-06 and there are no GitHub releases. The requirements file pins langchain at 0.0.101, torch at 1.13.1 and torchvision at 0.14.1, and the documented environment is Python 3.8, so the dependency set is from an earlier era than those dates. The README asks for community contributions and says issues are the way to get help.