Model or dataset
SkyworkAI/Skywork-R1V avatar
SkyworkAI/Skywork-R1V

Skywork-R1V: running the 38B multimodal reasoning model locally

Skywork-R1V is an advanced multimodal AI model series developed by Skywork AI, specializing in vision-language reasoning.

3,172 stars284 forksPythonMIT

At a glance

What is it?
Skywork-R1V is an MIT-licensed vision-language model series from Skywork AI, with R1V3-38B as the current release. The repository ships inference scripts and an evaluation harness, but the README leaves a lot about deployment to inference.
Who is it for?
Adopt Skywork-R1V if you need an open-weight vision-language model you can run on your own GPUs under MIT terms, and you are comfortable reading inference_with_transformers.py and setup.sh rather than following a step-by-step guide. Skip it if you need a hosted API, a documented serving stack, or a model small enough for a single consumer card.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 64 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Skywork-R1V is for, and who it is actually aimed at

Skywork-R1V is a vision-language model series built for multimodal reasoning: questions that require reading an image and then working through math, physics, logic or chart interpretation rather than just describing what is visible. The README frames the series around visual chain-of-thought, and the R1V3 release notes describe reinforcement learning in post-training as the main mechanism behind the reasoning gains. The audience is not application developers looking for an endpoint. It is research teams and ML engineers who want open weights they can run, fine-tune or evaluate on their own hardware, and who are willing to work from scripts instead of a managed service. The repository bundles model weights references, inference code, and an eval directory that the README says reproduces the published benchmark results. That combination, weights plus a reproduction path, is the actual product. The published comparison table puts R1V3-38B at 76.0 on MMMU and 77.1 on MathVista (mini), ahead of the listed QVQ-72B and InternVL-78B figures on those two rows, while trailing Claude 3.7 on EMMA and LogicVista. Those numbers come from the project's own evaluation, and the asterisk footnote marks results produced inside its framework, so treat cross-model rows as indicative rather than neutral.

The R1V3-38B architecture and where the reasoning comes from

The README states that Skywork-R1V3 uses InternVL3-38B as its base model, which is MIT-licensed. That matters for two reasons: the parameter count and memory footprint are inherited from a 38B vision-language backbone, and the licence chain stays permissive end to end. The reasoning behaviour is not a new architecture. According to the release notes, it comes mainly from reinforcement learning applied during post-training, which the repository topics label as GRPO. The practical consequence is that the model emits a chain of thought before an answer, and that chain is part of the token budget you pay for at inference time. The repository layout reflects this split between running and measuring: inference/ holds the two inference entry points, eval/ holds the vlmevalkit-based evaluation environment, and report/ and the PDF files at the top level hold the technical write-ups. There is also an r1v4/ directory and a Skywork_R1V4.pdf at the top level, which the README does not describe. If you are trying to understand what R1V4 changes, the README will not tell you; the PDF and that directory are the only starting points the repository offers.

Installing Skywork-R1V and running a first question

The README gives a local setup path built on conda. Clone the repository, then move into the inference directory, because the setup script is invoked from there.

bash
git clone https://github.com/SkyworkAI/Skywork-R1V.git
cd skywork-r1v/inference

The README then splits the environment in two. One conda environment is for the Transformers path, created at Python 3.10 and provisioned by setup.sh. A second, separate environment is for vLLM and evaluation, provisioned by the build script under eval/vlmevalkit. Keeping them apart avoids dependency conflicts between the two inference stacks.

bash
conda create -n r1-v python=3.10 && conda activate r1-v
bash setup.sh

For the vLLM path, the README gives a second environment and a different build script. Note that the two commands are presented as alternatives, not as sequential steps.

bash
conda create -n r1v-vllm python=3.10 && conda activate r1v-vllm
bash ./eval/vlmevalkit/build_env.sh

With the Transformers environment active, inference takes a model path, one or more image paths, and a question. The README's example pins two GPUs through CUDA_VISIBLE_DEVICES and passes a single image.

bash
CUDA_VISIBLE_DEVICES="0,1" python inference_with_transformers.py \
    --model_path path \
    --image_paths image1_path \
    --question "your question"

The vLLM variant accepts multiple images in one call and exposes tensor parallelism. The README's example uses a tensor parallel size of four, which implies four GPUs for that configuration.

bash
python inference_with_vllm.py \
    --model_path path \
    --image_paths image1_path image2_path \
    --question "your question" \
    --tensor_parallel_size 4

What you should see is a generated answer preceded by the model's reasoning trace. The README does not document the output format, the sampling defaults, or how to suppress the reasoning tokens, so expect to read the script to find those. The --model_path value is left as a literal placeholder in the README; you supply either a local checkpoint directory or a Hugging Face model identifier, and the README does not spell out which forms are accepted.

The deployment gap: what the README does not cover

The most honest limitation is documentation scope. The README covers cloning, two conda environments, and two inference scripts. It does not cover serving. There is no mention of an OpenAI-compatible endpoint, no Dockerfile in the top-level entries, no systemd unit, no batching guidance, and no memory table telling you which GPU configurations fit the 38B weights at which precision. For a model of this size that last omission is the one that will cost you time. The README does point to an AWQ quantized variant of the earlier R1V2 release, described as supporting single-card inference above 30GB, but that note is attached to R1V2, not to R1V3, and the README does not state that an equivalent quantized R1V3 build exists. If your hardware budget is one 24GB card, nothing in the documentation confirms this model will run for you. A second gap is versioning. The repository contains an r1v4 directory and a Skywork_R1V4.pdf, but the README's news section ends at the R1V3 release and never explains the relationship between the two. Anyone arriving from a search for the newer version will find the README silent. Finally, the README does not document rollback between releases, checkpoint compatibility, or how the eval harness handles the reasoning traces it produces.

How Skywork-R1V compares with Qwen2.5-VL and InternVL

The closest alternatives are the general-purpose open vision-language models, and the difference is in training objective rather than architecture. Qwen2.5-VL and InternVL are built for broad multimodal understanding: captioning, OCR, grounding, document parsing, visual question answering across many domains. Skywork-R1V starts from InternVL3-38B and then applies reinforcement learning specifically to sharpen multi-step reasoning, which is why the project's own table shows it leading on MathVista, MathVerse and PhyX while trailing on MMBench-en-1.1 and MMstar against InternVL-78B and Qwen-72B. If your workload is extracting fields from invoices or describing product photos, the reasoning training buys you nothing and the chain-of-thought output costs you latency and tokens. If your workload is a physics problem rendered as a diagram, or a chart where the answer requires several arithmetic steps, the reasoning trace is the point. There is also a size difference worth naming: the comparisons in the README table include 72B and 78B models, so R1V3-38B is competing against systems roughly twice its parameter count. That is a real argument for it on hardware grounds, but it also means your evaluation should be on your own task, not on the aggregate benchmark rows.

Licence position and what maintenance looks like

The repository is MIT-licensed, and the README states plainly that commercial use, modification and distribution are permitted with no liability. It also notes that the InternVL3-38B base model is MIT-licensed, so the chain from base weights to this repository does not introduce a copyleft obligation. The README does not, however, address the licence terms of the Hugging Face weight releases themselves, which are distributed separately from this code repository. If you plan to ship a product, check the model card on the weight release rather than assuming the repository licence covers it. This is not legal advice. On maintenance, the last push to the repository was on 2026-07-29, which is recent enough that the codebase is not stale, but the README's news section has not been updated past the R1V3 announcement despite the presence of an r1v4 directory. The upgrade cost is the thing to plan for: each release in this series has shipped as a separate weight set with its own environment, and the README gives no migration path between R1V, R1V2, R1V3 and whatever R1V4 turns out to be. Budget for re-running setup.sh and re-validating your prompts against each new checkpoint.

Reproducing the benchmark table with the eval harness

The README states that the code in the eval directory reproduces the published results, and the environment for it is the vLLM one built by eval/vlmevalkit/build_env.sh. That is the only reproduction path documented, and it is worth understanding what it does and does not establish. The harness is built on vlmevalkit, a third-party evaluation framework, so the numbers you get back depend on that framework's prompt templates and answer extraction as much as on the model. Several rows in the README's table carry an asterisk marking them as results from the project's own evaluation framework, which is a reasonable disclosure but also a reminder that the comparison is not apples to apples across every row. If you want to make a decision about this model, the useful move is to run the harness on a benchmark whose distribution resembles your own task, not to re-run MMMU. The README does not document how to add a custom benchmark to the harness, so expect to work from vlmevalkit's own documentation.

Editorial conclusion

Adopt Skywork-R1V if you need an open-weight vision-language model you can run on your own GPUs under MIT terms, and you are comfortable reading inference_with_transformers.py and setup.sh rather than following a step-by-step guide. Skip it if you need a hosted API, a documented serving stack, or a model small enough for a single consumer card. Before committing, verify three things: that setup.sh resolves cleanly on your CUDA and PyTorch versions, that your GPU memory fits the 38B weights at the precision you need, and that the eval directory's vlmevalkit harness reproduces a benchmark you care about on your own data.

Frequently asked questions

How do I install Skywork-R1V and run it locally?

Clone the repository, move into the inference directory, then create a Python 3.10 conda environment and run setup.sh for the Transformers path, or build the vLLM environment with bash ./eval/vlmevalkit/build_env.sh. Inference then runs through inference_with_transformers.py or inference_with_vllm.py with a model path, image paths and a question.

What is the difference between Skywork-R1V3 and Skywork-R1V2?

The README describes R1V2 as an open-source multimodal reasoning model evaluated on MMMU, MMMU-Pro, MathVista and OlympiadBench, and R1V3 as the later release that applies reinforcement learning in post-training and reports 76.0 on MMMU. R1V2 also has an AWQ quantized variant documented for single-card inference above 30GB, and the README does not state that an equivalent R1V3 quantized build exists.

Can Skywork-R1V be used commercially?

The README states the code repository is MIT-licensed, with commercial use, modification and distribution permitted and no liability, and it notes that the InternVL3-38B base model is also MIT-licensed. The README does not describe the licence terms of the separately distributed Hugging Face weight releases, so check those model cards before shipping.

Official sources

  1. Issues
  2. License: MIT
  3. Project website
  4. README
  5. SkyworkAI/Skywork-R1V on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/skyworkai-skywork-r1v.svg)](https://hysenlabs.com/projects/skyworkai-skywork-r1v)