Open-source project
jeremyipark/vision-demos avatar
jeremyipark/vision-demos

vision-demos: Four Runnable Computer Vision Demos in One Python Repo

Fun real-world computer vision demos!

747 stars115 forksPythonApache-2.0

At a glance

What is it?
Jeremy Park's vision-demos collects four self-contained Python projects (dance sync scoring, chin-up rep counting, bouldering hold sequencing, running cadence) that all call hosted models through a single VLM Run API key. The repo is small, the scope is narrow, and the API dependency is the whole story.
Who is it for?
Adopt vision-demos if you want a working reference for pose and segmentation pipelines on real footage and you are willing to obtain a VLM Run API key from https://app.vlm.run. Skip it if you need offline inference, a packaged library, or something with a documented upgrade path: there are no releases, and the README does not describe version pinning, rollback, or what happens when a hosted model changes.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 3 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 11, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What vision-demos actually is, and who it is for

This is not a library and not a framework. It is a repository of four separate demo applications, each in its own directory, each with its own README, and each aimed at a different physical activity: dance_sync compares two dancers performing the same choreography and returns a similarity metric; chin_ups counts repetitions from a clip and times the ascent and descent of each one; rock_climbing segments bouldering holds with sam3.1, reports which holds a climber used and in what order, and compares attempts on the same route; running measures cadence, times each foot strike, and averages knee shape at contact.

The audience is narrow and worth stating plainly. These demos suit an engineer who wants to see how pose estimation and segmentation get wired into a scoring or counting loop on real video, and who is comfortable reading per-project READMEs rather than a unified API reference. They do not suit someone looking for a drop-in package to import into a production service. Every project leans on hosted models, and three of the four use the same pose model, usyd-community-vitpose-plus-large, which tells you the interesting variation lives in the post-processing logic rather than the model choice.

One API key at the repo root, four projects that find it

The architecture is deliberately flat. There is no shared Python package, no common utility module documented at the top level, and no build system. What the projects share is credentials and a model gateway.

The .env.example file explains the mechanism: copy it to .env once at the repository root, and the loader walks up from each project directory to the nearest .env, so a single file serves all four demos. An exported VLMRUN_API_KEY takes precedence over the file. That is the entire configuration surface as documented.

Model access goes through the VLM Run gateway. The README links each model to a docs page: usyd-community-vitpose-plus-large for pose work in dance_sync, chin_ups, and running, and facebook-sam3.1 alongside the pose model for rock_climbing's hold segmentation. The data flow implied by the project descriptions is consistent: video in, model inference out, then project-specific logic that turns keypoints or masks into a number a human cares about, a sync score, a rep count, a hold sequence, a cadence. The repository layout shows no local model weights, no download scripts, and no inference server, which is consistent with the hosted-gateway approach.

Install and a first run

The top-level README does not give install steps. It states that each project has its own README with setup and instructions, so the commands below are the ones the repository files themselves document rather than a guessed sequence.

The only setup step visible at the root is credential configuration. The .env.example comments give the exact command and tell you where to obtain a key:

bash
cp .env.example .env

Then edit .env and set the value. The file ships with an empty assignment:

bash
VLMRUN_API_KEY=

The comment in .env.example points to https://app.vlm.run, Settings, then API keys. It also notes that .env is gitignored and that an exported VLMRUN_API_KEY overrides the file, which matters if you run these demos in CI or inside a container where writing a file is awkward.

After that, pick a project directory and follow its README. The top-level README does not list Python version requirements, dependency files, or entry-point scripts, so I cannot tell you what a first run looks like beyond the credential step. Read chin_ups/README.md or running/README.md before assuming a command exists.

The dependency on a hosted gateway is the main limitation

Three of the four demos route through the same hosted pose model, and rock_climbing adds a hosted segmentation model on top. That design keeps the repository small and the demos reproducible across machines, because nobody has to download weights or own a GPU. It also means the demos stop working without network access and without a valid key, and the top-level README does not describe an offline path, a local fallback, or a cached-inference mode.

The second limitation is subtler. When inference is hosted, the model behind a name can change without any commit in this repository. The README links model documentation but does not state version pinning, deprecation policy, or what a breaking change would look like. If you were considering adapting one of these pipelines for something you depend on, that gap is the thing to resolve first, not the Python code.

Finally, scope. These are demos of specific activities. If your problem is general object detection, OCR, or video summarization, nothing here addresses it. The projects are also not a benchmark suite: there are no accuracy figures, no comparison against ground truth, and no evaluation harness documented at the top level.

Compared with MediaPipe and other local pose stacks

The obvious alternative for this kind of work is a local pose estimation stack such as MediaPipe, or a self-hosted model you run yourself. The difference is not accuracy claims, which this repository does not make. The difference is where the compute lives.

A local stack means you install a runtime, manage model files, and accept whatever hardware you have. Inference is free at the margin, works offline, and cannot be changed out from under you by a vendor. The trade-off is setup cost and the fact that you own the model lifecycle.

vision-demos inverts that. Setup is one environment variable, and the heavy lifting happens on VLM Run's side. You get a working pipeline in minutes and you give up control over model versions, availability, and cost per call. Neither approach is better in the abstract. If you are prototyping an idea over a weekend, the hosted route removes a day of environment work. If you are building something that has to run in a gym with unreliable wifi, the hosted route is the wrong shape entirely.

Licence, maintenance, and what upgrading would cost

The repository is licensed Apache-2.0, with the licence text in LICENSE at the root. Apache-2.0 is permissive and includes an explicit patent grant, which is a meaningful difference from MIT for anything you might ship commercially. It does not, however, grant you rights to the hosted models, and the README says nothing about the terms under which VLM Run serves them. That distinction is worth checking before you treat these demos as a template for a product.

On maintenance: the last push was on 2026-09-16, one day before this writing, so the repository is being touched. There are no retrieved releases, which means there is no tagged version to pin to and no changelog to read. Upgrading therefore means pulling main and re-reading per-project READMEs, because the root README does not track changes. If a project breaks after a pull, the repository offers no documented rollback path beyond your own git history.

The practical upgrade cost is low while the demos stay small and high if you fork them into something load-bearing. Once you depend on a specific output format, a model change on the gateway side becomes your problem, and nothing in this repository is designed to insulate you from that.

Editorial conclusion

Adopt vision-demos if you want a working reference for pose and segmentation pipelines on real footage and you are willing to obtain a VLM Run API key from https://app.vlm.run. Skip it if you need offline inference, a packaged library, or something with a documented upgrade path: there are no releases, and the README does not describe version pinning, rollback, or what happens when a hosted model changes. Before committing, clone the repo, copy .env.example to .env, fill in VLMRUN_API_KEY, and run one project end to end on your own clip to see whether the output format matches what your pipeline expects.

Frequently asked questions

Are Vision Pro Demos free?

This repository is not related to Apple Vision Pro. vision-demos is a Python repository of four computer vision demos licensed under Apache-2.0, and its model inference runs through the VLM Run gateway, which requires an API key obtained from https://app.vlm.run.

What is an example of a demo in vision-demos?

The repository lists four. chin_ups counts chin-up reps from a clip and times the ascent and descent of each one, while rock_climbing segments bouldering holds, returns which holds the climber used and in what order, and compares attempts at the same route.

What is the purpose of vision-demos?

The README describes it as real-world computer vision demos. Each of the four projects applies pose estimation, and in one case segmentation, to a specific physical activity and turns the model output into a metric such as a sync score, a rep count, a hold sequence, or a cadence.

Are the demos in vision-demos usually free to run?

The repository itself is Apache-2.0, but the demos depend on hosted models through the VLM Run gateway and require a VLMRUN_API_KEY. The README does not state pricing for that gateway, so the cost of running the demos is not documented here.

Official sources

  1. Issues
  2. jeremyipark/vision-demos on GitHub
  3. License: Apache-2.0
  4. README
Community

Where developers are discussing it

Posts on Hacker News, dev.to and Lobsters that link to this repository, found by our scan of those communities. 1 in the last 30 days; newest first.

Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/jeremyipark-vision-demos.svg)](https://hysenlabs.com/projects/jeremyipark-vision-demos)