Model or dataset
hasaneyldrm/exercises-dataset avatar
hasaneyldrm/exercises-dataset

hasaneyldrm/exercises-dataset: a 1,324-exercise dataset with GIFs and 10-language instructions

1,324-exercise fitness dataset — animation GIFs, 180×180 thumbnails, muscle-group & equipment data, and step-by-step instructions in 6 languages. The exercise data layer behind the LogPress app.

22,403 stars2,859 forksHTMLNOASSERTION

At a glance

What is it?
The exercise data layer behind the LogPress app ships 1,324 records with animation GIFs, 180x180 thumbnails, muscle-group metadata and step-by-step instructions in ten languages. The data and code are MIT; the media is not.
Who is it for?
Adopt it if you are building a workout planner, a form-reference screen or a training-data pipeline and you can live with the Gym visual media terms. Do not adopt it if you need a maintained schema, a hosted API, or media you can redistribute without conditions.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 76 days ago.
What is it written in?
Mainly HTML, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What hasaneyldrm/exercises-dataset actually contains

Most exercise data on GitHub is a scraped CSV with a name column and nothing else. This repository is the opposite: a fixed snapshot of 1,324 exercises, each with a category, a body part, an equipment value, a target muscle, a list of synergist muscles, an animation GIF, a 180x180 thumbnail, and step-by-step instructions translated into ten languages (English, Spanish, Italian, Turkish, Russian, Chinese, Hindi, Polish, Korean and French). The README describes it as the exercise data layer behind LogPress, an AI-assisted workout tracker, and invites you to drop it into your own backend.

The intended audience is narrow and clear. You are building a fitness app, a workout planner, a form-reference screen, or a machine-learning pipeline over exercise metadata, and you would rather not license a commercial exercise API or hand-write 1,324 instruction sets. The repository also ships two browser tools: index.html, a client-side explorer with live search and filters, and setup.html, a developer guide that generates SQL and API client code. Neither needs a server.

The data model: one JSON array, one schema file, one media folder per type

The repository layout is flat and predictable. data/exercises.json holds a single JSON array of 1,324 objects. data/exercises.schema.json is a JSON Schema draft 2020-12 document describing every field, its type and its constraints. images/ holds 1,324 thumbnails at 180x180. videos/ holds 1,324 animation GIFs, also at 180x180. index.html, setup.html, README.md, NOTICE.md and LICENSE sit at the top level.

Inside each record, image and gif_url point at the local assets, media_id holds the original media reference id, and an attribution field carries the credit line for that record. The exercise id is a zero-padded numeric string such as "0001", which means sorting by id as a string and sorting as a number give the same order. That is a small thing, but it removes a class of bugs in pagination code.

Because the schema is a separate file rather than prose in the README, you can validate the dataset or your own additions with any standard JSON Schema validator before importing. That is the strongest engineering decision in the repository. It also means field names are a contract: if you rename target to target_muscle in your own database, you own that mapping.

Getting the dataset and running a first query

There is no package to install. The README's own instruction is that the HTML tools need no server, just a browser. The repository is a plain git repository, so the standard clone applies, and the file structure in the README shows what you get:

code
exercises-dataset/
├── data/
│   ├── exercises.json        # Full dataset — 1,324 exercise records (JSON array)
│   └── exercises.schema.json # JSON Schema (2020-12) describing every record
├── images/                  # 1,324 × 180×180 thumbnails  (© Gym visual)
├── videos/                  # 1,324 × 180×180 animation GIFs  (© Gym visual)
├── index.html               # Interactive exercise browser (client-side, no server needed)
├── setup.html               # Developer setup guide (DB import + API integration)
├── NOTICE.md                # Media attribution & license terms
└── README.md

After cloning, open index.html in a browser. You should see a grid of exercise cards with thumbnails, a live search box, and filters for category, equipment and target muscle. Clicking a card shows the full instructions in the ten languages the dataset carries.

The data file is the part most people want. It is a JSON array, and the README names the fields each record carries. The README's overview table lists them as follows:

code
| Field | Description |
|---|---|
| Unique ID | Numeric identifier (e.g. `"0001"`) |
| Name | Full descriptive exercise name |
| Category | Primary muscle group targeted |
| Target | Specific target muscle |
| Muscle Group | Supporting / synergist muscles |
| Equipment | Equipment required (or `body weight` for bodyweight) |
| Instructions | Step-by-step instructions for each exercise |
| Media | 180×180 thumbnail (`image`) + animation GIF (`gif_url`) per exercise |

The README also states that image and gif_url point at the local 180x180 assets, that each record carries an attribution field, and that media_id holds the original media reference id. Read one record from data/exercises.json to see the real shape, and use data/exercises.schema.json to check it.

If you want the data in a database instead, open setup.html. It generates CREATE TABLE SQL for SQL Server, PostgreSQL, MySQL and SQLite, and can build a ready-to-run .sql file containing all 1,324 INSERT statements entirely in your browser. The same page prints copy-paste API client code in JavaScript, Python, C#, Java, PHP, Go and cURL, where entering your base URL updates every example live.

The media licence is the constraint that shapes every deployment

The repository's licence badge reads MIT plus media terms, and that split is the single most important operational fact about it. The code and the data are MIT. The animation GIFs and thumbnails are credited to Gym visual and used with permission, and NOTICE.md carries the attribution and the media terms.

What this means in practice is that you cannot treat the repository as one uniform blob. A backend that serves the JSON is a different legal object from an app that re-hosts 1,324 GIFs on a CDN. The README points at NOTICE.md and the License section for the details, and it does not summarise them. Read the actual file rather than the badge. If your product needs to redistribute the media, or bundle it into an offline download, that is the question to resolve before you build anything on top of it. Hysen Labs does not give legal advice; the point is that the answer is not in the README.

There is a second-order effect. Because every record already carries an attribution field, the dataset has been designed with this constraint in mind. Whatever your UI does with that field, the data gives you what you need to display credit per exercise.

Where the dataset stops being the right tool

The dataset is a static snapshot. There are no releases in the repository, no versioned tags, and no changelog. The last push was on 2026-07-16. If you need a schema that evolves with a deprecation policy, or a hosted endpoint with an SLA, this is not that. You are adopting a file, and you own every future migration yourself.

The second limitation is scope. The README lists categories, body parts, equipment, target muscles and synergist muscles. That is enough to filter and to plan, but it is not a training-load model. There is no field for difficulty, tempo, range of motion, contraindications, or progression. If your product needs to tell a user how much weight to add next week, the dataset gives you the exercise vocabulary and nothing about the prescription.

The third is the media size. 1,324 GIFs at 180x180 is a lot of bytes for a mobile bundle, and the README does not state the total repository size. The animations are also fixed at 180x180, so a full-screen demonstration view will upscale them. For a list or a card, 180x180 is fine. For a hero view, it is not.

Finally, the instructions are translated into ten languages, but the README does not describe how those translations were produced or reviewed. Treat the non-English text as unverified until you have a native speaker check the strings you ship.

Alternatives, and what the difference really is

The obvious alternative is the wger exercise database, an open source workout and nutrition tracker that also publishes an exercise dataset. The difference in approach is architectural. wger is an application first: the exercise data lives inside a running Django service with a REST API, user accounts and its own database, and you consume it by calling that service or deploying it. This repository is a dataset first: a JSON file, a schema file and a media folder, with no runtime at all. If you want to own your data and query it locally, the file wins. If you want an API you do not have to build, the service wins.

The second alternative is a commercial exercise API, where you get hosted media, versioning and support, and you pay per call or per month. The trade is control and cost against maintenance. This repository sits between the two: more work than a hosted API, far less work than assembling the data yourself, and no vendor in the request path.

The third option, which people often reach for first, is asking a language model to generate an exercise list. That produces plausible names with no stable ids, no media, and no schema. The README's setup.html leans into LLM assistance, but for a different job: generating the REST API around the dataset, not the dataset itself. That distinction is worth keeping straight.

Maintenance and upgrade cost

The repository is not archived, and the last push was on 2026-07-16. There are no releases, so there is nothing to pin to and nothing to read release notes for. Upgrading means pulling the branch and diffing data/exercises.json against your copy.

That diff is the real cost. Because ids are stable zero-padded strings, you can match records across versions, but nothing in the repository tells you which records were added, removed or edited between two points in time. If you import the dataset into your own database, keep the upstream id as a primary key and keep your own additions in a separate table. That way a re-import is a merge rather than a rebuild.

The schema file is your upgrade safety net. Run data/exercises.schema.json against any new copy before importing it, and a field rename or type change surfaces as a validation error instead of a null in production. On the licence side, the MIT data is easy to carry forward; the media terms are not, so re-read NOTICE.md whenever you pull a new copy.

Editorial conclusion

Adopt it if you are building a workout planner, a form-reference screen or a training-data pipeline and you can live with the Gym visual media terms. Do not adopt it if you need a maintained schema, a hosted API, or media you can redistribute without conditions. Before writing any code, read NOTICE.md and the media section of LICENSE, and open data/exercises.schema.json to confirm the fields your app will query.

Frequently asked questions

Is hasaneyldrm/exercises-dataset free to use?

The code and data are MIT, but the animation GIFs and thumbnails are credited to Gym visual and governed by separate media terms described in NOTICE.md. The README points to the License section and NOTICE.md rather than summarising them, so read both before shipping.

How many exercises are in hasaneyldrm/exercises-dataset?

The dataset contains 1,324 exercises. Each record has an animation GIF, a 180x180 thumbnail, category, body part, equipment, target and muscle-group data, and step-by-step instructions in ten languages.

How do I open the exercise browser in hasaneyldrm/exercises-dataset?

Open index.html directly in any modern browser. The README states that no server is required, and the page gives live search across all 1,324 exercises plus filters by category, equipment and target muscle.

Does hasaneyldrm/exercises-dataset provide an API?

No API ships with the repository. setup.html instead generates CREATE TABLE SQL for SQL Server, PostgreSQL, MySQL and SQLite, prints API client code in several languages, and offers a prompt for generating your own REST API with an LLM.

How do I validate the data in hasaneyldrm/exercises-dataset?

Use data/exercises.schema.json, a JSON Schema draft 2020-12 document that describes every field, its type and constraints. The README suggests using it to validate the dataset or your own additions with any standard JSON Schema validator.

Official sources

  1. hasaneyldrm/exercises-dataset on GitHub
  2. Issues
  3. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/hasaneyldrm-exercises-dataset.svg)](https://hysenlabs.com/projects/hasaneyldrm-exercises-dataset)