Open-source project
google-ai-edge/mediapipe avatar
google-ai-edge/mediapipe

MediaPipe: Google's On-Device ML Framework for Streaming Media

Cross-platform, customizable ML solutions for live and streaming media.

37,118 stars6,171 forksC++Apache-2.0

At a glance

What is it?
MediaPipe is Google's open-source library for running machine learning inference directly on mobile, web, and desktop hardware, without routing data through a server. It provides pre-trained models and cross-platform APIs for vision, text, and audio tasks alongside a low-level C++ graph framework for custom pipelines.
Who is it for?
Developers building real-time gesture recognition, face landmarking, or pose detection for mobile or web applications will find MediaPipe Tasks a practical starting point: the pre-trained models remove the need for a training infrastructure and the on-device processing keeps user data local. The privacy trade-off is the metrics reporting built into the Tasks APIs, which requires informing users under applicable law.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly C++, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

What MediaPipe Solves: Real-Time Inference Without a Server

Most ML workflows send data to a server for inference and return results over a network. MediaPipe inverts that model: the README describes its goal as delivering ML solutions for live and streaming media that run entirely on the device. The target users are Android and iOS developers, web engineers building with JavaScript or TypeScript, desktop application authors, and engineers targeting edge or IoT hardware.

The practical benefits are latency and privacy. Inference on a local camera frame avoids a round-trip over the network, which matters for applications like sign-language translation or prosthesis control where the documentation cites real deployments. The privacy section of the README is explicit: when using the MediaPipe Tasks APIs, input data such as images and video is processed on device and is not sent to Google servers. The trade-off is that performance metrics from the APIs are sent to Google, and developers are responsible for obtaining user consent as required by applicable law.

The Architecture: Tasks, Models, Model Maker, and the C++ Framework

MediaPipe is structured in two distinct layers. The higher layer, called MediaPipe Tasks, provides cross-platform APIs targeting three domains: vision (object detection, pose estimation, face landmarking, hand tracking, image segmentation), text (text classification, language detection), and audio (audio classification). Each Task pairs a fixed API with a pre-trained model file that ships with the library.

The lower layer is MediaPipe Framework, the C++ component used to build the inference pipelines that power the Tasks. The framework organizes computation into three core concepts: Packets (the unit of data flowing through a pipeline, timestamped and typed), Graphs (directed acyclic graphs of connected Calculators declared in a protobuf config), and Calculators (individual processing nodes that consume and produce Packets). Building a custom Calculator requires familiarity with C++ and the Bazel build system; the repository root contains .bazelrc and MODULE.bazel files that configure the build.

Two additional tools round out the ecosystem. MediaPipe Model Maker allows developers to retrain or fine-tune models on their own data without writing training infrastructure from scratch. MediaPipe Studio is a browser-based tool for visualizing inference results, evaluating accuracy, and benchmarking a solution before shipping it.

Platforms and Language Bindings

The repository supports a broader surface area than typical ML libraries. The top-level files include Package.swift for iOS Swift integration, a pnpm-lock.yaml and tsconfig.json for the web TypeScript build, a setup.py for the Python package, build_android_examples.sh and build_ios_examples.sh for mobile builds, and a Dockerfile for Linux environments. The package.json shows the project name as medipipe-dev with TypeScript 5.3.3 and rollup in the dev dependencies, confirming the web path is a first-class build target.

The Python requirements are listed in requirements.txt: absl-py, certifi, numpy, sounddevice, flatbuffers, opencv-contrib-python, and matplotlib. These are the build-time dependencies; the actual runtime requirements for a given Task depend on which platform SDK you install.

Setting Up MediaPipe in Python

The README redirects primary setup documentation to developers.google.com/mediapipe, and the setup guides are separated by platform: Android, web apps, and Python each have their own page. For a Python developer, the README points to the Python setup guide at that external site.

Building from the repository source requires preparing the environment first. The Dockerfile provided in the repository shows the prerequisite Python packages:

bash
pip3 install --upgrade setuptools
pip3 install wheel

From there the repository's setup.py drives the MediaPipe Python package build. The Dockerfile also installs Clang 16 as the C++ compiler, which the build system requires on Linux. For developers who want pre-built binaries rather than a source build, the README links to the developer guides where binary installs are documented per platform.

The 2023 Legacy Solution Deprecation and Its Consequences

MediaPipe's history includes a significant architectural shift. The README announces that support ended for the MediaPipe Legacy Solutions on March 1, 2023. The legacy set covered solutions such as the older face detection, hand tracking, and holistic pipelines that were built directly on the C++ Framework without the Tasks abstraction layer. The README notes that all other legacy solutions will be upgraded to the new Tasks-based approach.

The consequence for existing users is that the legacy code and prebuilt binaries remain in the repository and on an as-is basis, meaning they are not patched for new OS versions, updated for API changes, or given new model weights. Projects that integrated the older solutions before 2023 and have not migrated to MediaPipe Tasks carry that technical debt. The new Tasks API is not a drop-in replacement; it has a different calling convention and a different model format.

MediaPipe versus OpenCV: Different Levels of Abstraction

OpenCV is a general computer vision library that provides image filters, feature detectors, video capture, and linear algebra utilities. It does not include trained ML models or a pipeline orchestration layer. MediaPipe, by contrast, bundles pre-trained models and a graph runtime on top of image handling.

The two are not competitors on the same level. The requirements.txt lists opencv-contrib-python as a direct dependency of MediaPipe, so MediaPipe builds on OpenCV's image processing primitives rather than replacing them. A developer who needs a custom convolution filter, a calibration routine, or a stereo depth algorithm should reach for OpenCV directly. A developer who needs to detect 21 hand landmarks in a live camera feed without training their own model should use MediaPipe Tasks, which wraps that model behind a stable API. The wrong choice is using the full MediaPipe Framework (graphs and calculators) for a workload that does not involve streaming media: the overhead of the pipeline machinery adds complexity without benefit for batch tasks.

Licence, Version History, and Ongoing Development

MediaPipe is released under the Apache-2.0 licence, which permits commercial use and redistribution with attribution. The first v1.0.0 release was tagged on 2026-07-28, marking a milestone after years of pre-1.0 versioning. The last repository push was on 2026-09-25, indicating current active development.

The privacy notice in the README was last modified in June 2026, confirming the project is being actively maintained and that its privacy obligations documentation is kept up to date. The community channels listed are a Slack workspace and a Google Groups discussion forum. Contributions follow GitHub issues for bug tracking and a Stack Overflow tag for user questions.

Editorial conclusion

Developers building real-time gesture recognition, face landmarking, or pose detection for mobile or web applications will find MediaPipe Tasks a practical starting point: the pre-trained models remove the need for a training infrastructure and the on-device processing keeps user data local. The privacy trade-off is the metrics reporting built into the Tasks APIs, which requires informing users under applicable law. Teams still relying on legacy solutions deprecated in March 2023 should plan a migration, as the code in the repository is preserved on an as-is basis with no further updates. Anyone needing custom pipeline logic below the Tasks API level should expect a steep ramp into Bazel builds and C++ calculator graphs.

Frequently asked questions

What is MediaPipe used for?

MediaPipe is used for running machine learning inference on images, video, and audio streams directly on Android, iOS, web browsers, and desktop systems. Common tasks include pose estimation, hand tracking, face landmarking, and audio classification.

Is MediaPipe owned by Google?

MediaPipe was created by Google and is maintained under the google-ai-edge organization on GitHub. It is released as open source under the Apache-2.0 licence.

Is MediaPipe still supported?

The repository received its last push on 2026-09-25 and released v1.0.0 in July 2026, so it is under active development. However, the older Legacy Solutions set was officially deprecated on March 1, 2023, and those components receive no further updates.

How do you install MediaPipe in Python?

The README points to a Python setup guide at developers.google.com/mediapipe/solutions/setup_python. Building from source requires setting up Bazel and Clang 16 as shown in the repository Dockerfile; binary installs follow the platform-specific guides on the developer documentation site.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
  4. Release notes
For maintainers

Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/google-ai-edge-mediapipe.svg)](https://hysenlabs.com/projects/google-ai-edge-mediapipe)
Community notes

Community notes