Library / SDK
Intent-Lab/VisionClaw avatar
Intent-Lab/VisionClaw

VisionClaw: a Gemini Live assistant that sees through Ray-Ban glasses

Real-time AI assistant for Meta Ray-Ban smart glasses -- voice + vision + agentic actions via Gemini Live and OpenClaw

2,555 stars511 forksTypeScriptNOASSERTION

At a glance

What is it?
A TypeScript reference app for Meta smart glasses that streams camera frames and microphone audio to Gemini Live over a WebSocket, then routes tool calls to an optional OpenClaw gateway.
Who is it for?
VisionClaw is worth reading as a wiring diagram rather than as a product. The gateway, the LiveKit worker and the glasses link are three separate failure points, and the README is unusually candid about all three, including the case where the app connects and then waits on a status message forever because the worker was never started.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 16 days ago.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 20, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What the app does with a camera stream

The project is a real-time assistant for Meta Ray-Ban smart glasses. The description in the repository says it sees what you see, hears what you say and takes actions on your behalf, all through voice, and the four example utterances underneath spell out what that means in practice: ask what you are looking at and Gemini describes the scene, ask to add milk to a shopping list and OpenClaw does it, ask to send a message to John about being late and it routes through WhatsApp, Telegram or iMessage, ask for the best coffee shops nearby and web search results come back spoken.

The transport specifics are the interesting part and the README states them plainly. The glasses camera streams at roughly one frame per second to Gemini for visual context, while audio flows bidirectionally in real time. The data path diagram in the README describes frames as JPEG at about one frame per second and microphone audio as PCM at 16kHz going up, with the response coming back as PCM at 24kHz for playback. That asymmetry is worth noticing if you have worked on voice models before, because it means the app is not doing speech recognition followed by a chat turn. The Gemini Live API is used in its native audio mode over a WebSocket.

The supported platforms line says iOS on iPhone and Android on Pixel, Samsung and similar devices. The build sits on the Meta Wearables DAT SDK for iOS and a DAT Android SDK, with a LICENSE file and a NOTICE file in the tree. The repository was last pushed on 2026-09-20 and is not archived. The license field on the GitHub side reads NOASSERTION, so the project does not present itself as carrying a single standard open source license in the metadata even though a LICENSE file is present in the tree.

The access code is a gateway token rather than a signup

The first screen asks for an access code, and the README spends a paragraph deflating what a reader might assume that means. The code is not issued by the author. It is a token from whichever gateway the app points at, and the gate verifies it against that gateway before unlocking.

Three defaults are worth knowing before you clone anything. The default gateway at api.visionagents.app is described as a small private pilot instance for a user study, not a public service, so its codes are not handed out. That sentence tells you what happens if you build the app and type in a code you invented: nothing good. Self hosting means running the gateway from the `gateway/` directory and setting `GATEWAY_TOKENS` in its `.env` file to a secret and name pair separated by a colon, where that secret becomes your access code. The app can then be pointed at your gateway either before building, by editing `gatewayBaseUrl` in `Secrets.swift` or `Secrets.kt`, or at runtime through the "Using your own gateway?" option on the access code screen.

The gateway is only half the requirement, and the README is direct about this. Voice needs the whole stack. The gateway mints room tickets and runs tasks, but calls also need the `agent/` worker running against a LiveKit Cloud project, with the free tier sufficient, and the same LiveKit credentials on both the gateway and the worker. Skip the worker and the app connects successfully, then sits on a "Waiting for agent" message indefinitely. A third prerequisite is hardware side: Ray-Ban Meta glasses need Developer Mode enabled in the Meta AI app, and without it the glasses link is rejected with "Error opening link".

Phone mode tests the pipeline before you touch the glasses

The most useful line in the README for anyone evaluating this project is that the full pipeline can be tested with the phone camera instead of glasses. Phone mode is listed as a first class feature rather than a debugging leftover, and the iOS instructions describe tapping "Start on iPhone", which uses the iPhone's back camera, then tapping the AI button to open a Gemini Live session.

That ordering is deliberate and it changes how you should read the repository. You can validate voice and vision, including the tool call path, with one Gemini API key and a phone, before touching Developer Mode, GitHub Packages or LiveKit. Anything that breaks at that point is genuinely in the Gemini integration. Anything that breaks later is in the wearable SDK plumbing.

The project also lists a WebRTC streaming feature, described as sharing your glasses point of view live to a browser viewer. The secrets files carry optional OpenClaw and WebRTC configuration alongside the required Gemini API key, which tells you the architecture treats OpenClaw and the browser viewer as add-ons to a core loop rather than as dependencies. OpenClaw itself is marked optional in the README and described as a local gateway that gives Gemini access to 56 tools and all your connected apps, with web search, messaging, smart home, notes and reminders given as examples.

iOS setup is four steps and one secrets file

The iOS quick start is short enough to reproduce exactly. Clone the repository, change into the sample project directory and open the Xcode project:

bash
git clone https://github.com/sseanliu/VisionClaw.git
cd VisionClaw/samples/CameraAccess
open CameraAccess.xcodeproj

Then the secrets file is copied from the example that ships with it:

bash
cp CameraAccess/Secrets.swift.example CameraAccess/Secrets.swift

You fill in the Gemini API key, which the README marks as required, plus optional OpenClaw and WebRTC configuration, build to your iPhone with Cmd+R, and then either run phone mode or pair the glasses. The glasses path has its own five step gate: open the Meta AI app on your iPhone, go to Settings via the gear icon in the bottom left, tap App Info, tap the App version number five times to unlock Developer Mode, then go back to Settings where a Developer Mode toggle now appears.

That tap-the-version-five-times ritual is a Meta account level setting rather than something VisionClaw controls, and it is worth knowing that it lives in the Meta AI app rather than in VisionClaw. The `samples/` directory in the tree holds this iOS project alongside the Android one and a `CLAUDE.md`, and the root also carries `CHANGELOG.md`, `CODE_OF_CONDUCT.md`, `CONTRIBUTING.md`, `assets/` for the teaser, cover and diagram images, plus the `agent/` and `gateway/` directories the README keeps referring to.

Android needs a GitHub Packages token before any code compiles

The Android quick start carries one obstacle the iOS path does not. The Meta DAT Android SDK ships through GitHub Packages, and GitHub Packages requires authentication even for public repositories. You need a classic Personal Access Token with the `read:packages` scope, created from the developer settings page, and it goes into `local.properties` in `samples/CameraAccessAndroid/`:

properties
github_token=YOUR_GITHUB_TOKEN

A 401 from Gradle therefore means the token is missing or wrong rather than the repository being private. The README suggests a shortcut if you already have the `gh` command line tool, running `gh auth token` to get a value, and `gh auth refresh -s read:packages` if the scope is absent. Getting those two facts straight saves most of the time people lose to this step, because the failure looks like a network or repository problem and is neither.

The secrets file for Android sits much deeper in the tree, under `app/src/main/java/com/meta/wearable/dat/externalsampleapps/cameraaccess/`, and is copied the same way:

bash
cd samples/CameraAccessAndroid/app/src/main/java/com/meta/wearable/dat/externalsampleapps/cameraaccess/
cp Secrets.kt.example Secrets.kt

After that it is a normal Android Studio loop: let Gradle sync and pull the DAT SDK from GitHub Packages, select a device, and run with Shift+F10. The README also documents wireless installation over ADB, which needs Wireless debugging enabled in Developer Options and then pairing with `adb pair`. Phone mode on Android is the mirror image of the iOS flow, using the "Start on Phone" control and the sparkle icon for the AI button.

A reference implementation with an unfinished service behind it

The honest read on VisionClaw is that it is a well documented integration sample pointed at a service that is not generally available. The GitHub side of the project describes the default gateway as a private pilot instance for a user study, which means the path a new user sees on first launch ends in a gate they cannot pass without running their own gateway. That is a deliberate choice rather than an oversight, and the README spends real effort making the self hosted path workable, including the exact environment variable and the runtime override on the access code screen.

For readers deciding whether to adopt it, the split is clean. What the repository genuinely settles is the integration surface: which SDKs are involved on each platform, the frame and audio rates, where secrets live, what the WebSocket carries, and which failure messages correspond to which missing prerequisite. What it does not settle is anything about the capabilities being exposed, since the 56 tools and connected apps belong to OpenClaw, and nothing about the data handling of a third party gateway you would not be running.

There are also no releases attached to the repository, so there is no version history to point at and nothing to pin. Combined with a push date of 2026-09-20, the practical read is that this is a moving target maintained alongside the underlying SDKs rather than a versioned dependency. Treat it as a working reference you copy from, not a library you depend on.

Editorial conclusion

VisionClaw is worth reading as a wiring diagram rather than as a product. The gateway, the LiveKit worker and the glasses link are three separate failure points, and the README is unusually candid about all three, including the case where the app connects and then waits on a status message forever because the worker was never started. What it settles for you is the transport details: JPEG frames at roughly one frame per second, PCM audio at 16kHz up and 24kHz back, and tool calls arriving as ordinary execute requests over the same socket. What it does not settle is anything about OpenClaw's 56 skills or about what a private pilot gateway is doing with your frames. Start in phone mode with a single Gemini key, then add the gateway and the glasses once the voice loop already works.

Frequently asked questions

What does VisionClaw actually do with Meta Ray-Ban glasses?

It streams the glasses camera and microphone to the Gemini Live API over a WebSocket so you can talk to an assistant that can see your surroundings, and it can route tool calls to an OpenClaw gateway for actions like messaging, web search and smart home control. The camera goes up as JPEG at roughly one frame per second and audio as 16kHz PCM, with 24kHz PCM coming back for playback.

Why does VisionClaw ask for an access code?

The code is not an account signup. It is a token belonging to whichever gateway the app points at, and the gate verifies it against that gateway before unlocking. The default gateway is a private pilot instance whose codes are not distributed, so you self host the gateway and set GATEWAY_TOKENS in its .env file to mint your own.

Why does the app hang on "Waiting for agent"?

The gateway only mints room tickets and runs tasks, so a voice session also needs the agent worker running against a LiveKit Cloud project with matching credentials on both sides. The app connects fine without the worker and then waits on that message indefinitely, which the README calls out as the main voice setup trap.

Can I test VisionClaw without owning Meta glasses?

Yes. Phone mode uses your iPhone or Android back camera in place of the glasses feed, and the README treats it as a first class way to try the project. You can exercise voice, vision and the tool call path with just a Gemini API key before dealing with Developer Mode, GitHub Packages or LiveKit.

What do I need before the Android build will compile?

A GitHub classic Personal Access Token with the read:packages scope, placed in local.properties as github_token, because the Meta DAT Android SDK is served from GitHub Packages and that service authenticates even public repositories. A 401 from Gradle almost always means the token or its scope is wrong rather than the repository being private.

Official sources

  1. Intent-Lab/VisionClaw on GitHub
  2. Issues
  3. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/intent-lab-visionclaw.svg)](https://hysenlabs.com/projects/intent-lab-visionclaw)