# Vision Agents: Building low-latency video and voice agents on Stream's edge network

> Vision Agents is an open-source Python framework for building AI agents that process video and audio in real-time. Built by Stream, it integrates with OpenAI, Gemini, and other LLMs, maintains sub-30ms latency via Stream's edge network, and includes vision processors, RAG, and phone integration.

**GetStream/Vision-Agents** — Open Vision Agents by Stream. Build voice and vision agents quickly with any model or video provider. Uses Stream's edge network for ultra-low latency.

- Repository: https://github.com/GetStream/Vision-Agents
- Website: https://visionagents.ai
- Stars: 8,141 · Forks: 683
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/getstream-vision-agents

## Real-time video understanding with pluggable LLM backends

Vision Agents provides building blocks for creating intelligent video experiences that combine video input, computer vision processing, and LLM intelligence. You define an agent with an LLM backend (OpenAI, Gemini, Claude, Qwen, or others), optional vision processors (YOLO, Roboflow, custom PyTorch models), and your instructions. The README shows a golf coaching example:

```python
agent = Agent(
    edge=getstream.Edge(),
    agent_user=agent_user,
    instructions="Read @golf_coach.md",
    llm=gemini.Realtime(fps=10),
    processors=[ultralytics.YOLOPoseProcessor(model_path="yolo11n-pose.pt", device="cuda")],
)
```

This agent runs a YOLO pose detector at 10 fps (10 frames per second), sends frames to Gemini Live, and returns coaching feedback. The agent reads instructions from a markdown file, allowing you to change behavior without code changes. The modularity allows you to swap any component: change the LLM provider, add different processors for pre-processing frames before LLM analysis or post-processing outputs, or use a custom edge network. Multi-modal capabilities mean the agent can process both video and audio, making it suitable for coaching, commentary, and surveillance scenarios.

## Ultra-low latency with Stream's edge network

The core advantage of Vision Agents is sub-30ms audio and video latency, critical for real-time interaction. Stream has deployed a global edge network that caches models close to users and handles media transmission with minimal delay. When you join a video call, participants join in under 500ms. This is critical for interactive applications like live coaching (instant feedback for swing correction), drone fire detection (response time matters), sports coaching (coaching must happen during play), physical therapy (form feedback must be immediate), workout coaching (rep counting and form guidance), and real-time sports commentary (commentary must match action). The latency advantage makes the difference between usable and unusable for applications requiring synchronous interaction. The framework is open and works with any video edge network, but Stream's is the primary deployment target and has built-in SDKs. The README emphasizes that latency is built into the architecture, not a trade-off or afterthought. Stream provides 333,000 participant minutes per month free for developers (substantial for testing), plus extra credits via the Maker Program for early-stage companies.

## Integrations with LLMs and speech services

Vision Agents ships with integrations for LLMs including OpenAI, Gemini, xAI, OpenRouter, Hugging Face, Kimi, MiniMax, and Telnyx. For realtime LLM conversation with low latency, it supports OpenAI Realtime, Gemini Live, AWS Nova Sonic, Qwen, and Inworld. Realtime models enable synchronous conversation without turn-taking latency, critical for interactive scenarios.

Speech-to-text (STT) providers include Deepgram, AssemblyAI, Fast-Whisper, Fish Audio, Wizper, Mistral Voxtral, and Telnyx. Text-to-speech (TTS) includes ElevenLabs, Cartesia, Deepgram, AWS Polly, Pocket, Kokoro, Inworld, Fish Audio, and Telnyx. This variety allows you to pick TTS voices and STT accuracy based on your language and dialect needs.

Vision processors include Ultralytics (YOLO for object detection), Roboflow (for custom vision models), Moondream (for visual reasoning), TwelveLabs (for video understanding), NVIDIA (for edge processing), and Decart (for video effects). This breadth of integrations lets you choose the best model for each task without locking into a single provider. You can also write custom processors as PyTorch or ONNX models.

## Features for production video agents

The framework includes several features for building production systems. Real-time WebRTC streams video directly to model providers, handling media transport. A pluggable processor pipeline lets you run vision models like YOLO before sending frames to the LLM or after receiving LLM output for post-processing. Turn detection with VAD (voice activity detection) and speaker diarization handles natural conversation flow, detecting when to interrupt and when to listen. Tool calling and MCP server integration allow the agent to execute code or fetch data mid-conversation, such as creating Linear issues, fetching weather data, making API calls, or calling telephony APIs.

Phone integration via Twilio or Telnyx allows inbound and outbound calls with bidirectional audio streaming, making agents accessible via traditional phone. RAG (retrieval-augmented generation) retrieval works with TurboPuffer, Qdrant vector search, or Gemini FileSearch for semantic search over documents, enabling agents to cite sources and answer questions grounded in your data. Memory via Stream Chat lets agents recall context across turns and sessions, building relationships with users. A text back-channel allows silent messages to the agent during a call for coaching overlays, remote instruction, or silent coaching tips without audio interruption.

## Examples for coaching, security, and multi-modal applications

Vision Agents includes eleven example projects in the repository demonstrating different use cases and integration patterns. Example 01 is a simple agent example. Example 02 is golf coaching with YOLO pose detection and Gemini Live, useful for sports coaching. Example 03 combines phone integration and RAG for document-aware phone agents. Example 04 is football commentary using video analysis. Example 05 is security camera monitoring for anomaly detection. Example 06 shows how to integrate Prometheus metrics for monitoring. Example 07 demonstrates Kubernetes deployment for production scaling. Example 08 is an agent server showing how to build a service. Example 09 is a sales assistant for video shopping. Example 10 is local transport monitoring. Example 11 is moderation for content filtering. Each example demonstrates composing agents from different processors and integrations. The README notes that combining a fast object detection model (like YOLO) with a full realtime LLM is useful for many video AI use cases including drone fire detection and physical therapy coaching.

## Production-ready deployment with metrics and Kubernetes

Vision Agents includes a built-in HTTP server, Prometheus metrics for monitoring, horizontal scaling support, and Kubernetes deployment examples. The examples directory includes a complete k8s deployment example showing how to containerize and scale agents. Metrics tracking includes request counts, latency histograms, and error rates. The framework is designed for production use with proper error handling, retry logic, and graceful shutdown. You can run multiple agents behind a load balancer for horizontal scaling. The examples show how to set up proper logging and monitoring.

## Installation with uv and configuring provider credentials

Install the framework and integrations using uv, a fast Python package installer:

```bash
uv add vision-agents
uv add "vision-agents[getstream, openai, elevenlabs, deepgram]"
```

You can install subsets of integrations by including them in brackets. Optional extras include getstream, openai, elevenlabs, deepgram, and many others for different providers. Configure API credentials in a .env file for Stream (API key and secret), OpenAI (API key and model), Gemini, Deepgram, ElevenLabs (voice ID), Cartesia, Anthropic, Roboflow, LemonSlice, and MiniMax. The repository includes .env.example documenting all required environment variables. Once installed, follow the quickstart guide at https://visionagents.ai/introduction/quickstart to build your first agent. The framework includes multiple example projects ranging from simple coaching agents to complete production deployments with RAG and memory. Install goreleaser for building distributions.

## When to use Vision Agents vs alternatives

Vision Agents excels at real-time interactive scenarios like live sports coaching, drone monitoring, security surveillance, and virtual shopping assistants. The low latency makes it suitable for applications where users expect immediate feedback. The framework requires integration with Stream's infrastructure or a compatible edge network for the latency guarantees. If you need purely on-premises processing without Stream's edge, you can still use the framework but will not achieve sub-30ms latency, making it less compelling compared to local alternatives. Additionally, costs depend on which LLM and speech providers you choose. Stream's free tier (333,000 participant minutes) covers experimentation but production deployments incur costs.

## Conclusion

Adopt Vision Agents if you need low-latency video or voice agent capabilities and can work with Stream's edge network infrastructure. The framework handles real-time processing pipelines, integrates with many LLM providers, and is production-ready with metrics and Kubernetes support. Avoid it if you need purely on-premises solutions or do not require sub-30ms latency. Before deploying, verify that Stream's API and your chosen LLM provider support your use case, that you can obtain necessary API credentials, and that Stream's edge network coverage reaches your deployment regions.

## FAQ

### How do you build an agent with Vision Agents?

Import the Agent class, initialize it with an edge (Stream or custom), an LLM (OpenAI, Gemini), optional processors (YOLO), and instructions. Call methods to stream video, handle audio, and invoke the LLM.

### Is Vision Agents free?

Vision Agents is open-source and free. Stream provides 333,000 participant minutes per month free for developers. LLM and speech service costs depend on your chosen providers.

### What integrations are available?

Vision Agents integrates with OpenAI, Gemini, Claude, Deepgram, ElevenLabs, Ultralytics, Roboflow, Twilio, Telnyx, and many others. See the integrations documentation for the full list.

### Can Vision Agents run on-premises without Stream's edge?

Yes, but you will not achieve sub-30ms latency guarantees. Vision Agents is designed to work with any edge network, but Stream's is the primary deployment target for low-latency scenarios.

### What are the example applications?

Examples include golf coaching, phone and RAG integration, football commentary, security camera monitoring, sales assistance, and Kubernetes deployment. Each shows different processor and integration combinations.

## Sources

- [GetStream/Vision-Agents on GitHub](https://github.com/GetStream/Vision-Agents)
- [License: Apache-2.0](https://github.com/GetStream/Vision-Agents/blob/main/LICENSE)
- [Project website](https://visionagents.ai)
- [README](https://github.com/GetStream/Vision-Agents/blob/main/README.md)
- [Releases](https://github.com/GetStream/Vision-Agents/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/getstream-vision-agents
