# RubyLLM: nineteen providers, thirteen verbs, one chat object

> RubyLLM is the self-described Ruby-native AI framework, offering chat, vision, audio, documents, OCR, image and video generation, embeddings, reranking, moderation, tools, agents and structured output through one consistent API in plain Ruby or Rails. Version 2.0.0 shipped 2026-09-18, the repository was pushed as recently as 2026-09-29, and the gem is MIT licensed.

**crmne/ruby_llm** — One delightful Ruby framework for every major AI provider. Build AI agents, chatbots, RAG apps, and multimodal workflows in beautiful, expressive code.

- Repository: https://github.com/crmne/ruby_llm
- Website: https://rubyllm.com/
- Stars: 4,423 · Forks: 509
- Language: Ruby
- License: MIT
- Published: 2026-09-09 · Updated: 2026-09-09 · Language: en
- Canonical page: https://hysenlabs.com/projects/crmne-ruby-llm

## Thirteen verbs, all module methods

The API's shape is a row of module-level verbs, each one AI capability in a single call. RubyLLM.chat.ask sends a prompt. RubyLLM.paint generates an image you can save with image.save. RubyLLM.animate generates a video the same way. RubyLLM.embed produces an embedding whose vectors come back from embedding.vectors. RubyLLM.rerank orders retrieval candidates, RubyLLM.transcribe turns audio into text, RubyLLM.speak turns text into audio, RubyLLM.ocr extracts a document as markdown, and RubyLLM.moderate returns flagged? on content. The simplest possible use is one line:

```ruby
RubyLLM.chat.ask "What's the best way to learn Ruby?"
```

and the learning curve from there is mostly learning which verb fits, not learning an object hierarchy. The framework tagline, build AI features the Ruby way, is a claim about this surface area.

## Files ride along with the with: parameter

Multimodal input is a keyword argument rather than a separate client. A chat backed by a model that supports the input types, the example uses gemini-3.7-flash, answers questions about files passed with with:, so one chat object handles an image, a video, a meeting recording, a contract PDF and a Ruby source file in consecutive asks. Multiple files go in one array:

```ruby
chat.ask "Analyze these files", with: ["diagram.png", "report.pdf", "notes.txt"]
```

The feature list names the capability Documents, questions about PDFs, text files and other supported formats, and Vision for images and videos, but the code shows they share a single mechanism. Streaming is symmetric, a block passed to ask receives chunks and prints chunk.content as it arrives, real time responses with blocks in the feature summary's words.

## Tools and agents are classes you declare

Giving the model your code means subclassing RubyLLM::Tool, declaring a description, and implementing execute with keyword arguments. The shipped example fetches weather from the open-meteo API inside execute and is attached with chat.with_tools(Weather).ask. Agents are the same idea held longer, a class inheriting RubyLLM::Agent declares its model, its instructions and its tools, and instances answer questions:

```ruby
class WeatherAssistant < RubyLLM::Agent
  model "gpt-5.6-luna"
  instructions "Be concise and always use tools for weather."
  tools Weather
end
```

Around this sit the control features, requires_approval parks a run until a human approves a tool call, the agentic loop of ask_later, step and complete? lets you drive iterations yourself, and with_provider_tools brings in web search, code execution and MCP connectors from the provider side.

## Structured answers via schemas and judges

Two mechanisms return typed results instead of prose. The first is schema-bound output, define a class inheriting Schematist::Schema with string, number and array fields, pass it with with_schema, and read the result from response.parsed, so the model's answer arrives as parsed Ruby data shaped like the declaration. The second is RubyLLM::Judge, for questions whose answer is a probability. A judge class declares named probabilities attached to questions, in the example a class Urgency declares probability :urgent with the question of whether something needs attention today, and calling Urgency.judge on a message returns urgent.probability. The feature summary groups these as structured output and judgments, probabilities, choices and scores, and both keep the framework's pattern of declaring intent in a class rather than configuring a client.

## Rails ties, a usage ledger and provider discounts

For Rails the integration list is specific, Active Record persistence, Active Storage attachments, Hotwire streaming and generators, so conversations and files live with the rest of the application and responses can stream over existing view machinery. Cost is treated as a first class concern, a per-attempt usage ledger sits behind chat.tokens and chat.cost. Prompt caching turns on with with_caching and marks the cacheable prefix with cache_until_here. Fallbacks retry on backup models via with_fallbacks and cancel stops a run. Batch processing goes through RubyLLM.batch for provider-side discounts, compaction lets providers condense long conversations with with_compaction, and workflows correlate multi-agent runs in telemetry through RubyLLM.workflow. Prompt templates are ERB files in app/prompts rendered by RubyLLM.render_prompt.

## Nineteen providers, local ones included

The framework advertises nineteen providers behind one API, with the promise that moving between hosted and local providers does not rewrite the application, and that any OpenAI-compatible endpoint can be connected. The repository's .env.example shows the concrete roster, Anthropic, AWS, Azure, Deepgram, DeepSeek, ElevenLabs, Gemini, Google Cloud, GPUStack, Hetzner, Mistral, Ollama, OpenAI, OpenRouter, Perplexity and xAI, with the local endpoints spelled out, Ollama at localhost:11434/v1 and GPUStack at localhost:11444/v1. Secrets are pulled from 1Password through op read substitutions rather than pasted values, and three debug switches exist, RUBYLLM_DEBUG, RUBYLLM_STREAM_DEBUG and RUBYLLM_LOG_FILE pointing at /tmp/ruby_llm.log. A model registry browses capabilities, limits and pricing across providers, and files upload once with RubyLLM.upload for reuse across chats.

## A 2.0 days old, with a conformance suite in the tree

The release cadence is fast, v2.0.0 on 2026-09-18 preceded by rc4 on 2026-09-16 and rc3 on 2026-09-14, and the repository's last push landed 2026-09-29. The README states the framework is battle tested at chatwithwork.com, a fully private work AI product, and invites users to share their story through a five minute form. Engineering hygiene is visible in the tree, a conformance directory, an Appraisals file with generated gemfiles for dependency matrices, overcommit and gitleaks configuration, RuboCop and RSpec wiring, plus AGENTS.md and CLAUDE.md for coding agents. Documentation splits by major version, the current 2.0 docs and a frozen 1.x set at rubyllm.com/v1, which is the honest way to handle a just-cut major.

## Conclusion

Choose RubyLLM when the application is already Ruby or Rails and AI features should read like the rest of the codebase, one chat object, plain classes for tools and agents, and a provider switch that does not rewrite call sites. Look elsewhere for a polyglot orchestration stack or for research-style agent graphs, this framework's bet is Ruby expressiveness rather than cross-language portability. Before committing, verify your target provider is among the nineteen and check its pricing in the model registry, confirm the 2.0 API against the current docs since the major version just landed, and read the Rails integration notes if you depend on Active Record persistence or Hotwire streaming.

## FAQ

### What is ruby llm?

RubyLLM is a MIT-licensed, Ruby-native AI framework for building agents, chatbots, RAG apps and multimodal workflows. It exposes chat, tools, agents, image and video generation, embeddings and more through one consistent API in plain Ruby or Rails, across nineteen providers including local ones like Ollama.

### Which providers does RubyLLM support?

The framework advertises nineteen providers behind one API. The repository's environment example lists Anthropic, AWS, Azure, Deepgram, DeepSeek, ElevenLabs, Gemini, Google Cloud, GPUStack, Hetzner, Mistral, Ollama, OpenAI, OpenRouter, Perplexity and xAI, and any OpenAI-compatible endpoint can be connected, with hosted and local providers interchangeable.

### How does RubyLLM integrate with Rails?

Rails integration includes Active Record persistence, Active Storage attachments, Hotwire streaming and generators. Prompt templates live as ERB files in app/prompts rendered with RubyLLM.render_prompt, so AI features sit alongside conventional Rails application code.

## Sources

- [crmne/ruby_llm on GitHub](https://github.com/crmne/ruby_llm)
- [License: MIT](https://github.com/crmne/ruby_llm/blob/main/LICENSE)
- [Project website](https://rubyllm.com/)
- [README](https://github.com/crmne/ruby_llm/blob/main/README.md)
- [Releases](https://github.com/crmne/ruby_llm/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/crmne-ruby-llm
