exllamav3
An optimized quantization and inference library for running LLMs locally on modern consumer-class GPUs
This is the broadest group: chat interfaces, SDKs and frameworks, prompt and evaluation tools, and applications built on large language models. Many work with models from several providers, and some can run models on your own hardware.
For anything you plan to ship, check three facts first: the licence, because some popular projects restrict commercial use; whether it supports the model provider you already use; and whether it is still maintained. The At a glance box on every project page answers the licence and maintenance questions from GitHub data.
An optimized quantization and inference library for running LLMs locally on modern consumer-class GPUs
open source grok bot. gawk bots automate your menial work via AI models and build you microapps to manage the outcome, so that you have a false sense of control.
Samurai-inspired multi-agent system for Claude Code. Orchestrate parallel AI tasks via tmux with shogun → karo → ashigaru hierarchy.
Qwen3.8-27B on a single RTX 3090 with vLLM: ~1,000 tok/s at 64 concurrent (int8 tensor-core GEMMs, fp16 DeltaNet state), ~114 tok/s single-user at default sampling / ~124 greedy (MTP drafts, own-output draft vocab, calibrated int4 lm_head, split-KV verify attention), 150k-262k context; patches, requant scripts, benchmarks
An IDA plugin for binary file analysis, powered by AI models such as OpenAI and DeepSeek.
Build production-ready LLM applications and advanced agents using Python, LangChain, and LangGraph. This is the companion repository for the book on generative AI with LangChain.
Fullstack SaaS Boilerplate built with tRPC, Fastify and React
🍦 Speech-AI-Forge is a project developed around TTS generation model, implementing an API Server and a Gradio-based WebUI.
Decision audit trail + persistent memory for AI trading agents. Outcome-weighted recall, tamper-evident SHA-256 chain with RFC 3161 anchoring, 20 MCP tools.
Universal Claude Code workflow plugin with agents, skills, hooks, and commands
A thin orchestration dashboard over Claude Code for managing context, automation, and developer headspace. You need tentacles. 🦑
Agent-driven media library for your cloud drives (Quark 夸克 / 115 / 光鸭 GuangYa / 123网盘 / 天翼 Tianyi)
BrowserWing turns your browser actions into MCP commands Or Claude Skill, allowing AI agents to control browsers efficiently and reliably. Say goodbye to slow, token-heavy LLM interactions — let agents call commands directly for faster automation. Perfect for AI-driven tasks, browser automation, and boosting productivity.
Agent-driven automated CVE discovery platform for source code auditing, vulnerability verification, and report generation.
🛸 Use pretrained transformers like BERT, XLNet and GPT-2 in spaCy
Fine-tune LLMs on your Mac with Apple Silicon. SFT, DPO, GRPO, Vision, TTS, STT, Embedding, and OCR fine-tuning — natively on MLX. Unsloth-compatible API.
An open-source, code-first Typescript toolkit for building, evaluating, and deploying sophisticated AI agents with flexibility and control.
Universal AI context generator. Saves thousands of tokens per conversation in Claude Code, Cursor, Copilot, Codex, and more.
The easiest way to use Ollama in .NET
💃 Dance with Intelligence in Your Code. Minuet offers code completion as-you-type from popular LLMs including OpenAI, Gemini, Claude, Ollama, Llama.cpp, Codestral, and more.
Trace-native CI/CD for AI agents — production failures become regression tests that block the PR. Auto-detect, cluster, freeze into hermetic cases, replay in CI for $0.
Portfolio of data science projects completed by me for academic, self learning, and hobby purposes.
A modular and comprehensive solution to deploy a Multi-LLM and Multi-RAG powered chatbot (Amazon Bedrock, Anthropic, HuggingFace, OpenAI, Meta, AI21, Cohere, Mistral) using AWS CDK on AWS
On-device LLM execution in React Native with Vercel AI SDK compatibility