rag-from-scratch
Demystify RAG by building it from scratch. Local LLMs, no black boxes - real understanding of embeddings, vector search, retrieval, and context-augmented generation.
C’est le groupe le plus large : interfaces de chat, SDK et frameworks, outils de prompts et d’évaluation, et applications reposant sur des grands modèles de langage. Beaucoup fonctionnent avec les modèles de plusieurs fournisseurs, et certains peuvent faire tourner des modèles sur votre propre matériel.
Pour tout ce que vous comptez mettre en production, vérifiez d’abord trois faits : la licence, car certains projets populaires restreignent l’usage commercial ; la prise en charge du fournisseur de modèles que vous utilisez déjà ; et la maintenance du projet. L’encadré « En bref » de chaque page projet répond aux questions de licence et de maintenance à partir des données GitHub.
Demystify RAG by building it from scratch. Local LLMs, no black boxes - real understanding of embeddings, vector search, retrieval, and context-augmented generation.
The classic Sillytavern, now has been rewritten in Tauri/Rust.
This repository provides an advanced Retrieval-Augmented Generation (RAG) solution for complex question answering. It uses sophisticated graph based algorithm to handle the tasks.
The unified multimodal backend for AI data apps. Database, orchestration, and serving in one Python file.
Open Source Implementation of Karpathy's LLM Wiki. Upload documents, connect your Claude account via MCP, and have it write your wiki !
Intel AutoRound est une boîte à outils d'optimisation de modèles pour les flux de travail de quantification qui réduit les coûts d'inférence tout en préservant la précision du déploiement de l'IA.
Fast Multimodal LLM on Mobile Devices
The official GitHub page for the survey paper "A Survey on Evaluation of Large Language Models".
Korea Investment & Securities Open API Github
🦭 会记忆、能持续推进目标、会动态编排多 Agent 的跨端桌面 AI 助手,也可服务化常驻 NAS / 云端 | A cross-device desktop AI agent with memory, autonomous goals, dynamic workflows, and headless deployment
Chat language model that can use tools and interpret the results
ktx is an executable context layer for data and analytics agents 🐙 Allow Claude Code, Codex, or other AI agents to query analytical databases accurately and with full context of your company
AI-powered OSINT agent with interactive REPL, MCP server, and CLI. 19 tools. Works with Claude, GPT-4, or local models. For authorized security research only.
Team memory for engineers and their AI agents. Lives in your repo. Shared through Git.
Strip AI-writing tells from papers and grant proposals (NSF/NIH), while keeping scholarly voice and tying claims to evidence. A skill for Claude Code, Codex, and MorphMind.
[EMNLP 2025 Oral] MemoryOS is designed to provide a memory operating system for personalized AI agents.
Claude reads its own source code — 17-chapter architectural deep-dive into Claude Code v2.1.88. EN/ZH bilingual.
Tiny Model, Big Logic: Diversity-Driven Optimization Elicits Large-Model Reasoning Ability in VibeThinker-1.5B
High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native MLX, continuous batching, multimodal models, MCP tool calling, and Claude Code support.
🤖 MLE-Agent: Your intelligent companion for seamless AI engineering and research. 🔍 Integrate with arxiv and paper with code to provide better code/research plans 🧰 OpenAI, Anthropic, Gemini, Ollama, etc supported. :fireworks: Code RAG
AgentEvolver: Towards Efficient Self-Evolving Agent System
A modular, stack-agnostic toolkit of security review skills for AI coding agents to autonomously find, reproduce, and patch vulnerabilities.
Learn LLM internals step by step - from tokenization to attention to inference optimization.
Nvidia GPU exporter for prometheus using nvidia-smi binary OR using NVML