SkillEvaluator
Multi-tier framework for evaluating AI agent skills with quality gates, semantic overlap detection, synthetic evaluation dataset generation, and live agent evaluation that measures how skills affect agent behavior.
An AI agent is a program that uses a language model to decide what to do next, calls tools such as a browser, a code interpreter or an API, and repeats until the task is done. The projects here range from frameworks for building your own agents to finished agents that write code, browse the web or run business workflows.
Compare three things. Which models it can use: a single provider, or any OpenAI-compatible endpoint including local ones. How it limits what the agent may do: sandboxing, approval steps and permission lists. And whether it keeps memory between steps and sessions. An agent that runs code or shell commands should run in a sandbox, so check that before giving it access to your machine.
Multi-tier framework for evaluating AI agent skills with quality gates, semantic overlap detection, synthetic evaluation dataset generation, and live agent evaluation that measures how skills affect agent behavior.
shadcn/ui, but for building agents. 🤖
Structural memory for AI coding agents. Bi-temporal graph, MCP-native, zero LLM calls. Cursor · Claude Code · Codex · DeepSeek Harness · Hermes · VS Code · Windsurf.
Aser is a lightweight, self-assembling AI Agent frame.
30 sec to give your AI agents persistent memory. Reduce 90% token consumption while also maintaining quality.
A curated registry of reusable Mercury Agent, Open Claw or Hermes Agent skills designed for real developer workflows, persistent memory, and token-efficient execution.
Agent-driven alpha factory — LLM autonomously designs, backtests, and submits factors to WorldQuant BRAIN
Explore the unknown, build the future, own your data.
A multithreaded 🕸️ web crawler that recursively crawls a website and creates a 🔽 markdown file for each page, designed for LLM RAG
Multi-harness control plane for Claude Code, Codex, Cursor, and OpenCode: quota-aware rotation across multiple Claude/Codex subscriptions, shared thread context, and cross-model review.
Three-layer audit enforcement framework for AI agents — hooks, skills, context recovery, and audit-driven daily reports
Flocks is an agentic SecOps platform.
Making ANY Software Skill-Native -- Auto-generate production-ready AI Agent Skills for Claude Code, OpenClaw, Codex, and more.
Companionship, chat, coding, and work share one memory and context framework — the kind of AI you see in science fiction: it keeps you company, and it gets things done with you.(这是一个基于上下文和注意力机制做的一个多元化的agent项目)
Project/Worktree manager, deeply integrated with AI agents to increase productivity and multitask efficiently
An autonomous AI scientist: a multi-agent loop over literature, experiments, self-critique and write-up, with deterministic guards against reward-hacking and hallucination.
CLI agent-readiness measurement, command-shape inference, and CI scorecards
Open-source template for a durable personal AI agent — web chat, Slack, Linear, and long-term memory with user-approved saves. Eve, Nuxt, Better Auth, Vercel Connect.
QuickDesk is the first AI-native remote desktop — an open-source, free application with a built-in MCP (Model Context Protocol) Server that lets any AI agent see and control remote computers.
ClearML - Auto-Magical CI/CD to streamline your AI workload. Experiment Management, Data Management, Pipeline, Orchestration, Scheduling & Serving in one MLOps/LLMOps solution
Local-first marketing operations control center for CRM, outreach, content, analytics, approvals, automations, and agent workflows.
给 Codex Desktop 一键换肤:OpenAI Codex/ChatGPT 桌面端主题工具,CDP 注入零修改应用,Miku/原神/鸣潮/火影/恋与深空 9 预设+自定义图片取色 | One-click theme & skin switcher for OpenAI Codex Desktop (macOS/Windows)
A tool-use-focused LLM plugin for neovim.
Living project docs for coding agents: keep guides, progress logs, change maps, and handoff context updated as your repo evolves.