SWE-smith
[NeurIPS 2025 D&B Spotlight] Scaling Data for SWE-agents
An AI agent is a program that uses a language model to decide what to do next, calls tools such as a browser, a code interpreter or an API, and repeats until the task is done. The projects here range from frameworks for building your own agents to finished agents that write code, browse the web or run business workflows.
Compare three things. Which models it can use: a single provider, or any OpenAI-compatible endpoint including local ones. How it limits what the agent may do: sandboxing, approval steps and permission lists. And whether it keeps memory between steps and sessions. An agent that runs code or shell commands should run in a sandbox, so check that before giving it access to your machine.
[NeurIPS 2025 D&B Spotlight] Scaling Data for SWE-agents
Learn any journal's writing conventions from its published papers, then revise your manuscript to match — section by section.
parrot.nvim 🦜 - the plugin that brings stochastic parrots to Neovim.
A Python library powered by Language Models (LLMs) for conversational data discovery and analysis.
An LLM-as-a-judge HTTP proxy to secure agents in production
SimWorld: An Open-ended Realistic Simulator for Autonomous Agents in Physical and Social Worlds
Codebase intelligence for AI. Detects patterns & conventions + remembers decisions across sessions. MCP server for any IDE. Offline CLI.
Fast, streaming indexing, query, and agentic LLM applications in Rust
Direct complete videos with multi-mode workflows for research, scripts, visual development, shot design, media generation, and delivery.
The retrieval layer for production AI systems. Lightning-fast (<10ms) search without vector databases. Built for browser, edge, on-device, and cloud.
Scalable RL for Any Agent and Sandbox.
An open-source platform for long-duration multi-agent workflows. This project automates long-running, multi-agent AI work by distilling each agent’s output into compact context for the next agent
open source agent engineering platform: traces, evals, and metrics to debug and improve your AI agents. Integrates with LangGraph, CrewAI, Claude Agent SDK, and more.
🍰 PromptLayer - Maintain a log of your prompts and OpenAI API requests. Track, debug, and replay old completions.
PentestCode - Multi-agent AI penetration testing system with persistent engagement state, strategic coordination, and parallel autonomous operations.
The open-source state layer for AI coding agents. Turn chaotic agent sessions into structured, traceable workflows with a local workspace for runs, events, and collections.
CheetahClaws: A Fast and Easy-to-Use Agent Harness Infrastructure for Long-Horizon, Multi-Model, and Tool-Using AI Systems
DaC is a dashboard-as-code tool. Build interactive dashboards using YAML and JSX. Built-in semantic layer. Get your agents to build standardized, reviewable dashboards.
MCP server that integrates the LINE Messaging API to connect an AI Agent to the LINE Official Account.
Turns corrections into Preferences, Project-specific skills, and Shared skills for Claude Code, Codex, and OpenCode.
Conversational voice AI agents
飞牛 fnOS NAS 第三方应用商店 — 115 款自托管应用的 .fpk 安装包 | Plex, Emby, Jellyfin, qBittorrent, Immich, Sonarr, Radarr, Vaultwarden, ZeroClaw AI, OpenClaw 等 | 每日自动同步上游版本
Web fetch, search, and crawl for AI agents. Built from scratch in Rust. No keys, no accounts. AGPL v3.
Executable, measurable, and reproducible AI4AI toward recursive self-improvement. Home of OpenMLE and Frontis-MA1.