TileRT
Tile-Based Runtime for Ultra-Low-Latency LLM Inference
HYSEN LABS DIRECTORY
Capability tag · curated repositories and deep analysis.
Tile-Based Runtime for Ultra-Low-Latency LLM Inference
A collection of original, innovative ideas and algorithms towards Advanced Literate Machinery. This project is maintained by the OCR Team in the Language Technology Lab, Tongyi Lab, Alibaba Group.
The classic Sillytavern, now has been rewritten in Tauri/Rust.
SkillsBench evaluates how well skills work and how effective agents are at using them.
Babysitter enforces obedience on agentic workforces and enables them to manage extremely complex tasks and workflows through deterministic, hallucination-free self-orchestration
Dive is an open-source MCP Host Desktop Application that seamlessly integrates with any LLMs supporting function calling capabilities. ✨
Practical course about Large Language Models.
Generate automated tests for your Node.js app via LLMs without developers having to write a single line of code.
A high-performance inference engine for AI models
Synthesizing Graphics Programs for Scientific Figures and Sketches with TikZ.
Cloud native, ultra-high performance AI&API gateway, LLM API management, distribution system, open platform, supporting all AI APIs.🦄云原生、超高性能 AI&API网关,LLM API 管理、分发系统、开放平台,支持所有AI API,不限于OpenAI、Azure、Anthropic Claude、Google Gemini、DeepSeek、字节豆包、ChatGLM、文心一言、讯飞星火、通义千问、360 智脑、腾讯混元等主流模型,统一 API 请求和返回,API申请与审批,调用统计、负载均衡、多模型灾备。一键部署,开箱即用。
Mellea is a library for writing generative programs.
Control panel for VLLM, Sglang, llama.cpp, exllamav3
Ship your code, on autopilot. An open source agent that lives on your machines 24/7 and keeps your apps running. 🦀
OpenGUI is an Android GUI agent framework for phone-use AI that can see, plan, and operate real mobile apps through the GUI.
🧡 The meta framework for code generation. Automate OpenAPI to type-safe TypeScript, Zod, and TanStack Query with a modular, plugin-based engine.
The AI-Engineering Foundation Framework for CRM/ERP and commerce: open-source TypeScript, with multi-tenancy, RBAC, events and domain modules already decided as conventions and specs, so Cursor, Claude Code and Codex build features instead of re-deciding architecture. Start with 80% done.
Community maintained hardware plugin for vLLM on Apple Silicon
Qwen3.8-27B on a single RTX 3090 with vLLM: ~1,000 tok/s at 64 concurrent (int8 tensor-core GEMMs, fp16 DeltaNet state), ~114 tok/s single-user at default sampling / ~124 greedy (MTP drafts, own-output draft vocab, calibrated int4 lm_head, split-KV verify attention), 150k-262k context; patches, requant scripts, benchmarks
Underthesea - AI Assistant