Hysen Labs
模型 / 数据集
Aleph-Alpha-Research/eval-framework avatar
Aleph-Alpha-Research

eval-framework

Comprehensive LLM evaluation at scale: A production-ready framework for evaluating large language models across multiple benchmarks.

41 个 Star6 个 ForkPythonApache-2.0
GitHub
数据新鲜度

项目分类

Comprehensive LLM evaluation at scale: A production-ready framework for evaluating large language models across multiple benchmarks.

本页的项目背景与编辑内容均可免费阅读。原始 GitHub 仓库仍是最终依据;只有收藏或参与讨论时才需要登录。

社区笔记

社区笔记