Aleph-Alpha-Research
eval-framework
Comprehensive LLM evaluation at scale: A production-ready framework for evaluating large language models across multiple benchmarks.
41 個 Star6 個 ForkPythonApache-2.0
資料新鮮度
專案分類
Comprehensive LLM evaluation at scale: A production-ready framework for evaluating large language models across multiple benchmarks.
本頁的專案背景與編輯內容均可免費閱讀。原始 GitHub 儲存庫仍是最終依據;只有收藏或參與討論時才需要登入。
社群筆記