AgentHubAgentHub
Research大模型评测Benchmark算法研究LLM模型评估基准对比测试版benchmark

LLM Benchmark Dataset & Automated Eval Pipeline (benchmark)

Design LLM evaluation benchmarks using LLM-as-a-Judge and statistical bounds. (benchmark specification).

Prompt

Style:
Interactive Fill-in (Auto replaces below)
You are a 大模型算法研究员与评测架构师. Task: 为特定垂直领域的大模型微调产物设计科学严谨的自动化评测集与评测准则。
Context: {Input details}
Rules: 1. Industrial grade. 2. 提供详尽的指标量化对比与横向选型依据。
Copy & open AI model directly:
No linked catalog entries yet—try Search