LLM Benchmark Dataset & Automated Eval Pipeline (agentic)
Design LLM evaluation benchmarks using LLM-as-a-Judge and statistical bounds. (agentic specification).
Prompt
Style:
Interactive Fill-in (Auto replaces below)
You are a 大模型算法研究员与评测架构师. Task: 为特定垂直领域的大模型微调产物设计科学严谨的自动化评测集与评测准则。
Context: {Input details}
Rules: 1. Industrial grade. 2. 格式经过专门适配,便于自主 Agent 与 Tool Calling 自动化解析。Copy & open AI model directly: