eval-report

仓库创建 2025年7月8日最近提交 18 天前SkillHot 收录 8 小时前
▸ 精选理由

把分散的评测信号综合成可执行的高层结论和优先级,便于决策。

这个 Skill 做什么

汇总多项评测结果,生成综合分析报告、趋势与优先整改建议。

把多项评测结果合成一份可读的报告:成熟度评估、跨评测信号、趋势判断和优先整改建议,最后给高管摘要。适合你跑了好几种评测后想整体把脉或检查评估体系健康时用。注重统计证据、历史数据和明确的生产门控,不会在证据不足时轻易下结论。

▸ 展开 SKILL.md 英文原文

Use when the user has run multiple evaluation skills and wants a comprehensive analysis — maturity assessment, cross-skill signals, trends, prioritized actions, and an executive summary. Also use when the user mentions eval health check, evaluation audit, ship readiness, evaluation maturity, or "how good is my evaluation system itself." This is a read-only analysis skill.

研究检索评估报告成熟度评估优先级建议通用
748
Stars
61
Forks
18
仓库内 Skill
+40
7 日增星
安装 / 使用
给你的 Agent 一句话(通用)
帮我安装这个 skill:https://raw.githubusercontent.com/agentscope-ai/OpenJudge/main/skills/eval_pipeline/04-eval-report/SKILL.md
或 curl 直取 SKILL.md
curl -fsSL "https://raw.githubusercontent.com/agentscope-ai/OpenJudge/main/skills/eval_pipeline/04-eval-report/SKILL.md"
SKILL.MD 节选查看完整文件 ↗
<HARD-GATE>
NO recommendation WITHOUT statistical evidence backing it.
NO "system ready" declaration WITHOUT all calibrated judges passing AND all production gates green.
NO trend analysis WITHOUT at least 2 data points in history.
</HARD-GATE>

# Eval Report

Synthesize everything from your evaluation journey into a comprehensive report.
This skill is read-only — it analyzes what exists, doesn't create new graders or datasets.

## When to Activate

- You've run 2+ evaluation skills and want the big picture
- You need to report evaluation status to non-technical stakeholders
- You're making a ship/no-ship decision and need evidence
- The evaluation system has been running for a while — time 
via SKILL·HOT · 数据来自 GitHub 公开信息 · 原文版权归作者所有