advanced-evaluation

仓库创建 2026年3月26日最近提交 19 天前SkillHot 收录 20 天前
▸ 精选理由

能快速搭建自动化评测管线,适合模型迭代验证。

这个 Skill 做什么

实现 LLM 驱动的输出比较与评估工作流,支持配对比较与打分。

帮你把不同模型或提示的输出用 LLM 当裁判做打分和配对比较,能自动产出评估报告和评分标准。适合要做模型比较、搭建评价流水线或想降低评估偏见时用。特别能处理 pairwise comparison 并保存原始工作流与溯源,方便复现和合并结果。

▸ 展开 SKILL.md 英文原文

Advanced Evaluation workflow skill. Use this skill when the user needs This skill should be used when the user asks to "implement LLM-as-judge", "compare model outputs", "create evaluation rubrics", "mitigate evaluation bias", or mentions direct scoring, pairwise comparison, position bias, evaluation pipelines, or automated quality assessment and the operator should preserve the upstream workflow, copied support files, and provenance before merging or handing off.

研究检索模型评估对比评测评分通用
96
Stars
23
Forks
40
仓库内 Skill
+25
7 日增星
安装 / 使用
给你的 Agent 一句话(通用)
帮我安装这个 skill:https://raw.githubusercontent.com/diegosouzapw/awesome-omni-skills/main/skills_omni/advanced-evaluation/SKILL.md
或 curl 直取 SKILL.md
curl -fsSL "https://raw.githubusercontent.com/diegosouzapw/awesome-omni-skills/main/skills_omni/advanced-evaluation/SKILL.md"
SKILL.MD 节选查看完整文件 ↗
# Advanced Evaluation

## Overview

This public intake copy packages `plugins/antigravity-awesome-skills-claude/skills/advanced-evaluation` from `https://github.com/sickn33/antigravity-awesome-skills` into the native Omni Skills editorial shape without hiding its origin.

Use it when the operator needs the upstream workflow, support files, and repository context to stay intact while the public validator and private enhancer continue their normal downstream flow.

This intake keeps the copied upstream files intact and uses `metadata.json` plus `ORIGIN.md` as the provenance anchor for review.

# Advanced Evaluation This skill covers production-grade techniques for evaluating LLM outputs using 
via SKILL·HOT · 数据来自 GitHub 公开信息 · 原文版权归作者所有