advanced-evaluation-v2
仓库创建 2026年3月26日最近提交 19 天前SkillHot 收录 20 天前
▸ 精选理由
适合需要系统化评测、生成评价标准和减偏的研究/工程团队。
这个 Skill 做什么
提供 LLM 作为评判者的评价流程,用于比较与评分模型输出。
把 LLM 拉来当评判者,帮你搭一套自动化的评测流程:写评分 rubrics、做模型输出的逐对比(pairwise comparison)和直接打分。适合要比较不同模型结果、实现 LLM-as-judge 或想降低评测偏差时用。特别会保留上游工作流、支持文件和溯源信息,便于合并或交接。
▸ 展开 SKILL.md 英文原文
Advanced Evaluation workflow skill. Use this skill when the user needs This skill should be used when the user asks to "implement LLM-as-judge", "compare model outputs", "create evaluation rubrics", "mitigate evaluation bias", or mentions direct scoring, pairwise comparison, position bias, evaluation pipelines, or automated quality assessment and the operator should preserve the upstream workflow, copied support files, and provenance before merging or handing off.
96
Stars
23
Forks
40
仓库内 Skill
+25
7 日增星
安装 / 使用
给你的 Agent 一句话(通用)
帮我安装这个 skill:https://raw.githubusercontent.com/diegosouzapw/awesome-omni-skills/main/skills_omni/advanced-evaluation-v2/SKILL.md或 curl 直取 SKILL.md
curl -fsSL "https://raw.githubusercontent.com/diegosouzapw/awesome-omni-skills/main/skills_omni/advanced-evaluation-v2/SKILL.md"SKILL.MD 节选查看完整文件 ↗
# Advanced Evaluation ## Overview This public intake copy packages `plugins/antigravity-awesome-skills/skills/advanced-evaluation` from `https://github.com/sickn33/antigravity-awesome-skills` into the native Omni Skills editorial shape without hiding its origin. Use it when the operator needs the upstream workflow, support files, and repository context to stay intact while the public validator and private enhancer continue their normal downstream flow. This intake keeps the copied upstream files intact and uses `metadata.json` plus `ORIGIN.md` as the provenance anchor for review. # Advanced Evaluation This skill covers production-grade techniques for evaluating LLM outputs using LLMs as
via SKILL·HOT · 数据来自 GitHub 公开信息 · 原文版权归作者所有