agent-evaluation-v3
仓库创建 2026年3月26日最近提交 19 天前SkillHot 收录 20 天前
▸ 精选理由
适合需要保留并移植上游评测流程和证据的场景。
这个 Skill 做什么
面向 LLM 代理的测试与基准评估工作流,含支持文件与来源信息。
提供面向 LLM 代理的测试与基准评估工作流,包含行为测试、度量定义以及保留上游工作流和支持文件的要求。用在需要对 agent 做系统性评测、复现实验或在合并/移交前保留溯源信息时。特别强调保留来源与可复查的证据链,方便审计与后续复现。
▸ 展开 SKILL.md 英文原文
Agent Evaluation workflow skill. Use this skill when the user needs Testing and benchmarking LLM agents including behavioral testing, and the operator should preserve the upstream workflow, copied support files, and provenance before merging or handing off.
96
Stars
23
Forks
40
仓库内 Skill
+25
7 日增星
安装 / 使用
给你的 Agent 一句话(通用)
帮我安装这个 skill:https://raw.githubusercontent.com/diegosouzapw/awesome-omni-skills/main/skills_omni/agent-evaluation-v3/SKILL.md或 curl 直取 SKILL.md
curl -fsSL "https://raw.githubusercontent.com/diegosouzapw/awesome-omni-skills/main/skills_omni/agent-evaluation-v3/SKILL.md"SKILL.MD 节选查看完整文件 ↗
# Agent Evaluation ## Overview This public intake copy packages `plugins/antigravity-bundle-agent-architect/skills/agent-evaluation` from `https://github.com/sickn33/antigravity-awesome-skills` into the native Omni Skills editorial shape without hiding its origin. Use it when the operator needs the upstream workflow, support files, and repository context to stay intact while the public validator and private enhancer continue their normal downstream flow. This intake keeps the copied upstream files intact and uses the `external_source` block in `metadata.json` plus `ORIGIN.md` as the provenance anchor for review. # Agent Evaluation Testing and benchmarking LLM agents including behavioral
via SKILL·HOT · 数据来自 GitHub 公开信息 · 原文版权归作者所有