agent-evaluation-v2

仓库创建 2026年3月26日最近提交 19 天前SkillHot 收录 20 天前
▸ 精选理由

保留上游完整工作流与支持文件,便于复现和迁移测试流程。

这个 Skill 做什么

用于对 LLM 代理进行行为测试与基准评估的工作流包。

为 LLM 代理做行为测试与基准评测的工作流包:设计测试场景、执行行为验证并产出评价指标与报告。适合需要严格验证代理表现、比较不同模型或替入生产前的质量把关时用。重点是保留上游工作流和支持文件、保全来源与可追溯性,便于合并或移交。

▸ 展开 SKILL.md 英文原文

Agent Evaluation workflow skill. Use this skill when the user needs Testing and benchmarking LLM agents including behavioral testing, and the operator should preserve the upstream workflow, copied support files, and provenance before merging or handing off.

研究检索代理评估基准测试行为测试通用
96
Stars
23
Forks
40
仓库内 Skill
+25
7 日增星
安装 / 使用
给你的 Agent 一句话(通用)
帮我安装这个 skill:https://raw.githubusercontent.com/diegosouzapw/awesome-omni-skills/main/skills_omni/agent-evaluation-v2/SKILL.md
或 curl 直取 SKILL.md
curl -fsSL "https://raw.githubusercontent.com/diegosouzapw/awesome-omni-skills/main/skills_omni/agent-evaluation-v2/SKILL.md"
SKILL.MD 节选查看完整文件 ↗
# Agent Evaluation

## Overview

This public intake copy packages `plugins/antigravity-awesome-skills/skills/agent-evaluation` from `https://github.com/sickn33/antigravity-awesome-skills` into the native Omni Skills editorial shape without hiding its origin.

Use it when the operator needs the upstream workflow, support files, and repository context to stay intact while the public validator and private enhancer continue their normal downstream flow.

This intake keeps the copied upstream files intact and uses `metadata.json` plus `ORIGIN.md` as the provenance anchor for review.

# Agent Evaluation Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliab
via SKILL·HOT · 数据来自 GitHub 公开信息 · 原文版权归作者所有