agent-evaluation
仓库创建 2026年1月15日最近提交 21 天前SkillHot 收录 20 天前
▸ 精选理由
帮助量化代理性能与回归监控,便于持续改进与部署决策。
这个 Skill 做什么
为 LLM 代理提供行为测试、能力评估与生产监控的基准工具。
为 LLM 代理做行为测试和基准评估,包含能力测评、可靠性指标、回归测试和生产监控,帮你量化代理在真实场景的表现。上线前做能力验证、回归或上线后监控都能用,能捕捉性能退化和高风险行为。特点是把主观表现转成可测指标,便于持续改进和设置告警。
▸ 展开 SKILL.md 英文原文
Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics, and production monitoring—where even top agents achieve less than 50% on real-world benchmarks
4.2w
Stars
6.8k
Forks
40
仓库内 Skill
+0
7 日增星
安装 / 使用
给你的 Agent 一句话(通用)
帮我安装这个 skill:https://raw.githubusercontent.com/sickn33/antigravity-awesome-skills/main/skills/agent-evaluation/SKILL.md或 curl 直取 SKILL.md
curl -fsSL "https://raw.githubusercontent.com/sickn33/antigravity-awesome-skills/main/skills/agent-evaluation/SKILL.md"SKILL.MD 节选查看完整文件 ↗
# Agent Evaluation Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics, and production monitoring—where even top agents achieve less than 50% on real-world benchmarks ## Capabilities - agent-testing - benchmark-design - capability-assessment - reliability-metrics - regression-testing ## Prerequisites - Knowledge: Testing methodologies, Statistical analysis basics, LLM behavior patterns - Skills_recommended: autonomous-agents, multi-agent-orchestration - Required skills: testing-fundamentals, llm-fundamentals ## Scope - Does_not_cover: Model training evaluation (loss, perplexity), Fairness and bias testing, User experience testin
via SKILL·HOT · 数据来自 GitHub 公开信息 · 原文版权归作者所有