agent-evaluation
仓库创建 2026年1月15日最近提交 3 小时前SkillHot 收录 3 小时前
▸ 精选理由
帮助量化代理弱点并建立回归与稳定性监测体系。
这个 Skill 做什么
为LLM代理提供行为测试、能力评估与生产可靠性监控的工具集。
给 LLM 代理做一套可量化的测试和基准评估:覆盖行为测试、能力评估、可靠性指标、回归测试和生产监控。适合在把 agent 投产前后做基准、找弱点或持续监控表现时使用。特别之处是把真实世界任务抽成可测指标,帮助发现上线后才会暴露的可靠性问题。
▸ 展开 SKILL.md 英文原文
Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics, and production monitoring—where even top agents achieve less than 50% on real-world benchmarks
4.4w
Stars
6.5k
Forks
40
仓库内 Skill
积累中
7 日增星
安装 / 使用
给你的 Agent 一句话(通用)
帮我安装这个 skill:https://raw.githubusercontent.com/sickn33/agentic-awesome-skills/main/skills/agent-evaluation/SKILL.md或 curl 直取 SKILL.md
curl -fsSL "https://raw.githubusercontent.com/sickn33/agentic-awesome-skills/main/skills/agent-evaluation/SKILL.md"SKILL.MD 节选查看完整文件 ↗
# Agent Evaluation Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics, and production monitoring—where even top agents achieve less than 50% on real-world benchmarks ## Capabilities - agent-testing - benchmark-design - capability-assessment - reliability-metrics - regression-testing ## Prerequisites - Knowledge: Testing methodologies, Statistical analysis basics, LLM behavior patterns - Skills_recommended: autonomous-agents, multi-agent-orchestration - Required skills: testing-fundamentals, llm-fundamentals ## Scope - Does_not_cover: Model training evaluation (loss, perplexity), Fairness and bias testing, User experience testin
via SKILL·HOT · 数据来自 GitHub 公开信息 · 原文版权归作者所有