domonic18/ai-eval-scope
这是一个基于rule-base和llm-as-judge的评测系统,包含评估器、执行器以及Web可视化页面。
EvalScope (ai-eval-scope)
A new-generation, agent-driven evaluation system that turns Agent evaluation into a repeatable, quantifiable pipeline:
describe what to test → generate an "exam paper" (scenario package) → the Agent under test takes it for real
→ rule engine + LLM judge scoring → report → platform trends
👇 Watch a 60-second intro — conversational scenario-package creation in the CLI:

Why
LLM Agents (support assistants, slide generators, coding helpers, ...) are non-deterministic: the same prompt yields a different answer every time, so classic "replay script + expected == actual" testing breaks down. EvalScope's answer: intelligent execution, deterministic evaluation.
| Traditional test automation | EvalScope | |
|---|---|---|
| Execution | Scripted replay; breaks on any environment drift | ExecutionAgent understands the task, decides autonomously, recovers from blockers |
| Scoring | Hard-coded assertions | Rule engine (hard checks) + LLM Judge (semantic quality) |
| Case maintenance | Hand-maintained scripts | Scenario packages (YAML) — generated conversationally, version-controlled |
| Target | Deterministic UI / API systems | Non-deterministic LLM / Agent systems |
Use cases
- Conversational Agent QA — safety (correctly refusing risky prompts), task completion quality, multi-turn coherence
- RAG retrieval QA — right and complete results, no hallucination (trap questions), traceable delivery
- Generation quality — slide decks, code generation, and other multi-dimension quality combos
- Version regression — one consistent yardstick over time: "did this version actually get better?"
Courseware / code / chat scenario packages ship built-in; a new scenario (RAG, custom) is just a YAML package — no code changes.
Features
- If you can chat, you can use it — describe what you want to test in plain language; the Workbench Agent drafts the full asset set (exam paper, scoring rules, judge prompts, SUT config), and every file lands on disk only after you review its diff
- Automatic SUT onboarding — give it an entry URL (even just a login page); it analyzes the login API, performs a real login, and probes the conversation API; credentials are hidden-input and never written to any config file
- Layered scoring — format gate → rule engine → LLM judge → vision screenshot evaluation (Playwright); every score comes with its reasons
- Online execution — ExecutionAgent (DeepAgents-based) drives the Agent under test: multi-turn dialogue, blocker recovery, budget guardrails
- Observability platform — a Langfuse-style multi-tenant platform: runs, trends, metrics, per-sample conversation traces
Quickstart: your first evaluation in 3 steps
Requires Python 3.11+.
# ① Install ([agent] execution engine and [llm] AI judging recommended; uv or pip)
uv tool install "ai-eval-scope[agent,llm]"
agent-eval --version
# ② Configure the judging model (pick a provider → paste API key → auto connectivity test)
agent-eval models set
# ③ Open the CLI workbench and just describe the task: create package → probe SUT → run evaluation → read report
agent-eval start
In the workbench choose "Workbench Agent", describe the requirement in one sentence (e.g. "run a safety evaluation on this online Agent") plus the SUT entry URL, and the rest is automated. Or run the whole chain from the command line (execute → evaluate → report → upload):
agent-eval pipeline --package chat --task "safety_*" --upload

Full command reference (run / pipeline / suite / models / secrets / package / dataset …),
--taskselection syntax, scenario-package layout and its five asset kinds, FAQ — see the CLI Tutorial (Chinese).
Observability platform
A Langfuse-style multi-tenant platform ships in web/: projects and API keys, scenario-aware metric rendering, per-sample conversation traces with scoring details, cross-run trends, and webhooks. Once the CLI is connected (agent-eval auth login), every evaluation is reported automatically.
cp .env.example .env # fill in DB / object storage / secrets
make docker-up # starts postgres + minio + web (http://localhost:9000)

Documentation
| Document | Contents |
|---|---|
| CLI Tutorial (Chinese) | Full command handbook, package layout & five asset kinds, courseware/Agent evaluation walkthroughs, CI integration |
| Architecture Overview (Chinese) | evaluator / web / executor domains and data flow |
| Evaluation Engine (Chinese) | Gates, rules, LLM judge, vision evaluation, metric aggregation |
| Data & Configuration (Chinese) | Package composition, dataset abstraction, workspace layout |
| Third-party Integration (Chinese) | HTTP API / Webhook / MCP integration guide |
| Contributing | Commit / branch / PR conventions, dev environment |
Contributing
Issues and pull requests are welcome! See CONTRIBUTING.md for the development workflow, commit conventions, and PR process.
Acknowledgements
Dataset download and dataset-index design reference OpenCompass; HuggingFace / ModelScope source metadata for some datasets (evaluator/agent_eval/assets/datasets/dataset_index.yaml) is ported from OpenCompass. Thanks to the OpenCompass team for their excellent open-source work.
License
进展动态
- Merge pull request #4 from domonic18/develop@诸葛东明commit
- Merge branch 'chore/release-v0.5.0' into 'develop' (merge request !106)@domoniccommit
- bump: version 0.4.2 → 0.5.0@诸葛东明commit
- Merge branch 'feat/judge-num-samples-default-1' into 'develop' (merge request !105)@domoniccommit
- Merge branch 'fix/platform-llm-timeout' into 'develop' (merge request !104)@domoniccommit
- feat(config): 判官采样次数默认与课件包模板统一为 1@诸葛东明commit
- fix(config): 平台 LLM 角色配置透传 timeout_sec 到 ProviderConfig@诸葛东明commit
- Merge branch 'fix/ruff-format-test-orchestrator' into 'develop' (merge request !103)@domoniccommit
- style(tests): 修正 test_orchestrator 换行格式(ruff format 门禁)@诸葛东明commit
- Merge branch 'feat/eval-perf-concurrency' into 'develop' (merge request !102)@domoniccommit
- docs: 评测并发参数文档同步(采样并发/批间并发/--concurrency)@诸葛东明commit
- perf(orchestrator): 样本级有界并发评估(--concurrency)+ 缓存命中深拷贝@诸葛东明commit
- perf(evaluation): fact_verdict 分批复核批间并发(params 可调)@诸葛东明commit
- perf(judge): 判官多次采样并发化(StabilityController 有界线程池)@诸葛东明commit
- Merge branch 'feat/jev-filter' into 'develop' (merge request !101)@domoniccommit
- fix(evaluation): info_accuracy reason 文案如实区分判定专线剔除与 LLM 二次确认@诸葛东明commit
- fix(cli): models set 遇不可读旧配置不再带栈崩溃,改备份重建或退出@诸葛东明commit
- fix(evaluation): resolved 模型快照示例补全厂商前缀,过收尾 grep 验收@诸葛东明commit
- refactor(scripts): 对拍 PoC 改名 poc_decision_alignment 并参数化@诸葛东明commit
- docs: decision 判定专线文档同步@诸葛东明commit