domonic18/ai-eval-scope

这是一个基于rule-base和llm-as-judge的评测系统,包含评估器、执行器以及Web可视化页面。

1forks 0PythonGitHub

EvalScope (ai-eval-scope)

English | 简体中文

CI PyPI License: MIT Python 3.11+ Code Style: ruff

A new-generation, agent-driven evaluation system that turns Agent evaluation into a repeatable, quantifiable pipeline:

describe what to test → generate an "exam paper" (scenario package) → the Agent under test takes it for real
→ rule engine + LLM judge scoring → report → platform trends

👇 Watch a 60-second intro — conversational scenario-package creation in the CLI:

Conversational scenario package creation in the CLI

Why

LLM Agents (support assistants, slide generators, coding helpers, ...) are non-deterministic: the same prompt yields a different answer every time, so classic "replay script + expected == actual" testing breaks down. EvalScope's answer: intelligent execution, deterministic evaluation.

Traditional test automationEvalScope
ExecutionScripted replay; breaks on any environment driftExecutionAgent understands the task, decides autonomously, recovers from blockers
ScoringHard-coded assertionsRule engine (hard checks) + LLM Judge (semantic quality)
Case maintenanceHand-maintained scriptsScenario packages (YAML) — generated conversationally, version-controlled
TargetDeterministic UI / API systemsNon-deterministic LLM / Agent systems

Use cases

  • Conversational Agent QA — safety (correctly refusing risky prompts), task completion quality, multi-turn coherence
  • RAG retrieval QA — right and complete results, no hallucination (trap questions), traceable delivery
  • Generation quality — slide decks, code generation, and other multi-dimension quality combos
  • Version regression — one consistent yardstick over time: "did this version actually get better?"

Courseware / code / chat scenario packages ship built-in; a new scenario (RAG, custom) is just a YAML package — no code changes.

Features

  • If you can chat, you can use it — describe what you want to test in plain language; the Workbench Agent drafts the full asset set (exam paper, scoring rules, judge prompts, SUT config), and every file lands on disk only after you review its diff
  • Automatic SUT onboarding — give it an entry URL (even just a login page); it analyzes the login API, performs a real login, and probes the conversation API; credentials are hidden-input and never written to any config file
  • Layered scoring — format gate → rule engine → LLM judge → vision screenshot evaluation (Playwright); every score comes with its reasons
  • Online execution — ExecutionAgent (DeepAgents-based) drives the Agent under test: multi-turn dialogue, blocker recovery, budget guardrails
  • Observability platform — a Langfuse-style multi-tenant platform: runs, trends, metrics, per-sample conversation traces

Quickstart: your first evaluation in 3 steps

Requires Python 3.11+.

# ① Install ([agent] execution engine and [llm] AI judging recommended; uv or pip)
uv tool install "ai-eval-scope[agent,llm]"
agent-eval --version

# ② Configure the judging model (pick a provider → paste API key → auto connectivity test)
agent-eval models set

# ③ Open the CLI workbench and just describe the task: create package → probe SUT → run evaluation → read report
agent-eval start

In the workbench choose "Workbench Agent", describe the requirement in one sentence (e.g. "run a safety evaluation on this online Agent") plus the SUT entry URL, and the rest is automated. Or run the whole chain from the command line (execute → evaluate → report → upload):

agent-eval pipeline --package chat --task "safety_*" --upload

CLI pipeline run and result summary

Full command reference (run / pipeline / suite / models / secrets / package / dataset …), --task selection syntax, scenario-package layout and its five asset kinds, FAQ — see the CLI Tutorial (Chinese).

Observability platform

A Langfuse-style multi-tenant platform ships in web/: projects and API keys, scenario-aware metric rendering, per-sample conversation traces with scoring details, cross-run trends, and webhooks. Once the CLI is connected (agent-eval auth login), every evaluation is reported automatically.

cp .env.example .env     # fill in DB / object storage / secrets
make docker-up           # starts postgres + minio + web (http://localhost:9000)

Platform run detail: scenario metrics + summary report + per-sample scores

Documentation

DocumentContents
CLI Tutorial (Chinese)Full command handbook, package layout & five asset kinds, courseware/Agent evaluation walkthroughs, CI integration
Architecture Overview (Chinese)evaluator / web / executor domains and data flow
Evaluation Engine (Chinese)Gates, rules, LLM judge, vision evaluation, metric aggregation
Data & Configuration (Chinese)Package composition, dataset abstraction, workspace layout
Third-party Integration (Chinese)HTTP API / Webhook / MCP integration guide
ContributingCommit / branch / PR conventions, dev environment

Contributing

Issues and pull requests are welcome! See CONTRIBUTING.md for the development workflow, commit conventions, and PR process.

Acknowledgements

Dataset download and dataset-index design reference OpenCompass; HuggingFace / ModelScope source metadata for some datasets (evaluator/agent_eval/assets/datasets/dataset_index.yaml) is ported from OpenCompass. Thanks to the OpenCompass team for their excellent open-source work.

License

MIT

进展动态

  1. Merge pull request #4 from domonic18/develop@诸葛东明commit
  2. Merge branch 'chore/release-v0.5.0' into 'develop' (merge request !106)@domoniccommit
  3. bump: version 0.4.2 → 0.5.0@诸葛东明commit
  4. Merge branch 'feat/judge-num-samples-default-1' into 'develop' (merge request !105)@domoniccommit
  5. Merge branch 'fix/platform-llm-timeout' into 'develop' (merge request !104)@domoniccommit
  6. feat(config): 判官采样次数默认与课件包模板统一为 1@诸葛东明commit
  7. fix(config): 平台 LLM 角色配置透传 timeout_sec 到 ProviderConfig@诸葛东明commit
  8. Merge branch 'fix/ruff-format-test-orchestrator' into 'develop' (merge request !103)@domoniccommit
  9. style(tests): 修正 test_orchestrator 换行格式(ruff format 门禁)@诸葛东明commit
  10. Merge branch 'feat/eval-perf-concurrency' into 'develop' (merge request !102)@domoniccommit
  11. docs: 评测并发参数文档同步(采样并发/批间并发/--concurrency)@诸葛东明commit
  12. perf(orchestrator): 样本级有界并发评估(--concurrency)+ 缓存命中深拷贝@诸葛东明commit
  13. perf(evaluation): fact_verdict 分批复核批间并发(params 可调)@诸葛东明commit
  14. perf(judge): 判官多次采样并发化(StabilityController 有界线程池)@诸葛东明commit
  15. Merge branch 'feat/jev-filter' into 'develop' (merge request !101)@domoniccommit
  16. fix(evaluation): info_accuracy reason 文案如实区分判定专线剔除与 LLM 二次确认@诸葛东明commit
  17. fix(cli): models set 遇不可读旧配置不再带栈崩溃,改备份重建或退出@诸葛东明commit
  18. fix(evaluation): resolved 模型快照示例补全厂商前缀,过收尾 grep 验收@诸葛东明commit
  19. refactor(scripts): 对拍 PoC 改名 poc_decision_alignment 并参数化@诸葛东明commit
  20. docs: decision 判定专线文档同步@诸葛东明commit