Head to head
Observatory vs Agent Eval
Same numbers we show everywhere else, lined up. Ratings come from users, the grade comes from our scanner, adoption comes from public registries.
No reviews yet
MCP security scanner. CI-native testing, attack simulation, health scoring, and SARIF.
Full listingNo reviews yet
Statistical regression testing for LLM agents: p-value, effect size, and CI on behavior change.
Full listingThe short version
Neither has been reviewed yet, so the comparison rests on the scan and the public usage numbers. The safety scan favours Agent Eval (85/100 against 77). Observatory has noticeably more public adoption.
| Observatory | Agent Eval | |
|---|---|---|
| User rating | No reviews yet | No reviews yet |
| Safety grade | B77/100 | A85/100 |
| Adoption | Established | Growing |
| GitHub stars | 139 | 0 |
| Downloads / week | 383 | 59 |
| Runs where | Your machine (local) | Your machine (local) |
| Gateway-ready | No | No |
| Tools | — | 1 |
| Write actions | — | 1 |
| Price per call | Free | Free |
| Licence | MIT | Apache-2.0 |
| Last release | 2026-09-22 | 2026-09-18 |
| Category | Developer tools | Developer tools |