HoneyHive
Free tierAgent Observability and Evaluation Platform
Free tier available·All audiences·API available
Key strengths
Agent tracing and observabilityOnline and offline evaluationsMonitoring with drift detection and alertsExperiment and regression testingAnnotation queues for human reviewPrompt management and versioningCI/CD integrationEnterprise self-hosting options
Free tier + paid plans
Self-hostable
No ratings yet
Documentation & Integration
HoneyHive exposes an API for programmatic access to tracing, evaluation, and prompt management functionality. Instrumentation is built on OpenTelemetry, making it straightforward to integrate with existing observability stacks.
CI/CD & testing integrations:
- GitHub Actions — Embed evaluation gates directly into your deployment pipelines.
- PyTest — Write evaluation assertions as standard Python tests.
- Vite — Frontend tooling integration for full-stack AI application workflows.
- Claude Code (MCP) — Integration with Anthropic's Claude Code via the Model Context Protocol.
Evaluation types supported:
- Automated evaluations (LLM-as-judge, heuristic, custom scorers)
- Human evaluations via annotation queues
Free tier includes:
- 10,000 events/month
- Up to 5 users
- 30-day data retention
- Automated & human evaluations
- Prompt versioning
- CI/CD integration
- Community support
Enterprise tier adds:
- Custom usage limits & unlimited users/workspaces
- SAML & custom SSO
- Custom data retention
- Uptime SLA
- Custom DPA and BAA
- Dedicated TAM & quarterly business reviews (QBRs)
- Self-hosting, hybrid, and single-tenant deployment options
