AI for LLM evaluation
Benchmark prompts, models, and agent outputs with datasets.
| # | Tool | Category | Pricing | Visit |
|---|---|---|---|---|
| 1 | Braintrust Enterprise AI evaluation platform — experiments, datasets, and production monitoring | LLM Frameworks | Freemium | Visit |
| 2 | Langfuse Open-source LLM observability — tracing, evaluation, and prompt management | LLM Frameworks | Free | Visit |
| 3 | Humanloop LLM evaluation and prompt management — improve and deploy AI features in production | LLM Frameworks | Paid | Visit |
| 4 | TruLens LLM app evaluation and monitoring — measure quality and detect hallucinations in production | LLM Frameworks | Free | Visit |
| 5 | Arize Phoenix Open-source AI observability | LLM Frameworks | Free | Visit |