AI for LLM evaluation
Benchmarks, eval suites, and quality gates.
| # | Tool | Category | Pricing | Visit |
|---|---|---|---|---|
| 1 | Promptfoo Open-source LLM testing and red-teaming — test prompts, evaluate outputs, catch regressions | LLM Frameworks | Free | Visit |
| 2 | RAGAS RAG evaluation framework — measure faithfulness, relevance, and recall of RAG pipelines | LLM Frameworks | Free | Visit |
| 3 | DeepEval Open-source LLM evaluation framework — 14+ metrics for RAG, agents, and chatbots | LLM Frameworks | Free | Visit |
| 4 | TruLens LLM app evaluation and monitoring — measure quality and detect hallucinations in production | LLM Frameworks | Free | Visit |
| 5 | Braintrust Enterprise AI evaluation platform — experiments, datasets, and production monitoring | LLM Frameworks | Freemium | Visit |
| 6 | Langfuse Open-source LLM observability — tracing, evaluation, and prompt management | LLM Frameworks | Free | Visit |
| 7 | LangSmith Debug, test, and monitor LLM applications built with LangChain or any framework | LLM Frameworks | Freemium | Visit |