Skip to main content

AI for LLM evaluation

Benchmarks, eval suites, and quality gates.

#ToolCategoryPricingVisit
1Promptfoo

Open-source LLM testing and red-teaming — test prompts, evaluate outputs, catch regressions

LLM FrameworksFreeVisit
2RAGAS

RAG evaluation framework — measure faithfulness, relevance, and recall of RAG pipelines

LLM FrameworksFreeVisit
3DeepEval

Open-source LLM evaluation framework — 14+ metrics for RAG, agents, and chatbots

LLM FrameworksFreeVisit
4TruLens

LLM app evaluation and monitoring — measure quality and detect hallucinations in production

LLM FrameworksFreeVisit
5Braintrust

Enterprise AI evaluation platform — experiments, datasets, and production monitoring

LLM FrameworksFreemiumVisit
6Langfuse

Open-source LLM observability — tracing, evaluation, and prompt management

LLM FrameworksFreeVisit
7LangSmith

Debug, test, and monitor LLM applications built with LangChain or any framework

LLM FrameworksFreemiumVisit