Skip to main content

AI for LLM evaluation

Benchmark prompts, models, and agent outputs with datasets.

#ToolCategoryPricingVisit
1Braintrust

Enterprise AI evaluation platform — experiments, datasets, and production monitoring

LLM FrameworksFreemiumVisit
2Langfuse

Open-source LLM observability — tracing, evaluation, and prompt management

LLM FrameworksFreeVisit
3Humanloop

LLM evaluation and prompt management — improve and deploy AI features in production

LLM FrameworksPaidVisit
4TruLens

LLM app evaluation and monitoring — measure quality and detect hallucinations in production

LLM FrameworksFreeVisit
5Arize Phoenix

Open-source AI observability

LLM FrameworksFreeVisit

See also