Skip to content
AI 도구

AI Evaluation Tools

AI evaluation tools는 AI 애플리케이션의 품질, 신뢰성, 안전성 및 성능을 테스트하고 측정하는 데 사용되는 프레임워크이자 플랫폼입니다.

shipped 2026년 8월 21일codefreemium
Domain rating95
code

핵심 포인트

1ai
2code
3product-hunt

overview

Overview

AI evaluation tools는 AI 애플리케이션의 품질, 신뢰성, 안전성 및 성능을 테스트하고 측정하는 데 사용되는 프레임워크이자 플랫폼입니다.

유사한 도구

대안 비교

고려해 볼 만한 다른 도구

1

DeepEval

Open-source LLM evaluation framework for unit testing LLM outputs, agents, RAG pipelines, and prompts with various metrics.

방문
2

Giskard

Open-source Python library for testing and evaluating LLM agents and models, offering automated checks, LLM judges, and red-teaming.

방문
3

LangSmith

Evaluation platform for LLM and AI agents, integrated with LangChain, supporting offline/online evals, human feedback, and prompt optimization.

Stork에서 보기
4

Arize AI

AI observability and evaluation platform with comprehensive LLM evaluation capabilities for production systems and agents.

Stork에서 보기

Stork에서 더 보기

관련 AI 도구

같은 카테고리의 다른 도구 — 공통 태그로 연결