CIOPages
DirectoryAI & ML PlatformsAI Governance & SafetyDeepEval

DeepEval

Open Source

About DeepEval

DeepEval provides evaluation tools that run in CI/CD pipelines or as Python scripts, allowing iteration on custom criteria in local environments.

Key Capabilities

  • Unit-testing framework for large language models
  • Native integration with Pytest for CI workflows
  • Support for multi-modal AI evaluation
  • Automated prompt optimization and synthetic data generation
  • 50+ research-backed evaluation metrics including G-Eval

Integrations

OpenAILangChainAnthropic

Related Buyer Guides

Independent evaluation frameworks for this category.

AI Agent & Agentic AI Platforms
Compare LangGraph, CrewAI, Microsoft Agent Framework, OpenAI Agents SDK, Google ADK, AWS Bedrock AgentCore, LlamaIndex, and Temporal — where production operability, not the slickest multi-agent demo, is the deciding criterion.
AI Governance & Responsible AI
Evaluate IBM watsonx.governance, Credo AI, Microsoft Purview, ServiceNow, Holistic AI, Fiddler, Arthur, and Monitaur — and decide first whether your gap is governance and compliance or ML observability, because the EU AI Act and NIST AI RMF reward the platform wired into how models are built and run.
Computer Vision & Visual AI
Evaluate Google Vision/Vertex AI, AWS Rekognition, Azure AI Vision, Landing AI, Roboflow, Cognex, Encord, and Ultralytics — with the pretrained-API vs. custom-model and cloud vs. edge decisions, not a generic feature list, as the deciding criteria.

This profile was compiled by CIOPages from public sources with AI assistance, and may be incomplete or out of date. It is informational only and not an endorsement. Represent this vendor? Claim this listing or .

Quick Facts

www.deepeval.com
CategoryAI & ML Platforms
SubcategoryAI Governance & Safety
PricingSubscription
DeploymentOpen Source, Cloud
Target SizeEnterprise