CIOPages
DirectoryData & AnalyticsData Governance & CatalogAmundsen (Lyft)

Amundsen (Lyft)

Open Source

About Amundsen (Lyft)

Amundsen is an open source data catalog designed to improve data discovery, metadata management, and governance within large enterprises. It enables data analysts, data scientists, and data engineers to efficiently find, understand, and trust data across the organization by providing a centralized platform for metadata aggregation and contextual insights. The platform leverages automated and curated metadata, including table descriptions, usage statistics, and data previews, to build confidence in data quality and relevance.

Built for enterprise-scale environments, Amundsen helps break down data silos and enhances collaboration by allowing users to share context, update metadata, and learn from others' data usage patterns. Its search capabilities are powered by a PageRank-inspired algorithm that recommends relevant data assets based on activity and metadata. The solution supports easy integration and deployment on Docker, EC2, and Kubernetes, making it adaptable to diverse infrastructure needs. Amundsen’s primary value lies in boosting productivity by reducing manual documentation efforts and enabling faster debugging and data pipeline trustworthiness.

How to evaluate Data Governance & Catalog

This is how the CIOPages Research Team evaluates this category. It is not an assessment of Amundsen (Lyft). The category covers two kinds of product, so both frameworks are here. Most buyers need one of them.

25%
Connector Coverage & Metadata Ingestion
Native connectors for your actual sources — warehouses, lakehouses, BI tools, ETL/ELT, streaming, NoSQL, files, and the awkward homegrown and SaaS systems; automated scanning and incremental crawl; depth of harvested metadata (schemas, profiles, usage); and how hard a custom connector is to build
20%
Lineage & Impact Analysis
Automated column-level lineage parsed from SQL, dbt, Spark, and orchestration; accuracy and freshness on YOUR pipelines, not a demo dataset; cross-system lineage that survives a hop between platforms; and upstream/downstream impact analysis a team would trust before a breaking change
20%
Active Metadata & AI Automation
Usage- and popularity-driven ranking from query logs; ML-assisted description generation, PII/PHI classification, and anomaly or staleness flags; metadata that writes back into BI tools, IDEs, and chat; and the ability to serve trusted context to retrieval/agentic systems (e.g. MCP)
15%
Governance, Stewardship & Compliance
Stewardship and certification workflow, business glossary tied to physical assets, policy and access-control integration, data-contract support, and audit-ready classification and lineage for GDPR, CCPA, HIPAA, and sector regimes — without forcing a separate governance suite
10%
Adoption, Collaboration & UX
Search relevance and natural-language query for non-technical users, embedded experiences in Slack/Teams, BI tools and IDEs, crowdsourced enrichment and certification badges, and the friction of getting a first-time business user to a trusted answer
10%
Architecture, Extensibility & Deployment
Open and extensible metadata model, API and event coverage for custom workflows, SSO/RBAC and fine-grained metadata access, deployment options (SaaS, customer-managed, on-prem), and total operational burden including upgrades and scaling
25%
Catalog, Lineage & Connectivity
Depth and freshness of automated metadata harvesting; column-level, end-to-end lineage parsed from SQL/ETL/BI (not hand-drawn); connector coverage across your warehouses, lakehouse, BI, SaaS, and on-prem sources; and how lineage survives transformations
20%
Active Metadata & Automation
AI-assisted classification and PII detection, auto-suggested glossary terms and descriptions, metadata that triggers actions (alerts, policy, tagging) rather than sitting passive, and open APIs that let agents and pipelines read and write metadata
20%
Stewardship & Business Usability
Business glossary and data-product workflows, steward task/approval routing, search and discovery non-technical users actually adopt, marketplace/shopping experience, and time-to-first-value for a real domain
15%
Policy, Privacy & Access Enforcement
Policy authoring tied to classifications, masking/row- and column-level controls, push-down enforcement into Snowflake/Databricks/BigQuery, audit trails, and alignment to GDPR, CCPA, and sector regimes
10%
Governance for AI & Data Quality
AI/model and agent registers, lineage from source data through training and inference, EU AI Act / NIST AI RMF mapping, and embedded or integrated data-quality rules, scoring, and observability
10%
Deployment, Scale & Ecosystem Fit
SaaS vs. self-managed options, scale across millions of assets, interoperability with lakehouse catalogs (Unity, Polaris/Horizon, Iceberg REST), identity/SSO, and total operational burden to keep it current

Related Buyer Guides

Our buyer guides across Data & Analytics. Each one compares the main vendors in its category and what buyers weigh up.

AI/ML Platforms
Compare Databricks Mosaic AI, AWS SageMaker, Azure Machine Learning, Google Vertex AI, Snowflake Cortex, Dataiku, DataRobot, and Weights & Biases on the question this category actually turns on — getting governed models into production and keeping them healthy, not the accuracy of a one-off notebook.
Business Intelligence & Analytics
Evaluate Power BI, Tableau, Qlik, Looker, ThoughtSpot, Sigma, Amazon QuickSight, Strategy, SAP Analytics Cloud, and Domo on the question that decides BI value — whether self-service freedom and a governed semantic layer can coexist, not whose charts look best.
Cloud Data Warehouse
Compare Snowflake, Databricks, BigQuery, Redshift, and Synapse across performance benchmarks, pricing models, ecosystem integrations, and governance capabilities for enterprise analytics workloads.

CIOPages put this listing together from public sources. It’s information, not an endorsement. How we build listings. Work here? Claim this listing or .

Quick Facts

www.amundsen.io
CategoryData & Analytics
SubcategoryData Governance & Catalog
FoundedNot on file
HeadquartersNot on file

We publish a detail only when we can point at the page it came from. Claim this listing to fill in the rest.