3 questions from the package
From the RFI round. The first shows part of the guide each question carries; the workbook adds follow-ups, how to verify the answer, a priority and a weight.
1. Complete a connector matrix for [source system list] giving each source one support status (native and generally available, native in preview, built and supported by a third party, or not supported), with a link to the documentation page of every native connector.
Why it matters. Connector lists on data sheets can mix generally available, preview and partner-built connectors. If the buyer finds the gaps after contract, it has to load metadata by hand, and those loads go stale.
Good answer
- Every row has exactly one status, with no blank rows.
- Supported source versions or editions are named for each native connector.
- Third-party connectors name the party that builds and supports them.
Red flags
- A generic database driver connection is listed as a native connector for a specific source.
- Rows are marked "roadmap" or "on request" with no release status.
- Supported source versions are not stated.
2. For each source in [source system list], state the finest lineage granularity your product derives automatically: column-level, table-level only, or none.
Why it matters. Lineage coverage often varies by source. A product that delivers column-level lineage for the warehouse but only table-level lineage for ETL or BI tools leaves gaps the buyer must fill by hand.
3. For each source type in [source system list], state whether your automated classification runs on table columns, on fields nested inside semi-structured columns, on unstructured files, or not at all.
Why it matters. Sensitive data in a source the classifier cannot read stays unlabeled. The buyer's inventory of personal data then has gaps that do not show in the catalog.
Capability areas
Connectivity & Metadata Harvesting (11)
Native connectors and scanners for warehouses, lakehouses, BI tools, ETL/ELT tools, SaaS applications, and on-prem or mainframe sources. Covers harvest depth, incremental refresh, freshness, schema-change detection, and connector maintenance. Out: generic integration APIs and SSO, which belong to the integration module.
Lineage & Impact Analysis (12)
Automated table- and column-level lineage parsed from SQL, stored procedures, ETL/ELT jobs, and BI semantic layers. Covers lineage across transformations and system boundaries, gaps and manual stitching, lineage versioning, and upstream and downstream impact analysis. Out: lineage into AI models, covered in AIG.
Classification & Sensitive-Data Detection (11)
Automated classification of columns and files, PII and sensitive-data detection, custom classifiers, sampling and profiling behavior, and how stewards review, correct, and tune classification results. Out: the policies and enforcement that act on classifications, covered in POL.
Active Metadata, Automation & Metadata APIs (9)
Auto-generated descriptions and suggested glossary terms, metadata events that trigger actions such as alerts, tagging, or tickets, and programmatic read and write access to the metadata model for pipelines and agents. Out: generic API availability and SDK quality, covered by the integration module.
Business Glossary & Metadata Model (10)
Business glossary structure, term hierarchies and relationships, linking terms to physical assets, domains and communities, custom asset types and attributes, and how the metamodel fits our operating model. Out: approval and task routing, covered in STW.
Stewardship Workflow & Ownership (11)
Ownership assignment, steward task queues, approval and certification workflows, change proposals from consumers, escalation, and adoption and curation metrics for the governance program. Out: end-user search experience, covered in DIS.
Search, Discovery & Data Products (11)
Search and browse for business users, trust signals shown on assets, data-product publishing, marketplace and access-request experience, and catalog context surfaced inside BI and query tools. Out: the enforcement of access once granted, covered in POL.
Policy, Privacy & Access Enforcement (12)
Policy authoring tied to classifications and glossary terms, masking and row- and column-level controls, push-down enforcement into warehouse and lakehouse platforms, drift between declared and enforced policy, and audit evidence for regulators. Out: the vendor's own handling of our data, covered by the data-protection-privacy module.
Data Quality & Observability (10)
Embedded or integrated data-quality rules, profiling, scoring, anomaly and freshness monitoring, and how quality results appear on catalog assets and along lineage. Out: master data management as a separate discipline.
Governance for AI (10)
Registers for AI models, agents, and use cases; lineage from source data through training, fine-tuning, and inference; mapping of registered systems to regulatory and risk frameworks; and serving governed metadata as context to AI consumers. Out: governance of the vendor's own built-in AI features, covered by the AI modules.
Platform Coexistence & Scale (12)
Interoperability with native warehouse, lakehouse, and open table-format catalogs, including bidirectional sync of tags, ownership, and policy and conflict handling. Also covers performance at large asset counts and the operational effort to keep scanners and metadata current. Out: hosting model and regions, covered by the deployment-hosting module.
Questions about this package
How many Data Governance & Catalog Platforms RFP questions are there?
119 solution questions in 11 capability areas: 23 for the RFI, 61 for the RFP and 35 deep-dive questions for the finalists. The workbook adds 80 due-diligence questions on security, integration, implementation and exit.
What comes with each question?
Why it matters, good-answer signals, red flags, follow-up questions, how to verify the answer (a demo step, a test or a document), and a suggested priority and weight for scoring.
Can I edit the questions?
Yes. The workbook is an ordinary Excel file. Change, add or remove questions, and change the weights; the scorecard recalculates.
Which license do I need?
The Enterprise License covers any number of evaluations inside one organization. The Consultancy License covers use with any number of clients. Neither allows reselling or republishing the questions.