Data Automation Framework.
Modern data platforms rely on a growing ecosystem of automation technologies. However, effective automation is about more than selecting the right tool. Rather than focusing on individual tools, our Data Automation Framework provides a structured approach for evaluating, comparing, and designing automation strategies across modern data platforms.
A Structured Approach to Automating Modern Data Platforms
Bringing Three Dimensions Together
Every data automation initiative can be evaluated through the same three questions:
Scope
What can be automated? Which data platform capabilities are in scope - modeling, movement, pipelines, testing, or platform automation?
Automation
How is the automation designed? Is the implementation manual, developer-assisted, metadata-driven, rule-driven, knowledge-driven, or autonomous?
Implementation
How is the automation implemented? Which cloud-native services, reusable components, or automation platforms deliver the solution?
Together, these dimensions provide a technology-agnostic framework for evaluating and designing automation across modern data platforms
Our framework focuses on automating the core capabilities required to build and manage modern data platforms. It applies regardless of whether the target is a:
Data Warehouse
Lakehouse
Data Lake
Data Fabric
Data Mesh
Data Products Platform
AI-ready Data Platform
The framework focuses on building the data platform. Semantic models, analytics, decision support, and AI consume the platform and are therefore outside the scope of the framework.
The Scope.
What can be automated?
The Scope dimension identifies the core data platform capabilities to which automation can be applied.
Data Modeling Automation
Defines the data structures and organization of the data platform, including logical and physical data models, Data Vault, dimensional models, and Medallion architectures, together with business entities, relationships, keys, business rules, and metadata.
Data Movement Automation
Automates connectivity, data ingestion - including batch, Change Data Capture, and streaming - and integration from operational, cloud, and external data sources.
Pipeline & Transformation Automation
Automates ELT and ETL pipelines, data transformation, and orchestration.
Test & Data Quality Automation
Automates data testing, validation, data quality, and monitoring across modern data platforms.
Data Platform Automation
Automates the design, generation, deployment, governance, documentation generation, lineage, and change management of data warehouses, lakehouses, data lakes, and data products.
Representative automation technologies
Different tools automate different scopes. Some specialize in a single capability, while broader platforms may span several areas.
| Automation Scope | Representative Tools |
|---|---|
| Data Modeling Automation | SAP PowerDesigner, erwin Data Modeler, ER/Studio, WhereScape 3D, VaultSpeed, Ellie.ai, Hackolade |
| Data Movement Automation | Fivetran, Airbyte, Azure Data Factory, Microsoft Fabric Data Factory, Informatica Cloud, Kafka Connect, Qlik Replicate |
| Pipeline & Transformation Automation | dbt, Coalesce, Matillion, Apache Airflow, Dagster, Azure Data Factory, Microsoft Fabric Data Factory, Informatica Cloud |
| Test & Data Quality Automation | BiG EVAL, Soda, Monte Carlo, Bigeye, dbt Tests |
| Data Platform Automation | WhereScape, VaultSpeed, TimeXtender, DataVault Builder |
This mapping is illustrative rather than exhaustive. It shows where technologies contribute within the framework rather than positioning every product as a direct competitor.
The Design.
How is the automation designed?
The Design dimension describes the pattern that drives the automation.
It is not an organizational maturity model. The levels describe how the automation itself is designed—from human-driven implementation to increasingly metadata-, rule-, knowledge-, and AI-driven automation.
| Level | Design Principle | Description |
| Level 0 | Manual | Human-driven implementation with little or no automation. |
| Level 1 | Developer-Assisted | Productivity tools, such as modeling or ETL tools, assist developers, but implementation remains largely manual. |
| Level 2 | Metadata-Driven | Metadata drives code generation and automation. |
| Level 3 | Rule-Driven | Reusable rules, patterns, and standards automate implementation decisions. |
| Level 4 | Knowledge-Driven | Business knowledge drives automation through semantic models, catalogs, business glossaries, and governance policies. The automation understands business context rather than relying only on technical metadata. |
| Level 5 | Autonomous | AI-assisted and AI-driven automation designs, builds, tests, optimizes, and operates data solutions. |
Applying the Design Principles
The same design principles can be applied across each automation scope.
Example: Designing a Data Pipeline
| Level | Example |
| Manual | Hand-written SQL and pipeline logic |
| Developer-Assisted | Visual ETL designer |
| Metadata-Driven | Metadata-driven pipeline generation |
| Rule-Driven | Reusable templates and governed pipeline patterns |
| Knowledge-Driven | Pipeline generation informed by catalog metadata and enterprise knowledge |
| Autonomous | AI generates, tests, and optimizes pipelines autonomously |
Example: Designing a Data Vault Model
| Level | Example |
| Manual | Whiteboard or diagram-based design |
| Developer-Assisted | Manual modeling in a data modeling tool |
| Metadata-Driven | Models generated or derived from metadata |
| Rule-Driven | Data Vault structures generated using predefined modeling rules |
| Knowledge-Driven | Models generated from business glossaries, catalogs, and enterprise semantics |
| Autonomous | AI creates and evolves models from business requirements |
AI-Assisted Automation Across the Framework
AI does not become a separate automation scope. It can enhance every existing scope.
| Automation Scope | AI-Assisted Automation |
| Data Modeling Automation | AI generates Data Vault or dimensional models from requirements. |
| Data Movement Automation | AI generates ingestion mappings and connector configurations. |
| Pipeline & Transformation Automation | AI generates ELT pipelines, SQL, orchestration, and mappings. |
| Test & Data Quality Automation | AI generates test cases and quality rules and detects anomalies. |
| Data Platform Automation | AI generates platform objects and documentation and recommends implementation optimizations. |
The Implementation.
How is the automation implemented?
Once the automation scope and design principle have been defined, the next decision is how the automation will be delivered.
The framework identifies three common implementation approaches.
Cloud-Native
Automation is built using native cloud platform services and capabilities.
Examples may include platform-native ingestion, transformation, orchestration, deployment, security, and monitoring services.
Component-Based
Automation is assembled from reusable components, connectors, templates, services, and accelerators.
This approach reduces repeated development while retaining flexibility over how individual components are combined.
Model-Driven Data Warehouse Automation
Automation is delivered through Data Warehouse Automation platforms that generate and manage data solutions from models, metadata, rules, and reusable patterns.
This approach is particularly relevant when organizations want to standardize the design, generation, deployment, documentation, and ongoing management of data warehouses and broader analytical data platforms.
The implementation approaches are not necessarily mutually exclusive. A solution may combine cloud-native services, reusable components, and a Data Warehouse Automation platform.
From Tools to an Automation Strategy
Most automation discussions begin with technology. The Data Automation Framework begins with the data platform capabilities that need to be automated. It then examines how those capabilities should be automated and which implementation approach is most appropriate. This enables architects and engineering teams to:
Evaluate automation opportunities consistently
Compare technologies based on the capabilities they automate
Identify gaps and overlaps in the technology landscape
Select an appropriate automation design principle
Combine implementation approaches where required
Develop a coherent, automation-first data platform strategy
The objective is not simply to automate individual tasks. It is to create a consistent approach for designing, implementing, and evolving automation across the modern data platform.