Etude
|
|
En Bref
Look closely at almost any enterprise technology function today, and you’ll see structural tension. For instance, a data platform that was modernized just a few years ago by migrating to Snowflake or BigQuery and having its business intelligence (BI) dashboards rebuilt is now being asked to do something very different: power machine learning (ML) models, feed generative AI knowledge assistants, and eventually support autonomous agents that make procurement or operational decisions without anyone in the loop. It wasn’t designed for this, and the seams are starting to show. Three forces have converged to bring this about.
Together, these shifts are stretching yesterday’s modern data warehouse well beyond its original design envelope. Integrating environments, not licensing, represents the true cost burden. For example, these overlaps could mean that finance's numbers don't match those of marketing or that the customer isn't defined consistently across the company’s financial and customer systems. Meanwhile, boards are asking when the company will have AI agents autonomously managing supply chain exceptions. The structural problem is compounded by a governance gap that’s rapidly becoming an existential risk. When a dashboard has a data quality issue, someone files a ticket. By comparison, when an autonomous agent makes procurement decisions based on bad data, the blast radius is orders of magnitude larger. Data quality, lineage, and access control are no longer limited just by compliance; they are operational safety requirements for organizations serious about AI. Underlying all of these challenges is a cultural one that technology investment alone cannot solve. Most companies don’t treat data as a shared asset or extend data literacy beyond specialist teams. Platform investments frequently deliver technical capabilities that the company isn’t ready to adopt, a pattern that derails more modernization programs than any architectural misstep. Five workloads, one platformEvery architecture decision starts with a clear-eyed assessment of what the platform must actually accomplish. Too often, organizations define the goal aspirationally, such as "we want to be an AI-native organization," instead of grounding it in current workloads and realistic expectations for the next 24 months. In practice, enterprise data platforms must support five distinct workloads, each with its own data requirements, consumer types, latency expectations, and governance profiles.
Different data architectures are good at supporting different types of workloads. And every architecture has a practical limit. Modern data warehouses now support far more than traditional reporting, extending into ML and many early GenAI use cases. lakehouse architectures, which are essentially a hybrid of warehouse and data lake capabilities, push that boundary further, enabling more complex real-time and agentic workloads. The critical question is not where the organization aspires to be but where its workloads will realistically sit over the next few years. Architecture decisions should follow workload requirements, not strategy rhetoric. Three architectures and their limitsKeeping the operating model as a separate decision, there are three distinct architectural patterns. Each represents a different set of trade-offs across workload coverage, operational complexity, talent requirements, and vendor exposure. Warehouse-centric with selective extensions: A single cloud data warehouse, such as Snowflake, BigQuery, or Redshift, serves as the center of gravity for all analytical workloads. The modern warehouse has converged substantially with lakehouse capabilities: Snowflake supports Apache Iceberg, BigQuery offers integrated Spark for ML, Redshift provides native SageMaker integration, and native vector search is emerging across all three platforms for basic GenAI use cases. These extensions have raised the ceiling meaningfully; the extended warehouse now reaches through moderate ML and initial GenAI workloads. It hits firmly at agentic AI and complex real-time requirements. This is the simplest architecture, with the broadest talent availability, fastest time to value, and lowest total cost of ownership. Adopting open table formats within the warehouse creates an on-ramp to lakehouse capabilities without committing to the full complexity from Day 1. Lakehouse as integration platform: An open lakehouse architecture converges batch and event origin data in a single governed storage layer. Its strongest advantage is ML integration: Data engineering and data science work on the same storage with shared catalogue, access controls, and lineage, creating a tighter feedback loop than architectures in which ML platforms operate on separate data copies. Because the lakehouse keeps historical versions of data, data scientists can reliably recreate the exact data sets used to train ML models, making results more auditable and repeatable. This architecture’s capabilities extend to support more complex workloads, with growing capabilities for generative AI and agentic workloads through loosely coupled AI infrastructure. The warehouse typically persists, however, as the optimized engine for instances in which large numbers of users will be running BI queries at the same time; lakehouse SQL engines have not yet matched warehouse performance for serving many users at the same time. Plan for this dual-engine pattern from the outset rather than discovering the gap halfway through implementation. Best of breed per workload: Different workloads get different engines, each optimized for its specific access pattern, connected by an event-streaming backbone and unified by a governance and semantic layer. There’s no workload that the architecture can’t support, but the trade-off is operational complexity: 6 to 10 distinct services plus the integration layer require sophisticated DevOps, cross-engine governance, and mature FinOps. Most organizations that believe they need this actually need the option above, lakehouse as integration, with better execution. This might be the right architecture for the most technically mature companies, a small minority of 5% to 10%.
Figure 1
The convergence between these architectures matters as much as their differences. For many enterprises—particularly those focused on BI, reporting, and early to moderate ML and GenAI—an extended warehouse is the pragmatic choice, offering the broadest workload coverage with the least operational complexity. Adopting open table formats (for example, Snowflake on Iceberg) creates a path to lakehouse capabilities without taking on the full complexity up front. For organizations with advanced data maturity, significant ML in production, real-time requirements, and the engineering capability to operate a multilayer platform, a lakehouse is the better fit. Even then, the warehouse often remains the preferred engine for high-concurrency BI, making a dual-engine model the likely end state. The semantic layer determines whether AI deliversAcross all three architectures, semantic infrastructure is both the highest-leverage investment and the hardest capability to retrofit. But what is commonly called "the semantic layer" is in fact three distinct capabilities, each serving different workloads and maturing on different timelines.
As foundation models commoditize, the differentiator shifts to the proprietary context an organization can provide. The depth of the ontological layer, not the choice of LLM, determines whether AI produces genuinely enterprise-specific insight and action. Organizations that build this infrastructure in the right sequence will compound its advantage over time; those that defer it will find it progressively harder to retrofit. Migration reality: Start where you are, not where you want to beFew companies are building from a green field, and the path from current to target state is messier than any architecture diagram suggests. For those with an existing cloud data warehouse, the best option is extension, not replacement. The warehouse continues to excel at BI workloads while a lakehouse layer is added for ML and AI workloads. The warehouse may never go away: It becomes the optimized BI engine within a broader architecture, continuing to operate as a component of the new lakehouse architecture. For organizations with multiple overlapping platforms, which are common after M&A, a data fabric overlay—namely, a virtualization layer that enables unified discovery and query across existing platforms—earns its keep as a bridge strategy. Deploy a virtualization layer to provide unified discovery, and query across existing platforms while building the target architecture underneath. Increasingly, capabilities such as Snowflake data sharing, Microsoft Fabric OneLake shortcuts, and the Delta Sharing protocol are bringing fabric-like access directly into the platforms, reducing dependence on standalone virtualization layers. The fabric gives users immediate cross-platform visibility without waiting for consolidation. Meanwhile, designate one platform as the strategic target, and begin routing all new workloads toward it. One principle applies regardless of starting point: Migrate workloads, not platforms. Move specific dashboards, pipelines, and ML models, not entire environments. Each workload has its own risks, stakeholders, and definition of success.
Delivering value early: The sequencing and funding imperativePlatform programs fail most often not because the architecture was wrong but because the organization ran out of patience or funding before value materialized. The governing principle is to deliver value early and incrementally so that each phase of investment is justified by the demonstrated outcomes of the previous phase. Three funding archetypes work in practice, often in combination depending on organizational content and CFO disposition.
The common failure across archetypes is building the platform before identifying the use cases that will fund it. Successful organizations start with the most valuable use case, build only what’s needed to deliver it, show the results, and use that credibility to fund the next priority. Architecture shapes what’s possible; the use case determines what gets built first. The cost of delay is not zeroThe temptation facing many CIOs is to delay platform decisions until the AI market matures. This ignores the compounding cost of inaction: Technical debt accumulates, and legacy platforms become harder to decommission. Additionally, the semantic infrastructure that every future workload depends on, from BI metrics to the ontological layer that both generative and agentic AI require, becomes harder to retrofit with every system built on inconsistent definitions. The good news is that the boundary between architecture options has blurred. The convergence between warehouse and lakehouse capabilities means organizations no longer face a binary choice between simplicity and readiness for the future. For most enterprises, extending the warehouse with open table formats and native AI capabilities is the pragmatic path. It delivers value quickly while creating an on-ramp to lakehouse capabilities as workloads mature. The organizations best positioned for the AI era are not those with the most advanced architecture today; they are those that start now, deliver value early, and build the organizational capability to evolve their platform as the demands of autonomous AI continue to unfold. |