Skip to content
Wide title banner: 'Composable CDP in 2026 — the data platform absorbed the CDP'. Isometric illustration of a data lake/warehouse platform with composable identity and activation blocks and orbiting AI agents.

Composable CDP in 2026: The Data Platform Absorbed the CDP

On 16 June 2026, Databricks shipped a CDP. Not a connector or a reference architecture — a native customer data platform called CustomerLake, built inside the lakehouse, with identity resolution and campaign orchestration run by AI agents. The detail that matters isn’t the product. It’s the launch-partner list. Adobe is on it. So is Twilio, which owns Segment. The two names you’d expect to compete hardest against a lakehouse-native CDP instead lined up to build on top of one.

For a decade the CDP was a place you copied your data to: ingest everything into a proprietary store, resolve identity there, assemble audiences there, sync them out. CustomerLake is the clearest signal yet that this model is being absorbed into the data platform. The CDP is becoming a set of composable capabilities that run on top of the customer data infrastructure you already own — and, increasingly, capabilities an agent operates rather than a marketer clicking through a UI. This is what “composable CDP” now means in practice, and it is no longer the challenger position. It’s the direction the whole market is moving in.

This piece maps that shift: what CustomerLake actually tells us, how Salesforce, Adobe, Oracle, Segment and mParticle are each responding, and — if you’re about to spend money on CDP technology this year — what the convergence should change about the decision.

CustomerLake is the tell, not just another product launch

Strip away the marketing and CustomerLake is a straightforward architectural statement. It runs inside Databricks. Customer data is not copied into a separate CDP database; it stays in the lakehouse and is governed by Unity Catalog, the same catalog governing everything else in the account. Identity resolution (“Profile Agents” doing what Databricks calls Agentic Identity Resolution — deterministic, probabilistic and agentic matching combined) happens against data in place — the same move toward AI-driven identity resolution that’s spreading across the category. Activation and orchestration (“Campaign Agents”) run from the same substrate. Through Lakehouse Federation it can read across Databricks, Snowflake, BigQuery and operational databases without moving the data at all.

That is the composable-CDP thesis made concrete by the vendor with the most to gain from it. The interesting part is the concession baked into the partner list. A lakehouse-native CDP could, in principle, try to replace the activation tools entirely. Instead the launch shipped with Adobe, Braze, Iterable, Bloomreach, LiveRamp, The Trade Desk, Meta and Twilio as partners. The message is that the lake becomes the profile store and the system of record, and the specialist tools compose on top of it. The data platform swallows the data layer of the CDP; the activation and channel layer stays a market.

If you only read one signal from 2026, read that one. The argument about whether customer data belongs in the warehouse or in a dedicated CDP store is over, and the warehouse won. What’s left to decide is which capabilities you compose on top and who supplies them.

The packaged incumbents are re-platforming onto the warehouse, not defending their database

Table of six CDP vendors and their 2026 direction: Databricks CustomerLake, Salesforce Data Cloud, Adobe Real-Time CDP and Segment all keep customer data in the warehouse; mParticle is partial; Oracle Unity, folded into Fusion CX, does not.
How the major vendors are moving in 2026 — every camp re-platforming around the customer’s warehouse.

The tell is stronger because the incumbents were already moving this way before CustomerLake made it obvious.

Salesforce Data Cloud — now the centre of its Data 360 story — leads with Zero Copy. Rather than ingesting your Snowflake, BigQuery or Databricks tables, it federates against them: the data stays in the warehouse and Data Cloud queries it in place, with clean rooms for collaboration on top. A few years ago Salesforce’s pitch was “bring your data into Salesforce.” Today the flagship feature is explicitly about not copying it. That is a packaged incumbent conceding that the customer’s warehouse, not Salesforce’s storage, is the substrate.

Adobe has done the same through Federated Audience Composition. Real-Time CDP will now compose audiences directly against data in Snowflake, Databricks and BigQuery without importing it first — Adobe’s own framing is “composable by design.” An organisation can build and activate a segment in Adobe while the underlying customer records never leave the warehouse. And Adobe is a CustomerLake launch partner, which means it is comfortable being the activation layer on someone else’s lake.

Oracle is the honest exception, and it is worth naming because the convergence is not uniform. Oracle folded Unity into its Fusion Cloud CX applications rather than pushing a warehouse-composable CDP; its centre of gravity is keeping customer data inside the Oracle apps suite, not federating against a neutral lake. If your stack is Oracle end to end that is coherent, but it is the one major direction that runs against the grain of everything else here — a reminder that “everyone is going composable” is a trend, not a law.

The collection layer is repositioning around the warehouse too

If the incumbents are conceding the store, the data-collection vendors are conceding the flow. Segment built its business on being the pipe you route events into — SDKs in, unified profile in Segment, activations out. Under Twilio it has spent the last few cycles adding warehouse-native capabilities: reading profile data from the warehouse rather than only holding it in Segment, reverse-ETL-style activation, a data graph that treats the warehouse as a first-class source. The pipe is becoming warehouse-first rather than warehouse-optional. Twilio being a CustomerLake partner is the same admission in a different register.

mParticle, now part of Rokt, has moved along the same axis — repositioning from a self-contained profile store toward feeding and reading the warehouse and leaning on AI-driven activation. The direction across the collection layer is consistent: stop asking customers to treat your platform as the source of truth, and instead be a well-behaved citizen around the warehouse that already is one.

None of these moves are charity. They happen because buyers stopped wanting a second copy of their customer data living in a vendor’s account, with the cost, latency and governance drift that a duplicate implies. The vendors followed the buyers.

It all converges on one architecture

Two-panel architecture diagram: a legacy packaged CDP copies web, app and CRM data into a proprietary store that holds identity resolution and audiences; the converged model leaves customer data in the lake or warehouse as one governed copy, with an identity service and activation layer composing on top and agents running across it.
Legacy CDP versus the converged model — copy your data in, or compose on the data you already hold.

Put the moves side by side and they describe a single shape. Data stays resident in the lake or warehouse. Identity resolution runs as a service against that data rather than inside a separate database. Activation and channel tools compose on top, pulling the resolved profile when they need it. Governance is the catalog you already run, not a second permissions model. And the newest layer — agents — sits across the top as the interface that assembles audiences and runs campaigns, in place of a human navigating screens.

That is worth seeing as a picture, because the contrast with the legacy model is the whole argument: the old CDP was a box you copied everything into and operated as its own system; the emerging one is a set of capabilities layered on the platform you already own. “Composable CDP” and “warehouse-native CDP” are two names for this same shape — the point is not the label but that the store, the identity layer and the activation layer have come apart into pieces you assemble, rather than one product you buy whole.

Where it heads next: agentic profiles, zero-copy collaboration, real-time

Three vectors are already visible in the 2026 announcements. The first is agentic operation: CustomerLake’s Profile and Campaign Agents are the leading edge, but the pattern — agents resolving identity and running “always-on” campaigns against the lake, the agentic CDP direction — is where the category is going, not a Databricks-only feature. The second is zero-copy collaboration: Salesforce clean rooms, warehouse-native data sharing and the broader move to match and enrich customer data across parties without any of them copying raw records. The third is real-time: agents act at machine speed, so the batch-oriented plumbing under most CDPs becomes the bottleneck, and streaming resolution and activation move up the priority list.

Each of these only works if the data is already in the platform. That’s the through-line: every forward-looking capability in the category now assumes the warehouse-resident, composable base rather than a copied-in store.

What this means if you’re buying a CDP right now

Decision diagram for CDP buyers: ask whether a CDP copies your data or composes on your warehouse. Copying is the side the market is leaving, justified only without a warehouse or data-engineering team; composing aligns with the trend — check federation, profile-store portability and agent-readiness.
The one question to ask any CDP you’re evaluating in 2026.

If you’re making a fresh CDP investment in 2026, the convergence changes the question you’re answering. The old question was “which CDP has the best profile store and audience builder.” The new one is “which activation and agent layer composes best on the data platform I already run” — because the store is increasingly your warehouse, whichever vendor’s logo is on the UI.

Practically, that means a few things. Anchor the decision on where your data already lives; if you have a warehouse or lakehouse, a CDP whose core value is holding a second copy of your data is buying you the losing side of this architecture. Evaluate any packaged CDP as a composable layer — its UI, its activation reach, its agents — sitting on top of your warehouse, and ask specifically whether it federates (Salesforce Zero Copy, Adobe Federated Audience Composition) or still insists on ingesting. Avoid locking your resolved profile inside a proprietary store you can’t easily leave; portability of the identity layer is now a first-order procurement question, not a footnote. Insist that whatever you buy is agent-ready and governed by a catalog you control, because that’s the direction every roadmap here is pointing.

And be honest about your own data-engineering capacity, because it decides how much of this you assemble versus buy as a managed layer — the composable model asks more of the team operating it, and for a lean team without warehouse maturity a packaged CDP is still the faster, safer call. Finally, mind the timing: CustomerLake is in private preview, and several of these federated capabilities are newer than their marketing implies. The direction is safe to bet on; specific 2026 products are not all shippable yet. That argues against signing a long “copy everything into us” contract now, and for choosing tools that keep your options open as the composable stack matures.

Conclusion

The debate that defined the CDP category — dedicated store versus warehouse — is effectively settled, and CustomerLake is the moment it became undeniable: when Databricks ships a CDP and Adobe and Twilio show up as partners, the architecture question has an answer. The customer data platform is dissolving into composable capabilities on the data platform you already own, with agents as the new operating layer.

That doesn’t make the buying decision easier, but it changes its shape. The job now is to invest on the right side of the shift rather than against it. On Monday, take whatever CDP you’re currently evaluating and ask one question before any feature comparison: does it copy your data, or does it compose on it? The answer tells you whether you’re buying into the architecture the whole market is converging on, or the one it’s leaving behind.

Take the next step

If you’re weighing a fresh CDP investment, the sharpest place to start is the cost and architecture comparison — what you actually pay, and what you actually own, under a packaged store versus a warehouse-native, composable approach. Our packaged-versus-warehouse-native CDP evaluation walks through the pricing mechanics and the ownership questions to put to any vendor before you sign.

Frequently Asked Questions

Do I still need a CDP if I already have a data warehouse or lakehouse?

You still need the capabilities a CDP provides — identity resolution, audience building, activation — but increasingly not as a separate system that copies your data. In 2026 the common pattern is to run those capabilities on top of the warehouse you already have, whether through a warehouse-native tool or a packaged CDP that federates rather than ingests. The question has shifted from “buy a CDP or not” to “which layer composes on my existing data platform.”

Is a composable CDP actually cheaper than a packaged one?

Not automatically. Composable removes the cost of duplicating and storing customer profiles in a second system and avoids per-profile SaaS pricing, but it moves work onto your own data-engineering team to assemble and operate the stack. For an organisation with a mature warehouse the economics usually favour composable; for a lean team without that foundation, a packaged CDP can be cheaper once you count the engineering time.

Is Databricks CustomerLake a real CDP or just an announcement?

It’s a real product, announced in June 2026 and available in private preview — a native CDP inside the Databricks lakehouse with agent-driven identity resolution and campaign orchestration. Private preview means it isn’t broadly shippable yet, so treat it as a strong signal of direction rather than something you can deploy at scale today. Its significance is as much about who partnered on it (Adobe, Twilio and others) as about the feature list.

Are “composable CDP” and “warehouse-native CDP” the same thing?

They describe the same architectural shape from slightly different angles. “Warehouse-native” emphasises that customer data stays in your warehouse or lakehouse; “composable” emphasises that the CDP is assembled from separate parts — storage, identity, activation — rather than bought as one product. In practice a composable CDP is warehouse-native, and the terms are often used interchangeably.

Will Salesforce and Adobe stop selling their packaged CDPs?

There’s no sign they’ll stop, but both have re-platformed around the warehouse: Salesforce Data Cloud leads with Zero Copy federation and Adobe Real-Time CDP with Federated Audience Composition, both of which let customer data stay in the warehouse instead of being ingested. The packaged product persists as a UI, activation and governance layer — increasingly one that composes on your data platform rather than replacing it.

What does an “agentic CDP” actually mean?

It’s a CDP where AI agents perform the work a person or a fixed pipeline used to do — resolving identities, assembling audiences, and running and adjusting campaigns continuously — rather than a marketer configuring each step by hand. CustomerLake’s Profile Agents and Campaign Agents are the current example. The practical implication is that the data has to be accessible, governed and real-time enough for an agent to act on safely, which is another reason the warehouse-resident model is winning.