Turning a governed data estate into a revenue line, without copies, connectors or compromise
The monetization gap
Ask most data leaders whether their organizations own data that someone else would pay for, and the answer is yes. Ask whether they are monetizing it, and the answer is usually “we tried.”
The gap is rarely the data. It is the delivery. Traditional monetization means building an extract pipeline for every customer, negotiating a bespoke transfer mechanism, shipping files to an SFTP site or maintaining an API layer nobody wants to own. Every new consumer adds engineering cost. Every copy adds compliance exposure. By the time the data lands it is already stale, and legal reviews stretch across quarters because nobody can prove what the recipient will actually be able to see. The unit economics collapse before the first invoice is raised.
Databricks has spent five years dismantling precisely these constraints, and the capability is now a connected stack rather than a single feature. Unity Catalog governs the estate, including data still sitting in Snowflake, AWS Glue or on-prem systems, and defines what a shareable asset is. OpenSharing, the open protocol that evolved from Delta Sharing, delivers those assets live to any recipient on any platform without copying them. Databricks Marketplace supplies discovery, listings and commerce, publicly or through private exchanges. Clean Rooms cover the cases where parties must compute together but nobody can expose raw records. Taken together, they turn launching a data product from an engineering program into a configuration decision.
Figure 1: The Databricks data monetization stack — from federated sources through Unity Catalog to external consumers.
Start with the catalog, not the pipeline
Unity Catalog is the control plane for the entire estate, covering tables, notebooks, models, and AI assets, with lineage, row- and column-level controls, and an audit trail on every access. Monetization starts here, because you cannot sell what you cannot govern or prove.
Importantly, you do not need to consolidate first. Lakehouse Federation and foreign catalog support allow Unity Catalog to govern data that still sits in Snowflake, AWS Glue, Hive Metastore or on-prem systems, and the Databricks Storage Ecosystem extends the same governance to on-prem and edge estates through partners such as MinIO and VAST Data. Product definition stops waiting on migration.
Distribute without copying: OpenSharing
Delta Sharing pioneered zero-copy sharing in 2021 and now serves more than 28,000 recipients. In June 2026 it evolved into OpenSharing, a Linux Foundation project and the first open, vendor-neutral protocol for sharing data and AI assets across any cloud, vendor or format. Three properties matter commercially:
• Zero copy. Recipients query live data in your storage. One source of truth, no reconciliation, no stale replicas, no re-delivery cost.
• No platform lock-in. Consumers read from Snowflake, Power BI, Tableau, Spark, Trino or any Iceberg-compatible client, authenticating by token or OIDC (OpenID Connect). They never need a Databricks account, so your addressable market is the whole market rather than one vendor’s installed base.
• Governance travels with the asset. Unity Catalog enforces permissions and audits every access, which is what shortens a six-month legal review to a few weeks.
Operational friction has been engineered out as well. SecureConnect replaces per-recipient firewall and IP allowlisting work with a managed proxy configured once, and Global Distribution replicates across regions and clouds so recipients read locally.
What OpenSharing actually costs
A monetization business case needs a cost line, and this one is unusually simple: there is no separate license or SKU for OpenSharing. It is part of the platform, and because nothing is duplicated you never pay to store a second copy per customer. Charges arise in three places:
• Compute. Databricks compute is billed on the provider side when sharing views, materialized views and streaming tables. Plain Delta and Iceberg tables are read straight from storage by the recipient, who pays for their own query compute.
• Storage. You are already paying for it. Zero-copy means the shared asset adds no incremental storage footprint.
• Egress. Sharing within a region incurs no egress cost at all. Cross-region and cross-cloud transfer is billed by your cloud vendor, or by Databricks if you use SecureConnect.
Egress is the one variable worth engineering, and there are well-trodden levers: Global Distribution and Deep Clone replicas place data in the recipient’s region, change data feed limits transfer to incremental updates rather than full reads, and Cloudflare R2 storage carries no egress fees at all. Critically, the OpenSharing Egress Pipeline notebooks on Marketplace attribute egress bytes and cost by share and by recipient. That is the difference between a platform cost and a product P&L: you can see gross margin per customer and price accordingly. See the Marketplace consumer billing documentation for details.
Sell it: Marketplace and private exchanges
A share is a delivery mechanism, not a business. Databricks Marketplace adds discovery and commerce. Publish a listing free or paid, either publicly to thousands of Databricks customers or restricted to a private exchange visible only to invited consumers. Private exchanges are the workhorse of B2B monetization: tailored datasets for named accounts, supplier networks, franchisees. Marketplace Commit Drawdown lets buyers pay using existing Databricks commitments, removing the procurement delay that quietly kills most data deals.
For a paid or commercialized offering, the common pattern is:
- Customer finds the listing in Marketplace.
- Customer requests access from their Databricks workspace.
- If the listing requires approval, the provider reviews the request in the Provider Console and contacts the customer using the provided email address.
- The commercial terms, contract, procurement and billing are handled outside Databricks Marketplace between provider and consumer.
- Once the agreement is complete, the provider can create and share the actual asset and approve access for the customer.
Collaborate where you cannot share: Clean Rooms
Some of your most valuable data can never leave the perimeter. Clean Rooms create an isolated environment where multiple parties contribute assets via OpenSharing and run approved notebook code across the combined set. Collaborators see column names and types, never rows. No raw data is exposed, and only agreed outputs are released. This is what enables retailer-brand measurement, insurer-provider risk analysis, and identity resolution against partners such as LiveRamp or Acxiom: revenue earned from insight rather than from handing over records.
Monetize AI assets, not just tables
The most commercially significant shift is the newest one. OpenSharing was deliberately extended beyond tables and files, because what buyers increasingly want is context and capability rather than raw rows. Four asset classes are now shareable under the same governance model:
• AI models. Models registered in MLflow and governed in Unity Catalog can be shared directly, letting you license a trained propensity, risk or forecasting model instead of the data behind it.
• Agent Skills. Packaged domain workflows and specialized knowledge that a partner's agents can call, a genuinely new product category.
• Unstructured data. A FILE type lets managed Delta and Iceberg tables govern PDFs, images, audio and video, so curated document corpora can be sold as retrieval-ready inputs for a customer’s own AI applications.
• Genie Agents. A governed conversational analytics experience, including semantic model, business metrics and curated logic, that is shared instead of the dataset.
Genie Agent Sharing is where the pricing model changes. Providers can hide proprietary instructions, restrict access to the agent alone, cap row exports and set daily prompt quotas. That makes usage-based pricing viable in place of an all-or-nothing data license: a lower entry price, a wider pool of buyers, higher realized value per asset and intellectual property that never leaves your control. It also raises switching costs, because the customer is buying your interpretation of the domain, not a file they can replicate.
OpenSharing vs. API-based approach
Finally, here is a point of view of using the Databricks OpenSharing vs. a traditional API-based approach.
|
Dimension |
Delta Sharing (OpenSharing) |
API-Based Approach |
POV / Recommendation |
|
Data Freshness & Latency |
Consumers query live tables directly (near real-time, no copy/ETL lag) |
Freshness depends on how often APIs/caches are refreshed; often batch-driven |
Choose Delta Sharing when consumers need near-real-time or large-volume analytical access |
|
Scalability & Cost |
No data duplication; consumer compute handles processing (pay-as-you-query model); avoids building and maintaining ETL pipelines to serve data out |
Requires dedicated API infra, caching layers, rate-limiting and scaling for concurrent consumers |
Delta Sharing shifts compute cost to consumer side, which is better for large datasets; APIs are better for lightweight, controlled payloads |
|
Governance & Control |
Fine-grained access via Unity Catalog, but less control over how data is used downstream (raw table access) |
Full control over exposed fields, business logic, rate limits and usage patterns per consumer |
If strict control over data shape/logic/exposure is critical (e.g., regulated fields), APIs win; if trust boundary is high, Delta Sharing is simpler |
|
Integration Effort / Ecosystem Fit |
Works out-of-box with any Delta Sharing client (Spark, Pandas, Power BI, etc.), with no custom code needed on either side |
Requires building, documenting, versioning and maintaining API contracts (OpenAPI specs, SDKs) |
Delta Sharing reduces engineering overhead significantly for analytics-native consumers; APIs are better when consumers are application/service layers, not analytics tools |
|
Use Case Fit |
Best for bulk data exchange, cross-org analytics, BI/reporting, data products |
Best for transactional, low-latency, application-embedded, or business-logic-heavy access patterns |
Use Delta Sharing for "data-as-a-product" analytics sharing; use APIs when serving operational/app-level consumption |
How Tredence helps
The technology is now the straightforward part. Getting it right, though, is less a platform question than a strategy one. It means deciding what to sell, to whom and at what price, and being confident the answer holds up commercially and legally. That's where Tredence works, as an advisor to your strategy, data and legal stakeholders, with your own platform teams owning the build.
We start with a monetization assessment. This includes an inventory of candidate assets, demand sizing with your commercial leadership, and a scored opportunity portfolio ranked on value, risk and effort. From the shortlist we shape the data product definition, including what the asset is, the semantic contract and quality SLAs it must meet, and how it should be packaged and priced. Alongside this we recommend a target governance and sharing model, covering which assets belong in a public listing versus a private exchange or a Clean Room, what the Unity Catalog entitlement and audit posture should look like, and where cost will accrue so margin can be modeled per recipient before anything is committed.
We then hand your teams an implementation blueprint and business case, and stay engaged as a review partner. This means validating design decisions, tracking realized value against the plan, and advising on the operating model and commercial governance that a data business needs once the first customers are live.
The infrastructure argument for data monetization is settled. What remains is a product and commercial question, and that is a far better problem to have.
LinkedIn