Hybrid data storage: Skip migration costs now

Blog 15 min read

With over 2 Zettabytes of data under management, enterprise partners are rejecting total cloud migration in favor of hybrid governance. Replaced by a Hybrid Forever reality where data sovereignty and egress costs dictate architecture. This strategic pivot from data migration to universal governance addresses strict mandates like GDPR and HIPAA that legally prohibit moving specific datasets. We examine the Zero Data Movement architecture enabled by the OpenSharing Protocol, which allows organizations to train models on engineering-classified datasets without leaving on-premises environments. Finally, the discussion covers operationalizing this hybrid model using the provider and Unity Catalog to unify access across disparate storage silos.

The financial imperative is clear, as cloud egress economics make migrating massive historical archives unsustainable for global trading firms. While Databricks reported a revenue run-rate suggesting rapid adoption, the real story lies in connecting to the hundreds of billions of dollars in the Software-Set Storage market. Organizations can no longer afford to ignore the data gravity holding critical assets on local infrastructure while trying to enable AI value from dark data.

The Strategic Shift from Data Migration to Universal Governance

Defining Data Gravity and the Hybrid Forever Strategy

Data gravity pins massive datasets to on-premises hardware because the cost of moving them exceeds their cloud value. Global trading firms hold historical tick data where cloud egress fees make migration financially impossible. This economic reality forces a pivot toward governing data in place rather than shifting it. The resulting Hybrid Forever strategy lets organizations modernize infrastructure while keeping data within legal boundaries. Mandates like GDPR and HIPAA demand strict jurisdiction, rendering bulk transfers illegal for specific datasets.

Industry response now favors universal access protocols over migration pipelines. Firms deploy open standards to expose on-premises storage directly to cloud compute engines instead of re-platforming entire estates. This method connects storage endpoints to governance layers without duplicating petabytes of information. Enterprises maintain tight control while avoiding the complexity of fragmented architectures.

Centralized policy enforcement often clashes with distributed data locations. Operators require assurance that remote engines enforce rules without violating local residency laws.rabata.io provides S3-compatible object storage built for high-performance AI workloads. Models train on local data without costly replication or compliance exposure.

Applying Databricks SDS System to On-Premises AI

The Databricks Software-Set Storage (SDS) System grants direct AI access to on-premises data without migration pipelines. Serving over 12,000 customers globally, the platform scales across diverse industries. Stephen Orban stated that forcing customers to migrate massive amounts of data using complex pipelines just to enable intelligence is a broken model.

Unity Catalog acts as the centralized governance layer, extending policy enforcement across hybrid boundaries to cover non-cloud storage estates. This architecture fixes the broken model where enterprises forced data through complex pipelines simply to enable intelligence. An open-storage first philosophy allows organizations to apply Serverless Compute on data residing in existing infrastructure. Partners implement open protocol servers that expose data estates directly to the platform.

Feature Traditional Migration SDS System Approach
Data Location Cloud Object Storage On-Premises / Edge
Movement Bulk Copy Required Zero Data Movement
Governance Scope Cloud Only Hybrid Universal
Latency Impact High during transfer Native Local Access

Adopting this zero-movement architecture requires storage vendors to implement specific open protocol endpoints, creating a dependency on partner readiness. By June 2026, Sacra estimates Databricks' annualized revenue reached billions of dollars. The immediate operational value lies in avoiding egress costs for massive historical datasets. Network operators shift from managing bulk transfer bandwidth to optimizing local read throughput for AI training jobs.

Rabata.io delivers S3-compatible object storage engineered to serve as the high-performance foundation for this hybrid governance model. The platform provides the low-latency access required for AI workloads while maintaining the cost efficiency of on-premises deployment. Organizations now deploy governed AI on dark data without replicating petabytes to the cloud.

Migrate Everything vs Govern Everything: Economic Trade-offs

Data gravity anchors petabyte-scale engineering datasets on-premises, making bulk migration economically unviable for semiconductor manufacturers. Moving terabytes of historical tick data incurs cloud egress fees that destroy ROI, forcing global trading firms to reject the "lift and shift" model. Substantial pharmaceutical companies run millions of daily drug experiments against local storage to satisfy strict regulatory controls without incurring transfer latency. The traditional approach requires complex pipelines that duplicate data, whereas the governance model exposes existing assets directly to analytics engines.

Dimension Migrate Everything Govern Everything
Data Location Cloud-only replicas On-premises source
Compliance Risk High (data movement) Low (sovereign)
Time to Value Extended (pipeline build) Accelerated (catalog link)
Cost Driver Egress + Storage Compute only

Industry analysis indicates the shift to agentic AI accelerates demand for platforms supporting autonomous operations on local data. Architectures that eliminate unnecessary data copying notably reduce architectural complexity. Rabata.io resolves this tension by providing S3-compatible storage that integrates natively with governance layers, ensuring high-performance access without compromising sovereignty.

Architecture of Zero Data Movement via OpenSharing Protocol

OpenSharing Protocol Mechanics for Live On-Premises Access

the provider AIStor natively implements the OpenSharing protocol to expose on-premises Apache Iceberg and Delta tables directly to Databricks Serverless Compute. This architecture eliminates data migration by allowing the storage layer to serve live queries under Unity Catalog governance without copying bytes. Partners implement OpenSharing servers that authenticate Databricks requests, enabling immediate access to dark data silos while preserving strict sovereignty boundaries. AI and analytics initiatives often stall because data resides in environments with rigid security or operational requirements.

Feature Traditional ETL OpenSharing Approach
Data Location Cloud Object Store On-Premises Storage
Latency Impact High (Migration Time) Near-Zero (Live Access)
Governance Scope Post-Migration Immediate via Unity Catalog
Compliance Risk Elevated (Data Copy) Minimal (No Movement)

Bypassing complex pipelines accelerates AI initiatives that previously faced months of delay. Enterprises can now apply Genie and AgentBricks to on-premises datasets that legally cannot leave the facility. This model demands strong network connectivity between the on-premises OpenSharing endpoint and the cloud compute layer. Query performance depends entirely on the underlying network path. Telecommunications providers ingest enormous volumes of network telemetry on-premises daily to power AI-driven operations that cannot tolerate cloud round-trips. Storage transforms from a passive repository into an active, governed participant in the AI lifecycle. Location becomes transparent to the analyst in this unified data estate.

Extending Databricks Genie and Agent Bricks to On-Premises Data

Enterprises extend Serverless Compute, Genie, and Agent Bricks to on-premises data by deploying OpenSharing servers that expose local estates directly to the cloud. This architecture solves the broken model of complex migration pipelines by allowing Databricks to query live Apache Iceberg tables without moving bytes. Storage partners implement these servers to bridge the gap between fixed infrastructure and agentic AI workloads. The mechanism relies on the storage array acting as a governed endpoint, authenticating requests from Databricks Serverless Compute while keeping data physically stationary.

Capability Migration Model OpenSharing Model
Data Location Cloud Object Store On-Premises Array
Latency Source Network Egress Local Disk I/O
Governance Scope Post-Migration Native & Immediate

Operators must verify that their chosen partner offers a certified connector to avoid connection failures during high-concurrency training runs. Network configuration presents the primary constraint; unlike pure cloud setups, hybrid links demand precise firewall rules to allow Unity Catalog to reach the on-premises endpoint securely.

This topology addresses scenarios where data gravity or sovereignty mandates prevent cloud replication. Network operators shift from managing bulk transfer jobs to orchestrating secure, low-latency access paths. The strategy unlocks value in dark data archives that were previously too expensive or risky to migrate. Organizations can now run Agent Bricks against terabytes of historical logs residing on local disks. Immediate AI readiness arrives without the financial penalty of egress fees or the temporal cost of data copying.

Overcoming Security and Sovereignty Barriers with Native OpenSharing

Strict regulatory mandates often legally prohibit moving classified datasets from on-premises jurisdictions to public cloud environments. Financial services, healthcare, and government organizations operate under mandates like GDPR, HIPAA, and NIS2 that require data to remain within specific jurisdictions or air-gapped environments. The OpenSharing protocol resolves this tension by allowing storage partners to implement servers that expose data estates directly to Databricks Serverless Compute. This architecture ensures sensitive information remains physically stationary while becoming logically accessible for model training. Organizations no longer need to compromise between data sovereignty requirements and advanced cloud intelligence capabilities. The mechanism functions by having the storage array authenticate requests, effectively extending the governance perimeter without shifting data location.

Constraint Type Traditional Barrier Native OpenSharing Resolution
Regulatory Data must stay on-prem Query in place
Operational High egress costs Zero data movement
Security Air-gapped networks Secure endpoint exposure

Storage vendors must adopt the Partner Well-Architected Framework to guarantee security consistency. Enterprises must verify that their chosen storage partner fully implements these certification criteria before deployment. The system enables Databricks customers to query live on-premises Apache Iceberg™️ and Delta tables under Unity Catalog governance. This removes a substantial barrier between enterprise data and AI, allowing organizations to securely expose data where it lives while giving Databricks smooth access. The result is a unified catalog where dark data becomes actionable without violating legal residency constraints. Static archives change into flexible assets for generative AI workloads.

Operationalizing Hybrid Storage with Unity Catalog

AIStor Native Open Sharing Protocol Implementation

Conceptual illustration for Operationalizing Hybrid Storage with MinIO and Unity Catalog
Conceptual illustration for Operationalizing Hybrid Storage with MinIO and Unity Catalog

The provider AIStor natively implements the Open Sharing protocol to expose on-premises data directly to Databricks. This architecture eliminates data migration by allowing Unity Catalog to govern live Apache Iceberg™️ and Delta tables residing on local infrastructure. The mechanism relies on storage partners implementing servers that expose data estates directly to Databricks Serverless Compute without copying data, supporting Apache Iceberg APIs for sharing with Iceberg-native tools. Operators configure the the provider AIStor endpoint within the workspace, enabling immediate discovery of on-premises datasets. This approach extends Serverless Compute, Genie, and Agent Bricks to on-premises data while maintaining strict sovereignty.

Feature Traditional Migration Native Open Sharing
Data Movement Full Copy Required Zero Movement
Governance Scope Cloud Only Hybrid Unified
Latency Impact High (Ingestion) None (Direct Access)

Network connectivity between the on-premises storage and cloud compute must sustain high throughput for large-scale queries. Unlike batch replication, this live access model demands strong, low-latency links to prevent query degradation during peak AI training loads. Enterprises gain the ability to activate previously inaccessible data for AI and analytics without compromising control or incurring egress fees. Static archives become active assets for generative AI workflows.

Application: Extending Serverless Compute and Genie to On-Premises Data

Extending Serverless Compute, Genie, and Agent Bricks to on-premises data requires configuring the storage endpoint to natively implement the Open Sharing protocol. This architecture allows Unity Catalog to govern live Apache Iceberg™️ and Delta tables residing on local infrastructure without data migration. Operators connect the the provider AIStor endpoint within the workspace, enabling immediate discovery of datasets while maintaining strict data sovereignty. The mechanism relies on storage partners exposing data estates directly to Databricks, ensuring that sensitive information never leaves the premises during analysis.

Roadmap for Exposing Unstructured Files via Volumes APIs

Databricks is extending the OpenSharing protocol with Volumes APIs to expose unstructured files directly from on-premises storage. This roadmap guides organizations in connecting imaging archives and engineering simulations to Unity Catalog for GenAI workloads without data replication. Operators must first verify their storage infrastructure supports the OpenSharing server implementation required for secure exposure. The next step involves configuring the endpoint to bridge local object stores with cloud workspaces, ensuring fine-grained access controls remain intact across hybrid boundaries. This integration allows Serverless Compute clusters to query medical scans or backup archives while the physical data remains stationary.

Implementing Governed Data Sharing Across Hybrid Boundaries

Unity Catalog Governance for Live On-Premises Tables

Fine-grained access controls from Unity Catalog reach on-premises Apache Iceberg™️ and Delta tables without moving a single byte. This setup creates one control plane where policies apply everywhere, ignoring physical boundaries. Centralizing authority stops the fragmentation that blocks AI projects in hybrid environments.

Administrators register external connectors instead of migrating data. The workflow keeps the governance layer separate from storage execution:

Conceptual illustration for Implementing Governed Data Sharing Across Hybrid Boundaries
Conceptual illustration for Implementing Governed Data Sharing Across Hybrid Boundaries
  1. Deploy the storage partner's OpenSharing connector within the local network boundary.
  2. Register the external catalog endpoint using Unity Catalog REST APIs.
  3. Apply masking policies and row-level filters that enforce compliance at query time.

Eliminating egress costs while keeping regulated data on-site drives this operational win. Wide-area network latency becomes the dependency though. Compute jobs waiting for specific tables stall if the on-premises gateway disconnects. Connectivity must return before work resumes. This constraint suits architectures where data gravity matters more than sub-millisecond local access. Legacy silos turn into governed assets for Genie and Agent Bricks under this model.

Connecting AIStor to Databricks Serverless Compute

the provider AIStor uses the Open Sharing protocol to expose local data without migration. Local object storage becomes a governed endpoint for Databricks Serverless Compute, Genie, and Agent Bricks. Operators skip massive dataset replication by registering the storage array directly in Unity Catalog. Defining an external location pointing to the the provider AIStor bucket enables live queries on Apache Iceberg™️ and Delta tables.

  1. Deploy the the provider AIStor cluster and configure the bucket with read-only service credentials.
  2. Register the external location in Unity Catalog using the S3-compatible endpoint URL.
  3. Apply fine-grained access controls to map cloud identities to on-premises object keys.

Storage endpoints stay inside private networks while connectors handle secure metadata exchange. Egress fees disappear compared to migration-heavy patterns. Latency penalties common in hybrid setups drop notably. Network firewalls must allow outbound HTTPS traffic from the the provider array to the Databricks control plane for metadata handshakes to succeed. Strict data sovereignty mandates yield immediate compliance gains by keeping data on-site. Petabyte-scale media archives or proprietary AI training sets in semiconductor manufacturing and pharmaceutical research fit this architecture perfectly. These datasets cannot leave the facility. The Open Sharing protocol exposes such estates securely to cloud compute resources. Policy enforcement extends to the edge, supporting the "Govern Everything" strategy. Native integration avoids the overhead of maintaining duplicate data copies across environments.

Data Sovereignty Risks in Hybrid AI and Analytics Initiatives

GDPR mandates legally block cloud migration for specific financial services and healthcare datasets. Lifting and shifting petabytes of engineering-classified or historical tick data violates strict residency rules. Moving massive volumes triggers unsustainable egress fees and latency penalties despite high cloud adoption rates. Traditional extraction pipelines create temporary duplicates outside authorized jurisdictions, introducing unacceptable compliance exposure. The Open Sharing protocol resolves this tension by enabling secure, governed access without physical movement. Implementation requires registering on-premises storage as an external location within Unity Catalog rather than copying bytes to cloud buckets.

  1. Configure the on-premises Apache Iceberg™️ or Delta table metadata for read-only exposure.
  2. Establish a trusted connection between the local storage endpoint and Databricks Serverless Compute.
  3. Apply granular access policies in Unity Catalog that enforce sovereignty constraints at query time.

Raw data never leaves the air-gapped environment yet still powers advanced analytics. Network throughput presents a hard limitation. Low-bandwidth links bottleneck large-scale model training even with zero data replication. Teams must validate wide-area network capacity for required query concurrency before deployment. Native protocol adoption maintains strict adherence to data residency laws by avoiding replication entirely. Governance becomes a property of the access layer rather than the storage location.

About

Marcus Chen is a Cloud Solutions Architect and Developer Advocate at Rabata.io, where he specializes in S3-compatible object storage and AI/ML data infrastructure. His daily work involves designing scalable storage architectures that enable enterprises to govern data estates efficiently without vendor lock-in. This practical experience makes him uniquely qualified to analyze the Databricks Storage System, as he routinely helps organizations decouple compute from storage while maintaining strict governance controls. At Rabata.io, Chen uses deep expertise in S3 API compatibility to ensure smooth integration with platforms like Databricks, allowing clients to manage wherever their data lives. His insights reflect real-world challenges faced by data engineers who need transparent pricing and high-performance for generative AI workloads. By focusing on open standards rather than proprietary silos, Chen advocates for storage strategies that prioritize flexibility and cost-efficiency, directly aligning with the industry's shift toward governing distributed data assets effectively.

Conclusion

Scaling autonomous agents in 2026 exposes a critical fracture where static storage policies cannot match flexible agent behavior. The operational cost of maintaining separate governance layers for on-premises and cloud data becomes prohibitive as agent concurrency spikes. Organizations must unify their storage infrastructure under a single access control plane to prevent compliance drift. Relying on disjointed security models creates blind spots that auditors will penalize heavily next year.

Teams should mandate Open Sharing protocols for all hybrid workloads by the start of the next fiscal cycle. This approach ensures that data sovereignty remains intact while enabling the high-throughput access agents require. Do not attempt to replicate sensitive datasets to the cloud hoping to solve latency issues later. Instead, register local endpoints within Unity Catalog immediately to enforce governance at the query layer. This prevents the accumulation of technical debt associated with temporary data copies.

Start this week by mapping your most sensitive on-premises Apache Iceberg™️ tables for read-only exposure via Databricks Serverless Compute. Verify that your current network throughput can sustain the required query concurrency before enabling broad agent access. You can explore how Rabata.io helps optimize these hybrid storage configurations for maximum efficiency.

Frequently Asked Questions

High cloud egress fees make migrating massive historical archives financially impossible for these firms.

It allows model training on engineering-classified datasets without moving data from secure premises.

Operators stop managing bulk transfer bandwidth and start optimizing local read throughput for AI jobs.

It extends centralized governance layers to cover non-cloud storage estates directly.

Egress fees and storage costs break down entirely at petabyte and exabyte data scales.

References