Microsoft Fabric Updates Blog

Microsoft’s vision of an open data lake ecosystem: Open lakes, not walled gardens

In today’s data-driven world, enterprise data estates contain many data sources for a variety of reasons, including differences in type of usage (operational vs. analytic), differences in ownership, and the presence of legacy infrastructure that is part of a corporate merger or acquisition. In addition, enterprises constantly acquire and refresh data from external sources. For analytics to be effective, we require a unified view across the entire data estate. However, creation and maintenance of data pipelines to aggregate data have consistently posed a significant hurdle.

With the maturation of cloud-native big data platforms and the exciting revolution in generative AI, the potential for data-driven decisions and operational optimizations has never been greater, raising the urgency of solving the longstanding problem of how to enable organizations to bring together estate-wide data for analytics.

Optimizing processes by simplifying data

We believe that the emergence of open, updatable table formats presents us with a unique opportunity to solve this problem by standardizing these formats across all analytic engines, and by simplifying data replication. In fact, as an increasing number of engines adopt open data formats, we can minimize data replication by instead using references to data sources.

Further, as the value of data is recognized, we are seeing corresponding emphasis on right-use and increasing regulation. Thus, it is important that we be able to govern the entire data estate in a compliant manner, and in particular, evolve current best practices for aggregating estate-wide data to reflect the emerging world of cloud-native data lakes that bring together a diverse range of analytic capabilities, from exploratory tools, to AI models, to tools for serving data and rich business reports reliably, securely, and at scale.

Shaping the future of data analytics

This vision of the future of analytics is at the heart of OneLake design in Microsoft Fabric. We have striven to make it the “one place to bring all data for analytics”, making it easy to virtualize and aggregate data from all sources. Fabric itself then democratizes access to the wealth of insights that can be unlocked, thanks to a Microsoft 356-like simplicity in bringing analytic tools to bear on the data through intelligent software as a service, and by infusing AI copilot experiences to assist with complex tasks in-stride. The entire life cycle of analytics, from aggregating data to unlocking rich insights for appropriately authorized users, can be managed using the data governance capabilities of Fabric and the integrated estate-wide governance capabilities of Microsoft Purview.

Read the whitepaper to learn more!

Kapcsolódó blogbejegyzések

Microsoft’s vision of an open data lake ecosystem: Open lakes, not walled gardens

április 11, 2024 Készítette Matthew Hicks

Microsoft OneLake is a unified data lake for all of your organization’s data. With OneLake shortcuts, you can reference data in different locations and have that data logically represented within OneLake, with no data movement or duplication. With the recent announcement of shortcuts to Google Cloud Storage, you can use shortcuts to seamlessly bring in … Continue reading “Public Preview of OneLake shortcuts to S3-compatible data sources”

április 11, 2024 Készítette Trevor Olson

We are excited to announce that you can now create OneLake shortcuts to your Google Cloud Storage (GCS) buckets. With the addition of GCS, you can now utilize cross-cloud shortcuts to analyze your data across all three major cloud platforms. Shortcuts in OneLake allow you to connect to your existing data through a single unified … Continue reading “Shortcuts to Google Cloud Storage, now available in Public Preview”