Sunday, October 9, 2016

Big Data Evolutions: The Data Lake

Discussion for EA874 Topic 3 > Data/Information Architecture Layer
Post # 2

A fairly recent buzzword in the field of big data is that of the  “data lake” –  a term that now appears to endure both the scrutiny of certain camps who warns of being bogged down in more of a data swamp than swimming on a lake. Taking note of where these new concepts come from (Apache Hadoop camps), the vision of those who see the potential of the concept may yet again have a profound impact on enterprise data architecture. Do we remember when legacy IT folks took a 3-second heads-up, and waved off the then buzz on "Big Data"?

A data lake is a storage repository that holds a vast amount of raw data in its native format until it is needed. In contrast, a hierarchical data warehouse stores data in files or folders, while a data lake uses a flat architecture to store data.

Gartner back in 2014 has issued a warning and said that "while the marketing hype suggests audiences throughout an enterprise will leverage data lakes, this positioning assumes that all those audiences are highly skilled at data manipulation and analysis, as data lakes lack semantic consistency and governed metadata.  I think this is understandable, as the so-called data lake begins to come off its hype cycle and is driven by the pressures of pragmatic IT and business stakeholders, the unraveling and demand for clear data lake definitions, use cases, and best practices will continue to grow. Although, I am not quite sure if we're past disillusionment and are now in the phase of enlightenment on the cycle.

While it's true that the data lake started out as a metaphor for the transformation of data architectures given the newfound data volume and velocity, I agree with other pundits that it is a useful metaphor that serves as a guide for now,  to evolve enterprise data management according to established principles, drivers, and best practices in dealing with big data -- the scale of increased volume, variety, and a velocity that has never before been seen in the past. Clearly IT folks and EA architects will need to wrap their heads around the new world that Big Data is creating. How do new tools and concepts like data lakes help with the challenges posed by big data? How is it related to the current enterprise data warehouse? How will the data lake and the enterprise data warehouse be used together? How can you get started on the journey of incorporating a data lake into your architecture?

I think those organizations doing data architecture at the edge, are now realizing that the data lake is as an evolution from existing data architecture patterns and information management.


































Image source: Hortonworks. http://www.slideshare.net/hortonworks/modern-data-architecture-for-a-data-lake-with-informatica-and-hortonworks-data-platform

Useful References:

Gartner. (July 28, 2014). Gartner Says Beware of the Data Lake Fallacy. Press Release. http://www.gartner.com/newsroom/id/2809117

Hortonworks. (2014). Putting the Data Lake to Work: A Guide to Best Practices. White Paper.
https://hortonworks.com/wp-content/uploads/2014/05/TeradataHortonworks_Datalake_White-Paper_20140410.pdf

No comments:

Post a Comment