Showing posts with label ea874-topic-3. Show all posts
Showing posts with label ea874-topic-3. Show all posts

Sunday, October 9, 2016

Big Data Evolutions: The Data Lake

Discussion for EA874 Topic 3 > Data/Information Architecture Layer
Post # 2

A fairly recent buzzword in the field of big data is that of the  “data lake” –  a term that now appears to endure both the scrutiny of certain camps who warns of being bogged down in more of a data swamp than swimming on a lake. Taking note of where these new concepts come from (Apache Hadoop camps), the vision of those who see the potential of the concept may yet again have a profound impact on enterprise data architecture. Do we remember when legacy IT folks took a 3-second heads-up, and waved off the then buzz on "Big Data"?

A data lake is a storage repository that holds a vast amount of raw data in its native format until it is needed. In contrast, a hierarchical data warehouse stores data in files or folders, while a data lake uses a flat architecture to store data.

Gartner back in 2014 has issued a warning and said that "while the marketing hype suggests audiences throughout an enterprise will leverage data lakes, this positioning assumes that all those audiences are highly skilled at data manipulation and analysis, as data lakes lack semantic consistency and governed metadata.  I think this is understandable, as the so-called data lake begins to come off its hype cycle and is driven by the pressures of pragmatic IT and business stakeholders, the unraveling and demand for clear data lake definitions, use cases, and best practices will continue to grow. Although, I am not quite sure if we're past disillusionment and are now in the phase of enlightenment on the cycle.

While it's true that the data lake started out as a metaphor for the transformation of data architectures given the newfound data volume and velocity, I agree with other pundits that it is a useful metaphor that serves as a guide for now,  to evolve enterprise data management according to established principles, drivers, and best practices in dealing with big data -- the scale of increased volume, variety, and a velocity that has never before been seen in the past. Clearly IT folks and EA architects will need to wrap their heads around the new world that Big Data is creating. How do new tools and concepts like data lakes help with the challenges posed by big data? How is it related to the current enterprise data warehouse? How will the data lake and the enterprise data warehouse be used together? How can you get started on the journey of incorporating a data lake into your architecture?

I think those organizations doing data architecture at the edge, are now realizing that the data lake is as an evolution from existing data architecture patterns and information management.


































Image source: Hortonworks. http://www.slideshare.net/hortonworks/modern-data-architecture-for-a-data-lake-with-informatica-and-hortonworks-data-platform

Useful References:

Gartner. (July 28, 2014). Gartner Says Beware of the Data Lake Fallacy. Press Release. http://www.gartner.com/newsroom/id/2809117

Hortonworks. (2014). Putting the Data Lake to Work: A Guide to Best Practices. White Paper.
https://hortonworks.com/wp-content/uploads/2014/05/TeradataHortonworks_Datalake_White-Paper_20140410.pdf

Information vs Data: Semantic Shifts

Discussion for EA874 Topic 3 > Data/Information Architecture Layer
Post # 1

I think the way to sort out enterprise architecture along the domains of information and data  is by going back and sorting out our semantics, i.e. what we mean by terms like information and data.  I admit that I am myself guilty of mix-ups, because there are many contexts that we can really get away with using one term for the other. However, by clearly defining what we mean by data, information and knowledge – and how they interact with one another – it should be much easier to develop our taxonomies to communicate our work in the EA discipline.

Information vs. Data vs. Knowledge

Neil Ingebrigtsen's blog at infogineering.net is an example of useful discussions on the differences between data, information and knowledge. Neil works backwards and starts with what knowledge is. Knowledge is not just what we know, but what we know based on our personal beliefs and expectations. This makes sense, and explains knowledge in the way we speak of the "giant network of ideas, memories, predictions, beliefs, etc." What are the sources of this knowledge? Data and Information.

Similar to many popular discussions, "data" are the basic facts of the world, physiologically perceived with the senses, and eventually processed by the brain. Thus, data are raw, unorganized facts that need to be processed. Data can be something simple and seemingly random and useless until it is organized, and combined with other data elements. When data is processed, organized, structured or presented in a given context so as to make it useful, it is called information. Simple example: We can have a notion of 6 feet, but it only remains as data until perhaps we associate it with say a person,  on which it becomes information about that person -- and so on.

So here's why I was drawn to Neil's discussion: "When people confuse data with information, they can make critical mistakes. Data is always correct (I can’t be 29 years old and 62 years old at the same time) but information can be wrong (there could be two files on me, one saying I was born in 1981, and one saying I was born in 1948). He keenly notes that information captures data at a single point in time, and that data changes over time. "The mistake people make is thinking that the information they are looking at is always an accurate reflection of the data." And I agree with him that by understanding the differences between these, we can better understand how to make better decisions based on the accuracy of [information].

In summary, we can thus find the following to be a useful working taxonomy.
Data: Raw factual descriptions of the World
Information: Captured Data and processed to provide useful context.
Knowledge: Our personal use of information to create a map/model of the World

Bellinger et al, in their articles at systems-thinking.org, elaborates on the following so-called DIKW extension that includes wisdom, and invokes the scholarship of folks like R. Ackoff and N. Sharma on the evolution of this model for information use.



 The DIKW framework is used by many organizations. The diagram below shows the adaptation of the DIKW pyramid by US Army Knowledge Managers.




Semantic Shifts.

Going deeper now into investigating the nuances of Information Architecture and Data Architecture, I found a truly fascinating compilation of historical notes from Resmini, et al, (2011) on how our semantics shifted for the way we communicate the notion of Ïnformation Architecture." Their article speaks of Wurman's original contribution and vision that many think stills holds today, and which places Information Architecture as follows:

a.) the organization of  the patterns inherent  in data, making the complex clear; b.) the creation of the structure or map of information which allows others to find their personal paths to knowledge; c.) the emerging 21st century occupation that addresses the needs of the age focused upon clarity, human understanding, and the science of the organization of information.

Resmini's article becomes truly insightful when it narrates how the Rosenfeld-Morville team shifted the notion Information Architecture later to emphasize the importance of structure of and organization in website design i.e. what they call "Pervasive IA". Quite interesting to note that this shows how the disconnect to Wurman's past scholarship allowed for this shift to happen. At any rate, I think although the emphasis has shifted, we can still see how scholarly intuition continues to place structure and organization -- in the conceptual and logical context of Wurman's classical IA, as the overarching notion associated with what we call Information Architecture.

What about Data Architecture?

Ultimately, I come back to this old NIST Enterprise Architecture reference diagram of the late-1980's, shown below  to serve as my own reminder that we can separate Information Architecture with what we call Data Architecture. The NIST Enterprise Architecture Model is a five-layered model with each layer are defined separately but are interrelated and interwoven. The model defined the interrelation as follows:
  • Business Architecture drives the information architecture
  • Information architecture prescribes the information systems architecture
  • Information systems architecture identifies the data architecture
  • Data Architecture suggests specific data delivery systems, and
  • Data Delivery Systems (Software, Hardware, Communications) support the data architecture.
The hierarchy in the model is based on the notion that an organization operates a number of business functions, each of which requires information from a number of sources, and each of these sources may come from any one or more operational systems, which in turn stores organized data in any number of data systems.

Mapping out these nuances onto our conceptual-logical-physical modeling constructs, we can perhaps intuitively associate the conceptual, semantic, and logical relationships at the Information Architecure layer, the processing logic, approach, and data systems at the Data Architecture layer, and leave most of the suppporting technology components for database engines, storage,  and I/O infrastructure to the realm of Infrastructure Architecture.

Thus for what we usually call Information or Data Architecture, we may perhaps find this deconstruction useful.



References:

Neil Ingebrigtsen (n.d.). The Differences Between Data, Information and Knowledge. http://www.infogineering.net/data-information-knowledge.htm

Gene Bellinger, Durval Castro, Anthony Mills. Data, Information, Knowledge, and Wisdom. http://www.systems-thinking.org/dikw/dikw.htm

Andrea Resmini, Luca Rosati (2011). A Brief History of Information Architecture. Journal of Information Architecture.  http://journalofia.org/volume3/issue2/03-resmini/


Data Virtualization and the Semantic Layer

Discussion for EA874 Topic 3 > Data/Information Architecture Layer
Post # 3

Ultimately, there are 2 key questions we ask as data architects in service to our business customers:
  1. How can business gain access to the data to gain insights?
  2. How do we provide consistency and quality -- while maintaining agility?
You can easily imagine the challenges of traversing traditional data warehouse solutions as depicted in the diagram below, in order to bring data to our internal enterprise customers.














source:
http://www.slideshare.net/Denodo/data-warehousesmdmdatavirtualizationdenodopackedlunchseriessession5
Implementing Data Virtualization for Data Warehouses and Master Data Management extensions
By Denodo Technologies, published on Jan 22, 2015

As Chris Daniels points out in his article, "inadequate reporting capabilities seem to be a common characteristic of both off-the-shelf and bespoke systems. It’s a real shame because it can prevent firms from exploiting the full potential of the information they hold." He points to several factors that result in computer systems that continue to be delivered without adequate reporting facilities, and that organisations wishing to make the most of their information will need to find ways to plug the gaps. Beyond the first-order solutions using Query-Reporting tools, he goes on to point out another solution which architects would also need to keep in mind: the provision of a Semantic Layer.

What is a Semantic Layer?
A semantic layer is a business representation of corporate data that helps end users access data using common business terms. The aim is to insulate users from the technical details of the data store and allow them to create queries in terms that are familiar and meaningful.

An increasing number of reporting tools allow users to make use of a “semantic layer.” Some vendors give the semantic layer a name that is specific to their particular product (for example, Business Objects calls it a “universe”), while others describe it as a business model or metadata layer, that should make it easier and intuitive of end-users to create their own reports by themselves.  However, it is typical that the semantic layer is still mainly used by developers -- and not business users.  Unfortunately, developers are SQL-savvy enough to perceive the semantic layer as a layer of nuisance i.e. another set of technical paradigm to deal with. As a result, it can be tempting for developers to bypass the semantic layer and revert to type and write queries by hand for each and every report.

As Chris notes in his conclusion, there are real benefits to be gained through the use of semantic layers. Semantic layers are increasingly becoming available within all types of reporting environments, and their use should be encouraged for the benefit of placing the power of harvesting information in the hands of the end-users themselves. Moreover, I think the technologies that create semantic layers prepares us to appreciate the more abstracted data access solution: Data virtualization.

What is Data virtualization?
Data virtualization is defined as an umbrella term used to describe any approach to data management that allows an application to retrieve and manipulate data without requiring technical details about the data, such as how it is formatted or where it is physically located i.e. it is another layering concept.

Data virtualization software may include functions for development, operation, and/or management.

Benefits include:
  • Reduce risk of data errors
  • Reduce systems workload through not moving data around
  • Increase speed of access to data on a real-time basis
  • Significantly reduce development and support time
  • Increase governance and reduce risk through the use of policies
  • Reduce data storage required 

Drawbacks include:
  • May impact Operational systems response time, particularly if under-scaled to cope with unanticipated user queries or not tuned early on
  • Does not impose a heterogeneous data model, meaning the user has to interpret the data, unless combined with Data Federation and business understanding of the data
  • Requires a defined Governance approach to avoid budgeting issues with the shared services
  • Not suitable for recording the historic snapshots of data - data warehouse is better for this
  • Change management "is a huge overhead, as any changes need to be accepted by all applications and users sharing the same virtualization kit" 

I won't be able to expand on virtualization in this post, but I think it's sufficient to say that data architects would need to be familiar with the never-ending solutions coming out that help fulfill mission for the business, and works well with the enterprise information principle of accessibility.

References:
Chris Daniels. (2010 March).  What is a Semantic Layer and Why Would I Want One? http://www.b-eye-network.co.uk/view/12763

Data Virtualization. https://en.wikipedia.org/wiki/Data_virtualization

Semantic Layer. https://en.wikipedia.org/wiki/Semantic_layer
/