Sunday, October 23, 2016

The Weak Links of Security in the Enterprise

Discussion for EA874 Topic 5 > Enterprise Security Architecture
Post # 3

In March 2015,  Cisco’s Talos Threat Intelligence and Research Team also reported a rise in tax-related phishing attempts in the US as the tax reporting season gets into full swing.  Some of these attempts are incredibly sophisticated and are very hard to spot as being fraudulent. We now also have cybercrime gangs placing adverts on legitimate websites and use them to inject malware into unsuspecting people browsing the ad. Cisco's 2015 Annual Security Report (CASR) suggested compromised users are often infected with malicious browser add-ons through the installation of bundled software (software distributed with another software package or product) via these sorts of malvertisments and usually without clear user consent. The CASR also highlighted how users’ careless behaviour when using the Internet, combined with targeted campaigns by adversaries, places many industry verticals at higher risk of web malware exposure. In 2014, the pharmaceutical and chemical industry emerged as the number one sectors to be targeted in this way.


In the case of the 2013 data breach at Target, the initial intrusion into its systems was traced back to network credentials that were stolen from a third party vendor which did contractual HVAC work at a number of locations at Target and other top retailers.

The reconnaissance that hackers conduct goes beyond mapping a company’s IT network, and would also be interested in gathering as much information as possible on their target, especially around how the business and its key personnel operate. These details will help attackers navigate around any technological or human barriers during an attack. To collect these details, hackers will use social media to learn where key members of your security team worked or went to college. Once an attack has penetrated your network, file access can be executed to review emails and calendar entries to learn when key security personnel are on vacation and attack when there’s a staffing gap. Hacking teams can also be specialized: usually one group dedicated to deception that creates a campaign that distracts the security team from the main attack operation. The distraction is meant to mitigate the risk of the campaign being discovered. DDoS attacks are usually employed as distractions that can be a challenge but can easily be detected. These decoy threats mask the real threat which can escape detection during the distraction.



Thus, these scenarios show that visibility across the whole corporate network is critical to managing security. It is not enough to just defend the threat coming into and out of the network; you have to be able to manage the threat across the whole continuum, before, during and after the attack. While better security technologies and solutions are needed, we should not discount the human factor as a key vulnerability that must be seriously considered when developing the enterprise security architecture.


Reference:
Graham Welch. (2015 Mar 11). People Remain the Weakest Link in Security.  http://www.cio.com/article/2895404/cybercrime/people-remain-the-weakest-link-in-security.html

Managing the Kill Chain

Discussion for EA874 Topic 5 > Enterprise Security Architecture
Post # 2

Enterprise security architects today have their hands full as they are tasked to develop security programs that will reduce their organization's attack surface and be able to rapidly respond to contain cyberattacks with security management solutions that enables total network visibility, and analytics which provide actionable security intelligence.  Security architects need to be familiar with the growing set of technology solutions that can combine firewall and network device data with vulnerability and threat intelligence, which enables effective security decisions for rapid response to bring threat management processes under control.

Securing the network against cyberattacks is not a straightforward task with clear-cut solutions. While it is true that enterprise security is currently being hardened with better programs and technologies, we see and hear alarming stories of more powerful and persistent attacks being developed by threat actors. Even more alarming are the data breach incidents that happened to those organizations who we easily perceive as highly advanced technology practitioners.

On September 22, 2016, Yahoo publicly disclosed a data breach on their systems that happened late 2014. Hackers stole information associated with at least 500 million Yahoo! user accounts, and is said to be the largest discovered in the history of the Internet. Specific details of material taken include names, email addresses, telephone numbers, encrypted or unencrypted security questions and answers, dates of birth, and encrypted passwords.  Yahoo alleged in its statement that the breach was carried out by "state-sponsored" hackers.

There is an alarming growth of such cyberattacks:

In June 2015, The Office of Personnel Management of the U.S. government suffered a data breach in which the records of 4 million current and former federal employees of the United States were hacked and stolen.

In September 2014, Home Depot suffered a data breach of 56 million credit card numbers.
In October 2014, Staples suffered a data breach of 1.16 million customer payment cards

In late November to early December 2013, Target Corporation announced that data from around 70 million credit and debit cards was stolen.

In October 2013, Adobe Systems revealed that their corporate data base was hacked and some 130 million user records were stolen. According to Adobe, "For more than a year, Adobe’s authentication system has cryptographically hashed customer passwords using the SHA-256 algorithm, including salting the passwords and iterating the hash more than 1,000 times. This system was not the subject of the attack we publicly disclosed on October 3, 2013. The authentication system involved in the attack was a backup system and was designated to be decommissioned. The system involved in the attack used Triple DES encryption to protect all password information stored.

The list is growing longer each month.

Many of these sophisticated cyber security attacks are generally attributed to groups classified as APT's or Advanced Persistent Threats, and are normally composed of organized syndicates or even state-sponsored actors.





























Image source: https://en.wikipedia.org/wiki/Advanced_persistent_threat

The attacks are sometimes described using the so-called "Kill Chain" model as a method to describe and analyze intrusions on a computer network. The model has it's share of critics, pointing out that the scope of recent intrusions extends far beyond that of the Cyber Kill Chain model. Nonetheless, I think the model can be useful for understanding an approach for developing defense or pre-emptive action against threats.

The following is a brief description of its seven steps.

Step 1: Reconnaissance. The attacker gathers information on the target before the actual attack starts. He can do it by looking for publicly available information on the Internet.

Step 2: Weaponization. The attacker uses an exploit and creates a malicious payload to send to the victim. This step happens at the attacker side, without contact with the victim.

Step 3: Delivery. The attacker sends the malicious payload to the victim by email or other means, which represents one of many intrusion methods the attacker can use.

Step 4: Exploitation. The actual execution of the exploit, which is, again, relevant only when the attacker uses an exploit.

Step 5: Installation. Installing malware on the infected computer is relevant only if the attacker used malware as part of the attack, and even when there is malware involved, the installation is a point in time within a much more elaborate attack process that takes months to operate.

Step 6: Command and control. The attacker creates a command and control channel in order to continue to operate his internal assets remotely. This step is relatively generic and relevant throughout the attack, not only when malware is installed.

Step 7: Action on objectives. The attacker performs the steps to achieve his actual goals inside the victim’s network. This is the elaborate active attack process that takes months, and thousands of small steps, in order to achieve.


In the case of the Target breach of 2013, an analysis using the kill chain model provides facility for deconstructing the attack.





















Image source: http://www.darkreading.com/attacks-breaches/leveraging-the-kill-chain-for-awesome/a/d-id/1317810


/

The Policy-Driven Security Program

Discussion for EA874 Topic 5 > Enterprise Security Architecture
Post # 1

I think enterprise architects can all find great value and use for the template provided by the Network Applications Consortium for Enterprise Security Architecture.  The document truly provides useful guidance for developing a policy-driven security program -- easily applicable in many organizations.

The three pillars of security -- comprised of Availability, Confidentiality and Integrity, also gives us guidance on the fundamental goals of security programs. Our security programs should keep these goals in mind as we identify and prioritize assets (what we need to protect), those that put these assets at risk (the threats we need to protecting our assets from),  and how we plan to protect it (our security architecture). Policies and its supporting processes are developed based on findings and recommendations to reduce the risks posed by the threats.

Among our reading selections, the Gartner article of McMillan and Sholtz (2013) warns us that in developing security architecture, one significant problem is the confusion that organizations sometimes create regarding three key functions: 1.) IT security governance, 2.) security management and 3.) security operations. Such confusion can easily create the kind of inefficiencies that result in the organization's overall security performance failure, conflict of interest,  and organizational dysfunction. The research article helps clarify the functional distinctions and argues the importance of establishing a security governance forum that stays out of operational issues, but rather gives direction and oversight for the organization in order to maintain the organization's focus on business outcomes.  The clear definitions of the three functions will ensure the necessary separation of roles within the security architecture and processes, such that conflicts of interest which can create vulnerabilities can be avoided.

The NAC expands these discussions with its framework for developing the security programs and architecture, and provides a template which describes and delineates the various elements that work together in the framework.  The diagram below, included in the white paper, shows the relationships of these elements with demarcations placed for Security Governance, Security Program Management, and Security Operations.








































The NAC framework is policy-centric and walks us through a discussion of clearly defined steps to develop a security program, starting from the development of policies, standards, procedures, and guidelines; all of which are in line with the organization's risk management practice, and their definition of security requirements.  I think the diagram provides a complete depiction of all key elements that needs to be considered for security program development.

By the way, I think the Trust Model from the Gartner toolkit selection also provides a good starting point for identifying security requirements that can be used to drive policy development. It has the trust level classifications of Low, Medium, and High trust, for each identified control area within each of the  13 control categories:   1.) Access Controls; 2.) Awareness and Training; 3.) Configuration Management; 4.) Maintenance; 5.) Media Protection; 6.) Physical Access Controls; 7.) Personnel Security; 8.) Risk Assessment; 9.) System Host Security; 10.) Systems and Information Interity; 11.) Workstation Protection; 12.) Wireless and Mobile; 13.) Audit and Accountability.

References:

Network Applications Consortium. (2004). Enterprise Security Architecture: A framework and template for a policy-driven security program. White paper. Prepared from NAC’s Strategic Interest Group (SIG) process.

Rob McMillan and Tom Scholtz. (2013 Jan 23) Security Governance, Management and Operations Are Not the Same. Gartner research article.

/


Sunday, October 9, 2016

Betting on Winners

Discussion for EA874 Topic 4 > Technology Infrastructure Architecture
Post # 3

Continuing with the previous blog's theme of risks in technology choices and acquisitions for infrastructure, Gartner's recommendation for installing an Advanced Technology Group (ATG) in the organization is a key step in minimizing such risks, and allows the organization to pursue their investments with better  odds. I would think that as part of a robust EA governance being in place in the organization, creating the strategic technology planning function that Gartner speaks of can truly bring the necessary defined process for prioritizing and transferring potentially high-impact emerging technologies. Doing so will strengthen the company for avoiding the potential disasters of personality-driven investment decisions.

Five Styles of an Advanced Technology Group
Working with various organizations in the past, I can relate to the five styles described for ATG's as presented by Gartner. I would also go further to say that even in companies without a formal ATG, the five styles can be present - particularly the Guerilla kind, with characteristics, strengths and challenges most akin to small to medium-sized enterprises. As pointed out by the article,  ATGs can combine qualities from more than one style, but a single style is usually dominant.  I would also add that the idea of an ATG can be a venue of sorts, in which the members can come and go to contribute, but usually with a core group to maintain control and direction. And I think in some places, I've seen that the notion of an ATG can actually be a process -- again an open model that allows the free exchange to deliberate  technology options and directions. Lastly, I think I have also observed that while a group can display a particular style as dominant, the styles can be fluid, and shift from one to another depending on context, and the particular posture and requirements of the organization to which the service of the ATG is called upon.  The following styles descriptions can be useful: 1.) Navigator: Determining the strategic business impact of emerging technologies, through tracking and evaluation. Usually challenged by: Finding a willing home for a technology at the end of a successful evaluation. 2.) Guerrilla: Tactical and pragmatic "SWAT" team helping business units deploy new technologies. Usually challenged by: Carving out time for strategic planning and tracking activities. 3.) Priest: Educating senior management on emerging technology issues and potential. Usually challenged by: Lack of hands-on evaluation activities. 4.) Conductor: Coordinating and leveraging advanced technology activities performed in other parts of the organization. Usually challenged by:  Making its recommendations a reality by working primarily through other groups. 5.) Scholar: Research and development group investigating technologies ahead of business need. Usually challenged by: Avoiding acquiring a reputation as an “ivory tower” out of touch with current business needs.

Timing-Decisions for Adoption Opportunities
I think one of the more challenging decisions for ATG's or any group that functions to provide strategic technology planning is when to jump in to adopt a new technology. A strong prioritization criteria should be in place, and a clear view of enterprise portfolio is requisite to arrive at timely technology decisions with the restraint to jump along on hype. In performing the prioritization process, Gartner points out that it is also important to identify, and thus avoid, the commonly occurring wrong reasons for which companies adopt technology. The high level of hype surrounding technology in the marketplace is one of the factors that frequently drives companies to a poorly timed adoption of technology (typically too early). As noted in our other discussions, Gartner's model of the Hype Cycle characterizes the typical progression of a technology, from over-enthusiasm through a period of disillusionment (because of the inevitable failures that arise from inappropriate application), to an eventual understanding of the technology’s relevance and role. If a company launches its efforts too soon, it will suffer unnecessarily through the painful and expensive lessons associated with deploying an immature technology. If it delays action for too long, it runs the even greater risk of being left behind by competitors that have succeeded in making the technology work to their advantage.

Reference:
Fenn, Linden, and Fairchok. (2003 July). Strategic Technology Planning: Picking the Winners. Gartner Research.

/

What should we expect from Infrastructure Architecture?

Discussion for EA874 Topic 4 > Technology Infrastructure Architecture
Post # 2

As one of the pillars of enterprise architecture, the technology infrastructure architecture should be developed and delivered to respond in support of the business mission and strategies. The technology infrastructure architecture as a service model is changing as it is driven by expected business outcomes which demand that information technology in the organization be an enabler  for the business innovation, with the business organization as a key influence in buying decisions, and where the business need for speed, agility, flexibility and cost-efficiencies are effectively provisioned .

Adaptive to  New Technologies
Gartner's key findings back in 2012, still resonates today, with continued indications of  accelerated spending on new technologies -- mobility, cloud, analytics and social computing services -- will be driven by the segment of buyers that seek untapped value from these technologies. Buyers have increasingly focused on speed to market, business impact of their enterprise applications and efficiency of their application portfolios; these choices guide and underlie new preferences for applications, application services and provider selection. Crowd and community sourcing continues to grow to create distinct "application management communities," which contribute on both the development and maintenance space, with outputs that clearly benefit all enterprises. The increased focus of providers in developing/maintaining reusable intellectual property (IP) assets is evidence of a future in which cloud-based platform development, assembly-based software strategies, IP leverage, and configurable solutions will be an option -- if not the norm, for some buyer solutions. The intentions of legacy application modernization efforts will shift conclusively from an IT-centric cost view to business enablement through new technologies -- such as analytics, mobility, and software as a service (SaaS).

Enabling Infrastructure
These trends and scenarios result in innovative solutions similar to GE's "Industrial Internet Platform,” which links GE’s products and sensors with analytic applications built with components from new ecosystem partners.  This essentially creates a platform that bridges operational technology (OT) sensor data with IT-based equipment service history, inspection details and other unstructured data for actionable insights. IT/OT alignment involves aligning and integrating traditional information technology with devices, sensors and software used to monitor and control physical equipment. Clearly these trends give us an  insight into how IT/OT alignment will impact device control technology, enterprise architecture and other mission-critical processes in key industries, including utilities, defense, transportation, healthcare, mining, manufacturing, media and telecommunications. Such an enabling infrastructure provides a platform where airlines can use big data analytics to predict maintenance issues, cut operating costs and improve on-time rates; where utility enterprises can analyze combustion turbine data with weather details to optimize asset performance against emissions constraints and fuel cost changes.

An architecture minimizing technology risks
Organizations will expect to have an infrastructure architecture that is designed to minimize technology risks for each life cycle stage. Organizations will always find the acquisition for technology infrastructure to be challenging to one degree or another,  and is only right to tread carefully. However, whatever buying strategy they come up with, organizations will need to ensure a sustainable infrastructure framework that can withstand the different degrees of risk for each of the Five-Stage Technology Adoption Life Cycle: Innovators, Early Adopters, Early Majority, Late Majority, Laggards;  a framework that can address technology risk areas such as technology death (too soon), vendor defocus, deplorable levels of vendor support, runaway maintenance costs that become prohibitive, availability of support expertise in the market.  Given all these risks, we should appreciate how the provision of this backbone for all our information systems -- i.e. the infrastructure for network, security, storage, process platforms, database, integration, and presentation models --  is truly a daunting task. A big part of Infrastructure architecture, because of it nature as a big-item acquisition, needs a strong governance process that can deal with heavy-duty decisions.  Such decisions should be designed to keep the pain of change to a minimum, hopefully strongly guided by EA principles,  with well-placed acquisitions of systems and products that will not cost the organization an arm and leg to replace or maintain.

References:
Mike Blechar. (2006 December). Managing Technology Life Cycle Risks. Gartner Research paper.

Kristian Steenstrup. (2011 July). IT and Operational Technology Alignment Innovation Key Initiative Overview. Gartner Research paper.

Gartner. (2012 November). Predicts 2013: Business Impact of Technology
Drives the Future Application Services Market. Gartner Research paper.

Sallam, Kart, and Rhodes. (2013 June). GE's Planned Industrial Internet Platform Shows Promise, Faces Hurdles. Gartner Research paper.

Infinite Data From Every Direction

Discussion for EA874 Topic 4 > Technology Infrastructure Architecture
Post # 1

As new web technologies have unleashed to bring in a tidal wave of data, the resulting increased requirements for handling "infinite data" has increased the possibilities for storage in every part of the Web. The Gartner article (Monroe, et al, 2012) reminds us that we are only beginning to see the enormous implications of that simple fact. As enterprise architects, we should be wise to consider these predictions in the course of work of designing systems, allocating resources, or selecting products for our organizations. The article relates to my own reflections on big data and data lakes in a previous blog.

As pointed out in their key findings, traditional content types, including simple unstructured user data, are seeing growth rates of up to 60% to 80% year over year as of 2012., and new enterprise infrastructure configurations associated with virtual servers are compounding this problem. And while Big data infrastructure and platform purchases are still driven by data management considerations, end users are now beginning to realize that underlying storage architectures are an important and often misunderstood cost center.  Good news is brought to consumers as Conditions in the generic cloud storage markets in both the consumer and enterprise arenas have opened the gates for severe price wars, and providers are seeking to differentiate themselves through platform/device partnerships and by targeting specific market segments. Moreover, hardware and software technologies have evolved to the point where it is now practical to avoid the high costs of proprietary solutions, at least for non-mission-critical data.

We should take action as part of our diligence and pass along the following advise to our respective organizations: 1.) Data must be managed according to policies that make sense for the business rather than by content type, owner, age or size alone. 2.) Do not remain complacent with your current storage technologies and big data implementation, and work to increase your organization's use of information management dimensions, such as volume, variety and velocity of data, through emerging and potentially disruptive storage architectures and designs; and 3.) Examine your relationship with your enterprise storage vendor to see which cloud storage providers it integrates with, and potentially seek to increase the serviceability and compliance of such an integration/partnership.

As part of our strategic planning, we should now have all included efficient, cost-effective information management as one of the top three measures of the business health for our respective organizations.  If we haven't joined the inescapable trend, we should make our organizations aware of the value of insights big data brings to the business, and that 80% of big data projects will use architectures that account for less than 20% of total storage spending today because of current cloud opportunities,  and the increasingly favorable buyer's market for cloud provisions. And as computing power and costs favorably continues to follow Moore's law, infrastructure architects should brush off their cobwebs and look away from  traditional storage arrays that are now increasingly replaced by more intelligent servers which bring a paradigm shift of what is now generally observed as gargantuan proportions.

Reference:
Monroe, Chandrasekaran, Childs, Deshpande, Filks, Zaffos, Unsworth. (2012 November). Predicts 2013: Managing Infinite Data From Every Direction. White paper. Gartner. ID G00245437.
/

Big Data Evolutions: The Data Lake

Discussion for EA874 Topic 3 > Data/Information Architecture Layer
Post # 2

A fairly recent buzzword in the field of big data is that of the  “data lake” –  a term that now appears to endure both the scrutiny of certain camps who warns of being bogged down in more of a data swamp than swimming on a lake. Taking note of where these new concepts come from (Apache Hadoop camps), the vision of those who see the potential of the concept may yet again have a profound impact on enterprise data architecture. Do we remember when legacy IT folks took a 3-second heads-up, and waved off the then buzz on "Big Data"?

A data lake is a storage repository that holds a vast amount of raw data in its native format until it is needed. In contrast, a hierarchical data warehouse stores data in files or folders, while a data lake uses a flat architecture to store data.

Gartner back in 2014 has issued a warning and said that "while the marketing hype suggests audiences throughout an enterprise will leverage data lakes, this positioning assumes that all those audiences are highly skilled at data manipulation and analysis, as data lakes lack semantic consistency and governed metadata.  I think this is understandable, as the so-called data lake begins to come off its hype cycle and is driven by the pressures of pragmatic IT and business stakeholders, the unraveling and demand for clear data lake definitions, use cases, and best practices will continue to grow. Although, I am not quite sure if we're past disillusionment and are now in the phase of enlightenment on the cycle.

While it's true that the data lake started out as a metaphor for the transformation of data architectures given the newfound data volume and velocity, I agree with other pundits that it is a useful metaphor that serves as a guide for now,  to evolve enterprise data management according to established principles, drivers, and best practices in dealing with big data -- the scale of increased volume, variety, and a velocity that has never before been seen in the past. Clearly IT folks and EA architects will need to wrap their heads around the new world that Big Data is creating. How do new tools and concepts like data lakes help with the challenges posed by big data? How is it related to the current enterprise data warehouse? How will the data lake and the enterprise data warehouse be used together? How can you get started on the journey of incorporating a data lake into your architecture?

I think those organizations doing data architecture at the edge, are now realizing that the data lake is as an evolution from existing data architecture patterns and information management.


































Image source: Hortonworks. http://www.slideshare.net/hortonworks/modern-data-architecture-for-a-data-lake-with-informatica-and-hortonworks-data-platform

Useful References:

Gartner. (July 28, 2014). Gartner Says Beware of the Data Lake Fallacy. Press Release. http://www.gartner.com/newsroom/id/2809117

Hortonworks. (2014). Putting the Data Lake to Work: A Guide to Best Practices. White Paper.
https://hortonworks.com/wp-content/uploads/2014/05/TeradataHortonworks_Datalake_White-Paper_20140410.pdf

Information vs Data: Semantic Shifts

Discussion for EA874 Topic 3 > Data/Information Architecture Layer
Post # 1

I think the way to sort out enterprise architecture along the domains of information and data  is by going back and sorting out our semantics, i.e. what we mean by terms like information and data.  I admit that I am myself guilty of mix-ups, because there are many contexts that we can really get away with using one term for the other. However, by clearly defining what we mean by data, information and knowledge – and how they interact with one another – it should be much easier to develop our taxonomies to communicate our work in the EA discipline.

Information vs. Data vs. Knowledge

Neil Ingebrigtsen's blog at infogineering.net is an example of useful discussions on the differences between data, information and knowledge. Neil works backwards and starts with what knowledge is. Knowledge is not just what we know, but what we know based on our personal beliefs and expectations. This makes sense, and explains knowledge in the way we speak of the "giant network of ideas, memories, predictions, beliefs, etc." What are the sources of this knowledge? Data and Information.

Similar to many popular discussions, "data" are the basic facts of the world, physiologically perceived with the senses, and eventually processed by the brain. Thus, data are raw, unorganized facts that need to be processed. Data can be something simple and seemingly random and useless until it is organized, and combined with other data elements. When data is processed, organized, structured or presented in a given context so as to make it useful, it is called information. Simple example: We can have a notion of 6 feet, but it only remains as data until perhaps we associate it with say a person,  on which it becomes information about that person -- and so on.

So here's why I was drawn to Neil's discussion: "When people confuse data with information, they can make critical mistakes. Data is always correct (I can’t be 29 years old and 62 years old at the same time) but information can be wrong (there could be two files on me, one saying I was born in 1981, and one saying I was born in 1948). He keenly notes that information captures data at a single point in time, and that data changes over time. "The mistake people make is thinking that the information they are looking at is always an accurate reflection of the data." And I agree with him that by understanding the differences between these, we can better understand how to make better decisions based on the accuracy of [information].

In summary, we can thus find the following to be a useful working taxonomy.
Data: Raw factual descriptions of the World
Information: Captured Data and processed to provide useful context.
Knowledge: Our personal use of information to create a map/model of the World

Bellinger et al, in their articles at systems-thinking.org, elaborates on the following so-called DIKW extension that includes wisdom, and invokes the scholarship of folks like R. Ackoff and N. Sharma on the evolution of this model for information use.



 The DIKW framework is used by many organizations. The diagram below shows the adaptation of the DIKW pyramid by US Army Knowledge Managers.




Semantic Shifts.

Going deeper now into investigating the nuances of Information Architecture and Data Architecture, I found a truly fascinating compilation of historical notes from Resmini, et al, (2011) on how our semantics shifted for the way we communicate the notion of Ïnformation Architecture." Their article speaks of Wurman's original contribution and vision that many think stills holds today, and which places Information Architecture as follows:

a.) the organization of  the patterns inherent  in data, making the complex clear; b.) the creation of the structure or map of information which allows others to find their personal paths to knowledge; c.) the emerging 21st century occupation that addresses the needs of the age focused upon clarity, human understanding, and the science of the organization of information.

Resmini's article becomes truly insightful when it narrates how the Rosenfeld-Morville team shifted the notion Information Architecture later to emphasize the importance of structure of and organization in website design i.e. what they call "Pervasive IA". Quite interesting to note that this shows how the disconnect to Wurman's past scholarship allowed for this shift to happen. At any rate, I think although the emphasis has shifted, we can still see how scholarly intuition continues to place structure and organization -- in the conceptual and logical context of Wurman's classical IA, as the overarching notion associated with what we call Information Architecture.

What about Data Architecture?

Ultimately, I come back to this old NIST Enterprise Architecture reference diagram of the late-1980's, shown below  to serve as my own reminder that we can separate Information Architecture with what we call Data Architecture. The NIST Enterprise Architecture Model is a five-layered model with each layer are defined separately but are interrelated and interwoven. The model defined the interrelation as follows:
  • Business Architecture drives the information architecture
  • Information architecture prescribes the information systems architecture
  • Information systems architecture identifies the data architecture
  • Data Architecture suggests specific data delivery systems, and
  • Data Delivery Systems (Software, Hardware, Communications) support the data architecture.
The hierarchy in the model is based on the notion that an organization operates a number of business functions, each of which requires information from a number of sources, and each of these sources may come from any one or more operational systems, which in turn stores organized data in any number of data systems.

Mapping out these nuances onto our conceptual-logical-physical modeling constructs, we can perhaps intuitively associate the conceptual, semantic, and logical relationships at the Information Architecure layer, the processing logic, approach, and data systems at the Data Architecture layer, and leave most of the suppporting technology components for database engines, storage,  and I/O infrastructure to the realm of Infrastructure Architecture.

Thus for what we usually call Information or Data Architecture, we may perhaps find this deconstruction useful.



References:

Neil Ingebrigtsen (n.d.). The Differences Between Data, Information and Knowledge. http://www.infogineering.net/data-information-knowledge.htm

Gene Bellinger, Durval Castro, Anthony Mills. Data, Information, Knowledge, and Wisdom. http://www.systems-thinking.org/dikw/dikw.htm

Andrea Resmini, Luca Rosati (2011). A Brief History of Information Architecture. Journal of Information Architecture.  http://journalofia.org/volume3/issue2/03-resmini/


Data Virtualization and the Semantic Layer

Discussion for EA874 Topic 3 > Data/Information Architecture Layer
Post # 3

Ultimately, there are 2 key questions we ask as data architects in service to our business customers:
  1. How can business gain access to the data to gain insights?
  2. How do we provide consistency and quality -- while maintaining agility?
You can easily imagine the challenges of traversing traditional data warehouse solutions as depicted in the diagram below, in order to bring data to our internal enterprise customers.














source:
http://www.slideshare.net/Denodo/data-warehousesmdmdatavirtualizationdenodopackedlunchseriessession5
Implementing Data Virtualization for Data Warehouses and Master Data Management extensions
By Denodo Technologies, published on Jan 22, 2015

As Chris Daniels points out in his article, "inadequate reporting capabilities seem to be a common characteristic of both off-the-shelf and bespoke systems. It’s a real shame because it can prevent firms from exploiting the full potential of the information they hold." He points to several factors that result in computer systems that continue to be delivered without adequate reporting facilities, and that organisations wishing to make the most of their information will need to find ways to plug the gaps. Beyond the first-order solutions using Query-Reporting tools, he goes on to point out another solution which architects would also need to keep in mind: the provision of a Semantic Layer.

What is a Semantic Layer?
A semantic layer is a business representation of corporate data that helps end users access data using common business terms. The aim is to insulate users from the technical details of the data store and allow them to create queries in terms that are familiar and meaningful.

An increasing number of reporting tools allow users to make use of a “semantic layer.” Some vendors give the semantic layer a name that is specific to their particular product (for example, Business Objects calls it a “universe”), while others describe it as a business model or metadata layer, that should make it easier and intuitive of end-users to create their own reports by themselves.  However, it is typical that the semantic layer is still mainly used by developers -- and not business users.  Unfortunately, developers are SQL-savvy enough to perceive the semantic layer as a layer of nuisance i.e. another set of technical paradigm to deal with. As a result, it can be tempting for developers to bypass the semantic layer and revert to type and write queries by hand for each and every report.

As Chris notes in his conclusion, there are real benefits to be gained through the use of semantic layers. Semantic layers are increasingly becoming available within all types of reporting environments, and their use should be encouraged for the benefit of placing the power of harvesting information in the hands of the end-users themselves. Moreover, I think the technologies that create semantic layers prepares us to appreciate the more abstracted data access solution: Data virtualization.

What is Data virtualization?
Data virtualization is defined as an umbrella term used to describe any approach to data management that allows an application to retrieve and manipulate data without requiring technical details about the data, such as how it is formatted or where it is physically located i.e. it is another layering concept.

Data virtualization software may include functions for development, operation, and/or management.

Benefits include:
  • Reduce risk of data errors
  • Reduce systems workload through not moving data around
  • Increase speed of access to data on a real-time basis
  • Significantly reduce development and support time
  • Increase governance and reduce risk through the use of policies
  • Reduce data storage required 

Drawbacks include:
  • May impact Operational systems response time, particularly if under-scaled to cope with unanticipated user queries or not tuned early on
  • Does not impose a heterogeneous data model, meaning the user has to interpret the data, unless combined with Data Federation and business understanding of the data
  • Requires a defined Governance approach to avoid budgeting issues with the shared services
  • Not suitable for recording the historic snapshots of data - data warehouse is better for this
  • Change management "is a huge overhead, as any changes need to be accepted by all applications and users sharing the same virtualization kit" 

I won't be able to expand on virtualization in this post, but I think it's sufficient to say that data architects would need to be familiar with the never-ending solutions coming out that help fulfill mission for the business, and works well with the enterprise information principle of accessibility.

References:
Chris Daniels. (2010 March).  What is a Semantic Layer and Why Would I Want One? http://www.b-eye-network.co.uk/view/12763

Data Virtualization. https://en.wikipedia.org/wiki/Data_virtualization

Semantic Layer. https://en.wikipedia.org/wiki/Semantic_layer
/