Showing posts with label bi-directional data virtualization. Show all posts
Showing posts with label bi-directional data virtualization. Show all posts

Friday, August 22, 2014

Cache as Cache Can

Caching is one of those afterthoughts, when you know you have a great solution, but you start wondering about performance. Since caching is about moving data at varying speeds, it is (or should be) an inherent feature and responsibility of any integration solution. You will find that a truly Agile Integration Software, such as Enterprise Enabler, makes it easy to configure a wide range of models of caching, and to adjust as your requirements change.

Agile Integration covers everything from ETL through near real time Bi-directional Data Virtualization (DV), all with federation at the core, so caching can be implemented anywhere, end-to-end, in the data flow cycle.

The Continuum of Caching
According to Wikipedia, cache is a “component that transparently stores data so that future requests for that data can be served faster.” I think of it as being any data store, however static or ephemeral, however Big or small, and whether the cached data is exactly in the source form, perhaps to be federated on the way out, or federated already as the endpoint needs or the Master Data form, ready to go on to its destination, or somewhere else in the flow of the data. The specific subset of data to be cached should be optimized to ensure the greatest efficiency, minimal size, and highest reusability. The transparency comes in because, in the big scheme of things, the destination, the consumer, or the workflow steps never need to know the data is not all coming live from the original sources.

This is where data federation and Data Virtualization add to the flexibility of caching. Agile Data Virtualization supports cache as one of the sources, so there could be DV involved to create the cache, whether in-memory, on disk or in a database, and then that cache can be used as one source in a federation that is delivered either on-demand or event-triggered.

Today, most people talk about cache as being refreshed as opposed to accumulating a history, however with all the options that can be configured, this is actually a  realistic and sometimes useful consideration. You can see that the possible combinations are many, clearly enough that one must be careful not to get tangled up, and not to lose sight of the original objectives of caching! 

One could easily argue that caching is more like ETL than like Data Virtualization, however DV often requires caching more than other integration patterns, since the uses generally expect rapid, “live” data, without latency. When the rubber meets the road, in many situations, caching is the only way to ensure that a DV solution with many users does not bring the source applications “to their knees.” This is why Agile Integration Software, which combines all the integration patterns, solves Data Virtualization problems better than pure DV platforms.

What do you need to determine before you configure caching?
·         Which data to cache
·         Why you selected caching this particular data
·         Where to cache – memory, disk, database, etc
·         How often to refresh – schedule, event, as soon as available
·         Where in its path to cache – directly from source, partially processed, before or after federation, endpoint ready, as part of a Master Data definition
·         When to release from cache- as soon as read, as soon as a particular set of consumers have read
·         Is the cache subject to bi-directional data flow

When should you plan to Cache?
First of all, keep in mind that if you don’t identify your caching needs up front, with Agile Integration Software, you can easily add it as your traffic grows and the parameters get to point where it’s needed.  Particularly when you are using Data Virtualization, and are hitting backend source systems live at each request, you should take a close look at the needs and best approaches to caching. You should consider caching in situations where:
·         You are concerned that too much traffic hitting mission critical or any sources could adversely impact the performance of those systems.
·         You are concerned about the response times for end users.
·         You need to have the same value throughout a process where you might be accessing it multiple times

What to Cache?
·         Data that doesn’t need to be real-time
·         Data that you want to ensure the same snapshot is used for different things
·         Data that changes so slowly that having it real-time doesn’t matter. You could refresh the cache once an hour or day or month, even.

Agile Caching
Agile Integration Software offers a wide range of options for caching, with ease of configuring even complex caching patterns without custom programming. With the ability to select full data sets, specific fields,  mixed in-memory and on-disk caching, and all combinations, including conditional full workflow-driven caches,  great architecting doesn’t have to be constrained by what is practical to implement.

Wednesday, June 5, 2013

What is Data Virtualization?

“Virtualization” is everywhere but nowhere. The term is virtually ubiquitous. The first use I remember when computers got into the picture was “virtual reality,” when we computationally rendered 3D worlds in 2D, complete with lighting models and all that. The good thing is that all those complicated algorithms are now encapsulated for cool things and used by beginner gamers. In those days, we had to actually calculate every pixel ourselves. But I digress.

First, let’s clarify that “virtualizing data” means putting in the cloud or elsewhere in order to eliminate some of the hassles of its existence and maintenance. That has nothing to do with data virtualization, which is a term that I believe is still evolving.

Data Virtualization, according to Rick van der Lans, who literally wrote the book, is “the technology that offers data consumers a unified, abstracted, and encapsulated view for querying and manipulating data stored in a heterogeneous set of data stores.**”  

As the discipline matures, he is expanding his view, as in his new white paper, Creating an AgileData Integration Platform Using Data Virtualization. Definitely recommended reading. 

The “unified, abstracted, and encapsulated view,” from his original definition is the core concept, in my opinion, of data virtualization. In other words, there is a mechanism to bring together, or “federate,” virtually, data from many sources in a way that is useful. This means that the data is federated without creating a physical or cached staging database, but is aligned, transformed, and made available for use. So, for example, you may have a SharePoint BCS application that needs data from SAP, Oracle, and Salesforce.com. Data virtualization will provide a mechanism to merge all of the data into the form necessary for the end user’s interaction in SharePoint. The data is federated “on the fly” and delivered virtually to a web page, on-demand upon refresh of the screen. Think about the security of the backend data that has been accessed…it never actually moves from its original source! Data virtualization also includes writeback to the sources (with end user security, but that’s for another blog) so that an end user can, for example correct his phone number or address, sending it as an update directly to the backend source. )See more examples at http://tinyurl.com/a3wkffc)

You can see that this description expands the definition to include any sources, not just data stores, although the focus of most data virtualization products is BI, in which case, that limitation it makes sense. The BI view of using data virtualization usually is with respect to federating relational databases for the sole purpose of querying. The tools that were designed assuming that constraint have some difficulty accommodating the expanding definition. 

In addition to evolving from federating data stores to federating any kinds of disparate sources, data virtualization is shedding the concept of “on-demand” only. Now federated data is not just available by web services, ado.net, ODBC, JDBC, etc, but for any type of data integration, such as ETL, EAI, etc. 

In fact, it is the “data virtualization” concept of federation that becomes the kingpin for “Convergence,” as Gartner is wont to say, of all integration modalities in a single toolset, sharing metadata and business rules across all. 

**Rick F. van der Lans, Data virtualization for Business Intelligence Systems, Morgan Kauyfmann, 2012

Sunday, October 7, 2012

2 Keys to Identifying Agile Integration Software

For years I’ve been talking about Agile Integration Software (AIS) as a class of products that enable very flexible and agile data flows across the enterprise, eliminating the clunky, expensive, time consuming infrastructure of the past. This paradigm is agnostic to integration patterns and shares metadata for all uses, such as Data Virtualization (DV), ETL, EAI, SOA, etc.

Over these years, I have come to understand that AIS is not a class of products. As it turns out, Enterprise Enabler® is the only product that has all of the characteristics necessary to fulfill the AIS vision in scope and flexibility. So much for a product-agnostic concept-oriented blog! Many of the features can be found in other products, but somehow there is always at least one critical feature missing that negates the possibility of agility for the solution.

As I think about what constitutes agility in this space, a few things come to mind that are imperatives in such a platform. The two most telling indicators lead the list:
  1. The product must have a transformation engine that aligns and transforms data a) from multiple disparate sources at once and b) in their native formats. Without this, complex integrations get little assistance from the platform itself, but are accomplished by extensive custom coding. A streamlined, high-performance data virtualization solution is impossible. Writing back to the sources becomes cumbersome at best, and live, real-time end user interaction with the endpoint simply cannot be effective. Older one-to-one transformation engines do not satisfy the needs; XSLT transformation engines also do not meet the two criteria, because all data must be converted to XML before it is transformed, and the XML output must then be converted to the destination format. Each of those conversions: to and from XML are effectively additional full transformations that generally must be accomplished with custom coding.
  2.  There must be a single Integrated Development Environment (IDE) that crosses the entire scope of functionality for all integration patterns. This IDE must incorporate the run-time engines in order to be able to design, develop, test, deploy, and monitor in the same environment. Leaving the environment for anything dramatically reduces the speed of implementation and of change, which is essential for agility, time to value, and minimizing tech debt.
Some things are only possible with AIS
The most recent feature that has been touted by all is in the Data Virtualization space, where data from multiple sources is brought together, cleaned, aligned and transformed without actually moving, staging, or copying the data anywhere, and then delivering it virtually as a “view” to an application or dashboard on demand, again without ever moving it. The most powerful functionality that simply cannot be done successfully with other technologies is data virtualization with write-back to the sources. Enterprise Enabler automatically generates data virtualization, including the ability to write back securely to the sources. These integrations are published for consumption in multiple formats for consumption, such as web services, ADO.Net driver, SharePoint External List, and others.

                          (Click picture to see full chart)

Bottom line value of Enterprise Enabler (AKA Agile Integration Software)
  • Time to Value is reduced by up to 90%
  • Tech debt, the cost of maintenance and change over time is similarly reduced
  • No need for expensive, specialized skill sets particular to a specific tool
  • Streamlined architecture inherently enables high performance
  • Lower security risks: with Data Virtualization data remains securely in the sources of record, without copies being made
  • With write-back to sources in Data Virtualization patterns, dashboards become interactive consoles instead of simply reporting tools
  • Without copies being made, a multitude of application specific databases become unnecessary, and synchronization activity is reduced
  • Latency issues are removed since applications and end users have always the most current (“live”) information.
  • Maintenance and change over time is no longer an overwhelming problem.
 We have observed that while Enterprise Enabler automatically generates bidirectional services, competing product companies tend to proliferate bi-directional lip service.

Monday, September 17, 2012

The Virtual Cycle: Bi-Directional Data Virtualization


With all the buzz about data virtualization (DV), it surprises me how little bi-directional data virtualization is discussed.  Without the ability to write back to sources, the use of DV is limited to Business Intelligence, and other reporting.  Of course, it’s hugely important for that, but when you add write-back to the sources, you are opening up a whole new world of possibilities for a new dimension of interaction with data. 

Suddenly all those dashboards become consoles where business operations can be performed, with end users not just viewing data, but correcting it, updating it, and taking action on decisions. Any application can leverage bi-directional DV to access federated data and to write back to the sources without having to know where it came from, treating it as a single entity. This capability goes a long way to reducing the time to value of many IT initiatives beyond reporting and analytics.

For those skeptics who are not already familiar with bi-directional data virtualization, the first questions are typically, “How do you handle the security to make sure users only write back when they have permission to do so?” and “What happens if  there’s a failure writing back to one of multiple sources?” 

The short answers are that end user security is handled for full CRUD capabilities using SSS or other models, and transaction rollback is managed using two phase commit or other modes.

Now, we can move on to the cool things you can do with this.  Suppose your training is frustrating and time consuming for new employees to learn how to navigate and use multiple systems that are necessary for them to handle their responsibilities. They need to log in to SAP, then the CRM, and then a scheduling system, plus a spreadsheet, all just to perform one task. You could build a browser based app that presents the relevant data from all these systems in one screen, aligned and meaningfully presented. This is the standard data virtualization, which is essentially a reporting tool. Now, turn on write capability for appropriate fields, and voila! That browser screen is a full-service, role-based application, interacting directly with backend systems and data stores just as if you were logged in to all of them. This, my friends, is the virtuous virtual cycle of bi-directional data virtualization.

Using an Agile Integration Software like Enterprise EnablerÒ you can leverage all the federated data services not just for BI, but also for Business Operations. These light weight rapidly deployed nuggets enable this third generation of data virtualization to make agile business a reality.