Showing posts with label Data virtualisation. Show all posts
Showing posts with label Data virtualisation. Show all posts

Tuesday, September 25, 2018

Ask not what AI can do for Data Virtualization. Ask what Data Virtualization can do for AI.


Artificial Intelligence (AI) can perform tasks that are considered “smart.” Combining AI with machine learning, it reduces programming and algorithm learning. AI has been exploding recently and taken a big leap in the last few years.

But what about Data Virtualization (DV) with AI? The first thought is usually AI optimizing queries on virtual data models. However, how can DV help AI? 

Why not leverage the virtues of Data Virtualization to streamline the AI data pipeline and process?

Example 1: Expedite the Machine Learning process when you need data from multiple sources

Suppose you want to train your Machine Learning (ML) scenario with composite data sets? Federating disparate data “on-the-fly” is a core competency of Data virtualization. Without data virtualization, of course, you could always add custom code here and there to do the federating and cleansing, but that complicates and certainly slows down the configuration of the AI as well as its execution. Even with huge amounts of data to deal with, Data Virtualization can quickly align and produce appropriate data for ML without incurring the time and risk of hand coding.

 Example 2: Define and Implement Algorithms

 With Data Virtualization a model or database is driving the iterations, DV secure write-back capability can continuously update the data sources to reflect the latest optimum values in real time.  

The learning occurs with many iterations through the logical model, usually with a huge data set. Each iteration brings more data (information) that is then used to adjust and fine tune the result set. That newest result becomes the source values of the next iteration, eventually converging on an acceptable solution.  That processing loop may be a set of algorithms that are defined in an automated workflow to calculate multiple steps and decisions. The workflow is configured in the Enterprise Enabler Process engine, eliminating the need for programming.

Example 3: DV + AI + Data Lake
Many companies are choosing to store their massive Big Data in Data Lakes. A data lake is usually treated as a dumping zone to catch any data (To learn more read - 5 Reasons to leverage EE BigDataNOW™ for your Big DataChallenges). Architects are still experimenting as to the best way to handle this paradigm, but the mentality is, “If you think we’ll be able to use it one day, throw it in.” With Data Virtualization, information can be fed in with a logical connection and can be defined across multiple data sources.



Enterprise Enabler® (EE) goes much further since it is not designed solely for Data Virtualization. The discovery features of EE can help after the fact to determine what’s been thrown in. EE can ensure that the data cleansing and management process executes flawlessly. Agile ETL™ moves data anywhere it is needed, without staging.

Since EE is 100% metadata-driven and fully extensible, it can morph to handle any integration and data management task that is needed. Enterprise Enabler is a single, secure platform that has pro-active configuration suggestions, maintains user access controls, versioning, monitoring, and logs.

From data access to product overview to data prep for business orders and reporting Enterprise Enabler is the single data integration platform that supports it all.



Tuesday, January 15, 2013

No Transformation Engine? No Agility.



Transformation Engine
     It keeps coming back to data transformation. That was the very first challenge with respect to integration that intrigued me years ago, because if all data at its various sources were in the same units of measure, had the same names, didn't need to be run through a formula to be in the same numeric context, were spelled the same, etcetera, etcetera, etcetera, then all you would have to do is get it all together where you wanted it.  I believe that the need to transform and manipulate data remains the single most important impediment to speedy, streamlined data flow throughout an enterprise. 
    The emergence of transformation engines aligned with the early Extract/Transform/Load (ETL) integration architecture pattern years ago. Unfortunately, it generally is not considered an essential function of element of other patterns, which continues to astound me. Regardless of whether the incumbent architecture is EAI, ESB, ETL, B2B, or Data Virtualization, the same issues are present, but the transformation engine is often not part of the solution. That means that all that data transformation is done by one-off coding or scripting, sometimes augmented by limited-scope conversion utilities. It seems like an “unmentionable” topic: people turning their heads the other way and pretending that it’s not important. In fact, it’s every bit as big a deal for the other patterns as it is for ETL!
     Without a transformation Engine, it is impossible to streamline the logic that makes the data work meaningfully. Without it, all data transformation yields brittle break points, impeding the ability to adjust quickly to changes in business or technical requirements and generally slows development and execution time and promotes the attitude of “it works. Don’t touch it.”  

What should you look for in a transformation engine?
·         Well, first of all, look for a transformation engine! If there is one, then consider…
·         A single transformation must handle
1.   Many sources to one (no staging required)
2.   Multiple source types (databases, Cloud apps, electronic instruments, web services, etc.etc.)
3.   Lookups, alignment, complex data manipulation, filtering, etc.
4.  “En route” data cleansing
·       Transformations must be completely metadata driven
·     All metadata should all be configured and modified in a single interface, without leaving to separate tools.

What are some typical characteristics of transformation engines that do not meet all these criteria?
·      XSLT engines, for example, operate only on data structured as XML. That means you must perform a separate transformation for each source that does not inherently handle its data in XML. Any transformation engine that requires incoming and /or output data to be in a particular format violates point 2 above.
o   Result: custom coding, or utilities that must be executed to handle the conversion at both ends of the transformation. More hand coding , and in the end to manually coded transformations
·   Classic ETL engines perform only one-to-one transformations, violating point 1 above. The only way to use these engines in a pattern that requires alignment across multiple sources is to stage the data physically or virtually.
o   Result: Development time includes designing the data model to align the sources, building the model, and implementing full transformations from each source to the model, and then from the model to the destination(s).
·    Many integration products have some amount of data conversion utilities built in.
o   Caution: These are always limited in scope and require leaving the environment to use a scripting or programming language to implement complex data manipulation.

     Just remember that without a transformation engine, you are looking at plenty of overhead and most likely the inability to merge, transform move data in real time from multiple sources. Not only will run-time performance be impacted, but the use of coding inevitably means dramatic reduction of agility in development and ability to adjust to changing requirements. 

     Now let’s take a brief look at data transformation in a data virtualization pattern. The typical picture of the concept of data virtualization is something like this:

     Each arrow represents the logic to transform the data from how it is in the particular source to how it needs to be represented in the virtual model. If there is no transformation engine, this logic will require a considerable amount of manual coding, although I concede that there may be a very small subset of the useful manifestations of this pattern that are simple enough to configure the data manipulation without any custom coding or transformation engine.  However, if you drop in the full transformation capability onto each arrow, you have a powerful, agile implementation that can be developed quickly and modified in seconds.

     A variation on this pattern brings the ability to write back to the sources: what we call bi-directional data virtualization. But that’s for another day. 
     You may want to check out Stone Bond Technologies’ Enterprise Enabler® if you want to see a product that has a transformation engine that meets all the mentioned criteria.




Sunday, October 7, 2012

2 Keys to Identifying Agile Integration Software

For years I’ve been talking about Agile Integration Software (AIS) as a class of products that enable very flexible and agile data flows across the enterprise, eliminating the clunky, expensive, time consuming infrastructure of the past. This paradigm is agnostic to integration patterns and shares metadata for all uses, such as Data Virtualization (DV), ETL, EAI, SOA, etc.

Over these years, I have come to understand that AIS is not a class of products. As it turns out, Enterprise Enabler® is the only product that has all of the characteristics necessary to fulfill the AIS vision in scope and flexibility. So much for a product-agnostic concept-oriented blog! Many of the features can be found in other products, but somehow there is always at least one critical feature missing that negates the possibility of agility for the solution.

As I think about what constitutes agility in this space, a few things come to mind that are imperatives in such a platform. The two most telling indicators lead the list:
  1. The product must have a transformation engine that aligns and transforms data a) from multiple disparate sources at once and b) in their native formats. Without this, complex integrations get little assistance from the platform itself, but are accomplished by extensive custom coding. A streamlined, high-performance data virtualization solution is impossible. Writing back to the sources becomes cumbersome at best, and live, real-time end user interaction with the endpoint simply cannot be effective. Older one-to-one transformation engines do not satisfy the needs; XSLT transformation engines also do not meet the two criteria, because all data must be converted to XML before it is transformed, and the XML output must then be converted to the destination format. Each of those conversions: to and from XML are effectively additional full transformations that generally must be accomplished with custom coding.
  2.  There must be a single Integrated Development Environment (IDE) that crosses the entire scope of functionality for all integration patterns. This IDE must incorporate the run-time engines in order to be able to design, develop, test, deploy, and monitor in the same environment. Leaving the environment for anything dramatically reduces the speed of implementation and of change, which is essential for agility, time to value, and minimizing tech debt.
Some things are only possible with AIS
The most recent feature that has been touted by all is in the Data Virtualization space, where data from multiple sources is brought together, cleaned, aligned and transformed without actually moving, staging, or copying the data anywhere, and then delivering it virtually as a “view” to an application or dashboard on demand, again without ever moving it. The most powerful functionality that simply cannot be done successfully with other technologies is data virtualization with write-back to the sources. Enterprise Enabler automatically generates data virtualization, including the ability to write back securely to the sources. These integrations are published for consumption in multiple formats for consumption, such as web services, ADO.Net driver, SharePoint External List, and others.

                          (Click picture to see full chart)

Bottom line value of Enterprise Enabler (AKA Agile Integration Software)
  • Time to Value is reduced by up to 90%
  • Tech debt, the cost of maintenance and change over time is similarly reduced
  • No need for expensive, specialized skill sets particular to a specific tool
  • Streamlined architecture inherently enables high performance
  • Lower security risks: with Data Virtualization data remains securely in the sources of record, without copies being made
  • With write-back to sources in Data Virtualization patterns, dashboards become interactive consoles instead of simply reporting tools
  • Without copies being made, a multitude of application specific databases become unnecessary, and synchronization activity is reduced
  • Latency issues are removed since applications and end users have always the most current (“live”) information.
  • Maintenance and change over time is no longer an overwhelming problem.
 We have observed that while Enterprise Enabler automatically generates bidirectional services, competing product companies tend to proliferate bi-directional lip service.

Monday, September 17, 2012

The Virtual Cycle: Bi-Directional Data Virtualization


With all the buzz about data virtualization (DV), it surprises me how little bi-directional data virtualization is discussed.  Without the ability to write back to sources, the use of DV is limited to Business Intelligence, and other reporting.  Of course, it’s hugely important for that, but when you add write-back to the sources, you are opening up a whole new world of possibilities for a new dimension of interaction with data. 

Suddenly all those dashboards become consoles where business operations can be performed, with end users not just viewing data, but correcting it, updating it, and taking action on decisions. Any application can leverage bi-directional DV to access federated data and to write back to the sources without having to know where it came from, treating it as a single entity. This capability goes a long way to reducing the time to value of many IT initiatives beyond reporting and analytics.

For those skeptics who are not already familiar with bi-directional data virtualization, the first questions are typically, “How do you handle the security to make sure users only write back when they have permission to do so?” and “What happens if  there’s a failure writing back to one of multiple sources?” 

The short answers are that end user security is handled for full CRUD capabilities using SSS or other models, and transaction rollback is managed using two phase commit or other modes.

Now, we can move on to the cool things you can do with this.  Suppose your training is frustrating and time consuming for new employees to learn how to navigate and use multiple systems that are necessary for them to handle their responsibilities. They need to log in to SAP, then the CRM, and then a scheduling system, plus a spreadsheet, all just to perform one task. You could build a browser based app that presents the relevant data from all these systems in one screen, aligned and meaningfully presented. This is the standard data virtualization, which is essentially a reporting tool. Now, turn on write capability for appropriate fields, and voila! That browser screen is a full-service, role-based application, interacting directly with backend systems and data stores just as if you were logged in to all of them. This, my friends, is the virtuous virtual cycle of bi-directional data virtualization.

Using an Agile Integration Software like Enterprise EnablerÒ you can leverage all the federated data services not just for BI, but also for Business Operations. These light weight rapidly deployed nuggets enable this third generation of data virtualization to make agile business a reality.


Tuesday, September 4, 2012

Agile Integration: Foundation for All That Hype



By definition, Agile Integration Software (AIS) is charged with accommodating all data sources, standards, etc. and moving data agilely throughout the enterprise, adjusting to changes over time. As it does, it captures information about all of the participating endpoints and modalities of data movement. The more an AIS like Enterprise EnablerÒ is used, the more it learns and the more metadata it maintains about the information that is involved in the company’s activities.  Perhaps it’s time to think about AIS as a central core of actionable metadata and a useful engine that can be leveraged for use in new initiatives as they are defined, as opposed to continuing to use AIS to adapt to whatever the new initiative independently demands.  

Consider Master Data Management (MDM) for example. Like all other hype waves, the need is there, and the initiative is fueled by the analysts in concert with a surge of new tech companies as well as the Big Players.  Your company decides it better get on the bandwagon; MDM is imperative to maintaining competitive advantage.  One of the top project architects is charged with establishing MDM, so he studies, attends symposia, and learns what it’s all about.

Next comes a technology selection phase, looking at the emerging companies, but mostly looking at the Big Guys, since that’s always safe, and besides the architect already knows the “Rep” really well. It seems the Big Guy has just bought one of the up-and-coming MDM companies, so they are off to the races.  After a couple of years, the architects begin contemplating questions like, “How does this MDM solution relate to the last decade’s SOA path?”   Of course, at this point it’s obvious that SOA is pretty closely related, or could be, had the two initiatives not each been addressed in its own blissful vacuum.  How does the MDM metadata relate to SOA, and how does it relate to the metadata that your multiple integration platforms require?  Some, with wishful thinking, fall for the Big Guys’ claims of interoperability across all the products of the companies they bought.  In the Quadrant or Wave, they cover all the requisite features, but Alas, poor Yorick! Those mighty features are compartmentalized and each discipline is a separate product with separate underpinnings unable to work cleanly together.  


Let’s look at this from a different perspective. With Agile Integration Software comes a complete flexible integrated metadata stack for use and reusability across all the historic and forward-looking integration schema and models.  Instead of integration adapting to the stand-alone fragmented hype solutions, leverage the power of an existing AIS platform  that brings together the disciplines of Data and Application Integration, Application Development, as well as all the special initiatives of SOA, MDM, B2B, Middleware, Virtualization, Federation, Cloud, Change Management and Big Data, all leveraging the AIS integrated metadata stack.

That means eliminating lots of steps to accomplish any task.  With visibility via the metadata stack, a Master Data definition can be combined with all the associated sanctioned sources and all the related business rules and security. Auto-generated bi-directional web services can handle security and rollback to federated sources. SOA is just another mechanism to make an integration or process available for consumption. Different Data Masters and integrations are chosen at run time based on the current state of any other activity throughout the system.

And finally, AIS monitors for change throughout the metadata stack, validating against the actual endpoints and determining the potential impact and remediating and reconciling as necessary across all those hype wave initiatives. This means stability with agility in your IT infrastructure.



Wednesday, July 25, 2012

You've Got Big Data! Use it in the Cloud… Virtualized!


I get why Big Data is capitalized. It's even getting to the point where it could rate all caps, "BIG DATA." The part I don't get is why this is such a new big deal. Big Data has been around for a long time. Just think about the huge bodies of data that scientists have been gathering and analyzing just fine over the last who-knows-how many years.  The explosion of social media certainly brings a new dimension on rapidly growing, potentially valuable data, and poses the challenge of making it as valuable as the Big Data you already have.

The most valuable information in your company is probably not social media, but rather your corporate data. What about all that Big Data lurking in your corporate systems, data warehouses and across your multiple divisions? Your corporate Big Data is not like the scientific nor the Social Media Big Data:

o It's not in one place like the scientific data or the Social Media Cloud.
o It's across different organizations and geographies
o It looks different each place

I think of the Big Data focus as being on how to make it useful, not just for business analytics, but more importantly, to make it actionable to your elastic and agile enterprise. This means your solutions needs to be something that can be configured in very short time frames, which means eliminating custom coding. 

In order to be effective at leveraging your Big Data assets, integrating Big Data sources must be treated with streamlined transformation, federation, and virtualization, in an extremely unified architecture. This is necessary to provide the high performance requirement and is achieved by eliminating the classic steps through multiple components for extraction, transformation, staging to federate, and transforming again to align to the destination. It also requires the ability to address complex data manipulation with inline quality checks and error management.

o Access data from multiple disparate systems
o Leverage premises based and cloud based data virtually together in your own cloud
o Federate as data is being accessed from multiple sources, lookup tables
o Virtualize your data. Make it consumable by applications on premise or in the cloud live, virtually. This means that there is never a copy made of your data. (Of course, if you should want to actually send data, that's fine, too!)

Now, about that cloud…

Leveraging the cloud brings a host of opportunities not usually found within the walls of a corporate enterprise. The cloud brings an elasticity to try new architectures, explore new applications and meld your legacy data with Social and New Media for breakthrough business tools such as Predictive Analytics.

The trouble with the cloud is that we are concerned about maintaining the security of our data.  We feel this new technology should bring a tectonic shift in the way we think about integration to the cloud and in the value it can bring, without paying the generally accepted price.

To the best of my knowledge, there is only one product on the market that actually can do this: Stone Bond Technologies' Enterprise Enabler. With Enterprise Enabler you have the solution for Federating, Virtualizing and Leveraging your Big Data in the cloud.

o Your Data remains your data. Your data is never physically copied or placed to the cloud or anywhere along the way. It is accessed, transformed, aligned and virtualized in a single execution
o The data integration is end-user aware. The end user of the cloud application will see only the data he has authority to see.
o Your data is immediately actionable. Users of the cloud applications can update and add data on their screens, and the data is passed back to the source for updating, provided the user has authorization to do that.

A use-case example is a financial institution using this technology to supplement the Salesforce.com data with on-premise information that needs to be seen by the loan personnel alongside the cloud data.


Friday, June 22, 2012

Inspiring New Patterns for Integration

Integration is an age-old problem that doesn't have much  opportunity to get people excited.  It's the same old problem, even if you have new ways to solve it. Over the years, we have tended to just keep massaging old "patterns" pretending like they are new  configurations of moving data around and stuffing it into databases only to take it out again. It takes totally new concepts for the underlying integration architecture to give birth to new patterns that can renovate solution concepts, and more importantly, can enable new business uses that have not been possible before.

Data federation is now finally coming of age enough to have at least a small number of products that can deliver live virtual data federated across disparate systems.

[[Wait..."Live virtual data?" If it doesn't exist, can it ever be dead? Or even stale? If it's live can it be virtual? 
Definition: Virtual
1. Existing in the mind, especially as a product of the imagination.
2. Computer Science Created, simulated, or carried on by means of a computer or computer network: virtual conversations in a chatroom

And I like the origins, from old English and Latin words  that mean effective, excellence, virtue.]]

But I digress!  And maybe it's a good thing, because we tend to throw all these terms around with great authority, completely confusing those who just want to understand what this new breed of software does.  To add to the confusion, "virtualization" is a pervasive term in hardware/operating system jargon, meaning essentially a stand-alone computer that is simulated on another, bigger computer or in the cloud. Now that I have said that, forget it, because that has nothing to do with virtualization with respect to integration.

 For the purpose of this discussion, let's say that in integration speak, "virtual" means that there is never a copy made of the data, and that it never moves anywhere. A virtual view of data allows you to look at data on a dashboard, web page, or the like, without storing it anywhere. When the screen is refreshed, it’s gone. A "federated" virtual view is a virtual view that is consolidated and aligned data from multiple sources. "Live" means that the data comes directly from the sources without staging it in any data store or virtual database along the way. "Bi-directional federation and virtualization" means that the virtual federated data can be interacted with. For example, an end user can update or correct the data on the screen, and it is sent back to the sources as updates, still without staging en route either direction.



With this new technology,  federated data no longer must be staged in order to transform and align it from multiple sources and send it to an ephemeral display endpoint, which presents the live data to the end user ("virtually").  This does wonders for Business Intelligence and Business Analytics.

But going beyond BI, there is one product that enables these displays to take end user's  changes and updates  and pass then securely back to the sources. This is what turns dashboards into control centers. http://tinyurl.com/bqfnzw9

Imagine, though, what a mind shift this requires, and the uncertainty and fear, even, of have a portal like SharePoint, for example, that can effectively become the a single window into all the applications a business user needs. He is not just looking at analytics and drilling down for detail input.


With the ability to interact with backend systems, the business user is no longer just a dead-end in the data flow,  absorbing, and absorbing, and only interacting with the business intelligence software to see more data to absorb. Instead, with bi-directional federation and virtualization, as he sees data that needs to be updated, and draws conclusions from the information, he  can enter his updates as if he were working directly with the backend source systems.

Imagine having a portal that presents you with exactly the information you need to perform your business role. This may include information from SAP, Salesforce,  and, let's say, a scheduling system. From that one screen on your portal, you may update a customer's address, look up prices, and place an order. All of the relevant information will be sent back into the appropriate backend systems automatically. You don't have to learn how to navigate in three separate systems, and you don’t have to figure out how the information is required to be handled differently with each system. Instead, all the work happens behind the scenes, and you only work with easy-to-use aliases for the virtual consolidated data.
  
While this new paradigm inspires new patterns for integration, it also brings pessimists, nay-sayers and even realists who find worrisome points to ponder.   For example, if you have everyone in the company hitting the backend systems directly, could this bring those systems to their knees? See my white paper, Inspiring New Patterns for Integration about a complex pattern that rationalizes and reduces the hits to the backend systems.

With this new technology, there's a whole new world of ideas and opportunities for streamlining integration implementations, which is what has been consistently the largest investment and risk with virtually every IT project since the beginning of time (as we know it).