Showing posts with label Agile Integration. Show all posts
Showing posts with label Agile Integration. Show all posts

Friday, August 22, 2014

Cache as Cache Can

Caching is one of those afterthoughts, when you know you have a great solution, but you start wondering about performance. Since caching is about moving data at varying speeds, it is (or should be) an inherent feature and responsibility of any integration solution. You will find that a truly Agile Integration Software, such as Enterprise Enabler, makes it easy to configure a wide range of models of caching, and to adjust as your requirements change.

Agile Integration covers everything from ETL through near real time Bi-directional Data Virtualization (DV), all with federation at the core, so caching can be implemented anywhere, end-to-end, in the data flow cycle.

The Continuum of Caching
According to Wikipedia, cache is a “component that transparently stores data so that future requests for that data can be served faster.” I think of it as being any data store, however static or ephemeral, however Big or small, and whether the cached data is exactly in the source form, perhaps to be federated on the way out, or federated already as the endpoint needs or the Master Data form, ready to go on to its destination, or somewhere else in the flow of the data. The specific subset of data to be cached should be optimized to ensure the greatest efficiency, minimal size, and highest reusability. The transparency comes in because, in the big scheme of things, the destination, the consumer, or the workflow steps never need to know the data is not all coming live from the original sources.

This is where data federation and Data Virtualization add to the flexibility of caching. Agile Data Virtualization supports cache as one of the sources, so there could be DV involved to create the cache, whether in-memory, on disk or in a database, and then that cache can be used as one source in a federation that is delivered either on-demand or event-triggered.

Today, most people talk about cache as being refreshed as opposed to accumulating a history, however with all the options that can be configured, this is actually a  realistic and sometimes useful consideration. You can see that the possible combinations are many, clearly enough that one must be careful not to get tangled up, and not to lose sight of the original objectives of caching! 

One could easily argue that caching is more like ETL than like Data Virtualization, however DV often requires caching more than other integration patterns, since the uses generally expect rapid, “live” data, without latency. When the rubber meets the road, in many situations, caching is the only way to ensure that a DV solution with many users does not bring the source applications “to their knees.” This is why Agile Integration Software, which combines all the integration patterns, solves Data Virtualization problems better than pure DV platforms.

What do you need to determine before you configure caching?
·         Which data to cache
·         Why you selected caching this particular data
·         Where to cache – memory, disk, database, etc
·         How often to refresh – schedule, event, as soon as available
·         Where in its path to cache – directly from source, partially processed, before or after federation, endpoint ready, as part of a Master Data definition
·         When to release from cache- as soon as read, as soon as a particular set of consumers have read
·         Is the cache subject to bi-directional data flow

When should you plan to Cache?
First of all, keep in mind that if you don’t identify your caching needs up front, with Agile Integration Software, you can easily add it as your traffic grows and the parameters get to point where it’s needed.  Particularly when you are using Data Virtualization, and are hitting backend source systems live at each request, you should take a close look at the needs and best approaches to caching. You should consider caching in situations where:
·         You are concerned that too much traffic hitting mission critical or any sources could adversely impact the performance of those systems.
·         You are concerned about the response times for end users.
·         You need to have the same value throughout a process where you might be accessing it multiple times

What to Cache?
·         Data that doesn’t need to be real-time
·         Data that you want to ensure the same snapshot is used for different things
·         Data that changes so slowly that having it real-time doesn’t matter. You could refresh the cache once an hour or day or month, even.

Agile Caching
Agile Integration Software offers a wide range of options for caching, with ease of configuring even complex caching patterns without custom programming. With the ability to select full data sets, specific fields,  mixed in-memory and on-disk caching, and all combinations, including conditional full workflow-driven caches,  great architecting doesn’t have to be constrained by what is practical to implement.

Friday, April 11, 2014

Agile Big Data is Coming

Agility with Big Data?
Shiny but not Agile, perhaps destined to never be Agile.
Just look at ETL. Never, never agile.
So Clunky. After all these years, why is it so Clunky? So Clunky it’s fragile.

Big Data follows in the footsteps.
Big feet. Big hype, Big opportunity
for discoveries otherwise unfathomed.

Great programmers performing great feats,
Coding, coding new horizons… Open Source assist.
No hope for Agility.
Where are the trailblazers of Agility?

Ah, yes!
They’ve conquered ETL, EAI.
They’ve conquered Data Virtualization.
Agile Big Data is coming.

Are great programmers actually an impediment to Agility in complex software problems?  I‘m not talking about Agile development, but rather Agile solutions. Great programmers like to be on the leading edge of the latest new shiny trend. I remember the rush of programmers to work on the Y2K-driven Enterprise Application Integration and ERP rage. Now we’re all over Big Data.

It takes time for the market, business, and technologists to evolve the thinking about what the Shiny thing really is for, what it means, and what the technology requirements are. So, who jumps in first? The really good programmers who want to blaze the trails and solve the problems first-hand. Unfortunately, the early players can’t benefit from the evolution and maturing of the Shiny and what the requirements will really be. By the time it’s clear what it takes to genericize the problem, the Great programmers have moved on to the next trend, and are certainly not going to step back and solve the problem again in a tool or platform that hides all the repetitive work behind the scenes. 

Product companies then step in to harvest the rest of the hype curve. Leveraging their Big Name, they make plenty of revenues on hard-coded, specialized solutions where custom coding is accepted as the only way to solve the problem. In my view, a timeless benchmark for Agility is the minimization of custom code. The more coding involved in generating a solution, the farther away that solution is from Agile.

The growing love affair with Open Source code is a huge step backward for the cause of Agility. Programmers love coding. I, for one, love coding, too. But I HATE having to code essentially the same thing over and over, with just a couple of tweaks difference each time. This is exactly what integration is all about: lots of small (very important) tweaks to the same code over and over.

That is why Enterprise Enabler exists. That is why we hide all the technical details behind the scenes of our agile integration software. Any time a programmer has to do essentially the same thing more than a handful of times, it is automated with tweaks configurable in a UI. Programmers should be programming exciting, innovative new things, not laboring over repetitive, boring scripts and code changes, and maintenance.

We’re applying the same philosophy to automate Big Data integration and analysis.

Monday, August 26, 2013

Convergence is, at Best, Asymptotic

No, not asymptomatic. Asymptotic. “Convergence” is a term we hear these days in IT. The convergence of Data Integration, in particular, is the one I care about. In the analysts’ vernacular, converged integration seems to mean a product, company, or platform that handles all modes of data integration – ETL, EAI, ESB, DV et al.

By definition, “convergence” means a coming together, which clearly implies the parts started from other places and came voluntarily or were coerced into coming together. Just looking at history, companies like IBM, Oracle, and Informatica absorbed outside companies and products to nominally have a product suite with all the modes. Here’s how it works: Need ETL? Buy Ascential, rename it, give it a massage with hot towels, then say, “Voilà! Voilà! Voilà!” and your product now covers ETL, too! Or, take a Data Virtualization product, write a bunch of code, and again, “Voilà!” and your product covers ETL, too.

With all due respect to the analysts, the word convergence may describe the reality of most companies incorporating more and more integration modes, but keep in mind though, that in mathematics convergence means getting closer and closer but never quite getting there: asymptotic [translation: Close, but no cigar!].

I am certain that the analysts do not mean Convergence in the mathematical sense. The term is quite useful for establishing a powerful vision of dramatically reduced time-to-value, clean architectures, flexible integration patterns, and highly streamlined change management and maintenance over time.


If you read my last blog http://tinyurl.com/lmwtzth you probably recognize that there is a significant difference with Enterprise Enabler®. It was designed from the ground up with a powerful core that handles all the common functionality of the range of modes, with implicit data federation across disparate sources (databases, electronic instruments, spreadsheets, data warehouses, ERP systems, cloud services and so on). 

 That core is the common root of all data integration modes, a bit like the trunk of a tree that has any number of branches and leaves, instead of trying to converge a bunch of branches by stuffing them all in an opaque vase and pretending like they have a single trunk. That spells trouble, and doesn't even come close to asymptotic. Probably not asymptomatic, either. 

Friday, November 2, 2012

Tech Debt Out-of-the-Box ... "And all the Ills of Integration-kind were Unleashed"

We often talk about having cool capabilities “out of the box,” which is a good thing. That means that you don’t have to do anything but a quick install and you can start using the feature. That is, unless you are talking about Legacy Integration Software (LIS), in which case, when you first “opened the box,” it began spewing Tech Debt before anything else happened. All the ills of integration-kind were unleashed. Years later, you are still prisoner to your Pandora’s Box.

You launch a new project using Legacy Integration Software. First open Pandora’s Box:
        1. Install new instance and all related tools
        2. Apply 64 patches; you can implement the work-arounds in a few months when you start actually developing the integrations.
        3. Send team to a few weeks of Legacy Integration University
        4. Better hire a few consultants, too.
Tech Debt abounds already!

Below is an actual post on a recent Integration Consortium’s LinkedIn Group discussion. The topic has to do with updating customer communications to a standard XML format as opposed to legacy file ftps. The Legacy Integration Software limits the options. Adding XML to the mix means that the LIS needs additional work.

**************
"We already have in house Informatica footprint. So we plan to use that for generating all outbound files. Here is the concern.
We have two options for generating these files.
1) DB -> Informatica -> Standard XML -> XSLT -> Custom File
2) DB -> Informatica -> Custom File
First option provides the benefit of standardization/canonical information model, long term migration path of custom files to standardized xml, and less development since only one Informatica process is required and all custom format are through XSLT.
Problem we see with this approach is the potential performance issue; outbound XML file size in some cases is more than 1GB due to XML tags while the corresponding custom file is a 50MB or so. Second issue is an additional hop that makes support/troubleshooting activities a little harder i.e. where/why a file generation process failed."
*******************

No one should have to think about these things. Agile Integration Software (AIS) like Stone Bond’s Enterprise Enabler would require only a single process for this solution. The differences in the mapping and destination format required would be handled by passing the customer ID, which determines either which map to run or passes variables directly into the transformation engine at run-time to modify the actions. The same process can alternatively step through a standard XML, although the value of doing that escapes me. Performance would not be an issue, and troubleshooting through the streamlined solution is simplified. Stone Bond customers implement such B2B transactions using DBAs as opposed to specially-skilled programmers.

Why, in the twenty-first century, do you have to jump through hoops to get data wherever you want it whenever you want it? Here we are, musing over the “leading edge” Big Data hype and allocating millions of dollars for pilot projects next year, when we can’t even get clean, quick, agile data exchange with our customers and business partners. Does that make sense? I don’t think so. Isn’t it time to embrace twenty-first century technology and start eliminating the Tech Debt you have accumulated instead of continuing on a path that parallels the national debt?

Agile Integration is easy to try out. You do owe it to your shareholders.



Friday, June 22, 2012

Inspiring New Patterns for Integration

Integration is an age-old problem that doesn't have much  opportunity to get people excited.  It's the same old problem, even if you have new ways to solve it. Over the years, we have tended to just keep massaging old "patterns" pretending like they are new  configurations of moving data around and stuffing it into databases only to take it out again. It takes totally new concepts for the underlying integration architecture to give birth to new patterns that can renovate solution concepts, and more importantly, can enable new business uses that have not been possible before.

Data federation is now finally coming of age enough to have at least a small number of products that can deliver live virtual data federated across disparate systems.

[[Wait..."Live virtual data?" If it doesn't exist, can it ever be dead? Or even stale? If it's live can it be virtual? 
Definition: Virtual
1. Existing in the mind, especially as a product of the imagination.
2. Computer Science Created, simulated, or carried on by means of a computer or computer network: virtual conversations in a chatroom

And I like the origins, from old English and Latin words  that mean effective, excellence, virtue.]]

But I digress!  And maybe it's a good thing, because we tend to throw all these terms around with great authority, completely confusing those who just want to understand what this new breed of software does.  To add to the confusion, "virtualization" is a pervasive term in hardware/operating system jargon, meaning essentially a stand-alone computer that is simulated on another, bigger computer or in the cloud. Now that I have said that, forget it, because that has nothing to do with virtualization with respect to integration.

 For the purpose of this discussion, let's say that in integration speak, "virtual" means that there is never a copy made of the data, and that it never moves anywhere. A virtual view of data allows you to look at data on a dashboard, web page, or the like, without storing it anywhere. When the screen is refreshed, it’s gone. A "federated" virtual view is a virtual view that is consolidated and aligned data from multiple sources. "Live" means that the data comes directly from the sources without staging it in any data store or virtual database along the way. "Bi-directional federation and virtualization" means that the virtual federated data can be interacted with. For example, an end user can update or correct the data on the screen, and it is sent back to the sources as updates, still without staging en route either direction.



With this new technology,  federated data no longer must be staged in order to transform and align it from multiple sources and send it to an ephemeral display endpoint, which presents the live data to the end user ("virtually").  This does wonders for Business Intelligence and Business Analytics.

But going beyond BI, there is one product that enables these displays to take end user's  changes and updates  and pass then securely back to the sources. This is what turns dashboards into control centers. http://tinyurl.com/bqfnzw9

Imagine, though, what a mind shift this requires, and the uncertainty and fear, even, of have a portal like SharePoint, for example, that can effectively become the a single window into all the applications a business user needs. He is not just looking at analytics and drilling down for detail input.


With the ability to interact with backend systems, the business user is no longer just a dead-end in the data flow,  absorbing, and absorbing, and only interacting with the business intelligence software to see more data to absorb. Instead, with bi-directional federation and virtualization, as he sees data that needs to be updated, and draws conclusions from the information, he  can enter his updates as if he were working directly with the backend source systems.

Imagine having a portal that presents you with exactly the information you need to perform your business role. This may include information from SAP, Salesforce,  and, let's say, a scheduling system. From that one screen on your portal, you may update a customer's address, look up prices, and place an order. All of the relevant information will be sent back into the appropriate backend systems automatically. You don't have to learn how to navigate in three separate systems, and you don’t have to figure out how the information is required to be handled differently with each system. Instead, all the work happens behind the scenes, and you only work with easy-to-use aliases for the virtual consolidated data.
  
While this new paradigm inspires new patterns for integration, it also brings pessimists, nay-sayers and even realists who find worrisome points to ponder.   For example, if you have everyone in the company hitting the backend systems directly, could this bring those systems to their knees? See my white paper, Inspiring New Patterns for Integration about a complex pattern that rationalizes and reduces the hits to the backend systems.

With this new technology, there's a whole new world of ideas and opportunities for streamlining integration implementations, which is what has been consistently the largest investment and risk with virtually every IT project since the beginning of time (as we know it).  

Wednesday, May 30, 2012

SAP Best Practices Streamlined by Agile Integration


With the augmentation of Business Object Data Integrator with ETL and the Rapid Marts, the orderliness of SAP reaches further out toward legacy sources.  SAP's ECC Master Data Structures, SAP AIO BP ("SAP All-in-One Best Practices" ), among other things, define and document IDOC structures for master data elements.

These tools work reasonably well as long as you are dealing with single source to the master definition.  Unfortunately, what lies beyond SAP in your environment is anybody's guess. If  more than half of these work for you "out-of-the-box,"  you must live a charmed life!   More than likely, your legacy sources impose a huge need for custom programming to align multiple data sources and transform them to meet the Best Practices requirements.  The variability and unknowns of your environment  probably will never be able to be addressed easily by SAP's tools, and yet that necessary custom programming consumes a good chunk of your budget and timeline.

Typical integration certification is for one source to SAP. Both the source and the destination have pre-defined schemas, which is virtually never true in real life. Again, with this, we see that this type of adapter seldom works without custom programming adjustments.


Whatever approach you are planning on using for your migration of legacy data to SAP,  you are most certainly facing a formidable, time consuming, and costly exercise.  SAP has greatly improved their tools and best practices over the years, and highly trained and experienced SAP consultants, architects, and programmers have contributed to streamlining the task of data migration. 

This is where Agile Integration Software (AIS), with its light-of-foot cross-application federation, can make a huge difference to your overall  success in meeting your objectives, and in ensuring that as changes occur, they can be accommodated quickly and easily.  That point where data migration becomes a nightmare is where AIS like Enterprise Enabler are the best solution.  The earlier it is introduced into the plan, the better, as some project steps may simply no longer be necessary.

Given the magnitude of a migration from Legacy systems to SAP, and of the ongoing integration required, it has to be a worthwhile effort to establish what the value could be to your project. With a high probability of a minimum of 50% reduction in manpower and elapsed time to implement the integrations, and a rapid, easy path to demonstrating this, why take the risk of NOT investigating?

You owe it to your shareholders, and to yourself to seriously consider AIS!