Tuesday, May 12, 2009

On Gartner's EPN Reference Architecture


Today is a holiday (for children, no vacation for adults..) called Lag Baomer, the highlight (besides not going to school) is that last night all children have gathered around bonfires, as seen in the picture. Fun.

Recently Gartner has published a report called "A Gartner Reference Architecture for Event Processing Networks".

On the positive side, it seems that the concept of EPN, as an underlying model for event processing is catching. The readers of the Blog may realize that I am in the opinion that we need an agreed upon conceptual and execution model for event processing (the same role that the relational model assumes in relational database, however, I never believed that the relational model per se, is appropriate also as the model behind event processing). The book I am writing now "Event Processing in Action" concentrates around the notion of EPN, and a deep dive into construction of EPN-based application.

Reading Gartner's report I found some slight differences between the way they describe EPN, and my own description. In the Gartner report they define a term called "dissemination network" that consists of event processing agents, channels and event flow among them, and then they define EPN to be a dissemination network + producers + consumers. I actually could not find any compelling reason to introduce the notion of dissemination network. According to the definition we are using, event processing network is a directed graph that has nodes for producers, channels, EPAs and consumers, and edges that determine the event flow among them. Another difference is that the Gartner report views event consumers and event producers as type of event processing agents. I have a slightly different opinions, I think that both event producers and consumers are not really event processing agents, since event processing agent is some software module that function events and may generate more events. Event consumer and producer have nodes representing them in the EPN in order to make the event flow from and to them explicitly, however, they are only proxies of the actual producer and consumer, for the event processing network, they are sources and sinks. The main difference is that EPA functionality is explicitly specified in the EPN definition, while what the producer and consumer do is "black box". We don't want to include their functionality, since we don't want to extend the event processing language ad infinitum,

Mentioning the EPIA book -- Chapter 3 is now on the Web, and can be obtained through the MEAP program, this is the last chapter in the introductory part, and deals with principles of programming with events. Chapter 4, the first in the deep dive will be sent to the publisher soon. It has been much more challenging to write, deals about what information we need to store about events -- I'll Blog about it soon.

Saturday, May 9, 2009

Beyond the horizon

Well - I did not come to Las Vegas just to enjoy the time, I have given two (identical) talks together with my colleague Kyle Brown, an IBM Distinguished Engineer, who belongs to the Websphere Services organizations. I talked about technology, Kyle introduced related application, fun staff. I have posted the presentation on slideshare. Enjoy.

Friday, May 8, 2009

On Smarter Planet


Still in Las Vegas. The icon above is the logo of IBM smarter planet, which is the main theme of the IMPACT conference. According to IBM, the smarter planet is instrumented, interconnected and intelligent. Event Processing is considered in all talks a key enabler. In the conference there has been a distribution of a mini-book, which includes the first few chapters from the upcoming book of Mani Chandy and Roy Schulte. Sandy Carter, IBM VP for Websphere marketing, has written a preface to this mini book, here is a quote: We at IBM are excited about event processing. As the planet becomes increasingly instrumented, interconnected , and intelligent, event processing is emerging as a key technology for mitigating risk, seizing opportunities, and achieving greater corporate agility. Smart cities, smart healthcare, some retail... these are just a few areas where event processing is adding more intelligent technology to our life. Under the smarter planet imitative there are vertical solutions, and today I went to listen to a talk by Honda Italy about smart manufacturing. Interesting stuff, more - later

Tuesday, May 5, 2009

On IMPACT 2009 - the first day


Las Vegas. I have arrived here yesterday to participate and give a talk (twice) in the IMPACT conference, IMPACT is the annual conference that IBM is holding for its customers, especially the Websphere brand, but also covers anything that is defined under the "SOA portfolio". James Taylor is reporting about it, and is doing much better work than I can do, so trace his Blog for more details, however I wanted to bring my personal angle into it. The first day started in a plenary session with a huge, besides a very funny performance of Billy Crystal, who put three IBM executives on the podium to do sound effects to a story he told, there have been presentations of the senior executives of IBM Software Group, Steve Mills and Tom Rosamillia on the main theme of this year's conference: Smarter planet. When I heard the presentation I could not help smiling to myself, practically all customer examples that have been described under "smarter planet" have included event processing component in a key role. One of the customer examples has been the one described by my colleague Guy Sharon in the 4th EPTS event processing symposium, about tracing people in industrial facility to verify authorization of people to be in a certain place accompanied or unaccompanied, since the customer was exposed in public, I can mention it has been British Petroleum, this is something that the IBM Haifa event processing team has been heavily involved in creating this application. Some more applications in different industries have also been presented. Event Processing is being used far beyond its early adopter, the financial markets industry. I view it both inside and outside IBM that people who have been somewhat skeptic about event processing 2-3 years ago, talk differently about it now. Actually, the interest will yield more investments and advance this area to cope with the various challenges, but this is another story.

The main business value that event processing is achieving, according to IBM studies, is business agility. A study done reveals that in today's environment, almost all CEO are trying to re-think and change their business processes, and often software is an obstacle. Event processing is considered as an enabler for such agility. I always believed (and wrote about it several times) that the main business value is agility and not high performance that is required by a certain segment of applications, while agility is required by many more... Reality has its own strange ways...

A comment about branding -- IBM decided more than a year ago to use the brand BEP (Business Event Processing) for the entire portfolio of capabilities in the various IBM products related to events, and not adopt theCEP acronym used widely in the industry . The idea has been to emphasize the "business user orientation". Anyway, I'll leave the branding discussion to the marketing guys -- the content is more important.


There are more than 20 sessions in the conference that talk about event processing from different angles -- from the business perspective, the technical perspective, various related products in IBM (e.g. enabling CICS applications to serve as producer and consumer of event processing), my own session later this week deals with the future, but I'll Blog about it at later time, it is late, and I need some rest, getting up early every day for some conference calls...

Saturday, May 2, 2009


Packing for a one week (net) travel to the USA. We had recently been informed that the paper entitled: A stratified approach for supporting high throughput event processing application
has been accepted for presentation in the DEBS 2009 conference that will take place in early July in Nashville. The paper written by Yuri Rabinovich, Geetika Lakshmanan and myself describes results obtained last year in our scalability project, the project is still going on, and its results will be flowing to IBM products.

Here is the abstract of this paper:

The quantity of events that a single application needs to process is constantly increasing, RFID related events have been doubled within the past year and reached 4 trillion events per day,
financial applications in large banks are processing 400 million events per day, and Massively Multiplayer Online (MMO) games are monitoring in peaks 1 million events per second. It is evident that scalability in event throughput is a major requirement for these types of pplications. While the first generation of event processing systems has been centralized, we see various solutions that attempt to use both scale-up and scale-out techniques. Alas, partitioning of the processing manually is difficult due to the semantic dependencies among various event rocessing agents. It is also difficult to tune up the partition dynamically in a manual way. Manual partitioning is typically vertical, i.e. there is a single partition set with centralized routing. This paper proposes a horizontal partition that is automatically created by analyzing the semantic dependencies among agents using a stratification principle. Each stratum contains a collection of independent agents, and events are always routed to subsequent strata. We also implement a profiling-based technique for assigning agents to nodes in each stratum with the goal of aximizing throughput. A complementary step is to distribute the load among the different execution nodes dynamically based on performance characteristics of nodes and agents and the event traffic model. Experimental results show significant improvement in the ability to process high hroughput of events relative to both centralized solutions as well as vertical partitions. We find this to be a promising approach to achieve high scalability without requiring difficult manual tuning, especially when the traffic model and the topology of the event processing network is often changed.


More about event processing distribution and parallelization will be discussed in subsequent postings.

DEBS has also issued recently a call for fast abstracts, posters and demos, an opportunity to share with the community work that is in less mature phase. show interesting demos, and discuss ideas.

Wednesday, April 29, 2009

On events and relativism


Today has been a holiday, the Independence day of Israel, and we spent some of the day in going to an exhibition called "Body World", in which there is an exhibition about the human body and its various functions using parts taken from dead people who contributed their body using some preservation method developed by somebody in the university of Heidelberg, actually I also spent some of the holiday working, since there was something that has become artificially urgent. These high-tech corporates makes you a slave, I am getting too old for this...

Anyway, in Israel the day before the "Independence Day" is the "Memorial Day" to remind us that the independence has its cost. It reminds me that in my first year in the USA (I lived in the USA for several years around 20 years ago), there were signs in the street saying "memorial day sale", we thought that somebody is making a joke in a bad taste, which does not really fit the famous "politically correctness" of the Americans, but found out after talking with some local people that the typical American does not attribute any semantics to the "memorial day", and it is just a long weekend with sales and travels like any other long weekend, well -- a cultural difference, since in Israel, memorial day is taken seriously.

Today, I wanted to say something about events and relativism. One of the questions about chapter 1 in the "Event Processing in Action" book on the forum came from Richard Veryard.
The question has been:
How do you count how many events? If you have a three-car pile-up, does that count as one collision or two, given that the third car hits a few seconds after the first two? Or three collisions, if the third car hits both of the first two cars?

My answer has been that the decision is relative for the application. From the insurance company or companies of the cars involved it may look as three different events, since the event refers to a single car; from the point of view of the traffic police it may be considered as a single event, where the number of cars involved is an attribute.

Another facet of relativism is whether an event is raw or derived. The event can be raw event from the point of view of a certain application, since it is provided from the outside, however, the consumer is sending an event that has been produced by another event processing application, and from the point of view of the producing application, this is a derived event. There are probably more example of relativism.

Monday, April 27, 2009

More on Revision


Long day today, I got to the office around 8AM and left around 9PM. Since we have holiday this week I am trying to condense the remaining days of the week and the result is long day with plenty of conference calls. The picture above is a glance (from below) on the IBM Haifa Lab (the pair of connected building on the right hand side of the picture), my office is in the back building (known as the "banana" due to its shape), and is not really in the nice part of the building -- the one with the view to the Haifa Bay -- well, one can have everything in life -).

I still need to complete the previous posting on revision. I gave some explanation about the concept of revision, and now I still need to discuss implementation of revision in event processing. To recall -- a revision in event processing is getting later knowledge that asserts that either a reported event did not really happen, or some information associated with the event was wrong.

Let's look at two separate cases, one in which the processing has not gone out of the event processing network, and second that the results of the processing have gone out to the "outside world".

In the first case, there may be an opportunity to revise the impact of the revised event by doing kind of undo-redo for all the event processing agents that it passes directly or indirectly. Direct ones are easy -- those that the revised event participate as an input in them, indirect is more tricky, since we need to trace the causality among events, in this case, an event that is an output of an event processing agent in which the revised event participate (relate to the same context) has a causality relation to the original event, thus, an event processing agent, in which this event participates as an input, also needs to do an undo/redo, and causality is a transitive relation, so it continues as far as the EPN arrived so far. It should be noted that the fact that there is a causality may not require a real undo/redo, take as an example that an event of type E1 designates a bid, and the event of type E2 designates the bid with maximal value arrived in a certain time interval. Let's assume that a certain bid has been revised, however, neither the revised bid, or the revising bid change the selection of E2.

The second case is that the revised event has consequences that have been sent to an external consumer, thus, it may have triggered an action, a collection of actions, or a workflow that has been carried out, and this may propagate further ("the butterfly effect"), in this case, either we can treat it as "too late" and do nothing, however, there may be a cases that it can be critical to undo/redo also the consequences, e.g. the revised event has some financial meaning. In this case we'll need to issue compensation for the triggered action, which may be impossible (the consumer does not support compensation) or difficult. I'll blog again about revising the history and its aspects at a later phase.