Friday, February 17, 2012

Timo Elliot's presentation on "Business in the moment - from reactive to proactive"





Timo Elliot from SAP gave a recent talk in the Gartner BI meeting in London entitled "Business in the moment -from reactive to proactive". You can download the presentation from a link in Timo's Blog posting.   In a following post on his Blog,  Timo refers to an FreshDirect explaining the proactive behavior:



“FreshDirect has an operations center that manages its fleet of delivery trucks. In a large metropolitan area like New York, traffic doesn’t always flow predictably. A traditional approach to BI would be to print a report showing the level of on-time deliveries (OTDs) the day before and then ask the transportation department what went wrong for the orders that were delivered late. FreshDirect uses analytics in a more impactful way.”
“The company monitors the delivery rate of every truck and enters that data into the BI system on an ongoing basis. Every hour, it uses the previous hour’s data to predict how many deliveries will be on-time in the next hour. If the predicted OTD rate is below FreshDirect’s target, the company sends out an auxiliary truck or trucks to help make deliveries. The company holds 10 trucks in reserve for just this purpose.”
I'll bring more proactive stories when I'll find out about them...

Tuesday, February 14, 2012

Killing elephants - is MapReduce dying?


My English teacher in the last grade of high school had an interesting taste in literature, and taught us the story on  "Shooting and Elephant" by Orwell.   I was not a very good student in English and forgot about it until reading Colin Clark's Blog posting entitled : "It's time to kill the elephant".   From time to time there are various people claiming that various things are dead or dying.   Some of the readers may still remember the discussion about whether SOA is dead.   Recently the Forbes Blog has announced the death of ERP.  Colin's contribution to the hunt is the observation that MapReduce is dying (or should be dying) and the batch processing should be replace by more real-time processing.  His evidence is that Google is dumping MapReduce and using Colossus for its search technology.   While this fact is certainly true, I think that there are still many types of analytic procedures that are done off-line using batch processes, so while the use of real-time analytics will substantially increase given supporting infrastructure, I am not sure that batch will die soon (the same goes for SOA and ERP)...   Old soldiers never die - they just fade away (s-l-o-w-l-y).

Sunday, February 12, 2012

Crash course to build simple EP application using Esper


A crash course claimed to take less than an hour entitled "A simple introduction to complex event processing" has been posted.  This is done by example, which seems to be indeed very simple, finding "decreasing" or "increasing" pattern over two consecutive events and setting the color as green or red.   The main emphasis is on the setting - how to obtain, define and use events, and configure the engine - threadpools, listeners etc...  However not much about what event processing actually can do -- this is probably the next lesson.


  Esper is contrasted with commercial products since its open source model allows developers to play with it, use it for toy examples, and for daily usage that is not necessarily a commercial application of big enterprise, in our days of enterprise computing this approach has certainly a role to play, it should be noted that Esper is not the only open source in this area, and that some of the commercial products allow free development version (not access to the source code, but enabling developers to use the product for these purposes for free).


Anyway -- if you wish to learn Esper, it is a good start. 

On revision and compensation in event processing

The event processing course that I taught in the Technion has ended (except for the make-up exam in March). This year the students' project has been different from the last couple of years (well, I am bored having the same type of project all over again).  The students were asked to select a research project.  My Teaching Assistant was very skeptic about the students' abilities to do such projects, but after the presentations they did in the last class summarizing what they did so far she changed her mind.    I'll write about these projects after they'll be submitted (due in early March).   One project is investigating the issue of compensation and revision in event processing,  I have written about revisions before,  both in general,  and specifically related to event processing.  
There are couple of motivations of why I have returned to be interested in it now. 


One of them is the investigation of uncertain events,  since more information may be acquired with time about events that are uncertain,  a revision might be needed.
The second is the work on future events,  since future events are obtained using a forecasting process, this forecasting process is not a one-time process, but can be sensitive to additional events that happen between the forecast and the occurrence of the future event, thus the forecast itself may be revised.  
The issue of revision may entail the need for compensation for decisions and actions already taken.

  • In some cases it is easy, example when no action was taken and it is still possible to take action
  • In some cases it is impossible, if an action has been taken, and this action cannot be retracted, 
  • In the remaining of the cases, it might be possible, however not always cost-effective, since it might have cascading effect of compensating for large amount of actions. 


After getting the students' work, I'll write again about this issue. 


Saturday, February 11, 2012

Uncertainty in event processing

This cartoon is taken from Cartoonsbyjosh.com indicates uncertainty about uncertainty. 
And indeed, there has been a lot of work about uncertainty in data over the years in the research community, but very little got into the products, the conception has been that while data may be noisy, there is a cleansing process that is applied before using the data.    Now with the "big data" trend, this assumption seems not to hold at all times,  the nature of data (streaming data that need to be processed online), the volume of the data, and the velocity of having also imply that the data, in many cases, cannot be cleansed before processing, and that decisions may be based on noisy, sometimes incomplete or uncertain data. Veracity (data in doubt) was thus added as one of the four Vs of big data. 
Uncertainty in event is not really different from uncertainty in data (that may represent either fact or event).
Some of the uncertainty types are:

  • Uncertainty whether the event occurred (or forecast to occur)
  • Uncertainty about when event occurred (or forecast to occur)
  • Uncertainty about where the event occurred (or forecast to occur) 
  • Uncertainty about the content of an event (attributes' value)


There are more uncertainties relate to the processing of events

  • Aggregation of uncertain events (where some of them might be missing)
  • Uncertainty whether a derived even matches the situation it needs to detect -- this is a crucial point, since the pattern indicates some situation that we wish to detect, but sometimes the situation is not well-defined by a single pattern.  Example:  a threshold oriented pattern such as:  "event E occurs at least 4 times during one hour".   There are false positives and false negatives.  Also if event E occurs 3 times during an hour,  it does not necessarily indicate that the situation did not happen.


We are planning to submit a tutorial proposal for DEBS'12 to discuss uncertainty in events, and now working on it.   I'll write more on that during the next few months

Wednesday, February 1, 2012

On "CEP and Big Data 2" - comments on Philip Howard's observations.

Philip Howard from Bloor Research has posted some observations on his Blog entitled "CEP and Big Data 2".   Here are some comments (actually nothing new - just summarizing things I have written about before).
Philip deals with three issues:

  • whether the name CEP is appropriate or should be changed? 
  • who should be credited as the pioneer of this area?   
  • whether CEP implies real-time processing?  
  •  who are the CEP big data platforms?

Here are summary of my views on each of this topics.

The name "Complex Event Processing"

Exactly four years ago I posted on this Blog an explanation about - "why I prefer to use the name event processing without any prefix, infix or suffix".   My particular dislike of the term "complex event processing" stems from the ambiguity in the name - some people (including David Luckham who coined this term) view it as processing of complex events, some interpret it as complex processing of events, and then debate of when something is complex enough, and what type of complexity is needed  to qualify as CEP.  Moreover some of the vendors use this term for products that are neither of the two options.   I think that two words is enough for the name of a discipline, examples: information retrieval, machine learning, image processing and much more....  Thus, from my point of view the term "event processing" subsumes all other terms like complex event processing, business event processing, event stream processing and more.

Who gets the pioneering credit

Philip as a good UK patriot wonders why the Wikipedia value about Wikipedia and other sources gives credit to David Luckham and forget the Apama work that came from Cambridge UK.    Looking at Wikipedia, it has one mention of David, as well as other references (like our EPIA book). It indeed does not mention Apama or any paper by John Bates, but being a Wikipedia, anybody can suggest additions.   
David Luckham had major influence on this area, since he was the first one who published a full book and exposed the young area to the general public.    An article in IEEE Computer, published in 2009,  made some investigation of the history of that area and determined that in the 1990-ies there were four parallel projects that can be classified as starting points in this area:  David Luckham's project in Stanford,  John Bates' project in Cambridge (UK, not Boston), Mani Chandy in Cal Tech,  and our Amit project in IBM Haifa Research Lab.    I share Philip's view that John Bates should have full credit as one of the pioneers, and still view David Luckham as the "elder statesman" of the community.

Is CEP necessarily associated with real-time?

I have written several times about this topic, last time in response to Chris Carlson, to whom Philip also responds.   There is some abuse of the term real-time in the industry, while its meaning is "within time constraints", many people interpret it as "with very low latency".   This is not the same,  anyway, event processing is a functionality with applications that require very low latency, applications which require to react within real-time constraints (which can be: 2 hours), some require both, and some require none.

Who are the CEP big data platforms?

I have taken upon myself the limitation not to state opinions on commercial products within this Blog  - leaving  it to analysts.   Thus will make one comment.  There is distinction between two types of software entities - 
which is sometimes confused in the language used by people.

  • Event Processing Platform is a software that enables the creation of event processing network, handle the routing of events among agents, management, and other common infrastructure issues.
  • Event Processing Engine is a software that enables the creation of the actual function - in the EPN term implementing agents.
This is similar to the difference between an application server and a single component (programming in the small vs. programming in the large).    Some of the available platforms for "event processing for big data" provide the first one -- it gives infrastructure, but not implementing any type of functionality, but enabling developers to create their own functionality, thus they don't do full-fledged event processing.   Seems that many people classify both under the same classification  (of course there are products that do both). 

Tuesday, January 31, 2012

On spime


According to Wikipedia: Spime is a neologism for a currently theoretical object that can be tracked through space and time throughout the lifetime of the object. The name “spime” for this concept was coined by author Bruce Sterling. 


Spime comes from the combination of the words space and time,  and is said to be enabled by the Internet of Things.  In the event processing terminology - spime is the collection of events that happened to a single entity during its life-span,  where each event has both time and space properties recorded as part of this event.   Any person may have a spime associated with this person, which can span from birth and actually last long time after the person's death, e.g. if I am writing now about Isaac Asimov, this can be considered an event in Asimov's spime, although he is not a living entity.  Spimes can relate to something with more limited length like a certain flight,  or the event processing course I taught this semester.


In some cases it make more sense to have Spime processing rather than individual event processing and have some patterns associated with Spimes, this, of course, has strong relationship to event processing -- I've recently started to look and spime processing and will write more about it in the future