Wednesday, February 1, 2012

On "CEP and Big Data 2" - comments on Philip Howard's observations.

Philip Howard from Bloor Research has posted some observations on his Blog entitled "CEP and Big Data 2".   Here are some comments (actually nothing new - just summarizing things I have written about before).
Philip deals with three issues:

  • whether the name CEP is appropriate or should be changed? 
  • who should be credited as the pioneer of this area?   
  • whether CEP implies real-time processing?  
  •  who are the CEP big data platforms?

Here are summary of my views on each of this topics.

The name "Complex Event Processing"

Exactly four years ago I posted on this Blog an explanation about - "why I prefer to use the name event processing without any prefix, infix or suffix".   My particular dislike of the term "complex event processing" stems from the ambiguity in the name - some people (including David Luckham who coined this term) view it as processing of complex events, some interpret it as complex processing of events, and then debate of when something is complex enough, and what type of complexity is needed  to qualify as CEP.  Moreover some of the vendors use this term for products that are neither of the two options.   I think that two words is enough for the name of a discipline, examples: information retrieval, machine learning, image processing and much more....  Thus, from my point of view the term "event processing" subsumes all other terms like complex event processing, business event processing, event stream processing and more.

Who gets the pioneering credit

Philip as a good UK patriot wonders why the Wikipedia value about Wikipedia and other sources gives credit to David Luckham and forget the Apama work that came from Cambridge UK.    Looking at Wikipedia, it has one mention of David, as well as other references (like our EPIA book). It indeed does not mention Apama or any paper by John Bates, but being a Wikipedia, anybody can suggest additions.   
David Luckham had major influence on this area, since he was the first one who published a full book and exposed the young area to the general public.    An article in IEEE Computer, published in 2009,  made some investigation of the history of that area and determined that in the 1990-ies there were four parallel projects that can be classified as starting points in this area:  David Luckham's project in Stanford,  John Bates' project in Cambridge (UK, not Boston), Mani Chandy in Cal Tech,  and our Amit project in IBM Haifa Research Lab.    I share Philip's view that John Bates should have full credit as one of the pioneers, and still view David Luckham as the "elder statesman" of the community.

Is CEP necessarily associated with real-time?

I have written several times about this topic, last time in response to Chris Carlson, to whom Philip also responds.   There is some abuse of the term real-time in the industry, while its meaning is "within time constraints", many people interpret it as "with very low latency".   This is not the same,  anyway, event processing is a functionality with applications that require very low latency, applications which require to react within real-time constraints (which can be: 2 hours), some require both, and some require none.

Who are the CEP big data platforms?

I have taken upon myself the limitation not to state opinions on commercial products within this Blog  - leaving  it to analysts.   Thus will make one comment.  There is distinction between two types of software entities
which is sometimes confused in the language used by people.

  • Event Processing Platform is a software that enables the creation of event processing network, handle the routing of events among agents, management, and other common infrastructure issues.
  • Event Processing Engine is a software that enables the creation of the actual function - in the EPN term implementing agents.
This is similar to the difference between an application server and a single component (programming in the small vs. programming in the large).    Some of the available platforms for "event processing for big data" provide the first one -- it gives infrastructure, but not implementing any type of functionality, but enabling developers to create their own functionality, thus they don't do full-fledged event processing.   Seems that many people classify both under the same classification  (of course there are products that do both). 

Tuesday, January 31, 2012

On spime


According to Wikipedia: Spime is a neologism for a currently theoretical object that can be tracked through space and time throughout the lifetime of the object. The name “spime” for this concept was coined by author Bruce Sterling


Spime comes from the combination of the words space and time,  and is said to be enabled by the Internet of Things.  In the event processing terminology - spime is the collection of events that happened to a single entity during its life-span,  where each event has both time and space properties recorded as part of this event.   Any person may have a spime associated with this person, which can span from birth and actually last long time after the person's death, e.g. if I am writing now about Isaac Asimov, this can be considered an event in Asimov's spime, although he is not a living entity.  Spimes can relate to something with more limited length like a certain flight,  or the event processing course I taught this semester.


In some cases it make more sense to have Spime processing rather than individual event processing and have some patterns associated with Spimes, this, of course, has strong relationship to event processing -- I've recently started to look and spime processing and will write more about it in the future

Monday, January 30, 2012

On Pecha Kucha

Back to presentation skills,   today, while working with one of my colleagues, Avi Yaeli, on a presentation, I've learned a new concept - Pecha Kucha.  This is a presentation pattern, in which the presenter presents a topic in 20 slides, and spends on each slide 20 seconds,  total of 6 minutes and 40 seconds per presentation. 
There is a youtube presentation containing Pecha Kucha style presentation about how to prepare Pecha Kucha style presentations.   I should try it once.  There are also Pech Kucha nights which seems to be marathon of Pecha Kucha presentations.   

Saturday, January 28, 2012

Is computer science a science or engineering?


I remember years ago a heated discussion in a conference whether computer science is a science or engineering, my daughter had a "science day" in the high school that she'll attend next year, and while they teach computer science they don't view it as a science, for them science consists of biology, chemistry, physics and some of their derivatives.   


Recently I came across an  article in "Scientific American",   about U.S. science degrees.   In this article, as you can see in the picture below,  computer science is neither classified as science nor as engineering,  it is actually classified as technology.   Interesting -- I think that computer science is not monolithic, and various sub-disciplines may be classified differently.





Sunday, January 22, 2012

On presentation skills



Somebody attracted my attention today that ACM  Membernet Europe in its last issue,  has written about me in the section "feature ACM European distinguished speaker".     In fact, several months ago somebody from ACM approached me to ask what is the meaning for me of being recognized as ACM Distinguished Speaker, 
The truth is that I intended to use this program to tour some exotic places in the universe, but did not have time yet to pursue it,  thus I answered that my action after this recognition is to coach and mentor young people about presentation skills.  Indeed I have added to courses and seminars I am teaching a pitch about presentations (I am a fan of Steve Jobs' style of presentation), while this is a "soft skill", it is very important in today's world, as the picture above shows - sometimes more than what you say.   In Israel we have a tendency to underestimate it, and believe that good content will sell itself,  this is also true on product packaging.   While some people are naturally good presenters, presentation skills is something that can be learned, and it is very rewarding to see young people catching quickly and producing great presentations (last week a students in a seminar I supervise did very creative presentations).  

Tuesday, January 17, 2012

Intelligent Business Operations - a medical use case


Within the recent year Gartner promotes the term "Intelligent Business Operations" (IBO)  - not to confuse with Business Intelligence (BI).  Roy Schulte from Gartner wrote about "operational IQ".   I am looking now at the concepts and facilities of IBO, in Gartner's view.    One way to study it is by looking on a recent post by Jim Sinur (also from Gartner).  Jim provides a success story in the medical domain, resource allocation in surgeries.   The ingredients of this scenario are:



  1.  Simulation-based optimization of scheduling and resource allocation in off-line for all surgeries planned for the next day.
  2. Real-time tracking of everything: physicians, nurses, equipment; monitor of procedure duration and status - using sensors, cameras and in Jim's terminology - exploiting the "Internet of Things".
  3. Determination of things already going wrong (not according to plan) or expected deviations from plan
  4. Re-applying the simulation based optimization (this time online!) and get updated resource allocation plan.


This may be instance of the "detect-forecast-decide-act" pattern we have identified as the basis of proactive computing, although in Jim's scenario it can also be reactive (the deviation from plan already occurred - there is no need to forecast anything).     


I'll write more about the IBO concept and some additional ingredients of it soon.  
Since the term "intelligence" is now back in fashion,  it would be nice to have metrics for the IQ of some operational process like the surgery management.
2. 

Monday, January 16, 2012

FFD - the distributed version


FFD (Fast Flower Delivery) is the example that accompanies the EPIA book.  Recently Phil Windley,  the CTO of Kyntex, started to teach a course in Brigham Young University, entitled "large scale Internet applications",  Phil is using the EPIA book as one of the textbooks for his course.   Today Phil posted in his Blog his variation of FFD as a totally distributed system, the illustration above demonstrates it.
This is interesting, indeed we see that event processing systems that started as centralized applications are getting more and more distributed.  In Phil's version, there is no single EPN, but a federation of EPNs, per individual player (driver, store etc..),  with pub/sub relations among them. 


BTW - an interesting phenomenon,  last week I have not written any posting in this Blog, have been busy in the last sprint of the EU proposal we submit (deadline tomorrow, so there is a light at the end of the tunnel), however, looking at the statistics, I think it was the week with the busiest traffic on my Blog ever both in page views and number of visitors, including a day which I think was the record high for this Blog for a single day ever,  very surprising, I actually cannot explain it -- maybe some people got back from vacation and are catching up?. I'll publish some statistics in the near future.