Saturday, September 20, 2008

On EPTS and coopetition



Hello from Bedford, MA, to where I have driven earlier today from Stamford (3.5 hours with one stop in the way).
Louis Lovas from Apama is a newcomer in the EPTS meetings (though other Progress/Apama persons participated in all the meetings). He summarizes his impressions in an amusing blog posting called "a truce in the CEP Front", where he finds out surprisingly that his competitors are persons and one can sit and drink wine with them (well - Louis drank beer, if my memory does not mislead me). I think that most of us have gone through this process earlier, when I have proposed that IBM will host the first symposium, I came to marketing and explained them the idea, they did not quite get it; The dialogue on the phone was something like:
- I give a 2 minutes speech explaining the idea of community building event
- The marketing guy says: you mean conference for customers, sure we can do it.
- I answer: there may be also customers, but also others - everybody that belongs to the community.
- The marketing guy says: Ah, now I understand you, you mean conference for IBM business partners.
- I answer: we can also have business partners, but we should also have companies like TIBCO, Oracle, Progress and some others...
- The marketing guy: I don't get it, you want to invite competitors? you must be kidding...
Well - I have been just a guy from a small province (IBM Haifa) and could not convince the almighty marketing to agree to fund this symposium, thus I had to wait a little bit more and find another sponsor (the IBM Academy of Technology eventually funded it)... in the first meeting, most people did not know each other, and there were low expectations about any concrete results.
However, now in the fourth meeting most people in the meeting know each other, and from the amount of work-groups that people suggested, it seems that there is now more confidence in the usefulness of EPTS as a community, and it seems that there will be step-function in the activities.
The co-opetition idea is interesting, the picture above illustrates, of course, co-opetition in the open source (the famous penguin), there are many co-opetitions in life, for example in Israel there is a central clearing-house to transfer money between the various competing banks, there are many other examples.
In the "Event Processing" area, we have barely scratched the surface of its potential, and if we'll succeed to help better understand the foundations, establish some standards, advance the awareness to this area, and better understand various ways to gain business value -- we'll all earn from that. I think that the activities done so far has already influence on individual vendors and customers.
Of course, in the daily life, the different vendors will continue to fight on deals and customers as usual - that is the beauty of co-opetition.
BTW - Louis, congratulations for your new role as Apama's representative in the EPTS steering committee, hope that you'll be as contributor to the cooperation side of the co-opetition as your two predecessors - Mark Palmer and John Bates.

On the EPTS 4th event processing symposium

Still in Connecticut, whose flag you can see above. The three days of the 4th event processing symposium are over - what a relief, took a lot of effort to organize it - but it seems that people were happy with the content; I'll let other blogs review the meeting, some already done it - Brenda Michelson had several postings (this one, for example), Paul Vincent also wrote some of his impressions. Needless to say that I am exhausted from these days, but I still need to summarize the (relatively big!) amount of action items and send the EPTS members, may be tomorrow morning. Anyway - some short impressions:




  • It seems that people feel more mature now to discuss issues like: benchmarks and languages, that they did not really want to discuss before, we have action items on both areas.


  • Brenda Michelson has given an interesting presentation in the business oriented panel, showing that EP is actually under-hyped relatively to some other hypes, the panel moderator, Alan Lundberg, asked the audience if anybody think that EP (or CEP) is over-hyped, and no one raised hand. It should be noted that audience consisted not only of vendors but from analysts, consultants, academic people and customers.


  • The use cases workgroup demonstrated some of the varaince we have under event processing (not sure that everything is CEP according to the glossary definition, but EPTS is about EP, not just CEP). The six has different use cases were: Alex Kozlenkov from Betfair talked about fraud detection in the betting business, although the audience did not agree that all his examples represented fraud; Arkady Godin from MITRE talked about large dissemination information (sophisticated pub/sub in the area of weather), Brian Connell from WestGlobal talked about performance monitoring of mobile operators (a monitoring system), Dieter Gawlick from Oracle talked on "first responders" in case of emergency - again a dissemination system, Guy Sharon from IBM talked about facility safety - a kind of monitoring system, and Richard Tibbetts from Streambase talked about alternative trading venues (operational system).


  • Susan Urban presented a list of research challenges, and there was report about various research projects, including two presentations from IBM Research which has very active research in this area; Fred Douglis talked about "system S", while my colleague from IBM Haifa Research Lab, Guy Sharon, talked about the history of AMiT, and about the current research projects we are conducting now, the one which attracted attention is the EPDL (Event Processing Description Language) project, and it seems that we may accelerate its exposure to the larger community.


  • Chris Ferris, the IBM CTO of Industry standards, explained the value of standards in general started with the chinese king who unified China, following by a standard panel - it seems that the community is now more mature to start talking about standards than in the previous years - I wonder what is the reason.


  • Today in the business meeting - quite a lot of activities and working groups have been proposed. We'll see how much of it will materialize, people are more enthusiastic to commit on investing time when they are out of the office, until we return to the grey reality... but it seems that the EPTS activity is accelerating.


Some Trivia:

  • We had 70 participants (the upper limit we set) - 41 vendors, and the rest - academic people, customers, analysts and consultants.



  • We had participants from 11 countries - see chart below (since the conference has taken place in the USA, we miss much of the European academic community).








      • More - Later.





      Wednesday, September 17, 2008

      On the Gartner EPS 2008


      Early morning in Stamford. The "hype cycle" has become Gartner's most known artifact, as well as their constant flow of TLAs - SOA, BAM, EDA, RTE, XTP - are all Gartners' words. The Gartner EPS has ended last night, and in less than two hours we'll start the EPTS 4th event processing symposium (so Roy Schulte will be able to relax and I'll start sweating).
      Some short impressions from the conference (besides the networking and meeting again old friends).
      • The Gartner analysts came with new slides, but very little new insights relative to their past messages.
      • Mani Chandy had interesting talk, saying that EDA is a natural continuation of SOA and not a paradigm shift (good topic to discuss). Also said that one of the big benefits of EP is - saving time for people by filtering and aggregation of flowing information.
      • David Luckham has talked about "holistic event processing" as the future -- which from technology point of view means - dynamic big event processing networks. Also talked about the need to have formally defined clear semantics of event processing language (well - that is what we are trying to do in EPDL).
      • Marc Adler has a great talk (in my opinion, the best one) - about his experience in developing CEP application, which seems to be a success story. He talked about criteria to select a vednor, difficulties in the development itself, and shortages in the state-of-the-art. One of his insights is that unlike what the "stream SQL" fans are saying that since people know SQL anyway, it is a good basis, he claims that while the syntax looks like SQL, it is a totally different type of thinking, and the knowledge of SQL does not help on getting it.
      • Richard Brown had a talk about collecting events from text - news, blogs etc... I think that getting events from unstructured data has a lot of potential, I also talked with somebody who told me about start-up that extracts events from video cameras - the scenario of tracing behavior of people that fit the shop-lifting pattern (was discussed recently in Blogs) - may become reality soon, it seems that the technology is getting there.

      That is all my time permits me to write today -- more later.

      Sunday, September 14, 2008

      On sporadic events


      I have never been a student in Stamford high school, but Stamford, CT is my home away from home for the next seven days. Starting tomorrow, I'll provide some impressions from the Gartner meeting and EPTS symposium, but I rely on other people in the blog-land to have a better coverage (e.g. Paul Vincent, with endnotes and references). I've Arrived earlier today, and resting before the busy week.


      One of the thoughts that came to mind when looking at some of the discussions around the Stream-SQL standards, is one more observation.


      While the claim that the difference between a "stream" and a "cloud" is that a stream is totally ordered and a cloud is partially ordered, I think that there are also some more distinctions. I'll discuss one of them --- sporadic events vs. known events. When dealing with "time series" type of input, then the timing of events are known, for each time unit (whatever it is) there is an event (or set of events) that are reported. This is true when the events are stock quotes provided periodically, or signals from sensors provide periodically. There are events that are not naturally organize themselves in time-series fashion, for example: bids for an auction, complaints from customers, irregular deposit, coffee machine failure etc...

      From the point of view of functionality, there is no much difference --- one can create time series which reports a possibly empty set of events for each time-unit, but if in most time-unit the reported set is empty, this will not be very efficient way to handle it. On the other hand a system that does not support time series as a primitive can view all events as sporadic events, but there may be some optimization merit to the knowledge that events are indeed expected at every time unit. This is just one dimension -- but this leads me to reinforce my conclusion that there are various ways to provide event processing functionality, and the most efficient way is probably hybrid of approaches based on the semantics of functions and characteristics of the input. So this is the observation of today; it will be interesting to see what will be the main discussion topics in the coming conferences.
      More later - from Stamford Hilton.




      Thursday, September 11, 2008

      On Occurrence time: a footnote to the UAL fiasco


      As a past Dungeon Master the word crawler always reminds me about the "carrion crawler", a monster you can see in the picture above, but recently a combination of the allmighty Google crawler, and automatic trading programs based on event processing has caused a fiasco that crashed the stock of United Airlines, some of the blogs have referred to it: Brenda Michelson in her Blog have talked about the butterfly that lead to the computer glitch. Mark Palmer thinks that news should be regulated (some people I know who were borne in countries were news are indeed regulated shiver to hear the idea that news - of any type - should be regulated).
      I will not go back to the story, but as a footnote - two issues come to mind - event validation and the issue of occurraece time. So I'll write today about occurance time since it is easier...
      The works in the temporal area are talking about several time dimensions - the bi-temporal model talks about: transaction time -- the time that a fact is recorded, and valid time -- the time interval in which the fact is valid. In event processing we also look at a bi-temporal time similar to this: detection time -- the time that the message that represents the event was detected by the processing system, and occurence time -- the time which the event happened in reality (occurrence time can be considered as the starting point of a valid time that ends when the event becomes irrelevant, but let's get it out of the scope and concentrate in occurrence time).
      Some of the implementation of event processing base the order of event on the detection time, some support occurance time, and some base the built-in temporal capabilities based on detection time, and enable defining times as an an attribute, but then the temporal operators have to be hand-coded as regular predicate.
      One of the common fallacies is that detection time is good enough as a metrics for temporal operations on event (e.g. trends), first - event from the past can suddenly pop up out of the blue (I know a person who has an habit to catch-up in Email every two weeks or so, and answer to the Email before realizing that there has been a whole thread of Emails that make answering the original Email quite obsolete), second - the order may not be kept even if the delay from the occurance time to the detection time is very small. The order of medical exams may not be consistent with the order of results reaching, and knowing the real order may be important for the differential diagnosis.
      Thinking about standard structures for events -- I would think that having "standard header" with some mandatory properties for each event - is a good candidate for having standard (I am less optimistic about standards for the content of the event), and in the header - the occurrence
      time should be a mandatory.
      Occurrence time has some inherent issues associated with it - but I'll discuss it another time.

      Wednesday, September 10, 2008

      On events about events


      Today, our entire family has travelled to the "instruction basis" of the Israeli Navy, where my (second) daughter had the ceremony of finishing her basic training, above is the logo of the Israeli Navy, the Navy is the only branch of the army who has ceremonial white uniforms, so it was all white day...

      This is also, somehow, the time in the year where many things related to event processing happen. Yesterday, IBM (the company who pays my salary) has done what was known internally as the "events event" in Boston for analysts and press, some report on it exist in the media, I think that David Berlind's report in "information week" is the most thorough one and includes video and slides. I have also seen several announcement by other vendors.

      Next week there will be the two back-to-back events, the Gartner EPS - second of its kind, follows by the EPTS event processing symposium - which is fourth of its kind, but first one after the formal EPTS launch. The first one to hear analysts reports and some other talks geared to customers, the second - highly interactive, discussion oriented meeting geared to the EP community. This is primarily for EPTS members, but we have invited some guests. The EPTS meeting will be video-taped by CITT, and will be posted on the Web, so additional people willl have access to the material.

      I'll blog more on these events next week. Looking forward to see the active those active in the community and also some new faces...





      Monday, September 8, 2008

      A footnote to the streamSQL paper

      The comment that my good friend Claudi (AKA Pattern Storm) made in the complexevents forum made me curious to actually read this paper; reading it I had the uncomfortable feeling that since people insist to use a language style that implies type of thinking about event processing, and this creates semantic problems which they try to solve by use the same type of thinking, with more complicated constructs.


      I'll use one simple example taken from the paper, which they had to deal with semantic problems that were caused by the way the language semantics.



      The scenario (translated to my language - without the "streams") -- Events are reported about cars that move through some segment of the road; each event consists of

      There are also simultaneous events, i.e. several events that happen in the same time unit (what ever the time granularity is). The inputs are events of this type, the output is - for each event, generate a derived event that include the original attributes of the events and the average speed of cars in the same time unit. If you want to see the types of problems that the SQL implementators see in this simple example, read the streamsql paper. Instead of discussing SQL, I would like to show an alternative way to think about the same problem.

      The slide below shows an alternative way to think about this problem - this is a very simple EPN (Event Processing Network) which has two functional agents, one producer (e.g. an event emitter that create events from video stream produced by a camera that looks at the road) and one consumer (whoever wants to see the output events)..




      The two agents work under the same temporal context (it can be spatio-temporal if we also want to group by road segment) - in this case, a temporal context is opened and closed every beginning and end of 1 time unit.

      • The raw event is called "car position event" and it goes to both agents.
      • The first agent is an aggregator which calculates (incrementally) the average, since it is bounded to the context, the average is of events from the same time unit, at the end of the time unit it produces a single event "speed-average-event" with the structure

      • The second agent is a "pattern detector" which takes two input events - the "car position event" again, and the derived event "speed-average-event"; the pattern that need to be identified is AND, and the "speed-average-event" for that agent has a consumption policy of "reuse" (which means that if an event can be used for multiple patterns). The agent produces a derived event - for each AND pattern that consists of the "output-event" whose structure is:

      This EPN does not involve "streams" - the thinking is "event oriented" and it attempts to provide natural thinking about event processing functionality.

      Comments:

      1. This is rather simple example, can also be solved by putting the average speed event on a global state (or event store/database) and then enrich it back - but the event-oriented is closer to the spirit of the original example which work on streams.

      2. Aggregator and pattern detector are type of agents, there are some (not many) more types. Typically, an event processing network consist of multiple types of agents.

      3. "Pattern Storm" claims that stream SQL ignore causality. One can view the relation between input events and output events of the same agent as a causality relation (he is using another scenario from the paper), and this can be set while defining the EPN.

      One general comment (not related to this posting) - to "anonymous" - I'll gladly answer your question if you'll send it back and identify yourself. I don't publish anonymous comments.

      I can post the solution to the rest of the examples in the stream SQL paper if anybody is interested...