oldmoe

NeverBlock, MySQL and MySQLPlus

Posted by oldmoe Thursday, August 28, 2008 8/28/2008 12:53:00 PM

Labels: mysql , neverblock , ruby , threads

I have great news for MySQL users. A very nice side effect has emerged from the development of the NeverBlock support for MySQL. I am glad to announce the release of a new MySQL driver for Ruby applications. It builds on top of the original Ruby MySQL driver but it comes with two notable additions:

Asynchronous query processing support

Threaded access support

Thanks to help from Roger Pack and Aman Gupta we were able to put the thing together that you can use and test right now (on Ruby1.8 and 1.9)

To install it please do:

sudo gem install espace-mysqlplus

Then you can use it in your code as follows:

require 'mysqlplus'
mysql = Mysql.real_connect(..)
mysql.query("select sleep(1)")

The test folder of the gem contains examples for threaded and evented implementations.

The announcement page in NeverBlock shows benchmark results for running the sleeping queries in normal(blocking), evented and threaded modes. The normal mode is 10X slower, which is normal due to its inability to run queries in parallel.

Now that Rails is becoming so-so-thread-safe this should show tremendous gains with Rails deployments that use MySQL (PostgreSQL already has such facilities).

ActiveRecord meets NeverBlock

Posted by oldmoe Monday, August 25, 2008 8/25/2008 03:18:00 AM

Labels: activerecord , neverblock , postgres , ruby , ruby 1.9

I happily announce the release of the first NeverBlock enabled activerecord adapter. The neverblock-postgresql-adapter. This is a beta release but I have been testing it for a while now with great results.

And while this is a big improvement it only requires you to replace the driver name in the connection to neverblock_postgresql instead of postgresql as described in the official neverblock blog

To make a long story short, this enables active record to issue queries in parallel, much like in a multi-threaded application. But this has several advantages over multi-threaded operations:

Fibers are cheaper than threads so this solution is theoretically faster.

NeverBlock does not require full thread safety, just avoid using globals and static variables for transient state.

It integrates nicely in evented programs thus eliminating the performance drop which occurs with the introduction of threads in such environments

I have benchmarked this against the plain postgresql adapter using different workloads categorized as follows

Very Light : A single count statement
Light : A single count and a create
Moderate : 2 counts and a create wrapped in a transaction that rolls back
Heavy : 3 counts, a create and an update wrapped in a transaction that commits
Very Heavy : 3 counts, one conditional count (on a non-indexed field), a create and two updates all wrapped in a transaction that commits

(if you are wondering why these queries in particular, they were extracted from some other code)

All were issued 1000 times

The results came as follows:

As you can see, NeverBlock::AR is persistently faster than vanilla AR. It appears that such work loads generate linear increase for both AR and NeverBlock::AR as the NeverBlock advantage was almost the same

Another benchmark was performed to test the effect of increasing the connection count for NeverBlock::AR. We tested with 2, 4, 8, 16 and 32 connections.

The benchmark consisted of first running "select 1" 5000 times and then running ~~"select sleep(10)"~~ "select sleep(1)" 20 times for each configuration.

As you can probably guess, increasing connection count has very little effect if the queries are all very fast (you cannot beat "select 1") but if the queries are all slow, you will be able to double the performance by simply doubling the connection count.

I hope this gives you a glimpse of what's coming next. Watch this space

101 Reasons Why PostgreSQL is a better fit for Rails than MySQL

Posted by oldmoe Wednesday, August 20, 2008 8/20/2008 07:07:00 PM

Labels: mysql , postgres , rails , ruby

1 - Indexing Support

MySQL cannot utilize more than one index per query. I believe this is worth repeating: MySQL CANNOT UTILIZE MORE THAN ONE INDEX PER QUERY. Wait till your tables get large enough and this will surely hit you. OTOH PostgreSQL can use multiple indices per query which come real handy.

2 - Full Text Indexing Support

MySQL can do full text indexing on MyISAM tables only, those working with InnoDB tables are out if luck. PostgreSQL has very advanced full text indexing capabilities wich enable you to control the tiniest details down to the stemming strategy.

3 - Asynchronous Interface

MySQL drivers are very unfriendly to the Ruby interpreter. Once a command is issued they take over until they come back with results. PostgreSQL sports a completely asynchronous interface where you can send queries to the database and then tend to other matters while the query is being processed by the server. The good news is that an Async ActiveRecord adapter for MySQL is being developed right now, as part of the rapidly growing NeverBlock library.

4 - Ruby Threading Aware

PostgreSQL dirvers enable the Ruby thread scheduler while IO requests are being processed (a nice side effect of the async interface). Which makes it much better suited for multithreaded Rails apps.

5 - Multistatements Per Query

Both MySQL and PostgreSQL support sending multiple statements separated by semi colons at once. But the returning result will be that of the last statement in the group. Now did you know that by using the async interface you can send multiple queries at once and then get back the results, one by one? One of the coolest features of the coming ActiveRecord (and Sequel btw) adapter is it's support for queuing queries to be consumed by a pool of connections. A trick we are contemplating working on is to group consequent selects together and send them in a single request to PostgreSQL and then later extract the results associated with each one of them. This is still very theoretical but should be verified soon.

Now that the 0b101 reasons are told I rest my case.

NeverBlock, much faster IO for Ruby

Posted by oldmoe 8/20/2008 06:54:00 PM

Labels: eventmachine , postgres , rev , ruby , ruby 1.9

At eSpace we have just released an alpha version of NeverBlock. A library that aims to bring evented IO to the masses. It does so by wrapping all IO in Fibers which handle all the async aspects and hides them totally from the developers.

Just as a teaser, here are some benchmarks of running PostgreSQL queries with and without NeverBlock

10x performance boost? how about that?

I am working on extending the NeverBlock library now, watch this space for great news soon

Document Matching In Ruby

Posted by oldmoe Wednesday, August 06, 2008 8/06/2008 09:55:00 PM

Labels: document matching , information retreival , ruby , tf-idf

With the debut of the new version of meOwns we have introduced several features that are concerned with how objects are related to each other. In the user profile page, you get a list of other user that are similar to him/her. And you are faced with the chemistry meter which tells you how much you are related to this user (if you are logged in of course). If you happen to be viewing your own profile page you will get a list of recommended items that you might be interested in. Last but not least, when you view an item, you get a list of similar items.

In this article I will be talking about the features from an implementation point of view. Naturally the first hurdle was to define the problem. What we needed was to find a way to match items and users. So first we needed to represent them in a way that can be matched. We started from the items first and looked at how to match two items together.

What is an item? In meOwns an item is simply a name, a description, a type and some tags. We decided to ignore photos and comments from the item. Each of those fields gets a weight which affects the value of terms found in it. Here is a sample product (encoded in Yaml):

Name : Fiat Sienna
Type : Car
Description : 1.6 HP, not bad for a sedan, relatively good performance for the price, best sedan i've bought
Tags : Cars Fiat Silver

The above fields are then processed to extract relevant terms from them, this is done in the following manner

Remove punctuation and non alphanumeric characters (replace them with spaces)

Collapse spaces and split the text on them

Match the generated list of terms to a stop word list to remove them (words like "on", "the" should not be considered in the index)

The remaining terms are converted to lower case and then converted to their stem representation (we use a snowball stemmer for now)

The above can be represented as follows:

Attribute : Value
===================
Name : fiat sienna
Type : car
Description : hp sedan performance price sedan buy
Tags : car fiat silver

After doing so, we create a list of terms with their frequencies, each field has a frequency multiplier according to its significance. Assumming a multiplier value of 1 for all the fields

Term : Frequency
=======================
fiat        : 2
car         : 2
sedan       : 2
sienna      : 1
performance : 1
silver      : 1
buy         : 1
hp          : 1

This Term-Frequency vector is the basis of doing item matching in meowns. Several different approaches can be implemented to reach an item representation. Even the details of a given approach can vary significantly. I have chosen to stick to the easiest approach in the initial implementation. Those willing to dig further are free to lookup more into document representation and indexing strategies.

User representation is just an aggregation of their item (and wished item) representations. This way user and items share the same term vector structure and hence we can match users to items as much as we can match items to items and users to users.

The matching process involves further encoding of the term frequency vector. Those in the academia refer to this as TF-IDF (Term Frequency, Inverse Document Frequency) representation. In lay man's terms this is a representation of how significant a term is to the certain document.

It is composed of the product of two parts. This first which is TF (Term Frequency) is simply the frequency of the term in the document divided by the total sum of frequency of terms in the same document.

The second (IDF) is the total number of documents in the corpse (the document store) divided by the number of documents which contain this term. If the term is found in all documents then the IDF will be equal to 1 and hence won't have an effect on the final product of the two parts. On the other hand, if we have a corpus of 1,000,000 documents and only one with the given term then it will multiply the TF value by 1,000,000 which is significant.

A slight variation (widely used) is to multiply the TF with the Logarithmic value of the IDF. After we are done calculating TF-IDF values for all the terms in term vectors matching can be done as follows

Term Vector A vs. Term Vector B = cos a = (A . B) / (|A|.|B|)
A.B = dot product for the two vectors of TF-IDF values 
|A|.|B| = scalar product of the magnitudes of the two vectors

The returned value is called the cosine similarity between the two items. It ranges between 0 and 1 where zero means no correlation and one means exact match. We are still experimenting with threshold values but these are essentially the figures you get when you see another user's chemistry meter for example. In another installment we will discuss how are we implementing behind the scenes matching of users and items in an efficient way.

Just to justify calling the post document matching in Ruby, here's a Ruby code to implement the above

# Monkey patch string to be able to extract terms from any string
Class String
  def to_terms(boost = 1, terms = {})
    # remove all non letters and reject stop words
    terms_list = self.gsub(/(\s|\d|\W)+/u,' ').rstrip.strip.split(' ').reject{|term|$stop_words.include?(term)}
    # transform to a hash with a frequency * boost value
    terms_list.each do|term|
      if terms[term]
        terms[term] = terms[term] + boost
      else
        terms[term] = boost
      end
    end
    terms
  end 
end

#our item class, which we match upon
Class Item 
  def to_terms(terms = {})
    #a hash of attributes to serialize 
    {:name => 10, 
     :description => 3, 
     :type_name => 1, 
     :tag_names => 1}.each_pair do |field, boost| 
      terms = send(field).to_terms(boost, terms)
    end
    terms
  end
end

The above methods enable us to extract the terms from different items. The first method refers to a global (bad me) stop word list which should be present.

Now that we have the items represented as term frequencies we can generate their tf-idf vectors (we will keep them as hashes though) and we can use them to do the matching

Class Item
  def to_tf_idf
    #assume we have a method that returns the df value for any term
    terms = self.to_terms
    total_frequency = terms.values.inject(0){|a,b|a+b}
    terms.each do |term,freq| 
      terms[term]= (freq / total_frequency) * self.df(term)
      magnitude = magnitude + terms[term]**2
    end
    terms, magnitude
  end
  def match(item)
    my_tf_idf,  my_magnitude  = self.to_tf_idf
    his_tf_idf, his_magnitude = item.to_tf_idf
    dot_product = 0
    my_tf_idf.each do |term,tf_idf|
      dot_product = dot_product + tf_idf * his_tf_idf[term] if his_tf_idf[term]
    end
    cosine_similarity = dot_product / (my_magnitude * his_magnitude)
  end
end

Pretty easy, now you get a value between 0 and 1 that represents how similar those two items are.

The case for a nonblocking Ruby stack

Posted by oldmoe Sunday, August 03, 2008 8/03/2008 12:48:00 AM

Labels: asynchronous , eventmachine , merb , nonblocking , performance , postgres , rails , ramaze , ruby , ruby 1.9 , sequel

In a previous post I talked about the problems that plauge the web based Ruby applications regarding processor and memory use. I proposed using non-blocking IO as a solution to this problem. In a follow up post I benchmarked nonblocking vs blocking performance using the async facilities in the Ruby Postgres driver in combination with Ruby Fibers. The results were very promising (up to 40% improvement) that I decided to take the benchmarking effort one step further. I monkey patched the ruby postgres driver to be fiber aware and was able to integrate it into sequel with little to no effort. Next I used the ~~unicycle~~ monorail server (the EventMachine HTTP server) in an eventmachine loop. I created a dumb controller which would query the db and render the results using the Object#to_json method.

As was done with the evented db access benchmark, a long query ran every n short queries (n belongs to {5, 10, 20, 50, 100}). The running application accepted 2 urls. One ran db operations in normal mode and the other ran in nonblocking mode (every action invocation was wrapped in a fiber in the latter case)

Here are the benchmark results

Full results

Comparing the number of requests/second fulfilled by each combination of blocking mode and conncurrency level. The first had the possible values of [blocking, nonblocking] the second had the possible values of [5, 10, 20, 50, 100]

Advantage Graph

Comparing the advantage gained for nonblocking over blocking mode for different long to short query ratios. Displaying the results for different levels of concurrency

And the full results in tabular form


                      Concurrent Requests
Ratio                 10     100    1000

1 To 100  Nonblocking 456.94 608.67 631.82
1 To 100  Blocking    384.82 524.39 532.26
          Advantage   18.74% 16.07% 18.71%
    
1 To 50   Nonblocking 377.38 460.74 471.89
1 To 50   Blocking    266.63 337.49 339.01
          Advantage   41.54% 36.52% 39.20%
    
1 To 20   Nonblocking 220.44 238.63 266.07
1 To 20   Blocking    142.6  159.7  141.92
          Advantage   54.59% 49.42% 87.48%
    
1 To 10   Nonblocking 130.87 139.76 195.02
1 To 10   Blocking    78.68  84.84  81.07
          Advantage   66.33% 64.73% 140.56%
    
1 To 5    Nonblocking 70.05  75.5   109.34
1 To 5    Blocking    41.48  42.13  41.77
          Advantage   68.88% 79.21% 161.77%

Conclusion

In accordance with my expectations. The nonblocking mode outperforms the blocking mode as long as enough long queries come into the play. If all the db queries are very small then the blocking mode will triumph mainly due to the overhead of fibers. But nevertheless, once there is a even single long query for every 100 short queries the performance is swayed into the nonblocking mode favor. There are still a few optimizations to be done, mainly complete the integrations with the EventMachine which should theoritically enhance performance. The next step is to integrate this into some framework and build a real application using this approach. Since Sequel is working now having Ramaze or Merb running in non-blocking mode should be a fairly easy task. Sadly Rails is out of the picture for now as it does not support Ruby 1.9 yet.

I reckon an application that does all its IO in an evented way will need much less processes per CPU core to make full use of it. Actually I am guessing that a single core can be maxed by a single process. If this is the case then I can live much happier if I can replace the 16 Thin processes running on my server with only 4. Couple that with the 30% memory savings we get from using RubyEE and we are talking about an amazing 82.5% memory foot print reduction without sacrificing performance.

Ruby Fibers Vs Ruby Threads

Posted by oldmoe 8/03/2008 12:12:00 AM

Labels: fibers , performance , ruby , ruby 1.9 , threads

Ruby 1.9 Fibers are touted as lightweight concurrency elements that are much lighter than threads. I have noticed a sizbale impact when I was benchmarking an application that made heavy use of fibers. So I wondered what If I switched to threads instead? After some time fighting with threads I decided I needed to write something specific for this comparison. I have written a small application that would spawn a number of fibers (or threads) and then would return the time went into this operation. I also recorded the VM size after the operation (all created fibers and threads are still reachable, hence, no garbage collection). I did not measure the cost of context switching for both approaches, may be in another time.

Here are the results for creation time:

And the results for memory usage:

Conclusion

Fibers are much faster to create than threads, they eat much less memory too. There is also a limit on the number of threads for 1.9 as I maxed on 3070 threads while fibers were not complaining when I created 100,000 of them (but they took 203 seconds and occuppied a whoping 500MB of RAM).

Final Veridect, No One To Blame!

Posted by oldmoe Sunday, July 27, 2008 7/27/2008 11:24:00 PM

More than a thousand humans drowned in the sea but apparently no one is to blame for this accident.

Apparently no one cares.

Are lives that cheap? Looks like some people believe so.

Faster IO for Ruby with Postgres

Posted by oldmoe Saturday, July 12, 2008 7/12/2008 02:55:00 AM

Labels: ansynchronous , eventmachine , fibers , nonblocking , postgres , ruby

Or 40% faster DB Access for your Ruby applications!

In a previous post I talked a bit about event based programming for Ruby. I mentioned the EventMachine/Asymy combo as a means of doing Asynchronous database operations hence freeing up the Ruby runtime to do other things while it is waiting on database I/O operations. Even more, the devs need not worry about using a different programming model, with the help of Ruby Fibers we will continue to program in the same old ways while Fibers will be doing all the twisted work underneath. Very promising indeed, but one big elephant in the room was the immaturity of the current solution. Asymy is still very infant and it is based on the super slow pure Ruby MySQL driver, not to mention that it is fairly incomplete as well.

So, what can we do about the elephant in the room? There is an Arabic proverb that basically says "Nothing can beat iron but iron" and this is exactly what we are going to do. Enter Postgres, the database with a realistic, unfriendly elephant mascot. Go away dolphins, a real elephant is in the room now.

Surprisingly Postgres happens to have an excellent asynchronous client API. It allows you to do almost all operations in a non blocking way. More surprisingly the Postgres driver for Ruby covers almost all those asynchronous API calls. The driver was originally written by Matz (yes the man himself) in 1997. It was later updated by ematsu in 1999 and now we have an update fresh from the oven in March 2008 by Jdavis. If you go through the C source code you will find many hidden gems. The methods are fairly well documented and you will discover that the driver has a blocking method that wraps the asynchronous calls inside but it does so in a Ruby threading friendly way. This way a threaded application will not block on the Postgres SQL commands. Good thing but I am more interested in the asynchronous side of the fence.

Let's walk through the API and see how can we use it to do non blocking database access. First you will need to install the gem ("sudo gem install pg"). Then you need to require 'pg' in your code.

One problem though before we start. The current driver has this nasty little bug that prevents you from setting the connection to nonblocking. It is actually a bug in the parameter count defined in the Ruby interface. A simple switch from 0 to 1 fixes this. To save you time and sweat I have provided a replacement gem with the modified sources (till the bug is fixed upstream). Now let's get back to the code.

Here's how to get things started

require 'pg'
# I have configured postgres to run in *trusted* mode
# so I don't need to supply a password
conn = PGconn.new({:host=>'localhost',:user=>'postgres',:dbname=>'evented'})
conn.setnonblocking(true)

This way our connection is ready for async operations. Now we need to start sending some sql commands to our connection. To do that we normally use the PGconn#exec method. But this method will block, waiting on postgres. So instead we will use the PGconn#send_query method. This method will return immediately, not waiting for Postgres to actually process the sql command. Here's how are going to use it.

conn.send_query("select * from users where name like '%am%'")
# the method will return immediately (or raise an exception in case of an error)

But wait, where are the results? Normally we expect the call to return with the data. Now where is my data? The results are being processed right now at the server side. We can continue to do other things till they come. But how do we know when they arrive? It turns out that this is easy as well. The PGconn instance provides a method that returns the connection's socket descriptor. PGconn#socket that is. We retrieve that socket descriptor and wrap it in a Ruby IO object by calling

io = IO.new(conn.socket)

Now have a nice IO object that we can get notified of its activity in a select call. For the uninitiated, event based programming is done by have a tight loop that runs forever. Within this loop we check if IO events happen and if so we respond to them. One efficient way of doing so is using the Ruby Kernel#select method (which is a wrapper to the UNIX select). The select method works that way: you provide it with three lists, one for sockets that you need to read from and one for sockets that you need to write to, the third is for errors that you are interested in. The call returns an array of the sockets that can be read/write or nil if none is ready.

We will use select as follows:

# the method that will be called if input is ready
def process_command(conn)
 # we will detail the implementation soon
end

loop do
 # we supply a list of sockets we need to read from.
 # Only our io object in this case. we nullify
 # the other lists and we set a timeout
 res = select([io],nil,nil, 0.001)
 # of course this needs to be done in a cleaner way
 process_command(conn) unless result.nil?
end

This way whenever there is info to read from the socket we will not get a nil (we will get an array actually) so we can call the process command. When the process command gets called it knows that there is data in the connection to be read so it calls the PGconn#consume_input method. After which it checks to see if the conn is busy or not. If it is still busy, it does nothing (it will do in a later event). On the other hand, if the connection is not busy then we start calling the PGconn#get_result method and append what we get to the result we got so far. We keep doing that till we get a nil result which indicates the end of the command and the readiness of the connection to accept further commands. Here is how the method will look like:

def process_command(conn)
 conn.consume_input
 unless conn.is_busy
  res, data = 0, []
  while res != nil
   res = get_result
   res.each {|d| data.push d}unless res.nil?  
  end
  #we are done, we need to put this data some where 
 end
end

Several things to be noted. First, one cannot process several commands using the same connection at once. You need several connections to achieve parallel command processing. Second, the model described above works in the twisted way, to get things working the normal way you can use Ruby Fibers (or continuations but they apparently leak memory)

I have put together a couple of Ruby classes that implement a nonblocking connection pool and a fiber pool. You can find them here Using those you can write code that looks like this:

require 'fiber_pool'
require 'fibered_connection_pool'

options = {:host=>'localhost',:user=>'postgres',:dbname=>'evented'}

cpool = FiberedC onnectionPool.new(options, 12)
# second param is the number of connections to spawn, defaults at 8
# note that one more connection than those will be spawned. This one
# will be used for processing blocking requests.

fpool = FiberPool.new(100)
# the number of fibers to spawn, defaults at 50

100.times do
 fpool.spawn do
  cpool.exec(some_sql_command, true) #true means async
  cpool.exec(some_other_sql_command, true)
  cpool.exec(yet_another_sql_command, true)
 end
end

# our event loop
loop do
 res = select(cpool.sockets,nil,nil,0) #check for something to read
 # IO is monkey patched to be able to hold a reference to the connection
 res.first.each{ |s|s.connection.process_command } if res
end

This works as follows, once a fiber calls cpool.exec the query is sent to the pool for processing and the fiber is halted, giving way for another one to start processing. The other one will halt as well once it hits a cpool.exec. Later during the event loop you will get notifications of completion of queries (in any order) and resume the fiber associated with the finished query. Note that commands issued in the same fiber will run sequentially while those issued from different fibers will interleave. This is effectively what is achieved by threading but without its costs.

Performance:

I am sure that my code might use some tweaking but I am getting very good results already. During benchmarking I found out that the cost on isntantiating fibers could be high (the cost of pausing and resuming is high as well, but unavoidable) So I created a pool of fibers that can be reused (a very naive implementation that can make use of lots of improvement).

I tested by issuing a group of long and short queries together. You actually provide the test program with the number of long queries and the multiplier it should use for short queries. i.e. ruby test.rb 10 20 will iterate 10 times and issue a long query then within the same iteration it will issue 20 short queries. It will do this in a blocking and then nonblocking way, reporting the time taken for each to complete and the percentage of performance increase/decrease.

I tested for 10, 50 and 100 long queries with the following multipliers (1, 2, 5, 10, 50, 100). The graph shows the performance gain for each number of queries vs the multiplier. For example 50 long queries with a multiplier of 10 (i.e. 500 short queries) achieves a 39.6% reduction in query execution time. I have repeated many of the tests several time (not all of them, too lazy to do that). The repeated tests showed consistent results so I am pretty confident of the presented results.

Here is the full list:


         Queries         Mode
Ratio  Long Short Blocking Non Blocking Advantage

:1/2   10   20    0.56     0.5          10.27%
       50   100   2.55     2.26         11.19%
       100  200   5.15     4.46         13.53%

:1/5   10   50    0.55     0.4          27.04%
       50   250   2.72     1.83         32.82%
       100  500   5.45     3.63         33.39%

:1/10  10   100   0.6      0.4          33.76%
       50   500   3.01     1.82         39.67%
       100  1000  5.9      3.65         38.13%

:1/20  10   200   0.72     0.45         38.12%
       50   1000  3.43     2.1          38.73%
       100  2000  6.83     4.33         36.53%

:1/50  10   500   0.98     0.62         36.57%
       50   2500  4.78     3.23         32.36%
       100  5000  9.74     8.68         10.93%

:1/100 10   1000  1.46     0.94         35.40%
       50   5000  7.42     5.17         30.31%
       100  10000 14.27    12.68        11.15%

The area I would like to focus on for performance tuning is the size of the fiber pool. The test is a bit sensitive to it so I believe I can gain a bit more performance with insane query counts if I optimize my fiber pool a bit. Setting the initial size too high certainly helps, but eats too much memory to make it usable.

A final note. I am playing with using this along side an EventMachine based http server. It works OK but is a cpu hog. Propably due to using select in next_tick calls withing EM's event loop. I would love to be able to provide EM with a list of IO objects and a call back instead of requiring me to use it to open the connection. Nevertheless, even though in many cases the nonblocking db implementation is slower than a blocking one in the http serving arena, I managed to get ~800 req/s vs ~500 req/s for a very typical use case, A request that runs a long query followed by many short ones. Impressive to say the least. I might be even try to hack EM to support the feature I need and then see what performance this could yield.

UPDATE

Apparently one can get more performance for the blocking requests if the fiber pool is initiated AFTER the blocking calls. Possibly due to the VM being impacted by the memory increase. Rerunning some of the tests showed fractional improvements for the blocking case. On the other hand, I tried some of the tests while another process was doing heavy I/O (RDoc generation). The performance gain jumped to an amazing 76% in one of the tests (it was generally between 51% and 76%).

Untwisting the Event Loop

Posted by oldmoe Tuesday, July 08, 2008 7/08/2008 04:19:00 AM

Labels: asymy , asynchronous , continuations , eventmachine , fibers , rails , ruby

Have you ever wondered why your Rails application is so memory hungry while it is not really trying to fully utilize your CPUs? To saturate your CPUs you have to have a large number of Thin (or Mongrel or whatever) instances. Why is that? We all know that the Ruby interpreter is not able to utilize more than one CPU (or no more than one CPU at a time in the case of 1.9). But why can't Ruby (or may be it's Rails?) utilize the processors efficiently? Let's look for an answer to this question.

First off, what happens in a typical Rails action? The Rails framework will be doing some request mapping and routing which is mostly CPU (if we consider memory latency negligible) Then a few requests will be sent to the database to retrieve some data after which a rendering process which is mostly CPU as well.

def show
  @user = User.find(params[id]) #db access
  @events = Events.find(:all) #another db access
  render :action => :show #rendering
end

The problem here comes with the database part of the action. Calls to the database will block processing till results get back from the DBMS. During that time, Rails will be frozen and not trying to do any thing else till the call ends. Good news is that threads can help here (even Ruby's green threads). A blocked thread will give way to other threads till it is back in the ready state. Thus filling those slots with some useful processing. Sounds good enough? NO!

Sadly Rails is NOT thread safe. You cannot use threads to do parallel processing in Rails. So why not something like Merb? I hear you say. Well Merb and threads will be able to interleave CPU operations and help with the time spent on IO in something like fetching data from some other service. But it won't save you when you do database IO. Simply because of the simple fact that calling C extensions blocks the whole Ruby interpreter. Yes, you read it correctly the first time. Nothing cannot be scheduled while a native call is being issued. Since database drivers are mostly C extensions they suffer from this. Your nice SELECT statement keeps the whole Ruby interpreter on hold till it is finished.

But there must be a solution to this. We cannot be all left high and dry with interpreters eating our memory and not really using our CPUs.

Enter EventMachine and AsyMy

For those who are not in the loop of events (bun intended) there happens to be another approach to this problem. Event based (read asynchronous) IO. In this mode of operation you request an IO operation and tell the event loop what to do when the request is fulfilled (either fully or partially). An excellent library for event handling exists for Ruby which is Francis' EventMachine (used internally by the Thin server and the evented flavour of Mongrel). But still, using EventMachine does not magically solve all our problems. The question that keeps popping up, what to do with database access? AsyMy to the rescue! AsyMy, written by Thomas Ptacek, is an evented driver for MySQL that operates in an asynchronous fashion. A quick example will look like:

connection.execute('SELECT * from events') do |headers,data|
  # do something with headers and data
  pp headers
  pp data
end

Asymy is still in a very early stage, the performance is horrible (as it is based on the darn slow pure Ruby MySQL driver) and it comes with many rough corners (I was not able to run INSERTs and UPDATEs without hacking it, and I am still not able to run the callbacks for those). Nevertheless, this is a formidable achievement on the road to a very fast single threaded implementation.

Here's how our action would look like if there was an Asymy adapter for ActiveRecord

#this is propably wrong but it can illustrate
#the twisted nature of evented programming
def show
  User.find(params[:id]) do |result_set|
      @user = result_set
      Events.find(:all) do |result_set|
          @events = result_set
          @events.each do |event|
              event.owner = @user
              if event != @events.last
                  event.save
              else
                  event.save do |ev|
                      render :action => :show
                  end
              end
          end 
      end
  end
end

We had to twist the function flow to be able to make use of the evented nature of the new driver. Instead of flow passing normally it is being scattered in the different callbacks. This is one of the areas where event based programming makes you change the way you think about program flow. A hurdle for many developers and a show stopper for some. No wonder the event library for Python is called Twisted

Why not untangle this with Fibers?

Fibers are lightweight concurrency primitives introduced in Ruby 1.9. How light weight? well they don't come at zero cost but in long running requests the weight they add can be negligible. Fibers provide some form of cooperative (rather than preemptive) concurrency inside a single thread (you cannot pass fibers between threads, you have been warned). Fibers enjoy the ability to pause and resume like continuations, but they don't suffer from the memory leaks the continuations have. When we use this feature wisely we can unwind the action code above to look like this:

def show
  @user = User.find(params[:id]
  @events = Events.find(:all).each do |event|
      event.owner = @user
      event.save
  end
end

Huh? this is the normal action code we are used to. Well, using fibers we can do this and still do things under the hood in an evented way.

To make things clear we need to illustrate Fibers with an example:

require 'fiber'

fiber = Fiber.new do
  #do something
  Fiber.yield another_thing
  #do yet another thing
end

yielded = fiber.resume # => runs the fiber till the yield,
                     # returns the yielded value
                     # and pauses the fiber where it is
fiber.resume #=> re-runs the fiber from the point it was paused.
fiber.resume #=> no more statements to run, raises an exception

Let's see how can this be useful for dispatching controller actions (this code will preferrably be in the server itself)

Fiber.new do
  Dispatcher.dispatch(controller,action,req,res)
  send_response res
end.resume

Inside the action we call the find method repeatedly. This method could be implemented like this:

class DataStore
  def find(*args)
      query = construct_query(*args)
      fiber = Fiber.current # grab the current fiber
      conn.execute(query) do |headers, data|
          fiber.resume convert_to_objects(data)
      end
      yield
  end
end

This way whenever the code passes a find method it will pass the query to the db driver, return immediately and pause, giving room for other requests to be processed. Once the data comes from the db server the call back is run and it resumes the fiber (passing to it the result of the query). The result gets passed back to the caller of the function and the original action method continues till completion (or till it is paused again by another find method)

Roger Pack has a nice writeup (with actual working code) on the Evented Fibered combo here.

Charles Jolley implemented a similar thing here. It is called Pipelined and while it is more obtrusive than the approach described above, it still has the advantage of being optional. Pipelined uses continuations and hence is available to Ruby 1.8 (and Rails).

I am still ironing out and tying things together (and doing lots of benchmarks) and I would like to tell you that I have ditched AsyMy for now for another alternative which I will attempt to discuss in detail in another blog post.

Meet Hamza

Posted by oldmoe Friday, June 06, 2008 6/06/2008 01:04:00 AM

Labels: baby , delivery , Hamza

وَلَقَدْ خَلَقْنَا الْإِنسَانَ مِن سُلَالَةٍ مِّن طِينٍ (12) ثُمَّ جَعَلْنَاهُ نُطْفَةً فِي قَرَارٍ مَّكِينٍ (13) ثُمَّ خَلَقْنَا النُّطْفَةَ عَلَقَةً فَخَلَقْنَا الْعَلَقَةَ مُضْغَةً فَخَلَقْنَا الْمُضْغَةَ عِظَاماً فَكَسَوْنَا الْعِظَامَ لَحْماً ثُمَّ أَنشَأْنَاهُ خَلْقاً آخَرَ فَتَبَارَكَ اللَّهُ أَحْسَنُ الْخَالِقِينَ (14) المؤمنون

Verily We created man from a product of wet earth (12); Then placed him as a drop (of seed) in a safe lodging (13); Then fashioned We the drop a clot, then fashioned We the clot a little lump, then fashioned We the little lump bones, then clothed the bones with flesh, and then produced it as another creation. So blessed be Allah, the Best of creators!(14) The Believers, Holy Quran

Hamza

Meet Hamza, my firstborn son. God's gift to me. I am so full of emotions that I am not able to express them adequately. It has been one very strange night when we left for the hospital. We wanted to have a natural childbirth but he refused to turn upside down. We decided to go for an operation and we had the luxury of presetting the date of birth. A bit less hassle than the shock that accompanies a surprising labor pain but still manages to make you very anxious.

We drove to the hospital, the doctor was there and my wife suddenly realized that there is no turning back and that she is going to actually be operated upon. She didn't panic or anything but she got her self one of the palest faces I ever saw on her. The doctor came to the room and recommended epidural anesthesia, which meant she will be awake during the operation. I gulped, as I was not sure if she will feel pain that way or not. Then twice gulped when the doctor asked me to change into an operation suit. "I will need your help there" he said.

I am not sure how did I change or how did the time pass when I was waiting in the doctor's area. Watching them washing their hands over and over (much like in the Panadol commercial). I sat there, silent, contemplating the moment. Thinking of what I might say to support her and if I am really capable of seeing her being cut. It was bizarre, ideas seemed to flash very fast in my mind and strangely it seemed totally empty at the same time. Something similar to that feeling when you try to memorize everything in your head minutes before an exam.

That state was preempted when they called for me. I went inside. My first time in an operation room. They were about to start. I sat on a chair so that my head was a bit higher than hers. I leaned forward to be level with her and started whispering in her ear. She was praying, she was repeating prayers and she was trying to remember the names of every person she knew. To pray for him/her. At that time I glanced over her. Just a glimpse as I ducked immediately after I saw that terrifying blood stream.

I tried as hard as I can to hide the effect of what I saw. I kept soothing her and helping her with her prayers. At that time we were both repeating: "My God, Nothing goes smooth but what you make smooth, and You are who makes hardship smooth" and "All power and might belong to Allah, the most High, the Great". Again I took another glimpse, and my God! I saw a little foot in the doctor's hand!.

It was pale, blueish white. The skin was wrinkled as wrinkled can be. The doctor was holding to it and pushing the rest of the baby out. My eyes were fixed on the scene in front of me. I was on complete awe. The miracle of birth. Exalted be Allah. I kept saying that between my words when I told my wife that I am seeing the baby. She was out of her breath. Firing questions as rapidly as she can manage. "Is he OK?", "Is he out already?". I tried to keep pace with her then resorted to "he's coming out, he's coming out"

It was a few minutes till the doctor manage to get him completely out. I saw him, very very small. Very vulnerable. Very weak. I couldn't help the tear forming in my eyes for the sight of him. He was crying as the nurses wrapped him and took him away for his first shower. My wife was tired beyond belief. But she asked me to run after him to take his first photos.

I ran for it and saw that they put him under some sort of a heater to keep him warm. I leaned on him, kissed him and recited the call to prayers in his both ears. They took him and showered him. And I watched the expression of disbelief on his face. It was funny to see such a thing. I wouldn't expect a different expression from an adult if was locked for nine months and then taken out like what happened to him.

I showered him myself with photos. Very funny ones. I hope the collection will grow. As I watch him grow.

I look at him now and see a much calmer expression on his face. Except when he is hungry of course. Getting him to that level means an inescapable scandal as he raises his voice loud enough to make sure that everyone in the neighborhood knows that we starve him.

Parenting is a new experience for me. I have always loved experiencing new things specially if I had a passion for them. In this case, I have more than a passion. I have love that I can't describe. I keep remembering how vulnerable I saw him that day. It affected me in a way that I have yet to understand.

How much horse power can your app engine deliver?

Posted by oldmoe Friday, April 18, 2008 4/18/2008 12:35:00 AM

Labels: appengine , appengines , benchmark , google , performance , scalability

I was wondering if the google app engine thing has the horse power to actually deliver the scalability experience that they are promising. I would expect that the app engines would be able to handle high loads and fulfill requests at decent rates.

I decided to benchmark the app engines and see if it can live to the promises that google are making.

For the benchmark I used a simple web page that visits the google data store, retrieves some data from it (small data set, so that the test wouldn't be bandwidth limited) then converts it to JSON (not using any library code for that, ugly hand written string manipulation code).

The web page (actually the currently super useless supergtd app) was written in accordance to the sample provided in the google app engine documentation. My Python skills are so immature but I can tell that everything is pretty straightforward. Route is determined, correct python file loaded and my get method is run. The method internally calls the data store and retrieves data then does some string processing and sends the result back.

I tested from a server at softlayer (this one has the most bandwidth and the least latency of all the servers that I have access too, and it's a dual quad core monster for the records)

Here are the results from the benchmark (click for a bigger version):

Frankly, I am very satisfied by the results. The app engine was capable of handling any load I threw at it. Even though the request I am testing is kinda simplistic it manages to pass through enough different parts of the stack to make it relevant. Many APIs will have resources that are only slightly more complex than the sample resource I created for this benchmark.

Bottom line, top notch performance. And you don't even need to bother with your Elastic Computing Cloud configuration. You don't manage instances or anything (or even pay by them) you only pay per usage which is a much sweeter deal if you ask me (if your app does not require that you cross the sandbox boundaries that is). Oh and before I forget, serving static content was pretty fast too. Not breathtakingly fast but very decent to say the least (you should never skip proper client caching though, at least to save the bandwidth)

If any one at Google is reading this I say "Great job, big hand for the big G". Now show me some Ruby love (and JavaScript too for what it's worth).

Did I mention that I need some Ruby love? Please.

Computing In The Cloud, Google App Engines

Posted by oldmoe Friday, April 11, 2008 4/11/2008 03:00:00 PM

Labels: appengines , google , python

One day we will plug our applications in the computing cloud as much as we plug our devices into the electricity network. Surprisingly, this days is today!

The trend towards offering computational infrastructure(storage, processing, bandwidth, etc) as a service is gaining momentum everyday. Services like Amazon's suite (S3, SimpleDB, EC2, SQS) and now Google AppEngines are manifestations of this trend. Infrastructure is becoming a commodity and large scale operations already invented their wheels, relieving the small to mid size player from having to invent wheels of their on. Stand on the shoulders of the giants and don't bother with performance, redundancy or data center setup overhead.

I have had a look at Google App Engines. In a nut shell, App Engines is an web application environment that is capable of running a subset of Python(self). Even though there is no write support to the disk; the available features cover more than 90% of web development needs. You are able to process HTTP requests, generate response, configure your routes, persist data and query it, and connect to other web services. One of the most cunning features is the ability to authenticate google accounts, plus one to whoever lobbied for this feature. Suddenly your application will be available to millions of users with no registration overhead (defining your user count gets tricky that way though).

One of the good things about Google App Engines is the choice of the programming language. Python(self) has dynamic and functional features that may be able to one day open the eyes of those who think that annotations are the next-big-thing!. I have been coding exclusively in Ruby and JavaScript (and of course GammaScript) for a while now. I have only experimented briefly with Python(self) earlier, so this was a good chance for a refresher. And Python(self) was fun except for some syntax quirks that I am not used to yet. I hope google extends their language support to cover Ruby and JavaScript (I stopped caring about Java as a web development tool long ago but others may be interested in including it as well). But kudos to Google for starting with Python(self).

I was over joyed with the ease of deployment. I faced a minor bug in the supplied local webserver but had an easy work around. I am still working off Google's web framework and didn't move to something like Django or Pylons yet.

Pros:

Platform comes with most of your development needs
Don't concern yourself with scaling, Google handles that for you, and is expected to do a very good job at it
Good choice for a starting language, Python(self) is a powerful daynamic language
Very good integration with Google accounts (brilliant!)
Ability to use full blown frameworks lik Django and Pylons

Cons:

Google lock in for your applications. Imagine that Google gets complaints against your application from governments (Chinese may be), will they shut it?
Sandboxes tend to constrain you. (that might be a pro rather than a con, it is your ticket for scaling)
For some reason Guido van Rossum thinks I should use "self" as the first argument in every method definition. I have read several explanations but I have yet to find something that convinces me.
I have been bitten with the white spaces thing a couple times. I attribute it to my acquired style though. Should be handled with time.

All in all, a very good move from Google, to them and everyone helping making cloud computing a reality. Thank you for bringing us the future!.

I am tempted to test Heroku now. I have been claiming that I was too busy before. This is no longer an excuse.

Preview of coming attractions, weNear

Posted by oldmoe Monday, April 07, 2008 4/07/2008 08:10:00 PM

Labels: android , iphone , LBS , location based services , wenear

coming soon, to a handset near you.

website, perview

enjoy

meOwns.com Site Tour

Posted by oldmoe 4/07/2008 08:08:00 PM

Labels: connection , interest , meowns , property

Here is a link to meOwns.com site tour. We are getting ready of an update in the coming weeks. Stay tuned.

An introduction to Ruby

Posted by oldmoe 4/07/2008 06:10:00 PM

Labels: programming language , ruby

I just gave a presentation on Ruby. Uploaded to scribd as usual (here)

Introducing GammaScript

Posted by oldmoe Sunday, March 23, 2008 3/23/2008 04:47:00 PM

Labels: CHAM , chemical programming , GAMMA , javascript

I have been experimenting with an implementation of the GAMMA formalism in Javascript. It is a VERY naive implementation that uses setTimeouts to mimick parallelism. Still it allows experimentation with the GAMMA and the chemical programming paradigm.

The current implementation can be found here. It provides a graphical representation of the multiset status (via the canvas element). A sample application is provided. An application that calculate the value for PI in a parallel way.

For the uninitiated, GAMMA provides a chemical like reaction model. For example, imagine that you have the following set (multiset) of numbers:

S -> 1,2,1,5,4,3,7,7,3,5,9,8,3,2,2,1,5

A GAMMA program to compute the max of this set would look something like this:

max(S):
   C: x >= y
   R: x,y -> x

or the more concise:
   x,y -> x for every x >= y

in GammaScript this can be written as:

var max = {
    condition : function(x,y){
        return x >=y;
    },
    reaction : function(x,y){
        return x.consume(y);
    }
}

As you can see, none of these dictates how the set should be traversed. It is totally left to the implementation. The current Javascript implementation provides pseudo (fake) parallelism.

I am considering a new implementation that utilizes Google Gears for true parallelism. And I need to refactor the view code from the core of the implementation.

*UPDATE*

To get the thing running you need to add data (comma separated) in the left box and click the (add) link. Then you should click the start button.

For the PI calculation programs that are supplied, you need to add a single element (the tuple [0,1]) and click add then start.

meOwns, a new way to express yourself

Posted by oldmoe Tuesday, March 18, 2008 3/18/2008 11:01:00 PM

Labels: espace , meowns

Have you noticed the widget to the left of this blog? This is a list of the items I have in meOwns, the new way to express yourself via your belongings. This is a Ruby on Rails application that is still in early beta. Registration is invitation only currently. If you would like an invitation then please send me an email with "meowns invitation required" in the subject to oldmoe at (g)mail dot com

Another JavaScript session

Posted by oldmoe 3/18/2008 05:27:00 PM

Labels: dom , javascript , prototype

I have uploaded my second session on JavaScript to scribd as well. Please find it here

My Latest Javascript Session

Posted by oldmoe Wednesday, March 12, 2008 3/12/2008 10:33:00 AM

Labels: javascript

I have uploaded the slides from my latest eSpace open session (this one about Javascript Internals) to scribd here. Please forgive the bad formatting. I hope this is useful (too many thanks for Crockford's writings, I would have been lost without those)