<?xml version="1.0" encoding="UTF-8"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
    <title>Francisco Claude-Faust</title>
    <link rel="self" type="application/atom+xml" href="https://recoded.cl/atom.xml"/>
    <link rel="alternate" type="text/html" href="https://recoded.cl"/>
    <generator uri="https://www.getzola.org/">Zola</generator>
    <updated>2020-05-05T11:00:16+00:00</updated>
    <id>https://recoded.cl/atom.xml</id>
    <entry xml:lang="en">
        <title>Binary search without division, shifts or multiplications</title>
        <published>2020-05-05T11:00:16+00:00</published>
        <updated>2020-05-05T11:00:16+00:00</updated>
        
        <author>
          <name>Unknown</name>
        </author>
        
        <link rel="alternate" type="text/html" href="https://recoded.cl/notes/binary-search-without-division-shifts-or-multiplications/"/>
        <id>https://recoded.cl/notes/binary-search-without-division-shifts-or-multiplications/</id>
        
        <summary type="html">&lt;p&gt;I subscribed to the &lt;a rel=&quot;external&quot; href=&quot;https://www.dailycodingproblem.com/&quot;&gt;daily coding problem&lt;/a&gt; a while ago by recommendation from a &lt;a rel=&quot;external&quot; href=&quot;https://www.linkedin.com/in/roberto-konow/&quot;&gt;friend&lt;/a&gt;. I only have the free version, just for fun. You get a problem per day, with no solution (unless you pay), the difficulty varies.&lt;/p&gt;
&lt;p&gt;Last week I got a problem that asks you to implement a search function that runs in $&lt;code&gt;O(\log n)&lt;/code&gt;$ time over a sorted array without using division, shifts, or multiplications.&lt;/p&gt;
&lt;p&gt;The problem is fun, and at least in my case, my first intuition was wrong. I will go through my solutions and compare them.&lt;/p&gt;</summary>
        
    </entry>
    <entry xml:lang="en">
        <title>Dicemath – a small Go service for keeping kids busy</title>
        <published>2020-04-19T15:00:00+00:00</published>
        <updated>2020-04-19T15:00:00+00:00</updated>
        
        <author>
          <name>Unknown</name>
        </author>
        
        <link rel="alternate" type="text/html" href="https://recoded.cl/notes/dicemat-go-service/"/>
        <id>https://recoded.cl/notes/dicemat-go-service/</id>
        
        <summary type="html">&lt;p&gt;I know this may not be the most fun activity for kids, but it turns out that my 5-year-old likes to do simple addition worksheets. Now that many of us are working from home with our kids, which is mostly trying to work from home while our kids prevent it, we have to take anything we can to keep the little ones busy and entertained.&lt;/p&gt;
&lt;p&gt;I can’t complain. Working from home with two little kids is hard, but the truth is that we are amazingly lucky that everyone close to us so far is healthy, and on top of that, we can work; I’m amazingly grateful for that!&lt;/p&gt;
&lt;p&gt;Back to the article, since my son likes to do worksheets where he can count the result of the additions, I built a small website that generates the worksheet. Every morning I print him one sheet, and he goes to either do the exercises or runs away from me because I’ll bug him with something he doesn’t want to do at the moment, win-win.&lt;/p&gt;
&lt;p&gt;The live site is no longer available (link removed). I’ll spend the rest of the article explaining how I built and set the site up. It’s a small example you can go through in one day, and it’s always fun to pause things and do something different. You can find the whole project &lt;a rel=&quot;external&quot; href=&quot;https://github.com/fclaude/dicemath&quot;&gt;here&lt;/a&gt; (no, I did not write any tests; this was a one-day fun project; that’s it).&lt;/p&gt;</summary>
        
    </entry>
    <entry xml:lang="en">
        <title>SICP Problem 2.6</title>
        <published>2020-04-15T03:00:00+00:00</published>
        <updated>2020-04-15T03:00:00+00:00</updated>
        
        <author>
          <name>Unknown</name>
        </author>
        
        <link rel="alternate" type="text/html" href="https://recoded.cl/notes/sicp-problem-2-6/"/>
        <id>https://recoded.cl/notes/sicp-problem-2-6/</id>
        
        <summary type="html">&lt;p&gt;This is a short &lt;a rel=&quot;external&quot; href=&quot;https://mitp-content-server.mit.edu/books/content/sectbyfn/books_pres_0/6515/sicp.zip/full-text/book/book-Z-H-14.html&quot;&gt;fun exercise&lt;/a&gt; I also went through while reading SICP. The problem gives a definition for zero and a function to compute the next integer (add 1):&lt;/p&gt;
&lt;pre class=&quot;giallo&quot; style=&quot;color-scheme: light dark; color: light-dark(#575279, #E0DEF4); background-color: light-dark(#FAF4ED, #232136);&quot; &gt;&lt;code data-lang=&quot;common-lisp&quot;&gt;&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;(define zero (&lt;/span&gt;&lt;span style=&quot;color: light-dark(#286983, #3E8FB0);&quot;&gt;lambda&lt;/span&gt;&lt;span&gt; (f) (&lt;/span&gt;&lt;span style=&quot;color: light-dark(#286983, #3E8FB0);&quot;&gt;lambda&lt;/span&gt;&lt;span&gt; (x) x)))&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;(define (add-1 n)&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;  (&lt;/span&gt;&lt;span style=&quot;color: light-dark(#286983, #3E8FB0);&quot;&gt;lambda&lt;/span&gt;&lt;span&gt; (f) (&lt;/span&gt;&lt;span style=&quot;color: light-dark(#286983, #3E8FB0);&quot;&gt;lambda&lt;/span&gt;&lt;span&gt; (x) (f ((n f) x)))))&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The problem asks you to write the function that adds two numbers without using add-1 iteratively.&lt;/p&gt;</summary>
        
    </entry>
    <entry xml:lang="en">
        <title>SICP Problem 1.28</title>
        <published>2020-03-27T14:00:34+00:00</published>
        <updated>2020-03-27T14:00:34+00:00</updated>
        
        <author>
          <name>Unknown</name>
        </author>
        
        <link rel="alternate" type="text/html" href="https://recoded.cl/notes/sicp-problem-1-28/"/>
        <id>https://recoded.cl/notes/sicp-problem-1-28/</id>
        
        <summary type="html">&lt;p&gt;In this post, I’ll go through some fun I had last week with problem 1.28 from &lt;a rel=&quot;external&quot; href=&quot;https://mitp-content-server.mit.edu/books/content/sectbyfn/books_pres_0/6515/sicp.zip/full-text/book/book-Z-H-11.html&quot;&gt;SICP&lt;/a&gt;. The problem statement asks you to implement the Miller-Rabin primality test in Scheme. I had not read this book and picked it up because of how many times I’ve seen it recommended. So far, I have to say it has been a fun ride; I’m enjoying it a lot!&lt;/p&gt;
&lt;p&gt;I started with that implementation, but then I sort of got into implementing my own ‘prime?’ function (a function that determines whether an integer is a prime number or not). I decided to limit it to implement a function that works for numbers up to 1 billion. This restriction is entirely arbitrary, but it’s big enough to have fun with it.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Important disclaimer&lt;/strong&gt;: the code in this article is a fun little exploration. If you need to test primality in real life, I recommend looking at &lt;a rel=&quot;external&quot; href=&quot;https://eprint.iacr.org/2018/749.pdf&quot;&gt;this paper&lt;/a&gt; before deciding on how to go with your implementation.&lt;/p&gt;</summary>
        
    </entry>
    <entry xml:lang="en">
        <title>Consistent Hashing (Part 1)</title>
        <published>2020-02-29T23:59:57+00:00</published>
        <updated>2020-02-29T23:59:57+00:00</updated>
        
        <author>
          <name>Unknown</name>
        </author>
        
        <link rel="alternate" type="text/html" href="https://recoded.cl/notes/consistent-hashing-part-1/"/>
        <id>https://recoded.cl/notes/consistent-hashing-part-1/</id>
        
        <summary type="html">&lt;p&gt;Consistent hashing is a well-known method for distributing stuff around multiple servers. We will implement a simple consistent hashing using Python and play a bit with it. This first part just focuses on having a way of computing hashes, we will build upon this in the later parts.&lt;/p&gt;</summary>
        
    </entry>
    <entry xml:lang="en">
        <title>My Old WordPress in docker</title>
        <published>2020-02-01T07:00:00+00:00</published>
        <updated>2020-02-01T07:00:00+00:00</updated>
        
        <author>
          <name>Unknown</name>
        </author>
        
        <link rel="alternate" type="text/html" href="https://recoded.cl/notes/old-wordpress-docker/"/>
        <id>https://recoded.cl/notes/old-wordpress-docker/</id>
        
        <content type="html" xml:base="https://recoded.cl/notes/old-wordpress-docker/">&lt;p&gt;This blog was offline for about 7 years. Before this, it was a self-hosted WordPress installation for which I had (somewhat buried in my backups) an up-to-date dump of the MySQL database. Once I managed to find the backup that survived to live in 5 houses, moving between 3 countries, and my messiness, I didn’t have anywhere to load it; the server hosting the blog was not there anymore. I was planning on using a one-click install in DigitalOcean, since I didn’t want to spend time setting up Apache, PHP, and all required things to run WordPress. I was not sure if I could plainly load this old database into a more modern WordPress, so I decided to experiment a bit running things in docker first.&lt;/p&gt;
&lt;p&gt;The first step was to load the database into a container; I had a file called fclaude_blog2.sql containing the old dump. I loaded the database (the default authentication plugin is because the default authentication of the latest version of MySQL is not supported, you can also get around by spinning an older version fo MySQL):&lt;/p&gt;
&lt;pre class=&quot;giallo&quot; style=&quot;color-scheme: light dark; color: light-dark(#575279, #E0DEF4); background-color: light-dark(#FAF4ED, #232136);&quot; &gt;&lt;code data-lang=&quot;plain&quot;&gt;&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;docker run --name wp-mysql -p 3306:3306  \&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;  -e MYSQL_ROOT_PASSWORD=nopass -d mysql \&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;  -default-authentication-plugin=mysql_native_password&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The default authentication plugin setting is because the default authentication of the latest version of MySQL is not supported. You can also get around by spinning an older version fo MySQL. Once the database was running, I just pushed the dump into it:&lt;/p&gt;
&lt;pre class=&quot;giallo&quot; style=&quot;color-scheme: light dark; color: light-dark(#575279, #E0DEF4); background-color: light-dark(#FAF4ED, #232136);&quot; &gt;&lt;code data-lang=&quot;plain&quot;&gt;&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;exec -i wp-mysql mysql -uroot -pnopass &lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;# (typed &amp;#39;create database fclaude_blog&amp;#39; and then closed the shell)&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;exec -i wp-mysql mysql -uroot -pnopass \&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;  fclaude_blog &amp;lt; Documents/fclaude_blog2.sql&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Now I could see my data, but I wanted the XML export to load into my new WordPress, and to do so, I needed to run WordPress and connect it to this database.&lt;/p&gt;
&lt;pre class=&quot;giallo&quot; style=&quot;color-scheme: light dark; color: light-dark(#575279, #E0DEF4); background-color: light-dark(#FAF4ED, #232136);&quot; &gt;&lt;code data-lang=&quot;plain&quot;&gt;&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;docker run -ti --name wp -p 8080:80 -e WORDPRESS_DB_USER=root \&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;  -e WORDPRESS_DB_PASSWORD=nopass -e WORDPRESS_DB_HOST=wp-mysql \&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;  -e WORDPRESS_DB_NAME=fclaude_blog --link wp-mysql -d wordpress&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Note that this is the settings of the MySQL we just set up, adding the –link option, so this container sees the container running the database.&lt;/p&gt;
&lt;p&gt;The most amazing part, at least to me, is that as soon as I accessed http://localhost:8080/wp-admin/ WordPress told me the database was in an old format and offered to upgrade it. It worked perfectly, so now I was ready to log into my blog, and get the file I needed! Except that I did not remember the password anymore!&lt;/p&gt;
&lt;p&gt;Following the WordPress documentation, I found that the password in the database was an md5. Per their recommendation, you could write the password in a text file, run the following command, and then delete the file (hopefully).&lt;/p&gt;
&lt;pre class=&quot;giallo&quot; style=&quot;color-scheme: light dark; color: light-dark(#575279, #E0DEF4); background-color: light-dark(#FAF4ED, #232136);&quot; &gt;&lt;code data-lang=&quot;plain&quot;&gt;&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;tr -d &amp;#39;\r\n&amp;#39; &amp;lt; pass.txt | md5sum | tr -d &amp;#39; -&amp;#39;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;I went into the MySQL container and updated the users’ table, setting the password field to the result from that command. And that was it! Then I could export my data and load it into this blog.&lt;/p&gt;
&lt;p&gt;After this, I just did:&lt;/p&gt;
&lt;pre class=&quot;giallo&quot; style=&quot;color-scheme: light dark; color: light-dark(#575279, #E0DEF4); background-color: light-dark(#FAF4ED, #232136);&quot; &gt;&lt;code data-lang=&quot;plain&quot;&gt;&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;docker kill wp wp-mysql&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;docker rm wp wp-mysql&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;And now, my machine is also not full of stuff I had no intention of permanently installing locally in the first place.&lt;/p&gt;
&lt;p&gt;Through all this experiment, the part I’m most surprised by is how WordPress handled the old database. I have no idea what changed, but still, now my database was up to date and working. Kudos to WordPress for this!&lt;/p&gt;
</content>
        
    </entry>
    <entry xml:lang="en">
        <title>Increasing the memory limit in Go</title>
        <published>2012-09-18T20:37:03+00:00</published>
        <updated>2012-09-18T20:37:03+00:00</updated>
        
        <author>
          <name>Unknown</name>
        </author>
        
        <link rel="alternate" type="text/html" href="https://recoded.cl/notes/increasing-memory-limit-in-go/"/>
        <id>https://recoded.cl/notes/increasing-memory-limit-in-go/</id>
        
        <content type="html" xml:base="https://recoded.cl/notes/increasing-memory-limit-in-go/">&lt;p&gt;Yeah, it&#39;s been a while, but instead of talking about why I haven&#39;t written anything, I&#39;ll just jump into the good stuff.&lt;/p&gt;
&lt;p&gt;I&#39;ve been developing in &lt;a rel=&quot;external&quot; href=&quot;http://golang.org&quot;&gt;Go&lt;/a&gt; lately. If you haven&#39;t tried it yet, I totally recommend it. Honestly, when you first look at the language, you tend to think that there is nothing new in there. That may be true. Many of the features are not unique, and you can find them in other languages, yet the combination is unique. After one week of using Go I didn&#39;t want to move back to C++ (what I usually use). There are a couple of things I miss from time to time, but the gain is so big that I don&#39;t mind.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The problem&lt;/strong&gt;. I&#39;ve had a couple of problems dealing with the kind of software I usually implement. Today I found one that did worry me a lot, I can&#39;t use more than 16GB of RAM! After that, my code will just crash with an &quot;out of memory&quot; error. In general this is not too bad, but for the particular algorithm I&#39;m running, I need more than 16GB. Luckily, I found out that the limitation is temporal and that you can lift it up by changing just a couple of lines in Go&#39;s source code. So here it goes.&lt;/p&gt;
&lt;p&gt;I will assume you installed Go in $HOME/bin/go (if not, don&#39;t worry, we are going to re-install it, right there).&lt;/p&gt;
&lt;p&gt;The first step is to get the source for Go:&lt;/p&gt;
&lt;pre class=&quot;giallo&quot; style=&quot;color-scheme: light dark; color: light-dark(#575279, #E0DEF4); background-color: light-dark(#FAF4ED, #232136);&quot; &gt;&lt;code data-lang=&quot;shellsession&quot;&gt;&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: light-dark(#797593, #908CAA);&quot;&gt;$&lt;/span&gt;&lt;span&gt; wget http://go.googlecode.com/files/go1.0.2.src.tar.gz&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Once that finishes, decompress the package:&lt;/p&gt;
&lt;pre class=&quot;giallo&quot; style=&quot;color-scheme: light dark; color: light-dark(#575279, #E0DEF4); background-color: light-dark(#FAF4ED, #232136);&quot; &gt;&lt;code data-lang=&quot;shellsession&quot;&gt;&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: light-dark(#797593, #908CAA);&quot;&gt;$&lt;/span&gt;&lt;span&gt; tar xvfz go1.0.2.src.tar.gz&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Now, using your favourite editor, change the following files as explained &lt;a rel=&quot;external&quot; href=&quot;http://code.google.com/p/go/issues/detail?id=2142#c11&quot;&gt;here&lt;/a&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;In file go/src/pkg/runtime/malloc.h&lt;br /&gt;
Change line 119, where it says &quot;MHeapMap_Bits = 22,&quot; for &quot;MHeapMap_Bits = 25,&quot;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;In file go/src/pkg/runtime/malloc.goc&lt;br /&gt;
Change line 309, where it says &quot;arena_size = 16LL&amp;lt;&amp;lt;30;&quot; for &quot;arena_size = 128LL&amp;lt;&amp;lt;30;&quot;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;And that&#39;s it, now lets install this.&lt;/p&gt;
&lt;pre class=&quot;giallo&quot; style=&quot;color-scheme: light dark; color: light-dark(#575279, #E0DEF4); background-color: light-dark(#FAF4ED, #232136);&quot; &gt;&lt;code data-lang=&quot;shellsession&quot;&gt;&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: light-dark(#797593, #908CAA);&quot;&gt;$&lt;/span&gt;&lt;span&gt; cd go/src&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: light-dark(#797593, #908CAA);&quot;&gt;$&lt;/span&gt;&lt;span&gt; ./all.bash&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: light-dark(#797593, #908CAA);&quot;&gt;$&lt;/span&gt;&lt;span&gt; cd ../..&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: light-dark(#797593, #908CAA);&quot;&gt;$&lt;/span&gt;&lt;span&gt; mkdir -p &lt;/span&gt;&lt;span style=&quot;font-style: italic;&quot;&gt;$&lt;/span&gt;&lt;span style=&quot;font-style: italic;&quot;&gt;HOME&lt;/span&gt;&lt;span&gt;/bin/go&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: light-dark(#797593, #908CAA);&quot;&gt;$&lt;/span&gt;&lt;span&gt; mv go &lt;/span&gt;&lt;span style=&quot;font-style: italic;&quot;&gt;$&lt;/span&gt;&lt;span style=&quot;font-style: italic;&quot;&gt;HOME&lt;/span&gt;&lt;span&gt;/bin/&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Now we are ready. Remember to set the environment variable in your .bashrc:&lt;/p&gt;
&lt;pre class=&quot;giallo&quot; style=&quot;color-scheme: light dark; color: light-dark(#575279, #E0DEF4); background-color: light-dark(#FAF4ED, #232136);&quot; &gt;&lt;code data-lang=&quot;shellscript&quot;&gt;&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: light-dark(#286983, #3E8FB0);&quot;&gt;export&lt;/span&gt;&lt;span style=&quot;font-style: italic;&quot;&gt; GOROOT&lt;/span&gt;&lt;span style=&quot;color: light-dark(#286983, #3E8FB0);&quot;&gt;=&lt;/span&gt;&lt;span style=&quot;font-style: italic;&quot;&gt;$&lt;/span&gt;&lt;span style=&quot;font-style: italic;&quot;&gt;HOME&lt;/span&gt;&lt;span&gt;/&lt;/span&gt;&lt;span style=&quot;font-style: italic;&quot;&gt;bin&lt;/span&gt;&lt;span&gt;/&lt;/span&gt;&lt;span style=&quot;font-style: italic;&quot;&gt;go&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: light-dark(#286983, #3E8FB0);&quot;&gt;export&lt;/span&gt;&lt;span style=&quot;font-style: italic;&quot;&gt; PATH&lt;/span&gt;&lt;span style=&quot;color: light-dark(#286983, #3E8FB0);&quot;&gt;=&lt;/span&gt;&lt;span style=&quot;font-style: italic;&quot;&gt;$&lt;/span&gt;&lt;span style=&quot;font-style: italic;&quot;&gt;PATH&lt;/span&gt;&lt;span&gt;:&lt;/span&gt;&lt;span style=&quot;font-style: italic;&quot;&gt;$&lt;/span&gt;&lt;span style=&quot;font-style: italic;&quot;&gt;GOROOT&lt;/span&gt;&lt;span&gt;/&lt;/span&gt;&lt;span style=&quot;font-style: italic;&quot;&gt;bin&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;You are ready to enjoy 128GB. If you want 64GB, change the 25 for 24, and the 128 by 64.&lt;/p&gt;
&lt;p&gt;This is supposed to be experimental, so please don&#39;t blame me if things crash. There has to be a reason for that limitation to be there. In my case, I was able to get away with this and run my code. In this particular case is the (expensive) construction of a compressed text index, so the thing runs, outputs, and that&#39;s it. Afterwards I process the output with other binaries. I&#39;m just happy if the program manages to go through the construction, without much care on the resources I&#39;m using at the time.&lt;/p&gt;
</content>
        
    </entry>
    <entry xml:lang="en">
        <title>SXSI</title>
        <published>2011-11-01T03:19:55+00:00</published>
        <updated>2011-11-01T03:19:55+00:00</updated>
        
        <author>
          <name>Unknown</name>
        </author>
        
        <link rel="alternate" type="text/html" href="https://recoded.cl/notes/sxsi/"/>
        <id>https://recoded.cl/notes/sxsi/</id>
        
        <content type="html" xml:base="https://recoded.cl/notes/sxsi/">&lt;p&gt;After a while I can happily say that our &lt;a href=&quot;https://recoded.cl/notes/the-sxsi-system/&quot;&gt;SXSI system&lt;/a&gt; is available for &lt;a rel=&quot;external&quot; href=&quot;http://www.lri.fr/~kn/files/sxsi/sxsi.tar.gz&quot;&gt;downloading&lt;/a&gt;. Kim created a package that can be easily installed and tested. In this post, I include a step by step tutorial on how to build the system and index a sample XML file.&lt;/p&gt;
&lt;p&gt;First we need to install the dependencies in our system. In this case, I&#39;ll show how to do it in Ubuntu 11.10.&lt;/p&gt;
&lt;pre class=&quot;giallo&quot; style=&quot;color-scheme: light dark; color: light-dark(#575279, #E0DEF4); background-color: light-dark(#FAF4ED, #232136);&quot; &gt;&lt;code data-lang=&quot;shellsession&quot;&gt;&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: light-dark(#797593, #908CAA);&quot;&gt;$&lt;/span&gt;&lt;span&gt; sudo apt-get install ocaml-ulex ocaml-findlib ocaml-nox libxml++2.6-dev camlp4&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Then, &lt;a rel=&quot;external&quot; href=&quot;http://www.lri.fr/~kn/files/sxsi/sxsi.tar.gz&quot;&gt;download&lt;/a&gt; the package. To decompress and compile the package do:&lt;/p&gt;
&lt;pre class=&quot;giallo&quot; style=&quot;color-scheme: light dark; color: light-dark(#575279, #E0DEF4); background-color: light-dark(#FAF4ED, #232136);&quot; &gt;&lt;code data-lang=&quot;shellsession&quot;&gt;&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: light-dark(#797593, #908CAA);&quot;&gt;$&lt;/span&gt;&lt;span&gt; tar xvf sxsi.tar.gz&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: light-dark(#797593, #908CAA);&quot;&gt;$&lt;/span&gt;&lt;span&gt; cd sxsi/libcds&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: light-dark(#797593, #908CAA);&quot;&gt;$&lt;/span&gt;&lt;span&gt; make&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: light-dark(#797593, #908CAA);&quot;&gt;$&lt;/span&gt;&lt;span&gt; cd ../xpathcomp/src/XMLTree&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: light-dark(#797593, #908CAA);&quot;&gt;$&lt;/span&gt;&lt;span&gt; make clean all&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: light-dark(#797593, #908CAA);&quot;&gt;$&lt;/span&gt;&lt;span&gt; cd ../../&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: light-dark(#797593, #908CAA);&quot;&gt;$&lt;/span&gt;&lt;span&gt; ./configure&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: light-dark(#797593, #908CAA);&quot;&gt;$&lt;/span&gt;&lt;span&gt; ./build&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: light-dark(#797593, #908CAA);&quot;&gt;$&lt;/span&gt;&lt;span&gt; cd xpathcomp&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;This generates a binary called &quot;main.native&quot;. This binary allows you to index and query xml files. Lets first generate an xml file with the standard &lt;a rel=&quot;external&quot; href=&quot;https://github.com/fclaude/xtrie/blob/master/gen_xml.c&quot;&gt;xmark&lt;/a&gt; tool&lt;/p&gt;
&lt;pre class=&quot;giallo&quot; style=&quot;color-scheme: light dark; color: light-dark(#575279, #E0DEF4); background-color: light-dark(#FAF4ED, #232136);&quot; &gt;&lt;code data-lang=&quot;shellsession&quot;&gt;&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: light-dark(#797593, #908CAA);&quot;&gt;$&lt;/span&gt;&lt;span&gt; ./gen_xml -f 1 &lt;/span&gt;&lt;span style=&quot;color: light-dark(#286983, #3E8FB0);&quot;&gt;&amp;gt;&lt;/span&gt;&lt;span&gt; sample.xml&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Now, in order to index this file we run:&lt;/p&gt;
&lt;pre class=&quot;giallo&quot; style=&quot;color-scheme: light dark; color: light-dark(#575279, #E0DEF4); background-color: light-dark(#FAF4ED, #232136);&quot; &gt;&lt;code data-lang=&quot;shellsession&quot;&gt;&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: light-dark(#797593, #908CAA);&quot;&gt;$&lt;/span&gt;&lt;span&gt; ./main.native -v -s sample.srx sample.xml &lt;/span&gt;&lt;span style=&quot;color: light-dark(#EA9D34, #F6C177);&quot;&gt;&amp;quot;&lt;/span&gt;&lt;span style=&quot;color: light-dark(#EA9D34, #F6C177);&quot;&gt;&amp;quot;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;Parsing XML Document : 84729.9ms&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;Mem use before: VmRSS:	    5380 kB&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;Mem use after: VmRSS:	  117652 kB&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;Building TextCollection : 189371ms&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;Mem use before: VmRSS:	  117916 kB&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;Mem use after: VmRSS:	  237264 kB&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;Building parenthesis struct : 303.961ms&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;Mem use before: VmRSS:	  237264 kB&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;Mem use after: VmRSS:	  237528 kB&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;Tags blen is 8&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;Building Tag Structure : 4715.13ms&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;Mem use before: VmRSS:	  237528 kB&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;Mem use after: VmRSS:	  214988 kB&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;Number of distinct tags 92&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;Building tag relationship table: 1209.822893ms&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;Parsing document: 280366.960049ms&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;Writing file to disk: 115.392923ms&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;character 0-0 Stream.Error(&amp;quot;illegal begin of query&amp;quot;)&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Yes, I gave an empty query, but I didn&#39;t want to search anything yet ;-). We have an indexed version of sample.xml with standard options saved as sample.srx. You can play with different indexing options, run main.native without parameters to see the full set of options supported at the moment. The result so far looks like this:&lt;/p&gt;
&lt;pre class=&quot;giallo&quot; style=&quot;color-scheme: light dark; color: light-dark(#575279, #E0DEF4); background-color: light-dark(#FAF4ED, #232136);&quot; &gt;&lt;code data-lang=&quot;shellsession&quot;&gt;&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: light-dark(#797593, #908CAA);&quot;&gt;$&lt;/span&gt;&lt;span&gt; ls -lh sample.&lt;/span&gt;&lt;span style=&quot;color: light-dark(#286983, #3E8FB0);&quot;&gt;*&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;-rw-r--r-- 1 fclaude fclaude 188M 2011-10-31 23:10 sample.srx&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;-rw-rw-r-- 1 fclaude fclaude 112M 2011-10-31 23:02 sample.xml&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Finally, as an example, we can count the number of results for the query &quot;/site/regions/africa&quot;:&lt;/p&gt;
&lt;pre class=&quot;giallo&quot; style=&quot;color-scheme: light dark; color: light-dark(#575279, #E0DEF4); background-color: light-dark(#FAF4ED, #232136);&quot; &gt;&lt;code data-lang=&quot;shellsession&quot;&gt;&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: light-dark(#797593, #908CAA);&quot;&gt;$&lt;/span&gt;&lt;span&gt; ./main.native -c -v sample.srx /site/regions/africa&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;Loading tag table: 4.344940ms&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;Loading parenthesis struct : 304.395ms&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;Mem use before: VmRSS:	    5380 kB&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;Mem use after: VmRSS:	    7492 kB&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;Loading tag names struct : 0.049ms&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;Mem use before: VmRSS:	    7492 kB&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;Mem use after: VmRSS:	    7492 kB&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;tags_blen is 8&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;11 MB for tag sequence&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;Loading tag struct : 11.366ms&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;Mem use before: VmRSS:	    7492 kB&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;Mem use after: VmRSS:	   27492 kB&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;Loading text bitvector struct : 11.738ms&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;Mem use before: VmRSS:	   27492 kB&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;Mem use after: VmRSS:	   29076 kB&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;Loading TextCollection : 144.782ms&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;Mem use before: VmRSS:	   29076 kB&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;Mem use after: VmRSS:	  194604 kB&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;Loading file: 478.782892ms&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;Parsing query: 0.061035ms&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;Parsed query:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;/child::site/child::regions/child::africa&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;Compiling query: 0.048876ms&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;Automaton (0) :&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;States {q₀ q₁ q₂ q₃ q₄}&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;Initial states: {q₀}&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;Marking states: {q₀ q₁ q₂ q₃}&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;Topdown marking states: {q₀ q₁ q₂ q₃}&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;Bottom states: {q₀ q₁ q₂ q₃ q₄}&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;True states: {q₄}&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;Alternating transitions&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;_________________________________&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;(q₀, {&amp;#39;&amp;#39; })         → ↓₂q₀ ∧ ↓₁q₁&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;(q₀, Σ)             → ↓₂q₀&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;(q₁, {&amp;#39;site&amp;#39; })     → ↓₂q₁ ∧ ↓₁q₂&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;(q₁, Σ)             → ↓₂q₁&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;(q₂, {&amp;#39;regions&amp;#39; })  → ↓₂q₂ ∧ ↓₁q₃&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;(q₂, Σ)             → ↓₂q₂&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;(q₃, {&amp;#39;africa&amp;#39; })   ⇒ ↓₂q₃ ∧ ↓₁q₄&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;(q₃, Σ)             → ↓₂q₃&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;(q₄, Σ)             → ↓₂q₄ ∧ ↓₁q₄&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;_________________________________&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;Execution time: 1.343012ms&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;Number of results: 1&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;Maximum resident set size: VmHWM:	  200084 kB&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;And that&#39;s it. Now you can play further with SXSI :-)&lt;/p&gt;
</content>
        
    </entry>
    <entry xml:lang="en">
        <title>Disk Crash, Recovering Files and Doing Backups</title>
        <published>2011-05-16T05:11:17+00:00</published>
        <updated>2011-05-16T05:11:17+00:00</updated>
        
        <author>
          <name>Unknown</name>
        </author>
        
        <link rel="alternate" type="text/html" href="https://recoded.cl/notes/disk-crash-and-backups/"/>
        <id>https://recoded.cl/notes/disk-crash-and-backups/</id>
        
        <summary type="html">&lt;p&gt;About one and a half weeks ago I had a disk crash. I didn&#39;t lose anything, but was pretty close, mainly because I deleted by hand an important file :-(.&lt;/p&gt;
&lt;p&gt;It is interesting to talk about my disk crash because I faced many problems to bring my computer back. Luckily my notebook has two hard drives, so I&#39;m up and running with the secondary one now, but to do so I had to clone the recovery partition. Then I copied my backups to my home directory, and in the process I deleted an important file, which took me days to recover. And finally, I wrote a small script to back up things. This script is not a program (there is no error checking or anything like that), but it shows how to keep an encrypted backup.&lt;/p&gt;</summary>
        
    </entry>
    <entry xml:lang="en">
        <title>Land of LISP and Lazy Evaluation</title>
        <published>2011-05-01T10:00:24+00:00</published>
        <updated>2011-05-01T10:00:24+00:00</updated>
        
        <author>
          <name>Unknown</name>
        </author>
        
        <link rel="alternate" type="text/html" href="https://recoded.cl/notes/land-of-lisp-and-lazy-evaluation/"/>
        <id>https://recoded.cl/notes/land-of-lisp-and-lazy-evaluation/</id>
        
        <summary type="html">&lt;p&gt;One of my recent read was &lt;a rel=&quot;external&quot; href=&quot;http://landoflisp.com/&quot;&gt;&quot;Land of LISP&quot; by Conrad Barski&lt;/a&gt;. It&#39;s an unconventional programming book, packed with comics. Most examples are tiny games you can code in a couple of pages. I personally liked the book a lot, it is fun to read and presents a language that, in my opinion, is also fun.&lt;/p&gt;
&lt;p&gt;My experience with LISP is quite limited, so most of the things I found in the book were new to me (I knew how to define functions and basic stuff, but nothing about macros, only a couple of built-in functions, etc.). One of the things I liked the most was one of the examples for macros where the author presents a simple solution to get lazy evaluation. In the book, the author codes a game with some weird rules, and I don&#39;t think I would learn much by just copying that example, therefore, I will use the same idea here but with our old friend tic-tac-toe. I have to warn you that the implementation I&#39;m going to post is not a good tic-tac-toe implementation, you can probably find more challenging opponents. The main goal of this exercise is to illustrate the lazy evaluation. Having said that, lets get started.&lt;/p&gt;</summary>
        
    </entry>
    <entry xml:lang="en">
        <title>Finding the most significant bit</title>
        <published>2011-04-13T23:59:49+00:00</published>
        <updated>2011-04-13T23:59:49+00:00</updated>
        
        <author>
          <name>Unknown</name>
        </author>
        
        <link rel="alternate" type="text/html" href="https://recoded.cl/notes/computing-the-most-significant-bit/"/>
        <id>https://recoded.cl/notes/computing-the-most-significant-bit/</id>
        
        <summary type="html">&lt;p&gt;In one of the problems I&#39;ve been working on with Diego Seco and Patrick Nicholson, we needed the most-significant-bit (msb) function as a primitive for our solution. As Diego pointed out today, this function was one of the bottlenecks of our structure, consuming a considerable amount of time.&lt;/p&gt;
&lt;p&gt;In this post I&#39;ll go through the solutions we tested in our data structure.&lt;/p&gt;</summary>
        
    </entry>
    <entry xml:lang="en">
        <title>Seven Languages in Seven Weeks</title>
        <published>2011-04-04T07:45:17+00:00</published>
        <updated>2011-04-04T07:45:17+00:00</updated>
        
        <author>
          <name>Unknown</name>
        </author>
        
        <link rel="alternate" type="text/html" href="https://recoded.cl/notes/seven-languages-in-seven-weeks/"/>
        <id>https://recoded.cl/notes/seven-languages-in-seven-weeks/</id>
        
        <content type="html" xml:base="https://recoded.cl/notes/seven-languages-in-seven-weeks/">&lt;p&gt;I read this book, by &lt;a rel=&quot;external&quot; href=&quot;http://blog.rapidred.com/&quot;&gt;Bruce Tate&lt;/a&gt;, some weeks ago and totally recommend it. You can buy it from &lt;a rel=&quot;external&quot; href=&quot;http://www.amazon.com/gp/product/193435659X&quot;&gt;amazon&lt;/a&gt; or &lt;a rel=&quot;external&quot; href=&quot;https://secure.pragprog.com/titles/btlang/seven-languages-in-seven-weeks/&quot;&gt;the pragmatic bookshelf&lt;/a&gt;.&lt;/p&gt;
&lt;h3 id=&quot;what-s-in-it&quot;&gt;&lt;strong&gt;What&#39;s in it?&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;From the title of the book is easy to get an idea. This book explores seven different programming languages, the idea is that you spend one week using each programming language and get an idea of what&#39;s out there to offer alternatives to the standards we are used to. By the standard I mean what I consider the standard, based on my experience most people program in Python, PHP, C++, C# or Java**.**&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The languages covered are&lt;/strong&gt;:  &lt;a rel=&quot;external&quot; href=&quot;http://www.ruby-lang.org/en/&quot;&gt;Ruby&lt;/a&gt;, &lt;a rel=&quot;external&quot; href=&quot;http://iolanguage.com/&quot;&gt;Io&lt;/a&gt;, &lt;a rel=&quot;external&quot; href=&quot;http://en.wikipedia.org/wiki/Prolog&quot;&gt;Prolog&lt;/a&gt;, &lt;a rel=&quot;external&quot; href=&quot;http://www.scala-lang.org/&quot;&gt;Scala&lt;/a&gt;, &lt;a rel=&quot;external&quot; href=&quot;http://www.erlang.org/&quot;&gt;Erlang&lt;/a&gt;, &lt;a rel=&quot;external&quot; href=&quot;http://clojure.org/&quot;&gt;Clojure&lt;/a&gt;, and &lt;a rel=&quot;external&quot; href=&quot;http://www.haskell.org/haskellwiki/Haskell&quot;&gt;Haskell&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;The book has one chapter per language (plus intro and wrap-up chapters). Each language is presented in three parts, called days. Usually the first one shows the very basics, like input/output or math operations. The second day is usually used to present something that is different for this language, and the third day to present a harder example where the language shows to be useful and superior to others. For example, in Prolog, you can solve Sudoku puzzles by day 3 with little effort.&lt;/p&gt;
&lt;h3 id=&quot;why-do-i-recommend-it&quot;&gt;&lt;strong&gt;Why do I recommend it?&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;I liked several things about this book.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;The third day shows you an example where the language actually helps you.&lt;/strong&gt; This is in contrast to what my programming languages course was,  where we went through 4 different languages. In that course we implemented pretty much the same things in each language. Maybe someone would disagree with my following statement, but in my opinion, implementing QuickSort in Prolog just makes you think you are wasting your time. You know how to implement it in C, and it works faster, so what&#39;s the point? In this book you don&#39;t do that kind of things. As I mentioned before, one of the examples for Prolog is solving a Sudoku puzzle. This shows you something great about the language. You can do this with almost no effort. If you try implementing it in another language, say Java, it is certainly going to take more effort.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The presentation is clear.&lt;/strong&gt; It is easy to follow and you don&#39;t feel lost at some point looking at code you don&#39;t understand. For instance, in the basics, you usually go through basic math operations which are really similar to at least one language you know, yet the author takes the time to go through them. The explanations are of the right length to not be boring either, which is also a great plus. Another point regarding the basics, the author goes into things like typing with examples of that part, which also makes it much more interesting.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;You are likely to find lots of things you didn&#39;t know.&lt;/strong&gt; In my case, I haven&#39;t been exposed to that many programming languages, most of them imperative and object oriented. The only language I knew besides that was Scheme and some Common Lisp. I had a great time reading about these other languages that are somehow different and offer powerful constructs that make it easy to do things I know to be hard. For example, the future objects in Io, or the way you can build a service monitor in Erlang.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The author is not selling you the languages.&lt;/strong&gt; Well, maybe some of them more than others. But the important point is that I didn&#39;t find the book to be too biased. In fact, for every language the author presents it and finishes listing advantages and disadvantages. This is a great thing, not only you get information about when to use a language, but also when not to.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;You get to hear from the people that invented the languages (or important users of it).&lt;/strong&gt; In every chapter you get to see an interview with someone that can be considered as a worthy representative of the language. I found this really interesting, and adds external opinions to the book, which I think, increases its value.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;It is important to make clear that you will not learn the seven languages by reading this book. I would not recommend it to someone trying to learn how to program or trying to learn one of those seven languages.&lt;/p&gt;
&lt;p&gt;If you know how to program, and enjoy it, you should consider buying it.&lt;/p&gt;
</content>
        
    </entry>
    <entry xml:lang="en">
        <title>The SXSI System</title>
        <published>2010-04-02T04:49:48+00:00</published>
        <updated>2010-04-02T04:49:48+00:00</updated>
        
        <author>
          <name>Unknown</name>
        </author>
        
        <link rel="alternate" type="text/html" href="https://recoded.cl/notes/the-sxsi-system/"/>
        <id>https://recoded.cl/notes/the-sxsi-system/</id>
        
        <content type="html" xml:base="https://recoded.cl/notes/the-sxsi-system/">&lt;p&gt;I promised a post about the &lt;del&gt;sexy&lt;/del&gt; SXSI system in my last post. Its awesome name stands for Fast In-Memory XPath Search over Compressed Text and Tree Indexes, which says much more about what it does.&lt;/p&gt;
&lt;p&gt;The main idea for representing the XML file is as follows: the shape of the tree is represented using succinct balanced parentheses, the labels are represented with structures that support rank/select and access queries, and the text with a self-index for the collection of texts inside the xml file.&lt;/p&gt;
&lt;p&gt;Let&#39;s go over each part:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;The tree: the tree is represented using balanced parentheses, by doing a DFS over the tree and writing a &#39;(&#39; every time you start visiting a subtree and a &#39;)&#39; when you leave the subtree. This representation allows for very fast navigation over the tree, for a good reference about it see the &lt;a rel=&quot;external&quot; href=&quot;http://www.dcc.uchile.cl/~gnavarro/ps/alenex10.pdf&quot;&gt;paper&lt;/a&gt; by Arroyuelo, Canovas, Navarro and Sadakane.  In the same order we keep another sequence that corresponds to the ids of each node in the tree, we write down the ids in DFS order. This representation allows to move though the tree fast and access node ids quickly. Besides that, something that is really interesting with this combination is the ability to jump to a node with a given id, we just need to perform a select query over the sequence of tags (recall that select(c,i) over a sequence obtains the position where the i-th c occurs). This last operation allows for traversing the tree only looking at the nodes that have a given id.&lt;/li&gt;
&lt;li&gt;The self-index, representing every text in the XML tree, is based on the FM-Index. Every node that contains textual data introduces a text into the collection. All the elements are concatenated introducing delimiters and every text in the collection gets an id (in our representation the texts are at the leaves of the tree so it&#39;s easy to map back a forward). The id is given by the position of the delimiter in the BWT, and using a range searching data structure we can map to original leaf-id. This index allows for fast pattern matching among the text data present in the XML file.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;This an overview, you can get more details looking the original &lt;a rel=&quot;external&quot; href=&quot;http://www.cs.uwaterloo.ca/~fclaude/docs/icde10.pdf&quot;&gt;paper&lt;/a&gt;. The nice thing is that by using this structures we allow for a flexible navigation of the structure, and an automaton-based search procedure achieves great practical results over real XML files (more details look at the experimental results in the paper). This makes SXSI an interesting option as engine for representing and querying XML files. The idea of the automaton was formalized a couple of days ago, you can see the &lt;a rel=&quot;external&quot; href=&quot;http://arxiv.org/abs/1003.4353&quot;&gt;paper&lt;/a&gt; by Maneth and Nguyen.&lt;/p&gt;
&lt;p&gt;What&#39;s next? Well, we are not done, there are some operations we would like SXSI to support that aren&#39;t implemented yet. Another issue is that the structure is static. Making it dynamic in practice is quite tricky, as it usually is with succinct representations.&lt;/p&gt;
&lt;p&gt;At this moment we are working on a first release (cleaning code, adding documentation, etc), and we hope to have something soon :-).&lt;/p&gt;
</content>
        
    </entry>
    <entry xml:lang="en">
        <title>MFCS Talk</title>
        <published>2009-08-24T20:58:41+00:00</published>
        <updated>2009-08-24T20:58:41+00:00</updated>
        
        <author>
          <name>Unknown</name>
        </author>
        
        <link rel="alternate" type="text/html" href="https://recoded.cl/notes/mfcs-talk/"/>
        <id>https://recoded.cl/notes/mfcs-talk/</id>
        
        <content type="html" xml:base="https://recoded.cl/notes/mfcs-talk/">&lt;p&gt;As I said in the previous post the talk is ready and online. I&#39;ll use this space to tell a little bit about the result we achieved with &lt;a rel=&quot;external&quot; href=&quot;http://www.dcc.uchile.cl/~gnavarro/&quot;&gt;Gonzalo Navarro&lt;/a&gt; for the paper &lt;a rel=&quot;external&quot; href=&quot;http://www.cs.uwaterloo.ca/~fclaude/docs/mfcs09.pdf&quot;&gt;&quot;Self-Indexed Text Compression using Straight-Line Programs&quot;&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;The main problem we address is indexing straight-line programs, which are a special kind of free-context grammars that represent a language composed of one single text. This can be used as a compressed representation of a text, for example, LZ78 and Re-Pair map to an SLP without much effort. The bad new is that finding the minimum grammar that generates the text is NP-Complete, so approximations have to be used.&lt;/p&gt;
&lt;p&gt;The formal definition is: A &lt;em&gt;Straight-Line Program (SLP)&lt;/em&gt; $&lt;code&gt;\mathcal{G}=(X=\{X_1,\ldots,X_n\},\Sigma)&lt;/code&gt;$ is a grammar that defines a single finite sequence $&lt;code&gt;T[1,u]&lt;/code&gt;$, drawn from an alphabet $&lt;code&gt;\Sigma =[1,\sigma]&lt;/code&gt;$ of terminals. It has $&lt;code&gt;n&lt;/code&gt;$ rules, which must be of the following types:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;$&lt;code&gt;X_i \rightarrow \alpha&lt;/code&gt;$, where $&lt;code&gt;\alpha \in \Sigma&lt;/code&gt;$. It represents string $&lt;code&gt;\mathcal{F}(X_i)=\alpha&lt;/code&gt;$.&lt;/li&gt;
&lt;li&gt;$&lt;code&gt;X_i \rightarrow X_l X_r&lt;/code&gt;$, where $&lt;code&gt;l,r &amp;lt; i&lt;/code&gt;$. It represents string $&lt;code&gt;\mathcal{F}(X_i) = \mathcal{F}(X_l)\mathcal{F}(X_r)&lt;/code&gt;$.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;We call $&lt;code&gt;\mathcal{F}(X_i)&lt;/code&gt;$ the &lt;em&gt;phrase generated&lt;/em&gt; by nonterminal $&lt;code&gt;X_i&lt;/code&gt;$, and $&lt;code&gt;T=\mathcal{F}(X_n)&lt;/code&gt;$.&lt;/p&gt;
&lt;p&gt;Given a pattern P and a text represented as an SLP, we focus on the following problems:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Count the number of occurrences of P in the text generated by the SLP&lt;/li&gt;
&lt;li&gt;Locate the positions where the pattern occurs.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;In the past this problem was approached from an online-searching point of view, our goal was to provide an index that supports searching in sublinear time on the size of the SLP. For the indexing approach we work on representing the SLP as a labeled binary relation. We consider a binary relation $&lt;code&gt;\mathcal{R} \subset A\times B&lt;/code&gt;$ where $&lt;code&gt;A=[n_1]&lt;/code&gt;$ and $&lt;code&gt;B=[n_2]&lt;/code&gt;$, and a labeling function $&lt;code&gt;\mathcal{L}:A\times B \rightarrow L&lt;/code&gt;$, where $&lt;code&gt;L=[\ell]\cup \{ \perp\}&lt;/code&gt;$ is the set of labels and $&lt;code&gt;\perp&lt;/code&gt;$ represents no label/relation between the pairs. We want to answer the following queries:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Label of a given pair.&lt;/li&gt;
&lt;li&gt;Elements related to an element of B.&lt;/li&gt;
&lt;li&gt;Elements related to an element of A.&lt;/li&gt;
&lt;li&gt;Elements related trough a given label.&lt;/li&gt;
&lt;li&gt;Given to contiguous ranges in A and B respectively, find the pair that belong to the relation where $&lt;code&gt;a\in A&lt;/code&gt;$ is in the range for A and $&lt;code&gt;b\in B&lt;/code&gt;$ is in it&#39;s range too.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The first theorem we include shows that each of this queries can be answered paying close to $&lt;code&gt;O(\log n_2)&lt;/code&gt;$ time per element retrieved (element in the answer). Using that result we show how to modify an SLP so that the rules are sorted by lexicographical order and then represent this variant of the SLP in $&lt;code&gt;n\log n+o(n\log n)&lt;/code&gt;$ bits of space, and support queries as:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Access to rules.&lt;/li&gt;
&lt;li&gt;Reverse access to rules (given two symbols, which rules generates them in a given order).&lt;/li&gt;
&lt;li&gt;Rules using a given left/right symbol.&lt;/li&gt;
&lt;li&gt;Rules using a contiguous range of symbols.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;All the queries require $&lt;code&gt;O(\log n)&lt;/code&gt;$ time per datum delivered, as a direct consequence of our labeled binary relation representation. Finally, including some extra structures we get our final result (I&#39;m not going to go into the details on how to use the SLP representation, or where the different trade-offs come from):&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Theorem:&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Let $&lt;code&gt;T[1,u]&lt;/code&gt;$ be a text over an effective alphabet $&lt;code&gt;[\sigma]&lt;/code&gt;$ represented by an SLP of $&lt;code&gt;n&lt;/code&gt;$ rules and height $&lt;code&gt;h&lt;/code&gt;$. Then there exists a representation of $&lt;code&gt;T&lt;/code&gt;$ using $&lt;code&gt;n(\log u + 3\log n + O(\log\sigma + \log h) + o(\log n))&lt;/code&gt;$ bits, such that any substring $&lt;code&gt;T[l,r]&lt;/code&gt;$ can be extracted in time $&lt;code&gt;O((r-l+h)\log n)&lt;/code&gt;$, and the positions of $&lt;code&gt;occ&lt;/code&gt;$ occurrences of a pattern $&lt;code&gt;P[1,m]&lt;/code&gt;$ in $&lt;code&gt;T&lt;/code&gt;$ can be found in time $&lt;code&gt;O((m(m+h)+h\,occ)\log n)&lt;/code&gt;$. By removing the $&lt;code&gt;O(\log h)&lt;/code&gt;$ term in the space, search time raises to $&lt;code&gt;O((m^2+occ)h\log n)&lt;/code&gt;$. By further removing the $&lt;code&gt;O(\log\sigma)&lt;/code&gt;$ term in the space, search time raises to $&lt;code&gt;O((m(m+h)\log n+h\,occ)\log n)&lt;/code&gt;$. The existence problem is solved within the time corresponding to $&lt;code&gt;occ=0&lt;/code&gt;$.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;There are some open problems we were stuck at the end of the paper included in the article and we are now working on a journal version with some interesting extensions.&lt;/p&gt;
</content>
        
    </entry>
    <entry xml:lang="en">
        <title>Text Partitioning</title>
        <published>2009-06-27T04:34:34+00:00</published>
        <updated>2009-06-27T04:34:34+00:00</updated>
        
        <author>
          <name>Unknown</name>
        </author>
        
        <link rel="alternate" type="text/html" href="https://recoded.cl/notes/text-partitioning/"/>
        <id>https://recoded.cl/notes/text-partitioning/</id>
        
        <content type="html" xml:base="https://recoded.cl/notes/text-partitioning/">&lt;p&gt;Today I saw this article posted in arXiv:&lt;/p&gt;
&lt;p&gt;&lt;a rel=&quot;external&quot; href=&quot;http://arxiv.org/abs/0906.4692&quot;&gt;On optimally partitioning a text to improve its compression&lt;/a&gt;&lt;br /&gt;
Written by Paolo Ferragina, Igor Nitto and Rossano Venturini&lt;/p&gt;
&lt;p&gt;I found this article really interesting and nicely presented. The problem they focus on is:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Problem:&lt;/strong&gt; Given a compressor $&lt;code&gt;C&lt;/code&gt;$ and a text $&lt;code&gt;T&lt;/code&gt;$ of length $&lt;code&gt;n&lt;/code&gt;$, drawn from an alphabet $&lt;code&gt;\Sigma&lt;/code&gt;$ of size $&lt;code&gt;\sigma&lt;/code&gt;$. Find the optimal partition of $&lt;code&gt;T=T_1T_2\ldots T_k&lt;/code&gt;$ such that $&lt;code&gt;|C(T_1)C(T_2)\ldots C(T_k)|&lt;/code&gt;$ is minimized.&lt;/p&gt;
&lt;p&gt;This means that we want to cut the text in $&lt;code&gt;k&lt;/code&gt;$ pieces, with $&lt;code&gt;k&lt;/code&gt;$ unknown, such that applying the compressor $&lt;code&gt;C&lt;/code&gt;$ over each piece achieves the best compression possible over the text. This is ignoring possible permutations of the text, such as the Burrows-Wheeler Transform (BWT).&lt;/p&gt;
&lt;p&gt;A simple solution is to transform this problem into a shortest path problem, every position in the text is a node in the graph, and every node $&lt;code&gt;i&lt;/code&gt;$ is connected with nodes $&lt;code&gt;i+1,i+2,\ldots, n&lt;/code&gt;$. The cost of going from node $&lt;code&gt;i&lt;/code&gt;$ to node $&lt;code&gt;j &amp;lt; i&lt;/code&gt;$ is $&lt;code&gt;|C(T_{i,j})|&lt;/code&gt;$. It is easy to see that the best partition obtains a total size equal to the minimum path from $&lt;code&gt;1&lt;/code&gt;$ to $&lt;code&gt;n&lt;/code&gt;$. Here I include a figure of the graph (using IPE :-)).&lt;/p&gt;
&lt;figure&gt;
&lt;img src=&quot;text.png&quot; alt=&quot;Example shortest path for text partitioning&quot;&gt;
&lt;figcaption&gt;Example shortest path for text partitioning&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;The main problem is that just building this graph takes $&lt;code&gt;O(n^3)&lt;/code&gt;$ time. Assume that $&lt;code&gt;C&lt;/code&gt;$ takes linear time to compress a sequence, then building the graph takes:&lt;/p&gt;
&lt;pre class=&quot;giallo&quot; style=&quot;color-scheme: light dark; color: light-dark(#575279, #E0DEF4); background-color: light-dark(#FAF4ED, #232136);&quot; &gt;&lt;code data-lang=&quot;math&quot;&gt;&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;\sum_{i=1}^{n-1} \sum_{j=i+1}^n j-i =\sum_{i=1}^{n-1}\sum_{j=1}^{n-i} j = \sum_{i=1}^{n-1}\sum_{j=1}^{n-1} j&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;pre class=&quot;giallo&quot; style=&quot;color-scheme: light dark; color: light-dark(#575279, #E0DEF4); background-color: light-dark(#FAF4ED, #232136);&quot; &gt;&lt;code data-lang=&quot;math&quot;&gt;&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;= \sum_{i=1}^{n-1} O(n^2) = O(n^3)&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;In the paper they present an algorithm to approximate the problem, they require $&lt;code&gt;O(n\log_{\epsilon+1}n)&lt;/code&gt;$ time and achieve an $&lt;code&gt;(1+\epsilon)-&lt;/code&gt;$approximation. The main idea behind the approach is to approximate the graph in such a way that, by storing less edges, they can still approximate the cost of the minimum path. They show how to run the algorithm without building the approximated version of the graph, to keep the space consumption low. They also show how to estimate the size of the compression for 0-order and k-order compressors during the process.&lt;/p&gt;
</content>
        
    </entry>
</feed>
