<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
	
	<title type="text" xml:lang="en">Applied Data Analysis Lab | News from our lab</title>
	<link type="application/atom+xml" href="http://adalab.icm.edu.pl/feeds/news-atom.xml" rel="self"/>
 	<link type="text" href="http://paulstamatiou.com" rel="alternate"/>
	<updated>2018-03-02T08:54:49+00:00</updated>
	<id>http://adalab.icm.edu.pl/blog/</id>
	<author>
		<name>ADA Lab, ICM UW</name>
	</author>
	<rights>CC-BY-SA 3.0</rights>
	
	
	
	<entry>
		<title>Sparkling-ferns for #ApacheSpark (Part 1: The Algorithm)</title>
		<link href="http://adalab.icm.edu.pl/blog/news/2015/11/17/spark-summit-eu-p1.html"/>
		<updated>2015-11-17T12:00:00+00:00</updated>
		<id>http://adalab.icm.edu.pl/blog/news/2015/11/17/spark-summit-eu-p1</id>
		<content type="html">&lt;p&gt;Two weeks ago together with Mateusz Fedoryszak I attended the first european Spark Summit (#SparkSummitEU). What did we find there and how did we enrich Spark Community? Let me tell you the story of the summit and Sparkling Ferns...&lt;/p&gt;

&lt;!-- more --&gt;

&lt;p&gt;First of all, yes, Matei Zaharia was excited as usual during his presentation. Second, yes, everyone had been super enthusiastic on every occasion during the Summit. Third, yes, you must attend Spark Summit if you haven’t done it yet. If you have - you know nothing else is even close to this experience. Anyway, we shall meet there next year!&lt;/p&gt;

&lt;p&gt;Many attendees asked us about Random Ferns and our implementation of it for Apache Spark: &lt;a href=&quot;https://github.com/CeON/sparkling-ferns&quot;&gt;Sparkling-ferns&lt;/a&gt;. Let’s go through Random Ferns FAQ. BTW: You may enjoy watching our talk:&lt;/p&gt;

&lt;div style=&quot;text-align: center&quot;&gt;&lt;iframe width=&quot;560&quot; height=&quot;315&quot; src=&quot;https://www.youtube.com/embed/333PVMcluvw&quot; frameborder=&quot;0&quot; allowfullscreen&gt;&lt;/iframe&gt;&lt;/div&gt;

&lt;p&gt;Our &lt;a href=&quot;http://www.slideshare.net/SparkSummit/sparkling-random-ferns-by-p-dendek-and-m-fedoryszak&quot;&gt;slides&lt;/a&gt; are publically available as well.&lt;/p&gt;

&lt;h2&gt;1. What are Random Ferns?&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Random Ferns&lt;/strong&gt; is a supervised learning classification algorithm.&lt;/p&gt;

&lt;p&gt;In classification algorithms items are described with a feature vector &lt;strong&gt;f&lt;/strong&gt;. We want to label each item using a class &lt;strong&gt;c&lt;/strong&gt; from the set of classes &lt;strong&gt;C&lt;/strong&gt;. To do so we create a model, which takes a vector &lt;strong&gt;f&lt;/strong&gt; and returns &lt;strong&gt;c&lt;/strong&gt;, the most suitable class for the vector &lt;strong&gt;f&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Now, there are many ways to do so. We may use Naive Bayes, Random Forests, SVM, etc. or Random Ferns. &lt;/p&gt;

&lt;h2&gt;2. When should I be interested with Random Ferns?&lt;/h2&gt;

&lt;p&gt;There are some natural indicators to use a specific algorithm.&lt;/p&gt;

&lt;p&gt;With Random Ferns these indicators are:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;requirement to train a model in linear time against the number of items&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;requirement to train a model in linear time against the number of features&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;plenty of memory to use (Random Ferns model can be quite big, I am going to explain it later)&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;3. I’ve heard about Random Forests - can you compare it with Random Ferns?&lt;/h2&gt;

&lt;p&gt;Sure!&lt;/p&gt;

&lt;h3&gt;3.1. Random Forests&lt;/h3&gt;

&lt;p&gt;With &lt;strong&gt;Random Forests&lt;/strong&gt; you have many decision trees. Each is trained on some subset of data. Let’s investigate one tree within the forest (see &lt;strong&gt;Fig.1.1.&lt;/strong&gt;). In each node one or more features of an item are tested to choose which child node should be chosen. When you arrive at a leaf node you obtain a class &lt;strong&gt;c&lt;/strong&gt; or list of probabilities for each class from &lt;strong&gt;C&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The important things to note are:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;one or more features&lt;/strong&gt; may be used to choose child node,&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;each node has its own test&lt;/strong&gt; - sibling nodes can use completely different features and/or different thresholds.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The final class &lt;strong&gt;c&lt;/strong&gt; assigned by a model is chosen e.g. by voting on results from particular trees.&lt;/p&gt;

&lt;p&gt;&lt;img style=&quot;margin-left: auto; margin-right: auto&quot; class=&quot;img-responsive&quot; src=&quot;/assets/2015-11-17-spark-summit-eu-p1/fig.1.1.png&quot;/&gt;
&lt;strong&gt;Fig.1.1.&lt;/strong&gt; The example of a &lt;strong&gt;decision tree&lt;/strong&gt; within a &lt;strong&gt;Random Forest&lt;/strong&gt; model.&lt;/p&gt;

&lt;h3&gt;3.2. Random Ferns&lt;/h3&gt;

&lt;p&gt;Now, when we use &lt;strong&gt;Random Ferns&lt;/strong&gt; two things are different:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;We are using perfect binary trees, where all levels of a tree are filled (they are balanced), and on each level only one feature is checked.&lt;/li&gt;
&lt;li&gt;Each fern has a threshold set for each feature. As a result on each level the decision if an item&amp;#39;s feature passes a threshold can be encoded with a binary value: 0 or 1.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;As you can see in &lt;strong&gt;Fig.1.2.&lt;/strong&gt; we have exactly 2^N leafs, where N is the number of features used. We also have C classes. What is more important is that we can enumerate all leafs using bits collected from the root to the chosen leaf. By going from the root to the bottom-left leaf in &lt;strong&gt;Fig.1.2.&lt;/strong&gt; we collect bits 000, which represents the integer number 0. By going to the bottom-right leaf we collect bits 111, which represents number 7.&lt;/p&gt;

&lt;p&gt;The natural next step is to switch from the tree representation to the 2D array representation,
where leaf code is the first coordinate and a class number is the second coordinate. A value obtained by passing a leaf code and a class number &lt;strong&gt;c&lt;/strong&gt; is probability of correctly labeling an item as a matching to a class &lt;strong&gt;c&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;img style=&quot;margin-left: auto; margin-right: auto&quot; class=&quot;img-responsive&quot; src=&quot;/assets/2015-11-17-spark-summit-eu-p1/fig.1.2.png&quot;/&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Fig.1.2.&lt;/strong&gt; The example of a &lt;strong&gt;fern&lt;/strong&gt; (represented as a tree) within a &lt;strong&gt;Random Ferns&lt;/strong&gt; model.&lt;/p&gt;

&lt;p&gt;&lt;img style=&quot;margin-left: auto; margin-right: auto&quot; class=&quot;img-responsive&quot; src=&quot;/assets/2015-11-17-spark-summit-eu-p1/fig.1.3.png&quot;/&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Fig.1.3.&lt;/strong&gt; The example of a &lt;strong&gt;fern&lt;/strong&gt; (represented as a 2D array) within a &lt;strong&gt;Random Ferns&lt;/strong&gt; model.&lt;/p&gt;

&lt;h2&gt;4. How big can be a Random Ferns model?&lt;/h2&gt;

&lt;p&gt;&lt;img style=&quot;margin-left: auto; margin-right: auto&quot; class=&quot;img-responsive&quot; src=&quot;/assets/2015-11-17-spark-summit-eu-p1/eq.1.png&quot;/&gt;&lt;/p&gt;

&lt;p&gt;Now let’s use some numbers.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Example 1&lt;/strong&gt;
&lt;img style=&quot;margin-left: auto; margin-right: auto&quot; class=&quot;img-responsive&quot; src=&quot;/assets/2015-11-17-spark-summit-eu-p1/eq.2B.png&quot;/&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Example 2&lt;/strong&gt;
&lt;img style=&quot;margin-left: auto; margin-right: auto&quot; class=&quot;img-responsive&quot; src=&quot;/assets/2015-11-17-spark-summit-eu-p1/eq.3B.png&quot;/&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Observation 1&lt;/strong&gt;: Adding 10 more features to a model makes it 1024 times bigger than the previous one.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Observation 2&lt;/strong&gt;: Doubling the number of ferns or classes make a new model 2 times bigger than the previous one.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;5. How fast can I train Random Ferns model?&lt;/h2&gt;

&lt;p&gt;During tests checking model creation time the following empirical dependency had been established: &lt;img style=&quot;margin-left: auto; margin-right: auto&quot; class=&quot;img-responsive&quot; src=&quot;/assets/2015-11-17-spark-summit-eu-p1/eq.4B.png&quot;/&gt;&lt;/p&gt;

&lt;p&gt;It is easier to think about this dependency in terms of &lt;img style=&quot;margin-left: auto; margin-right: auto&quot; class=&quot;img-responsive&quot; src=&quot;/assets/2015-11-17-spark-summit-eu-p1/eq.5B.png&quot;/&gt;&lt;/p&gt;

&lt;p&gt;So with a fixed number of features used &lt;strong&gt;f&lt;/strong&gt; a model creation time grows linearly with a growth of a dataset size &lt;strong&gt;D&lt;/strong&gt;.
Conversely, with a fixed size of a dataset &lt;strong&gt;D&lt;/strong&gt; a model creation time grows linearly with a growth of a number of features &lt;strong&gt;f&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;What does it mean? When you are training your model, you can easily predict when it will be ready. &lt;/p&gt;

&lt;h2&gt;6. How can I “plug” Random Ferns package into Spark?&lt;/h2&gt;

&lt;p&gt;Use the command:&lt;/p&gt;

&lt;p&gt;&lt;img style=&quot;margin-left: auto; margin-right: auto&quot; class=&quot;img-responsive&quot; src=&quot;/assets/2015-11-17-spark-summit-eu-p1/c.1.png&quot;/&gt;&lt;/p&gt;

&lt;p&gt;and enjoy Random Ferns right away!&lt;/p&gt;

&lt;h2&gt;7. How can I use Random Ferns code in Spark?&lt;/h2&gt;

&lt;p&gt;&lt;img style=&quot;margin-left: auto; margin-right: auto&quot; class=&quot;img-responsive&quot; src=&quot;/assets/2015-11-17-spark-summit-eu-p1/c.2.png&quot;/&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;You should import mllib Vectors and LabeledPoint as well as all classes from the package pl.edu.icm.sparkling_ferns.&lt;/li&gt;
&lt;li&gt;Next you should cast your data into instances of the LabeledPoint class. &lt;/li&gt;
&lt;li&gt;Finally, you should feed the method FernForest.train with LabeledPoint instances. Please pass along also numberOfFerns, numberOfFeatures and the mapping from feature ID to the number of possible values of a feature. 

&lt;ol&gt;
&lt;li&gt;Having two features, e.g. &amp;quot;doors&amp;quot; and &amp;quot;persons&amp;quot;, with possible values (respectively): List(&amp;quot;2&amp;quot;, &amp;quot;3&amp;quot;, &amp;quot;4&amp;quot;, &amp;quot;more&amp;quot;) and List(&amp;quot;2&amp;quot;, &amp;quot;4&amp;quot;, &amp;quot;more&amp;quot;), the map passed should be: Map(0 -&amp;gt; 4, 1 -&amp;gt; 3)&lt;/li&gt;
&lt;li&gt;If feature values are continuous you should pass an empty map (Map.empty).&lt;/li&gt;
&lt;/ol&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;8. Where I can find out more about Random Ferns?&lt;/h2&gt;

&lt;p&gt;For more details consult &lt;a href=&quot;https://github.com/CeON/sparkling-ferns&quot;&gt;GitHub&lt;/a&gt; or &lt;a href=&quot;http://spark-packages.org/package/CeON/sparkling-ferns&quot;&gt;Spark-packages.org&lt;/a&gt;.
The complete example of the package use is in the &lt;a href=&quot;https://github.com/CeON/sparkling-ferns/blob/master/src/test/scala/pl/edu/icm/sparkling_ferns/FernForestIntegrationSuite.scala&quot;&gt;test segment&lt;/a&gt; of a code.&lt;/p&gt;
</content>
	</entry>
	
	
	
	<entry>
		<title>PhD defense of ADA Laber Mateusz Kobos</title>
		<link href="http://adalab.icm.edu.pl/blog/news/2015/06/18/phd-defense-of-mateusz-kobos.html"/>
		<updated>2015-06-18T12:00:00+00:00</updated>
		<id>http://adalab.icm.edu.pl/blog/news/2015/06/18/phd-defense-of-mateusz-kobos</id>
		<content type="html">&lt;p&gt;On 2015-06-11, I defended my PhD thesis entitled &amp;quot;Multiresolution classification using combination of density estimators&amp;quot; in the &lt;a href=&quot;http://www.ibspan.waw.pl/glowna/en&quot;&gt;Systems Research Inistitute of the Polish Academy of Sciences&lt;/a&gt;.&lt;/p&gt;

&lt;!-- more --&gt;

&lt;p&gt;&lt;img style=&quot;margin-left: auto; margin-right: auto&quot; class=&quot;img-responsive&quot; src=&quot;/assets/2015-06-18-phd-defense-of-mateusz-kobos/kdes.png&quot;/&gt;&lt;/p&gt;

&lt;p&gt;In the thesis, we introduce a classification algorithm based on an idea of &amp;quot;multiple-resolution&amp;quot; (or &amp;quot;multiscale&amp;quot;) approach to data analysis. In practice, the method uses an average of kernel density estimators where each estimator corresponds to a different data &amp;quot;resolution&amp;quot; (see the figure above for estimation of density generated by &lt;a href=&quot;https://en.wikipedia.org/wiki/Kernel_density_estimation&quot;&gt;kernel density estimators&lt;/a&gt; with different smoothing parameters which can be interpreted as a multiple-resolution view of given five data points). First, we examine theoretical properties of this method; next, we propose a practical implementation of such algorithm with parameters of the density estimators and their number adjusted to minimize the misclassification probability. Subsequently, we test the algorithm on artificial data sets characterized by a multiple-resolution property. The tests show that the introduced algorithm is superior to the basic version based on one estimator per class. We also test the algorithm on benchmark data sets and compare the results obtained with the results of the basic version and other popular classification algorithms. The method is shown to fare better than the basic version and to be on a par with other popular algorithms.&lt;/p&gt;

&lt;p&gt;&lt;img style=&quot;margin-left: auto; margin-right: auto&quot; class=&quot;img-responsive&quot; src=&quot;/assets/2015-06-18-phd-defense-of-mateusz-kobos/error_functions.png&quot;/&gt;&lt;/p&gt;

&lt;p&gt;The image above is an intuitive data-based justification of why the proposed algorithm using many kernel density estimators is better for certain data sets than the basic version of the method using a single density estimator per class. The image shows the mean classification error computed for two benchmark data sets: BUPA liver disorders and Pima Indians diabetes. Two density estimators per class with the same value of smoothing parameter per class were used. The more more the color of the point resembles blue, the smaller the value of given function in this point. The global minima of the function lying outside of the diagonal are marked with triangles while the minimum for points lying on the diagonal is marked with a circle. The basic version of the method can achieve only the values lying on the diagonal; however, the proposed algorithm can achieve all shown values. The value next to the ∆ symbol above the plots is the difference between the minimal value of the function for points lying on the diagonal and the value of the global minimum. One can see that in case of the BUPA liver disorders dataset, the difference is nonzero thus we can expect a smaller classification error when applying the proposed algorithm. In case of the Pima Indians diabetes the difference is zero, so no gain is expected.&lt;/p&gt;

&lt;p&gt;The thesis is mainly an extension of paper &lt;a href=&quot;http://www.tandfonline.com/doi/abs/10.1080/09540091.2011.631166&quot;&gt;M. Kobos, J. Mańdziuk. Multiple-resolution classification with combination of density estimators. Connection Science, 23(4):219–237, 2011&lt;/a&gt;.&lt;/p&gt;
</content>
	</entry>
	
	
	
	<entry>
		<title>CERMINE wins award at ESWC 2015</title>
		<link href="http://adalab.icm.edu.pl/blog/news/2015/06/17/cermine-wins-award-at-eswc-2015.html"/>
		<updated>2015-06-17T09:00:00+00:00</updated>
		<id>http://adalab.icm.edu.pl/blog/news/2015/06/17/cermine-wins-award-at-eswc-2015</id>
		<content type="html">&lt;p&gt;ADA Lab&amp;#39;s &lt;a href=&quot;http://cermine.ceon.pl/&quot;&gt;CERMINE&lt;/a&gt;
participated in the &lt;a href=&quot;https://github.com/ceurws/lod/wiki/SemPub2015&quot;&gt;Semantic Publishing challenge&lt;/a&gt;
during the recent &lt;a href=&quot;http://2015.eswc-conferences.org/&quot;&gt;Extended Semantic Web Conference (ESWC 2015)&lt;/a&gt; in Portorož, Slovenia
and we won the Best Performing Approach Award!&lt;/p&gt;

&lt;!-- more --&gt;

&lt;p&gt;&lt;img style=&quot;margin-left: auto; margin-right: auto&quot; class=&quot;img-responsive&quot; src=&quot;/assets/awards/eswc2015.jpg&quot;/&gt;&lt;/p&gt;

&lt;p&gt;Extended Semantic Web Conference gathers researchers interested in various semantic technologies.
The week between May 31st and June 4th
was filled with workshops, tutorials, challenges, poster sessions, networking events and, of course, regular presentations.
Apart from our participation in the Semantic Publishing Challenge, I came to ESWC to learn state of the art
in knowledge representation (taking notes of ontologies and tools),
and to meet researchers interested in machine-friendly scholarly communication.&lt;/p&gt;

&lt;p&gt;Each of the three days of the main conference was kicked off by an excellent keynote:
Viktor Mayer-Schönberger spoke about Big Data,
Lise Getoor about Statistical Relational Learning
and Massimo Poesio about Games with a Purpose.
Two posters caught my attention: &lt;a href=&quot;http://aksw.org/Projects/GERBIL.html&quot;&gt;GERBIL&lt;/a&gt;, a system for benchmarking semantic annotations,
and &lt;a href=&quot;http://www.organicdatascience.org/&quot;&gt;ODSF&lt;/a&gt; for managing data-intensive scientific collaboration.&lt;/p&gt;

&lt;p&gt;The Semantic Publishing challenge, in which our CERMINE took place, gathered 9 teams, which worked on two tasks:
one for extracting information from HTML pages and one for mining scholarly PDFs.
The challenge was a good opportunity to meet some old friends (hello Christoph and Stefan!) and to make new ones (hello Angelo, Bahar, Francesco, Silvio and others!)
Also, thank you Mendeley and Springer for sponsoring the awards!&lt;/p&gt;

&lt;p&gt;&lt;img style=&quot;margin-left: auto; margin-right: auto&quot; class=&quot;img-responsive&quot; src=&quot;/assets/portoroz-short.jpg&quot;/&gt;&lt;/p&gt;

&lt;p&gt;This year the conference took place in the lovely Portorož, Slovenia — right by the Adriatic Sea.
It was the 12th edition of the event.
As a newcomer, I was enchanted by the friendly and relaxed atmosphere,
both pre-organized and spontaneous social events were a testimony that the community is well-integrated.
I&amp;#39;m looking forward to the next year&amp;#39;s edition!&lt;/p&gt;
</content>
	</entry>
	
	
	
	<entry>
		<title>Introducing ADA Lab Open Science APIs</title>
		<link href="http://adalab.icm.edu.pl/blog/news/2015/05/11/introducing-apis.html"/>
		<updated>2015-05-11T12:00:00+00:00</updated>
		<id>http://adalab.icm.edu.pl/blog/news/2015/05/11/introducing-apis</id>
		<content type="html">&lt;p&gt;Having our roots in the Centre for Open Science (&lt;a href=&quot;http://www.ceon.pl/en/&quot;&gt;CeON&lt;/a&gt;) we&amp;#39;re very keen on making sure anybody interested can take advantage of algorithms we design. Today we are making another step in that direction: we introduce &lt;a href=&quot;/api&quot;&gt;ADA Lab Open Science APIs&lt;/a&gt;.&lt;/p&gt;

&lt;!-- more --&gt;

&lt;p&gt;&lt;code&gt;APIs&lt;/code&gt; is the section of our website that will allow you to quickly see our technology in action. It contains demonstrators, each showcasing a small part of methods that we have designed. Although experimental for the time being, &lt;code&gt;RESTful&lt;/code&gt; &lt;code&gt;API&lt;/code&gt; is also provided so that you can use it in your apps.&lt;/p&gt;

&lt;p&gt;To begin with we provide two &lt;a href=&quot;http://cermine.ceon.pl/&quot;&gt;CERMINE&lt;/a&gt;-based demonstrators: citation and affiliation parsers. Stay tuned as we&amp;#39;ll regularly extend this section.&lt;/p&gt;
</content>
	</entry>
	
	
	
	<entry>
		<title>Text Mining Services in OpenAIRE</title>
		<link href="http://adalab.icm.edu.pl/blog/news/2015/02/16/openaire.html"/>
		<updated>2015-02-16T11:00:00+00:00</updated>
		<id>http://adalab.icm.edu.pl/blog/news/2015/02/16/openaire</id>
		<content type="html">&lt;p&gt;Recently in Athens there was an impressive kick-off of the OpenAIRE2020 project, during which we presented OpenAIRE’s plans in the area of text and data mining of scholarly publications. Publications contain all kinds of rich information, which, although understandable to a human reader, are not machine-readable and thus cannot be used directly for indexing and recommending purposes. Authors’ affiliations, document classifications, references to biological and chemical databases, acknowledgements to research funding agencies are all valuable pieces of information for OpenAIRE’s scholarly communication services.&lt;/p&gt;

&lt;!-- more --&gt;

&lt;p&gt;&lt;img src=&quot;/assets/openaire/text-analysis.jpg&quot; alt=&quot;Text analysis&quot;&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;(Photo credit: Wouter Vandenneucker, license: CC BY-SA 2.0)&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;Enriching content&lt;/h2&gt;

&lt;p&gt;At ADA Lab we are primarily interested in mining scholarly publications and extracting from them information that will help making OpenAIRE even more attractive to our users. On the one hand, knowledge which we’re going to extract will shed more light on the state of scholarly communication. Thanks to quantitative indicators, policy-makers and enthusiasts alike will be able to take the pulse of open science and observe the latest trends in research. On the second hand, researchers will have better means to find the research outputs they need, as the extracted knowledge will allow our indexing services to better “understand” the content and come up with more relevant search results and recommendations.&lt;/p&gt;

&lt;h2&gt;How will we do this?&lt;/h2&gt;

&lt;p&gt;Together with our colleagues from CNR in Pisa and ARC in Athens, in the work package devoted to Knowledge Extraction Services (WP10), we will improve and expand the Information Inference Service created in the OpenAIREplus project. We will make improvements to the inference infrastructure: add visual workflow management and improve quality assurance. We will extend existing document content analysis functionality to extract information about structure of the document, affiliation of the authors, and sentiment of the citations. We will also enhance our automatic document classification functionality and introduce functionality of creating clusters of similar documents. We will also search for new types of links to outside knowledge bases, i.e., 3rd party, domain-specific repositories describing genes, chemicals, organisms, etc. Some solutions will be built from scratch, other will be based on software developed by the partners, like &lt;a href=&quot;http://cermine.ceon.pl/&quot;&gt;CERMINE&lt;/a&gt; and &lt;a href=&quot;https://code.google.com/p/madis/&quot;&gt;MadIS&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;Code on Github&lt;/h2&gt;

&lt;p&gt;Finally, we will work on better uptake of the project’s deliverables by making our results even more discoverable and usable by the general public. To that end, we plan to migrate our code to GitHub (star &lt;a href=&quot;https://github.com/openaire/openaire-mining&quot;&gt;our repository&lt;/a&gt; now) and to publish our data sets on Zenodo. Both the source codes and the data sets will be available on open licenses, of course! First deliverables in our work package are scheduled for August 2015. We’ll keep you up-to-date about our research on this blog, so stay tuned!&lt;/p&gt;

&lt;p&gt;&lt;em&gt;This blog post has been simultaneously published on &lt;a href=&quot;http://blogs.openaire.eu/&quot;&gt;the official OpenAIRE blog&lt;/a&gt; and on the ADA Lab blog. It is available under the CC BY 4.0 license.&lt;/em&gt;&lt;/p&gt;
</content>
	</entry>
	
	
	
	<entry>
		<title>Let's join FORCEs and make a difference in scholarly communication</title>
		<link href="http://adalab.icm.edu.pl/blog/news/2015/02/02/force15.html"/>
		<updated>2015-02-02T09:00:00+00:00</updated>
		<id>http://adalab.icm.edu.pl/blog/news/2015/02/02/force15</id>
		<content type="html">&lt;p&gt;Two weeks ago I participated in &lt;a href=&quot;https://www.force11.org/meetings/force2015/&quot;&gt;FORCE2015&lt;/a&gt; in Oxford. It was a third conference organized by &lt;a href=&quot;https://www.force11.org/&quot;&gt;FORCE11 community&lt;/a&gt; and a must-attend event for people interested in scholarly communication, and in particular its problems and various ways of addressing them.&lt;/p&gt;

&lt;!-- more --&gt;

&lt;p&gt;One great thing about FORCE11 conferences is that they gather together people from a wide variety of backgrounds and professions: publishers, funders, librarians, researchers, programmers, and so on. This made FORCE2015 a great place to discuss various groups&amp;#39; needs and expectations, gain collaborators, advertise own work to potential consumers, exchange ideas, and provide and receive feedback about various initiatives.&lt;/p&gt;

&lt;p&gt;The day before the main conference I attended &lt;a href=&quot;http://contentmine.org/&quot;&gt;ContentMine&lt;/a&gt; workshop. ContentMine is a community-driven initiative aiming at extracting facts (eg. species, molecules, particles) from scientific literature and making them accessible and reusable. The project is still young and in the process of building the community, but definitely worth taking a closer look at.&lt;/p&gt;

&lt;p&gt;Another young but already interesting initiative I came across during the conference is &lt;a href=&quot;http://libraccess.org/&quot;&gt;Libraccess&lt;/a&gt; - a project which aims at collecting, aggregating, deduplicating and making available all kinds of open access scientific resources. Since Libraccess has a lot in common with our &lt;a href=&quot;http://comac.ceon.pl/&quot;&gt;COMAC&lt;/a&gt; project, we decided to join forces and use this great opportunity to achieve common goals collaboratively, making use of individual complementary strengths. There aren&amp;#39;t a lot of details yet, but stay tuned!&lt;/p&gt;

&lt;p&gt;During the conference I was also presenting a demo of &lt;a href=&quot;http://cermine.ceon.pl&quot;&gt;CERMINE&lt;/a&gt; - our Java library for extracting metadata and bibliography from scientific literature. Many thanks to all interested people, it was really great to meet you all!&lt;/p&gt;

&lt;p&gt;All the interesting presentations and discussions at FORCE2015 painted a clear picture of the current state of scolarly communication, its problems and efforts made to solve them. For me the most important (and very optimistic) issue is increasing understanding in the community that academic data is in fact not only text, and therefore simply putting paper publications into computers is not enough. Data sets, code and images should become first-class citizens - properly identified, shared and cited. So from one side more and more tools and platforms for managing scientific artifacts other than text emerge, and from the other - a lot of effort is dedicated to automatically process huge volume of already existing unstructured scientific text in order to reverse engineer the process of creating them, mine the knowledge burried in them and transform into machine-readable formats. The latter is exactly what we are passionate about in ADA Lab.&lt;/p&gt;

&lt;p&gt;FORCE2015 was a very interesting and unique experience for me. The event proved without a doubt that there are a lot of enthusiasts interested in the future of scolarly communication and the ways of improving it. Instead of attacking the same problems separately by individual people and teams, we should start organizing in larger groups and collaborate across teams, organizations and countries. If we manage to do so, the FORCE will definitely be with us!&lt;/p&gt;
</content>
	</entry>
	
	
	
	<entry>
		<title>Kraków – where AI meets the law</title>
		<link href="http://adalab.icm.edu.pl/blog/news/2014/12/22/krakow-where-ai-meets-the-law.html"/>
		<updated>2014-12-22T10:00:00+00:00</updated>
		<id>http://adalab.icm.edu.pl/blog/news/2014/12/22/krakow-where-ai-meets-the-law</id>
		<content type="html">&lt;p&gt;Recently, I was lucky enough to participate in the &lt;a href=&quot;http://conference.jurix.nl/2014/&quot;&gt;JURIX
2014&lt;/a&gt; conference, taking place in Kraków,
10-12 December 2014.  This was an event aimed at injecting the advancements of
computer science into the legal domain.  I must admit that the Organizers really
achieved their goal. At least from my strongly &lt;em&gt;computer-scientish&lt;/em&gt;
perspective...&lt;/p&gt;

&lt;p&gt;During the conference, I presented a proof-of-concept study on how to detect and
analyze topical trends in public procurement judgments. You can have a look at
my &lt;a href=&quot;http://dx.doi.org/10.6084/m9.figshare.1272026&quot;&gt;poster&lt;/a&gt;,
&lt;a href=&quot;http://arxiv.org/abs/1412.5212&quot;&gt;preprint&lt;/a&gt; or
&lt;a href=&quot;http://dx.doi.org/10.3233/978-1-61499-468-8-131&quot;&gt;paper&lt;/a&gt;.
Being able to present your work and gather feedback is great (by the way, I am very
grateful for all questions and comments during poster session). However,
listening to other talks is even fancier! Especially that JURIX 2014 provided
loads of interesting stuff for me...&lt;/p&gt;

&lt;p&gt;At the heart of each conference there are the invited talks.
JURIX was no exception. Both of them were stunning.&lt;/p&gt;

&lt;p&gt;On the first day, &lt;a href=&quot;http://researcher.watson.ibm.com/researcher/view.php?person=il-NOAMS&quot;&gt;Noam
Slonim&lt;/a&gt;
presented the research related to the &lt;a href=&quot;http://researcher.watson.ibm.com/researcher/view_group.php?id=5443&quot;&gt;IBM Debating Technologies
Project&lt;/a&gt;.
Following the previous endeavour, that is
&lt;a href=&quot;http://en.wikipedia.org/wiki/Watson_(computer)&quot;&gt;WATSON&lt;/a&gt;, IBM comes up with a
new challenge. WATSON was created, roughly speaking, to answer sophisticated
questions formulated  in the natural language. Now IBM wants to teach the
machine to search for claims pro or against a given topic, together with the
evidence supporting it. Typical topics could be banning violent video games or
permitting performance enhancing drugs in sports. To get the feeling, what is
it like to debate with the machine just spare 3 minutes to watch &lt;a href=&quot;https://www.youtube.com/watch?v=7g59PJxbGhY&quot;&gt;this
video&lt;/a&gt;.  If you are interested in
the science behind, read the very fresh papers of the IBM Debator group – &lt;a href=&quot;http://acl2014.org/acl2014/W14-21/W14-21-2014.pdf#page=76&quot;&gt;ACL
Argumentation Mining Workshop 2014
paper&lt;/a&gt; or &lt;a href=&quot;http://www.aclweb.org/anthology/C/C14/C14-1141.pdf&quot;&gt;COLING
2014 paper&lt;/a&gt;. For me the
most amazing thing is that this debating technology works on the basis of a
large body of raw text (e.g., Wikipedia). You basically make the computer read,
understand and find only the very relevant information for you. As Noam pointed
out, this is not another search engine, this is a &lt;em&gt;research engine&lt;/em&gt;!&lt;/p&gt;

&lt;p&gt;Second talk, despite very difficult task, was a great match to the first one.
&lt;a href=&quot;http://www.pieter-adriaans.com/&quot;&gt;Pieter Adriaans&lt;/a&gt; talked about measures of
information present in the data. This talk addressed very fundamental
questions, which, sadly, are not asked frequently enough in the age of the Big
Data fuss.  The roots of this subject date back to the giants – Shanon, Fisher
and Kolmogorov.  You can have a look at this very interesting
&lt;a href=&quot;http://arxiv.org/pdf/1203.2245v1.pdf&quot;&gt;paper&lt;/a&gt; full of insights and further
references.&lt;/p&gt;

&lt;p&gt;Except keynotes, there were a lot of interesting talks involving a large
variety of subjects such, as linked data, legal information interchange
standards/datasets, Bayesian networks, legal information retrieval systems,
computer aided analysis of legislation, etc. Browse the (unfortunately
pay-walled)
&lt;a href=&quot;http://ebooks.iospress.nl/volume/legal-knowledge-and-information-systems-jurix-2014-the-twenty-seventh-annual-conference&quot;&gt;proceedings&lt;/a&gt;,
if you are hungry for more.  The conference was accompanied by four workshops
and doctoral consortium.  The Organizers decided for parallel sessions
scenario. Therefore, it was impossible for see all the interesting stuff. For
me the definite highlights were the semantic workshop
&lt;a href=&quot;http://www.cs.unibo.it/sw4law2014&quot;&gt;SW4LAW&lt;/a&gt; and the network analysis
&lt;a href=&quot;http://www.leibnizcenter.org/%7Ewinkels/NAiL2014.html&quot;&gt;NAiL2014&lt;/a&gt; workshop.
Luckily proceedings from both are freely available on-line
&lt;a href=&quot;http://ceur-ws.org/Vol-1296/&quot;&gt;here&lt;/a&gt; and
&lt;a href=&quot;http://www.leibnizcenter.org/%7Ewinkels/NAiL2014-pre-proceedings.pdf&quot;&gt;there&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Altogether, JURIX 2014 was very fruitful conference for me. I have learnt a lot
about AI, law, and the intersection of both domains. Big &amp;quot;thank you&amp;quot; for
the Organizers! I hope to make it to Braga in 2015!&lt;/p&gt;
</content>
	</entry>
	
	
	
	<entry>
		<title>Spark, D3, data visualization and Super Cow Powers</title>
		<link href="http://adalab.icm.edu.pl/blog/news/2014/11/26/spark-d3-super-cow-powers.html"/>
		<updated>2014-11-26T15:00:00+00:00</updated>
		<id>http://adalab.icm.edu.pl/blog/news/2014/11/26/spark-d3-super-cow-powers</id>
		<content type="html">&lt;p&gt;Did you know that the amount of milk given by a cow depends on the number of days since its last calving? A plot of this correlation is called a lactation curve. Read on to find out how do we use &lt;a href=&quot;http://spark.apache.org&quot;&gt;Apache Spark&lt;/a&gt; and &lt;a href=&quot;http://d3js.org&quot;&gt;D3&lt;/a&gt; to find out how much milk we can expect on a particular day.&lt;/p&gt;

&lt;!-- more --&gt;

&lt;p&gt;&lt;img src=&quot;/assets/spark-d3-super-cow-powers/lactation-all-data.png&quot; alt=&quot;Milk yield per day.&quot;&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Fig. 1.&lt;/strong&gt; Milk yield per day.&lt;/p&gt;

&lt;h3&gt;Background — a lot of cows&lt;/h3&gt;

&lt;p&gt;Recently, ADA Lab has started a cooperation with the Polish Federation of Cattle Breeders and Dairy Farmers (&lt;a href=&quot;http://www.pfhb.pl/english/&quot;&gt;PFHBiPM&lt;/a&gt;). One of the goals of our project is as simple as that: predict how much milk will a particular cow produce on a particular day. It turns out PFHBiPM was doing &lt;em&gt;big data&lt;/em&gt; before it was cool: they have gathered 80M records of test milkings from 3M cows over two decades. This created a great opportunity for data analysis.&lt;/p&gt;

&lt;p&gt;While drilling through the data, we thought it would be interesting to visualize lactation curves passing certain points on a chart. Just drawing them wouldn&amp;#39;t tell us much: they were too numerous. Our goal then was to create an interactive 2D histogram of data points with respect to days after calving and the amount of obtained milk.&lt;/p&gt;

&lt;h3&gt;Our technology toolbox&lt;/h3&gt;

&lt;p&gt;We&amp;#39;ve decided to harness &lt;code&gt;Spark&lt;/code&gt; to choose interesting points and group them.
The first idea was to create an app with GUI which at some point would trigger computation on a cluster. However, we have come across a &lt;a href=&quot;http://spark-summit.org/2014/talk/spark-job-server-easy-spark-job-management&quot;&gt;Spark Summit talk&lt;/a&gt; about &lt;a href=&quot;https://github.com/spark-jobserver/spark-jobserver&quot;&gt;Spark Job Server&lt;/a&gt;. It&amp;#39;s a piece of software which allows you to talk with your cluster via &lt;code&gt;REST&lt;/code&gt;. It would be a shame not to make use of that!&lt;/p&gt;

&lt;p&gt;Other pieces were easy to fit: we&amp;#39;ve crafted a website which uses &lt;code&gt;AJAX&lt;/code&gt; to fire Spark jobs and uses D3 to present the results. Working with bleeding edge technologies previously taught us they tend to have sharp edges. Actually none of them were very serious: during compilation &lt;code&gt;Job Server&lt;/code&gt; didn&amp;#39;t pass all the tests (so we&amp;#39;ve turned them off...) and &lt;code&gt;AJAX&lt;/code&gt; refused to send requests to the remote domain (so we&amp;#39;ve hacked the &lt;code&gt;Job Server&lt;/code&gt; to include &lt;code&gt;Access-Control-Allow-Origin: *&lt;/code&gt; HTTP header).&lt;/p&gt;

&lt;p&gt;As of &lt;code&gt;D3&lt;/code&gt;, we haven&amp;#39;t used any off-the-shelf 2D histogram function. Instead we&amp;#39;ve used range of lower level &lt;code&gt;D3&lt;/code&gt; features: &lt;code&gt;AJAX&lt;/code&gt; requests handling, &lt;code&gt;SVG&lt;/code&gt; manipulation and chart axis drawing.&lt;/p&gt;

&lt;h3&gt;The results&lt;/h3&gt;

&lt;p&gt;&lt;img src=&quot;/assets/spark-d3-super-cow-powers/lactation-one-marker.png&quot; alt=&quot;Lactation curves passing a given point.&quot;&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Fig. 2.&lt;/strong&gt; Lactation curves passing a given point. &lt;/p&gt;

&lt;p&gt;Above are the results of our work. Once again we&amp;#39;ve experienced that a picture is worth a thousand words. &lt;code&gt;X-axis&lt;/code&gt; represents days after calving , while &lt;code&gt;Y-axis&lt;/code&gt; corresponds to the milk yield. The darker the rectangle, the more data points in it. Red line is arithmetic mean. We&amp;#39;ll keep you informed about our milky-project. Moo!&lt;/p&gt;
</content>
	</entry>
	
	
	
	<entry>
		<title>Affiliation parsing in CERMINE</title>
		<link href="http://adalab.icm.edu.pl/blog/news/2014/11/13/cermine-affiliations.html"/>
		<updated>2014-11-13T09:00:00+00:00</updated>
		<id>http://adalab.icm.edu.pl/blog/news/2014/11/13/cermine-affiliations</id>
		<content type="html">&lt;p&gt;&lt;a href=&quot;http://cermine.ceon.pl/&quot;&gt;CERMINE&lt;/a&gt; is our Java library for extracting metadata from scientific literature. Among other information, CERMINE extracts the authors of the input document, their affiliations, and also associates authors with affiliations. Recently new functionality has beed added: affiliation parsing.&lt;/p&gt;

&lt;!-- more --&gt;

&lt;p&gt;The goal of affiliation parsing is to recognize affiliation string fragments related to institution, address and country. Additionally, country names are decorated with their ISO codes. Here follows an example of a parsed affiliation string (it conforms to the &lt;a href=&quot;http://jats.nlm.nih.gov/archiving/&quot;&gt;JATS&lt;/a&gt; document description format):&lt;/p&gt;

&lt;pre&gt;
&amp;lt;aff id=&quot;id&quot;&amp;gt;
  &amp;lt;label&amp;gt;id&amp;lt;/label&amp;gt;
  &amp;lt;institution&amp;gt;Interdisciplinary Centre for Mathematical and Computational Modelling, University of Warsaw&amp;lt;/institution&amp;gt;,
  &amp;lt;addr-line&amp;gt;ul. Prosta 69, 00-838 Warsaw&amp;lt;/addr-line&amp;gt;,
  &amp;lt;country country=&quot;PL&quot;&amp;gt;Poland&amp;lt;/country&amp;gt;
&amp;lt;/aff&amp;gt;
&lt;/pre&gt;

&lt;p&gt;Affiliations are parsed with the use of Conditional Random Fields classifier. First the affiliation string is tokenized, then each token is classified as &lt;i&gt;institution&lt;/i&gt;, &lt;i&gt;address&lt;/i&gt;, &lt;i&gt;country&lt;/i&gt; or &lt;i&gt;other&lt;/i&gt;, and finally neighbouring tokens with the same label are concatenated. The main feature used in CRFs is the classified word itself. Additional features are all binary: whether the token is a number, whether it is all uppercase/lowercase word, whether it is a lowercase word that starts with an uppercase letter, whether the token is contained by dictionaries of countries or words commonly appearing in institutions or addresses. Additionally, the token&amp;#39;s feature vector contains not only features of the token itself, but also features of two preceding and two following tokens.&lt;/p&gt;

&lt;p&gt;Affiliation parser was evaluated by a 5-fold cross validation with the use of 8,000 affiliations from &lt;a href=&quot;ftp://ftp.ncbi.nlm.nih.gov/pub/pmc&quot;&gt;PubMed Central Open Access Subset&lt;/a&gt;. Labelled affiliation fragment (institution, address or country) was considered correct only if the entire string was identical to the ground truth. The following results were obtained:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;institution was correctly recognized in 92.4% of cases,&lt;/li&gt;
&lt;li&gt;address was correctly recognized in 92.3% of cases,&lt;/li&gt;
&lt;li&gt;country was correctly recognized in 99.5% of cases,&lt;/li&gt;
&lt;li&gt;92.1% of affiliations were entirely correctly parsed.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Affiliation parser can be used via REST service. It can be accessed using &lt;font face=&quot;Courier New,Courier,monospace&quot;&gt;cURL&lt;/font&gt; tool:&lt;/p&gt;

&lt;pre&gt;
$ curl -X POST --data &quot;affiliation=the text of the affiliation&quot; http://cermine.ceon.pl/parse.do
&lt;/pre&gt;

&lt;p&gt;For more information about the usage, visit &lt;a href=&quot;https://github.com/CeON/CERMINE&quot;&gt;CERMINE&amp;#39;s GitHub page&lt;/a&gt;.&lt;/p&gt;
</content>
	</entry>
	
	
	
	
	
	<entry>
		<title>Paperity chooses CERMINE as its content extraction engine</title>
		<link href="http://adalab.icm.edu.pl/blog/news/2014/10/23/paperity-chooses-cermine.html"/>
		<updated>2014-10-23T12:00:00+00:00</updated>
		<id>http://adalab.icm.edu.pl/blog/news/2014/10/23/paperity-chooses-cermine</id>
		<content type="html">&lt;p&gt;&lt;em&gt;This is a guest post by Selcuk Ayguney and Marcin Wojnarski, creators of Paperity.  We invited the authors to share their reasons for choosing ADA Lab&amp;#39;s (&lt;a href=&quot;/blog/news/2014/04/11/cermine-at-das.html&quot;&gt;recently awarded&lt;/a&gt;) CERMINE as their content extraction engine.  Here&amp;#39;s their story.&lt;/em&gt;&lt;/p&gt;

&lt;!-- more --&gt;

&lt;p&gt;Paperity (&lt;a href=&quot;http://www.paperity.org&quot;&gt;www.paperity.org&lt;/a&gt;) is the first multi-disciplinary aggregator of peer-reviewed Open Access journals and papers, &amp;quot;gold&amp;quot; and &amp;quot;hybrid&amp;quot;. It was launched in the beginning of October 2014 to facilitate access to scholarly literature across all different fields and already now includes nearly 200,000 articles from over 2,000 journals. While developing Paperity, we encountered the problem of extracting full text from PDF documents. Vast majority of academic papers are published as PDFs, and we wanted to unlock their contents and make them searchable in Paperity.&lt;/p&gt;

&lt;p&gt;Extracting text from a PDF document is one of the hardest practical problems that seem easy on the first sight. It should not be much different than using a word processor, right? Absolutely wrong. PDF format is designed for laying out pages and faithfully reproducing the same visual layout everywhere, be it a screen or a printer. Therefore, it does not consist of a continuous stream of letters, words, and sentences; but of pages and objects with specific sizes and coordinates relative to the page. This is a very low-level representation that must be thoroughly preprocessed before it can be analyzed as a complete text. Moreover, PDF authoring tools apply different typographical tricks while converting the text to PDF. For example, letters &amp;quot;f&amp;quot; and &amp;quot;i&amp;quot; are typically joined in a single &amp;quot;glyph&amp;quot;, to make them look better when printed, so that, for instance, the word &amp;quot;justification&amp;quot; in your word processor becomes &amp;quot;justiﬁcation&amp;quot; (note the single character that is a combination of &amp;quot;f&amp;quot; and &amp;quot;i&amp;quot;) when converted to PDF. This adds another level of complexity while extracting text. Split words at the end of lines pose another problem.&lt;/p&gt;

&lt;p&gt;We evaluated several toolkits designed for text extraction from PDF documents. In this article, we will share our findings and the rationale behind our final choice of CERMINE – the extraction tool developed by ADA Lab and CeON in ICM UW.
After reviewing many packages, we shortlisted the following three open source tools:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;a href=&quot;http://mstamy2.github.io/PyPDF2/&quot;&gt;PyPDF2&lt;/a&gt;: Written in Python, PyPDF2 is the successor of pyPDF and mainly focuses on document manipulation.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://pdfbox.apache.org/&quot;&gt;Apache PDFBox&lt;/a&gt;: Written in Java, it allows creation of new PDF documents, manipulation of existing documents and the ability to extract contents from documents.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;http://cermine.ceon.pl/&quot;&gt;CERMINE&lt;/a&gt;: Written in Java, designed specifically for analysis of scholarly articles; it is both a library and a web service for extracting metadata and content from scientific papers.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;We will use a sample Open Access paper (Sulfolobus chromatin proteins modulate strand displacement by DNA polymerase B1, &lt;em&gt;Nucleic Acids Research&lt;/em&gt;, 2013, Vol. 41, No. 17, accessible at &lt;a href=&quot;http://paperity.org/p/34709961&quot;&gt;http://paperity.org/p/34709961&lt;/a&gt;) to demonstrate some of the test cases we were concerned with. An image of the first page is included here for understanding the test results:&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/assets/paperity/test-pdf.png&quot; alt=&quot;Original&quot;&gt;&lt;/p&gt;

&lt;h2&gt;Content Selection Comparison&lt;/h2&gt;

&lt;p&gt;The first problem that arises during analysis of scholarly PDFs is how to correctly detect different blocks of text. A scholarly paper is not a continuous  stream of words. Rather, it contains many different sections that can be arranged in very different ways on the page and inside entire document. Each block plays a different role, therefore it is important to correctly detect all of them and discover what roles they play in the document. Only CERMINE was able to do this job. Below we give a preview of outputs of all the three tools.&lt;/p&gt;

&lt;p&gt;PDFMiner:&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/assets/paperity/test1-pdfminer.png&quot; alt=&quot;PDFMiner output&quot;&gt;&lt;/p&gt;

&lt;p&gt;PDFBox:&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/assets/paperity/test1-pdfbox.png&quot; alt=&quot;PDFBox output&quot;&gt;&lt;/p&gt;

&lt;p&gt;CERMINE:&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/assets/paperity/test1-cermine.png&quot; alt=&quot;CERMINE output&quot;&gt;&lt;/p&gt;

&lt;p&gt;(Please note that the red arrows denote line wraps.)&lt;/p&gt;

&lt;p&gt;At the beginning of the text extracted by PDFMiner and PDFBox, you can see meta data (title, journal name, authors) and article abstract, all combined into one stream of text. Moreover, paragraphs are most often broken into many separate lines in the output stream. 
CERMINE, on the other hand, can properly concatenate lines that belong to the same paragraph. It also detects the type of each block and starts extraction directly with the main body of the article. When necessary, CERMINE can also extract meta data fields as separate items in the output XML file.&lt;/p&gt;

&lt;p&gt;All the three tools did a good job eliminating the download link written vertically on the right hand side of the pages.&lt;/p&gt;

&lt;h2&gt;Paragraph Structure and Formatting Comparison&lt;/h2&gt;

&lt;p&gt;Below is a more detailed example of how paragraphs are processed by the evaluated tools.&lt;/p&gt;

&lt;p&gt;Original:&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/assets/paperity/test2-pdf.png&quot; alt=&quot;Original&quot;&gt;&lt;/p&gt;

&lt;p&gt;PDFMiner:&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/assets/paperity/test2-pdfminer.png&quot; alt=&quot;PDFMiner output&quot;&gt;&lt;/p&gt;

&lt;p&gt;PDFBox:&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/assets/paperity/test2-pdfbox.png&quot; alt=&quot;PDFBox output&quot;&gt;&lt;/p&gt;

&lt;p&gt;CERMINE:&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/assets/paperity/test2-cermine.png&quot; alt=&quot;CERMINE output&quot;&gt;&lt;/p&gt;

&lt;p&gt;We can see that both PDFMiner and PDFBox produce line breaks wherever the original PDF has one, while CERMINE successfully joins lines into paragraphs, retaining the document structure. In other experiments, we have also found that CERMINE can actually represent the paragraph structure with high accuracy.&lt;/p&gt;

&lt;p&gt;Moreover, only CERMINE was able to successfully merge split words: like “com-pacted” in the example above, merged to “compacted”. This is a very important feature, especially in documents with multi-column layout, where split words are very common.
While PDFMiner was unable to parse typographical shorthands (like joined f and i), both PDFBox and CERMINE were able to interpret them correctly.&lt;/p&gt;

&lt;h2&gt;Conclusions&lt;/h2&gt;

&lt;p&gt;All in all, we found CERMINE to be a very sophisticated and reliable tool for analysis of scholarly PDFs. It does an excellent job in detecting different types of blocks, preserving the structure of paragraphs and decoding characters correctly. What is also important, CERMINE is open source, which enables unconstrained use and guarantees that the tool will be easy to customize when necessary and will have a growing community of users and contributors. That is why we decided to use it in Paperity. &lt;/p&gt;

&lt;p&gt;We hope that CERMINE will be further developed in the future and one of possible directions that we might suggest is the recognition of national characters – a difficult task, given the peculiar encoding used in many PDFs, but very important for scholarly community and surely the one that can be successfully solved by CERMINE team.&lt;/p&gt;
</content>
	</entry>
	
	
	
	<entry>
		<title>Summer internship at ADA Lab</title>
		<link href="http://adalab.icm.edu.pl/blog/news/2014/10/06/summer-internship.html"/>
		<updated>2014-10-06T09:00:00+00:00</updated>
		<id>http://adalab.icm.edu.pl/blog/news/2014/10/06/summer-internship</id>
		<content type="html">&lt;p&gt;My name is Jan Lasek and I was an intern at ICM ADA Lab team in the summer time. And I need to say that it was a great experience to work here!&lt;/p&gt;

&lt;!-- more --&gt;

&lt;p&gt;I cooperated with Dominika Tkaczyk on implementing a new functionality to &lt;a href=&quot;http://cermine.ceon.pl/&quot;&gt;CERMINE&lt;/a&gt; project. Our goal was to extract table of contents from scholarly publications. This is, like many other problems in PDF processing, a challenging task. It is quite simple for a human, or even straightforward, however, the machines are still struggling with tasks of such type.&lt;/p&gt;

&lt;p&gt;To be more precise: a given PDF is a sequence of consecutive lines of text. The goal is to extract the header lines into structured table of content. We divided the main task into two parts:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;extracting outlying lines in a document that are presumably a header line and&lt;/li&gt;
&lt;li&gt;grouping the lines according to similar formatting to arrive with sections, subsections and possible subsubsection headers.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Our initial approach to the first task was to perform clustering on lines according to their different properties. We wanted to identify outliers, i.e. the lines that did not fit into main group of lines. This solution resulted in retrieving most of the headers, however, also many false positives, which were other lines looking &amp;quot;suspicious&amp;quot; (e.g. equations, part of the tables, the last line in a column). We decided to give a supervised classifier a try, which learns to select header lines. To this end, we labeled all lines in 150 documents as regular or header lines. For the task of classification we employed Random Forest algorithm as &amp;quot;the best off-the-shelf classifier&amp;quot;. This gave us promising results with about 0.95 of F-score statistic. &lt;/p&gt;

&lt;p&gt;As far as second task is concerned, we used clustering algorithm to choose the best division of the lines (previously identified as &amp;quot;headers&amp;quot; in the first step) into one, two or three groups. In initial experiments, when we applied the clustering algorithm to the &amp;quot;clear data&amp;quot; (that is, manually labeled lines) we were able to extract 3 out of 4 table of contents without any mistake. When we merged both steps of the project we obtained a solution that retrieves every second table of content without an error. There is still some work to be done to arrive at 9 out of 10 accuracy!&lt;/p&gt;

&lt;p&gt;Summing up, my internship at ICM UW was time well spent and I really appreciate that experience. I had an opportunity to work on an interesting and challenging task and cooperate with people passionate about their work. See you soon!&lt;/p&gt;
</content>
	</entry>
	
	
	
	<entry>
		<title>Impressions from PolTAL 2014</title>
		<link href="http://adalab.icm.edu.pl/blog/news/2014/09/30/impressions-from-poltal.html"/>
		<updated>2014-09-30T20:00:00+00:00</updated>
		<id>http://adalab.icm.edu.pl/blog/news/2014/09/30/impressions-from-poltal</id>
		<content type="html">&lt;p&gt;A couple of days ago, members of our lab participated in
&lt;a href=&quot;http://poltal.ipipan.waw.pl/&quot;&gt;PolTAL 2014&lt;/a&gt;, a conference bringing together
linguists, computer scientists, and other researchers involved 
in computational linguistics and natural language processing.&lt;/p&gt;

&lt;!-- more --&gt;

&lt;p&gt;After &lt;a href=&quot;http://icetal.ru.is/&quot;&gt;Island&lt;/a&gt;, 
&lt;a href=&quot;http://lang.cs.tut.ac.jp/japtal2012/&quot;&gt;Japan&lt;/a&gt; 
and numerous other distant places, TAL conference made it to Warsaw this year.
Therefore, the two of the ADALabers 
(Michał Jungiewicz and &lt;a href=&quot;/people/lopuszynski&quot;&gt;Michał Łopuszyński&lt;/a&gt;)
used this opportunity to present a poster on 
&lt;a href=&quot;http://poltal.ipipan.waw.pl/files/2614/1111/5047/P3.9_-_Unsupervised_keyword_extraction_from_Polish_legal_texts.pdf&quot;&gt;Unsupervised Keyword Extraction from Polish Legal Texts&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The conference was full of interesting and informative research reports.
All the &lt;a href=&quot;http://poltal.ipipan.waw.pl/index.php/programme&quot;&gt;materials&lt;/a&gt; were generously made
available by the participants with the encouragement and help from the
organisers. Springer published the LNCS volume with &lt;a href=&quot;http://link.springer.com/book/10.1007/978-3-319-10888-9&quot;&gt;PolTAL
Proceedings&lt;/a&gt;.
They promise to keep it available for free within a few  weeks 
after the closing of the conference. So do not hesitate to grab your copy!&lt;/p&gt;

&lt;p&gt;Just to wet your appetite before browsing the above materials, let us
mention some of our personal highlights.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;http://www.rug.nl/staff/johan.bos/&quot;&gt;Johan Bos&lt;/a&gt; presented an excellent
keynote on 
&lt;a href=&quot;http://poltal.ipipan.waw.pl/files/3514/1098/5369/I1_-_Adventures_in_Meaning_Banking.pdf&quot;&gt;Adventures in Meaning Banking&lt;/a&gt;, 
where he gave an overview of the experiences from building 
&lt;a href=&quot;http://gmb.let.rug.nl/&quot;&gt;Groningen Meaning Bank (GMB)&lt;/a&gt;.
GMB is free! Interestingly, not only you can 
&lt;a href=&quot;http://gmb.let.rug.nl/explorer/explore.php&quot;&gt;browse it&lt;/a&gt;,
&lt;a href=&quot;http://gmb.let.rug.nl/data.php&quot;&gt;download it&lt;/a&gt;, 
but you may also &lt;strong&gt;improve the GMB&lt;/strong&gt;. This last thing you do by
playing a game called &lt;a href=&quot;http://www.wordrobe.org&quot;&gt;wordrobe&lt;/a&gt; (sic!).
This is an excellent application example of
&lt;a href=&quot;http://en.wikipedia.org/wiki/Human-based_computation_game&quot;&gt;game with a purpose&lt;/a&gt;
– a recently popular strategy to attract volunteers to a project. &lt;/p&gt;

&lt;p&gt;Melanie Reiplinger presented a very tasty talk on &lt;a href=&quot;http://poltal.ipipan.waw.pl/files/1214/1112/0026/O6.1_-_Relation_Extraction_for_the_Food_Domain_without_Labeled_Training_Data_-_Is_Distant_Supervision_the_Best_Solution.pdf&quot;&gt;Relation Extraction for the
Food Domain&lt;/a&gt;.
This was mostly focused on relations &lt;em&gt;substitutedBy&lt;/em&gt; and &lt;em&gt;suitsTo&lt;/em&gt;. The
analysis was carried out on unlabelled data from the
&lt;a href=&quot;http://www.chefkoch.de/&quot;&gt;chefkoch.de&lt;/a&gt; community forum.  Melanie and her
colleagues even put up a 
&lt;a href=&quot;http://www.lsv.uni-saarland.de/personalPages/michael/relFood.html&quot;&gt;web page with data&lt;/a&gt;
together with a little 
&lt;a href=&quot;http://ws45lx.lsv.uni-saarland.de/startForm.html&quot;&gt;demo in German&lt;/a&gt; 
to let you play with their results...&lt;/p&gt;

&lt;p&gt;There were also many interesting posters, e.g., 
&lt;a href=&quot;http://poltal.ipipan.waw.pl/files/6414/1103/6933/P3.10__Evaluation_of_IR_Strategies_for_Polish.pdf&quot;&gt;Evaluation of IR Strategies for Polish&lt;/a&gt; (this is always interesting, if you work with search
engines for documents in Polish, like we do),
&lt;a href=&quot;http://poltal.ipipan.waw.pl/files/8114/1104/6749/P3.1_-_NER_in_tweets_using_a_small_crowdsourced_dataset.PDF&quot;&gt;Named Entity Recognition in Tweets&lt;/a&gt; 
(the form of micropost renders many traditional linguistic tools useless, 
so you have to develop new), or 
&lt;a href=&quot;http://poltal.ipipan.waw.pl/files/5814/1111/6909/P3.6_-_Slovak_Web_Discussion_Corpus.pdf&quot;&gt;Slovak Web Discussion Corpus and accompanying NLP tools&lt;/a&gt;
(in terms of the NLP resources, Slavic languages are unpopular and difficult to 
work with, so we appreciate the efforts of our colleagues from Košice).&lt;/p&gt;

&lt;p&gt;The above highlights are by no means exhaustive, we definitely 
encourage you to fish for your own favourites 
&lt;a href=&quot;http://poltal.ipipan.waw.pl/index.php/programme&quot;&gt;here&lt;/a&gt;
and 
&lt;a href=&quot;http://link.springer.com/book/10.1007/978-3-319-10888-9&quot;&gt;there&lt;/a&gt;!&lt;/p&gt;

&lt;p&gt;In the meanwhile, we countdown to xTAL 2016, wherever x turns out to be ... &lt;/p&gt;
</content>
	</entry>
	
	
	
	<entry>
		<title>Mind the gap! – DL2014</title>
		<link href="http://adalab.icm.edu.pl/blog/news/2014/09/23/ada-lab-at-dl2014.html"/>
		<updated>2014-09-23T09:30:00+00:00</updated>
		<id>http://adalab.icm.edu.pl/blog/news/2014/09/23/ada-lab-at-dl2014</id>
		<content type="html">&lt;p&gt;Recently a few people from our lab visited London to participate in the &lt;a href=&quot;http://www.dl2014.org/&quot;&gt;Digital Libraries 2014&lt;/a&gt; which was a conjunction of TPDL and JCDL – two best-known conferences on digital libraries. &lt;/p&gt;

&lt;!-- more --&gt;

&lt;p&gt;Łukasz was co-chairing the &lt;a href=&quot;http://lcpd2014.research-infrastructures.eu&quot;&gt;2nd Workshop on Linking and Contextualizing Publications and Datasets&lt;/a&gt; while Dominika and Mateusz were presenting their results on the &lt;a href=&quot;http://core-project.kmi.open.ac.uk/dl2014/&quot;&gt;3rd International Workshop on Mining Scientific Publications&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;LCPD2014 had a number of great presentations and — even more importantly — lively discussions about managing, sharing, peer-reviewing and citing research data sets.  There was a good mix of theory and practice, including an excellent case study of the CRAWDAD repository presented by Tristan Henderson, and a comprehensive list of &amp;quot;dos&amp;quot; and &amp;quot;don&amp;#39;ts&amp;quot; for data repositories by Sarah Callaghan, based on her peer review of Earth sciences data sets.&lt;/p&gt;

&lt;p&gt;WOSP2014 was an amazing opportunity to meet people from all over the world interested in scholarly communication, including the gurus like C. Lee Giles. Having discovered several very promising possible cooperation areas, we can&amp;#39;t wait to unveil more details soon. Meanwhile, here are our presentation slides:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Dominika Tkaczyk, Pawel Szostek and Łukasz Bolikowski: &lt;a href=&quot;http://www.slideshare.net/dtkaczyk/tkaczyk-grotoap2slides&quot;&gt;GROTOAP2 - The methodology of creating a large ground truth dataset of scientific articles&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Mateusz Fedoryszak and Łukasz Bolikowski: &lt;a href=&quot;http://www.slideshare.net/mfedoryszak/wosp-2014-39398147&quot;&gt;Efficient blocking method for a large scale citation matching&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Proceedings from both events will soon appear as special issues in the &lt;a href=&quot;http://www.dlib.org/&quot;&gt;D-Lib Magazine&lt;/a&gt;.&lt;/p&gt;
</content>
	</entry>
	
	
	
	
	
	
	
	
	
	
	
	
	
	
	
	<entry>
		<title>At an OpenAIREplus technical meeting in Pisa</title>
		<link href="http://adalab.icm.edu.pl/blog/news/2014/06/06/at-an-openaireplus-technical-meeting.html"/>
		<updated>2014-06-06T12:00:00+00:00</updated>
		<id>http://adalab.icm.edu.pl/blog/news/2014/06/06/at-an-openaireplus-technical-meeting</id>
		<content type="html">&lt;p&gt;Last week, three of us (Mateusz, Marek, Paweł) attended a technical meeting of the OpenAIREplus project in Pisa.&lt;/p&gt;

&lt;!-- more --&gt;

&lt;p&gt;Our lab takes part in technical development in the &lt;a href=&quot;http://www.openaire.eu/&quot;&gt;OpenAIREplus project&lt;/a&gt;. Thus, from time to time, all of the technical partners meet in a single place to discuss the current status and future plans. This time around, partners from four institututions were present: &lt;a href=&quot;http://www.isti.cnr.it/&quot;&gt;CNR&lt;/a&gt; (the host of the meeting), &lt;a href=&quot;http://en.uoa.gr/&quot;&gt;NKUA&lt;/a&gt;, &lt;a href=&quot;http://www.uni-bielefeld.de/&quot;&gt;UNIBI&lt;/a&gt;, and ICM (i.e. us).&lt;/p&gt;

&lt;p&gt;Here we are working hard:&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/assets/2014-06-06-at-an-openaireplus-technical-meeting/work.jpg&quot;&gt;&lt;/p&gt;

&lt;p&gt;and hardly working ;)&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/assets/2014-06-06-at-an-openaireplus-technical-meeting/dinner.jpg&quot;&gt;&lt;/p&gt;

&lt;p&gt;As a bonus, here's a compulsory photo of the Leaning Tower; this time from an unusual perspective, at night:&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/assets/2014-06-06-at-an-openaireplus-technical-meeting/tower.jpg&quot;&gt;&lt;/p&gt;
</content>
	</entry>
	
	
	
	<entry>
		<title>CERMINE wins Best Student Paper Award at DAS conference</title>
		<link href="http://adalab.icm.edu.pl/blog/news/2014/04/11/cermine-at-das.html"/>
		<updated>2014-04-11T12:00:00+00:00</updated>
		<id>http://adalab.icm.edu.pl/blog/news/2014/04/11/cermine-at-das</id>
		<content type="html">&lt;p&gt;&lt;a href=&quot;http://cermine.ceon.pl&quot;&gt;CERMINE&lt;/a&gt; system was presented yesterday at this year&amp;#39;s &lt;a href=&quot;http://das2014.sciencesconf.org/&quot;&gt;Document Analysis Systems conference&lt;/a&gt;. Our article entitled &amp;quot;CERMINE - automatic extraction of metadata and references from scientific literature&amp;quot; won &lt;a href=&quot;http://das2014.sciencesconf.org/resource/page/id/27&quot;&gt;ITESOFT Best Student Paper Award&lt;/a&gt;.&lt;/p&gt;

&lt;!-- more --&gt;

&lt;p&gt;The slides presented at DAS can be found &lt;a href=&quot;http://cermine.ceon.pl/static/docs/slides.pdf&quot;&gt;here&lt;/a&gt;.&lt;/p&gt;
</content>
	</entry>
	
	
	
	<entry>
		<title>On big trouble with big data at TOK FM</title>
		<link href="http://adalab.icm.edu.pl/blog/news/2014/04/10/on-big-trouble-with-big-data-at-tok-fm.html"/>
		<updated>2014-04-10T12:00:00+00:00</updated>
		<id>http://adalab.icm.edu.pl/blog/news/2014/04/10/on-big-trouble-with-big-data-at-tok-fm</id>
		<content type="html">&lt;p&gt;Yesterday at TOK FM (a popular Polish talk radio)
I discussed with Cezary Łasiczka about the recent article in FT.com by Tim Harford titled &amp;quot;&lt;a href=&quot;http://on.ft.com/P0PVBF&quot;&gt;Big data: are we making a big mistake?&lt;/a&gt;&amp;quot;.&lt;/p&gt;

&lt;!-- more --&gt;

&lt;p&gt;For me, the text is not so much a critique of &amp;quot;big data approaches&amp;quot; to modern problems
(although the hype associated with big data is, admittingly, unbearable)
as a warning that statistics are difficult and often counter-intuitive.
This stresses the need for good statistics education,
especially when talking about the skills required for &lt;a href=&quot;hbr.org/2012/10/data-scientist-the-sexiest-job-of-the-21st-century/&quot;&gt;the sexiest job of the 21st century&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Our conversation (in Polish) can be found in &lt;a href=&quot;http://audycje.tokfm.pl/audycja/67&quot;&gt;TOK FM&amp;#39;s archive&lt;/a&gt;.&lt;/p&gt;
</content>
	</entry>
	
	
	
	<entry>
		<title>Scoobi, Scalding, Spark, Stratosphere – ICM at Scalar 2014</title>
		<link href="http://adalab.icm.edu.pl/blog/news/2014/04/05/icm-at-scala-conference.html"/>
		<updated>2014-04-05T12:00:00+00:00</updated>
		<id>http://adalab.icm.edu.pl/blog/news/2014/04/05/icm-at-scala-conference</id>
		<content type="html">&lt;p&gt;We had a talk about Scala in ADA Lab at the &lt;a href='http://scalar-conf.com'&gt;Scalar 2014 conference&lt;/a&gt;.&lt;/p&gt;

&lt;!-- more --&gt;

&lt;p&gt;We&amp;#39;ve shown:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;how a word count in Scala should look like,&lt;/li&gt;
&lt;li&gt;how it looks in Scoobi and Scalding,&lt;/li&gt;
&lt;li&gt;a new hot concept in the Hadoop world: crunching data &lt;em&gt;in–memory&lt;/em&gt; on a cluster,&lt;/li&gt;
&lt;li&gt;Spark — framework featuring this concept.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Here are our slides:&lt;/p&gt;

&lt;iframe  src=&quot;http://www.slideshare.net/slideshow/embed_code/33266496&quot; width=&quot;427&quot; height=&quot;356&quot; frameborder=&quot;0&quot; marginwidth=&quot;0&quot; marginheight=&quot;0&quot; scrolling=&quot;no&quot; style=&quot;border:1px solid #CCC; border-width:1px 1px 0; margin: 10px 0; max-width: 100%;&quot; allowfullscreen&gt; &lt;/iframe&gt;
</content>
	</entry>
	
	
	
	<entry>
		<title>Perfect Data Analysis for Every Moment – ICM at Spotify 2014</title>
		<link href="http://adalab.icm.edu.pl/blog/news/2014/03/26/icm-at-spotify.html"/>
		<updated>2014-03-26T12:00:00+00:00</updated>
		<id>http://adalab.icm.edu.pl/blog/news/2014/03/26/icm-at-spotify</id>
		<content type="html">&lt;p&gt;Since monday we have started our one week in-house cooperation with &lt;a href='http://spotify.com'&gt;Spotify&lt;/a&gt; at its Stockholm HQ.&lt;/p&gt;

&lt;!-- more --&gt;

&lt;p&gt;We had meeting with people from virtually all IT divisions. It is always a pleasure to meet passionate people who believe in what they do and basically dream in code.&lt;/p&gt;

&lt;p&gt;On Monday evening we gave a talk (see below) about data analysis at ICM, Scoobie and Spark.&lt;/p&gt;

&lt;iframe src=&quot;http://prezi.com/embed/g7cxzsmxkt2g/?bgcolor=ffffff&amp;amp;lock_to_path=0&amp;amp;autoplay=0&amp;amp;autohide_ctrls=0&amp;amp;features=undefined&amp;amp;disabled_features=undefined&quot; width=&quot;427&quot; height=&quot;356&quot; frameBorder=&quot;0&quot; webkitAllowFullScreen mozAllowFullscreen allowfullscreen&gt;&lt;/iframe&gt;

&lt;p&gt;Since yesterday we&amp;#39;ve been cracking Spotify data (the privilage which is not given to everyone :-) )
It is no surprise that the data analysis tools, which we employ on the regular basis for scholarly communication is also useful for Spotify&amp;#39;s data.&lt;/p&gt;

&lt;p&gt;We hope that the cooperation will be fruitfull for both sides
and that using out heatmaps we can elevate the temperature in Stockholm even more.&lt;/p&gt;

&lt;p&gt;P.S. Sincere thanks for Adam Kawa, Boxun Zhang and many others, for inviting us, exchanging experiences and creating extraordinary environment. Spotify, you are the best!&lt;/p&gt;
</content>
	</entry>
	
	
	
	<entry>
		<title>Visit at ScraperWiki</title>
		<link href="http://adalab.icm.edu.pl/blog/news/2014/03/26/icm-at-scraperwiki.html"/>
		<updated>2014-03-26T12:00:00+00:00</updated>
		<id>http://adalab.icm.edu.pl/blog/news/2014/03/26/icm-at-scraperwiki</id>
		<content type="html">&lt;p&gt;Last week I spent in Liverpool visiting &lt;a href=&quot;https://scraperwiki.com/&quot;&gt;ScraperWiki&lt;/a&gt;. ScraperWiki provides tools for extracting, cleaning, analysing and managing data coming from various sources.&lt;/p&gt;

&lt;!-- more --&gt;

&lt;p&gt;I spent most of my time on conversations, getting familiar with the variety of ScraperWiki&amp;#39;s excellent extraction utilities that can deal with web pages, Twitter data, PDFs, and more. Since my work is partly related to processing PDF files, I was mostly interested in &lt;a href=&quot;https://blog.scraperwiki.com/2013/07/pdftables-a-python-library-for-getting-tables-out-of-pdf-files/&quot;&gt;pdftables&lt;/a&gt; - a library for extracting tabular data from PDFs. During a &lt;a href=&quot;http://en.wikipedia.org/wiki/Brown_bag_seminar&quot;&gt;brown bag session&lt;/a&gt; I also presented &lt;a href=&quot;http://cermine.ceon.pl&quot;&gt;CERMINE&lt;/a&gt; - ADA Lab&amp;#39;s system for extracting metadata and content from scientific literature (which currently barely touches tables in articles). The solutions have different purposes and scopes and they cannot be directly compared. Currently the only common part is parsing the PDF content, which is done by &lt;a href=&quot;http://poppler.freedesktop.org/&quot;&gt;Poppler&lt;/a&gt; in pdftables and by &lt;a href=&quot;http://itextpdf.com/&quot;&gt;iText&lt;/a&gt; in CERMINE.&lt;/p&gt;

&lt;p&gt;Another interesting difference is that pdftables (and I believe other ScraperWiki&amp;#39;s tools as well) strongly focuses on providing perfect results for specific, known cases, while CERMINE is a generic solution designed to deal with a wide variety of layouts. As a result, pdftables and other extraction tools will most likely need some manual work to adapt to documents and sources they&amp;#39;ve never seen before, but once this is done, they simply WORK. No excuses. CERMINE, on the other hand, will work out of the box for many different cases and document layouts, but the extraction results may not be perfect.&lt;/p&gt;

&lt;p&gt;Apart from the technical aspects, I was also very interested to see the differences between the &amp;quot;academic&amp;quot; and &amp;quot;commercial&amp;quot; work place (I am mostly familiar with the former). Unfortunately, my hopes to experience the work environment different from what I am used to have been shattered very quickly. It turned out the atmosphere in ScraperWiki is not that far from the academic world I know. I felt it even before I entered the building for the first time - ScraperWiki is located in the university campus among university buildings. Warm and welcoming atmosphere in the company&amp;#39;s room only added to this feeling. From day one, and in particular from the first stand-up meeting (held every morning), I was treated as a member of the family. The tea (no milk!) miraculously appeared on my desk every day, and I never even had to wash the tea mug (sorry, guys!).&lt;/p&gt;

&lt;p&gt;The visit was an awesome experience. I met a lot of wonderful people, smart and passionate about their work. It seems that no matter where you come from, you will always feel good among other &amp;quot;computer people&amp;quot;.&lt;/p&gt;
</content>
	</entry>
	
	
	
	
	
	<entry>
		<title>Mathematical modelling workshop for talented youth</title>
		<link href="http://adalab.icm.edu.pl/blog/news/2014/01/31/workshops.html"/>
		<updated>2014-01-31T12:00:00+00:00</updated>
		<id>http://adalab.icm.edu.pl/blog/news/2014/01/31/workshops</id>
		<content type="html">&lt;p&gt;Every year for almost 20 years, in collaboration with &lt;a href=&quot;http://fundusz.org&quot;&gt;Polish Childrens&amp;#39; Fund&lt;/a&gt;, ICM organizes weekly workshops for talented youth. This year&amp;#39;s edition has just finished.&lt;/p&gt;

&lt;!-- more --&gt;

&lt;p&gt;Polish Childrens&amp;#39; Fund (KFnrD) is a non-profit NGO which helps the most talented Polish children to grow their passion for art or science.
In case of high school students interested in mathematics and computer science, KFnrD organizes workshops at the best Polish research institutes.
ICM has had the privilege of collaborating with KFnrD and organining every year a weekly workshop on mathematical modelling.
Personally, I have been co-ordinating these workshops for the last couple of years (I have also participated as a student in one of the first workshops).&lt;/p&gt;

&lt;p&gt;Our idea for the workshops is to give the participants a taste of the research adventure.
In our work, we pose much more interesting questions and problems than we are able to solve.
Some of these problems can be attacked, with proper supervision, by the workshop participants.
Our guests usually don&amp;#39;t know all the &amp;quot;best&amp;quot; practices of solving particular problems,
so their approaches are fresh and often surprising!
We get an extra pair of hands and a fresh view on a problem, they — an interesting real-world challenge
and opportunity to work with experienced researchers. A clear win-win.&lt;/p&gt;

&lt;p&gt;This year, the participants formed three teams working on three problems from the following areas:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;image analysis,&lt;/li&gt;
&lt;li&gt;statistical machine learning,&lt;/li&gt;
&lt;li&gt;text mining / network analysis.&lt;/li&gt;
&lt;/ul&gt;
</content>
	</entry>
	
	
	
	
	
	
	
	
</feed>
