KB Research

Research at the National Library of the Netherlands

Author: lottewilms (page 3 of 3)

Europeana Newspapers Refinement & Aggregation Workshop

The KB participates in the Europeana Newspapers project that has started in February 2012. The project will enrich 18 million pages of digitised newspapers with Optical Character Recognition (OCR), Optical Layout Recognition (OLR) and Named Entity Recognition (NER) from all over Europe and deliver them to Europeana. The project consortium consists of 18 partners from all over Europe: some will provide (technical) support, while other will provide their digitised newspapers. The KB has two roles: we will not only deliver 2 million of our newspaper pages to Europeana, but we will also enrich ours and the newspapers of other partners with NER.

Untitled

Europeana Newspapers Workshop in Belgrade

In the last months, the project has welcomed 11 new associated partners and to make sure they can benefit as much as possible from the experiences of the project partners the University Library of Belgrade and LIBER jointly organised a workshop on refinement and aggregation on 13 and 14 June. Here, the KB (Clemens Neudecker and I) presented the work that is currently being done to make sure that we will have Named Entities for several partners. To make sure that the work that is being done in the project also benefits our direct colleagues, we were joined by someone from our Digitisation department.

The workshop started with a warm welcome in Belgrade by the director of the library, Prof. Aleksandar Jerkov. After a short introduction into the project by the project leader Hans-Jörg Lieder from the State Library Berlin, Clemens Neudecker from the KB presented the refinement process of the project. All presentations will be shared on the project’s Slideshare account. The refinement of the newspapers has already started and is being done by the University of Innsbruck and the company CCS in Hamburg. However, it was still a big surprise when Hans-Jörg Lieder announced a present for the director of the University Library Belgrade; the first batch of their processed newspapers!

Giving a gift of 200,000 digitised and refined newspapers to our Belgrade hosts

Giving a gift of 200,000 digitised and refined newspapers to our Belgrade hosts

The day continued with an introduction into the importance of evaluation of OCR and OLR and a demonstration of the tools used for this by Stefan Pletschacher and Cristian Clausner from the University of Salford. This sparked some interesting discussions in the break-out sessions on methods of evaluation in the libraries digitising their collections. For example, do you tell your service provider what you will be checking when you receive a batch? You could argue that the service provider would then only fix what you check. On the other hand if that is what you need to reach your goal it would save a lot of time and rejected batches.

After a short getting-to-know-each-other session the whole workshop party moved to the Nikola Tesla Museum nearby where we were introduced to their newspaper clippings project. All newspaper clippings collected by Nikola Tesla are now being digitised for publication on the museum’s website. A nice tour through the museum followed with several demonstrations (don’t worry, no one was electrocuted) and the day was concluded with a dinner in the bohemian quarter.

Breakout groups at the Belgrade Workshop

The second day of the workshop was dedicated solely to refinement. I kicked off the day with the question ‘What is a named entity?’. This sounds easy, but can provide you with some dilemmas as well. For example, a dog’s name is a name, but do you want it to be tagged as a NE? And what do you do with a title such as Romeo and Juliet? Consistency is key in this and as long as you keep your goal in mind while training your software you should end up with the results you are looking for.

Claus Gravenhorst followed me with his presentation on OLR at CCS, by using docWorks, with which they will process 2 million pages. It was then again our turn with a hands-on session about the tools we’re using, which are also available on Github. The last session of the workshop was a collaboration between Claus Gravenhorst from CCS and Günter Mühlberger from the University of Innsbruck who gave us a nice insight into their tools and the considerations made when working with digitised newspapers. For example, how many categories would you need to tag every article?

Group photo from the Europeana Newspapers workshop in Belgrade

All in all, it was a very successful workshop and I hope that all participants enjoyed it as much as I have. I at least am happy to have spoken to so many interesting people with new experiences from other digitisation projects. There is still much to learn from each other and projects like Europeana Newspapers contribute towards a good exchange of knowledge between libraries to ensure our users get the best experience when browsing through the rich digital collections.

Digital Humanities at the National library

About two months ago the Journal for Library Associations published an issue completely about Digital Humanities in libraries. Enthusiastically I printed all the open access articles (I know, not very nature conscious of me…) and put them on my desk. As it often goes with papers on desks, they’ve been lying there ever since. This changed this morning as I took out the stack and started reading them. And I loved it! Article after article I took out my highlighter and marked sentences and paragraphs that sounded too familiar to me, working in a research library with an interest in Digital Humanities.

The KB has started to look at Digital Humanities (DH) as a topic not so long ago, but has been involved with DH related projects for quite some time, although there were not called DH at the time. Examples are the CATCH projects that started in 2004, but since the beginning of our digitising-days our material is used by a variety of people and institutions. However, the KB is special in a way when it comes to doing DH research. We are the national library of the Netherlands and are thus not connected to a specific university or research institution. This means that we do not employ our own researchers. We do have a Research department, but most of the people here do not dive into our content, but do research to ensure the public can do this.

catchplus

The continuation of CATCH, CATCHPlus, where prototypes are converted to reliable tools

Although the JLA I read only discusses university libraries, and their associated researchers, this does not mean that the articles from for example Miriam Posner or Bethany Nowviskie are not relevant for us. The, often mentioned, lack of flexibility that is apparently inherent to a library also exist here and the desire to only publish something once it is perfect is something I too can relate to. Working with digitised material is never perfect. The software is not perfect, so how can the outcomes be? Nonetheless, the KB chose to show these imperfections in our OCR by opening up the texts to the public, including all mistakes and an estimate of accuracy.

kbnewspaper

A KB newspaper article. The OCR quality of this article is estimated at 84,9% character accuracy.

Being a research institute, with a large digital corpus that we are more than happy to share, without our own researchers (apart from the occasional research fellow), the KB not only faces the challenges of the university libraries as mentioned by Miriam Posner in her article (i.e. inflexibility, lack of time, authority, and incentive, overcautionesness, etc.), but I believe another crucial element can be added to this list: No affiliated researchers. Until not so long ago, when a researcher wanted to use (a section) of our digitised sets he/she would find someone from within the library who could help them get it. There was no official route to obtain the data or one contact person for a specific set, so it could be possible that people left the KB with hard disks full of images or that they tracked one of our employees down at a conference and badgered them until they got an e-mail with instructions on how to harvest a collection. Luckily, this has changed with the creation of the Data Services team.

The Data Services team are the go-to guys when it comes to our digital sets. They have taken up the responsibilities of advertising our datasets on our (unfortunately only in Dutch) website, at events and conferences and on the Dutch Open Data community, such as Open Data Nederland. We hope these efforts will lead to interesting use of our data and perhaps even some enrichments that we might implement in the future (OCR correction anyone?). But how can we be sure that our data does indeed gets used and that we reach the people who might be interested? And how do we know if our methods is in fact what they are looking for?

photo-1

The KB at the CLIN2013 conference.

This issue is one that I would imagine is easier to solve when you can simply walk to the other side of the building, knock on some doors and talk to researchers of whom you know their interests, because they teach Data Mining at your university. Unfortunately, we are not in that position, apart from the people who have asked for our data and those that will come to our (currently a work in progress) KB Lab. Now that we have the  instructions to harvest our sets available on the website, less and less people will probably be doing this, leaving us more in the dark about what interesting things are happening with the digitised Early Dutch Books Online or the ANP radio bulletins.

So, how do we get and stay in touch with interested parties that might contribute to the enrichment of our collections? How can we be sure that what we are doing is in fact what researchers need? How much can and do we want to adapt our methods to fit the need of researchers? For example, do we want to offer all possible data formats if there is a demand for it or is that something that the scholars might be able to tackle themselves? (Solutions and ultimate answers of course always welcome in the comment section below!)

We are undertaking several activities to try to find answers and also our place in the wonderful world of Digital Humanities. The establishment of our own KB Lab, where we will to work with scholars who wish to do something with our data, is one such activity. Another is the poster session that we will present together with the BL Labs project at the DH2013 conference this summer. Our aim there is to talk to different researchers about our collections and their ways of working. What types of collections they would like, what data format they would love to see, but also what they would like to do with our data. So, if you’re around in Nebraska, please come and find me at the posters and let’s talk this through!

MOOCs in the Netherlands by Surf Academy

The SurfAcademy, a program set up to encourage knowledge exchange between higher education institutions in the Netherlands, organised a seminar on MOOCs, Massive Open Online Courses, on 26 February. Several Dutch institutions have started with MOOCs on various platforms and subjects, so the special interest group Open Educational Resources (OER) of Surf thought it was time to share experiences and open up the discussion for institutions that wish to jump on this fast moving train.

The Koninklijke Bibliotheek does not normally provide education as the National Library of the Netherlands, but we do work together with the Dutch universities (of applied sciences) and we are happy to share knowledge with our colleagues and users. Also, as one of the founding members of the impact Centre of Competence in text digitisation, we were asked to think about how we can best share the knowledge that was gathered in the 4 year research project IMPACT. Perhaps a MOOC would be a good idea?

The afternoon has an ambitious program, but is filled with experiences and interesting observations. I thought the most interesting parts of the afternoon were the presentations of the universities that are currently working with MOOCs in the Netherlands. Those were LeidenUniversity, presented by Marja Verstelle, the University of Amsterdam, presented by Frank Benneker and Willem van Valkenburg on the work the Technical University Delft is doing with their MOOC.

[slideshare id=16787746&w=427&h=356&sc=no]

It is interested to see the different choices each institution made for their own implementation of a MOOC. Leiden chose to work with Coursera and TU Delft joined EdX, while Amsterdam built their own platform (forever beta) in only two months and just 20k euro with a private partner. Each have their own reasons for these choices, such as flexibility (Amsterdam), openness (Delft) or ease (Leiden). Amsterdam is the only university that has started its MOOC already with great success (4800 participants in the first week), Leiden plans to start in May 2013 and Delft follows in September.

Another interesting presentation was the one by Timo Kos, both from KahnAcademy and Capgemini Consulting. He shared the results of two projects he did on OER, including MOOCs. As he showed us that MOOCs are not a technical hype, because they use no new technologies, merely combine existing ones for a new purpose. However, MOOCs can be indicated as a disruptive innovation, but as he says in the panel discussion at the end of the day we do not have to fear that real-life universities will be pushed out by MOOCs.

[slideshare id=16790022&w=427&h=356&sc=no]

All in all, I thought it was a very educative day with lots of food for thought. Most presentations are unfortunately in Dutch, but can be found on the website of the Surf Academy, where you will also find the videos made during the seminar. The English presentations have been embedded or linked to in this post.

Some of the questions and insights I took home with me:

  • Leiden and Amsterdam chose to create shorter videos for their MOOCs, while Delft will record regular classes. When do you choose which approach?
  • Do you want to use a platform of your own or will you sign up with one of the existing ones? (Examples: Coursera, EdX, Udacity, canvas.net)
  • Coursera takes 80-90% of the money made in a MOOC and they sell their user’s data to third parties. (Do have to say that I did not did a fact-check on this one!)
  • Do you want to get involved in the world of MOOCs as a non-top-50 university or even as a non-educational institute? The BL will do so, by joining FutureLearn.
  • PR of your MOOC is very important, especially if you use your own platform. However, getting a news item on the Dutch 8 o’clock news will probably mean one server is not enough for the first class.
  • The success of a MOOC also depends on the reputation of your institution.
  • Do students feel they are studying at an institute/university or at i.e. Coursera?
  • Using a MOOC towards your own degree is possible when you take the exam in/with a certified testing centre, such as Pearson or ProctorU.
  • If you plan to go into online education, when do you consider it a MOOC and when is it simply an online course?
Newer posts

© 2018 KB Research

Theme by Anders NorenUp ↑