KB Research Lab – KB Research http://blog.kbresearch.nl Research at the National Library of the Netherlands Fri, 24 Aug 2018 13:17:55 +0000 en-US hourly 1 https://wordpress.org/?v=4.4.2 Farewell; Work on Discerning Journalistic Styles continues! http://blog.kbresearch.nl/2017/01/17/farewell-work-on-discerning-journalistic-styles-continues/ http://blog.kbresearch.nl/2017/01/17/farewell-work-on-discerning-journalistic-styles-continues/#respond Tue, 17 Jan 2017 11:42:40 +0000 http://blog.kbresearch.nl/?p=2055 At the end of December our current researcher-in-residence dr. Frank Harbers of Groningen University ended his project ‘Discerning Journalistic Styles’. In this blogpost he describes the outcomes and plans for the future.

It is January 2017, meaning my period as researcher-in-residence at the KB has come to an end. It also means that my project Discerning Journalistic Styles (DJS) has come to an end. It was a really nice and valuable experience and a fruitful project in which we (I couldn’t have done it without the expertise of KB programmer Juliette Lonij) have managed to create a classification tool that automatically determines the genre of news articles. You can try the tool yourself at: http://www.kbresearch.nl/genre. Just paste a Dutch news article in the text box, press the button below and the result will appear on the right side; simple as that!

genre-classifier

Currently the tool predicts the correct genre in 65% of the cases. This might not seem that high at face value. However, we need to take into account 1) that genres are ideal types that never manifest themselves in their pure form and boundaries between different genres are fluid; 2) that genres are dynamic concepts that change over time, and 3) that genre is a typical example of a ‘latent content’ category, meaning that determining a genre involves a considerable amount of interpretation. It is therefore unsurprising that classifying genres manually is also difficult and that human coders also regularly disagree on what the correct genre of a text is. In fact, it is not unusual that 20 to 30% of the time, human coders disagree on what the right genre of a (historical) news article is. With that in mind, 65% is a solid result – which is not to say that it doesn’t need to be improved.

In that sense the research has only just begun. Not only because in the coming period we will keep concern ourselves with presenting the results on conferences and in academic articles, but also because we are developing research projects that follow-up on DJS. I therefore hope this won’t be the last time I visit The Hague to delve into the historical newspaper collection. If you are curious about the tool, please keep a close eye on the new Lab website of the KB that will be launched soon. On that website we aim to give much more detailed information on the tool and its background.

]]>
http://blog.kbresearch.nl/2017/01/17/farewell-work-on-discerning-journalistic-styles-continues/feed/ 0
Call for proposals KB Researcher-in-residence 2017 http://blog.kbresearch.nl/2016/06/08/call-for-proposals-kb-researcher-in-residence-2017/ http://blog.kbresearch.nl/2016/06/08/call-for-proposals-kb-researcher-in-residence-2017/#comments Wed, 08 Jun 2016 09:49:00 +0000 http://blog.kbresearch.nl/?p=1798 The Koninklijke Bibliotheek (KB), National Library of the Netherlands is seeking proposals for its Researcher-in-residence program to start in 2017. This program offers a chance to early career researchers to work in the library with the Digital Humanities team and KB data. In return, we learn how researchers use the data of the KB. Together we will address your research question in a 6 month project using the digital collections of the KB and computational techniques. The output of the project will be incorporated in the KB Research Lab and is ideally beneficial for a larger (scholarly) community.

The KB and digitisation

The Koninklijke Bibliotheek (KB), National Library of the Netherlands  is a research library with a broad collection in the fields of Dutch history, culture and society, and as a national library collects and stores all (digital) publications that appear in the Netherlands, as well as a part of the international publications about the Netherlands. The KB has planned to have digitised and OCRed its entire collection of books, periodicals and newspapers from 1470 onward by the year 2030. Already in 2016, about 15% of this enormous task was completed, either from the KB itself or via public-private partnerships as Google Books and ProQuest. Over 20 million book-, newspaper- and magazine papers are currently available via the search portal www.delpher.nl. The project will be carried out in the Research Department of the KB and there will be two consecutive placements in 2017.

Who are we looking for?

Early career researchers who are:

  • PhD-students that are in their final stages of their PhD project or researchers that have obtained their PhD between 2011 and 2016
  • Employed at a university or research institute in the EU,
  • Interested in using one (or more) of the digital collections of the KB,
  • Available for 0.5 fte over a period of 6 months (Jan – Jun 2017 or Jul – Dec 2017) and able to spend at least 1 day a week at the KB.

What can we offer you?

  • A secondment with the KB for 0,5 fte for a period of 6 months based on your current salary
  • Access to all data sets of the KB,
  • An office space,
  • Travel costs within the Netherlands,
  • Support from a programmer, collection and data specialists.

Which collections do we have?

You can use any digital collection of the KB and even combine it with an external collection, if copyright allows. Several of our digitised collections are described in more detail on our website, such as the parliamentary papers and the medieval illuminated manuscripts.

You can also browse through our collection of more than 1 million newspapers, magazines, radio bulletins and books on Delpher.nl.

What kind of projects are we looking for?

We’re open to all kinds of projects that use our data and benefit your research and other users of the KB and/or the KB Research Lab. The KB Research Department currently focuses on research projects that improve, enrich, connect and analyse our data by using techniques and methods from the domains of Information Retrieval (IR), Natural Language Processing (NLP) and Machine Learning (ML). We encourage you to define your project by:

  1. formulating a fundamental research question that stems from your field of expertise and that can be linked to the applied techniques at the KB Research Department,
  2. formulating a project that is different from the previous executed Researcher in Residence projects that can be found on our blog.

For more inspiration also take a look at the previously submitted proposals on our blog: here, here, here and here.

How do I apply?

Fill out this form before 31 August 2016 to submit your project, after having read carefully our terms and conditions. The form contains the following elements: details, project description (including research question, theoretical background and applied methods and techniques), outcomes, work plan, personal background, your availability in 2017 and a checkbox on our terms and conditions.

Before you start working on your proposal, we encourage you take a look at the form so you will be able to fill it out in the most efficient manner.

Don’t forget to read the terms and conditions of this call and agree to them.

All proposals will first be reviewed by an internal KB committee and then forwarded to an external committee of representative experts from several Dutch universities and institutions that consists of:

  • prof. dr. Franciska de Jong, Erasmus University Rotterdam & Clarin
  • prof. dr. Sally Wyatt, eHumanities & Maastricht University
  • prof. dr. Karina van Dalen-Oskam, Huygens ING & University of Amsterdam
  • prof. dr. Joris van Eijnatten, Utrecht University
  • prof. dr. Maarten de Rijke, University of Amsterdam
  • prof. dr. Marcel Broersma, University of Groningen
  • Prof. dr. Emiel Krahmer, University of Tilburg
  • Prof. dr. Hilde de Weerdt, University of Leiden
  • Prof. dr. Arjen de Vries, Radboud University

All entries will be judged on:

  • Originality and quality
  • Link with techniques and methods currently applied at the KB Research Department (Information Retrieval, Natural Language Processing and Machine Learning)
  • Feasibility (technically, legally and practically)
  • How the KB data will be showcased and used
  • Whether the end results are of use for a wider community

You will be notified of the outcome of this call in October 2016.

For answer to more questions, read our FAQ. Please also read the terms of this call and placement.

Respondents are strongly advised to contact dh@kb.nl in advance of proposal submission to discuss eligibility, project details, prerequisites, and KB support with the Digital Humanities team, consisting of Lotte Wilms, Steven Claeyssens, Martijn Kleppe, Juliette Lonij and Willem Jan Faber.

]]>
http://blog.kbresearch.nl/2016/06/08/call-for-proposals-kb-researcher-in-residence-2017/feed/ 2
FAQ Call for Proposals Researcher-in-Residence http://blog.kbresearch.nl/2016/06/08/faq-call-for-proposals-researcher-in-residence-2017/ http://blog.kbresearch.nl/2016/06/08/faq-call-for-proposals-researcher-in-residence-2017/#respond Wed, 08 Jun 2016 09:48:50 +0000 http://blog.kbresearch.nl/?p=1813 Updated 04 June 2018

I don’t live or work in the Netherlands. Can I apply? 
Probably! Contact us at dh@kb.nl and we’ll discuss your options.

I want to use my own dataset. Is that possible?
Sure! As long as you also use one of the datasets of the KB and it doesn’t limit the publication of the project end results.

I don’t know how to code, is that a problem?
Not at all. We have skilled programmers who can help you with your project or we will try to find a match for you if you prefer someone else. This would mean submitting as a team and will cut the budget in half. Reach out to us to discuss the options.

I don’t speak Dutch. Is your content still interesting to me?
That depends on your research question :) It might not be so appealing to linguists, but could offer an novel collection for computer scientists. Contact us to see which collections we have and we can discuss what might be the most interesting set for you.

Why will you publish my abstract?
We want to show others what types of proposals we have received to offer future researchers an insight into the selection process and to prevent them from entering a similar project.

Can I submit a project I’ve submitted previously (at another institution)?
We’d like you to submit an original idea. It can be one you have had lying around for some time, but we’d appreciate projects that haven’t been done before. Projects that have been previously entered into a similar program should be changed significantly before resubmitting.

Can I also work fulltime on my project for a period of 3 months?
We prefer you to work part-time so you can spend a total of 6 months with us. This also allows you to continue your research or teaching obligations at your university.

Will you be able to reimburse any housing or hotel costs?
Unfortunately, when you come from outside the Netherlands, we are not able to find and fund your housing or pay for your travel expenses to the KB. However, we do fund travel costs within the Netherlands allowing you to come to and work in the KB, the Hague wherever you are based in the Netherlands.

I want to use my own programmer, can I?
Yes, you can. We even encourage you to bring in extra people when you want to address a subject we’re not experts in (such as multimedia). However, the budget remains the same, so it will have to be split between you. We do ask that the whole team is available in the KB for at least one day a week. If you want to know whether we can help you or if you should bring someone in, please contact us at dh@kb.nl.

I don’t know if my idea is what you’re looking for. What can I do?
You are welcome to contact us at dh@kb.nl to discuss your ideas and the possibilities.

Can I submit more than one project?
Please focus your efforts on one great project.

Who will be judging the entries?
The entries will be judged by an internal committee and then forwarded to an external committee of representative experts from several Dutch universities and institutions.

What will you judge my project on?
We will judge the entries on criteria such as feasibility (technically, legally and practically), how the KB data will be showcased and used and whether the end results are of use for a wider community. Next to this, we will also look at the originality and quality of the proposal and the amount of support needed (and in this case, more is not necessarily worse!).

What happens if you submit a plagiarized project?
When we notice your your project is plagiarized, we will not consider your application for placement. You are responsible for the originality and authenticity of the project, but we will keep our eyes open.

What happens to any software I write for my project?
All software in the projects, whether you or we write it, will be made available on the KB Lab and Github page under an open source license.

What happens to the data I collect/produce in my project?
At the KB Lab we try to be as open as possible. All data produced in the programme is to be made available for research purposes, either through the KB Lab, KB Data Services or DANS, and where possible will receive a CC-license.

Can I publish any papers about the project?
Yes, we even encourage you to do so. If necessary, we’re happy to help.

]]>
http://blog.kbresearch.nl/2016/06/08/faq-call-for-proposals-researcher-in-residence-2017/feed/ 0
Terms and conditions of the KB Researcher-in-residence programme 2017 http://blog.kbresearch.nl/2016/06/08/terms-and-conditions-of-the-kb-researcher-in-residence-programme-2017/ http://blog.kbresearch.nl/2016/06/08/terms-and-conditions-of-the-kb-researcher-in-residence-programme-2017/#respond Wed, 08 Jun 2016 09:48:41 +0000 http://blog.kbresearch.nl/?p=1822 This programme as detailed at the KB-website (“Programme”) is operated by the Koninklijke Bibliotheek, National Library of the Netherlands (“KB”), Prins Willem-Alexanderhof 5 (2509 LK) Den Haag, The Netherlands.

  1. General
    1. If you enter the Programme you agree to abide by all of the following Terms and Conditions under which the Programme is run, and to be bound by them. We may disqualify you without prior notice if you are in breach of any of these Terms and Conditions.
    2. KB reserves the right to cancel the Programme at any stage if KB deems this necessary or circumstances arise that are outside of our control.
  2. Entries
    1. The Programme is open to all excluding employees or contractors of the KB or their direct family members.
    2. The Programme is only open to Ph.D.-students, or applicants who have obtained a Ph.D.-degree between 2011 and 2016, and who are aged 18 years or older.
    3. The Programme is only open to researchers employed by a university or a research institute within the EU.
    4. Entry is free.
    5. To enter the Programme, you must:
      1. Submit your proposal in accordance with the Call for proposals via the KB website; and
      2. Ensure that (i) your proposal is original (you can submit a proposal you have submitted previously as long as it has not been used in another programme or fellowship) and (ii) your proposal does not in any way infringe the copyright or other intellectual property rights, or other rights of any third party.
      3. Ensure that your proposal is aimed at the use of KB-data that benefits your research and other users of KB and/or the KB Research Lab.
      4. Tick the box to accept these Terms and Conditions when you submit your entry.
    6. All entries must be received by midnight on 8 November 2015. Any entries received after this will not be considered.
    7. The entries will be judged by an internal committee and then forwarded to an external committee of representatives experts from several Dutch universities and institutions.
    8. You may be contacted by email or telephone to answer further questions about your entry up to 15 December 2015. Two entries and two backup-entries will be chosen for the final stage. The best two entries will be contacted personally by email to assess whether those applicants are available for the secondment-period.
  3. Rights and permissions
    1. All entrants accept and agree that KB may publish the title and abstract of your proposal on the KB Research Blog at the KB’s sole discretion by January 2016. This blog is harvested for the Dutch web archive.
    2. In the event that you are offered a secondment:
      1. you agree to enter in a written agreement with KB and the university or research institute where you are employed on, among other things, intellectual property rights, and the obligations of the parties concerned including the obligations as set out in these Terms and Conditions.
      2. you agree to participate in any publicity planned by KB, if required;
      3. you will be expected to work on your project at least one day per week in residence at the KB in Den Haag, between 1 January 2017 and 30 June 2017 or 1 July 2017 and 31 December 2017, for 0.5 fte.
      4. you agree to present your research results to employees of the KB in the final month of your secondment.
      5. you will be expected to publish a blog about your project on the KB Research Blog.
      6. you agree to submit a Certificate of Conduct for Natural Persons (‘VOG NP’).
      7. you agree to mention the secondment KB offered you in any publication your secondment will give rise to.
    3. If any software is produced for your research, it will be developed on the principles of open source software and it will be made available on the KB Research Lab and Github website under a GPLv3
    4. KB offers you access to all data sets of the KB and support from a programmer, collection specialists and data specialists.
    5. During the secondment, KB offers you an office space and payment of all travel expenses within the Netherlands regarding the secondment.
    6. In case you do not live in the Netherlands, you are responsible to find and fund your own housing and pay for travel expenses to the KB in Den Haag, the Netherlands.
  4. Personal data
    1. Other than as expressly permitted by these Terms and Conditions, KB will only use your contact details for the purposes of administering this Programme, and will not publish them or provide them to anyone without your permission.
    2. You consent to KB holding and processing data relating to you for legal, administrative and management purposes.
    3. Any personal data relating to you will be used solely in accordance with the current data protection legislation in the Netherlands (Wet bescherming persoonsgegevens) and will not be disclosed to another party – except for the external committee of representatives experts from several Dutch universities and institutions – without your prior consent. Please see the Privacy statement KB for further details.
    4. Data relating to you will be retained by KB for a reasonable period after the closing date specified in clause 2.6 to assist in the administration of the Programme in a consistent manner and to deal with any queries on the Programme.
  5. Applicable Law
    1. These Terms and Conditions and any dispute or claim arising out of or in connection with them shall be governed by and construed in accordance with Dutch law.
    2. The parties will attempt in good faith to resolve any dispute or claim arising out of or relating to these Terms and Conditions promptly by negotiation. If the dispute cannot be resolved by negotiation, you hereby agree that the sole jurisdiction and venue for any actions that may arise in relation to the subject matter hereof shall be the Dutch Court in Den Haag, the Netherlands
]]>
http://blog.kbresearch.nl/2016/06/08/terms-and-conditions-of-the-kb-researcher-in-residence-programme-2017/feed/ 0
KB at DHBenelux 2016 http://blog.kbresearch.nl/2016/06/07/kb-at-dhbenelux-2016/ http://blog.kbresearch.nl/2016/06/07/kb-at-dhbenelux-2016/#respond Tue, 07 Jun 2016 07:27:07 +0000 http://blog.kbresearch.nl/?p=1774 This week, the annual DHBenelux conference will take place in Belval, Luxembourg. It will bring together practically all DH scholars from Belgium (BE), the Netherlands (NE) and Luxembourg (LUX). You can read the full program and all abstracts on the website. Two presentations are by members of our DH team (Steven Claeyssens & Martijn Kleppe) and one presentation is by our current researcher in residence (Puck Wildschut – Radboud University Nijmegen). Please find the first paragraphs of their abstracts below:

  • Puck Wildschut – Roles, relations and references: Towards a computation-based distant reading of narrative-semantic roles in large datasets in Dutch

Since the rise of Russian Formalism in the early 19th century, literary theorists have been interested in finding ways to detect actants (characters) in narratives. The recent rise of computational methods within the humanities offers new ways of tackling this issue. As researcher-in-residence at the National Library of the Netherlands (KB) I am currently involved in a research project1 that aims to develop a computation-based model for analyzing narrative-semantic roles in large datasets in Dutch. The premise of the project is that actantial roles are not only to be detected on the higher level of motive- and theme-building, but also at the linguistic level of semantic roles. Furthermore, the project aims to not only develop a tool for the detection of actantial roles, but also, and more importantly, for discovering the relationships between those roles as they are encoded in language. (…) The poster presentation will show the most up-to-date version of the tool we are developing at the KB and the preliminary results of its implementation. Special prominence will be given to how the issues mentioned are integrated in the tool’s development. The poster aims to show how literarylinguistic theory and computational practice encourage each other in the development process

Full abstract (in pdf) here.

  • Steven Claeyssens – The Ideal Corpus. Towards a Critique of Large Digital Libraries from a Digital Humanities Perspective

Large scale digitisation of historical paper publications enables analyses of vast amounts of digital surrogates using machines, algorithms and software to ‘read’ the texts. However, continuously expanding collections of such texts, like the Google Books corpora, HathiTrust or – closer to home – Delpher, combined with a proliferation of computational approaches to study textual data and the growing number of scholarly disciplines taking part in the Digital Turn, calls for a renewed reflection on two of the key aspects of the research process: source selection and source criticism.  (…) This paper argues that a digital source criticism is urgently needed to tackle the questions raised by these opposing expectations. Researchers and librarians should collaborate closely on this and join forces to define the limits and fits of ‘the ideal corpus’. Inspiration for this definition can, amongst other things, be found in textual criticism, corpus linguistics and analytical bibliography

Full abstract (in pdf) here.

  • Martijn Kleppe & Desmond Elliott – Doing Visual Big Data – Creating the KBK-1M Dataset Containing 1,6 Million Newspaper Images Available for Researchers

The visualisation of news through photographs has exploded since the second half of the 20th century (Kester & Kleppe 2015). However, methods that are employed to analyse the (re)use of visual materials are labour-intensive because Humanities researchers tend to analyse their sources manually (Burke 2001). To estimate the increase in the use of pressphotographs in Dutch newspapers, Kester & Kleppe (2015) e.g manually analysed a sample of 385 newspapers and 5.877 press photographs over the period 1870-2013. To find the recurring use of photographs in Dutch history textbooks, Kleppe (2012) followed a same approach by manually analysing over 5.000 photographs in 400 history textbooks, creating the ‘Foto’s in Nederlandse Geschiedenisschoolboeken (FiNGS) (Photos in Dutch History textbooks) dataset (Kleppe 2013b). Even though manually created and annotated datasets such as FiNGS contain rich & well-annotated data, their scope remains limited given its labour-intensive creation and analyses process. Therefor this poster presents the KBK-1M dataset, that was created specifically for (Digital) Humanities researchers. This dataset contains a collection of 1.603.395 captioned images extracted from Dutch digitised newspapers stored in the Dutch National Library (KB) Newspaper archive of the period 1922-1994. On our poster, we will describe how we obtained the images, what types of research questions it could tailor and how researchers can obtain the dataset for their research purposes.

Full abstract (in pdf) here. See poster below.KBK-1M Poster A1When you are at DHBenelux and if you would like to meet our colleagues, please feel free to approach them during the meeting or at one of the sessions Steven (Digital Textual Analysis II) and Martijn (Digital Art & Culture I) are chairing.

Finally, we are also very happy that during DHBenelux we will launch the call for our Researcher in Residence program 2017 that allows young Digital Humanities researchers to come and work with us in 2017. All details on the call are now online at http://blog.kbresearch.nl/2016/06/08/call-for-proposals-kb-researcher-in-residence-2017/ 

 

 

]]>
http://blog.kbresearch.nl/2016/06/07/kb-at-dhbenelux-2016/feed/ 0
Dataset KBK-1M containing 1.6 Million Newspaper Images available for researchers http://blog.kbresearch.nl/2016/05/26/dataset-kbk-1m-containing-1-6-million-newspaper-images-available-for-researchers/ http://blog.kbresearch.nl/2016/05/26/dataset-kbk-1m-containing-1-6-million-newspaper-images-available-for-researchers/#respond Thu, 26 May 2016 08:46:21 +0000 http://blog.kbresearch.nl/?p=1763 Each year the KB invites two academics to come and work with us as researchers in residence: early career researchers who work in the library with our Digital Humanities team and KB Data.  Together we address their research questions in a 6 month project using our digital collection and computational techniques. The output of the project will be incorporated in the KB Research Lab. Today we are happy to announce the output of the PhoCon project (‘Photos in and Out of Context’) by dr. Martijn Kleppe and dr. Desmond Elliott: the KBK-1M Dataset containing 1.6 Million Newspapers Images

During their residency, Kleppe & Elliott worked on ways to study the reuse of images in historical newspapers. Up until now, most scholars who worked with our historical newspapers focussed on their textual content. However, since we scanned the full lay-out of the pages we also have the images available. To find these images within delpher.nl you can e.g. filter the results of the newspapers via the facet ‘Illustratie met onderschrift’ (‘Illustration with caption’). To enable Kleppe and Elliott to deploy computer vision techniques to find recurring images in our newspapers, we worked hard with them to create a dedicated dataset that contain all images and captions as they are published in our newspapers.

Since we think this dataset can be of use for other types of research questions, we are happy to make this dataset also available to other researchers. The dataset is called ‘KBK-1M’ (‘Koninklijke Bibliotheek Kranten – 1 Miljoen’) and is a collection of 1.603.396 images and accompanying captions (in Dutch) of the period 1922-1994. It contains photographs (black&white and colour), comic strips, political cartoons and weather-forecasts.

The coming months, Kleppe and Elliot will present their dataset during the International Conference on Language Resources and Evaluation (LREC) (paper here), Digital Humanities Benelux (abstract in pdf here) and Digital Humanities 2016 (short abstract here). In their LREC paper they describe which use they foresee of this dataset in several domains. Humanities scholars can e.g. use it to analyse photographic style changes, the representation of people and societal issues and/or the creation of new tools for exploring photographic reuse via image-similarity-based search. Computer scientist can e.g. use the dataset for experiments in automatic image captioning, image-article matching, object recognition, and data-to-text generation for weather forecasting.

More information about the dataset can be found on its page in our Lab: http://lab.kbresearch.nl/static/html/KBK-1M.html

KBK-1MLabPic

We hope several researchers from different domains are interested in our dataset and we are happy to collaborate with you! If you are interest in using this dataset please send an e-mail with a request for access to dataservices@kb.nl. A representative of the KB will contact you and can provide you access to the dataset for scientific or scholarly purposes only after a contract has been signed.

If you are interested in other Researcher in Residence projects, please take a look at the page on the program on our blog or at the KB website. Please note that within a couple of weeks we will open the call for proposals for the new researchers in residence that will join us in 2017. If you would like to stay updated on the call, please keep an eye on our blog or send an email to express your interest to dh@kb.nl

]]>
http://blog.kbresearch.nl/2016/05/26/dataset-kbk-1m-containing-1-6-million-newspaper-images-available-for-researchers/feed/ 0
Call for proposals KB Researcher-in-residence http://blog.kbresearch.nl/2015/09/30/call-for-proposals-kb-researcher-in-residence/ http://blog.kbresearch.nl/2015/09/30/call-for-proposals-kb-researcher-in-residence/#comments Wed, 30 Sep 2015 11:49:03 +0000 http://blog.kbresearch.nl/?p=1487 The Koninklijke Bibliotheek (KB), National Library of the Netherlands is seeking proposals for its Researcher-in-residence program. This program offers a chance to early career researchers to work in the library with the Digital Humanities team and KB data. In return, we learn how researchers use the data of the KB. Together we will address your research question in a 6 month project using the digital collections of the KB and computational techniques. The output of the project will be incorporated in the KB Research Lab and is ideally beneficial for a larger (scholarly) community.

The KB and digitisation

The Koninklijke Bibliotheek (KB), National Library of the Netherlands  is a research library with a broad collection in the fields of Dutch history, culture and society, and as a national library collects and stores all (digital) publications that appear in the Netherlands, as well as a part of the international publications about the Netherlands. The KB has planned to have digitised and OCRed its entire collection of books, periodicals and newspapers from 1470 onward by the year 2030. Already in 2013, 10% of this enormous task was completed, resulting in 73 million digitised pages, either from the KB itself or via public-private partnerships as Google Books and ProQuest. Over 1 million books, newspapers and magazines are currently available via the search portal www.delpher.nl.

Researcher-in-residence

The project will be carried out in the Research Department of the KB and there will be two consecutive placements in 2016.

Who are we looking for?

Early career researchers who are:

  • PhD-students or have obtained their PhD between 2010 and 2015,
  • Employed at a university or research institute in the EU,
  • Interested in using one (or more) of the digital collections of the KB,
  • Available for 0.5 fte over a period of 6 months (Jan – Jun 2016 or Jul – Dec 2016) and able to spend at least 1 day a week at the KB.

What can we offer you?

  • A secondment with the KB,
  • Access to all data sets of the KB,
  • An office space,
  • Travel costs within the Netherlands,
  • Support from a programmer, collection and data specialists.

Which collections do we have?

You can use any digital collection of the KB and even combine it with an external collection, if copyright allows. Several of our digitised collections are described in more detail on our website, such as the parliamentary papers and the medieval illuminated manuscripts.

You can also browse through our collection of more than 1 million newspapers, magazines, radio bulletins and books on Delpher.nl.

What kind of projects are we looking for?

We’re open to all kinds of projects that use our data and benefit your research and other users of the KB and/or the KB Research Lab. Read our blog for more inspiration.

One of the previous Researchers-in-residence has worked on a best practice method for concept searching using keyword generation. Another team has worked on creating a data set that makes image similarity search a real possibility for all photos in our digitised historical newspapers.

For answer to more questions, read our FAQ. Please also read the terms of this call and placement. Respondents are urged to contact dh@kb.nl in advance of proposal submission to discuss eligibility, project details, prerequisites, and KB support with the Digital Humanities team.

How do I apply?

Fill out this form before 1 8 November 2015 to submit your project. Don’t forget to read the terms and conditions of this call and agree to them. You will be notified of the outcome in December.

]]>
http://blog.kbresearch.nl/2015/09/30/call-for-proposals-kb-researcher-in-residence/feed/ 4
FAQ KB Researcher-in-residence http://blog.kbresearch.nl/2015/09/30/faq-kb-researcher-in-residence/ http://blog.kbresearch.nl/2015/09/30/faq-kb-researcher-in-residence/#comments Wed, 30 Sep 2015 11:45:15 +0000 http://blog.kbresearch.nl/?p=1477 I don’t live or work in the Netherlands. Can I apply?
Probably! Contact us at dh@kb.nl and we’ll discuss your options.

I want to use my own dataset. Is that possible?
Sure! As long as you also use one of the sets of the KB and it doesn’t limit the publication of the project end results.

I don’t know how to code, is that a problem?
Not at all. We have skilled programmers who can help you with your project or we will try to find a match for you if you prefer someone else. This would mean submitting as a team and will cut the budget in half. Reach out to us to discuss the options.

I don’t speak Dutch. Is your content still interesting to me?
That depends on your research question :) It might not be so appealing to linguists, but could offer an novel collection for computer scientists. Contact us to see which collections we have and we can discuss what might be the most interesting set for you.

Why will you publish my proposal?
We want to show others what types of proposals we have received to offer future researchers an insight into the selection process and to prevent them from entering a similar project.

Can I submit a project I’ve submitted previously (at another institution)?
We’d like you to submit an original idea. It can be one you have had lying around for some time, but we’d appreciate projects that haven’t been done before. Projects that have been previously entered into a similar program should be changed significantly before resubmitting.

I want to use my own programmer, can I?
Yes, you can. We even encourage you to bring in extra people when you want to address a subject we’re not experts in (such as multimedia). However, the budget remains the same, so it will have to be split between you. We do ask that the whole team is available in the KB for at least one day a week. If you want to know whether we can help you or if you should bring someone in, please contact us at dh@kb.nl.

I don’t know if my idea is what you’re looking for. What can I do?
You are welcome to contact us at dh@kb.nl to discuss your ideas and the possibilities.

Can I submit more than one project?
Yes, you can submit as many as you like.

Who will be judging the entries?
The entries will be judged by an internal committee and then forwarded to an external committee of representative experts from several Dutch universities and institutions.

What will you judge my project on?
We will judge the entries on criteria such as feasibility (technically, legally and practically), how the KB data will be showcased and used and whether the end results are of use for a wider community. Next to this, we will also look at the originality and quality of the proposal and the amount of support needed (and in this case, more is not necessarily worse!).

What happens if you receive a plagiarized project?
That person or team will not be considered for placement. You are responsible for the originality and authenticity of the project, but we will keep our eyes open.

What happens to any software I write for my project?
All software in the projects, whether you or we write it, will be made available on the KB Research Lab and Github page under an open source license (at least GPLv3).

Can I publish any papers about the project?
Yes, of course. If necessary, we’re happy to help.

]]>
http://blog.kbresearch.nl/2015/09/30/faq-kb-researcher-in-residence/feed/ 1
Supporting History Research with Temporal Topic Previews at Querying Time http://blog.kbresearch.nl/2015/04/20/supporting-history-research-with-temporal-topic-previews-at-querying-time/ http://blog.kbresearch.nl/2015/04/20/supporting-history-research-with-temporal-topic-previews-at-querying-time/#comments Mon, 20 Apr 2015 13:24:05 +0000 https://researchkb.wordpress.com/?p=1208 This post is written by Dr. Jiyin He – Researcher-in-residence at the KB Research Lab from June – October 2014.

Being able to study primary sources is pivotal to the work of historians. Today’s mass digitisation of historical records such as books, newspapers, and pamphlets now provides researchers with the opportunity to study an unprecedented amount of material without the need for physical access to archives. Access to this material is provided through search systems, however, the effectiveness of such systems seems to lag behind the major web search engines. Some of the things that make web search engines so effective are redundancy of information, that popular material is often considered relevant material, and that the preferences of other users may be used to determine what you would find relevant. These properties do not hold or are unavailable for collections of historical material. In the past 3 months I have worked at the KB as a guest researcher. Together with Dr. Samuël Kruizinga, a historian, we explored how we can enhance the search system at KB to assist the search challenges of the historian. In this blogpost, I will share our experience of working together, the system we have developed, as well as lessons learnt during this project.

A Historian at Work

Samuël summizes the research approach of a historian as a 4-stage procedure.

(1) Exploring. In this stage, a researcher has an initial research idea. With this initial idea in mind, he explores the target domain and literature in order to arrive at a preliminary research question. At this stage, the researcher also starts to explore possible primary data sources that can be used.

(2) Contextualisation. Given the preliminary research question, the researcher conducts historiographical analyses and further explores the availability of data sources. By the end of this stage, the researcher arrives at a refined research question.

(3) Operationalisation. Given the refined research question, the researcher formulates his theory or model, and decides on the primary sources from which he will search for materials to answer his research question. Based on his theory or model, the researcher formulates a set of sub research questions. These sub research questions form the sufficient or necessary conditions to answer the original research question.

(4) Execution. In this stage, the researcher actually starts to execute search in the data sources that have been selected from the previous stage. For each sub research question, the following steps are taken:

  • Search each data source with source-specific queries. That is, these queries are specific with respect to the type of content, metadata, organisation, etc., of the source collection. Note that this is not necessarily done with a single search system, or even with digital tools.
  • By analysing the retrieved materials, the researcher attempts to answer the sub research questions, and evaluates its impact on the original research question.
  • If results of this stage are not satisfying,  the researcher goes back to stage (3).

The stages sketched above vary in the search style, information needs, and material sought for, hence different digital tools might be developed to best support the historian at each stage. For instance, in the first two stages, there is a need to explore, overview, and compare different data sources to assist the researcher to discover possible data sources, and to select the ones that may contain useful information. In later stages, the focus moves towards locating detailed information objects (e.g., articles, images), relevant to solve specific research questions. In this project, we focus on the latter, i.e., to provide support in exploring and access information in a particular data source, namely the historical Dutch newspapers.

The prototype system

One misconception about the role of digital tools in humanities research is that they should hide the complexity of the data selection or analysis techniques from the user. Samuel describes the role of digital tools as: supporting researchers to locate and gain access to potentially relevant materials, while allowing the researcher to select, digest, and interpret the materials. He stresses the risk of blindly following the results and analyses generated by digital tools and his request for our system’s functionality can be summarised as: control and transparency. That is, to have control over the ways to search and explore a collection, and simple but understandable operations are preferred over blackbox complex algorithms.

Data collection

The data used in our project included the KB historical collection during the first world war (WWI) period (1914 – 1940). As this material aligns with Samuël’s research interest of comparing collective memories of WWI as represented in the newspapers of different regional groups, e.g., “Landelijke” and “Nationale/Lokaal”.

Elasticsearch (ES) provides the basic framework for our retrieval system. News articles were retrieved from the KB data API, processed and then stored in an ES index. In total 55,639,628 articles were indexed. More details about the indexing process with ES are provided later in this post.

Query interface

On way of providing users with greater control over their search is to provide a richer querying language than keywords. In particular we decided to implement the following features: (1) Boolean query operation. This include specifying terms that “must occur”, “should occur but not necessarily occur”, or “must not occur” in the retrieved articles. (2) Wildcard queries and proximity queries. (3) Filtering on multiple time ranges and newspaper selections.

Note that features such as wildcard screenshot_jiyin1and proximity queries are already  supported by Lucene, the underlying search engine of KB’s Delpher system (as well as the basis of Elasticsearch). However, it is rather implicit, as users may not be aware of what is possible, and may not be able to use the query language defined by Lucene.  Here, we make the possible operations explicit, including explicit instructions on how wildcard and proximity queries should be constructed, as shown in the screenshot on the right.

Result operation

To provide greater control over the results displayed, we decided on two additional sorting options in addition to the default ranking criterion (i.e., by relevance scores), namely sorting by date, and by article length. Sorting by date provides a quick means to locate articles on specific dates. The argument for sorting by article length is that longer articles are likely to contain important content while news articles consisting of few lines are generally of less value.

screenshot_jiyin2

Interestingly, in many retrieval tasks (e.g., microblog search, ad-hoc search), document length and date have been combined with the relevance scores for the ranking of retrieved results. In our case, however, it was prefered that different ranking criteria are kept separated, and that the control of ranking criteria is transparent and flexible to the researcher.

Query preview

In Web search, query suggestion is a common means to assists users in issuing better keyword queries. In our case, users are allowed to issue queries with constraints such as time range and selected newspapers. That is, in an effort to support users in formulating their queries, suggestions of query words combined with suggestions of appropriate values of these constraints are needed. To this end, we implemented a temporal topic preview widget. That is, while typing a query, users can see named entities from the documents in the form of term clouds along a timeline. This widget is intended to help users in three ways:

  1. To determine the interesting time periods.
  2. To identify entities related to the original query which can be used for query reformulation.
  3. To compare topics discussed in different types of newspapers.

The methods to generate entity clouds for a specific year and for a given selection of newspapers are described next.

Entities.  For each entity cloud, we select the top 10 most significant entities from the articles within that period and from the selected newspapers.  Entities were extracted using KB’s named entity recognizer and prestored in the index.

Entity selection. To select the top 10 entities, we take the following steps.

  • We consider a foreground and a background document set. The foreground set consists of articles that contain the query words and are in the selected period and newspapers. The background set consists of all articles in the collection. Our goal is to select entities that are representative (e.g., frequently occur) in the foreground set in contrast to the background set.
  • We compute two types of conditional probabilities: the probability that the given entity is “generated” by the foreground set (p), and the probability that it is generated by the background set (q), using a language modeling approach. We then compute the Kullback-Leibler divergence between the two probability distributions KL(p||q).
  • Entities within the foreground document set are ranked in descending order of their KL divergence scores.

Preview updates. When the user types in query words or changes newspaper selections, the preview updates. To prevent updates on incomplete query input, we wait for 500ms after the user stops typing before updating.

An illustrative example

The following example is generated by typing in the query word “beurskrach”, referring to the economic crisis around 1930. In the screenshot we see two rows of entity clouds. The top row is generated from “Landelijke” newspapers, and the bottom row is generated for the “Nationale / Lokaal” newspapers.

screenshot_jiyin3

We have the following observations.

  1. This word starts to appear in news after 1929. This is correct, as the crisis starts in that year. The entity clouds in the previous years were absent as no articles containing this word were retrieved.
  2. We can see entities relevant to the crisis were selected, e.g., New York.
  3. If we compare the “Landelijke” newspapers to the “Nationale / Locaal” newspapers, we find that in local newspapers the crisis is hardly discussed.

UI wrap up

Finally, the resulting user interface looks as follows. While the user is formulating his or her query, the temporal topic overview is shown to assist this query formulation process (left). After the user has submitted the query, search results are shown, with possible result operations (right).

screenshot_jiyin4 screenshot_jiyin5

Lessons learnt

Finally, I would like to discuss some of the lessons learnt during this interdisciplinary project to design experimental search tools for historical research.

Support for an iterative design process.  In this project, I started with a standard search engine setup, i.e., to support keyword searching in document content. Later, after Samuel jointed the project, we started to add additional features to the system. This is an iterative process consisting of discussion – implementation – testing – new discussion. During this process new requested features kept emerging.

The updates of features fall into two categories: at the UI level and at the index level. Updates at the UI level are relatively simple: it adds additional access or means of interactions with existing (indexed) data. Updates at the indexing level enables additional data to be searchable, which is more complicated — often it means reindexing the collection.

The updates of index were necessary in two situations: metadata that existed but was not in the same collection (e.g., the page number of a news articles existed, but in a different versions of the KB news collection than the one previously indexed); derived data (e.g., article length — while it is possible to compute it at querying time, in order to allow efficient sorting of results document length was included at indexing time).

With respect to the design and development of tools for historical research my view is as follows:

  • From a system perspective, we need systems that allow flexible index updates. It was helpful that Elasticsearch allows adding additional fields without reindexing the data. In addition, when reindexing has to happen (e.g., when the data schema has changed), it can be done in the background without bringing down the whole system.
  • From a user perspective, it may be useful to provide supporting tools that allow the target users to explore the availability of data as well as possible derived data in the early stage of the design process.

Memory issues. It was the first time I used Elasticsearch. Before, I have always been using academic search systems such as Lemur and Terrier. I decided to experiment with Elasticsearch mainly because of its rich aggregation functions.

One of the issues I have been struggling with was the memory usage. Given the focus of the project as well as its short duration, I did not experiment with different configurations of ES, but simply used default configuration.  The indexed collection consists of 55,639,628 documents, resulting in an index of 260G. It seems that the memory usage can easily go over 10G, which is rather surprising. The machine we used has 30G RAM. While it was fine to perform simple keyword search, operations such as sorting on a specific field or more complexed queries can lead to out of memory exceptions. Unfortunately, I did not encounter this problem until the end of the project when all the data were indexed and all functionalities were implemented. The resulting system is therefore rather unstable.

Here are what I learnt with respect to the use of Elasticsearch: (1) It is not trivial to set the appropriate configuration for the elasticsearch system. Careful study and experiments are needed. (2) In order to properly configure the system, it is important to have an estimation of the size of the collection before hand. In my case, two factors make the estimation difficult. On the one hand, the documents were retrieved from a data service API, which were processed and indexed on the fly.  On the other hand, we kept updating the index with additional data fields with respect to newly emerged functionality requests throughout the project.

Outlook

In this project, we have discussed the research practice of historians and explored possible ways to support this process with novel search tool features. While a prototype system has been developed, much is left for further exploration, e.g., do the research practices as well as requirements for search systems found in this project generalise to that of other historians?

Dr. He’s tool is available for download on KB Research’s Github: https://github.com/KBNLresearch/spatio-temporal-topics  

]]>
http://blog.kbresearch.nl/2015/04/20/supporting-history-research-with-temporal-topic-previews-at-querying-time/feed/ 1
Researcher-in-residence at the KB http://blog.kbresearch.nl/2015/04/09/researcher-in-residence-at-the-kb/ http://blog.kbresearch.nl/2015/04/09/researcher-in-residence-at-the-kb/#comments Thu, 09 Apr 2015 11:58:45 +0000 https://researchkb.wordpress.com/?p=1217 At DH2013, we presented a poster to ask researchers what they need from a National Library. The responses varied from ‘Nothing, just give us your data’ to ‘We’d like to be fully supported with tools and services’, showing once again that different users have different requirements. In order to accommodate all groups of researchers, the Collections department of the KB, who ‘own’ the data, and the Research department, where tools and services are developed, combined efforts and spoke to scholars to  discuss the best method of supporting their work. However, we noticed that it was still quite difficult to get a good idea of how they used our data and in what way our actions and decisions would benefit them. Also, it seemed that researchers were often not aware of what activities the we undertake in this respect, which led to work being done twice.

In order to bridge this gap between our services and the needs of the scholars, a researcher-in-residence program was set up in the summer of 2014. In this program, we invite researchers into the library to work with our data, programmers and in our virtual research lab to learn how the scholars conduct their research, but also to benefit from their expertise to further our services. The program is aimed at young researchers in the early stages of their career (PhD or postdoc), in order to provide new opportunities for young scholars, and consists of a short research project (3-6 months) in the KB. The researcher is supported by programmers and experts from the KB. The results of the projects feed back into the KB, either by means of a post on this blog, or by deploying a prototype in the KB Research Lab, and are often accompanied with a publication.

The first three placements of the program were set up as pilot projects. For these projects, we invited three Dutch universities that currently use our data to join the program. This resulted in the selection of two historians, one media researcher and two computer scientists from the University of Amsterdam, Utrecht University, Erasmus University Rotterdam and CWI. The latest researcher is about to start his work at the KB and as we are very enthusiastic about the program, we have decided to start blogging on the work that is done by the researchers. We will evaluate the program in the summer of 2015 and hope to successfully evolve from pilot phase to full-blown residencies.

Our first blog post is written by Dr. Jiyin He, a computer scientist from the University of Amsterdam, who was a researcher-in-residence at the KB from July – October 2014, and will be posted next week.

The researchers that have joined the program so far are:

Do you want to know more about our efforts to work with researchers and this program? Come visit us at our poster presentation at DH2015 or come find the KB at the poster sessions of DHBenelux!

]]>
http://blog.kbresearch.nl/2015/04/09/researcher-in-residence-at-the-kb/feed/ 2