digital scholarship – KB Research http://blog.kbresearch.nl Research at the National Library of the Netherlands Fri, 24 Aug 2018 13:17:55 +0000 en-US hourly 1 https://wordpress.org/?v=4.4.2 Call for proposals KB Researcher-in-residence 2017 http://blog.kbresearch.nl/2016/06/08/call-for-proposals-kb-researcher-in-residence-2017/ http://blog.kbresearch.nl/2016/06/08/call-for-proposals-kb-researcher-in-residence-2017/#comments Wed, 08 Jun 2016 09:49:00 +0000 http://blog.kbresearch.nl/?p=1798 The Koninklijke Bibliotheek (KB), National Library of the Netherlands is seeking proposals for its Researcher-in-residence program to start in 2017. This program offers a chance to early career researchers to work in the library with the Digital Humanities team and KB data. In return, we learn how researchers use the data of the KB. Together we will address your research question in a 6 month project using the digital collections of the KB and computational techniques. The output of the project will be incorporated in the KB Research Lab and is ideally beneficial for a larger (scholarly) community.

The KB and digitisation

The Koninklijke Bibliotheek (KB), National Library of the Netherlands  is a research library with a broad collection in the fields of Dutch history, culture and society, and as a national library collects and stores all (digital) publications that appear in the Netherlands, as well as a part of the international publications about the Netherlands. The KB has planned to have digitised and OCRed its entire collection of books, periodicals and newspapers from 1470 onward by the year 2030. Already in 2016, about 15% of this enormous task was completed, either from the KB itself or via public-private partnerships as Google Books and ProQuest. Over 20 million book-, newspaper- and magazine papers are currently available via the search portal www.delpher.nl. The project will be carried out in the Research Department of the KB and there will be two consecutive placements in 2017.

Who are we looking for?

Early career researchers who are:

  • PhD-students that are in their final stages of their PhD project or researchers that have obtained their PhD between 2011 and 2016
  • Employed at a university or research institute in the EU,
  • Interested in using one (or more) of the digital collections of the KB,
  • Available for 0.5 fte over a period of 6 months (Jan – Jun 2017 or Jul – Dec 2017) and able to spend at least 1 day a week at the KB.

What can we offer you?

  • A secondment with the KB for 0,5 fte for a period of 6 months based on your current salary
  • Access to all data sets of the KB,
  • An office space,
  • Travel costs within the Netherlands,
  • Support from a programmer, collection and data specialists.

Which collections do we have?

You can use any digital collection of the KB and even combine it with an external collection, if copyright allows. Several of our digitised collections are described in more detail on our website, such as the parliamentary papers and the medieval illuminated manuscripts.

You can also browse through our collection of more than 1 million newspapers, magazines, radio bulletins and books on Delpher.nl.

What kind of projects are we looking for?

We’re open to all kinds of projects that use our data and benefit your research and other users of the KB and/or the KB Research Lab. The KB Research Department currently focuses on research projects that improve, enrich, connect and analyse our data by using techniques and methods from the domains of Information Retrieval (IR), Natural Language Processing (NLP) and Machine Learning (ML). We encourage you to define your project by:

  1. formulating a fundamental research question that stems from your field of expertise and that can be linked to the applied techniques at the KB Research Department,
  2. formulating a project that is different from the previous executed Researcher in Residence projects that can be found on our blog.

For more inspiration also take a look at the previously submitted proposals on our blog: here, here, here and here.

How do I apply?

Fill out this form before 31 August 2016 to submit your project, after having read carefully our terms and conditions. The form contains the following elements: details, project description (including research question, theoretical background and applied methods and techniques), outcomes, work plan, personal background, your availability in 2017 and a checkbox on our terms and conditions.

Before you start working on your proposal, we encourage you take a look at the form so you will be able to fill it out in the most efficient manner.

Don’t forget to read the terms and conditions of this call and agree to them.

All proposals will first be reviewed by an internal KB committee and then forwarded to an external committee of representative experts from several Dutch universities and institutions that consists of:

  • prof. dr. Franciska de Jong, Erasmus University Rotterdam & Clarin
  • prof. dr. Sally Wyatt, eHumanities & Maastricht University
  • prof. dr. Karina van Dalen-Oskam, Huygens ING & University of Amsterdam
  • prof. dr. Joris van Eijnatten, Utrecht University
  • prof. dr. Maarten de Rijke, University of Amsterdam
  • prof. dr. Marcel Broersma, University of Groningen
  • Prof. dr. Emiel Krahmer, University of Tilburg
  • Prof. dr. Hilde de Weerdt, University of Leiden
  • Prof. dr. Arjen de Vries, Radboud University

All entries will be judged on:

  • Originality and quality
  • Link with techniques and methods currently applied at the KB Research Department (Information Retrieval, Natural Language Processing and Machine Learning)
  • Feasibility (technically, legally and practically)
  • How the KB data will be showcased and used
  • Whether the end results are of use for a wider community

You will be notified of the outcome of this call in October 2016.

For answer to more questions, read our FAQ. Please also read the terms of this call and placement.

Respondents are strongly advised to contact dh@kb.nl in advance of proposal submission to discuss eligibility, project details, prerequisites, and KB support with the Digital Humanities team, consisting of Lotte Wilms, Steven Claeyssens, Martijn Kleppe, Juliette Lonij and Willem Jan Faber.

]]>
http://blog.kbresearch.nl/2016/06/08/call-for-proposals-kb-researcher-in-residence-2017/feed/ 2
FAQ Call for Proposals Researcher-in-Residence http://blog.kbresearch.nl/2016/06/08/faq-call-for-proposals-researcher-in-residence-2017/ http://blog.kbresearch.nl/2016/06/08/faq-call-for-proposals-researcher-in-residence-2017/#respond Wed, 08 Jun 2016 09:48:50 +0000 http://blog.kbresearch.nl/?p=1813 Updated 04 June 2018

I don’t live or work in the Netherlands. Can I apply? 
Probably! Contact us at dh@kb.nl and we’ll discuss your options.

I want to use my own dataset. Is that possible?
Sure! As long as you also use one of the datasets of the KB and it doesn’t limit the publication of the project end results.

I don’t know how to code, is that a problem?
Not at all. We have skilled programmers who can help you with your project or we will try to find a match for you if you prefer someone else. This would mean submitting as a team and will cut the budget in half. Reach out to us to discuss the options.

I don’t speak Dutch. Is your content still interesting to me?
That depends on your research question :) It might not be so appealing to linguists, but could offer an novel collection for computer scientists. Contact us to see which collections we have and we can discuss what might be the most interesting set for you.

Why will you publish my abstract?
We want to show others what types of proposals we have received to offer future researchers an insight into the selection process and to prevent them from entering a similar project.

Can I submit a project I’ve submitted previously (at another institution)?
We’d like you to submit an original idea. It can be one you have had lying around for some time, but we’d appreciate projects that haven’t been done before. Projects that have been previously entered into a similar program should be changed significantly before resubmitting.

Can I also work fulltime on my project for a period of 3 months?
We prefer you to work part-time so you can spend a total of 6 months with us. This also allows you to continue your research or teaching obligations at your university.

Will you be able to reimburse any housing or hotel costs?
Unfortunately, when you come from outside the Netherlands, we are not able to find and fund your housing or pay for your travel expenses to the KB. However, we do fund travel costs within the Netherlands allowing you to come to and work in the KB, the Hague wherever you are based in the Netherlands.

I want to use my own programmer, can I?
Yes, you can. We even encourage you to bring in extra people when you want to address a subject we’re not experts in (such as multimedia). However, the budget remains the same, so it will have to be split between you. We do ask that the whole team is available in the KB for at least one day a week. If you want to know whether we can help you or if you should bring someone in, please contact us at dh@kb.nl.

I don’t know if my idea is what you’re looking for. What can I do?
You are welcome to contact us at dh@kb.nl to discuss your ideas and the possibilities.

Can I submit more than one project?
Please focus your efforts on one great project.

Who will be judging the entries?
The entries will be judged by an internal committee and then forwarded to an external committee of representative experts from several Dutch universities and institutions.

What will you judge my project on?
We will judge the entries on criteria such as feasibility (technically, legally and practically), how the KB data will be showcased and used and whether the end results are of use for a wider community. Next to this, we will also look at the originality and quality of the proposal and the amount of support needed (and in this case, more is not necessarily worse!).

What happens if you submit a plagiarized project?
When we notice your your project is plagiarized, we will not consider your application for placement. You are responsible for the originality and authenticity of the project, but we will keep our eyes open.

What happens to any software I write for my project?
All software in the projects, whether you or we write it, will be made available on the KB Lab and Github page under an open source license.

What happens to the data I collect/produce in my project?
At the KB Lab we try to be as open as possible. All data produced in the programme is to be made available for research purposes, either through the KB Lab, KB Data Services or DANS, and where possible will receive a CC-license.

Can I publish any papers about the project?
Yes, we even encourage you to do so. If necessary, we’re happy to help.

]]>
http://blog.kbresearch.nl/2016/06/08/faq-call-for-proposals-researcher-in-residence-2017/feed/ 0
Terms and conditions of the KB Researcher-in-residence programme 2017 http://blog.kbresearch.nl/2016/06/08/terms-and-conditions-of-the-kb-researcher-in-residence-programme-2017/ http://blog.kbresearch.nl/2016/06/08/terms-and-conditions-of-the-kb-researcher-in-residence-programme-2017/#respond Wed, 08 Jun 2016 09:48:41 +0000 http://blog.kbresearch.nl/?p=1822 This programme as detailed at the KB-website (“Programme”) is operated by the Koninklijke Bibliotheek, National Library of the Netherlands (“KB”), Prins Willem-Alexanderhof 5 (2509 LK) Den Haag, The Netherlands.

  1. General
    1. If you enter the Programme you agree to abide by all of the following Terms and Conditions under which the Programme is run, and to be bound by them. We may disqualify you without prior notice if you are in breach of any of these Terms and Conditions.
    2. KB reserves the right to cancel the Programme at any stage if KB deems this necessary or circumstances arise that are outside of our control.
  2. Entries
    1. The Programme is open to all excluding employees or contractors of the KB or their direct family members.
    2. The Programme is only open to Ph.D.-students, or applicants who have obtained a Ph.D.-degree between 2011 and 2016, and who are aged 18 years or older.
    3. The Programme is only open to researchers employed by a university or a research institute within the EU.
    4. Entry is free.
    5. To enter the Programme, you must:
      1. Submit your proposal in accordance with the Call for proposals via the KB website; and
      2. Ensure that (i) your proposal is original (you can submit a proposal you have submitted previously as long as it has not been used in another programme or fellowship) and (ii) your proposal does not in any way infringe the copyright or other intellectual property rights, or other rights of any third party.
      3. Ensure that your proposal is aimed at the use of KB-data that benefits your research and other users of KB and/or the KB Research Lab.
      4. Tick the box to accept these Terms and Conditions when you submit your entry.
    6. All entries must be received by midnight on 8 November 2015. Any entries received after this will not be considered.
    7. The entries will be judged by an internal committee and then forwarded to an external committee of representatives experts from several Dutch universities and institutions.
    8. You may be contacted by email or telephone to answer further questions about your entry up to 15 December 2015. Two entries and two backup-entries will be chosen for the final stage. The best two entries will be contacted personally by email to assess whether those applicants are available for the secondment-period.
  3. Rights and permissions
    1. All entrants accept and agree that KB may publish the title and abstract of your proposal on the KB Research Blog at the KB’s sole discretion by January 2016. This blog is harvested for the Dutch web archive.
    2. In the event that you are offered a secondment:
      1. you agree to enter in a written agreement with KB and the university or research institute where you are employed on, among other things, intellectual property rights, and the obligations of the parties concerned including the obligations as set out in these Terms and Conditions.
      2. you agree to participate in any publicity planned by KB, if required;
      3. you will be expected to work on your project at least one day per week in residence at the KB in Den Haag, between 1 January 2017 and 30 June 2017 or 1 July 2017 and 31 December 2017, for 0.5 fte.
      4. you agree to present your research results to employees of the KB in the final month of your secondment.
      5. you will be expected to publish a blog about your project on the KB Research Blog.
      6. you agree to submit a Certificate of Conduct for Natural Persons (‘VOG NP’).
      7. you agree to mention the secondment KB offered you in any publication your secondment will give rise to.
    3. If any software is produced for your research, it will be developed on the principles of open source software and it will be made available on the KB Research Lab and Github website under a GPLv3
    4. KB offers you access to all data sets of the KB and support from a programmer, collection specialists and data specialists.
    5. During the secondment, KB offers you an office space and payment of all travel expenses within the Netherlands regarding the secondment.
    6. In case you do not live in the Netherlands, you are responsible to find and fund your own housing and pay for travel expenses to the KB in Den Haag, the Netherlands.
  4. Personal data
    1. Other than as expressly permitted by these Terms and Conditions, KB will only use your contact details for the purposes of administering this Programme, and will not publish them or provide them to anyone without your permission.
    2. You consent to KB holding and processing data relating to you for legal, administrative and management purposes.
    3. Any personal data relating to you will be used solely in accordance with the current data protection legislation in the Netherlands (Wet bescherming persoonsgegevens) and will not be disclosed to another party – except for the external committee of representatives experts from several Dutch universities and institutions – without your prior consent. Please see the Privacy statement KB for further details.
    4. Data relating to you will be retained by KB for a reasonable period after the closing date specified in clause 2.6 to assist in the administration of the Programme in a consistent manner and to deal with any queries on the Programme.
  5. Applicable Law
    1. These Terms and Conditions and any dispute or claim arising out of or in connection with them shall be governed by and construed in accordance with Dutch law.
    2. The parties will attempt in good faith to resolve any dispute or claim arising out of or relating to these Terms and Conditions promptly by negotiation. If the dispute cannot be resolved by negotiation, you hereby agree that the sole jurisdiction and venue for any actions that may arise in relation to the subject matter hereof shall be the Dutch Court in Den Haag, the Netherlands
]]>
http://blog.kbresearch.nl/2016/06/08/terms-and-conditions-of-the-kb-researcher-in-residence-programme-2017/feed/ 0
Supporting History Research with Temporal Topic Previews at Querying Time http://blog.kbresearch.nl/2015/04/20/supporting-history-research-with-temporal-topic-previews-at-querying-time/ http://blog.kbresearch.nl/2015/04/20/supporting-history-research-with-temporal-topic-previews-at-querying-time/#comments Mon, 20 Apr 2015 13:24:05 +0000 https://researchkb.wordpress.com/?p=1208 This post is written by Dr. Jiyin He – Researcher-in-residence at the KB Research Lab from June – October 2014.

Being able to study primary sources is pivotal to the work of historians. Today’s mass digitisation of historical records such as books, newspapers, and pamphlets now provides researchers with the opportunity to study an unprecedented amount of material without the need for physical access to archives. Access to this material is provided through search systems, however, the effectiveness of such systems seems to lag behind the major web search engines. Some of the things that make web search engines so effective are redundancy of information, that popular material is often considered relevant material, and that the preferences of other users may be used to determine what you would find relevant. These properties do not hold or are unavailable for collections of historical material. In the past 3 months I have worked at the KB as a guest researcher. Together with Dr. Samuël Kruizinga, a historian, we explored how we can enhance the search system at KB to assist the search challenges of the historian. In this blogpost, I will share our experience of working together, the system we have developed, as well as lessons learnt during this project.

A Historian at Work

Samuël summizes the research approach of a historian as a 4-stage procedure.

(1) Exploring. In this stage, a researcher has an initial research idea. With this initial idea in mind, he explores the target domain and literature in order to arrive at a preliminary research question. At this stage, the researcher also starts to explore possible primary data sources that can be used.

(2) Contextualisation. Given the preliminary research question, the researcher conducts historiographical analyses and further explores the availability of data sources. By the end of this stage, the researcher arrives at a refined research question.

(3) Operationalisation. Given the refined research question, the researcher formulates his theory or model, and decides on the primary sources from which he will search for materials to answer his research question. Based on his theory or model, the researcher formulates a set of sub research questions. These sub research questions form the sufficient or necessary conditions to answer the original research question.

(4) Execution. In this stage, the researcher actually starts to execute search in the data sources that have been selected from the previous stage. For each sub research question, the following steps are taken:

  • Search each data source with source-specific queries. That is, these queries are specific with respect to the type of content, metadata, organisation, etc., of the source collection. Note that this is not necessarily done with a single search system, or even with digital tools.
  • By analysing the retrieved materials, the researcher attempts to answer the sub research questions, and evaluates its impact on the original research question.
  • If results of this stage are not satisfying,  the researcher goes back to stage (3).

The stages sketched above vary in the search style, information needs, and material sought for, hence different digital tools might be developed to best support the historian at each stage. For instance, in the first two stages, there is a need to explore, overview, and compare different data sources to assist the researcher to discover possible data sources, and to select the ones that may contain useful information. In later stages, the focus moves towards locating detailed information objects (e.g., articles, images), relevant to solve specific research questions. In this project, we focus on the latter, i.e., to provide support in exploring and access information in a particular data source, namely the historical Dutch newspapers.

The prototype system

One misconception about the role of digital tools in humanities research is that they should hide the complexity of the data selection or analysis techniques from the user. Samuel describes the role of digital tools as: supporting researchers to locate and gain access to potentially relevant materials, while allowing the researcher to select, digest, and interpret the materials. He stresses the risk of blindly following the results and analyses generated by digital tools and his request for our system’s functionality can be summarised as: control and transparency. That is, to have control over the ways to search and explore a collection, and simple but understandable operations are preferred over blackbox complex algorithms.

Data collection

The data used in our project included the KB historical collection during the first world war (WWI) period (1914 – 1940). As this material aligns with Samuël’s research interest of comparing collective memories of WWI as represented in the newspapers of different regional groups, e.g., “Landelijke” and “Nationale/Lokaal”.

Elasticsearch (ES) provides the basic framework for our retrieval system. News articles were retrieved from the KB data API, processed and then stored in an ES index. In total 55,639,628 articles were indexed. More details about the indexing process with ES are provided later in this post.

Query interface

On way of providing users with greater control over their search is to provide a richer querying language than keywords. In particular we decided to implement the following features: (1) Boolean query operation. This include specifying terms that “must occur”, “should occur but not necessarily occur”, or “must not occur” in the retrieved articles. (2) Wildcard queries and proximity queries. (3) Filtering on multiple time ranges and newspaper selections.

Note that features such as wildcard screenshot_jiyin1and proximity queries are already  supported by Lucene, the underlying search engine of KB’s Delpher system (as well as the basis of Elasticsearch). However, it is rather implicit, as users may not be aware of what is possible, and may not be able to use the query language defined by Lucene.  Here, we make the possible operations explicit, including explicit instructions on how wildcard and proximity queries should be constructed, as shown in the screenshot on the right.

Result operation

To provide greater control over the results displayed, we decided on two additional sorting options in addition to the default ranking criterion (i.e., by relevance scores), namely sorting by date, and by article length. Sorting by date provides a quick means to locate articles on specific dates. The argument for sorting by article length is that longer articles are likely to contain important content while news articles consisting of few lines are generally of less value.

screenshot_jiyin2

Interestingly, in many retrieval tasks (e.g., microblog search, ad-hoc search), document length and date have been combined with the relevance scores for the ranking of retrieved results. In our case, however, it was prefered that different ranking criteria are kept separated, and that the control of ranking criteria is transparent and flexible to the researcher.

Query preview

In Web search, query suggestion is a common means to assists users in issuing better keyword queries. In our case, users are allowed to issue queries with constraints such as time range and selected newspapers. That is, in an effort to support users in formulating their queries, suggestions of query words combined with suggestions of appropriate values of these constraints are needed. To this end, we implemented a temporal topic preview widget. That is, while typing a query, users can see named entities from the documents in the form of term clouds along a timeline. This widget is intended to help users in three ways:

  1. To determine the interesting time periods.
  2. To identify entities related to the original query which can be used for query reformulation.
  3. To compare topics discussed in different types of newspapers.

The methods to generate entity clouds for a specific year and for a given selection of newspapers are described next.

Entities.  For each entity cloud, we select the top 10 most significant entities from the articles within that period and from the selected newspapers.  Entities were extracted using KB’s named entity recognizer and prestored in the index.

Entity selection. To select the top 10 entities, we take the following steps.

  • We consider a foreground and a background document set. The foreground set consists of articles that contain the query words and are in the selected period and newspapers. The background set consists of all articles in the collection. Our goal is to select entities that are representative (e.g., frequently occur) in the foreground set in contrast to the background set.
  • We compute two types of conditional probabilities: the probability that the given entity is “generated” by the foreground set (p), and the probability that it is generated by the background set (q), using a language modeling approach. We then compute the Kullback-Leibler divergence between the two probability distributions KL(p||q).
  • Entities within the foreground document set are ranked in descending order of their KL divergence scores.

Preview updates. When the user types in query words or changes newspaper selections, the preview updates. To prevent updates on incomplete query input, we wait for 500ms after the user stops typing before updating.

An illustrative example

The following example is generated by typing in the query word “beurskrach”, referring to the economic crisis around 1930. In the screenshot we see two rows of entity clouds. The top row is generated from “Landelijke” newspapers, and the bottom row is generated for the “Nationale / Lokaal” newspapers.

screenshot_jiyin3

We have the following observations.

  1. This word starts to appear in news after 1929. This is correct, as the crisis starts in that year. The entity clouds in the previous years were absent as no articles containing this word were retrieved.
  2. We can see entities relevant to the crisis were selected, e.g., New York.
  3. If we compare the “Landelijke” newspapers to the “Nationale / Locaal” newspapers, we find that in local newspapers the crisis is hardly discussed.

UI wrap up

Finally, the resulting user interface looks as follows. While the user is formulating his or her query, the temporal topic overview is shown to assist this query formulation process (left). After the user has submitted the query, search results are shown, with possible result operations (right).

screenshot_jiyin4 screenshot_jiyin5

Lessons learnt

Finally, I would like to discuss some of the lessons learnt during this interdisciplinary project to design experimental search tools for historical research.

Support for an iterative design process.  In this project, I started with a standard search engine setup, i.e., to support keyword searching in document content. Later, after Samuel jointed the project, we started to add additional features to the system. This is an iterative process consisting of discussion – implementation – testing – new discussion. During this process new requested features kept emerging.

The updates of features fall into two categories: at the UI level and at the index level. Updates at the UI level are relatively simple: it adds additional access or means of interactions with existing (indexed) data. Updates at the indexing level enables additional data to be searchable, which is more complicated — often it means reindexing the collection.

The updates of index were necessary in two situations: metadata that existed but was not in the same collection (e.g., the page number of a news articles existed, but in a different versions of the KB news collection than the one previously indexed); derived data (e.g., article length — while it is possible to compute it at querying time, in order to allow efficient sorting of results document length was included at indexing time).

With respect to the design and development of tools for historical research my view is as follows:

  • From a system perspective, we need systems that allow flexible index updates. It was helpful that Elasticsearch allows adding additional fields without reindexing the data. In addition, when reindexing has to happen (e.g., when the data schema has changed), it can be done in the background without bringing down the whole system.
  • From a user perspective, it may be useful to provide supporting tools that allow the target users to explore the availability of data as well as possible derived data in the early stage of the design process.

Memory issues. It was the first time I used Elasticsearch. Before, I have always been using academic search systems such as Lemur and Terrier. I decided to experiment with Elasticsearch mainly because of its rich aggregation functions.

One of the issues I have been struggling with was the memory usage. Given the focus of the project as well as its short duration, I did not experiment with different configurations of ES, but simply used default configuration.  The indexed collection consists of 55,639,628 documents, resulting in an index of 260G. It seems that the memory usage can easily go over 10G, which is rather surprising. The machine we used has 30G RAM. While it was fine to perform simple keyword search, operations such as sorting on a specific field or more complexed queries can lead to out of memory exceptions. Unfortunately, I did not encounter this problem until the end of the project when all the data were indexed and all functionalities were implemented. The resulting system is therefore rather unstable.

Here are what I learnt with respect to the use of Elasticsearch: (1) It is not trivial to set the appropriate configuration for the elasticsearch system. Careful study and experiments are needed. (2) In order to properly configure the system, it is important to have an estimation of the size of the collection before hand. In my case, two factors make the estimation difficult. On the one hand, the documents were retrieved from a data service API, which were processed and indexed on the fly.  On the other hand, we kept updating the index with additional data fields with respect to newly emerged functionality requests throughout the project.

Outlook

In this project, we have discussed the research practice of historians and explored possible ways to support this process with novel search tool features. While a prototype system has been developed, much is left for further exploration, e.g., do the research practices as well as requirements for search systems found in this project generalise to that of other historians?

Dr. He’s tool is available for download on KB Research’s Github: https://github.com/KBNLresearch/spatio-temporal-topics  

]]>
http://blog.kbresearch.nl/2015/04/20/supporting-history-research-with-temporal-topic-previews-at-querying-time/feed/ 1
Researcher-in-residence at the KB http://blog.kbresearch.nl/2015/04/09/researcher-in-residence-at-the-kb/ http://blog.kbresearch.nl/2015/04/09/researcher-in-residence-at-the-kb/#comments Thu, 09 Apr 2015 11:58:45 +0000 https://researchkb.wordpress.com/?p=1217 At DH2013, we presented a poster to ask researchers what they need from a National Library. The responses varied from ‘Nothing, just give us your data’ to ‘We’d like to be fully supported with tools and services’, showing once again that different users have different requirements. In order to accommodate all groups of researchers, the Collections department of the KB, who ‘own’ the data, and the Research department, where tools and services are developed, combined efforts and spoke to scholars to  discuss the best method of supporting their work. However, we noticed that it was still quite difficult to get a good idea of how they used our data and in what way our actions and decisions would benefit them. Also, it seemed that researchers were often not aware of what activities the we undertake in this respect, which led to work being done twice.

In order to bridge this gap between our services and the needs of the scholars, a researcher-in-residence program was set up in the summer of 2014. In this program, we invite researchers into the library to work with our data, programmers and in our virtual research lab to learn how the scholars conduct their research, but also to benefit from their expertise to further our services. The program is aimed at young researchers in the early stages of their career (PhD or postdoc), in order to provide new opportunities for young scholars, and consists of a short research project (3-6 months) in the KB. The researcher is supported by programmers and experts from the KB. The results of the projects feed back into the KB, either by means of a post on this blog, or by deploying a prototype in the KB Research Lab, and are often accompanied with a publication.

The first three placements of the program were set up as pilot projects. For these projects, we invited three Dutch universities that currently use our data to join the program. This resulted in the selection of two historians, one media researcher and two computer scientists from the University of Amsterdam, Utrecht University, Erasmus University Rotterdam and CWI. The latest researcher is about to start his work at the KB and as we are very enthusiastic about the program, we have decided to start blogging on the work that is done by the researchers. We will evaluate the program in the summer of 2015 and hope to successfully evolve from pilot phase to full-blown residencies.

Our first blog post is written by Dr. Jiyin He, a computer scientist from the University of Amsterdam, who was a researcher-in-residence at the KB from July – October 2014, and will be posted next week.

The researchers that have joined the program so far are:

Do you want to know more about our efforts to work with researchers and this program? Come visit us at our poster presentation at DH2015 or come find the KB at the poster sessions of DHBenelux!

]]>
http://blog.kbresearch.nl/2015/04/09/researcher-in-residence-at-the-kb/feed/ 2
What’s happening with our digitised newspapers? http://blog.kbresearch.nl/2015/04/08/whats-happening-with-our-digitised-newspapers/ http://blog.kbresearch.nl/2015/04/08/whats-happening-with-our-digitised-newspapers/#respond Wed, 08 Apr 2015 13:29:25 +0000 https://researchkb.wordpress.com/?p=1164 The KB has about 10 million digitised newspaper pages, ranging from 1650 until 1995. We negotiated rights to make these pages available for research and this has happened more and more over the past years. However, we thought that many of these projects might be interested in knowing what others are doing and we wanted to provide a networking opportunity for them to share their results. This is why we organised a newspapers symposium focusing on the digitised newspapers of the KB, which was a great success!

Prof. dr. Huub Wijfjes (RUG/UvA) showing word clouds used in his research.

Prof. dr. Huub Wijfjes (RUG/UvA) showing word clouds used in his research.

On 24 March, we opened up our auditorium to 10 research projects that use our digitised newspapers. The presenters came from many areas of DH research as we had historians, linguists, computer programmers, a linked data specialist, and media researchers. Next to this, we also asked colleagues to present what they are working on with regards to our newspapers. This resulted in presentations about copyright, digitisation (OCR and future plans), a Wikipedia project, the Research Lab and Europeana Newspapers. All in all we had three blocks of presentations spanning the whole day, and an auditorium filled with interested researchers, journalists, librarians, students and colleagues.

The original idea of the day was to provide a networking opportunity for those researchers who used our newspapers in their projects, but while organising it we noticed that not only those researchers were interested in hearing what others are doing, but that a wider group of people wanted to know what was possible with the data set. When we opened up the registration, we quickly outgrew the room we booked and had to transfer to the auditorium. Ultimately, we had 145 people who visited the symposium and who were responsible for lively discussions after the presentations and during the breaks.

Discussions in between presentations.

Discussions in between presentations.

The varied audience and speakers of the day also meant having different levels of expertise in Digital Humanities and working with data, but this provided very valuable insights into working with the newspapers. The differences in skills between, for example, a historian and a programmer provides challenges when working together in research projects, but also provides a very valuable learning experience for both.

Working with new computational methods might be a bit scary for an inexperienced researcher, but can mean very interesting results, according to Dr. Martijn Kleppe (EUR). But how does a researcher know if these results are trustworthy? If the software actually does its job? Dr. Antske Fokkens (VU) explains that it is important to show researchers that they should be aware of what software can and cannot do and how you should interpret results. Another common technical issue that researchers came across is the lack of computational capacity. Humanities faculties often do not have the setup to process large amounts of data.

Dr. Antske Fokkens on the methods used at the Vrije Universiteit.

Dr. Antske Fokkens on the methods used at the Vrije Universiteit.

Next to these technical challenges, the material itself also provides problems when using it, such as OCR issues, lack of access due to the copyright, changes in the meaning of words and historical spelling and of course the ‘problem’ that we have not yet digitised every newspaper page in our collection. Working with this data and around these challenges is something that we can work together in. The KB works hard to improve the access to and usability of the newspapers and with the KB Research Lab, we can also provide support in the use of the data.

Martijn Kleppe listening to his colleague Laura Hollink

Martijn Kleppe listening to his colleague Laura Hollink

Because of this, many researchers have already used the newspapers in great research projects, such as Polimedia. Here, our newspapers are linked to the parliamentary papers and the radio bulletins, providing a search engine that shows how parliamentary debates are mentioned in the media. Or the Dutch Ships and Sailors-project where maritime data is linked together (and to our newspapers) to provide information about Dutch ships in the 18th and 19th century. Mentioning all success stories might result in a very long blog, but all presentations (in Dutch) are available via our website (scroll down for the URLs).

We were very happy with the whole symposium, the great speakers and the wonderful audience and we hope to be able to organise something similar in the future. If you are interested in our newspapers, other data sets, or our Research Lab, feel free to contact us!

]]>
http://blog.kbresearch.nl/2015/04/08/whats-happening-with-our-digitised-newspapers/feed/ 0
Workshop topic modelling with MALLET at KB http://blog.kbresearch.nl/2015/01/27/workshop-topic-modelling-with-mallet-at-kb/ http://blog.kbresearch.nl/2015/01/27/workshop-topic-modelling-with-mallet-at-kb/#respond Tue, 27 Jan 2015 09:59:44 +0000 https://researchkb.wordpress.com/?p=1101

[A] topic model is a type of statistical model for discovering the abstract “topics” that occur in a collection of documents (Wikipedia).

Topic modelling is a very popular method in the Digital Humanities to discover more about a large set of data and is also used by many researchers working on data of the KB. Unfortunately, not all topic modelling tools are as easy to access, due to a lack of technical skills or a lack of access to the data for example. The current guest researcher at the KB (Dr. Samuël Kruizinga) came across such problems while doing his research into the memory of the First World War in the KB newspapers. Not only was it difficult for him to select a corpus to work with, he was also unfamiliar with the go-to tool MALLET. Luckily, his university (Universiteit van Amsterdam) wanted to help and provided funds to organise a workshop, not only for him, but also for other academics interested in topic modelling.

The workshop was focused on topic modelling for researchers interested in the historical newspapers of the KB and was taught by Dr. Marijn Koolen, assistent professor of Digital Humanities at the UvA. The 15 participants were all academics or supporting staff with little to no experience with topic modelling or digital humanities.

Marijn Koolen at MALLET workshop

Dr. Marijn Koolen at MALLET workshop

 

The afternoon workshop consisted of an hour of theory where Marijn explained how topic modelling worked, what can be expected of the method when using the KB newspapers and also what the limits are of this method. For this, he prepared some examples using the corpus that Samuël is using in his research, namely the newspapers from the Interbellum. (If you are interested in this corpus, please contact dataservices@kb.nl for more information.). Marijn’s presentation is available on his website.

Participants of MALLET workshop

Participants of MALLET workshop

After a short break, the practical part of the workshop could start! We asked all participants to install the software beforehand to make sure our time was spent topic modelling and not installing software. The Programming Historian has a very useful guide for this and all participants were able to install everything on their laptops. We made sure there was technical backup in the hour before the workshop for any questions, but this proved unnecessary.

With the help of a few short exercises and a sample set of KB newspapers (and Marijn of course), we were able to create a collection of topics related to the First World War. Working together was the key to the exercises, as all participants ran into a problem at one time or another. In the exercises we learned the difference between working with 10 or 100 topics, and having a set of 100, 1000 or 10.000 articles. This gave us a very good insight into the workings of the tool and what we could expect when working with it. Of course, it is then up to the academics to use the output in their research!

The afternoon ended with drinks and interesting conversations. We hope to organise more of such meetings to encourage researchers to work with KB data and to learn more about the way academics use our material for their research. If you are interested to join a similar workshop, please let us know at research@kb.nl and we’ll keep you updated!

]]>
http://blog.kbresearch.nl/2015/01/27/workshop-topic-modelling-with-mallet-at-kb/feed/ 0
OCR improvement: helping and hindering researchers http://blog.kbresearch.nl/2014/08/19/ocr-improvement-helping-and-hindering-researchers/ http://blog.kbresearch.nl/2014/08/19/ocr-improvement-helping-and-hindering-researchers/#comments Tue, 19 Aug 2014 08:50:53 +0000 http://researchkb.wordpress.com/?p=979 Author: Tineke Koster

As I am writing this, volunteers are rekeying our 17th century newspapers articles. Optical character recognition of the gothic text type in use at the time has yielded poor results, making this part of our digital collection nearly inaccessible for full-text search. The Meertens institute, who have an excellent track record when it comes to crowdsourcing, has developed the editor (Dutch). Together with them we are working towards a full update of all newspaper issues from 1618 to 1700 that are available in our website Delpher.

Great news and, for some researchers, an eagerly awaited development. A bright future beckons in which our digital text corpus is 100% correct, just waiting to be mined for dynamic phenomena and paradigm shifts.

But we have to realize that without the proper precautions, correcting digital texts may also hinder researchers in their work. How so? These texts may have been used (browsed, mined, cited, etc.) by researchers in their earlier form. The improvement or enrichment may have consequences for the reproducibility of their research results.

For all researchers the need to reproduce research results is growing, with new guidelines due to new laws. There is also a specific group of researchers that need sustained access to older versions of digital text. The need is highest for research where the goal is to develop an algorithm and to assess its quality relative to previous versions of the same algorithm or to other algorithms. Without sustained access to older versions, these people cannot do their work.

Is it our role to provide this access? How the National Library of the Netherlands is thinking about this issue, I hope to explain in a later blogpost (soon!). Meanwhile, I would be very interested to hear your experiences. How is this subject discussed in your organization? Does your organization have a policy in place to deal with this?

]]>
http://blog.kbresearch.nl/2014/08/19/ocr-improvement-helping-and-hindering-researchers/feed/ 9
Roles and responsibilities in guaranteeing permanent access to the records of science – at the Conference for Academic Publishers (APE) 2014 http://blog.kbresearch.nl/2014/03/04/roles-and-responsibilities-in-guaranteeing-permanent-access-to-the-records-of-science-the-conference-for-academic-publishers-ape-conference-2014/ http://blog.kbresearch.nl/2014/03/04/roles-and-responsibilities-in-guaranteeing-permanent-access-to-the-records-of-science-the-conference-for-academic-publishers-ape-conference-2014/#respond Tue, 04 Mar 2014 00:01:41 +0000 http://researchkb.wordpress.com/?p=568 On Tuesday 28 and Wednesday 29 January the annual Conference for Academic Publishers Europe was held in Berlin. The title of the conference: Redefining the Scientific Record. – Report by Marcel Ras (NCDD) and Barbara Sierman (KB)

Dutch politics set on “golden road” to Open Access

During the fist day the focus was on Open Access, starting with a presentation by the Dutch State Secretary for Education, Culture and Science on Open Access. In his presentation called “Going for Gold” Sander Dekker outlined his policy with regards to the practice of providing open access to research publications and how that practice will continue to evolve. Open access is “a moral obligation” according to Sander Dekker. Access to scientific knowledge is for everyone. It promotes knowledge sharing and knowledge circulation and is essential for further development of society.

OA "gold road" supporter and State Secretary Sander Dekker (right) during a recent visit to the K

“Golden road” open access supporter and State Secretary Sander Dekker (right) during a recent visit to the KB – photo KB/Jacqueline van der Kort

Open access means having electronic access to research publications, articles and books (free of charge). This is an international issue. Every year, approximately two million articles appear in 25,000 journals that are published worldwide. The Netherlands account for some 33,000 articles annually. Having unrestricted access to research results can help disseminate knowledge, move science forward, promote innovation and solve the problems that society faces.

The first steps towards open access were taken twenty years ago, when researchers began sharing their publications with one another on the Internet. In the past ten years, various stakeholders in the Netherlands have been working towards creating an open access system. A wide variety of rules, agreements and options for open access publishing have emerged in the research community. The situation is confusing for authors, readers and publishers alike, and the stakeholders would like this confusion to be resolved as quickly as possible.

The Dutch Government will provide direction so that the stakeholders know what to expect and are able to make arrangements with one another. It will promote “golden” open access: publication in journals that make research articles available online free of charge. The State Secretary’s aim is fully implement the golden road to open access within ten years, in other words by 2024. In order to achieve this, at least 60 per cent of all articles will have to be available in open access journals in five years’ time. A fundamental changeover will only be possible if we cooperate and coordinate with other countries.

Further reading: http://www.government.nl/issues/science/documents-and-publications/parliamentary-documents/2014/01/21/open-access-to-publications.html orhttp://www.rijksoverheid.nl/ministeries/ocw/nieuws/2013/11/15/over-10-jaar-moeten-alle-wetenschappelijke-publicaties-gratis-online-beschikbaar-zijn.html

Do researchers even want Open Access?

The two other keynote speakers, David Black and Wolfram Koch presented their concerns on the transition from the current publishing model to open access. Researchers are increasingly using subject repositories for sharing their knowledge. There is an urgent need for a higher level of organization and for standards in this field. But who will take the lead? Also, we must not forget the systems for quality assurance and peer review. These are under pressure as enormous quantities of articles are being published and peer review tends to take place more and more after publication. Open access should lower the barriers for access to research for the users, but what about the barriers for scholars publishing on their research? Koch stated that the traditional model worked fine for researchers. They don’t want to change. However, there do not seem to be any figures to support this assertion.

It is interesting to note that in almost all presentations on the first day of APE digital preservation was mentioned one way or the other. The vocabulary was different, but it is acknowledged as an important topic. Accessibility of scientific publications for the long term is a necessity, regardless of the publishing model.

KB and NCDD workshop on roles and responsibilities

The 2nd day of the conference the focus was on innovation (the future of the article, dotcoms) and on preservation!

The National Library of The Netherlands (KB) and the Dutch Coalition for Digital Preservation (NCDD) organized a session on preservation of scientific output: “Roles and responsibilities in guaranteeing permanent access to the scholarly record”. The session was chaired byMarcel Ras, program manager for the NCDD.

The trend towards e-only access for scholarly information is increasing at a rapid pace, as well as the volume of data which is ‘born digital’ and has no print counterpart. As for scholarly publications, half of all serial publications will be online-only by 2016. For researchers and students there is a huge benefit, as they now have online access to journal articles to read and download, anywhere, any time. And they are making use of it to an increasing extend. However, the downside is that there is an increasing dependency on access to digital information. Without permanent access to information scholarly activities are no longer possible. For libraries there are many benefits associated with publishing and accessing academic journals online. E-only access has the potential to save the academic sector a considerable amount of money. Library staff resources required to process printed materials can be reduced significantly. Libraries also potentially save money in terms of the management and storage of and end user access to print journals. While suppliers are willing to provide discounts for e-only access.

Publishers may not share post-cancellation and preservation concerns

However, there are concerns that what is now available in digital form may not always be available due to rapid technological developments or organisational developments within the publishing industry; these concerns and questions about post-cancellation access to paid-for content are key barriers to institutions making the move to e-only. There is a danger that e-journals become “ephemeral” unless we take active steps to preserve the bits and bytes that increasingly represent our collective knowledge. We are all familiar with examples of hardware becoming obsolete; 8 inch and 5.25 inch floppy discs, Betamax video tapes, and probably soon cd-roms. Also software is not immune to obsolescence.

In addition to this threat of technical obsolescence there is the changing role of libraries. Libraries have in the past assumed preservation responsibility for the resources they collect, while publishers have supplied the resources libraries need. This well-understood division of labour does not work in a digital environment and especially so when dealing with e-journals. Libraries buy licenses to enable their users to gain network access to a publisher’s server. The only original copy of an issue of an e-journal is not on the shelves of a library, but tends to be held by the publisher. But long-term preservation of that original copy is crucial for the library and research communities, and not so much for the publisher.

Can third-party solutions ensure safe custody?

So we may need new models and sometimes organizations to ensure safe custody of these objects for future generations. A number of initiatives have emerged in an effort to address these concerns. Research and development efforts in digital preservation issues have matured. Tools and services are being developed to help plan and perform digital preservation activities. Furthermore third-party organizations and archiving solutions are being established to help the academic community preserve publications and to advance research in sustainable ways. These trusted parties can be addressed by users when strict conditions (trigger events or post-cancellation) are met. In addition, publishers are adapting to changing library requirements, participating in the different archiving schemes and increasingly providing options for post-cancellation access.

In this session the problem was presented from the different viewpoints of the stakeholders in this game, focussing on the roles and responsibilities of the stakeholders.

Neil Beagrie explained the problem in depth, both in a technical, organisational and financial sense. He highlighted the distinction between perpetual access and digital preservation. In the case of perpetual access, organisations have a license or subscription for an e-journal and either the publisher discontinues the journal or the organisation stops its subscription – keeping e-journals available in this case is called “post-cancellation” . This situation differs from long-term preservation, where the e-journal in general is preserved for users whether they ever subscribed or not. Several initiatives for the latter situation were mentioned as well as the benefits organisations like LOCKSS, CLOCKSS, Portico and the e-Depot of the KB bring to publishers.  More details about his vision can be read in the DPC Tech Watch report Preservation, Trust and Continuing Access to e-Journals . (Presentation: APE2014_Beagrie)

Susan Reilly of the Association of European Research Libraries  (LIBER) sketched the changing role of research libraries. It is essential that the scholarly record is preserved, which encompasses e-journal articles, research data, e-books, digitized cultural heritage and dynamic web content. Libraries are a major player in this field and can be seen as an intermediary between publishers and researchers. (Presentation: APE2014_Reilly)

Eefke Smit of the International Association of Scientific, Technical and Medical Publishers (STM) explained to the audience why digital preservation was especially important in the playing field of STM publishers. Many services are available but more collaboration is needed. The APARSEN project is focusing of some aspects like trust, persistent identifiers and cost models, but there are still a wide range of challenges to be solved as the traditional publication models will continually change, from text and documents to “multi-versioned, multi-sourced and multi-media”. (Presentation: APE2014_Smit)

As Peter Burnhill from EDINA, University of Edinburgh, explained, continued access to the scholarly record is under threat as libraries are no longer the custodians of the scholarly record in e-journals. As he phrased it nicely: libraries no longer have e-collections but only e-connections. His KEEPERS registry is a global registry of e-journal archiving and offers an overview of who is preserving what. Organisations like LOCKSS, CLOCKSS, the e-Depot, the Chinese National Science Library and, recently, the US Library of Congress submit their holding information to this KEEPERS Registry. However nice, it was also emphasized that the registry only contains a small percentage of existing e-journals (currently about 19% of the e-journals with an ISSN assigned). More support for the preserving libraries and more collaboration with publishers is needed to preserve the e-journals of smaller publishers and improve coverage. (Presentation: APE2014_Burnhill)

(Reblogged with slight changes from http://www.ncdd.nl/blog/?p=3467)

]]>
http://blog.kbresearch.nl/2014/03/04/roles-and-responsibilities-in-guaranteeing-permanent-access-to-the-records-of-science-the-conference-for-academic-publishers-ape-conference-2014/feed/ 0
On-line scholarly communications: vd Sompel and Treloar sketch the future playing field of digital archives http://blog.kbresearch.nl/2014/01/22/on-line-scholarly-communications-and-the-role-of-digital-archives/ http://blog.kbresearch.nl/2014/01/22/on-line-scholarly-communications-and-the-role-of-digital-archives/#comments Wed, 22 Jan 2014 10:33:11 +0000 http://researchkb.wordpress.com/?p=434 The Dutch data archive DANS invited two ‘great thinkers and doers’ (quote by Kevin Ashley on Twitter) in scholarly communications to do some out-of-the-box thinking about the future of scholarly communications – and the role of the digital archive in that picture. The joint efforts of DANS visiting fellows Herbert van de Sompel (Los Alamos) and Andrew Treloar (ANDS) made for a really informative and inspiring workshop on 20 January 2014 at DANS. Report & photographs by Inge Angevaare, KB Research

(a copy of) Rembrandt's 17th-century scholar Dr. Tulp overseeing Herbert van de Sompel outlining the research world of the 21st century

Rembrandt’s 17th-century scholar Dr. Tulp overseeing Herbert van de Sompel outlining the research world of the 21st century (the painting is a copy …)

Life used to be so simple. Researchers would do their research and submit their results in the form of articles to scholarly journals. The journals would filter out the good stuff, print it, and distribute it. Libraries around the world would buy the journals and any researcher wishing to build upon the published work could refer to it by simple citation. Years later and thousands of miles away, a simple citation would still bring you to an exact copy of the original work.

Van de Sompel and Treloar [the link brings you to their workshop slides] quoted Roosendaal & Geurts (1998) in summing up the functions this ‘journal system’ effectively performed:

  • Registration: allows claims of precedence for a scholarly finding (submission of manuscript)
  • Certification: establishes validity of claim (peer review, and post-publication commentary)
  • Awareness: allows actors in the system to remain aware of new claims (discovery services)
  • Archiving: preserves the scholarly record (libraries for print; publishers and special archives like LOCKSS, Portico and the KB for e-journals).
  • (A last function, that of academic recognition and rewards, was not discussed during this workshop.)

So far so good.

But then we went digital. And we created the world-wide web. And nothing was the same ever again.

Andrew Treloar (at the back) captivating his audience

Andrew Treloar (at the back) captivating his audience

Future scholarly communications: diffuse and ever-changing

Van de Sompel and Treloar went online to discover some pointers to what the future might look like – and found that the future is already here, ‘just not evenly distributed’. In other words: one discipline is moving into the digital reality at a faster pace than another, and geographically there are many differences too. But van de Sompel and Treloar found many pointers to what is coming and grouped them in Roosendaal & Geurts’s functional framework:

  • Registration is increasingly done on (discipline-specific) online platforms such as BioRxiv, ideacite (where one can register mere ‘ideas’!) and Github, a collaborative platform for software developers (also used by the KB research team).
    Common characteristics include:
    – Decoupling registration from certification
    – Timestamping, versioning
    – Registration of various types of objects
    – Machines also function as creators and contributors.
    (We’ll discuss below what these features mean for digital archiving)
  • Certification is also moving to lots of online platforms, such as PubMed Commons, PubPeer, ZooUniverse and even Slideshare, where the number of views and downloads is an indication of the interest generated by the contents.
    Common characteristics include:
    – Peer-review is decoupled from the publication process
    – Certification of various types of objects (not just text)
    – Machines carry out some of the validating
    – Social endorsement
  • Awareness is facilitated by online platforms such as the Dutch ‘gateway to scholarly information’ NARCIS, myExperiment and a really advanced platform such as eLabNotebook RSS where malaria research is being documented as it happens and completely in the open.
    Common characteristics include:
    – Awareness for various types of objects (not just text)
    – Real time awareness
    – Awareness support targeted at machines
    – Awareness through social media.
  • Archiving is done by library consortia such as CLOCKSS, data archives such as DANS Easy, and, although not mentioned during the presentation I may add our own KB e-Depot.
    Common characteristics include:
    – Archiving for various types of objects
    – Distributed archives
    – Archival consortia
    – Audit for trustworthiness (see, e.g., the European Framework for Audit and Certification of Digital Repositories).
Very few places remained unused

Very few seats remained unoccupied

Fundamental changes

Here’s how van de Sompel and Treloar summarise the fundamental changes going on. (The fact that the arrows point both ways is, to my mind, slightly confusing. The changes are from left to right, not the other way around.)

vdSompelTreloar32

Huge implications for digital libraries and archives

The above slide merits some study, because the implications for libraries and digital archives are huge. In the words of vd Sompel and Treloar:

vdSompelTreloar33

From the ‘journal system’ we are moving towards what van de Sompel and Treloar call a ‘Web of Objects’ which is much more difficult to organise in terms of archiving, especially because the ‘objects’ now include ever-changing software & operating systems, as well as data which are not properly handled and thus prone to disappear (Notice on student cafe door: ‘If you have stolen my laptop, you may keep it if you just let me download my PHD-thesis’).

Why archiving is more difficult in the Web of Objects

Why archiving is more difficult in the Web of Objects (if print is too small, check out Slideshare original)

It’s like web archiving – ‘but we have to do better’

Van de Sompel and Treloar compared scholarly communications to websites – ever-changing content, lots of different objects (software, text, video, etc.), links that go all over the place. Plus, I may add, a enormous variety of producers on the internet. Van de Sompel and Treloar concluded: ‘We have to do better than present web-archiving methods if we are to preserve the scholarly record in any meaningful way.’

Two 'great thinkers and doers' confer - Herbert van de Sompel (left) and Andrew Treloar

Two ‘great thinkers and doers’ confer – Herbert van de Sompel (left) and Andrew Treloar

‘The web platforms that are increasingly used for scholarship (Wikis, GitHub, Twitter, WordPress, etc.) have desirable characteristics, such as versioning, timestamping and social embedding. Still, they record rather than archive: they are short-term, without guarantees, read/write and reflect the scholarly process, whereas archiving concerns longer terms, is trying to provide guarantees, is read-only and results in the scholarly record.’

The slide below sums it all up – and it is with this slide that van de Sompel and Treloar turned the discussion over to their audience of some 70 digital data experts, mostly from the Netherlands:

A work in progress: the scholarly communications arena of the future

A work in progress: the scholarly communications arena of the future

Group discussions about the digital archive of the future

So, what does all of this mean for digital libraries and digital archives? One afternoon obviously was not enough to analyse the situation in full, but here are some of the comments reported from the (rather informal) break-out sessions:

  • One thing is certain: it is a playing field full of uncertainties. Velocity, variety and volume are the key characteristics of the emerging landscape. And everybody knows how difficult these are to manage.
  • The ‘document-centred’ days, where only journal and book publications were rated as First Class Scholarly Objects are over. Treloar suggested a move to a ‘researcher-centric’ approach, where First Class Objects include publications and data and software.
  • To complicate matters: the scholarly record is not all digital – there are plenty of physical objects to deal with.
  • How do we get stuff from the recording platforms to the archives? Van de Sompel suggested a combination of approaches. Some of it we may be able to harvest automatically. Some of it may come in because of rules and regulations. But Van de Sompel and Treloar both figured that rules and regulations would not be able to cover all of it. That is when Andrea Scharnhorst (workshop moderator, DANS) suggested that we will have to allow for a certain degree of serendipity (‘toeval’ in Dutch).
Andrea Scharnhorst (DANS): 'Perhaps we have to allow for a certain degree of serendipity'

Andrea Scharnhorst (DANS): ‘Perhaps we have to allow for a certain degree of serendipity’

  • Whatever libraries and archives do, time-stamped versioning will become an essential feature of any archival venture. This is the only way to ensure that scientists can adequately cite anything and verify any research (‘I used version X of software Y at time Z – which can be found in a fixed form in Archive D’).
  • The archival community introduced the concept of persistent identifiers (PID’s) to manage the uncertainties of the web. But perhaps the concept’s usefulness will be limited to the archival stage. Should we distinguish between operational use cases and archival use cases?
  • Lots of questions remain about roles and responsibilities in this new picture, and who is to pay for what. Looking at the Netherlands, the traditional distribution of tasks between the KB National Library (books, journals) and the data archives (research data) certainly merits discussion in the framework of the NCDD (Netherlands Organisation for Digital Preservation); the NCDD’s new programme manager, Marcel Ras, attended the workshop with interest.
Breakout discussions about infrastructure implications

Breakout discussions about infrastructure implications

  • Who or what will filter the stuff that is worth keeping from the rest?
  • Interoperability is key in this complex picture. And thus we will need standards and minimal requirements (as, e.g., in the Data Seal of Approval)
  • Perhaps baffled by so much uncertainty in the big picture, some attendants suggested that we first concentrate on what we have now and/or are developing now, and at least get that right. In other words, let’s not forget that there are segments of the scientific landscape that are being covered even now. The rest of the scholarly communications landscape was characterised by Laurents Sesink (DANS) as ‘the Wild West’.
In this breakout session, clearly discussions focussed on the role of the archive.

In this breakout session, clearly discussions focussed on the role of the archive. Selection: when and by whom? Roles and responsibilities?

  • What if the Internet fails? What if it succumbs to hacks and abuse? This possibility is not wholly unimaginable. But the workshop decided not to go there. At least not today.

In his concluding remarks Peter Doorn, Director of DANS, admitted that there had been doubts about organising this workshop. Even Herbert van de Sompel and Andrew Treloar asked themselves: ‘Do we know enough?’ Clearly, the answer is: no, we do not know what the future will bring. And that is maybe our biggest challenge: getting our minds to accept that we will never again ‘know enough’ at any time. While yet having to make decisions every day, every year, on where to go next. DANS is to be commended for creating a very open atmosphere and for allowing two great minds to help us identify at least some major trends to inspire our thinking.

See also:

  • tweets #rtwsaf (after the official name of the workshop, Riding the Wave and the Scholarly Archive of the Future – the title referring to the 2010 European Commission Report on Scholarly Communications which was the last major report on the issue available).
  • Blog post by Simon Hodson
Where do we go from here? Peter Doorn asked his two visiting fellows

Where do we go from here? Peter Doorn asked his two visiting fellows in Alice-in-Wonderland fashion

]]>
http://blog.kbresearch.nl/2014/01/22/on-line-scholarly-communications-and-the-role-of-digital-archives/feed/ 1