KB Research

Research at the National Library of the Netherlands

Tag: KB Research Lab (page 2 of 2)

What’s happening with our digitised newspapers?

The KB has about 10 million digitised newspaper pages, ranging from 1650 until 1995. We negotiated rights to make these pages available for research and this has happened more and more over the past years. However, we thought that many of these projects might be interested in knowing what others are doing and we wanted to provide a networking opportunity for them to share their results. This is why we organised a newspapers symposium focusing on the digitised newspapers of the KB, which was a great success!

Prof. dr. Huub Wijfjes (RUG/UvA) showing word clouds used in his research.

Prof. dr. Huub Wijfjes (RUG/UvA) showing word clouds used in his research.

Continue reading

How to maximise usage of digital collections

Libraries want to understand the researchers who use their digital collections and researchers want to understand the nature of these collections better. The seminar ‘Mining digital repositories’ brought them together at the Dutch Koninklijke Bibliotheek (KB) on 10-11 April, 2014, to discuss both the good and the bad of working with digitised collections – especially newspapers. And to look ahead at what a ‘digital utopia’ might look like. One easy point to agree on: it would be a world with less restrictive copyright laws. And a world where digital ‘portals’ are transformed into ‘platforms’ where researchers can freely ‘tinker’ with the digital data. – Report & photographs by Inge Angevaare, KB.

Mining Digital Repositories Conference 2014

Hans-Jorg Lieder of the Berlin State Library (front left) is given an especially warm welcome by conference chair Toine Pieters (Utrecht), ‘because he was the only guy in Germany who would share his data with us in the Biland project.’

Libraries and researchers: a changing relationship

‘A lot has changed in recent years,’ Arjan van Hessen of the University of Twente and the CLARIN project told me. ‘Ten years ago someone might have suggested that perhaps we should talk to the KB. Now we are practically in bed together.’

But each relationship has its difficult moments. Researchers are not happy when they discover gaps in the data on offer, such as missing issues or volumes of newspapers. Or incomprehensible transcriptions of texts because of inadequate OCR (optical character recognition). Conference organisers Toine Pieters and Jaap Verheul (University of Utrecht) invited Hans-Jorg Lieder of the Berlin State Library to explain why he ‘could not give researchers everything everywhere today’.

Lieder & Thomas: ‘Digitising newspapers is difficult’

Both Deborah Thomas of the Library of Congress and Hans-Jorg Lieder stressed how complicated it is to digitise historical newspapers. ‘OCR does not recognise the layout in columns, or the “continued on page 5”. Plus the originals are often in a bad state – brittle and sometimes torn paper, or they are bound in such a way that text is lost in the middle. And there are all these different fonts, e.g., Gothic script in German, and the well-known long-s/f confusion.’ Lieder provided the ultimate proof of how difficult digitising newspapers is: ‘Google only digitises books, they don’t touch newspapers.’

Mining Digital Repositories Damaged Newspapers

Thomas: ‘The stuff we are digitising is often damaged’

Another thing researchers should be aware of: ‘Texts are liquid things. Libraries enrich and annotate texts, versions may differ.’ Libraries do their best to connect and cluster collections of newspapers (e.g., in the Europeana Newspapers), but ‘the truth of the matter is that most newspapers collections are still analogue; at this moment we have only bits and pieces in digital form, and there is a lot of bad OCR.’ There is no question that libraries are working on improving the situation, but funding is always a problem. And the choices to be made with bad OCR are sometimes difficult: Should we manually correct it all, or maybe retype it, or maybe even wait a couple of years for OCR technology to improve?’

Mining Digital Repositories Conference Claeyssens Van Hessen Kenter

Librarians and researchers discuss what is possible and what not. From the left, Steven Claeyssens, KB Data Services, Arjan van Hessen, CLARIN, and Tom Kenter, Translantis.

Researchers: how to mine for meaning

Researchers themselves are debating how they can fit these new digital resources into their academic work. Obviously, being able to search millions of newspaper pages from different countries in a matter of days opens up a lot of new research possibilities. Conference organisers Toine Pieters and Jaap Verheul (University of Utrecht) are both involved in the HERA Translantis project which is taking a break from traditional ‘national’ historical research by looking at transnational influences of so-called ‘reference cultures’:

Mining digital repositories - Definition of reference cultures

Definition of Reference Cultures in the Translantis project which mines digital newspaper collections

In the 17th century the Dutch Republic was such a reference culture. In the 20th century the United States developed into a reference culture and Translantis digs deep into the digital newspaper archives of the Netherlands, the UK, Belgium and Germany to try and find out how the United States is depicted in public discourse:

Mining Digital Repositories Jaap Verheul Translantis

Jaap Verheul (Translantis) shows how the US is depicted in Dutch newspapers

Joris van Eijnatten introduced another transnational HERA project, ASYMENC, which is exploring cultural aspects of European identity with digital humanities methodologies.

All of this sounds straightforward enough, but researchers themselves have yet to develop a scholarly culture around the new resources:

  • What type of research questions do the digital collections allow? Are these new questions or just old questions to be researched in a new way?
  • What is scientific ‘proof’ if the collections you mine have big gaps and faulty OCR?
  • How to interpret the findings? You can search words and combinations of words in digital repositories, but how can you assess what the words mean? Meanings change over time. Also: how can you distinguish between irony and seriousness?
  • How do you know that a repository is trustworthy?
  • How to deal with language barriers in transnational research? Mere translations of concepts do not reflect the sentiment behind the words.
  • How can we analyse what newspapers do not discuss (also known as the ‘Voldemort’ phenomenon)?
  • How sustainable is digital content? Long-term storage of digital objects is uncertain and expensive. (Microfilms are much easier to keep, but then again, they do not allow for text mining …)
  • How do available tools influence research questions?
  • Researchers need a better understanding of text mining per se.

Some humanities scholars have yet to be convinced of the need to go digital

Rens Bod, Director of the Dutch Centre for Digital Humanities enthusiastically presented his ideas about the value of parsing (analysing parts of speech) for uncovering deep patterns in digital repositories. If you want to know more: Bod recently published a book about it.

Rens Bod

Professor Rens Bod: ‘At the University of Amsterdam we offer a free course in working with digital data.’

But in the context of this blog, his remarks about the lack of big data awareness and competencies among many humanities scholars, including young students, was perhaps more striking. The University of Amsterdam offers a crash course in working with digital data to bridge the gap. The one-week, free course, deals with all aspects of working with data, from ‘gathering data’ to ‘cooking data’.

As the scholarly dimensions of working with big data are not this blogger’s expertise, I will not delve into these further but gladly refer you to an article Toine Pieters and Jaap Verheul are writing about the scholarly outcomes of the conference [I will insert a link when it becomes available].

Mining Digital Repositories Jaap Verheul Toine Pieters

Conference hosts Jaap Verheul (left) and Toine Pieters taking analogue notes for their article on Mining Digital Repositories. And just in case you wonder: the meeting rooms are probably the last rooms in the KB to be migrated to Windows 7

More data providers: the ‘bad’ guys in the room

It was the commercial data providers in the room themselves that spoke of ‘bad guys’ or ‘bogey man’ – an image both Ray Abruzzi of Cengage Learning/Gale and Elaine Collins of DC Thomson Family History were hoping to at least soften a bit. Both companies provide huge quantities of digitised material. And, yes, they are in it for the money, which would account for their bogeyman image. But, they both stressed, everybody benefits from their efforts:

Value proposition of DC Thomson Family History

Value proposition of DC Thomson Family History

Cengage Learning is putting 25-30 million pages online annually. Thomson is digitising 750 million (!) newspaper & periodical pages for the British Library. Collins: ‘We take the risk, we do all the work, in exchange for certain rights.’ If you want to access the archive, you have to pay.

In and of itself, this is quite understandable. Public funding just doesn’t cut it when you are talking billions of pages. Both the KB’s Hans Jansen and Rens Bod (U. of Amsterdam) stressed the need for public/private partnerships in digitisation projects.

And yet.

Elaine Collins readily admitted that researchers ‘are not our most lucrative stakeholders’; that most of Thomson’s revenue comes from genealogists and the general public. So why not give digital humanities scholars free access to their resources for research purposes, if need be under the strictest conditions that the information does not go anywhere else? Both Abruzzi and Collins admitted that such restricted access is difficult to organise. ‘And once the data are out there, our entire investment is gone.’

Libraries to mediate access?

Perhaps, Ray Abruzzi allowed, access to certain types of data, e.g., metadata, could be allowed under certain conditions, but, he stressed, individual scholars who apply to Cengage for access do not stand a chance. Their requests for data are far too varied for Cengage to have any kind of business proposition. And there is the trust issue. Abruzzi recommended that researchers turn to libraries to mediate access to certain content. If libraries give certain guarantees, then perhaps …

Mining Digital Repositories Toine Pieters

You think OCR is difficult to read? Try human handwriting!

What do researchers want from libraries?

More data, of course, including more contemporary data (… ah, but copyright …)

And better quality OCR, please.

What if libraries have to choose between quality and quantity?  That is when things get tricky, because the answer would depend on the researcher you question. Some may choose quantity, others quality.

Should libraries build tools for analysing content? The researchers in the room seemed to agree that libraries should concentrate on data rather than tools. Tools are very temporary, and researchers often need to build the tools around their specific research questions.

But it would be nice if libraries started allowing users to upload enrichments to the content, such as better OCR transcriptions and/or metadata.

Mining Digital Repositories 2014

Researchers and libraries discussing what is desirable and what is possible. In the front row, from the left, Irene Haslinger (KB), Julia Noordegraaf (U. of Amsterdam), Toine Pieters (Utrecht), Hans Jansen (KB); further down the front row James Baker (British Library) and Ulrich Tiedau (UCL). Behind the table Jaap Verheul (Utrecht) and Deborah Thomas (Library of Congress).

And there is one more urgent request: that libraries become more transparent in what is in their collections and what is not. And be more open about the quality of the OCR in the collections. Take, e.g., the new Dutch national search service Delpher. A great project, but scholars must know exactly what’s in it and what’s not for their findings to have any meaning. And for scientific validity they must be able to reconstruct such information in retrospect. So a full historical overview of what is being added at what time would be a valuable addition to Delpher. (I shall personally communicate this request to the Delpher people, who are, I may add, working very hard to implement user requests).

American newspapers

Deborah Thomas of the US Library of Congress: ‘This digital age is a bit like the American Wild West. It is a frontier with lots of opportunities and hopes for striking it rich. And maybe it is a bit unruly.’

New to the library: labs for researchers

Deborah Thomas of the Library of Congress made no bones about her organisation’s strategy towards researchers: We put out the content, and you do with it whatever you want. In addition to API’s (Application Protocol Interfaces), the Library is also allowing for downloads of bulk content. The basic content is available free of charge, but additional metadata levels may come at a price.

The British Library (BL) is taking a more active approach. The BL’s James Baker explained how the BL is trying to bridge the gap between researchers and content by providing special labs for researchers. As I (unfortunately!) missed that parallel session, let me mention the KB’s own efforts to set up a KB lab where researchers are invited to experiment with KB data making use of open source tools. The lab is still in its ‘pre-beta phase’ as Hildelies Balk of the KB explained. If you want the full story, by all means attend the Digital Humanities Benelux Conference in the Hague on 12-13 June, where Steven Claeyssens and Clemens Neudecker of the KB are scheduled to launch the beta-version of the platform. Here is a sneak preview of the lab, a scansion machine built by KB Data Services in collaboration with phonologist Marc van Oostendorp (audio in Dutch):

https://www.youtube.com/watch?v=FcTufco9P3A

Europeana: the aggregator

“Portals are for visiting; platforms are for building on.”

Another effort by libraries to facilitate transnational research is the aggregation of their content in Europeana, especially Europeana Newspapers. For the time being the metadata are being aggregated, but in Alistair Dunning‘s vision, Europeana will grow from an end-user portal into a data brain, a cloud platform that will include the content and allow for metadata enrichment:

Alistair Dunning: 'Europeana must grow into

Alistair Dunning: ‘Europeana must grow into a data brain to bring disparate data sets together.’

Dunning's vision of Europeana in the future

Dunning’s vision of Europeana 3.0

Dunning also indicated that Europeana might develop brokerage services to clear content for non-commercial purposes. In a recent interview Toine Pieters said that researchers would welcome Europeana to take such a role, ‘because individual researchers should not be bothered with all these access/copyright issues.’ In the United States, the Library of Congress is not contemplating a move in that direction, Deborah Thomas told her audience. ‘It is not our mission to negotiate with publishers.’ And recent ‘Mickey Mouse’ legislation, said to have been inspired by Disney interests, seems to be leading to less rather than more access.

Dreaming of digital utopias

What would a digital utopia look like for the conference attendees? Jaap Verheul invited his guests to dream of what they would do if they were granted, say, €100 million to spend as they pleased.

Deborah Thomas of the Library of Congress would put her money into partnerships with commercial companies to digitise more material, especially the post-1922 stuff (less restrictive copyright laws being part and parcel of the dream). And she would build facilities for uploading enrichments to the data.

James Baker of the British Library would put his money into the labs for researchers.

Researcher Julia Noordegraaf of the University of Amsterdam (heritage and digital culture) would rather put the money towards improving OCR quality.

Joris van Eijnatten’s dream took the Europeana plans a few steps further. His dream would be of a ‘Globiana 5.0’ – a worldwide, transnational repository filled with material in standardised formats, connected to bilingual and multilingual dictionaries and researched by a network of multilingual, big data-savvy researchers. In this context, he suggested that ‘Google-like companies might not be such a bad thing’ in terms of sustainability and standardisation.

Joris van Eijnatten

Joris van Eijnatten: ‘Perhaps – and this is a personal observation – Google-like companies are not such a bad thing after all in terms of sustainability and standardisation of formats.’

At the end of the two-day workshop, perhaps not all of the ambitious agenda had been covered. But, then again, nobody had expected that.

Agenda for Mining Digital Repositories 2014

Mining Digital Repositories 2014 – the ambitious agenda

The trick is for providers and researchers to keep talking and conquer this ‘unruly’ Wild West of digital humanities bit by bit, step by step.

And, by all means, allow researchers to ‘tinker’ with the data. Verheul: ‘There is a certain serendipity in working with big data that allows for playfulness.’

See also:

 

KB at DH2013

So, how do you summarise a 4-day conference with 159 papers, 52 posters, 13 workshops and 9 panels in one blogpost? You don’t… But I am going to try anyway!

I had the pleasure to attend and present a poster at the DH2013 conference this year, which took place two weeks ago in Lincoln, Nebraska. It was my first time attending the event and I was not disappointed! After a 14 hour trip (and a good night’s sleep) I started off my DH2013 experience with a wonderful workshop about Voyant, a web-based reading and analysis environment for digital text. All the material from the workshop is available online: http://hermeneuti.ca/workshop/dh13.

DH2013 logo

DH2013 logo

After some introductions of the people there, but also the tool, we formed groups to discuss how Voyant would be of use in our work. I was happy to see that we had quite a big group of librarians there, so of course we discussed how we could either show our own data to the users, but also how we can introduce the tool to students or professors for the university libraries. I’d love to see what our data looks like, so that’s a nice task for the coming months!

The main part of the conference started on my second day and I mainly spent my hours in the various short paper sessions that were held in the conference hotel. There were five papers in a sessions, each with a similar topic. And, being a text junkie, I listened to a lot of text analysis and stylometry, but I also found the time to visit papers on how to best serve researchers with tools, environments and other helpful equipment.

Some of my highlights included the paper of Anna Jobin and Fredric Kaplan on Google’s adwords lexicon, which not only consists of expensive and cheap words, but also of misspelled and non-existents words. The team at the DHLab of EPFL is undertaking a case study on the linguistic effects of autocompletion alghoritms and keyword bidding.

During the poster presentation

During the poster presentation

Another paper that stuck with me was that of our colleague from the British Library, Nora McGregor, who introduced their Digital Scholarship Training Programme. The BL has set up a training programme, consisting of 15 courses, all about anything digital. Their (not so digital?) curators can take a class on, for example, HTML, metadata formats, the BL’s digital collections or linked data. These low level courses can be taken by everyone in the library with an interest in the digital world.

Most of the people took about three courses, but there was also someone who completed the entire programme. Food for thought here at the KB! We have, what we call, kennis sessies (knowledge exchange sessions) and also offer some excellent courses in-house on copyright and digital preservation, but we have never looked at these as an entire course load. Perhaps we should!

KB and BL Poster

KB and BL Poster

And then the absolute top highlight of my DH2013 experience, our poster! The organisation arranged for a very nice and spacious room where we had our poster up on our own board, giving us the opportunity to have a big crowd of people around us. Luckily, this was also the case for us! I spoke to many people about our data (there actually ís an interest in Dutch data in the US!) and about what they would expect from a national library like ours or the BL. Want to leave your feedback as well? Please fill out our survey and help us improve!

So, with this blogpost I have NOT done justice to the conference at all, but have given you a very short overview of some of my wonderful experiences while in Nebraska. Please do read other posts to learn about the rest of the talks, such as the wonderful keynotes by David Ferriero, Willard McCarthy, and my favourite, Isabel Galina.

I met many very interesting people while there and was happy to find out I was accompanied by a lot of librarians. However, most of them were from university libraries. So, national librarians, we need you! Want to share experiences on how we work with digital humanists? Please do get in touch! And researchers? I just want to mention our survey once more.

Digital Humanities at the National library

About two months ago the Journal for Library Associations published an issue completely about Digital Humanities in libraries. Enthusiastically I printed all the open access articles (I know, not very nature conscious of me…) and put them on my desk. As it often goes with papers on desks, they’ve been lying there ever since. This changed this morning as I took out the stack and started reading them. And I loved it! Article after article I took out my highlighter and marked sentences and paragraphs that sounded too familiar to me, working in a research library with an interest in Digital Humanities.

The KB has started to look at Digital Humanities (DH) as a topic not so long ago, but has been involved with DH related projects for quite some time, although there were not called DH at the time. Examples are the CATCH projects that started in 2004, but since the beginning of our digitising-days our material is used by a variety of people and institutions. However, the KB is special in a way when it comes to doing DH research. We are the national library of the Netherlands and are thus not connected to a specific university or research institution. This means that we do not employ our own researchers. We do have a Research department, but most of the people here do not dive into our content, but do research to ensure the public can do this.

catchplus

The continuation of CATCH, CATCHPlus, where prototypes are converted to reliable tools

Although the JLA I read only discusses university libraries, and their associated researchers, this does not mean that the articles from for example Miriam Posner or Bethany Nowviskie are not relevant for us. The, often mentioned, lack of flexibility that is apparently inherent to a library also exist here and the desire to only publish something once it is perfect is something I too can relate to. Working with digitised material is never perfect. The software is not perfect, so how can the outcomes be? Nonetheless, the KB chose to show these imperfections in our OCR by opening up the texts to the public, including all mistakes and an estimate of accuracy.

kbnewspaper

A KB newspaper article. The OCR quality of this article is estimated at 84,9% character accuracy.

Being a research institute, with a large digital corpus that we are more than happy to share, without our own researchers (apart from the occasional research fellow), the KB not only faces the challenges of the university libraries as mentioned by Miriam Posner in her article (i.e. inflexibility, lack of time, authority, and incentive, overcautionesness, etc.), but I believe another crucial element can be added to this list: No affiliated researchers. Until not so long ago, when a researcher wanted to use (a section) of our digitised sets he/she would find someone from within the library who could help them get it. There was no official route to obtain the data or one contact person for a specific set, so it could be possible that people left the KB with hard disks full of images or that they tracked one of our employees down at a conference and badgered them until they got an e-mail with instructions on how to harvest a collection. Luckily, this has changed with the creation of the Data Services team.

The Data Services team are the go-to guys when it comes to our digital sets. They have taken up the responsibilities of advertising our datasets on our (unfortunately only in Dutch) website, at events and conferences and on the Dutch Open Data community, such as Open Data Nederland. We hope these efforts will lead to interesting use of our data and perhaps even some enrichments that we might implement in the future (OCR correction anyone?). But how can we be sure that our data does indeed gets used and that we reach the people who might be interested? And how do we know if our methods is in fact what they are looking for?

photo-1

The KB at the CLIN2013 conference.

This issue is one that I would imagine is easier to solve when you can simply walk to the other side of the building, knock on some doors and talk to researchers of whom you know their interests, because they teach Data Mining at your university. Unfortunately, we are not in that position, apart from the people who have asked for our data and those that will come to our (currently a work in progress) KB Lab. Now that we have the  instructions to harvest our sets available on the website, less and less people will probably be doing this, leaving us more in the dark about what interesting things are happening with the digitised Early Dutch Books Online or the ANP radio bulletins.

So, how do we get and stay in touch with interested parties that might contribute to the enrichment of our collections? How can we be sure that what we are doing is in fact what researchers need? How much can and do we want to adapt our methods to fit the need of researchers? For example, do we want to offer all possible data formats if there is a demand for it or is that something that the scholars might be able to tackle themselves? (Solutions and ultimate answers of course always welcome in the comment section below!)

We are undertaking several activities to try to find answers and also our place in the wonderful world of Digital Humanities. The establishment of our own KB Lab, where we will to work with scholars who wish to do something with our data, is one such activity. Another is the poster session that we will present together with the BL Labs project at the DH2013 conference this summer. Our aim there is to talk to different researchers about our collections and their ways of working. What types of collections they would like, what data format they would love to see, but also what they would like to do with our data. So, if you’re around in Nebraska, please come and find me at the posters and let’s talk this through!

What Do Scholars Want? British Library Labs launched

On 25 March 2013 the BL launched their Labs-project. As KB Research is also setting up a Lab we follow whatever happens at the BL in this area with keen interest.

Image

The main objective for BL in the Labs is to engage with users of the digital collections, says Aly Conteh, who heads the digital research and curator team at BL; ‘Humanities researchers are now able to work with new types of resources, using new technologies, and the BL wants to understand what is required from us’. The scholarly  landscape is in transformation and will continue to change. Libraries must change with this, not only in their services but also in the capabilities of their staff. As all curators in the BL are to be digital curators, a training program has been set up to take curators through a new digital scholarship curriculum, from text mining on large datasets to the use of social media. The BL aims to develop new ways of working with scholars– but first they to need to know what it is these scholars want.

The Labs provide the following:

  • A wiki space where scholars in the humanities and developers can meet
  • Access to available collections
  • Developer support for research in the digital collections
  • Opportunities for developers to make tools or apps on the digital collections
  • Hackathons and workshops

Competition

The Labs are kicked off with a competition for projects that explore the BL resources – there’s 3.000 GBP plus a summer residency at BL for the researcher and/or developer with the best project idea. The BL is looking for cross collection search/analysis and the use of novel techniques. ‘The best idea’, says recently appointed Labs Manager Mahendra Mahey, ‘is the one that also helps the BL learn how to support scholars and developers’ .

The launch was a low key affair, mainly testing the water with the digital humanities community. And a very sensible thing to do too – whatever you build without involving this very intelligent and discriminating crowd will not be used. There were thirty to forty people from organisations like the Open Knowledge Foundation; partner institutions like the BBC, and UK digital humanities groups at universities like Kings College, UCL and University of Hertfordshire. We were shown examples of Digital Humanities projects, BL content and tools and techniques for working with datasets.

A few lines of code

All presentations will come online in the next days I expect so I will not bother to repeat them here. I just wish to finish with the Do’s and Don’ts  learned from this launch:

  1. Involve users before , during and always in everything you do. Whatever you think of without them, you might as well not think of- it will not be used. Very wisely, BL has formed an advisory board of partner institutions and leading figures in the digital humanities to help them shape the lab
  2. There was feedback from the friendly but critical crowd on all details of the plan, and most of it was very relevant. The best one: on top of offering an overview of collections, make available to us a dataset of ten pages per collection, with available metadata, OCR etc – so we can judge the quality of the material before proposing any research on this
  3. Do not bother to develop too many (or any?) tools or services yourself. Tony Hirst of The Open university  gave a dazzling overview of tools and techniques that are already out there – you just need ‘few lines of code’ to connect this to your database with content you have picked up from BL
  4. To help researchers fit tools to the data , to write these ‘few lines of code’  , make development capacity available in the lab for your users
  5. Partner up with other content holders to foster cross collection research.

It was an inspiring day in snowy London – cannot wait until we have something to show!

The speakers at launch were: @pmgooding @marcgalexander @MappingMetaphor @psychemedia @DigiPalProject @noeL_maS @pj_webster

Newer posts

© 2018 KB Research

Theme by Anders NorenUp ↑