KB Research

Research at the National Library of the Netherlands

Page 10 of 13

10 Tips for making your OCR project succeed

(reblogged from http://www.digitisation.eu/community/blog/article/article/10-tips-for-making-your-ocr-project-succeed/)

This year in November, it has been exactly 10 years that I have been more or less involved with digital libraries and OCR. In fact, my first encounter with OCR even predates the digital library: during my student days, one of my fellow students was blind, and I was helping him out with his studies by scanning and OCR-ing the papers he needed, so their contents could be read out to him using Text2Speech software or printed on a braille display. Looking back, OCR technology has evolved significantly in many areas since then. Projects like MetaE and IMPACT have greatly improved the capabilities of OCR technology to recognize historical fonts, and open source tools such as Google’s Tesseract or those offered by the IMPACT Centre of Competence are getting closer and closer to the functionalities and success rates offered by commercial products.

Accordingly, I would like to take this opportunity to present you some thoughts and recommendations that I’ve derived from my personal experience of 10+ years with OCR processing.

A final caveat: while this is a very interesting discussion, I will not say a single word here about whether to perform OCR as an in-house activity or via out-sourcing. My general assumption is that below considerations can provide useful information for both scenarios.

1.    Know your material

The more you know about the material / collection you are aiming to OCR, the better. Some characteristics are essential for the configuration of the OCR, like e.g. the language of a document and the fonts (Antiqua, Gothic, Cyrillic, etc.) present. While such information is typically not available in library catalogues, sending documents in French language to an OCR engine configured to recognize English will yield equally poor results as trying to OCR a Gothic typeface with Antiqua settings.

Fortunately there are some helpful tools available – e.g. Apache Tika can detect the language of a document quite reliably. You may consider running such or similar characterization software in a pre-processing step to gather additional information about the content for a more fine-grained configuration of the OCR software.

Some more features in the running text the presence and frequency of which could influence your OCR setup are: tables and illustrations, paragraphs with rotated text, handwritten annotations, foldouts.

2.    Capture high quality – INPUT

Once you are ready to proceed to the image capture step it is important to think about how to set this up. While recent experiments have shown that (on simple documents) there is no apparent loss in recognition quality from using e.g. compressed JPEG images for OCR, my recommendation still remains to scan with the highest optical resolution (typically 300 or 400 ppi) and store the result in an uncompressed format like TIFF or PNG (or even the RAW data directly from the scanner).

While this may result in huge files and storage costs (btw, did you know that the cost per GB of hard drive space drop by 48% every year?), keep in mind that any form of post-processing or compression does essentially reduce the amount of information available in the image for subsequent processing – and it turns out that OCR engines are becoming more and more sophisticated in using this information (e.g. colour) to improve recognition. However, once gone, this information can never be retrieved again without rescanning. If you binarize (=convert to black-and-white) your images immediately after scanning, you won’t be able to leverage the benefits of the next-generation OCR system that requires greyscale or colour documents.

It may also be worthwhile mentioning that while this has never been made very explicit, the classifiers in many OCR engines are optimized for an optical resolution of 300 ppi, and deliver the best recognition rates with documents in that particular resolution. Only in the case of very small characters (as e.g. found on large newspaper pages) can it make sense to scale the image up to 600 ppi for better OCR results.

3.    Capture high quality – OUTPUT

OCR is still a costly process – from preparation to execution, costs can easily amount to between .5 up to .50 € per page. Thus you want to make sure that you derive the most possible value from it. Don’t be satisfied with plain text only! Nowadays some form of XML with (at least) basic structuring and most importantly positional information on the level of blocks / regions, or even better line and word or sometimes even glyph level, should always be available after OCR. ALTO is one commonly used standard for representing such information in an XML format, but also TEI or other XML-based formats can be a good choice.

Not only does the coordinate information enable greatly enhanced search and display of search results (hit term highlighting), there are also many further application scenarios such as the automated generation of table of contents, the production of eBooks, the presentation on mobile devices etc. that rely heavily on structural and layout information being available from OCR processing.

4.    Manage expectations

No matter how modern and in pristine condition your documents are, or whether you use the most advanced scanning equipment and highly configured OCR software, it is quite unrealistic to expect anything more than 90 – 95 % word accuracy from automatic processing. Most of the times though you will be happy to even come anywhere near that range.

Note that most commercial OCR engines calculate error rates based on characters and not words. This can be very misleading, since users will want to search for words. Given there are only 30 errors across a single page with 3000 characters, the character error rate (30/3000, 0,01%) seems exceptionally low. But now assume the 3000 characters boil down to only around 600 words – and the 30 erroneous characters are well distributed across different words. We arrive at an actually much higher (5x) error rate (30/600, 0,05%). To make things worse, OCR engines typically report a “confidence score” in the output. This however only means that the software believes with a certain threshold to have recognized a character or word correctly or incorrectly. These “assumptions”, despite conservative, are unfortunately often found not to be true. That is why the only possible way to derive absolutely reliable OCR accuracy scores is by the use of ground truth-driven evaluation, which is expensive and cumbersome to perform.

Obviously all of this has implications on the quality of any service based on the OCR result. These issues must be made transparent to the organization, and should in all cases also be communicated to the end user.

5.    Exploit full text to the fullest

Once you derive full text from OCR processing, it can be the first stepping stone for a wide array of further enhancements of your digital collection. Life does not stop with (even good) OCR results!

Full text gives you the ability to exploit a multitude of tools for natural language processing (NLP) on the content. Named entity recognition, topic modelling, sentiment analysis, keyword extraction etc. are just a few of the possibilities to further refine and enrich the full text.

6.    Tailor the workflow

The enemy of large-scale automated processing, it can nevertheless often be worthwhile investing some more time and tailor the OCR processing flow to the characteristics of the source material. There are highly specialized modules and engines for particular pre- and post-processing tasks, and integrating these with your workflow for a very particular subset of a collection can often yield surprising improvements in the quality of the result.

7.    Use all available resources

One of the important findings of the IMPACT project was that the use of additional language technologies can boost OCR recognition by an amount than cannot realistically be expected from even major breakthroughs in pattern recognition algorithms. Especially in dealing with historical material there is a lot of spelling variation, and it gets extremely difficult for the OCR software to correctly detect these old words. Making the OCR software aware of historical spelling by supplying it with a historical dictionary or word list can deliver dramatic improvements here. In addition, new technologies can detect valid historical spelling variants and distinguish them from common OCR errors. This makes it much quicker and easier to correct those OCR mistakes while retaining the proper historical word forms (i.e. no normalization is applied).

8.    Try out different solutions

There is a surprisingly large number of OCR software available, both freely and commercially. The Succeed project compiled information about all OCR and related software tools in a huge database that you can search here.

Also quite useful in this are the IMPACT Framework and Demonstrator Platform – these tools allow you to test different solutions for OCR and related tasks online, or even combine distinct tools into comprehensive document recognition workflows and compare those using samples of the material you have to process.

9.    Consult experts

All over the world people are applying, researching and sometimes re-inventing OCR technology. The IMPACT Centre of Competence provides a great entry point to that community. eMOP is another large OCR project currently run in the US. Consult with the community to find out about others who may have done projects similar to yours in the past and who can share findings or even technology.

Finally, consider visiting one of the main conferences in the field, such as ICDAR or ICPR and look at the relevant journal publications by IAPR etc. There is also a large community of OCR and pattern recognition experts in the Biosciences, e.g. in iDigBioHackathons like for example the ones organized by Succeed can provide you with hands-on experience with the tools and technologies being available for OCR.

10.    Consider post-correction

When all other things fail and you just can’t obtain the desired accuracy using automated processing methods, post-correction is often the only possible way to increase the quality of the text to a level suitable for scientific study and text mining. There are many solutions offered to adopt OCR post-correction, from simple-to-use crowdsourcing efforts to rather specialized tools for experts. Gamification of OCR correction has also been explored by some. And as a side effect you may also learn to interact more closely with your users and understand their needs.

With this I hope to have given you some points to take into consideration when planning your next OCR project and wish you much success in doing so. If you would like to comment on any of the points mentioned or maybe share your personal experience with an OCR project, we would be very happy to hear from you!

ANADP II in Barcelona

De tweede bijeenkomst van Aligning National Approaches to Digital Preservation (ANADP) vond afgelopen week in Barcelona plaats. De eerste bijeenkomt, in Tallinn in Estonia in 2011, resulteerde in een interessante publikatie http://www.educopia.org/publications met een overzicht van de laatste stand van zaken. En  een reeks aanbevelingen voor verdere discussie (6 in de verkorte versie en 47 in de uitgebreide versie). Om een korte indruk te geven, noem ik enkele belangrijke topics die tijdens deze drie dagen steeds opnieuw onderwerp van discussie waren tijdens de panelsessies,  de actieve werkgroepen en de lezingen van Clifford Lynch (Coalition of Networked Information) die de openingslezing hield en Adam Farquhar (British Library) , die de slotlezing verzorgde.

Clifford Lynch blikte terug wat er sinds 2011 bereikt was. Veel ‘collaboration” (“often just a lot of talking”) dat wel,  maar hij waarschuwde dat deze samenwerking ook tot onderlinge afhankelijkheid kon leiden (“interdependency”) wat een risico kan vormen: gaat het bij een ander mis, dan heb jij daar ook last van. Denk dus van te voren goed na hoe ver de samenwerking moet gaan. Een ander punt betrof de grenzen van digitale duurzaamheid. Zijn die wellicht te nauw? Zouden we ons niet over meer druk moeten maken dan alleen de veilige opslag. Bijvoorbeeld over nieuwe toegangsmogelijkheden, zoals Europeana die biedt. Over informatie die verloren gaat als wij niets doen. Over gewijzigd gebruik en een ander verwachtingspatroon bij gebruikers.  Adam Farquhar constateerde dat de meeste systemen die we nu voor digitale duurzaamheid gebruiken, zijn ingericht op opvraging van één object per keer, maar de nieuwe onderzoekers zien onze collecties als “big data” en willen onderzoek doen op grote aantallen objecten.

Niet alleen in de VS werd een “devaluation of public goods” gevoeld,  nog versterkt  door de krimpende budgetten. “Making the case for digital preservation “ zal steeds belangrijker worden.  Dat kan op verschillende manieren, niet alleen door aan te tonen wat we allemaal bewaren, maar ook door aandacht te vragen voor wat er nu (ongemerkt) verloren gaat. Weten de beleidsmakers wel wat er op het spel staat? Wie maakt zich druk om kleine, lokale krantjes? Of om het bewaren van “public broadcasting”, dat in sommige landen nauwelijks gebeurt, terwijl dat een essentiële bron voor toekomstige onderzoekers is. Welke onderzoeken zijn in de toekomst niet meer mogelijk? Als voorbeeld werd genoemd: hoe komt iemand er over 10 jaar achter hoe lang het reizen van A naar B duurde? Er zijn geen papieren spoorboekjes meer, en niemand bewaart de databases van de spoorwegen.  Het kan ons helpen dat het algemene publiek langzamerhand ook begint te beseffen dat de traditionele manier van overdracht van eigendom voor digitale objecten niet meer werkt. Je bent geen eigenaar meer van je favoriete muziek op Spotify of je favoriete boeken op je Kindle en je kunt ze niet aan je kinderen nalaten.

Tegenwerping is vaak dat we gehinderd worden door de copyrightwetgeving. Dat gaf Lynch direct toe, maar als “digital preservation community” zouden we overeenstemming moeten zien te bereiken over “some sweeping statements” , waarmee we direct de noodzaak voor wijzigingen kunnen aantonen, in plaats van ons in details te verliezen.  

En hoe tonen we aan dat we onze beloften waar maken? Kleine organisaties zeggen soms dat ze “sustainable” zijn voor een bepaalde periode, maar wie controleert dat? Lynch merkte op dat in alle branches sprake is van data verlies, maar dat dit in onze (library) wereld niet lijkt voor te komen. Meermalen is tijdens de conferentie gesproken over het opzetten van een “registry of failures”. Maar er is al een plaats waar de “horror stories” van verloren digitaal materiaal verteld kunnen worden: www.atlasofdigitaldamages.info  

“Economics, the nightmare of sustainability”, (waarbij  “sustainability” volgens Lynch maar al te vaak uitgelegd werd als “somebody else need to pay for his”) was een ander terugkerend onderwerp. Ons antwoord hierop kan gerelateerd zijn aan het feit dat we “public goods’ bewaren:  het is met publieke middelen gemaakt, men heeft er recht op om er toegang tot te houden, en het is een enorme desinvestering als dit verloren zou gaan. Aan de andere kant is het de vraag of we erg veel energie moeten steken in gedetailleerde kostenmodellen.

Luciane Duranti (InterPARES/CICRA) wees er op dat het belangrijk is om de juiste bondgenoten te vinden (de cloud storage providers bijvoorbeeld zouden ook tot onze digital preservation community moeten horen, evenals  leveranciers van systemen en services) en dat we op de juiste plekken moeten zijn, bij UNESCO en bij de conferenties van leveranciers om ons verhaal te vertellen en elkaar te versterken. Ook Chris Greer (Research Data Alliance) pleitte voor meer aansluiting bij andere disciplines en noemde als voorbeeld bio medici die nu beginnen hun collecties duurzaam op te slaan. Zij zouden kunnen profiteren van onze kennis.

Adam Farquhar vatte de trends samen in zijn slotlezing. We zullen overspoeld worden met data en toch moeten we er in slagen digitale duurzaamheid te integreren in onze dagelijkse activiteiten. Dat kunnen we niet meer alleen en zal leiden tot samenwerkingsverbanden en (gezonde) concurrentie met externe partijen die services verlenen. Onderzoekers zullen onze digitale collecties op een andere manier gebruiken, dat vergt aanpassingen in onze systemen (en m.i. mogelijk ook van het OAIS model). Maar bovenal zal de digitale duurzaamheid gemeenschap één consistente boodschap uitstralen; onze activiteiten zijn niet alleen gericht op het gebruik van het digitale materiaal in de toekomst maar ook in het heden.  

Hoe nu verder? Men vond unaniem dat het niet nodig was weer een nieuwe organisatie op te richten om “alignment” te bevorderen, er zijn vele samenwerkingsverbanden die we kunnen gebruiken om bovengenoemde punten verder uit te werken. (Kijk maar eens op  cdb.io/17laZbO  voor samenwerkingen). Wel was er behoefte aan om over enkele jaren weer op deze  strategische wijze over digital preservation te praten. Daar kijk ik naar uit!

Linked Open Data and the STCN

Author: Fernie Maas, VU University f.g.t.maas@vu.nl

In 2012, several projects were funded at VU University, the University of Amsterdam and the Royal Netherlands Academy of Arts and Sciences (KNAW), all under the umbrella of the Centre for Digital Humanities. The short and intensive research projects (approximately 9 months) combined methodologies from the traditional humanities disciplines with tools provided by computing and digital publishing. One of these projects, Innovative Strategies in a Stagnating Market. Dutch Book Trade 1660-1750, was based at VU University. Historians worked together with computer scientists of the Knowledge Representation and Reasoning Group in dealing with a specific dataset: the Short Title Catalogue, Netherlands (STCN). The project description and the research report (plus appendix) can be found here.

The project was set up within the research focus of the Golden Age cultural industries and dealt with the way early modern book producers interacted with the market, especially in times of stagnating demand and increasing competition. Book historians and economic historians, as well as scholars dealing with modern day cultural industries, have described several strategies that often occur when times are getting tough. A common denominator seems to be the constant search for a balance between on the one hand inventing new products, and on the other hand appealing to recognizable concepts. In short: differentiating, rather than revolutionizing, was (and is) seen as a key to survival. A case study was set up around the fictitious imprint of Marteau, an imprint used to cover up the provenance of controversial books. Contemporary book producers and authors had already noticed that the prohibition of, or suspicion around, certain books could spark a desire for exactly those books, eventually influencing sales.

Records in the STCN [http://bit.ly/1aHtBWs]

Records in the STCN [http://bit.ly/1aHtBWs]

The STCN is an important dataset for studying the early modern Dutch book trade and production, offering information about 200,000 titles in the period 1540-1800 (see ill. 1). The project team was provided with a bulk download of the STCN data, to work and play around with. This dataset was converted into a Resource Description Framework (RDF). RDF is a set of W3C specifications designed as a metadata data model. It is used as a conceptual description method in computing: described entities of the world are represented with nodes (e.g. “Dante Alighieri” or “The Divine Comedy”), while the relationships between these nodes are represented with edges connecting them (e.g. “Dante Alighieri” “wrote” “The Divine Comedy”). The redactiebladen (i.e. records) of the STCN have a very specific syntax of KMC’s (kenmerkcode), which contain information about author, title, place of publication, year of publication, etc. This syntax is interpreted in a program that reads the redactiebladen and gets the relevant properties about authors, titles, publishers, places, and the like out of them. Then it generates the RDF graph, linking all these entities together conveniently, and writes the results in a file. This file is exposed online, and it can be queried live by users using the query language SPARQL.

Size of titles under imprint of Marteau in the STCN [http://bit.ly/1bdzDsR]

Size of titles under imprint of Marteau in the STCN [http://bit.ly/1bdzDsR]

The RDF conversion makes it possible to query the data independently from the interface the STCN is offering. The regular interface of the STCN offers multiple ways of querying the data, especially in the ‘advanced search’ setting of the interface. However, the possibilities to filter and sort the data by using different properties are limited to a number of three fields, in combination with filtering on years of publication. A question as: in which size were publications under the Marteau-imprint mostly published, has to be broken down in several steps in the STCN, namely retrieving a list (and consequently a number) of Marteau-publications for each size used, separately. By querying the RDF-graph, this output can be retrieved in one go (see ill. 2). Also, this query structure allows for information to be visualized quite fast, for example the occurrence of Marteau-titles in the STCN, over time (see ill.3).

Titles with the fictitious imprint of Marteau in the STCN [http://bit.ly/1dJFDzq]

Titles with the fictitious imprint of Marteau in the STCN [http://bit.ly/1dJFDzq]

Publishing structured data by means of RDF is a component of the Linked Open Data approach, which means the converted STCN-dataset can be linked to other datasets. In linking the datasets, the provenance of the data stays intact, allowing for example to integrate updates of the dataset. Lists of forbidden, prohibited and condemned books (e.g. Knuttel) are in the process of being connected to the STCN, a link that could answers questions about the actual amount of Marteau-titles under investigation or suspicion. Also, combining and comparing the information about years and reasons of prohibition from the lists of forbidden books, with the information about date and place of publication in the STCN, could reconstruct a timeline of prohibition and publication, revealing a publishers’ strategy when the date of prohibition proceeds the date of publication.

The report mentioned above describes more examples, queries, and overall the rather exploratory course of the project. The pilot character of the project has allowed the team to explore the (im)possibilities of the dataset, to become aware of the importance of expert knowledge and to strengthen the collaboration between humanities researchers and computer scientists. Further research and collaboration with the STCN and book historians will be aimed at improving the infrastructure of the dataset, a better understanding of the statistical relevance of our queries, and a conceptualization of the relation between the publications, its producers, and its settings and editions.

Preserving e-journals

Last Thursday a new DPC Technolgy Watch report was presented in London. Neil Beagrie wrote Preservation, Trust and Continuing Access for e-Journals . In a lively setting at RIBA almost all major players were present with representatives from Portico, CLOCKSS , the KB International e-Depot and the Keepers registry to celebrate the launch of this publication and to discuss a variety of challenges and complexities related to preserving this material.

The DPC report gives a good overview of the current state of affairs, the terminology used in this area, the way organizations acquire e-journals (either directly from the publishers or via web harvesting the publisher sites)  and the reasons why organisations like the above mentioned are undertaking this task. E-journals are seen as the basic for scholarly communication. But the publishing model has changed the situation for libraries: instead of having the paper copies on the stacks, they need  “preserving a connection”  – this phrase is from Peter Burnhill- . This is what most subscribing organisations do: they don’t own the content, only the right to distribute the subscription to their members . To avoid loss of this material, one should start preserving the collection and negotiate with publishers  the rights to preserve this. Six use cases illustrate the challenges in preserving this material and they are not so much technical challenges as well as “organisational challenges”, like publishers ceasing operation or transferring part of their collection to another publisher without notifying. One chapter is about Trust, and in this case it is not about the trust in the sense of one repository certified by the  ISO 16363 standard for Trustworthy Repositories. But it is more about how to trust that these e-journals in general will be available in the future. The total sum of the participating and in future participating organisations that preserve e-journals should lead to trusting them to have a complete set that is accessible for the community.

In contrast to websites, where nobody expects to preserve the whole Word Wide Web, with e-journals we strive ‘ to have them all’, at least to preserve all e-journals that are relevant for the scholarly communication.  To monitor this, the KEEPERS registry is there to show us who is preserving which e-journal. In his talk Peter Burnhill tried on the hand to be optimistic about the registry but showed on the other hand that we are not there. Although the e-journals of big publishers like Elsevier and Springer are preserved for example by Portico and the International e-Depot of the KB , they represent only a small part of the total. It is far more difficult to collect the rest of the e-journals, the  “long tail”, as these are often called:  small publishers with only a few e-journals .  Collecting these is costly. One need to search for them, negotiate the terms with each publisher individually and design an ingest flow, which is as time consuming for one small publisher as it is for big publishers.  Some statistics here,  from the 100.000 serials with an e-ISSN, only 21.000 are mentioned in the Keepers registry. So 79.000 are in danger, not to mention the amount without a e-ISSN (more statistics in Burnhills blog).

For preservation the challenge lies also in the technical developments around e-journals, what is exactly the “digital object” ? This topic is less represented in the DPC Tech Watch Report, but a growing problem for collecting organisations.  The time lies behind us that a publication was simply a pdf article. Nowadays it is often accompanied by supplemental material (this can still be seen as part of the article) and “context information”, like websites, altmetrics, data etc. Can this be seen as part of the object and should this also be preserved? The same discussion takes place related to “enhanced publications”. And this is different from the analogue world, where no one expected a library to preserve all the literature referred to in the footnotes of scientific publications! Preserving organisations will need to publish their policies in this respect, to manage the expectations of their user community.

Beagrie writes that “ This makes e-journals one of the most dynamic and challenging areas of digital preservation” . But how about e-books and websites, are they less challenging? Let’s not categorize the objects to preserve (“who is doing the toughest job”), time will show that all digital genres will offer us similar challenges!

The KB, Big data and digital humanities at the kick off of the Dutch weekend of Science

The KB, Big data and digital humanities at the kick off of the Dutch weekend of  Science

The KB gave a presentation at  the Science dinner, the official kick off of the Dutch weekend of Science. Main theme of the walking dinner was digital treasures.

The Science dinner at the Van Nelle fabriek

The Science dinner at the Van Nelle fabriek

In between courses there were presentations which all related to this theme.  The first presentation was delivered by the KB.

The future of the KB is digital. Material is being digitized at a fast pace and important progress is made in the area of digital services. The aim is to increase the outreach and actively encourage the use of the rich KB collection.

To show what can be done with all this new data the KB invited three guests to give their vision on the use of big data in their field of work:

What is their relationship with Big Data and Digital Humanities? How do they see the future of  Digital Humanities and the use of Big data? What fascinates them when it comes to new possibilities?

To illustrate their relationship with Big data  introductory films have been made:

(English subtitles available by clicking  the Watch on Youtube button)

“For heritage research Big Data is a whole new and exciting field”

Julia Noordegraaf, Professor of Heritage and Digital Culture, University of Amsterdam

 

 “Science asks the question: what is knowledge? Art approaches this theme poetically by speculating and creating things.” Geert Mul, Media artist

“I have a love-hate relationship with the use of computers for language research” Professor Marc van Oostendorp of  Leiden University and the first digital humanities fellow of the KB

Save our preservation tool kit!

Author: Barbara Sierman
Originally posted on: http://digitalpreservation.nl/seeds/save-our-preservation-tool-kit/

Jan Luyken Tea and coffy tool kit. Courtesey Rijksmuseum, Netherlands

The recent US Government shut down should make all people involved in digital preservation thinking, if not worrying. I gave some feedback in Simon Tanners blog post  , but the weekend helped to ponder a bit more about this topic.

We have always said, digital preservation is an international activity and we act like that, by having international collaboration in various areas. Sometimes one organisation starts a very good initiative and we all like to make use of the results, like PRONOM (TNA), FITS (Harvard), JHOVE, OPF, NDSA, DCC, DPC,  PREMIS  at the Library of Congress. Oops… due to the US government shut down this one was no longer available via the well known URL. Although the LoC website has recently be restored (5-10-2013), many more sites are still not available, like for example data.gov . So we could say that the digital preservation community is affected by the US Government shut down, and not only because we can’t have our regular meetings with the preservation people of the Library of Congress.

Have we been naïve as digital preservationists? This is not the first time the US Government shuts down, it also happened in  1995 and 1996 and before. But at that time the web was less influential on our daily activities and we were less dependent of it. Things have changed and we work with the web all day. But we might have been a little bit naïve in expecting things to be there, while our daily job is based on the expectation that things will not always be there. We try to save things. But we don’t have a rescue plan for the information we are dependent on in our processing activities. We might need registries at ingest and transformations and reference works when doing risk assessment of file formats and new object types. But we don’t have an overview of these vital sources that together make our digital preservation tool kit: standards,registries, software, reference works etc. All things that are only accessible from one place are principally in danger – the same rule we’ll apply for our preserved digital collections. What happened in the US can also happen somewhere else.

I would suggest to create an overview of our digital preservation tool kit, the vital elements without which we cannot work,  and be reassured that their availability will no longer be restricted to one place. They should be present on various mirror sites. This activity really requires collaboration, but we already have the infrastructure for that. It is now important to make use of it. We should see this US Government shut down as a wake up call, it is not too late yet.

KB director in This Week in Libraries TWIL #103: The European Library

Watch KB director general Bas Savenije on This Week in Libraries talk on how KB, a national library, will integrate its infrastructure with that of public libraries. Also good stuff on why Europe needs to work together  in The European Library : ‘by working together on developing tools and services you can all share, you free up efforts in your library for other services you can offer your users!’

 

Presenting European Historic Newspapers Online

As was posted earlier on this blog, the KB participates in the European project Europeana Newspapers. In this project, we are working together with 17 other institutions (libraries, technical partners and networking partners) to make 18 million European newspapers pages available via Europeana on title level. Next to this, The European Library is working on a specifically built portal to also make the newspapers available as full-text. However, many of the libraries do not have OCR for their newspapers yet, which is why the project is working together with the University of Innsbruck, CCS Content Conversion Specialists GmbH from Hamburg and the KB to enrich these pages with OCR, Optical Layout Recognition (OLR), and Named Entity Recognition (NER).

Hans-Jörg Lieder

Hans-Jorg Lieder of the Berlin State Library presents the Europeana Newspapers Project at our September 2013 workshop in Amsterdam.

In June, the project had a workshop on refinement, but it was now time to discuss aggregation and presentation. This workshop took place in Amsterdam on 16 September, during The European Library Annual Event. There was a good group of people, not only from the project partners and the associated partners, but also from outside the consortium. After the project, TEL hopes to be able to also offer these institutions a chance to send in their newspapers for Europeana, so we were very happy to have them join us.

The workshop kicked off with an introduction from Marieke Willems of LIBER and Hans-Joerg Lieder of the Berlin State Library.. They were followed by Markus Muhr from TEL, who introduced the aggregation plan and the schedule for the project partners. With so many partners, it can be quite difficult to find a schedule that works well, to ensure everyone has their material sent in on time. After the aggregation, TEL will then have to do some work on the metadata to convert it to the Europeana Data Model. Markus was followed by a presentation from Channa Veldhuijsen from the KB, who unfortunately, could not be there in person. However, her elaborate presentation on usability testing provided some good insights on how to get your website to be the best it can be and how to find out what your users really think when they are browsing your site.

[slideshare id=26336264&style=border: 1px solid #CCC; border-width: 1px 1px 0; margin-bottom: 5px;&sc=no]

It was then time for Alastair Dunning from TEL to showcase the portal that they have been preparing for Europeana Newspapers. Unfortunately, the wifi connection was not up to so many visitors and only some people could follow his presentation along on their own devices. However, there were some valuable feedback points which TEL will use to improve the portal. Unfortunately, the portal is not yet available from outside, so people who missed the presentation need to wait a bit longer to be able to see and browse the European newspapers.

But what we do already can see, are some websites of partners that have already been online for some time. It was very interesting to see the different choices each partner made to showcase their collection. We heard from people from the British Library, the National and University Library of Iceland, the National and University Library of Slovenia, the National Library of Luxembourg and the National Library of the Czech Republic.

P1100058

Yves Mauer from the National Library of Luxembourg presenting their newspaper portal

The day ended with a lovely presentation by Dean Birkett of Europeana, who, partly with Channa’s notes, went to all the previously presented websites and offered comments on how to improve them. The videos he used in his talk are available on Youtube. His key points were:

  1. Make the type size large: 16px is the recommended size.
  2. Be careful of colours. Some online newspapers sites use red to highlight important information but red is normally associated with warning signals and errors in the user’s mind.
  3. Use words to indicate language choices (eg. ‘english’, ‘français’) not flags. The Spanish flag won’t necessarily be interpreted to mean ‘click here for spanish’ if the user is from Mexico.
  4. Cut down on unnecessary text. Make it easy for users to skim (eg. though the use of bullet points).

All in all, it was a very useful afternoon in which I learned a lot about what users want from a website. If you want to see more, all presentations can be found at the Slideshare account of Europeana Newspapers or join us at one of the following events:

  • Workshop on Newspapers in Europe and the Digital Agenda. British Library, London. September 29-30th, 2014.
  • National Information Days.
    • National Library of Austria. March 25-26th, 2014.
    • National Library of France. April 3rd, 2014.
    • British Library. June 9th, 2014.

1st Succeed hackathon @ KB

Throughout recent weeks, rumors spread at KB National Library of the Netherlands that there would be a party of programmers coming to the library to participate in a so-called “hackathon”. In the beginning, especially the IT department was rather curious: will we have to expect port scans being done from within the National Library’s network? Do we need to apply special security measures? Fortunately, none of that was necessary.

A “hackathon” is nothing to be afraid of, normally. On the contrary: the informal gatherings of software developers to work collaboratively on creating and improving new or existing software tools and/or data have emerged as a prominent pattern in recent years – in particular the hack4Europe series of hack days that is organized by Europeana has shown that this model can also be successfully applied in the context of cultural heritage digitization.

After that was sorted, a network switch with static IP addresses was deployed by the facilities department of the KB, thereby ensuring that participants of the event had a fast and robust internet connection at all times and allowing access to the public parts of the internet and the restricted research infrastructure of the KB at the same time – which received immediate praise from the hackers. Well done, KB!

So when the software developers from Austria, England, France, Poland, Spain and the Netherlands gathered at the KB last Thursday, everyone already knew they were indeed here to collaboratively work on one of the European projects the KB is involved in: the Succeed project. The project had called in software developers from all over Europe to participate in the 1st Succeed hackathon to work on interoperability of tools and workflows for text digitization.

There was a good mix of people from the digitization as well as digital preservation communities, with some additional Taverna expertise tossed in. While about half of the participants had participated in either Planets, IMPACT or SCAPE, the other half of them were new to the field and eager to learn about the outcomes of these projects and how Succeed will address them.

And so after some introduction followed by coffee and fruit, the 15 participants immersed straight away into the various topics that were suggested prior to the event as needing attention. And indeed, the results that were presented by the various groups after 1.5 days (but only 8 hours of effective working time) were pretty impressive…

hack
Hackers at work @ KB Succeed hackathon

The developers from INL were able to integrate some of the servlets they created in IMPACT and Namescape with the interoperability-framework – although also some bugs were uncovered while doing so. They will be fixed asap, rest assured!  Also, with the help of the PSNC digital libraries team, Bob and Jesse were able to create a small training set for Tesseract, outperforming the standard dictionary despite some problems that were found in training Tesseract version 3.02. Fortunately it was possible to apply the training to version 3.0and then run the generated classifier in Tesseract version 3.02, which is the current stable(?) release.

Even better: the colleagues from Poznań (who have a track record of successful participation at hackathons) had already done some training with Tesseract earlier and developed some supporting tools for it. Quickly Piotr created a tool description for the “cutouts” tool that automatically creates binarized clippings of characters from a source image. On the second day another feature of the cutouts application was added: creating an artificial image suitable for training Tesseract from the binarized character clippings. When finally wrapping the two operations in a Taverna workflow time eventually ran out, but given only little work remained we look forward to see the Taverna workflow for Tesseract training becoming available shortly! Certainly this is also of interest to the eMOP project in the US, in which the KB is a partner as well.

Meanwhile, another colleague from Poznań was investigating the process of creating packages for Debian-based Linux operating systems from existing (open source) tools. And despite using a laptop with OSX Mountain Lion, Tomasz managed to present a valid Debian package (including even icon and man page) – kudos! Certainly the help of Carl from the Open Planets Foundation was also partly to blame for that…next steps will include creating a change log straight off github. To be continued!

psnc
Two colleagues from PSNC-dl working on a Tesseract training workflow

Another group attending the event were the team from LITIS lab at the University of Rouen. Thierry demonstrated the newest PLaIR tools such as the newspaper segmenter capable of automatically separating articles in scanned newspaper images.  The PLaIR tools use GEDI as the encoding format, so some work was immediately invested by David to also support the PAGE format, the predominant format for document encoding used in the IMPACT tools, thereby in principle establishing interoperability between IMPACT and PLaIR applications. In addition, since the PLaIR tools are mostly already available as web services, Philippine started with creating Taverna workflows for these methods. We look forward to complement the existing IMPACT workflows with those additional modules from PLaIR!

plairScreenshot of the PLaIR system for post-correction of newspaper OCR

All this was done without requiring any help from the PRImA group at the University of Salford, Greater Manchester, who are maintaining the PAGE format and a number of tools to support it. So with some free time on his hand, Christian from PRImA instead had a deeper look at Taverna and the PAGE serialization of the recently released open source OCR evaluation tool from the University of Alicante, the technical lead of the Centre of Competence, and found it to be working quite fine. Good to finally have an open source community tool for OCR evaluation with support for PAGE – and more features shall be added soon: we’re thinking word accuracy rate, bag-of-words evaluation and more – send us your feature requests (or even better: pull request).

We were particularly glad also that some developers beyond the usual MLA community suspects have found the way to the KB on those 2 days: a team from the Leiden University Medical Centre was also attending, keen on learning how they could use the T2-Client for their purposes. Initially slowed down by some issues encountered in deploying Taverna 2 Server on a Windows machine (don’t do it!), eventually Reinout and Eelke were able to resolve it simply by using Linux instead. We hope a further collaboration of Dutch Taverna users will arise from this!

Besides all the exciting new tools and features it was good to also see some others getting their hands dirty with (essential) engineering tasks – work progressed well on several issues from the interoperability-framework’s issue tracker: support for output directories is close to being fully implemented thanks to Willem Jan, and a good start was made on future MTOM support. Also Quique from the Centre of Competence was able to improve the integration between IMPACT services and the website Demonstrator Platform.

Without the help of experienced developers Carl from the Open Planets Foundation and Sven from the Austrian National Library (who had just conducted a training event for the SCAPE project earlier in the same week in London, and quickly decided to cross the channel for yet one more workshop), this would not have been so easily possible. While Carl was helping out everywhere at once, Sven found some time to fit in a Taverna training session after lunch on Friday, which was hugely appreciated from the audience.

sven
Sven Schlarb from the Austrian National Library delivering Taverna training

After seeing all the powerful capabilities of Taverna in combination with the interoperability-framework web services and scripts in a live demo, no one needed further reassurance that it was well worth spending the time to integrate this technology and work with the interoperability-framework and it’s various components.

Everyone said they really enjoyed the event and found plenty of valuable things that they had learned and wanted to continue working with. So watch out for the next Succeed hackathon in sunny Alicante next year!

Preservation at Scale: workshop report

Digital preservation practitioners from Portico and from the National Library of The Netherlands (KB) organized a workshop on “Preservation at Scale” as part of iPres2013. This workshop aimed to articulate and, if possible, to address the practical problems institutions encounter as they collect, curate, preserve, and make content accessible at Internet scale.

Preservation at scale has entailed continual development of new infrastructure. In addition to preservation of digital documents and publications, data archives are collecting a vast amount of content which must be ingested, stored and preserved. Whether we have to deal with nuclear physics materials, social science datasets, audio and video content, or e-books and e-journals, the amount of data to be preserved is growing at a tremendous pace.

The presenters at this workshop each spoke from the experience of organizations in the digital preservation space that are wrestling with the issues introduced by large scale preservation. Each of these organizations has experienced annual increases in throughput of content, which they have had to meet, not just with technical adaptations (increases in hardware and software processing power), but often also with organizational re-definition, along with new organizational structures, processes, training, and staff development.

There were a number of broad categories addressed by the workshop speakers and participants:

  1. Technological adaptations
  2. Institutional adaptations
  3. Quality assurance at scale and across scale
  4. The scale of the long tail
  5. Economies and diseconomies of scale

Technological Adaptations
Many of the organizations represented at this workshop have gone through one or more cycles of technological expansion, adaption, and platform migration to manage the current scale of incoming content, to take advantage of new advances in both hardware and software, or to respond to changes in institutional policy with respect to commercial vendors or suppliers.

These include both optimizations and large-scale platform migrations at the Koninklijke Bibliotheek, Harvard University Library, the Data Conservancy at Johns Hopkins University, and Portico, as well as the development by the PLANETS and SCAPE projects of frameworks, tools and test beds for implementing computing-intensive digital preservation processes such as the large-scale ingestion, characterization, and migration of large (multi-terabyte) and complex data sets.

A common challenge was reaching the limits of previous-generation architectures (whether those limits are those of capacity or of the capability to handle new digital object types), with the consequent need to make large-scale migrations both of content and of metadata.

Institutional Adaptations
For many of the institutions represented at this workshop, the increasing scale of digital collections has resulted in fundamental changes to those institutions themselves, including changes to an institution’s own definition of its mission and core activities. For these institutions, a difference in degree has meant a difference in kind.

For example, the Koninklijke Bibliotheek, the British Library, and Harvard University Library have all made digital preservation a library level mandate. This shift from relegating the preservation of digital content to an organizational sub-unit to ensuring that digital preservation is an organization-wide endeavor is challenging, as it requires changing the mindsets of many in each organization. It has meant reallocation of resources from other activities. It has necessitated strategic planning and budgeting for long-term sustainability of digital assets, including digital preservation tools and frameworks – a fundamental shift from one-time, project-based funding. It has meant making choices; we cannot do everything. It has meant comprehensive review of organizational structures and procedures, and has entailed equally comprehensive training and development of new skill sets for new functions.

Quality Assurance at Scale and Across Scales
A challenge to scaling up the acquisition and ingest of content is the necessity for quality assurance of that content. Often institutions are far downstream from the creators of content. This brings along many uncertainties and quality issues. There was much discussion of how institutions define just what is “good enough,” and how those decisions are reflected in the architecture of their systems. Some organizations have decided to compromise on ingest requirements as they have scaled up, while other organizations have remained quite strict about the cleanliness of content entering their archives. As the amount of unpreserved digital content continues to grow, this question of “what is sufficient” will persist as a challenge, as will the challenge of moving QA capabilities further upstream, closer to the actual producers of data.

The Scale of the Long Tail
As more and more content is both digitized and born digital, institutions are finding they must scale for increases in both resource access requests and expectations for completeness of collections.

The number of e-journals in the world that are not preserved was a recurrent theme. The exact number of journals that are not being preserved is unknown, but some facts are:

  • 79% of the 100,000 serials with ISSN are not being known to be preserved anywhere. It is not know how many serials that do not have ISSNs are being preserved.
  • In 2012, Cornell and Columbia University Libraries (2CUL) estimated that about 85% of e-serial content is unpreserved.

This digital “dark matter” is dwarfed in scope by existing and anticipated scientific and other research data, including that generated by sensor networks and by rich multimedia content.

Economies and Diseconomies of Scale
Perhaps the most important question raised at this workshop was the question as to whether we as a community are really at scale yet? Can we yet leverage true economies of scale? David Rosenthal noted that as we centralize more and more preserved content in fewer hands, we will be able to better leverage economies of scale, but we will also be increasing risk of a single point of failure.

Next Steps
The consensus of the group seemed to be that, as a whole, the digital preservation community is not yet truly at scale. However, the organizations in the room have moved beyond a project mentality and into a service oriented mentality, and are actively seeking ways to avoid wasteful duplication of effort, and to engage in active cooperation and collaboration.

Workshop presentations and notes on each presentation are available at: https://drive.google.com/folderview?id=0B1X7I2IVBtwzcGVhWUF0TmJIUms&usp=sharing

« Older posts Newer posts »

© 2018 KB Research

Theme by Anders NorenUp ↑