KB Research

Research at the National Library of the Netherlands

Page 8 of 13

The Research Data Alliance in Amsterdam and the KB

Prof. C. Borgman Image: Inge Angevaare KB-NL

Our colleagues from DANS organized the 4th Plenary Meeting of the Research Data Alliance ( RDA: research data sharing without barriers) in Amsterdam, held this past three days. I was there, representing the KB, one of the few national libraries present. The concept that national libraries have “research data” is a concept that needs some explanation. There are repositories that collect data sets that are a result of research, often underpinning an article. DANS and 3TU are good examples of this. But there are also repositories that have “collections” to facilitate research, like sensor data, astronomical data, climate data. This is similar to what the KB offers the researchers: a vast amount of digitized historical texts and a (restricted accessible) web archive. Researchers use these sets, see for example the Webart project. With the growing attention for digital scholarship or e-Humanities, we can expect more use. And to make the process complete, the results of research done on KB collections might end up as a publication in the KB and a data set at DANS. An NCDD working group on Enhanced Publications is looking into ways to present both outputs smoothly as an integral entity to the user. In short, there are good reasons for libraries to be at RDA! The opening of the conference had several speakers from the European Commission. Both Robert Jan Smits (Director General DG Research) and Neelie Kroes, Vice president of the European Commision via video, stressed that the European Commission expects RDA to contribute to the growing importance of sharing and preserving research data, as open access to research data is a key message in the Horizon 2020 Programme. With a new cohort of EU politicians, some canvassing work to convince them of the ins and outs of this and the role of RDA will be necessary. Prof. dr. Barend Mons from Leiden University and founder of the “fair data” initiative was asked to give his views on the matter. FAIR data being: Findable, Accessible Interoperable and Re-usable, for both humans and computers. With the motto “Bringing data to Broadway” he pleaded for professionalism in data publishing by a good infrastructure for data and a rewarding system for researchers (data should have the same “status” as a publication) and for real data stewardship. Difficulties in hiring and keeping competent data scientists, for example are a barrier. Are publishers ready for data publishing or will the data end up in a black hole? Despite the trend of putting data central, he believes that there will always be “a narrative” explaining the findings (read: articles, books). To improve professional data stewardship, he pleaded to reserve 5% of research budgets to achieve the goals of FAIR data. Prof. Christine Borgman of UCLA gave an interesting talk in which she criticized some assumptions related to research data. For example data sharing: this is not common practice in every discipline and (again) as long as researchers are not rewarded for it, it will not happen. The emphasis on data might not be fair, publications are not simply “containers for data” but are arguments, supported by the data. The carefully designed process in publications (for example the order of appearance of the authors) is not even designed yet for data sets. More of this will be described in her new book, to be published by the end of the year. The rest of the work these days was done in a variety of Interest Groups (IGs) and Working Groups (WGs). The KB participates in the activities on Certification of Digital Repositories and Data Publishing (about workflows in data publishing, of interest for our (inter-) national e-Depot, about costs for data centers). All information is available from the RDA website. At the final meeting an interesting announcement was made: in December a follow up of the Riding the Wave report will be published, with the working title Harvesting the Data. Knowing the immense impact the Riding the Wave report had, this is something to look forward to. The Research Data Alliance started as a small group and has now over 2500 members, with a large range of Interest Groups and Working Groups. Time has come to streamline the activities more in order to integrate the results and to think about the sustainability of the RDA itself. The results of this process will be discussed in the next Plenary Meeting in San Diego 9-11 March 2015.

“Materials contain the seeds of their own destruction”. A preservation handbook.

Author: Barbara Sierman
Originally posted on: http://digitalpreservation.nl/seeds/materials-contain-the-seeds-of-their-own-destruction-a-preservation-handbook/

harvey-204x300Regularly I have discussions whether digital material and analogue material can be treated the same way or whether the digital aspect requires special treatments, sometimes even resulting in different working processes, staffing and policies. Quite too often this discussion takes place with participants that are either representatives of the digital or of the analogue view. The polite way of trying to understand each other by finding analogies often lead to simplified views and unsatisfying outcomes and nobody gets the wiser. Therefore I was triggered when a new digital preservation handbook exactly raised this issue by stating “This book is based on the philosophy that there are preservation principles that apply to all kinds of materials, whether digital or not.” For a preservation handbook this is a realistic perspective, as organisations have both kinds of materials. The authors present this book as the first example of ” the essential tools and principles of a preservation management programme in the 21st century – one that addresses the realities of diverse collections and materials and embraces the challenges of working with both analogue and digital collections.”

This being stated, the authors start addressing the different issues related to digital versus analogue and refer to the fact that digitization in the past led to destruction of the related physical objects by assuming that “the information” was saved in the new digital object, a debatable point of view nowadays (see Nicholson Baker’s Double Fold. Libraries and the assault on paper. 2001) . They come with a set of shared preservation principles, for both analogue and digital material.

harvey-21

Four principles describe the context and aims of preservation, amongst which the needs of the user is mentioned (a point of view we also see in the OAIS model) as well as “Preservation is the responsibility of all, from the creators of objects to the users of objects“. A set of 8 general principles focus on “collaboration”, “advocacy, “active, managed care” and the preference for actions “that address large quantities of material over actions that focus on individual objects” [ although this is highly dependent on the value of these objects I would say] . The following principle describes the key of preservation: “Understanding the structure of material is the key to understanding what preservation actions to take, as materials contain the seeds of their own destruction (inherent vice)”.

This set of Preservation principles and practices is the red line for the rest of the book, which contains a wealth of information. I can recommend this book to both the digital and the analogue preservationists, as it will contribute to mutual understanding so desperately needed! And don’t complain about the price (90 dollars) : this book might be expensive, but a one day course is more expensive and almost all the rest you want to know about digital preservation is freely available on the internet!

The preservation management handbook: a 21st-century guide for libraries, archives and museums. [Edited by] Ross Harvey and Martha R. Mahard. Rowman and Littlefield, 2014. ISBN 978-0-7591-2315-1 (also available as e-book)

Meetup “Digitaal Geheugenverlies” van de VPRO

Als vervolg op de Tegenlicht uitzending van de VPRO over “Digitaal geheugenverlies” was er afgelopen dinsdag een “meetup” in Pakhuis De Zwijger in Amsterdam, waar ruim 100 mensen op afkwamen. De teloorgang van de bibliotheek van het KIT was voor regisseur Bregtje van der Haak aanleiding om deze documentaire te maken. Hans van Harteveld, voormalig directeur van deze bibliotheek, zei aanvankelijk verbijsterd te zijn geweest over de plannen met de bibliotheek. Zijn roman “De verkwanseling van een kroonjuweel”  (Uitgeverij In de Knipscheer ISBN 978-90-6265-862-6) werd ter plekke gepresenteerd en beschrijft de hele gang van zaken.

Bart Krul (Instituut Maatschappelijke Innovatie) interviewde Bas Savenije (namens KB)  en Marcel Ras (namens de NCDD). Leidraad van hun gezamenlijk betoog was dat niet alles bewaard kan worden, dat degene die het verzamelt niet noodzakelijkerwijze ook degene is die het moet bewaren, en dat de uiteindelijke gebruiker de reden is dat we het doen. Maar weten we wat die toekomstige gebruiker wil en maken de ontwikkelingen niet een herdefinitie van kernbegrippen noodzakelijk? Is de definitie van “publicatie” niet te strikt, als daardoor beleidsnotities van Jan Pronk buiten de boot zouden vallen? Door een intensieve nationale samenwerking, zoals nagestreefd door de NCDD (de Nationale Coalitie Digitale Duurzaamheid) kunnen we in Nederland tot een evenwichtig bewaarbeleid komen, om zo het gevreesde verlies te voorkomen.

Hoe helder en bekend dit ook in de oren klinkt van de mensen in de zaal die “in het vak zitten”, mij werd vooral duidelijk dat er een enorme kloof ligt tussen wat men denkt dat we doen als erfgoedsector en wat we feitelijk doen. Begripsverwarring over digitaal geheugenverlies bleek al in de uitzending, die voor een groot deel over de verdwijning van een papieren collectie ging. Onbekendheid bij de VPRO met erfgoedsector in Nederland zorgde er voor dat voornamelijk initiatieven uit de Verenigde Staten (Brewster Khale van Internet Archive, Jason Scott van The Archive Team en ingenieurs die NASA tapes redden) in de uitzending getoond werden. En steeds weer werd de vraag gesteld: waarom moeten we alles bewaren? Het is aan ons daar een antwoord op te hebben, net als Brewster Khale, Jason Scott en Ismail Serageldin, directeur van de bibliotheek van Alexandrië.

Als erfgoedsector doen we veel, maar het is goed te weten van particuliere initiatieven om digitaal erfgoed te behouden en te gebruiken. In de zaal zaten mensen die meededen met The Archive Team, wakker geschud toen Hyves werd opgeheven. Richard Vijgen vertelde hoe hij als grafisch ontwerper gebruik maakt van de geredde website Geocities en  een nieuwe “installatie” maakte van dit materiaal door het opnieuw tot leven te brengen.

De VPRO heeft ons een goede dienst bewezen met deze documentaire en “meetup” (het is vast geen toeval dat de NRC onlangs ook over Internet Archive schreef), hoog tijd om ons eigen verhaal eraan toe te voegen!

Meetup VPRO Amsterdam

Jpylyzer software finalist voor digitale duurzaamheidsprijs

Vandaag maakte de Britse Digital Preservation Coalition de finalisten bekend die in de race zijn voor de Digital Preservation Awards 2014. Deze prijs is in 2004 in het leven geroepen om aandacht te vestigen op initiatieven die een belangrijke bijdrage leveren aan het toegankelijk houden van digitaal erfgoed.

In de categorie Research and Innovation is een op de KB door de afdeling Onderzoek ontwikkelde softwaretool genomineerd: jpylyzer. Met jpylyzer kun je op een eenvoudige manier controleren of JP2 (JPEG 2000) beeldbestanden technisch in orde zijn. Binnen de KB wordt de tool onder meer ingezet bij de kwaliteitscontrole van gedigitaliseerde boeken, kranten en tijdschriften. Jpylyzer wordt ook gebruikt door diverse internationale collega-instellingen.

Jpylyzer is deels ontwikkeld binnen het Europese project SCAPE, waarin de KB projectpartner is. De uiteindelijke winnaars worden op 17 november bekendgemaakt.

Meer informatie over de nominatie van jpylyzer is te vinden op de website van de Digital Preservation Coalition:

http://www.dpconline.org/newsroom/latest-news/1271-dpa-2014finalists

Het volgende artikel is interessant voor wie meer wil weten over jpylyzer, en waarom we zo’n tool eigenlijk nodig hebben:

http://www.kb.nl/research/kb-onderzoek-het-internationale-succes-van-de-jpylyzer-en-wat-is-dat-eigenlijk-voor-ding

Ten slotte is hier de jpylyzer homepage:

 http://openplanets.github.io/jpylyzer/

[vimeo 53693082 w=500 h=281]

Linked Open Data at the National library of the Netherlands

Authors: Theo van Veen and Sieta Neuerburg

According to the National library of the Netherlands, libraries are positioned at the very core of the Semantic Web. In the words of Hans Jansen, Head of the library’s Innovation and Development department, “Linking data is the way forward for libraries. Any cultural heritage institution that does not invest in linking data will become obsolete.”

[vimeo 36752317 w=400 h=300]

Linked Open Data video by Europeana

The National library completed several successful Linked Data projects based on metadata linking. For example, for our newspaper app Here was the news (available in Dutch only), we added latitude and longitude data to our Dutch historical newspaper articles. The app allows users to search the newspaper collection by location. Moreover, the library has added links between its own journal collection and the TV and radio recordings of the Netherlands Institute for Sound and Vision. (This feature is not yet available on our website.)

While the work on linking our metadata continues, the library’s Research department is currently involved in an effort to link named entities in our full text collections. The purpose of this project is to contribute to a fully Linked Open Data-enabled library, entailing an enhanced user experience and improved discovery based on semantic relations. To further contribute to the progress of the semantic web, we will offer our full text enrichments as open data.

Linking named entities

To achieve our objectives, we need to be able to identify and link relevant named entities in our text collections. As part of the Europeana Newspapers project, we programmed a machine-learning tool to identify named entities in our full text collections. This software will allow us to extract all named entities from our full text collections and link them to related resources and resource descriptions. The software and documentation are available on GitHub.

The Research department created an enrichment database to collect information about the named entities and to store links to external resources. We are currently linking the named entities to DBpedia, while simultaneously storing links to related resource descriptions in Freebase and VIAF. This will be further extended to other resources, such as genealogy databases. Additionally, we will develop software for ‘socially enhanced linking’, i.e. tools allowing users to validate or reject links that were obtained automatically and to create new links for resources.

Challenges

Within the project, we still face some important challenges. A first problem is the issue of incomplete external coverage. Not all named entities are covered by resource description databases such as DBpedia and Freebase. Historical entities are especially neglected, and international databases such as DBpedia frequently omit even well-known Dutch figures. Moreover, there is no single global identifier for a resource. Resource description databases – such as DBpedia, Freebase and Geonames – all use their own identifiers, resulting in single resources leading to multiple resource descriptions.

Another issue that has to be dealt with is that of intellectual property rights. Ownership issues can hinder progress towards openness. Then there are problems of textual recognition difficulties. This is mostly related to OCR issues, but it also applies to historical language variations, name variants and other types of ambiguity. And finally, manual intervention is indispensable. We will need crowdsourcing to check, validate and correct links which have been automatically generated.

Expected results

The National library expects to achieve the most important results in two areas: first, enhanced discoverability and second, data enrichments. Both are essential to keep the library relevant in the digital age.

Linked Open Data is a powerful method to link digital heritage at a national or international level, especially because the data is openly available. Relations between resources transcend organisational, national and language barriers. The identification and linking of data helps to transform library collections into (machine and human readable) information and knowledge. This will allow for much richer search and discovery opportunities.

The National library’s Senior researcher Theo van Veen sees the future of Linked Open Data as “a single worldwide resource description database which will replace most or all bibliographic thesauri, with all resources mentioned in metadata or text linking to the same single identifier”.

OCR improvement: helping and hindering researchers

Author: Tineke Koster

As I am writing this, volunteers are rekeying our 17th century newspapers articles. Optical character recognition of the gothic text type in use at the time has yielded poor results, making this part of our digital collection nearly inaccessible for full-text search. The Meertens institute, who have an excellent track record when it comes to crowdsourcing, has developed the editor (Dutch). Together with them we are working towards a full update of all newspaper issues from 1618 to 1700 that are available in our website Delpher.

Great news and, for some researchers, an eagerly awaited development. A bright future beckons in which our digital text corpus is 100% correct, just waiting to be mined for dynamic phenomena and paradigm shifts.

But we have to realize that without the proper precautions, correcting digital texts may also hinder researchers in their work. How so? These texts may have been used (browsed, mined, cited, etc.) by researchers in their earlier form. The improvement or enrichment may have consequences for the reproducibility of their research results.

For all researchers the need to reproduce research results is growing, with new guidelines due to new laws. There is also a specific group of researchers that need sustained access to older versions of digital text. The need is highest for research where the goal is to develop an algorithm and to assess its quality relative to previous versions of the same algorithm or to other algorithms. Without sustained access to older versions, these people cannot do their work.

Is it our role to provide this access? How the National Library of the Netherlands is thinking about this issue, I hope to explain in a later blogpost (soon!). Meanwhile, I would be very interested to hear your experiences. How is this subject discussed in your organization? Does your organization have a policy in place to deal with this?

National Library of the Netherlands participates in Digitisation Days, Madrid, 19-20 May

On 19 and 20 May, the National Library of the Netherlands (KB) visited the Digitisation Days which were held at the Biblioteca Nacional in Madrid. The conference was supported by the European Commission, and organised by the Support Action Centre of Competence in Digitisation (Succeed) project  and the IMPACT Centre of Competence (IMPACT CoC) with the cooperation of Biblioteca Nacional de España.

For the National Library, being a collection holder, the Succeed awards ceremony was one of the highlights of the conference, because it showed the application of technology to actual collections. The Succeed awards aim to recognise successful digitisation programmes in the field of historical texts, especially those using the latest technology.

Two prizes went to the Hill Museum and Manuscript Library and the Centre d’Études Supérieures de la Renaissance, while two Commendations of Merit were awarded to the London Metropolitan Archives/ University College London  and to Tecnilógica.

In her role of member of the IMPACT CoC executive board, the KB’s Head of Research, Hildelies Balk, took part in the ceremony and awarded the Commendation of Merit to the London Metropolitan Archives/ University College London for their Great Parchment Book project. You will find a short video about the project here.[youtube=http://www.youtube.com/watch?v=WDD2cVT7PeU]

Moreover, the KB hosted an interesting and fruitful Round table workshop on the future of research and funding in digitisation and the possible roles of Centres of Competence on 20 May. Some 30 librarians and researchers joined this workshop, and discussed the below topics:

  • What research is needed to further the development of the Digital Library?
  • How can Centres of Competence assist your research or development?
  • In digitisation, are we ready to move the focus from quantity to quality?
  • What enrichments, e.g. in Named Entity Recognition, Linked Data services, or crowdsourcing for OCR correction, would be most beneficial for digitisation?
  • What’s your take on Labs and Virtual Research Environments?
  • What would you like to do in these types of research settings?
  • What do you expect to get out of them?

The preliminary outcomes of the workshop show that the main goal for institutions is to give users unrestricted access to data. During the workshop, the participants discussed the many layered aspects of these three topics, i.e. ‘users’, ‘access’, and ‘data’. Moreover, the participants gave their view on the following questions in relation to these topics:

  • What stops us from making progress?
  • What helps us to make progress?
  • And what role could CoCs play in this?

The outcomes of the workshop have been documented and will be used as a starting point for the roadmap to further development of digitisation and the digital library, which will be produced within the Succeed project. This roadmap will serve to support the European Commission in preparing the 2014–2020 Work Programme for Research and Innovation.

 

Too early for audits?

Author: Barbara Sierman
Originally posted on: http://digitalpreservation.nl/seeds/too-early-for-audits/

I never realized that the procedure of getting to an ISO standard could take several years, but this is true for two standards related to audit and certification of trustworthy digital repositories.  Although we have the ISO 16363 standard on Audit and Certification since 2012, official audits cannot take place against this standard until the related standard Requirements for bodies providing Audit and Certification (ISO 16919) is approved, regulating the appointment of auditors. This standard, similar to the ISO 16363 compiled by the PTAB group in which I participate, was already finished a few years ago, but the ISO review procedure, especially when revisions need to be made, takes long. The latest prediction is that this summer (2014) the ISO 16919 will be approved, after which national standardization bodies can train the future (official) auditors.  How many organizations will then apply for an official certification against the ISO standard is not yet clear, but if you’re planning to do so, it might be worthwhile to have a look at the recent report of the European 4C project  Quality and trustworthiness as economic determinants in digital curation.

The 4C project (Collaboration to Clarify the Cost of Curation) is looking at the costs and benefits of digital curation. Trustworthiness is one of the “economic determinants” of the 15 they distinguish. As quality is seen as a precondition for trustworthiness, the 4C project focusses in this report on the costs and benefits of “standards based quality assurance” and looks at the 5 current standards related to audit and certification: DSA, Drambora, DIN 31644 of the German nestor group, TRAC and TDR. The first part of the report gives an overview of the current status of these standards. Woven in this overview are some interesting thoughts about audit and certification. It all starts with the Open Archival Information System (OAIS) Reference Model. The report suggests that the OAIS model is there to help organisations to create processes and workflows (page 18), but I think this does not right to the OAIS model. If one really reads the OAIS standard from cover to cover (and should not we all do that regularly?) one will recognize that the OAIS model expects a repository to do more than designing workflows and processes. Instead, a repository needs to develop a vision on how to do digital preservation and the OAIS model gives directions. But the OAIS model is not a book of recipes and we all are trying to find the best way to translate OAIS into practice. It is this lack of evidence which approach will offer the best preserved digital objects, that made the authors in the report wonder whether an audit that will take place now might lead to a risky outcome (either too much confidence in the repository or too little). They use the phrase “dispositional trust” . “It is the trustor’s belief that it will have a certain goal B in the future and, whenever it will have such a goal and certain conditions obtain, the trustee will perform A and thereby will ensure B.”(p. 22). We expect that our actions will lead to a good result in the future, but this is uncertain as we don’t have an agreed common approach with evidence that this approach will be successful.  This is a good point to keep in mind I think as well as the fact that there are many more standards applicable for digital preservation then only the above mentioned. Security standards, record management standards and standards related to the creation of the digital object, to name just a few.

Based on publicly available audit reports (mainly TRAC and DSA, and test audits on TDR) the report describes the main benefits of audits for organisations as

  • to improve the work processes,
  • to meet a contractual obligation and
  • to provide a publicly understandable statement of quality and reliability (p. 29).

These benefits are rather vague but one could argue that these vague notions might lead to more tangible benefits in the future like more (paying) depositors, more funding, etc. By the way, one of the benefits recognized in the test audits was the process of peer review in itself and the ability for the repository management to discuss the daily practices with knowledgeable people.

The authors also tried to get more information about costs related to audit and certification, but had to admit in the end that currently there is hardly any information about the actual costs of an audit and/or get certified (why they mention on page 23 financial figures of 2 specific audits without any context is unclear to me) and base themselves mainly on information that was collected during the test audits that the APARSEN project performed and the taxonomy of costs that was created. For costs we need to wait for more audits and for repositories that are willing to publish all their costs in relation to this exercise.

Reading between the lines,  one could easily conclude that it is not recommended to perform audits yet. But especially now the DP community is working hard to discover the best way to protect digital material, it is important for any repository to protect their investments and to avoid that current funding organizations (often tax payers) will back off because of costly mistakes. The APARSEN trial audits were performed by experts in the field and the audited organizations (and these experts) found the discussions and recommendations valuable. As standards are evolving and best practices and tools are developed, a regular audit by experts in the field can certainly safeguard organizations to minimize the risk for the material. These expert auditors need to be aware of the current state of digital preservation, the uncertainties, the risks, the lack of tools and the best practices that are there. The audit results  will help the community to understand the issues encountered by the audited organizations, as audit results will be published.

As I noticed while reading a lot of preservation policies for SCAPE, many organisations want to get certified and put this aim in their policies. Publishers want to have their data and publications in trustworthy, certified repositories. But all stakeholders (funders, auditors, repository management) should realise that the outcomes of an audit should be seen in the light of the current state of digital preservation: that of pioneering.

‘We learn so much from each other’ – Hildelies Balk about the Digitisation Days (19-20 May)

The Digitisation Days will take place in Madrid on 19-20 May. What can you expect from them and why should you go? In order to get answers to these questions we interviewed Hildelies Balk of the National Library of the Netherlands (KB), who is also a member of the executive board of the organizing insitution, the IMPACT Centre of Competence (IMPACT CoC). – Interview and photo by Inge Angevaare (see below for Dutch version)

Hildelies Balk Reading room National Library

Hildelies Balk in the National Library’s Reading Rooms

The Digitisation Days will be of interest to …?

‘Anyone who is working with digitised historical texts. These are often difficult to use because the software cannot decipher damaged originals or illegible characters. For example:

example OCR historical text

‘The software used to ‘read’ this (Dutch) text produces the following result:

VVt Venetien den 1.Junij, Anno 1618.
DJgn i f paffato te S’ aö’Jifeert mo?üen/bah
.)etgi’uotbciraetail)i.r/JtmelchontDecht
te / sbnbe bele btr felbrr geiufttceert baer bnber
eeniglje jprant o^fen/bie ftcb .met beSpaenfcbeu
enbeeemgljen bifet Cbeiiupcen berbonbru befe

‘The Dutch National Library and many other libraries are striving to make these types of historical text more usable to researchers by enhancing the quality of the OCR (optical character recognition). Since 2008, we have been involved in European projects set up to improve the usability of OCR’d texts – preferably automatically. The IMPACT Centre of Competence as well as the Digitisation Days are quite unique in that they bring together three interest groups:

  • institutions with digitised collections (libraries, archives, museums)
  • researchers working on means to improve access to digitised text (image recognition, pattern recognition, language technology)
  • companies providing products and services in the field of digitisation and OCR.

‘Representatives of all of these groups will be taking part in the Digitisation Days and they offer participants a complete overview of the state of the art in document analysis, language technology and post-correction of OCR.’

What are the most important benefits from the Centre of Competence and the Digitisation Days, in your opinion?

‘The IMPACT Centre of Competence assists heritage institutions in taking important decisions. We evaluate available tools and report about them. Evaluation software of good quality is available as well. We also provide institutions with guidance and advice in digitisation issues by answering questions such as: what would be the best tools and methods for this particular institution? What quality can you expect from a solution? And what will it cost?’

‘The Digitisation Days offer a perfect opportunity for heritage institutions to get together and share experience and knowledge on issues such as: how to embed digitisation in your institution? How to deal with providers? Also: how do we start up new projects? Where do we find funding? On the second day, those who are interested are invited to join a workshop on the topic of the research agenda for digitisation. What should be the focus for the coming years? Should we focus on quantity or quality? How can we help shapeEuropean plans and budgets?’

Now that you mention Europe: IMPACT, IMPACT Centre of Competence, SUCCEED – the announcement of the Digitisation Days is packed with acronyms. Can you give us a bit of help here??

‘IMPACT was the first European research project aimed at improving access to historical texts. It started in 2008, at the initiative of, among others, the Dutch KB. When the project ended, a number of IMPACT partners set up the IMPACT Centre of Competence to ensure that the project results would be supported and developed. The Centre is not a project, but a standing organisation.’

Succeed is another European project, and, by definition, temporary. The objectives are in line with the IMPACT CoC, and the project involves some of the same partners. The aim is raise awareness about the results of European projects related to the digital library and to stimulate implementation. Before the CoC, it was not uncommon for prototypes to be left as they were after completion of a project. Thus the investments did not pay off.’

Will you really turn theory into practice?

‘Yes, most definitely! It is our prime focus for the conference. This is why we instituted the Succeed awards which will be handed out during the Digitisation Days; the Succeed awards recognise the best implementations of innovative technologies. The board has recently announced the winners.’

What do you personally look forward to most during the Digitisation Days?

‘To meeting everybody, to bringing together all these different parties. Colleagues from other institutions, researchers – this is exactly the right kind of meeting for generating exciting ideas and solutions.’

‘We kunnen zoveel van elkaar leren’ – Hildelies Balk over de Digitisation Days (19-20 mei)

Op 19-20 mei worden in Madrid de Digitisation Days gehouden. Wat valt er te beleven en waarom zou je erheen gaan? We vroegen het Hildelies Balk van de Koninklijke Bibliotheek, die voorzitter is van het bestuur van de organisator, het IMPACT Centre of Competence (IMPACT CoC). – interview en foto Inge Angevaare

Hildelies Balk leeszaal KB

Hildelies Balk in de leeszalen van de KB

Voor wie zijn de Digitisation Days interessant?

‘Voor iedereen die te maken heeft met gedigitaliseerde, historische teksten. Die zijn vaak moeilijk bruikbaar omdat de leessoftware veel fouten maakt. Dat komt bij voorbeeld omdat het originele drukwerk zelf al slecht was, of omdat de drukletter slecht leesbaar is:

voorbeeld OCR historische tekst

‘De software die de plaatjes moet omzetten in leesbare tekst maakt daarvan:

VVt Venetien den 1.Junij, Anno 1618.
DJgn i f paffato te S’ aö’Jifeert mo?üen/bah
.)etgi’uotbciraetail)i.r/JtmelchontDecht
te / sbnbe bele btr felbrr geiufttceert baer bnber
eeniglje jprant o^fen/bie ftcb .met beSpaenfcbeu
enbeeemgljen bifet Cbeiiupcen berbonbru befe

‘De KB en andere bibliotheken willen dit soort teksten in bruikbare vorm aanbieden aan wetenschappers. Dus zoeken we al sinds 2008 in Europees verband naar methoden om de teksten te verbeteren, liefst automatisch. Het unieke aan het IMPACT Centre of Competence én van de Digitisation Days is dat daar drie belangengroepen bij elkaar komen die elkaar versterken:

  • instellingen met collecties die gedigitaliseerd zijn (bibliotheken, archieven, musea)
  • onderzoekers die methoden ontwikkelen om gedigitaliseerde tekst te verbeteren (beeldherkenning en – verbetering, patroonherkenning, taaltechnologie)
  • leveranciers van producten en diensten voor digitalisering en OCR (optical character recognition).

‘Door de aanwezigheid van al deze mensen krijgt de bezoeker in twee dagen tijd een compleet overzicht van wat er momenteel allemaal mogelijk is – op het gebied van documentanalyse, taaltechnologie en post-correctie van OCR.’

Wat zie jij als het grootste nut van het Centre of Competence en de Digitisation Days?

‘Het IMPACT Centre of Competence helpt erfgoedinstellingen belangrijke beslissingen te nemen. We evalueren bestaande tools en publiceren daarover. Er is zelfs heel goede evaluatiesoftware. En we leveren begeleiding; als een instelling wil gaan digitaliseren kunnen wij ze van advies dienen. Wat zijn de beste tools en methoden in hun specifieke geval? Wat voor kwaliteit mag je verwachten? Wat gaat het kosten?’

‘De Digitisation Days zijn een perfecte manier voor erfgoedinstellingen om elkaar te ontmoeten, uitgebreid ervaringen en kennis te delen. Bijvoorbeeld: Hoe ga je om met leveranciers? Hoe geef je digitalisering een plek in je organisatie? Maar ook: hoe zetten we nieuwe projecten op? Hoe vinden we geldstromen? Op de tweede dag is er een workshop waarin we met belangstellenden gaan praten over de onderzoeksagenda voor digitalisering. Waar moeten we de nadruk op leggen? Meer kwantiteit of meer kwaliteit? Hoe kunnen we de plannen en budgetten van Europa beïnvloeden?’

Nu je het over Europa hebt: IMPACT, IMPACT Centre of Competence, SUCCEED – de aankondiging van de Digitisation Days staat vol met afkortingen. Kun je een beetje orde scheppen in die chaos?

‘IMPACT was het eerste Europese onderzoeksproject voor verbetering van toegang tot historische teksten dat mede op initiatief van de KB in 2008 is gestart. Toen het project afgelopen was, hebben een aantal IMPACT-partners de handen ineengeslagen om ervoor te zorgen dat de resultaten van het project onderhouden en verder ontwikkeld zouden worden. Dat is het IMPACT Centre of Competence. Geen project, maar een staande organisatie.’

Succeed is weer een Europees project en dus tijdelijk. De doelstellingen liggen helemaal in lijn met het IMPACT CoC, en daarom zijn er deels dezelfde partners bij betrokken. Doel is om te zorgen dat eindresultaten van Europese projecten op het gebied van de digitale bibliotheek goed onder de aandacht worden gebracht zodat ze gebruikt gaan worden in de praktijk. In het verleden bleven prototypes nog wel eens op de plank liggen. Dat is zonde van de investering.’

Wordt de stap van theorie naar praktijk echt gezet?

‘Jazeker! Die willen we juist alle aandacht geven. Daarom reiken we tijdens de Digitisation Days de Succeed awards uit – prijzen voor de beste toepassingen van innovatieve oplossingen. De jury heeft onlangs de kandidaten en de winnaars bekend gemaakt.’

Waar verheug jijzelf je het meest op tijdens de Digitisation Days?

‘Op de ontmoeting, het bij elkaar brengen van al die belanghebbenden. Collega’s van andere instellingen, de onderzoekers – juist uit de ontmoeting komen vaak spannende ideeën en oplossingen voort.’

« Older posts Newer posts »

© 2018 KB Research

Theme by Anders NorenUp ↑