KB Research

Research at the National Library of the Netherlands

Category: Digital Preservation (page 5 of 7)

Roles and responsibilities in guaranteeing permanent access to the records of science – at the Conference for Academic Publishers (APE) 2014

On Tuesday 28 and Wednesday 29 January the annual Conference for Academic Publishers Europe was held in Berlin. The title of the conference: Redefining the Scientific Record. – Report by Marcel Ras (NCDD) and Barbara Sierman (KB)

Dutch politics set on “golden road” to Open Access

During the fist day the focus was on Open Access, starting with a presentation by the Dutch State Secretary for Education, Culture and Science on Open Access. In his presentation called “Going for Gold” Sander Dekker outlined his policy with regards to the practice of providing open access to research publications and how that practice will continue to evolve. Open access is “a moral obligation” according to Sander Dekker. Access to scientific knowledge is for everyone. It promotes knowledge sharing and knowledge circulation and is essential for further development of society.

OA "gold road" supporter and State Secretary Sander Dekker (right) during a recent visit to the K

“Golden road” open access supporter and State Secretary Sander Dekker (right) during a recent visit to the KB – photo KB/Jacqueline van der Kort

Open access means having electronic access to research publications, articles and books (free of charge). This is an international issue. Every year, approximately two million articles appear in 25,000 journals that are published worldwide. The Netherlands account for some 33,000 articles annually. Having unrestricted access to research results can help disseminate knowledge, move science forward, promote innovation and solve the problems that society faces.

The first steps towards open access were taken twenty years ago, when researchers began sharing their publications with one another on the Internet. In the past ten years, various stakeholders in the Netherlands have been working towards creating an open access system. A wide variety of rules, agreements and options for open access publishing have emerged in the research community. The situation is confusing for authors, readers and publishers alike, and the stakeholders would like this confusion to be resolved as quickly as possible.

The Dutch Government will provide direction so that the stakeholders know what to expect and are able to make arrangements with one another. It will promote “golden” open access: publication in journals that make research articles available online free of charge. The State Secretary’s aim is fully implement the golden road to open access within ten years, in other words by 2024. In order to achieve this, at least 60 per cent of all articles will have to be available in open access journals in five years’ time. A fundamental changeover will only be possible if we cooperate and coordinate with other countries.

Further reading: http://www.government.nl/issues/science/documents-and-publications/parliamentary-documents/2014/01/21/open-access-to-publications.html orhttp://www.rijksoverheid.nl/ministeries/ocw/nieuws/2013/11/15/over-10-jaar-moeten-alle-wetenschappelijke-publicaties-gratis-online-beschikbaar-zijn.html

Do researchers even want Open Access?

The two other keynote speakers, David Black and Wolfram Koch presented their concerns on the transition from the current publishing model to open access. Researchers are increasingly using subject repositories for sharing their knowledge. There is an urgent need for a higher level of organization and for standards in this field. But who will take the lead? Also, we must not forget the systems for quality assurance and peer review. These are under pressure as enormous quantities of articles are being published and peer review tends to take place more and more after publication. Open access should lower the barriers for access to research for the users, but what about the barriers for scholars publishing on their research? Koch stated that the traditional model worked fine for researchers. They don’t want to change. However, there do not seem to be any figures to support this assertion.

It is interesting to note that in almost all presentations on the first day of APE digital preservation was mentioned one way or the other. The vocabulary was different, but it is acknowledged as an important topic. Accessibility of scientific publications for the long term is a necessity, regardless of the publishing model.

KB and NCDD workshop on roles and responsibilities

The 2nd day of the conference the focus was on innovation (the future of the article, dotcoms) and on preservation!

The National Library of The Netherlands (KB) and the Dutch Coalition for Digital Preservation (NCDD) organized a session on preservation of scientific output: “Roles and responsibilities in guaranteeing permanent access to the scholarly record”. The session was chaired byMarcel Ras, program manager for the NCDD.

The trend towards e-only access for scholarly information is increasing at a rapid pace, as well as the volume of data which is ‘born digital’ and has no print counterpart. As for scholarly publications, half of all serial publications will be online-only by 2016. For researchers and students there is a huge benefit, as they now have online access to journal articles to read and download, anywhere, any time. And they are making use of it to an increasing extend. However, the downside is that there is an increasing dependency on access to digital information. Without permanent access to information scholarly activities are no longer possible. For libraries there are many benefits associated with publishing and accessing academic journals online. E-only access has the potential to save the academic sector a considerable amount of money. Library staff resources required to process printed materials can be reduced significantly. Libraries also potentially save money in terms of the management and storage of and end user access to print journals. While suppliers are willing to provide discounts for e-only access.

Publishers may not share post-cancellation and preservation concerns

However, there are concerns that what is now available in digital form may not always be available due to rapid technological developments or organisational developments within the publishing industry; these concerns and questions about post-cancellation access to paid-for content are key barriers to institutions making the move to e-only. There is a danger that e-journals become “ephemeral” unless we take active steps to preserve the bits and bytes that increasingly represent our collective knowledge. We are all familiar with examples of hardware becoming obsolete; 8 inch and 5.25 inch floppy discs, Betamax video tapes, and probably soon cd-roms. Also software is not immune to obsolescence.

In addition to this threat of technical obsolescence there is the changing role of libraries. Libraries have in the past assumed preservation responsibility for the resources they collect, while publishers have supplied the resources libraries need. This well-understood division of labour does not work in a digital environment and especially so when dealing with e-journals. Libraries buy licenses to enable their users to gain network access to a publisher’s server. The only original copy of an issue of an e-journal is not on the shelves of a library, but tends to be held by the publisher. But long-term preservation of that original copy is crucial for the library and research communities, and not so much for the publisher.

Can third-party solutions ensure safe custody?

So we may need new models and sometimes organizations to ensure safe custody of these objects for future generations. A number of initiatives have emerged in an effort to address these concerns. Research and development efforts in digital preservation issues have matured. Tools and services are being developed to help plan and perform digital preservation activities. Furthermore third-party organizations and archiving solutions are being established to help the academic community preserve publications and to advance research in sustainable ways. These trusted parties can be addressed by users when strict conditions (trigger events or post-cancellation) are met. In addition, publishers are adapting to changing library requirements, participating in the different archiving schemes and increasingly providing options for post-cancellation access.

In this session the problem was presented from the different viewpoints of the stakeholders in this game, focussing on the roles and responsibilities of the stakeholders.

Neil Beagrie explained the problem in depth, both in a technical, organisational and financial sense. He highlighted the distinction between perpetual access and digital preservation. In the case of perpetual access, organisations have a license or subscription for an e-journal and either the publisher discontinues the journal or the organisation stops its subscription – keeping e-journals available in this case is called “post-cancellation” . This situation differs from long-term preservation, where the e-journal in general is preserved for users whether they ever subscribed or not. Several initiatives for the latter situation were mentioned as well as the benefits organisations like LOCKSS, CLOCKSS, Portico and the e-Depot of the KB bring to publishers.  More details about his vision can be read in the DPC Tech Watch report Preservation, Trust and Continuing Access to e-Journals . (Presentation: APE2014_Beagrie)

Susan Reilly of the Association of European Research Libraries  (LIBER) sketched the changing role of research libraries. It is essential that the scholarly record is preserved, which encompasses e-journal articles, research data, e-books, digitized cultural heritage and dynamic web content. Libraries are a major player in this field and can be seen as an intermediary between publishers and researchers. (Presentation: APE2014_Reilly)

Eefke Smit of the International Association of Scientific, Technical and Medical Publishers (STM) explained to the audience why digital preservation was especially important in the playing field of STM publishers. Many services are available but more collaboration is needed. The APARSEN project is focusing of some aspects like trust, persistent identifiers and cost models, but there are still a wide range of challenges to be solved as the traditional publication models will continually change, from text and documents to “multi-versioned, multi-sourced and multi-media”. (Presentation: APE2014_Smit)

As Peter Burnhill from EDINA, University of Edinburgh, explained, continued access to the scholarly record is under threat as libraries are no longer the custodians of the scholarly record in e-journals. As he phrased it nicely: libraries no longer have e-collections but only e-connections. His KEEPERS registry is a global registry of e-journal archiving and offers an overview of who is preserving what. Organisations like LOCKSS, CLOCKSS, the e-Depot, the Chinese National Science Library and, recently, the US Library of Congress submit their holding information to this KEEPERS Registry. However nice, it was also emphasized that the registry only contains a small percentage of existing e-journals (currently about 19% of the e-journals with an ISSN assigned). More support for the preserving libraries and more collaboration with publishers is needed to preserve the e-journals of smaller publishers and improve coverage. (Presentation: APE2014_Burnhill)

(Reblogged with slight changes from http://www.ncdd.nl/blog/?p=3467)

On-line scholarly communications: vd Sompel and Treloar sketch the future playing field of digital archives

The Dutch data archive DANS invited two ‘great thinkers and doers’ (quote by Kevin Ashley on Twitter) in scholarly communications to do some out-of-the-box thinking about the future of scholarly communications – and the role of the digital archive in that picture. The joint efforts of DANS visiting fellows Herbert van de Sompel (Los Alamos) and Andrew Treloar (ANDS) made for a really informative and inspiring workshop on 20 January 2014 at DANS. Report & photographs by Inge Angevaare, KB Research

(a copy of) Rembrandt's 17th-century scholar Dr. Tulp overseeing Herbert van de Sompel outlining the research world of the 21st century

Rembrandt’s 17th-century scholar Dr. Tulp overseeing Herbert van de Sompel outlining the research world of the 21st century (the painting is a copy …)

Life used to be so simple. Researchers would do their research and submit their results in the form of articles to scholarly journals. The journals would filter out the good stuff, print it, and distribute it. Libraries around the world would buy the journals and any researcher wishing to build upon the published work could refer to it by simple citation. Years later and thousands of miles away, a simple citation would still bring you to an exact copy of the original work.

Van de Sompel and Treloar [the link brings you to their workshop slides] quoted Roosendaal & Geurts (1998) in summing up the functions this ‘journal system’ effectively performed:

  • Registration: allows claims of precedence for a scholarly finding (submission of manuscript)
  • Certification: establishes validity of claim (peer review, and post-publication commentary)
  • Awareness: allows actors in the system to remain aware of new claims (discovery services)
  • Archiving: preserves the scholarly record (libraries for print; publishers and special archives like LOCKSS, Portico and the KB for e-journals).
  • (A last function, that of academic recognition and rewards, was not discussed during this workshop.)

So far so good.

But then we went digital. And we created the world-wide web. And nothing was the same ever again.

Andrew Treloar (at the back) captivating his audience

Andrew Treloar (at the back) captivating his audience

Future scholarly communications: diffuse and ever-changing

Van de Sompel and Treloar went online to discover some pointers to what the future might look like – and found that the future is already here, ‘just not evenly distributed’. In other words: one discipline is moving into the digital reality at a faster pace than another, and geographically there are many differences too. But van de Sompel and Treloar found many pointers to what is coming and grouped them in Roosendaal & Geurts’s functional framework:

  • Registration is increasingly done on (discipline-specific) online platforms such as BioRxiv, ideacite (where one can register mere ‘ideas’!) and Github, a collaborative platform for software developers (also used by the KB research team).
    Common characteristics include:
    – Decoupling registration from certification
    – Timestamping, versioning
    – Registration of various types of objects
    – Machines also function as creators and contributors.
    (We’ll discuss below what these features mean for digital archiving)
  • Certification is also moving to lots of online platforms, such as PubMed Commons, PubPeer, ZooUniverse and even Slideshare, where the number of views and downloads is an indication of the interest generated by the contents.
    Common characteristics include:
    – Peer-review is decoupled from the publication process
    – Certification of various types of objects (not just text)
    – Machines carry out some of the validating
    – Social endorsement
  • Awareness is facilitated by online platforms such as the Dutch ‘gateway to scholarly information’ NARCIS, myExperiment and a really advanced platform such as eLabNotebook RSS where malaria research is being documented as it happens and completely in the open.
    Common characteristics include:
    – Awareness for various types of objects (not just text)
    – Real time awareness
    – Awareness support targeted at machines
    – Awareness through social media.
  • Archiving is done by library consortia such as CLOCKSS, data archives such as DANS Easy, and, although not mentioned during the presentation I may add our own KB e-Depot.
    Common characteristics include:
    – Archiving for various types of objects
    – Distributed archives
    – Archival consortia
    – Audit for trustworthiness (see, e.g., the European Framework for Audit and Certification of Digital Repositories).
Very few places remained unused

Very few seats remained unoccupied

Fundamental changes

Here’s how van de Sompel and Treloar summarise the fundamental changes going on. (The fact that the arrows point both ways is, to my mind, slightly confusing. The changes are from left to right, not the other way around.)

vdSompelTreloar32

Huge implications for digital libraries and archives

The above slide merits some study, because the implications for libraries and digital archives are huge. In the words of vd Sompel and Treloar:

vdSompelTreloar33

From the ‘journal system’ we are moving towards what van de Sompel and Treloar call a ‘Web of Objects’ which is much more difficult to organise in terms of archiving, especially because the ‘objects’ now include ever-changing software & operating systems, as well as data which are not properly handled and thus prone to disappear (Notice on student cafe door: ‘If you have stolen my laptop, you may keep it if you just let me download my PHD-thesis’).

Why archiving is more difficult in the Web of Objects

Why archiving is more difficult in the Web of Objects (if print is too small, check out Slideshare original)

It’s like web archiving – ‘but we have to do better’

Van de Sompel and Treloar compared scholarly communications to websites – ever-changing content, lots of different objects (software, text, video, etc.), links that go all over the place. Plus, I may add, a enormous variety of producers on the internet. Van de Sompel and Treloar concluded: ‘We have to do better than present web-archiving methods if we are to preserve the scholarly record in any meaningful way.’

Two 'great thinkers and doers' confer - Herbert van de Sompel (left) and Andrew Treloar

Two ‘great thinkers and doers’ confer – Herbert van de Sompel (left) and Andrew Treloar

‘The web platforms that are increasingly used for scholarship (Wikis, GitHub, Twitter, WordPress, etc.) have desirable characteristics, such as versioning, timestamping and social embedding. Still, they record rather than archive: they are short-term, without guarantees, read/write and reflect the scholarly process, whereas archiving concerns longer terms, is trying to provide guarantees, is read-only and results in the scholarly record.’

The slide below sums it all up – and it is with this slide that van de Sompel and Treloar turned the discussion over to their audience of some 70 digital data experts, mostly from the Netherlands:

A work in progress: the scholarly communications arena of the future

A work in progress: the scholarly communications arena of the future

Group discussions about the digital archive of the future

So, what does all of this mean for digital libraries and digital archives? One afternoon obviously was not enough to analyse the situation in full, but here are some of the comments reported from the (rather informal) break-out sessions:

  • One thing is certain: it is a playing field full of uncertainties. Velocity, variety and volume are the key characteristics of the emerging landscape. And everybody knows how difficult these are to manage.
  • The ‘document-centred’ days, where only journal and book publications were rated as First Class Scholarly Objects are over. Treloar suggested a move to a ‘researcher-centric’ approach, where First Class Objects include publications and data and software.
  • To complicate matters: the scholarly record is not all digital – there are plenty of physical objects to deal with.
  • How do we get stuff from the recording platforms to the archives? Van de Sompel suggested a combination of approaches. Some of it we may be able to harvest automatically. Some of it may come in because of rules and regulations. But Van de Sompel and Treloar both figured that rules and regulations would not be able to cover all of it. That is when Andrea Scharnhorst (workshop moderator, DANS) suggested that we will have to allow for a certain degree of serendipity (‘toeval’ in Dutch).
Andrea Scharnhorst (DANS): 'Perhaps we have to allow for a certain degree of serendipity'

Andrea Scharnhorst (DANS): ‘Perhaps we have to allow for a certain degree of serendipity’

  • Whatever libraries and archives do, time-stamped versioning will become an essential feature of any archival venture. This is the only way to ensure that scientists can adequately cite anything and verify any research (‘I used version X of software Y at time Z – which can be found in a fixed form in Archive D’).
  • The archival community introduced the concept of persistent identifiers (PID’s) to manage the uncertainties of the web. But perhaps the concept’s usefulness will be limited to the archival stage. Should we distinguish between operational use cases and archival use cases?
  • Lots of questions remain about roles and responsibilities in this new picture, and who is to pay for what. Looking at the Netherlands, the traditional distribution of tasks between the KB National Library (books, journals) and the data archives (research data) certainly merits discussion in the framework of the NCDD (Netherlands Organisation for Digital Preservation); the NCDD’s new programme manager, Marcel Ras, attended the workshop with interest.
Breakout discussions about infrastructure implications

Breakout discussions about infrastructure implications

  • Who or what will filter the stuff that is worth keeping from the rest?
  • Interoperability is key in this complex picture. And thus we will need standards and minimal requirements (as, e.g., in the Data Seal of Approval)
  • Perhaps baffled by so much uncertainty in the big picture, some attendants suggested that we first concentrate on what we have now and/or are developing now, and at least get that right. In other words, let’s not forget that there are segments of the scientific landscape that are being covered even now. The rest of the scholarly communications landscape was characterised by Laurents Sesink (DANS) as ‘the Wild West’.
In this breakout session, clearly discussions focussed on the role of the archive.

In this breakout session, clearly discussions focussed on the role of the archive. Selection: when and by whom? Roles and responsibilities?

  • What if the Internet fails? What if it succumbs to hacks and abuse? This possibility is not wholly unimaginable. But the workshop decided not to go there. At least not today.

In his concluding remarks Peter Doorn, Director of DANS, admitted that there had been doubts about organising this workshop. Even Herbert van de Sompel and Andrew Treloar asked themselves: ‘Do we know enough?’ Clearly, the answer is: no, we do not know what the future will bring. And that is maybe our biggest challenge: getting our minds to accept that we will never again ‘know enough’ at any time. While yet having to make decisions every day, every year, on where to go next. DANS is to be commended for creating a very open atmosphere and for allowing two great minds to help us identify at least some major trends to inspire our thinking.

See also:

  • tweets #rtwsaf (after the official name of the workshop, Riding the Wave and the Scholarly Archive of the Future – the title referring to the 2010 European Commission Report on Scholarly Communications which was the last major report on the issue available).
  • Blog post by Simon Hodson
Where do we go from here? Peter Doorn asked his two visiting fellows

Where do we go from here? Peter Doorn asked his two visiting fellows in Alice-in-Wonderland fashion

The Elephant Returns to the Library…with a Pig!

Hadoop Driven Digital Preservation Hackathon in Vienna

Organised by: SCAPE project  and  Open Planets Foundation

By Clemens Neudecker & René van der Ark

These days, libraries are no longer exclusively collecting physical books and publications, but are also investing in digitisation on a massive scale. At the same time they are harvesting born-digital publications and websites alike. The problem of how to preserve all that digital information in the (very) long run has received a lot of attention. So called “preservation risks” pose a severe threat to the long-term availability of these digital assets. To name but a few: bitrot, format obsolescence and lack of open tools and frameworks.

To tackle these problems now and in the future, beefing up server performance and storage space is no longer a viable option. The growth of data is simply too fast to keep up. Therefore, in recent years, computer scientists tend to opt for scaling out, as opposed to scaling up. This is where buzzwords like the cloud and big data come in. To demystify: instead of scaling up with expensive hardware, scale out by setting up a cluster of cheap machines and doing distributed parallel calculations on them. Scaling out is now often the solution for fast processing of big data, but the principle might just as well be applied to safe (redundant) storage.

The European research project SCAPE (SCAlable Preservation Environments) has been set up in order to help the GLAM community lift their preservation technology to the big data needs of the 21st century digital library. One of the key ideas in SCAPE is to leverage big data technologies like Hadoop, and to apply them in order to scale out the preservation tools and technologies currently in use.

The hackathon on “Hadoop Driven Digital Preservation” at the Austrian National Library (ONB) in Vienna therefore provided a great opportunity for the KB to further its understanding of Hadoop and all its applications. Especially because Hadoop guru Jimmy Lin from the University of Maryland joined us not only as keynote, but also as teacher and co-hacker. Jimmy Lin has past experience working for Twitter and Cloudera and shared many insights in using Hadoop on a productive scale, several dimensions greater than what libraries are currently struggling with. One of his recent projects was to implement a web archive browser on Hadoop’s HBase called WarcBase. A great initiative which might just turn into the next generation Wayback Machine.

Besides, Vienna in December is always worth a visit!

At the event

The event started out with an introduction to the two use cases that were provided upfront by the organisers:

1) Web-Archiving: File Format Identification/Characterisation

2) Digital Books: Quality Assurance, text mining (OCR Quality)

However, participants were free to dive into either of these issues, continue developing their own projects or just investigate completely fresh ideas. So it was not a big surprise when soon after Jimmy’s first presentation introducing Pig as an alternative to writing MapReduce jobs “for lazy people”, many of the participants decided to work on creating small Pig scripts for various preservation related tools.

The nice thing about Pig is that it somewhat resembles common query languages like SQL. This makes it quite readable for most IT savy people. Also it is extensible with custom functions, which can be implemented in Java. Writing some of these user defined functions (UDF) is what we decided to focus on.

This event distinguished itself by the great amount of collaboration. As code reuse was greatly encouraged we decided to fork Jimmy Lin’s WarcBase project on github and extend it with UDF’s for language detection and MIME-type detection using Apache TIKA. The UDF’s we wrote were then in turn again used by many of the other participants’ projects.

The rest of the time we used to get more familiar with writing Pig scripts to apply on actual ARC/WARC files. While unfortunately there was a lack of publicly available ARC/WARC files for testing our MIME-type and language detection UDF’s, we were lucky that colleague Per Møldrup-Dalum from the SB in Aarhus had a cluster and a large collection from the Danish web archive  available for us to test these on:

HadoopVersion   PigVersion            UserId    StartedAt               FinishedAt
2.0.0-cdh4.0.1     0.9.2-cdh4.0.1      scape      14:19:53                 14:22:22

JobId:    job_201308301115_0151

Maps      Reduces
172         19

MaxMapTime       MinMapTime      AvgMapTime
106                         18                           77

MaxReduceTime MinReduceTime  AvgReduceTime
19                           17                           19

Alias                       Feature
a,b,c,d,e,f,raw      GROUP_BY,COMBINER

Input(s):

Successfully read 547093 records (3985472612 bytes) from:
“hdfs://zone1.isilon.sblokalnet/user/scape/arc-files/97-9-2005*”

Output(s):

Successfully stored 90 records (1613 bytes) in:
“hdfs://zone1.isilon.sblokalnet/user/scape/clemens-rene-1”

Counters:

Total records written : 90
Total bytes written : 1613
Spillable Memory Manager spill count : 0
Total bags proactively spilled: 0
Total records proactively spilled: 0

Honestly, we were a bit surprised ourselves when we got these numbers from Per – did our simple Pig script of 7 lines really just process the almost 550,000 ARC/WARC records in only 2:29 minutes? Indeed it did!

To learn more about the various results and outcomes of the event, make sure to check the blog post by Sven Schlarb. It must be mentioned that the colleagues from the ONB and OPF did an outstanding job in terms of event preparation – next to two real-life use cases, they also provided a virtual machine with a pseudo-distributed Hadoop and some more helpful tools from the Hadoop ecosystem that could be used for experimentation. In fact, there was even a cluster with some data ready to execute MapReduce jobs against and really test out how well they would scale. Thanks again!

Outlook

While the KB has been a member of PLANETS, a predecessor to SCAPE, as well as a member of the OPF and the SCAPE project, so far we have only had little time to experiment with Hadoop in our library.

Currently we are looking into using Hadoop for migrating around 150 TB of TIF images from the Metamorfoze Programme to the JPEG2000 format. Following the example of the British Library, we started experimenting with our own implementation of a TIF → JP2 workflow using Hadoop. Will Palmer from the British Library Digital Preservation team has already successfully built such a workflow and published it on github.

However, the TIF → JP2 migration is just about the hardest scenario to optimally implement on Hadoop. The encoding algorithm is very complex and would actually need to be rewritten entirely for parallel processing to make use of the power of Hadoop. Nevertheless, we believe that Hadoop has serious potential  – so the KB is also investigating some more use cases for Hadoop, currently in at least three different areas:

1. Webarchiving

The KB is one of the partners in the WebART project, where together with researchers from the UvA and the CWI new tools and methods to maximize the web archive’s utility for research are being created. Hadoop and HBase are also amongst the applications used here. Together with CWI and UvA the KB hopes to start up a new project for establishing an instance of WarcBase running on top of a scalable HBase cluster – this would really enable new, scalable ways of researching the Dutch web history. Colleague Thaer Sammar from CWI also participated in the hackathon, and the results of his efforts were quite convincing. Also, Jimmy is again one of the collaborators – we keep your fingers crossed!

2. Content enrichment

In the Europeana Newspaper project, the KB is currently creating a framework for named entity recognition (NER) on historical newspapers from all over Europe. Around 10 million pages of full-text will be created in the project, and the KB will provide named entities software for materials in Dutch, German and French language. It is expected that at least 2 million pages will be processed with NER, but the total collection of digitised newspapers at the KB is 8.5 million. Plus there are other collections (books, journals, radiobulletins) for which OCR exists. The KB aims to have all the entities in its digital collections detected, disambiguated and linked within the next 5 years. Given that all of this is text, and the sofware for NER is in Java, it would be interesting to see in how far Hadoop could be used to scale out the processing of all this data.

3. Business processes

From the organization, we are aware of a few particular scalability issues with some of the business processes, such as:

  • Generating business reports quickly: i.e. counting all the KB’s newspapers per publisher is now a painstakingly slow and somewhat unreliable process.
  • Acting as a data provider: harvesting the newspapers’ metadata sequentially takes a week now, and will only take longer when the collection is expanded. Harvesting records in parallel (16 requests / second) created serious stress on the current middleware and even came close to crashing it.
  • Parallelizing (pre-) ingest processes, like metadata validation, file characterization and checksum validation.

Finally, we have also looked at HDFS for scalable, but durable storage. However, for preservation purposes storing files on the Hadoop cluster might introduce some risk, because the files would be partitioned and scattered across the cluster. Then again, the concept of scaling out can still be useful, for example to:

  • prevent bitrot: by replication of files and using agents which repair corrupt replica’s on a regular basis;
  • applying file migrations: here processing data on one machine in sequential batches is again not viable in the long run.

As you see, while there is no Hadoop cluster running in the KB (yet?), it seems there are sufficient ideas and use cases for continuing to work with these technologies, and to build up expertise with Hadoop, HBase etc. We are also very interested in exchanging ideas and use cases with other libraries that are already using Hadoop productively. Last but not least, Hadoopsummit Europe will be held in Amsterdam again next year – so, we’ll see you there perhaps?

Useful links:

Github:

Twitter:

ANADP II in Barcelona

De tweede bijeenkomst van Aligning National Approaches to Digital Preservation (ANADP) vond afgelopen week in Barcelona plaats. De eerste bijeenkomt, in Tallinn in Estonia in 2011, resulteerde in een interessante publikatie http://www.educopia.org/publications met een overzicht van de laatste stand van zaken. En  een reeks aanbevelingen voor verdere discussie (6 in de verkorte versie en 47 in de uitgebreide versie). Om een korte indruk te geven, noem ik enkele belangrijke topics die tijdens deze drie dagen steeds opnieuw onderwerp van discussie waren tijdens de panelsessies,  de actieve werkgroepen en de lezingen van Clifford Lynch (Coalition of Networked Information) die de openingslezing hield en Adam Farquhar (British Library) , die de slotlezing verzorgde.

Clifford Lynch blikte terug wat er sinds 2011 bereikt was. Veel ‘collaboration” (“often just a lot of talking”) dat wel,  maar hij waarschuwde dat deze samenwerking ook tot onderlinge afhankelijkheid kon leiden (“interdependency”) wat een risico kan vormen: gaat het bij een ander mis, dan heb jij daar ook last van. Denk dus van te voren goed na hoe ver de samenwerking moet gaan. Een ander punt betrof de grenzen van digitale duurzaamheid. Zijn die wellicht te nauw? Zouden we ons niet over meer druk moeten maken dan alleen de veilige opslag. Bijvoorbeeld over nieuwe toegangsmogelijkheden, zoals Europeana die biedt. Over informatie die verloren gaat als wij niets doen. Over gewijzigd gebruik en een ander verwachtingspatroon bij gebruikers.  Adam Farquhar constateerde dat de meeste systemen die we nu voor digitale duurzaamheid gebruiken, zijn ingericht op opvraging van één object per keer, maar de nieuwe onderzoekers zien onze collecties als “big data” en willen onderzoek doen op grote aantallen objecten.

Niet alleen in de VS werd een “devaluation of public goods” gevoeld,  nog versterkt  door de krimpende budgetten. “Making the case for digital preservation “ zal steeds belangrijker worden.  Dat kan op verschillende manieren, niet alleen door aan te tonen wat we allemaal bewaren, maar ook door aandacht te vragen voor wat er nu (ongemerkt) verloren gaat. Weten de beleidsmakers wel wat er op het spel staat? Wie maakt zich druk om kleine, lokale krantjes? Of om het bewaren van “public broadcasting”, dat in sommige landen nauwelijks gebeurt, terwijl dat een essentiële bron voor toekomstige onderzoekers is. Welke onderzoeken zijn in de toekomst niet meer mogelijk? Als voorbeeld werd genoemd: hoe komt iemand er over 10 jaar achter hoe lang het reizen van A naar B duurde? Er zijn geen papieren spoorboekjes meer, en niemand bewaart de databases van de spoorwegen.  Het kan ons helpen dat het algemene publiek langzamerhand ook begint te beseffen dat de traditionele manier van overdracht van eigendom voor digitale objecten niet meer werkt. Je bent geen eigenaar meer van je favoriete muziek op Spotify of je favoriete boeken op je Kindle en je kunt ze niet aan je kinderen nalaten.

Tegenwerping is vaak dat we gehinderd worden door de copyrightwetgeving. Dat gaf Lynch direct toe, maar als “digital preservation community” zouden we overeenstemming moeten zien te bereiken over “some sweeping statements” , waarmee we direct de noodzaak voor wijzigingen kunnen aantonen, in plaats van ons in details te verliezen.  

En hoe tonen we aan dat we onze beloften waar maken? Kleine organisaties zeggen soms dat ze “sustainable” zijn voor een bepaalde periode, maar wie controleert dat? Lynch merkte op dat in alle branches sprake is van data verlies, maar dat dit in onze (library) wereld niet lijkt voor te komen. Meermalen is tijdens de conferentie gesproken over het opzetten van een “registry of failures”. Maar er is al een plaats waar de “horror stories” van verloren digitaal materiaal verteld kunnen worden: www.atlasofdigitaldamages.info  

“Economics, the nightmare of sustainability”, (waarbij  “sustainability” volgens Lynch maar al te vaak uitgelegd werd als “somebody else need to pay for his”) was een ander terugkerend onderwerp. Ons antwoord hierop kan gerelateerd zijn aan het feit dat we “public goods’ bewaren:  het is met publieke middelen gemaakt, men heeft er recht op om er toegang tot te houden, en het is een enorme desinvestering als dit verloren zou gaan. Aan de andere kant is het de vraag of we erg veel energie moeten steken in gedetailleerde kostenmodellen.

Luciane Duranti (InterPARES/CICRA) wees er op dat het belangrijk is om de juiste bondgenoten te vinden (de cloud storage providers bijvoorbeeld zouden ook tot onze digital preservation community moeten horen, evenals  leveranciers van systemen en services) en dat we op de juiste plekken moeten zijn, bij UNESCO en bij de conferenties van leveranciers om ons verhaal te vertellen en elkaar te versterken. Ook Chris Greer (Research Data Alliance) pleitte voor meer aansluiting bij andere disciplines en noemde als voorbeeld bio medici die nu beginnen hun collecties duurzaam op te slaan. Zij zouden kunnen profiteren van onze kennis.

Adam Farquhar vatte de trends samen in zijn slotlezing. We zullen overspoeld worden met data en toch moeten we er in slagen digitale duurzaamheid te integreren in onze dagelijkse activiteiten. Dat kunnen we niet meer alleen en zal leiden tot samenwerkingsverbanden en (gezonde) concurrentie met externe partijen die services verlenen. Onderzoekers zullen onze digitale collecties op een andere manier gebruiken, dat vergt aanpassingen in onze systemen (en m.i. mogelijk ook van het OAIS model). Maar bovenal zal de digitale duurzaamheid gemeenschap één consistente boodschap uitstralen; onze activiteiten zijn niet alleen gericht op het gebruik van het digitale materiaal in de toekomst maar ook in het heden.  

Hoe nu verder? Men vond unaniem dat het niet nodig was weer een nieuwe organisatie op te richten om “alignment” te bevorderen, er zijn vele samenwerkingsverbanden die we kunnen gebruiken om bovengenoemde punten verder uit te werken. (Kijk maar eens op  cdb.io/17laZbO  voor samenwerkingen). Wel was er behoefte aan om over enkele jaren weer op deze  strategische wijze over digital preservation te praten. Daar kijk ik naar uit!

Preserving e-journals

Last Thursday a new DPC Technolgy Watch report was presented in London. Neil Beagrie wrote Preservation, Trust and Continuing Access for e-Journals . In a lively setting at RIBA almost all major players were present with representatives from Portico, CLOCKSS , the KB International e-Depot and the Keepers registry to celebrate the launch of this publication and to discuss a variety of challenges and complexities related to preserving this material.

The DPC report gives a good overview of the current state of affairs, the terminology used in this area, the way organizations acquire e-journals (either directly from the publishers or via web harvesting the publisher sites)  and the reasons why organisations like the above mentioned are undertaking this task. E-journals are seen as the basic for scholarly communication. But the publishing model has changed the situation for libraries: instead of having the paper copies on the stacks, they need  “preserving a connection”  – this phrase is from Peter Burnhill- . This is what most subscribing organisations do: they don’t own the content, only the right to distribute the subscription to their members . To avoid loss of this material, one should start preserving the collection and negotiate with publishers  the rights to preserve this. Six use cases illustrate the challenges in preserving this material and they are not so much technical challenges as well as “organisational challenges”, like publishers ceasing operation or transferring part of their collection to another publisher without notifying. One chapter is about Trust, and in this case it is not about the trust in the sense of one repository certified by the  ISO 16363 standard for Trustworthy Repositories. But it is more about how to trust that these e-journals in general will be available in the future. The total sum of the participating and in future participating organisations that preserve e-journals should lead to trusting them to have a complete set that is accessible for the community.

In contrast to websites, where nobody expects to preserve the whole Word Wide Web, with e-journals we strive ‘ to have them all’, at least to preserve all e-journals that are relevant for the scholarly communication.  To monitor this, the KEEPERS registry is there to show us who is preserving which e-journal. In his talk Peter Burnhill tried on the hand to be optimistic about the registry but showed on the other hand that we are not there. Although the e-journals of big publishers like Elsevier and Springer are preserved for example by Portico and the International e-Depot of the KB , they represent only a small part of the total. It is far more difficult to collect the rest of the e-journals, the  “long tail”, as these are often called:  small publishers with only a few e-journals .  Collecting these is costly. One need to search for them, negotiate the terms with each publisher individually and design an ingest flow, which is as time consuming for one small publisher as it is for big publishers.  Some statistics here,  from the 100.000 serials with an e-ISSN, only 21.000 are mentioned in the Keepers registry. So 79.000 are in danger, not to mention the amount without a e-ISSN (more statistics in Burnhills blog).

For preservation the challenge lies also in the technical developments around e-journals, what is exactly the “digital object” ? This topic is less represented in the DPC Tech Watch Report, but a growing problem for collecting organisations.  The time lies behind us that a publication was simply a pdf article. Nowadays it is often accompanied by supplemental material (this can still be seen as part of the article) and “context information”, like websites, altmetrics, data etc. Can this be seen as part of the object and should this also be preserved? The same discussion takes place related to “enhanced publications”. And this is different from the analogue world, where no one expected a library to preserve all the literature referred to in the footnotes of scientific publications! Preserving organisations will need to publish their policies in this respect, to manage the expectations of their user community.

Beagrie writes that “ This makes e-journals one of the most dynamic and challenging areas of digital preservation” . But how about e-books and websites, are they less challenging? Let’s not categorize the objects to preserve (“who is doing the toughest job”), time will show that all digital genres will offer us similar challenges!

Save our preservation tool kit!

Author: Barbara Sierman
Originally posted on: http://digitalpreservation.nl/seeds/save-our-preservation-tool-kit/

Jan Luyken Tea and coffy tool kit. Courtesey Rijksmuseum, Netherlands

The recent US Government shut down should make all people involved in digital preservation thinking, if not worrying. I gave some feedback in Simon Tanners blog post  , but the weekend helped to ponder a bit more about this topic.

We have always said, digital preservation is an international activity and we act like that, by having international collaboration in various areas. Sometimes one organisation starts a very good initiative and we all like to make use of the results, like PRONOM (TNA), FITS (Harvard), JHOVE, OPF, NDSA, DCC, DPC,  PREMIS  at the Library of Congress. Oops… due to the US government shut down this one was no longer available via the well known URL. Although the LoC website has recently be restored (5-10-2013), many more sites are still not available, like for example data.gov . So we could say that the digital preservation community is affected by the US Government shut down, and not only because we can’t have our regular meetings with the preservation people of the Library of Congress.

Have we been naïve as digital preservationists? This is not the first time the US Government shuts down, it also happened in  1995 and 1996 and before. But at that time the web was less influential on our daily activities and we were less dependent of it. Things have changed and we work with the web all day. But we might have been a little bit naïve in expecting things to be there, while our daily job is based on the expectation that things will not always be there. We try to save things. But we don’t have a rescue plan for the information we are dependent on in our processing activities. We might need registries at ingest and transformations and reference works when doing risk assessment of file formats and new object types. But we don’t have an overview of these vital sources that together make our digital preservation tool kit: standards,registries, software, reference works etc. All things that are only accessible from one place are principally in danger – the same rule we’ll apply for our preserved digital collections. What happened in the US can also happen somewhere else.

I would suggest to create an overview of our digital preservation tool kit, the vital elements without which we cannot work,  and be reassured that their availability will no longer be restricted to one place. They should be present on various mirror sites. This activity really requires collaboration, but we already have the infrastructure for that. It is now important to make use of it. We should see this US Government shut down as a wake up call, it is not too late yet.

KB director in This Week in Libraries TWIL #103: The European Library

Watch KB director general Bas Savenije on This Week in Libraries talk on how KB, a national library, will integrate its infrastructure with that of public libraries. Also good stuff on why Europe needs to work together  in The European Library : ‘by working together on developing tools and services you can all share, you free up efforts in your library for other services you can offer your users!’

 

Preservation at Scale: workshop report

Digital preservation practitioners from Portico and from the National Library of The Netherlands (KB) organized a workshop on “Preservation at Scale” as part of iPres2013. This workshop aimed to articulate and, if possible, to address the practical problems institutions encounter as they collect, curate, preserve, and make content accessible at Internet scale.

Preservation at scale has entailed continual development of new infrastructure. In addition to preservation of digital documents and publications, data archives are collecting a vast amount of content which must be ingested, stored and preserved. Whether we have to deal with nuclear physics materials, social science datasets, audio and video content, or e-books and e-journals, the amount of data to be preserved is growing at a tremendous pace.

The presenters at this workshop each spoke from the experience of organizations in the digital preservation space that are wrestling with the issues introduced by large scale preservation. Each of these organizations has experienced annual increases in throughput of content, which they have had to meet, not just with technical adaptations (increases in hardware and software processing power), but often also with organizational re-definition, along with new organizational structures, processes, training, and staff development.

There were a number of broad categories addressed by the workshop speakers and participants:

  1. Technological adaptations
  2. Institutional adaptations
  3. Quality assurance at scale and across scale
  4. The scale of the long tail
  5. Economies and diseconomies of scale

Technological Adaptations
Many of the organizations represented at this workshop have gone through one or more cycles of technological expansion, adaption, and platform migration to manage the current scale of incoming content, to take advantage of new advances in both hardware and software, or to respond to changes in institutional policy with respect to commercial vendors or suppliers.

These include both optimizations and large-scale platform migrations at the Koninklijke Bibliotheek, Harvard University Library, the Data Conservancy at Johns Hopkins University, and Portico, as well as the development by the PLANETS and SCAPE projects of frameworks, tools and test beds for implementing computing-intensive digital preservation processes such as the large-scale ingestion, characterization, and migration of large (multi-terabyte) and complex data sets.

A common challenge was reaching the limits of previous-generation architectures (whether those limits are those of capacity or of the capability to handle new digital object types), with the consequent need to make large-scale migrations both of content and of metadata.

Institutional Adaptations
For many of the institutions represented at this workshop, the increasing scale of digital collections has resulted in fundamental changes to those institutions themselves, including changes to an institution’s own definition of its mission and core activities. For these institutions, a difference in degree has meant a difference in kind.

For example, the Koninklijke Bibliotheek, the British Library, and Harvard University Library have all made digital preservation a library level mandate. This shift from relegating the preservation of digital content to an organizational sub-unit to ensuring that digital preservation is an organization-wide endeavor is challenging, as it requires changing the mindsets of many in each organization. It has meant reallocation of resources from other activities. It has necessitated strategic planning and budgeting for long-term sustainability of digital assets, including digital preservation tools and frameworks – a fundamental shift from one-time, project-based funding. It has meant making choices; we cannot do everything. It has meant comprehensive review of organizational structures and procedures, and has entailed equally comprehensive training and development of new skill sets for new functions.

Quality Assurance at Scale and Across Scales
A challenge to scaling up the acquisition and ingest of content is the necessity for quality assurance of that content. Often institutions are far downstream from the creators of content. This brings along many uncertainties and quality issues. There was much discussion of how institutions define just what is “good enough,” and how those decisions are reflected in the architecture of their systems. Some organizations have decided to compromise on ingest requirements as they have scaled up, while other organizations have remained quite strict about the cleanliness of content entering their archives. As the amount of unpreserved digital content continues to grow, this question of “what is sufficient” will persist as a challenge, as will the challenge of moving QA capabilities further upstream, closer to the actual producers of data.

The Scale of the Long Tail
As more and more content is both digitized and born digital, institutions are finding they must scale for increases in both resource access requests and expectations for completeness of collections.

The number of e-journals in the world that are not preserved was a recurrent theme. The exact number of journals that are not being preserved is unknown, but some facts are:

  • 79% of the 100,000 serials with ISSN are not being known to be preserved anywhere. It is not know how many serials that do not have ISSNs are being preserved.
  • In 2012, Cornell and Columbia University Libraries (2CUL) estimated that about 85% of e-serial content is unpreserved.

This digital “dark matter” is dwarfed in scope by existing and anticipated scientific and other research data, including that generated by sensor networks and by rich multimedia content.

Economies and Diseconomies of Scale
Perhaps the most important question raised at this workshop was the question as to whether we as a community are really at scale yet? Can we yet leverage true economies of scale? David Rosenthal noted that as we centralize more and more preserved content in fewer hands, we will be able to better leverage economies of scale, but we will also be increasing risk of a single point of failure.

Next Steps
The consensus of the group seemed to be that, as a whole, the digital preservation community is not yet truly at scale. However, the organizations in the room have moved beyond a project mentality and into a service oriented mentality, and are actively seeking ways to avoid wasteful duplication of effort, and to engage in active cooperation and collaboration.

Workshop presentations and notes on each presentation are available at: https://drive.google.com/folderview?id=0B1X7I2IVBtwzcGVhWUF0TmJIUms&usp=sharing

iPRES 2013

Author: Barbara Sierman

The iPRES2013 conference took place in beautiful Lisbon, together with the Dublin Core 2013 conference. In total there were around 400 people, from 38 countries.  Each conference had its own program. But the three (shared) key note speakers draw the attention from both the bibliographic people and the digital preservation in the room and sketched their views on important challenges we need to work on collaboratively. Gildas Illien (BnF) strongly advocated that  bibliographic people and digital preservation people would be more cooperative, as they both are trying to make the collections accessible but from a different angle. The user expectations should be leading in both fields and, if so, will require more collaboration in the organizations. Management need to be convinced of this. Paul Bertone from the European Bioinformatics Institute explained the recent breakthrough in storage: storage in DNA, which might be a solution for massive storage of data. And finally Carlos Morais Pires, from the European Commission, talked about Horizon 2020 and data infrastructures (and here – as libraries we need to point this out again and again: data is not only restricted to scientific data generated by instruments, but also the big data collections in libraries and data centres for social sciences ! Carlos Morais Pires immediately agreed on this and changed his slide.)

IMG_3486

Barbara Sierman speaking at iPRES

All presentations can be found on http://purl.pt/24107 ,  covering a wide range of aspects. There are simply so many aspects related to digital preservation ( webarchiving, preservation policies, open source preservation systems, trust, storage, and so on…). I can only advise you to have a look at the above mentioned URL.

Is there a trend to be discovered in all these presentations? To me, they demonstrate there is a lot of national and international collaboration nowadays. The European projects like Blog4Ever, SCAPE, APARSEN, ENSURE and Timbus, national initiatives like Goportis and international collaboration in the 4C project,  they all bring together people from various disciplines . No longer is it only about libraries, archives and data centers, but institutional repositories, health care and business are now also tackling the problem and are presenting their views.  The presentations reflect a greater self-confidence of the digital preservation community; we don’t have the answers to all challenges but we are developing a methodological way to deal with them: the development of standards, life cycle models, cost models, monitoring of the environment, lending from other communities to create tools etc. And most important of all, we know how to find each other.

IMG_2866

Organiser José Borbinha with all varietes of the conference badge on his shirt

But there was also another topic, mainly raised in discussions and during breaks. Our own organisations. The elephant in the room is the fact that our own organisations will need to deal with both analogue and digital material, while the expertise in dealing with analogue material is far more developed in the organisation then the competence of dealing with digital material. Someone said to me “these are different people”.  May be that is the case. Look at the sometimes heated debates  around reading e-books or preferring the paper ones. I like both and don’t think the paper book will disappear. So as a reader I will integrate both worlds and sometimes prefer a paper book above an e-book. This is the world we need to deal with, and organizations need to integrate both worlds. It will require training to have employees that are both  familiar with digital as well as print collections. This is a management challenge, but as digital preservation people we cannot close our eyes for it. We need to convince our management and as the keynote speaker Gildas Illien said (paraphrased by me): “ We need to show our added value. Use the rest of the world to convince your management.” This is how we as digital preservation people can exploit  our  existing collaboration structures!

A deadly sin

Author: Barbara Sierman
Originally posted on: http://digitalpreservation.nl/seeds/a-deadly-sin/

At last week’s iPRES2013 conference in Lisbon, a talk was given about an experiment on the migration of WARC files, done by Tessella, called Studies on the scalability of web preservation. One remark in the talk caused some rumour, namely the fact that the presenter suggested to adapt the WARC file and deviate from the standard. Why did they? Because – as we were told –  the current version of the Wayback Machine software, that enables you to render the WARC file format, is not optimal for rendering WARC files with conversion records. But tweaking the format of the Archival Information Package and store this for the long term is not the way we should go. We preserve information for long term. Our future custodians will not understand this (unless they are told so via metadata and even then) and will assume if they see a WARC format, all the rules in the standard are taken into account. Deviating from this is wrong, in fact it is almost a deadly sin.

After reading the corresponding publication (the conference papers are published by the Portugese National Library as a free e-book), I saw that things were less straight forward. The approach Tessella chose was to create two WARC files: a correct WARC according to the standards and an adapted  WARC for access. From the article:

 This required the development of two different workflows for creating migrated WARC files: one, which is formally correct according to the WARC standard, and maintains the integrity of the WARC schema, and a second which is more pragmatic,  and produces a file that can be displayed correctly by current WARC viewers. This pragmatic workflow can also be used for the migration of container formats  which do not support conversion records, such as ARC files.

So what should one do in the case an ISO  standard does not meet ones requirements? In this case the WARC standard is maintained by the BnF , which can easily be seen if one looks for the standard itself. This is especially mentioned on the internet so that people can get in touch. Another approach is to look for interested parties in the Wayback Machine software, which every one who is involved in web archiving knows, is the Internet Archive. And there is the IIPC, the International Internet Preservation Coalition that is currently initiating a developers working group to improve the Wayback Machine software. So if you have some problems with the standards, think about the millions of precious digital objects  that need to be preserved in that format and get in touch with the community. But don’t tweak the format itself!

Older posts Newer posts

© 2018 KB Research

Theme by Anders NorenUp ↑