KB Research http://blog.kbresearch.nl Research at the National Library of the Netherlands Fri, 24 Aug 2018 13:17:55 +0000 en-US hourly 1 https://wordpress.org/?v=4.4.2 KB Research blog now in archive mode http://blog.kbresearch.nl/2018/08/21/kb-research-blog-now-in-archive-mode/ http://blog.kbresearch.nl/2018/08/21/kb-research-blog-now-in-archive-mode/#respond Tue, 21 Aug 2018 14:04:48 +0000 http://blog.kbresearch.nl/?p=2086 As of mid-2017, the KB research blog has been discontinued. A static archive of the blog will remain available here.

]]>
http://blog.kbresearch.nl/2018/08/21/kb-research-blog-now-in-archive-mode/feed/ 0
Whitts cure for preservationists despair? http://blog.kbresearch.nl/2017/02/04/whitts-cure-for-preservationists-despair/ http://blog.kbresearch.nl/2017/02/04/whitts-cure-for-preservationists-despair/#respond Sat, 04 Feb 2017 22:03:25 +0000 http://blog.kbresearch.nl/?p=2079 (This blogpost was first posted by Barbara Sierman at www.digitalpreservation.nl) 

alice_par_john_tenniel_04After Christmas I tried to reduce my digital pile of recent articles, conference papers, presentations etc. on digital preservation. Interesting initiatives (“a pan European AIP” in the e-Ark project:  wow!) could not prevent that after a few days of reading I ended up slightly in despair: so many small initiatives but should not we march together in a shared direction to get the most out of these initiatives? Where is our vision about this road? David Rosenthals blog post offered a potential medicine for my mood.

He referred to the article of Richard Whitt “Through A Glass, Darkly” Technical, Policy, and Financial Actions to Avert the Coming Digital Dark Ages.” 33 Santa Clara High Tech. L.J. 117 (2017).  http://digitalcommons.law.scu.edu/chtlj/vol33/iss2/1

Mr Whitt works for Google and we have seen on several occasions in the last years that Google finally is interested in digital preservation. Why should Google be interested in digital preservation? Well: if the results of the searches in Google are no longer there – the digital objects -, the fundament of their activity is gone. As Richard Whitt puts it “there is a viable argument that the total value of  the Internet and the World Wide Web in particular, declines significantly in a world without digital preservation”(p. 225).

The challenges

Whitss article is more a booklet about digital preservation than an article  (144 pages!) and gives an up to date overview of the broad range of topics and challenges related to our profession. And he did a good job there!

Read Davids blogpost for some additions and corrections and further thoughts. I would like to add to this that Whitts view is often an American one,  as are many of his sources.  For example the certification tool TRAC (p. 164) is succeeded by the ISO standard 16363 and in Europe this is part of a pyramid  (the European framework) of auditing instruments (DSA, DIN/nestor, ISO), which does not seem to  work in the same manner in the US.  The chapter on copyright describes the US situation and is of less use for Europeans and other parts of the world.

Apart from that he gives a thorough overview of the current state of affairs: an overwhelming and complex set of challenges.

The cure

But the main part of the article is a description of a possible cure to decrease the amount of despair and to help bringing digital preservation forward.  A plea for an organized approach, to create what is called “a deep infrastructure”. To avoid what is often described as the digital dark age.

Whitts ideas are inspired by how the Internet works. He combines the elements in the digital life cycle with an extended OSI model (version Nemeth) in which not only software and hardware layers are present but a Political and Financial Layer were added to it. This results in a model “that helps us uncover the full complexity, so as to better understand and work effectively with it.”

whitss-model

Richard Whitts model describing where preservation challenges need to be solved

In order to build this “deep infrastructure” we need to collaborate and to get organized. The last chapter describes potential areas of collaboration and partners, without summing up concrete organisations.

But could not the existing preservation organisations initiate some activity? Like the Open Preservation Foundation, the Digital Preservation Coalition, nestor, the NCDD,  the IIPC, the NDSA just to name a few (apologies for the European flavour in it). Is not there a challenge for these groups to discuss this framework and start collaborating on different areas to bring digital preservation forward?  A suggestion for iPRES 2017 perhaps?

]]>
http://blog.kbresearch.nl/2017/02/04/whitts-cure-for-preservationists-despair/feed/ 0
Bridging the gap between quantitative and qualitative research in digital newspaper archives http://blog.kbresearch.nl/2017/01/25/bridging-the-gap/ http://blog.kbresearch.nl/2017/01/25/bridging-the-gap/#comments Wed, 25 Jan 2017 08:26:10 +0000 http://blog.kbresearch.nl/?p=2061 This blog post is written by Thomas Smits, KB Researcher-in-residence from May 2017

One of the central and most far-reaching promises of the so-called Digital Humanities has been the possibility to analyse large datasets of cultural production, such as books, periodicals, and newspapers, in a quantitative way. Since the early 2000s, humanities 3.0, as Rens Bod has called it, was posited as being able to discover new patterns, mostly over long periods of time, that were overlooked by traditional qualitative approaches.[1] In the last couple of weeks a study by a team of academics led by Professor Nello Christianini of the University of Bristol made headlines: “This AI found trends hidden in British history for more than 150 years” (Wired) and “What did Big Data find when it analysed 150 years of British history? (Phys.org). Did Big Data and Humanities 3.0 finally deliver on its promise? And could the KB’s collection of digitised newspapers be used for similar research?

The study, “Content analysis of 150 years of British periodicals”, is based on a corpus of 28.6 billion words, contained in 35.9 million articles of 120 regional, or local British newspapers from the period 1800-1950. [2] Focussing on six spheres – values and beliefs, UK politics, technology, economy, social change, and popular culture – the study is mostly based on the ‘use frequency’ of n-grams: the number of times a (combination of) word(s) appears in relation to all the words of the corpus in a specific year. For example, if a corpus consists of 100 words and the 2-gram, a combination of two words, “Digital Humanities” appears three times, the use frequency of this 2-gram is 0,03. In short: by applying n-grams the researchers were able to measure the relative importance of certain words, or combinations of words.

By using this method, the researchers were able to pinpoint specific historic events in their corpus, such as coronations, the election of a new pope, and outbreaks of several contagious diseases. More importantly, they used their method to test the validity of certain long-held notions about the nineteenth century. For example, the study suggests a very clear timeline in the emergence of the concept of “Britishness” in the popular imagination. While recent studies have posited that ‘national identity’ has deep historical roots, predating the nineteenth century, Christianini and his colleagues found that “British” overtook “English” only in the late-nineteenth century, supporting the close connection between the production of national identity, modernity, and the rise of mass media.[3]

Bias

While the researchers carefully composed their corpus, further contextualisation could make future research less biased. First of all, while the 120 newspaper titles studied in this research represent roughly 14% of all published titles, it remains unclear to what part of the press landscape these titles belonged. For example, the study neglects to discuss the fact that newspapers predominantly aimed to reach middle class readers. This bias is further enhanced by two factors: publications directed at lower classes, such as those of the chartist movement or the so-called penny press, were often deemed to be unworthy of archiving.[4] This process is enhanced by digitisation: well-known nineteenth century titles are more likely to be digitised than lesser-known, but not necessarily less influential, cheaper and/or radical publications.[5] Furthermore, by focussing on regional newspapers, the study aspires to mitigate the London-centric bias of research based on newspaper coverage. However, it neglects to account for the fact that in the first half of the nineteenth century many regional newspapers copied articles from London-based newspapers on a large scale, while syndication of content achieved, to some extent, the same result in the final decades of the century.[6]

Most crucially, the researchers seem to equate attention given in newspapers to historical significance. By doing so, they run the risk of failing to acknowledge the most important aspect of the medial form of the newspaper: its focus on ‘newsworthy’ events. The importance of certain long-term developments, which were never perceived as being radically new by contemporary commentators, can only be recognised with the benefit of hindsight. This leads to a somewhat paradoxical situation: digital newspaper archives are used to discover long-term trends, while newspaper discourse is mostly centred on short-term developments.

Delpher

What can researchers using the Delpher corpus learn from this study? The research department of the KB has already made important steps in the large-scale analysis of digitsed newspapers. Its open access n-gram viewer, developed by the University of Amsterdam, enables any user to replicate important parts of the British research. For example, a search for ‘nieuwe paus [new pope]’, yields the same results as the British study. In addition, recent projects of the KB’s fellows and researchers-in-residence use Dutch digitised newspapers in innovative ways. I hope that my own project, which applies two computer vision techniques to images in Dutch newspapers, can continue this tradition.

One could even ask the question why British digitised newspapers are not used more frequently for research similar to that of Christianini. Two private companies, Gale and FindMyPast, provide access to the collection of digitised historical newspapers, originally archived by the British Library. In a recent article, I compare these companies to Trove, the digital collection of Australian newspapers maintained by the government.[7] While public digital collections, such as Trove and Delpher, encourage users to tweak the archive, providing them with access to ‘raw’ data and API’s, private companies, such as FindMyPast, focus on a specific kind of use, in this case amateur historians studying their family history. This results in the fact that access to raw data is expensive, which makes it relatively hard, especially for junior researchers, to use it. In my opinion, we should take a critical look at the network of actors involved in the digitisation of newspapers, books, and other sources. Private companies, such as Google and FindMyPast, increasingly shape our access to the past and our use of historical sources. As I argue in my article, we should continue to discuss how this influences the ways that both researchers and the general public are able to interpret the past and relate to it in ways that are meaningful to them.

Bridging the gap

(Media) historians using qualitative methods would be wise to take note of the results of this study. It opens up a new world of possible research and shows how quantitative analysis can be used to substantiate existing theories. More importantly, the article raises the question if the strict separation between qualitative and quantitative research, or distant and close reading, is useful in distinguishing ‘traditional’ methods from their ‘digital’ counterparts. As the study amply shows, insights from traditional research are essential in defining the questions and contextualising both the corpus and the results of this kind of data-driven research. The project points to the importance of interdisciplinary research teams and, hopefully, will further undermine the trenches into which practitioners of the ‘digital’ and the ‘traditional’ humanities have grouped themselves.

[1] R. Bod, “Who is afraid of patterns? The Particular versus the Universal and the Meaning of Humanities 3.0” BMGN 128, no. 4 (2013): 171-80,  175.

[2] Landall-Welfare et al, “Content analysis of 150 years of British periodicals” PNAS (published ahead of print January 9, 2017). doi:10.1073/pnas.1606380114.

[3] This body of scholarship is mostly connected to Benedict Anderson’s concept of the imagined community. B. Anderson, Imagined Communities. London: Verso, 1983.

[4] For an explanation of what French press historian Jean-Pierre Bacot has called the ‘downward spiral of popularity’ of nineteenth-century newspapers and periodicals see Andrew King’s work on the London Journal: J.P. Bacot, La presse illustrée au XIXe siècle: une histoire oublié. Limoges: PULIM, 2005, 75 : A. King, The London Journal 1845-83: Periodicals, Production, and Gender. Aldershot: Ashgate, 2004, 16.

[5] A. Hobbs, “The Deleterious Dominance of The Times in Nineteenth-Century Scholarship” Journal of Victorian Culture 18, no. 4 (2013): 472-497.

[6] M. Beals, “Musings on a Multimodal Analysis of Scissors-and-Paste Journalism (Part 1),” accessed November 22, 2016, http://mhbeals.com/musings-on-a-multimodal-analysis-of-scissors-and-paste-journalism; B. Nicholson, “‘You Kick the Bucket; We Do the Rest!’: Jokes and the Culture of Reprinting in the Transatlantic Press” Media History 17, no. 3 (2012): 277-278.

[7] T. Smits, “Making the News National: Using Digitized Newspapers to Study the Distribution of the Queen’s Speech by W. H. Smith & Son, 1846–1858” Victorian Periodicals Review 49, no. 4 (2016): 598-625. DOI: 10.1353/vpr.2016.0041

]]>
http://blog.kbresearch.nl/2017/01/25/bridging-the-gap/feed/ 1
Farewell; Work on Discerning Journalistic Styles continues! http://blog.kbresearch.nl/2017/01/17/farewell-work-on-discerning-journalistic-styles-continues/ http://blog.kbresearch.nl/2017/01/17/farewell-work-on-discerning-journalistic-styles-continues/#respond Tue, 17 Jan 2017 11:42:40 +0000 http://blog.kbresearch.nl/?p=2055 At the end of December our current researcher-in-residence dr. Frank Harbers of Groningen University ended his project ‘Discerning Journalistic Styles’. In this blogpost he describes the outcomes and plans for the future.

It is January 2017, meaning my period as researcher-in-residence at the KB has come to an end. It also means that my project Discerning Journalistic Styles (DJS) has come to an end. It was a really nice and valuable experience and a fruitful project in which we (I couldn’t have done it without the expertise of KB programmer Juliette Lonij) have managed to create a classification tool that automatically determines the genre of news articles. You can try the tool yourself at: http://www.kbresearch.nl/genre. Just paste a Dutch news article in the text box, press the button below and the result will appear on the right side; simple as that!

genre-classifier

Currently the tool predicts the correct genre in 65% of the cases. This might not seem that high at face value. However, we need to take into account 1) that genres are ideal types that never manifest themselves in their pure form and boundaries between different genres are fluid; 2) that genres are dynamic concepts that change over time, and 3) that genre is a typical example of a ‘latent content’ category, meaning that determining a genre involves a considerable amount of interpretation. It is therefore unsurprising that classifying genres manually is also difficult and that human coders also regularly disagree on what the correct genre of a text is. In fact, it is not unusual that 20 to 30% of the time, human coders disagree on what the right genre of a (historical) news article is. With that in mind, 65% is a solid result – which is not to say that it doesn’t need to be improved.

In that sense the research has only just begun. Not only because in the coming period we will keep concern ourselves with presenting the results on conferences and in academic articles, but also because we are developing research projects that follow-up on DJS. I therefore hope this won’t be the last time I visit The Hague to delve into the historical newspaper collection. If you are curious about the tool, please keep a close eye on the new Lab website of the KB that will be launched soon. On that website we aim to give much more detailed information on the tool and its background.

]]>
http://blog.kbresearch.nl/2017/01/17/farewell-work-on-discerning-journalistic-styles-continues/feed/ 0
Detecting broken ISO images: introducing Isolyzer http://blog.kbresearch.nl/2017/01/13/detecting-broken-iso-images-introducing-isolyzer/ http://blog.kbresearch.nl/2017/01/13/detecting-broken-iso-images-introducing-isolyzer/#respond Fri, 13 Jan 2017 15:36:18 +0000 http://blog.kbresearch.nl/?p=2053

In my previous blog post I addressed the detection of broken audio files in an automated workflow for ripping audio CDs. For (data) CD-ROMs and DVDs that are imaged to an ISO image, a similar problem exists: how can we be reasonably sure that the created image is complete? In this blog post I will discuss some possible ways of doing this using existing tools, along with their limitations. I then introduce Isolyzer, a new tool that might be a useful addition to the existing methods.

Checksums

A number of techniques exist to verify a newly created ISO image. A seemingly obvious solution would be to do a checksum comparison on both the ISO image and the physical carrier. For instance, the following will work on any Linux system:

md5sum myimage.iso
md5sum /dev/sr0

The first line computes an MD5 checksum from the ISO image; the second line repeats this for the physical carrier. This method is not completely fail-safe. In some tests I did over a year ago, I ran into a a very strange issue where my attempts to image a CD would sometimes result in incomplete reads, and, as a result, truncated ISO images. The problem was most likely caused by faulty hardware (the machine on which I ran those tests more or less died shortly afterwards). Most worryingly, the machine would sometimes return incomplete data, both while creating the ISO image as well as during the subsequent checksum calculation on the physical carrier. The result of this was that the computed checksums were identical in both cases, which meant that the image passed the checksum quality check, even though it was incomplete!

Isovfy

The popular cdrtools library includes a tool called isovfy. Its man page describes it as follows:

isovfy is a utility to verify the integrity of an iso9660 image. Most of the tests in isovfy were added after bugs were discovered in early versions of mkisofs. It isn’t all that clear how useful this is anymore, but it doesn’t hurt to have this around.

I already commented on this tool in an earlier blog post:

The documentation of the tool isn’t very clear about what specific checks it performs. In one of my tests I fed it an ISO image that had its last 50 MB missing (truncated). This did not result in any error or warning message! Most of the reported isovfy errors that I came across in my tests simply reflected the file system on the physical CD not conforming to ISO 9660 (this seems to be pretty common).

You can try this yourself by running isovfy on the following two ISO images:

I ran both images through isovfy (version 3.02a06); both resulted in the following output:

Root at extent 17, 2048 bytes
[0,0]
No errors found

This demonstrates that isovfy is not very useful for detecting truncated ISO files.

Digging into the specs

At this point I decided it was time to start digging into some specs. The ISO 9660 page on the OSDev Wiki gives a good explanation of the internal organisation of an ISO 9660 image. From this I learnt that the Primary Volume Descriptor (which is a data structure that is present on all ISO images) contains two interesting fields:

  • Volume Space Size, which is the "number of Logical Blocks in which the volume is recorded";
  • Logical Block Size, which is "the size in bytes of a logical block".

In theory, multiplying both figures should give the expected size of the ISO image, and this would provide a useful way to check if data are missing. To test this, I wrote a Python script that parses an ISO’s Primary Volume Descriptor fields, calculates the expected file size and then compares this against the actual file size. Running the script against some 20 ISO images I had lying around showed that for 7 files the expected size was indeed identical to the actual file size. For most images, the actual size turned out to be marginally larger than expected (typically about 300-600 kB). For 3 images, the actual size was about twice the expected size. Digging deeper, I found out that these were hybrid images that contain an Apple partition on top of the ISO 9660 file system. According to this Wikipedia article, these hybrid discs come in two varieties:

  1. Hybrid discs that contain an Apple Partition Map (located at 512 bytes into the disc/image).
  2. Hybrid discs without a Partition Map. These contain a Master Directory Block (located at 1024 bytes into the disc/image).

In my case all of the 3 hybrid images turned out to be of the first category. Using the information here and here I was able to add detection of such hybrid images to my code, as well as a simple parser for the ‘zero block’ structure that contains two fields that define the partition’s size: Block Size and Block Count. For my hybrid images, multiplying both figures resulted in a value that was close to (but again marginally smaller than) the actual file size.

Finally, I also added detection of the second hybrid disc category (no Partition Map, but Master Directory Block). The Master Directory Block also contains Block Size and Block Count fields that allow one to calculate the size of the file system.

Isolyzer

I wrapped up the results of the above analyses into Isolyzer, which is a dedicated (Python) tool for checking the size of an ISO image. What it does is this:

  1. Locate the image’s Primary Volume Descriptor (PVD).
  2. From the PVD, read the Volume Space Size (number of sectors/blocks) and Logical Block Size (number of bytes for each block) fields.
  3. Calculate the expected file size as ( Volume Space Size x Logical Block Size ).
  4. If the image contains an Apple Partition Map, read the Block Size and Block Count fields from the ‘zero block’
  5. Calculate the expected file size as ( Block Size x Block Count )
  6. If the image contains an Apple Master Directory Block, read its Block Size and Block Count fields
  7. Calculate the expected file size as ( Block Size x Block Count )
  8. Calculate the final expected file size as the largest value out of any of the above 3 values
  9. Compare this against the actual size of the image files.

In addition to this, Isolyzer also extracts and reports technical metadata from the Primary Volume Descriptor and the Zero Block.

Currently the test results are reported in the following format (this may well change in upcoming releases):

<tests>
    <containsISO9660Signature>True</containsISO9660Signature>
    <containsApplePartitionMap>False</containsApplePartitionMap>
    <containsAppleHFSHeader>False</containsAppleHFSHeader>
    <containsAppleMasterDirectoryBlock>False</containsAppleMasterDirectoryBlock>
    <parsedPrimaryVolumeDescriptor>True</parsedPrimaryVolumeDescriptor>
    <sizeExpected>358400</sizeExpected>
    <sizeActual>358400</sizeActual>
    <sizeDifference>0</sizeDifference>
    <sizeAsExpected>True</sizeAsExpected>
    <smallerThanExpected>False</smallerThanExpected>
</tests>

In the above example the sizeExpected field is the size as calculated from the ISO/Apple headers, and sizeActual is the actual size. In this case both are identical. Below some output for a truncated ISO:

<tests>
    <containsISO9660Signature>True</containsISO9660Signature>
    <containsApplePartitionMap>False</containsApplePartitionMap>
    <containsAppleHFSHeader>False</containsAppleHFSHeader>
    <containsAppleMasterDirectoryBlock>False</containsAppleMasterDirectoryBlock>
    <parsedPrimaryVolumeDescriptor>True</parsedPrimaryVolumeDescriptor>
    <sizeExpected>358400</sizeExpected>
    <sizeActual>49157</sizeActual>
    <sizeDifference>-309243</sizeDifference>
    <sizeAsExpected>False</sizeAsExpected>
    <smallerThanExpected>True</smallerThanExpected>
</tests>

So, in this case sizeDifference is negative, and flag smallerThanExpected equals ‘True’ (which indicates a damaged image).

Feedback wanted

At this stage Isolyzer is a bit experimental and pretty rough around the edges, and I wouldn’t recommend it for production use. Nevertheless I’m curious about any feedback on the tool. Do others find this useful? Are things missing (i.e. other hybrid disc types I’m not aware of), or did I get anything completely wrong?

One thing that puzzles me a bit is that for the majority of ISO images I’ve come across, the expected size as calculated by Isolyzer is marginally smaller than the actual size. The difference is typically in the order of about 300-600 kB. I’m not quite sure what’s causing this, although this article mentions that some CD writing software packages add padding bytes when writing a CD. I wasn’t able to verify if this, although this SuperUser answer on validating a burnt DVD suggests it as well. If anyone knows more about this, please let me know!

Isolyzer can be found here on Github. It can be installed using pip; see the instructions here. For Windows users who cannot/don’t want to install Python I also provided stand-alone Windows binaries, which are available for download here.

]]>
http://blog.kbresearch.nl/2017/01/13/detecting-broken-iso-images-introducing-isolyzer/feed/ 0
Breaking WAVEs (and some FLACs) http://blog.kbresearch.nl/2017/01/04/breaking-waves-and-some-flacs/ http://blog.kbresearch.nl/2017/01/04/breaking-waves-and-some-flacs/#respond Wed, 04 Jan 2017 14:55:01 +0000 http://blog.kbresearch.nl/?p=2048 At the KB we have a large collection of offline optical media. Most of these are CD-ROMs, but we also have a sizeable proportion of audio CDs. We’re currently in the process of designing a workflow for stabilising the contents of these materials using disk imaging. For audio CDs this involves ‘ripping’ the tracks to audio files. Since the workflow will be automated to a high degree, basic quality checks on the created audio files are needed. In particular, we want to be sure that the created audio files are complete, as it is possible that some hardware failure during the ripping process could result in truncated or otherwise incomplete files.

To get a better idea of what software tool(s) are best suitable for this task, I created a small dataset of audio files which I deliberately damaged. I subsequently ran each of these files through a set of candidate tools, and then looked which tools were able to detect the faulty files. The first half of this blog post focuses on the WAVE format; the second half covers the FLAC format (at the moment we haven’t decided on which format to use yet).

WAVE dataset

For the WAVE dataset I started out with a small, intact WAVE file. Using a Hex editor I then made the following derivatives of this file:

Candidate tools, WAVE

The candidate tools I used to analyse the WAVE files are:

  • jhove includes a WAVE validation module, which makes it an obvious choice. The tested version is 1.14.6, 2016-05-12.
  • shntool is a "multi-purpose WAVE data processing and reporting utility". It was first released in 2000. The tested version is 3.0.7.
  • ffmpeg is a popular conversion tool for audio and video formats. The tested version is 3.2.2.
  • mediainfo is a widely-used feature extraction tool for audiovisual files. The tested version is v0.7.81.

Note that of the above tools, only Jhove and Shntool are designed to detect problems in WAVE files. Both Ffmpeg and Mediainfo were primarily designed for other purposes (format conversion and technical metadata extraction), and they were not designed to detect defective files! I included these tools here mainly because they are widely used, and I was curious whether they would throw up anything interesting in case of defective files1. I ran the tools with the following command-line arguments (replacing "foo.wav" with the actual file name):

Jhove

jhove -m WAVE-hul foo.wav

Shntool

shntool info foo.wav

Ffmpeg

ffmpeg -v error -i foo.wav -f null -

Mediainfo

mediainfo foo.wav

I automated this using a simple shell script that runs each tool on all files, and then writes the output to a set of text files.

Results, WAVE

The full output results of each tool can be found here.

Jhove

The ‘Status’ field in Jhove’s output summarises the validation outcome. Here are the results for each file:

File Result
frogs-01.wav Status: Well-Formed and valid
frogs-01-last-byte-missing.wav Status: Well-Formed and valid
frogs-01-last-2032-bytes-missing.wav Status: Well-Formed and valid
frogs-01-byte-missing-at-offset-811537.wav Status: Well-Formed and valid

So, Jhove was unable to detect any of the damaged files at all!

Shntool

Shntool checks a WAVE on six criteria, which are listed in its output under ‘Possible problems’:

Possible problems:
  File contains ID3v2 tag:    no
  Data chunk block-aligned:   yes
  Inconsistent header:        no
  File probably truncated:    no
  Junk appended to file:      no
  Odd data size has pad byte: n/a

The thing to watch here is the ‘File probably truncated’ item:

File Result
frogs-01.wav File probably truncated: no
frogs-01-last-byte-missing.wav File probably truncated: yes (missing 1 byte)
frogs-01-last-2032-bytes-missing.wav File probably truncated: yes (missing 2032 bytes
frogs-01-byte-missing-at-offset-811537.wav File probably truncated: yes (missing 1 byte)

So, Shntool was able to detect all damaged files.

Ffmpeg

For our Ffmpeg call we monitor any errors that are sent to the standard error stream. The results:

File result
frogs-01.wav
frogs-01-last-byte-missing.wav [pcm_s16le @ 0x3545380] Invalid PCM packet, data has size 3 but at least a size of 4 was expected
Error while decoding stream #0:0: Invalid data found when processing input
frogs-01-last-2032-bytes-missing.wav
frogs-01-byte-missing-at-offset-811537.wav [pcm_s16le @ 0x2768380] Invalid PCM packet, data has size 3 but at least a size of 4 was expected
Error while decoding stream #0:0: Invalid data found when processing input

Interestingly, Ffmpeg reports an error for both files that have 1 byte missing, but it doesn’t for the file that has 2023 bytes missing. This suggests that Ffmpeg is not suitable for detecting broken WAVE files.

Mediainfo

Mediainfo didn’t report errors or warnings for any of these files. This is not surprising, but it does confirm that Mediainfo cannot be used for detecting broken WAVE files.

FLAC dataset

Analogous to the WAVE dataset, I started out with a small, intact FLAC file, which I then butchered into the following derivative files:

Candidate tools, FLAC

The set of candidate tools is identical to the one used for the WAVE analysis, with two exceptions:

  • flac is the reference implementation of the FLAC format. The tested version is 1.3.0.
  • Since Jhove does not include a FLAC module, it was not used.

Flac

The Flac tool is able to encode audio to FLAC, and decode and analyze FLAC files. For this tests I ran it with the * -t* (or –test) option:

flac -t foo.flac

This decodes a FLAC without writing the decoded data to a file. Any errors during the decoding process are reported to the standard error stream.

Results, FLAC

The full output results of each tool can be found here.

Shntool

Even though Shntool supports FLAC, it was not able to detect the missing data in any of the files:

File Result
frogs-01.flac File probably truncated: no
frogs-01-last-byte-missing.flac File probably truncated: no
frogs-01-last-1000-bytes-missing.flac File probably truncated: no
frogs-01-byte-missing-at-offset-651202.flac File probably truncated: no

So, Shntool does not provide any meaningful information on whether a FLAC is damaged.

Ffmpeg

Here are the results for Ffmpeg:

File Result
frogs-01.flac
frogs-01-last-byte-missing.flac [flac @ 0x294b860] overread: 1
Error while decoding stream #0:0: Invalid data found when processing input
frogs-01-last-1000-bytes-missing.flac [flac @ 0x3c5d860] overread: 1
Error while decoding stream #0:0: Invalid data found when processing input
frogs-01-byte-missing-at-offset-651202.flac [flac @ 0x279faa0] overread: 1
Error while decoding stream #0:0: Invalid data found when processing input

So, Ffmpeg was able to identify all damaged FLACs.

Mediainfo

Similar to the WAVE results, Mediainfo again didn’t report errors or warnings for any of these files.

Flac

Finally the results for the Flac tool:

File Result
frogs-01.flac
frogs-01-last-byte-missing.flac ERROR while decoding data
state = FLAC__STREAM_DECODER_END_OF_STREAM| |frogs-01-last-1000-bytes-missing.flac|ERROR while decoding data
state = FLAC__STREAM_DECODER_END_OF_STREAM| |frogs-01-byte-missing-at-offset-651202.flac|ERROR while decoding data
state = FLAC__STREAM_DECODER_READ_FRAME|

So, the Flac tool was able to identify all defective files2.

Conclusion

Out of the candidate tools considered here, only Shntool was able to identify all damaged WAVE files in this experiment. As a result, this (ancient!) tool still appears to be the best choice for detecting damaged WAVE files. Surpringly, Jhove was unable to detect any of the damaged files at all, and is probably best avoided for this particular purpose. For FLAC, both the Flac tool (FLAC reference implementation) and Ffmpeg were able to detect all damaged files, and both appear to be suitable tools.

Dataset and scripts

All example files, scripts and raw tool output are available here:

https://github.com/KBNLresearch/detectDamagedAudio

Post scriptum: update on MediaInfo and MediaConch

In response to this post the developers of MediaInfo added support for detecting truncated WAVE files. This should cover all of the damaged WAVE files presented here. Moreover, their Twitter account announced that detection of FLAC flaws is planned for the MediaConch tool, but that they are looking for sponsors for this.


  1. Also, this thread on superuser.com recommends Ffmpeg for checking the integrity of video files.

  2. On a side note, I noticed that the error stream of the Flac tool sometimes contained a sequence of 21 non-printable ‘0x08’ (backspace) characters. This is probably a bug.

]]>
http://blog.kbresearch.nl/2017/01/04/breaking-waves-and-some-flacs/feed/ 0
Two Dutch DPC Preservation Awards: what is it all about? http://blog.kbresearch.nl/2016/12/11/two-dutch-preservation-awards-what-is-it-all-about/ http://blog.kbresearch.nl/2016/12/11/two-dutch-preservation-awards-what-is-it-all-about/#respond Sun, 11 Dec 2016 10:06:52 +0000 http://blog.kbresearch.nl/?p=2043 Accompanied by traditional festival tunes of Scottish bagpipes the finalists of the 2016 Digital Preservation Awards and their colleagues “celebrated digital preservation”, as William Kilbride called this event last week in London. And in the audience the proud Dutch group of attendees celebrated even more as we won both the Award for Research and Innovation sponsored by the Software Sustainability Institute and the award for Safeguarding the digital legacy sponsored by The National Archives. The 17 international judges looked at 33 submissions, from 10 different countries.  What was the magical ingredient that helped the Netherlands submitting 3 projects, two of them worthwhile to receive the trophees?

With the help of Rijksmuseum digitization

One of the reasons is the following. As reported earlier on iPRES and in several other places, the Ministry of Education, Culture and Science started in 2015  (and financed) a program called Network Cultural Heritage, with a focus on exploring more the Dutch digital cultural heritage, by making our digital collections more visible, more connected with each other and more sustainable. While the NCDD had already started some projects in 2014, this initiative gave a boost to these projects and two of the nominees can be related directly to this Network Cultural Heritage.

A nationwide network

The Award for Research and Innovation went to the project Constructing a network of nationwide facilities together, project-lead Joost van der Nat. Although the Netherlands is a small country, we have a large amount of organisations with a mandate to preserve digital heritage. Not everyone is aware of what others are doing, nor are we always aware how we can benefit from each other. More collaboration in all kinds of areas related to digital preservation can improve efficiency, effectiveness and professionalism for both small and large organisations. Based on common sense, desk research, lots of discussions, interviews and an intense feedback meeting  in 2014 – where for the first time over 80 Dutch preservationists were together, the Analytical Framework for the infrastructure was created. This framework combined requirements from OAIS, legal requirements, policies and quality assurance, R&D, Training and ICT, thus giving an overview of the areas that were important for exercising digital preservation and finding partners to collaborate on this task. But we are all different organisations, as is our digital material and our legal mandate. So it was also important to pay attention in the model to the differences between the various domains.

The second part of the project was a reality check: interviews with the stakeholders in different domains helped identifying the current status of collaboration, whether people were willing to collaborate more by using shared services for digital preservation activities and what activities they thought were so organisation or domain specific that they always will do them on their own or with people in their own domain. The resulting diagram is the starting point for developing a real network of collaboration in digital preservation. Currently an inventory is created of existing services, which can be candidates to be incorporated in this model.

Web archaeology The Digital City

Quite a different project in the Network of Cultural Heritage activities is the Digital City revives of which “web archaeologist” Tjarda de Haan is project lead. In 1994  Amsterdam was the first city in the world with a digital counterpart, when The Digital City (De Digitale Stad) was started, offering its citizens free access to the internet, enabling them to have an email address and to build their own virtual environment. This web presence lasted until 2001 when the website was taken down. Initiated by the Amsterdam Museum, numerous paths have been explored since 2011 to find traces of this Digital City and to make a reconstruction. This project offers an interesting case to raise awareness for digital preservation. It not only shows how “digital archaeology” is able to retrieve a digital environment from the past and what can be learnt of the efforts, tools and technical skills  that need to be in place to enable this, but the results of this archaeology activity also need to be safely preserved for the future. This is a good example of a local initiative that had a need to be served by a “network of nationwide facilities”.  As we want to keep the result accessible too, there will be a close connection with the visibility and usability parts of the Network Cultural Heritage.

The 3rd Dutch nominee was the submission of Research Data Netherlands (a collaboration of DANS and 4TU). They developed a training course for those responsible for preserving research data: Essentials 4 Data Support.

These initiatives by the partners in the Network Cultural Heritage under project management of the NCDD and by Research Data Netherlands all contributed to more preservation awareness in the Netherlands. Winning the awards are the cherries on the cake!

]]>
http://blog.kbresearch.nl/2016/12/11/two-dutch-preservation-awards-what-is-it-all-about/feed/ 0
NCDD Studiedag: Een web van webarchieven http://blog.kbresearch.nl/2016/11/18/ncdd-studiedag-een-web-van-webarchieven/ http://blog.kbresearch.nl/2016/11/18/ncdd-studiedag-een-web-van-webarchieven/#respond Fri, 18 Nov 2016 11:41:19 +0000 http://blog.kbresearch.nl/?p=2036 Nederland mag dan een klein land zijn, maar we staan wereldwijd wel op nummer 3 wat betreft het aantal uitgereikte domeinnamen – meer dan 5 miljoen. Ruim 14.000 daarvan worden nu door de KB verzameld en gearchiveerd in onze Web Collectie. Gisteren hield de NCDD een studiedag bij het Instituut voor Beeld en Geluid onder de titel Een web van web archieven om de Nederlandse samenwerking bij het bouwen van web collecties te bevorderen.

nominet

Uitsnede van: http://nominet-prod.s3.amazonaws.com/wp-content/uploads/2016/03/Map-Of-The-Online-World.jpg

Dat web archivering nuttig is bleek meteen al uit de door Peter de Bode (KB) samengestelde presentatie van web sites die niet meer online te zien zijn maar wel in ons KB web archief aanwezig zijn. Ook de gastheer Tom de Smet van Beeld en Geluid zei dat hun audiovisuele collectie niet meer los te koppelen is van wat er op het web gebeurt. En dat geldt voor veel organisaties in Nederland. Niet alleen worden web collecties aangelegd omdat het gezien wordt als cultureel erfgoed (zoals door de KB) maar er zijn ook veel organisaties die aan web archivering doen om te voldoen aan de Archiefwet. De NDE werkgroep Web Archivering heeft een onderzoek onder 22 Nederlandse instellingen gedaan om te kijken of er al onderling samengewerkt wordt (de N van NDE is immers “netwerk”). Conclusie is dat dit nog aanzienlijk verbeterd kan worden. De overgrote meerderheid doet pas sinds de laatste 3 jaar iets aan web archivering (de KB sinds 2007), waarbij dan een groot gedeelte dit heeft uitbesteed aan commerciële bedrijven (waardoor ze bijvoorbeeld “geen idee” hadden van de tools die gebruikt werden om de websites binnen te halen: een cruciaal en zeer kwetsbaar onderdeel). Naast het verzamelen van web sites, komt ook de (technisch moeilijke) wens om sociale media te gaan bewaren – waar overigens weer allerlei privacy aspecten aan zitten.

Keynote spreker Herbert van de Sompel (onderzoeker bij Los Alamos National Laboratory, tijdelijk werkzaam bij DANS) demonstreerde een aantal technische oplossingen om ook web archieven die alleen ter plekke en dus niet online raadpleegbaar zijn, toch met elkaar te verbinden, zodat gebruikers weten wie wat bewaart en zelfs een overzicht in de tijd kunnen krijgen op basis van verschillende data waarop een snapshot van de website is gemaakt (memento, memgator ). Ook incomplete pagina’s (het is niet altijd mogelijk alles van een website binnen te halen) zou je op deze wijze kunnen “reconstrueren”, waarbij je je natuurlijk wel moet realiseren in hoeverre dit “authentiek” is.

Filip Boudrez van het Stadsarchief Antwerpen doet al sinds de jaren 90 aan web archivering, waarbij bewijsvorming en verantwoording het uitgangspunt vormen. Hij beschreef de overwegingen bij verschillende stappen en de manieren waarop aan kwaliteitscontrole gedaan wordt, waaronder handmatige controle door vrijwilligers. Documentatie over alle besluitvorming wordt nauwkeurig per web site vastgelegd.

Naast web archivering kennen we tegenwoordig ook web archeologie, dankzij het NDE project van Tjarda de Haan, waarbij De Digitale Stad van Amsterdam gereconstrueerd wordt aan de hand van verzamelde informatie op oude dragers in combinatie met nieuwe technieken.

Rafaël Rozendaal, een web kunstenaar gaf zijn visie op hoe zijn kunst bewaard zou moeten blijven en welke maatregelen hij neemt om zijn kunst zo lang mogelijk werkend te houden, onder het motto “als het niet meer werkt, wordt het archief” (bijvoorbeeld omdat de gebruikte tools niet langer in combinatie werken of het beoogde effect bereiken).

Kees Teszelszky (KB afdeling Onderzoek) toonde de eerste Nederlandse website en vertelde over zijn onderzoek naar de context van de KB Web Collectie in relatie tot wat anderen doen en wat er wellicht meer zou kunnen om het Nederlandse web in web collecties vast te leggen. Samenwerking is ook hier weer het toverwoord.

Dat erfgoed instellingen en overheidsinstellingen elkaar daarbij moeten vinden, vertelde Nick Chapel (Erfgoedinspectie) op basis van het binnenkort te verschijnen rapport: hoewel wettelijk verplicht, schiet de Centrale Overheid nog schromelijk tekort wat betreft het archiveren van hun eigen websites en de bijbehorende sociale media.

In een afsluitend panel werden de NCDD Expert groep Webarchivering gepresenteerd, en werden mogelijke initiatieven met het publiek in de zaal bediscussieerd, onder deskundige leiding van de dagvoorzitter Jantje Steenhuis, stadsarchivaris van Rotterdam.  Aan de slag dus maar, de tijd is er rijp voor!

Barbara Sierman

]]>
http://blog.kbresearch.nl/2016/11/18/ncdd-studiedag-een-web-van-webarchieven/feed/ 0
DH Clinics – librarians unite! http://blog.kbresearch.nl/2016/11/11/dh-clinics-librarians-unite/ http://blog.kbresearch.nl/2016/11/11/dh-clinics-librarians-unite/#comments Fri, 11 Nov 2016 09:52:27 +0000 http://blog.kbresearch.nl/?p=2005 You might have heard someone from @KBNLResearch mention DH Clinics, or a colleague at the libraries of the Vrije Universiteit or Universiteit Leiden, but what are they, why do we need them and who are they for?

The DH Clinics are our attempt of spreading the DH-word amongst our Dutch colleagues. We wanted to set up a community of librarians who were involved in DH, in order to learn from each other and discuss new methods and initiatives. However, we soon learned that a lot of academic libraries in the Netherlands were still thinking about DH and how to implement it in their organisations. We’re speaking early 2015 now and luckily, a lot has happened since, but we believe a small impulse is needed to speed everything along.

And that is why we are now organising DH Clinics for Dutch academic librarians (and possibily also some archivists). The idea of the clinics is that we tackle several major themes of DH over six full-day sessions. The mornings are dedicated to lectures about the theme (think of, for example, text and data mining) and are open to a public that is interested in DH. In the afternoon, we’ll have hands-on workshops with specific applications. This part of the session is meant for people who are actually working with DH researchers, or want to do so.

We’re not setting out to re-train librarians into programmers or data crunchers, but we do want to provide them with the basics of DH with which they should increase their knowledge level to such an extent that they are able to follow the (online) discussions in the field, give tips to beginning researchers, perhaps even use some of the tools for their own work and ideally to engage with the very rich online content to learn more, such as The Programming Historian or Library Carpentry (both of which are used as inspiration for our clinics).

We’re developing the clinics with the Working Out Loud-principles, which means we will regularly share what we are doing and invite you to comment on it. Since we’re working together with the VU and UBL, this blog is not the only one you should keep an eye on, but we’ll announce everything via Twitter as well, so if you’re not already following us, now is the perfect time to start! @KBNLresearch

]]>
http://blog.kbresearch.nl/2016/11/11/dh-clinics-librarians-unite/feed/ 2
Tackling problems and making progress http://blog.kbresearch.nl/2016/11/03/tackling-problems-and-making-progress/ http://blog.kbresearch.nl/2016/11/03/tackling-problems-and-making-progress/#respond Thu, 03 Nov 2016 15:57:11 +0000 http://blog.kbresearch.nl/?p=1994 Our current Researcher-in-Residence, Frank Harbers, is well under way with his project “Discerning Journalistic Styles. Exploring Automated Analysis of Journalism’s Modes of Expression”. In this blogpost he gives an update on his project and its progress.

Frank Harbers

It has been several months since I wrote the first blog about my work as researcher-in-residence and the research project is in full swing by now. The first phase of the project , connecting the metadata from my own database to the historical newspaper data (and metadata) in Delpher is finished and we are fully enveloped in the main part of the project: training a classifier to automatically determine the genre of historical newspaper articles.

The first phase was not as successful as we hoped, but we have managed to create a – modest – dataset to train the classifier. Initially, we hoped to be able to connect the metadata about approximately 33.000 Dutch newspaper articles to the data in Delpher. A crucial factor in the success of this attempt was the extent to which the segmentation of newspaper articles in Delpher matched the way the newspaper articles were segmented for the content analysis that resulted in the set of metadata about the historical newspaper articles. Unfortunately, it was far from a perfect match. For that reason the newspapers before 1945 could not be included – basically half of the metadata. Furthermore, De Volkskrant after the Second World War has not been digitized by the KB. In addition, the segmentation of De Telegraaf in the postwar period was so different that we couldn’t include that either. In the end, this meant that we could only use the data of Algemeen Handelsblad/NRC Handelsblad in the postwar period. So quickly we saw our dataset shrink from the potential 33.000 articles to a modest 2000 articles. A bit of a setback, but fortunately we can still use this smaller dataset to train a genre classifier. This experience does make clear how crucial segmentation is for the creation of datasets that can be fruitfully used for digital humanities research into journalism history.

At the moment, we are working on the second phase of the project. We have identified several genres that we would like to classify. These genres, such as news reports, reportages, interviews, opinion articles, reviews, news analyses, can shed light on the way journalism developed from a reflective, opinion-oriented way of doing journalism to a more event-centered and fact-oriented journalism practice. At the core of this part is the translation of the genre definitions to clear linguistic markers that can be identified automatically. Take for instance the news report, a genre that is defined by the use of the inverted pyramid (a story structure in which typical journalistic questions, like Who, What, Where and When, are answered in the first paragraph. Moreover, it often contains direct quotes from sources and is generally a fairly concise article written in a depersonalized, objective style. Question is how you can recognize these features automatically in the text. In this case, the quotes can be recognized by the presence of quotation marks (for which a high quality OCR is crucial) and we will attempt to identify the inverted pyramid structure by using named identity recognition to see whether questions concerning who was involved and where and when it happened are answered. We hope the depersonalized style can be captured by looking at the lack of a first person perspective (the use of the pronoun ‘I’ or ‘We’) and the lack of adjectives that create a colorful and subjective account.

Juliette Lonij, programmer on this project, is currently developing the Python software to extract the features on which the classifier will run. She looked into different natural language processing software packages to pre-process the article texts and chose to use FROG for tokenization and Part-of-Speech tagging, which facilitates our research needs quite well (other packages might be added in the future). And today, we have just run a first exploratory test with the classifier, which showed promising results. In the coming weeks we will keep on testing and refining the classifier. So wish us luck!

 

]]>
http://blog.kbresearch.nl/2016/11/03/tackling-problems-and-making-progress/feed/ 0