Uncategorized – KB Research http://blog.kbresearch.nl Research at the National Library of the Netherlands Fri, 24 Aug 2018 13:17:55 +0000 en-US hourly 1 https://wordpress.org/?v=4.4.2 KB Research blog now in archive mode http://blog.kbresearch.nl/2018/08/21/kb-research-blog-now-in-archive-mode/ http://blog.kbresearch.nl/2018/08/21/kb-research-blog-now-in-archive-mode/#respond Tue, 21 Aug 2018 14:04:48 +0000 http://blog.kbresearch.nl/?p=2086 As of mid-2017, the KB research blog has been discontinued. A static archive of the blog will remain available here.

]]>
http://blog.kbresearch.nl/2018/08/21/kb-research-blog-now-in-archive-mode/feed/ 0
Bridging the gap between quantitative and qualitative research in digital newspaper archives http://blog.kbresearch.nl/2017/01/25/bridging-the-gap/ http://blog.kbresearch.nl/2017/01/25/bridging-the-gap/#comments Wed, 25 Jan 2017 08:26:10 +0000 http://blog.kbresearch.nl/?p=2061 This blog post is written by Thomas Smits, KB Researcher-in-residence from May 2017

One of the central and most far-reaching promises of the so-called Digital Humanities has been the possibility to analyse large datasets of cultural production, such as books, periodicals, and newspapers, in a quantitative way. Since the early 2000s, humanities 3.0, as Rens Bod has called it, was posited as being able to discover new patterns, mostly over long periods of time, that were overlooked by traditional qualitative approaches.[1] In the last couple of weeks a study by a team of academics led by Professor Nello Christianini of the University of Bristol made headlines: “This AI found trends hidden in British history for more than 150 years” (Wired) and “What did Big Data find when it analysed 150 years of British history? (Phys.org). Did Big Data and Humanities 3.0 finally deliver on its promise? And could the KB’s collection of digitised newspapers be used for similar research?

The study, “Content analysis of 150 years of British periodicals”, is based on a corpus of 28.6 billion words, contained in 35.9 million articles of 120 regional, or local British newspapers from the period 1800-1950. [2] Focussing on six spheres – values and beliefs, UK politics, technology, economy, social change, and popular culture – the study is mostly based on the ‘use frequency’ of n-grams: the number of times a (combination of) word(s) appears in relation to all the words of the corpus in a specific year. For example, if a corpus consists of 100 words and the 2-gram, a combination of two words, “Digital Humanities” appears three times, the use frequency of this 2-gram is 0,03. In short: by applying n-grams the researchers were able to measure the relative importance of certain words, or combinations of words.

By using this method, the researchers were able to pinpoint specific historic events in their corpus, such as coronations, the election of a new pope, and outbreaks of several contagious diseases. More importantly, they used their method to test the validity of certain long-held notions about the nineteenth century. For example, the study suggests a very clear timeline in the emergence of the concept of “Britishness” in the popular imagination. While recent studies have posited that ‘national identity’ has deep historical roots, predating the nineteenth century, Christianini and his colleagues found that “British” overtook “English” only in the late-nineteenth century, supporting the close connection between the production of national identity, modernity, and the rise of mass media.[3]

Bias

While the researchers carefully composed their corpus, further contextualisation could make future research less biased. First of all, while the 120 newspaper titles studied in this research represent roughly 14% of all published titles, it remains unclear to what part of the press landscape these titles belonged. For example, the study neglects to discuss the fact that newspapers predominantly aimed to reach middle class readers. This bias is further enhanced by two factors: publications directed at lower classes, such as those of the chartist movement or the so-called penny press, were often deemed to be unworthy of archiving.[4] This process is enhanced by digitisation: well-known nineteenth century titles are more likely to be digitised than lesser-known, but not necessarily less influential, cheaper and/or radical publications.[5] Furthermore, by focussing on regional newspapers, the study aspires to mitigate the London-centric bias of research based on newspaper coverage. However, it neglects to account for the fact that in the first half of the nineteenth century many regional newspapers copied articles from London-based newspapers on a large scale, while syndication of content achieved, to some extent, the same result in the final decades of the century.[6]

Most crucially, the researchers seem to equate attention given in newspapers to historical significance. By doing so, they run the risk of failing to acknowledge the most important aspect of the medial form of the newspaper: its focus on ‘newsworthy’ events. The importance of certain long-term developments, which were never perceived as being radically new by contemporary commentators, can only be recognised with the benefit of hindsight. This leads to a somewhat paradoxical situation: digital newspaper archives are used to discover long-term trends, while newspaper discourse is mostly centred on short-term developments.

Delpher

What can researchers using the Delpher corpus learn from this study? The research department of the KB has already made important steps in the large-scale analysis of digitsed newspapers. Its open access n-gram viewer, developed by the University of Amsterdam, enables any user to replicate important parts of the British research. For example, a search for ‘nieuwe paus [new pope]’, yields the same results as the British study. In addition, recent projects of the KB’s fellows and researchers-in-residence use Dutch digitised newspapers in innovative ways. I hope that my own project, which applies two computer vision techniques to images in Dutch newspapers, can continue this tradition.

One could even ask the question why British digitised newspapers are not used more frequently for research similar to that of Christianini. Two private companies, Gale and FindMyPast, provide access to the collection of digitised historical newspapers, originally archived by the British Library. In a recent article, I compare these companies to Trove, the digital collection of Australian newspapers maintained by the government.[7] While public digital collections, such as Trove and Delpher, encourage users to tweak the archive, providing them with access to ‘raw’ data and API’s, private companies, such as FindMyPast, focus on a specific kind of use, in this case amateur historians studying their family history. This results in the fact that access to raw data is expensive, which makes it relatively hard, especially for junior researchers, to use it. In my opinion, we should take a critical look at the network of actors involved in the digitisation of newspapers, books, and other sources. Private companies, such as Google and FindMyPast, increasingly shape our access to the past and our use of historical sources. As I argue in my article, we should continue to discuss how this influences the ways that both researchers and the general public are able to interpret the past and relate to it in ways that are meaningful to them.

Bridging the gap

(Media) historians using qualitative methods would be wise to take note of the results of this study. It opens up a new world of possible research and shows how quantitative analysis can be used to substantiate existing theories. More importantly, the article raises the question if the strict separation between qualitative and quantitative research, or distant and close reading, is useful in distinguishing ‘traditional’ methods from their ‘digital’ counterparts. As the study amply shows, insights from traditional research are essential in defining the questions and contextualising both the corpus and the results of this kind of data-driven research. The project points to the importance of interdisciplinary research teams and, hopefully, will further undermine the trenches into which practitioners of the ‘digital’ and the ‘traditional’ humanities have grouped themselves.

[1] R. Bod, “Who is afraid of patterns? The Particular versus the Universal and the Meaning of Humanities 3.0” BMGN 128, no. 4 (2013): 171-80,  175.

[2] Landall-Welfare et al, “Content analysis of 150 years of British periodicals” PNAS (published ahead of print January 9, 2017). doi:10.1073/pnas.1606380114.

[3] This body of scholarship is mostly connected to Benedict Anderson’s concept of the imagined community. B. Anderson, Imagined Communities. London: Verso, 1983.

[4] For an explanation of what French press historian Jean-Pierre Bacot has called the ‘downward spiral of popularity’ of nineteenth-century newspapers and periodicals see Andrew King’s work on the London Journal: J.P. Bacot, La presse illustrée au XIXe siècle: une histoire oublié. Limoges: PULIM, 2005, 75 : A. King, The London Journal 1845-83: Periodicals, Production, and Gender. Aldershot: Ashgate, 2004, 16.

[5] A. Hobbs, “The Deleterious Dominance of The Times in Nineteenth-Century Scholarship” Journal of Victorian Culture 18, no. 4 (2013): 472-497.

[6] M. Beals, “Musings on a Multimodal Analysis of Scissors-and-Paste Journalism (Part 1),” accessed November 22, 2016, http://mhbeals.com/musings-on-a-multimodal-analysis-of-scissors-and-paste-journalism; B. Nicholson, “‘You Kick the Bucket; We Do the Rest!’: Jokes and the Culture of Reprinting in the Transatlantic Press” Media History 17, no. 3 (2012): 277-278.

[7] T. Smits, “Making the News National: Using Digitized Newspapers to Study the Distribution of the Queen’s Speech by W. H. Smith & Son, 1846–1858” Victorian Periodicals Review 49, no. 4 (2016): 598-625. DOI: 10.1353/vpr.2016.0041

]]>
http://blog.kbresearch.nl/2017/01/25/bridging-the-gap/feed/ 1
Detecting broken ISO images: introducing Isolyzer http://blog.kbresearch.nl/2017/01/13/detecting-broken-iso-images-introducing-isolyzer/ http://blog.kbresearch.nl/2017/01/13/detecting-broken-iso-images-introducing-isolyzer/#respond Fri, 13 Jan 2017 15:36:18 +0000 http://blog.kbresearch.nl/?p=2053

In my previous blog post I addressed the detection of broken audio files in an automated workflow for ripping audio CDs. For (data) CD-ROMs and DVDs that are imaged to an ISO image, a similar problem exists: how can we be reasonably sure that the created image is complete? In this blog post I will discuss some possible ways of doing this using existing tools, along with their limitations. I then introduce Isolyzer, a new tool that might be a useful addition to the existing methods.

Checksums

A number of techniques exist to verify a newly created ISO image. A seemingly obvious solution would be to do a checksum comparison on both the ISO image and the physical carrier. For instance, the following will work on any Linux system:

md5sum myimage.iso
md5sum /dev/sr0

The first line computes an MD5 checksum from the ISO image; the second line repeats this for the physical carrier. This method is not completely fail-safe. In some tests I did over a year ago, I ran into a a very strange issue where my attempts to image a CD would sometimes result in incomplete reads, and, as a result, truncated ISO images. The problem was most likely caused by faulty hardware (the machine on which I ran those tests more or less died shortly afterwards). Most worryingly, the machine would sometimes return incomplete data, both while creating the ISO image as well as during the subsequent checksum calculation on the physical carrier. The result of this was that the computed checksums were identical in both cases, which meant that the image passed the checksum quality check, even though it was incomplete!

Isovfy

The popular cdrtools library includes a tool called isovfy. Its man page describes it as follows:

isovfy is a utility to verify the integrity of an iso9660 image. Most of the tests in isovfy were added after bugs were discovered in early versions of mkisofs. It isn’t all that clear how useful this is anymore, but it doesn’t hurt to have this around.

I already commented on this tool in an earlier blog post:

The documentation of the tool isn’t very clear about what specific checks it performs. In one of my tests I fed it an ISO image that had its last 50 MB missing (truncated). This did not result in any error or warning message! Most of the reported isovfy errors that I came across in my tests simply reflected the file system on the physical CD not conforming to ISO 9660 (this seems to be pretty common).

You can try this yourself by running isovfy on the following two ISO images:

I ran both images through isovfy (version 3.02a06); both resulted in the following output:

Root at extent 17, 2048 bytes
[0,0]
No errors found

This demonstrates that isovfy is not very useful for detecting truncated ISO files.

Digging into the specs

At this point I decided it was time to start digging into some specs. The ISO 9660 page on the OSDev Wiki gives a good explanation of the internal organisation of an ISO 9660 image. From this I learnt that the Primary Volume Descriptor (which is a data structure that is present on all ISO images) contains two interesting fields:

  • Volume Space Size, which is the "number of Logical Blocks in which the volume is recorded";
  • Logical Block Size, which is "the size in bytes of a logical block".

In theory, multiplying both figures should give the expected size of the ISO image, and this would provide a useful way to check if data are missing. To test this, I wrote a Python script that parses an ISO’s Primary Volume Descriptor fields, calculates the expected file size and then compares this against the actual file size. Running the script against some 20 ISO images I had lying around showed that for 7 files the expected size was indeed identical to the actual file size. For most images, the actual size turned out to be marginally larger than expected (typically about 300-600 kB). For 3 images, the actual size was about twice the expected size. Digging deeper, I found out that these were hybrid images that contain an Apple partition on top of the ISO 9660 file system. According to this Wikipedia article, these hybrid discs come in two varieties:

  1. Hybrid discs that contain an Apple Partition Map (located at 512 bytes into the disc/image).
  2. Hybrid discs without a Partition Map. These contain a Master Directory Block (located at 1024 bytes into the disc/image).

In my case all of the 3 hybrid images turned out to be of the first category. Using the information here and here I was able to add detection of such hybrid images to my code, as well as a simple parser for the ‘zero block’ structure that contains two fields that define the partition’s size: Block Size and Block Count. For my hybrid images, multiplying both figures resulted in a value that was close to (but again marginally smaller than) the actual file size.

Finally, I also added detection of the second hybrid disc category (no Partition Map, but Master Directory Block). The Master Directory Block also contains Block Size and Block Count fields that allow one to calculate the size of the file system.

Isolyzer

I wrapped up the results of the above analyses into Isolyzer, which is a dedicated (Python) tool for checking the size of an ISO image. What it does is this:

  1. Locate the image’s Primary Volume Descriptor (PVD).
  2. From the PVD, read the Volume Space Size (number of sectors/blocks) and Logical Block Size (number of bytes for each block) fields.
  3. Calculate the expected file size as ( Volume Space Size x Logical Block Size ).
  4. If the image contains an Apple Partition Map, read the Block Size and Block Count fields from the ‘zero block’
  5. Calculate the expected file size as ( Block Size x Block Count )
  6. If the image contains an Apple Master Directory Block, read its Block Size and Block Count fields
  7. Calculate the expected file size as ( Block Size x Block Count )
  8. Calculate the final expected file size as the largest value out of any of the above 3 values
  9. Compare this against the actual size of the image files.

In addition to this, Isolyzer also extracts and reports technical metadata from the Primary Volume Descriptor and the Zero Block.

Currently the test results are reported in the following format (this may well change in upcoming releases):

<tests>
    <containsISO9660Signature>True</containsISO9660Signature>
    <containsApplePartitionMap>False</containsApplePartitionMap>
    <containsAppleHFSHeader>False</containsAppleHFSHeader>
    <containsAppleMasterDirectoryBlock>False</containsAppleMasterDirectoryBlock>
    <parsedPrimaryVolumeDescriptor>True</parsedPrimaryVolumeDescriptor>
    <sizeExpected>358400</sizeExpected>
    <sizeActual>358400</sizeActual>
    <sizeDifference>0</sizeDifference>
    <sizeAsExpected>True</sizeAsExpected>
    <smallerThanExpected>False</smallerThanExpected>
</tests>

In the above example the sizeExpected field is the size as calculated from the ISO/Apple headers, and sizeActual is the actual size. In this case both are identical. Below some output for a truncated ISO:

<tests>
    <containsISO9660Signature>True</containsISO9660Signature>
    <containsApplePartitionMap>False</containsApplePartitionMap>
    <containsAppleHFSHeader>False</containsAppleHFSHeader>
    <containsAppleMasterDirectoryBlock>False</containsAppleMasterDirectoryBlock>
    <parsedPrimaryVolumeDescriptor>True</parsedPrimaryVolumeDescriptor>
    <sizeExpected>358400</sizeExpected>
    <sizeActual>49157</sizeActual>
    <sizeDifference>-309243</sizeDifference>
    <sizeAsExpected>False</sizeAsExpected>
    <smallerThanExpected>True</smallerThanExpected>
</tests>

So, in this case sizeDifference is negative, and flag smallerThanExpected equals ‘True’ (which indicates a damaged image).

Feedback wanted

At this stage Isolyzer is a bit experimental and pretty rough around the edges, and I wouldn’t recommend it for production use. Nevertheless I’m curious about any feedback on the tool. Do others find this useful? Are things missing (i.e. other hybrid disc types I’m not aware of), or did I get anything completely wrong?

One thing that puzzles me a bit is that for the majority of ISO images I’ve come across, the expected size as calculated by Isolyzer is marginally smaller than the actual size. The difference is typically in the order of about 300-600 kB. I’m not quite sure what’s causing this, although this article mentions that some CD writing software packages add padding bytes when writing a CD. I wasn’t able to verify if this, although this SuperUser answer on validating a burnt DVD suggests it as well. If anyone knows more about this, please let me know!

Isolyzer can be found here on Github. It can be installed using pip; see the instructions here. For Windows users who cannot/don’t want to install Python I also provided stand-alone Windows binaries, which are available for download here.

]]>
http://blog.kbresearch.nl/2017/01/13/detecting-broken-iso-images-introducing-isolyzer/feed/ 0
Tackling problems and making progress http://blog.kbresearch.nl/2016/11/03/tackling-problems-and-making-progress/ http://blog.kbresearch.nl/2016/11/03/tackling-problems-and-making-progress/#respond Thu, 03 Nov 2016 15:57:11 +0000 http://blog.kbresearch.nl/?p=1994 Our current Researcher-in-Residence, Frank Harbers, is well under way with his project “Discerning Journalistic Styles. Exploring Automated Analysis of Journalism’s Modes of Expression”. In this blogpost he gives an update on his project and its progress.

Frank Harbers

It has been several months since I wrote the first blog about my work as researcher-in-residence and the research project is in full swing by now. The first phase of the project , connecting the metadata from my own database to the historical newspaper data (and metadata) in Delpher is finished and we are fully enveloped in the main part of the project: training a classifier to automatically determine the genre of historical newspaper articles.

The first phase was not as successful as we hoped, but we have managed to create a – modest – dataset to train the classifier. Initially, we hoped to be able to connect the metadata about approximately 33.000 Dutch newspaper articles to the data in Delpher. A crucial factor in the success of this attempt was the extent to which the segmentation of newspaper articles in Delpher matched the way the newspaper articles were segmented for the content analysis that resulted in the set of metadata about the historical newspaper articles. Unfortunately, it was far from a perfect match. For that reason the newspapers before 1945 could not be included – basically half of the metadata. Furthermore, De Volkskrant after the Second World War has not been digitized by the KB. In addition, the segmentation of De Telegraaf in the postwar period was so different that we couldn’t include that either. In the end, this meant that we could only use the data of Algemeen Handelsblad/NRC Handelsblad in the postwar period. So quickly we saw our dataset shrink from the potential 33.000 articles to a modest 2000 articles. A bit of a setback, but fortunately we can still use this smaller dataset to train a genre classifier. This experience does make clear how crucial segmentation is for the creation of datasets that can be fruitfully used for digital humanities research into journalism history.

At the moment, we are working on the second phase of the project. We have identified several genres that we would like to classify. These genres, such as news reports, reportages, interviews, opinion articles, reviews, news analyses, can shed light on the way journalism developed from a reflective, opinion-oriented way of doing journalism to a more event-centered and fact-oriented journalism practice. At the core of this part is the translation of the genre definitions to clear linguistic markers that can be identified automatically. Take for instance the news report, a genre that is defined by the use of the inverted pyramid (a story structure in which typical journalistic questions, like Who, What, Where and When, are answered in the first paragraph. Moreover, it often contains direct quotes from sources and is generally a fairly concise article written in a depersonalized, objective style. Question is how you can recognize these features automatically in the text. In this case, the quotes can be recognized by the presence of quotation marks (for which a high quality OCR is crucial) and we will attempt to identify the inverted pyramid structure by using named identity recognition to see whether questions concerning who was involved and where and when it happened are answered. We hope the depersonalized style can be captured by looking at the lack of a first person perspective (the use of the pronoun ‘I’ or ‘We’) and the lack of adjectives that create a colorful and subjective account.

Juliette Lonij, programmer on this project, is currently developing the Python software to extract the features on which the classifier will run. She looked into different natural language processing software packages to pre-process the article texts and chose to use FROG for tokenization and Part-of-Speech tagging, which facilitates our research needs quite well (other packages might be added in the future). And today, we have just run a first exploratory test with the classifier, which showed promising results. In the coming weeks we will keep on testing and refining the classifier. So wish us luck!

 

]]>
http://blog.kbresearch.nl/2016/11/03/tackling-problems-and-making-progress/feed/ 0
FAQ Call for Proposals Researcher-in-Residence http://blog.kbresearch.nl/2016/06/08/faq-call-for-proposals-researcher-in-residence-2017/ http://blog.kbresearch.nl/2016/06/08/faq-call-for-proposals-researcher-in-residence-2017/#respond Wed, 08 Jun 2016 09:48:50 +0000 http://blog.kbresearch.nl/?p=1813 Updated 04 June 2018

I don’t live or work in the Netherlands. Can I apply? 
Probably! Contact us at dh@kb.nl and we’ll discuss your options.

I want to use my own dataset. Is that possible?
Sure! As long as you also use one of the datasets of the KB and it doesn’t limit the publication of the project end results.

I don’t know how to code, is that a problem?
Not at all. We have skilled programmers who can help you with your project or we will try to find a match for you if you prefer someone else. This would mean submitting as a team and will cut the budget in half. Reach out to us to discuss the options.

I don’t speak Dutch. Is your content still interesting to me?
That depends on your research question :) It might not be so appealing to linguists, but could offer an novel collection for computer scientists. Contact us to see which collections we have and we can discuss what might be the most interesting set for you.

Why will you publish my abstract?
We want to show others what types of proposals we have received to offer future researchers an insight into the selection process and to prevent them from entering a similar project.

Can I submit a project I’ve submitted previously (at another institution)?
We’d like you to submit an original idea. It can be one you have had lying around for some time, but we’d appreciate projects that haven’t been done before. Projects that have been previously entered into a similar program should be changed significantly before resubmitting.

Can I also work fulltime on my project for a period of 3 months?
We prefer you to work part-time so you can spend a total of 6 months with us. This also allows you to continue your research or teaching obligations at your university.

Will you be able to reimburse any housing or hotel costs?
Unfortunately, when you come from outside the Netherlands, we are not able to find and fund your housing or pay for your travel expenses to the KB. However, we do fund travel costs within the Netherlands allowing you to come to and work in the KB, the Hague wherever you are based in the Netherlands.

I want to use my own programmer, can I?
Yes, you can. We even encourage you to bring in extra people when you want to address a subject we’re not experts in (such as multimedia). However, the budget remains the same, so it will have to be split between you. We do ask that the whole team is available in the KB for at least one day a week. If you want to know whether we can help you or if you should bring someone in, please contact us at dh@kb.nl.

I don’t know if my idea is what you’re looking for. What can I do?
You are welcome to contact us at dh@kb.nl to discuss your ideas and the possibilities.

Can I submit more than one project?
Please focus your efforts on one great project.

Who will be judging the entries?
The entries will be judged by an internal committee and then forwarded to an external committee of representative experts from several Dutch universities and institutions.

What will you judge my project on?
We will judge the entries on criteria such as feasibility (technically, legally and practically), how the KB data will be showcased and used and whether the end results are of use for a wider community. Next to this, we will also look at the originality and quality of the proposal and the amount of support needed (and in this case, more is not necessarily worse!).

What happens if you submit a plagiarized project?
When we notice your your project is plagiarized, we will not consider your application for placement. You are responsible for the originality and authenticity of the project, but we will keep our eyes open.

What happens to any software I write for my project?
All software in the projects, whether you or we write it, will be made available on the KB Lab and Github page under an open source license.

What happens to the data I collect/produce in my project?
At the KB Lab we try to be as open as possible. All data produced in the programme is to be made available for research purposes, either through the KB Lab, KB Data Services or DANS, and where possible will receive a CC-license.

Can I publish any papers about the project?
Yes, we even encourage you to do so. If necessary, we’re happy to help.

]]>
http://blog.kbresearch.nl/2016/06/08/faq-call-for-proposals-researcher-in-residence-2017/feed/ 0
Terms and conditions of the KB Researcher-in-residence programme 2017 http://blog.kbresearch.nl/2016/06/08/terms-and-conditions-of-the-kb-researcher-in-residence-programme-2017/ http://blog.kbresearch.nl/2016/06/08/terms-and-conditions-of-the-kb-researcher-in-residence-programme-2017/#respond Wed, 08 Jun 2016 09:48:41 +0000 http://blog.kbresearch.nl/?p=1822 This programme as detailed at the KB-website (“Programme”) is operated by the Koninklijke Bibliotheek, National Library of the Netherlands (“KB”), Prins Willem-Alexanderhof 5 (2509 LK) Den Haag, The Netherlands.

  1. General
    1. If you enter the Programme you agree to abide by all of the following Terms and Conditions under which the Programme is run, and to be bound by them. We may disqualify you without prior notice if you are in breach of any of these Terms and Conditions.
    2. KB reserves the right to cancel the Programme at any stage if KB deems this necessary or circumstances arise that are outside of our control.
  2. Entries
    1. The Programme is open to all excluding employees or contractors of the KB or their direct family members.
    2. The Programme is only open to Ph.D.-students, or applicants who have obtained a Ph.D.-degree between 2011 and 2016, and who are aged 18 years or older.
    3. The Programme is only open to researchers employed by a university or a research institute within the EU.
    4. Entry is free.
    5. To enter the Programme, you must:
      1. Submit your proposal in accordance with the Call for proposals via the KB website; and
      2. Ensure that (i) your proposal is original (you can submit a proposal you have submitted previously as long as it has not been used in another programme or fellowship) and (ii) your proposal does not in any way infringe the copyright or other intellectual property rights, or other rights of any third party.
      3. Ensure that your proposal is aimed at the use of KB-data that benefits your research and other users of KB and/or the KB Research Lab.
      4. Tick the box to accept these Terms and Conditions when you submit your entry.
    6. All entries must be received by midnight on 8 November 2015. Any entries received after this will not be considered.
    7. The entries will be judged by an internal committee and then forwarded to an external committee of representatives experts from several Dutch universities and institutions.
    8. You may be contacted by email or telephone to answer further questions about your entry up to 15 December 2015. Two entries and two backup-entries will be chosen for the final stage. The best two entries will be contacted personally by email to assess whether those applicants are available for the secondment-period.
  3. Rights and permissions
    1. All entrants accept and agree that KB may publish the title and abstract of your proposal on the KB Research Blog at the KB’s sole discretion by January 2016. This blog is harvested for the Dutch web archive.
    2. In the event that you are offered a secondment:
      1. you agree to enter in a written agreement with KB and the university or research institute where you are employed on, among other things, intellectual property rights, and the obligations of the parties concerned including the obligations as set out in these Terms and Conditions.
      2. you agree to participate in any publicity planned by KB, if required;
      3. you will be expected to work on your project at least one day per week in residence at the KB in Den Haag, between 1 January 2017 and 30 June 2017 or 1 July 2017 and 31 December 2017, for 0.5 fte.
      4. you agree to present your research results to employees of the KB in the final month of your secondment.
      5. you will be expected to publish a blog about your project on the KB Research Blog.
      6. you agree to submit a Certificate of Conduct for Natural Persons (‘VOG NP’).
      7. you agree to mention the secondment KB offered you in any publication your secondment will give rise to.
    3. If any software is produced for your research, it will be developed on the principles of open source software and it will be made available on the KB Research Lab and Github website under a GPLv3
    4. KB offers you access to all data sets of the KB and support from a programmer, collection specialists and data specialists.
    5. During the secondment, KB offers you an office space and payment of all travel expenses within the Netherlands regarding the secondment.
    6. In case you do not live in the Netherlands, you are responsible to find and fund your own housing and pay for travel expenses to the KB in Den Haag, the Netherlands.
  4. Personal data
    1. Other than as expressly permitted by these Terms and Conditions, KB will only use your contact details for the purposes of administering this Programme, and will not publish them or provide them to anyone without your permission.
    2. You consent to KB holding and processing data relating to you for legal, administrative and management purposes.
    3. Any personal data relating to you will be used solely in accordance with the current data protection legislation in the Netherlands (Wet bescherming persoonsgegevens) and will not be disclosed to another party – except for the external committee of representatives experts from several Dutch universities and institutions – without your prior consent. Please see the Privacy statement KB for further details.
    4. Data relating to you will be retained by KB for a reasonable period after the closing date specified in clause 2.6 to assist in the administration of the Programme in a consistent manner and to deal with any queries on the Programme.
  5. Applicable Law
    1. These Terms and Conditions and any dispute or claim arising out of or in connection with them shall be governed by and construed in accordance with Dutch law.
    2. The parties will attempt in good faith to resolve any dispute or claim arising out of or relating to these Terms and Conditions promptly by negotiation. If the dispute cannot be resolved by negotiation, you hereby agree that the sole jurisdiction and venue for any actions that may arise in relation to the subject matter hereof shall be the Dutch Court in Den Haag, the Netherlands
]]>
http://blog.kbresearch.nl/2016/06/08/terms-and-conditions-of-the-kb-researcher-in-residence-programme-2017/feed/ 0
KB at DHBenelux 2016 http://blog.kbresearch.nl/2016/06/07/kb-at-dhbenelux-2016/ http://blog.kbresearch.nl/2016/06/07/kb-at-dhbenelux-2016/#respond Tue, 07 Jun 2016 07:27:07 +0000 http://blog.kbresearch.nl/?p=1774 This week, the annual DHBenelux conference will take place in Belval, Luxembourg. It will bring together practically all DH scholars from Belgium (BE), the Netherlands (NE) and Luxembourg (LUX). You can read the full program and all abstracts on the website. Two presentations are by members of our DH team (Steven Claeyssens & Martijn Kleppe) and one presentation is by our current researcher in residence (Puck Wildschut – Radboud University Nijmegen). Please find the first paragraphs of their abstracts below:

  • Puck Wildschut – Roles, relations and references: Towards a computation-based distant reading of narrative-semantic roles in large datasets in Dutch

Since the rise of Russian Formalism in the early 19th century, literary theorists have been interested in finding ways to detect actants (characters) in narratives. The recent rise of computational methods within the humanities offers new ways of tackling this issue. As researcher-in-residence at the National Library of the Netherlands (KB) I am currently involved in a research project1 that aims to develop a computation-based model for analyzing narrative-semantic roles in large datasets in Dutch. The premise of the project is that actantial roles are not only to be detected on the higher level of motive- and theme-building, but also at the linguistic level of semantic roles. Furthermore, the project aims to not only develop a tool for the detection of actantial roles, but also, and more importantly, for discovering the relationships between those roles as they are encoded in language. (…) The poster presentation will show the most up-to-date version of the tool we are developing at the KB and the preliminary results of its implementation. Special prominence will be given to how the issues mentioned are integrated in the tool’s development. The poster aims to show how literarylinguistic theory and computational practice encourage each other in the development process

Full abstract (in pdf) here.

  • Steven Claeyssens – The Ideal Corpus. Towards a Critique of Large Digital Libraries from a Digital Humanities Perspective

Large scale digitisation of historical paper publications enables analyses of vast amounts of digital surrogates using machines, algorithms and software to ‘read’ the texts. However, continuously expanding collections of such texts, like the Google Books corpora, HathiTrust or – closer to home – Delpher, combined with a proliferation of computational approaches to study textual data and the growing number of scholarly disciplines taking part in the Digital Turn, calls for a renewed reflection on two of the key aspects of the research process: source selection and source criticism.  (…) This paper argues that a digital source criticism is urgently needed to tackle the questions raised by these opposing expectations. Researchers and librarians should collaborate closely on this and join forces to define the limits and fits of ‘the ideal corpus’. Inspiration for this definition can, amongst other things, be found in textual criticism, corpus linguistics and analytical bibliography

Full abstract (in pdf) here.

  • Martijn Kleppe & Desmond Elliott – Doing Visual Big Data – Creating the KBK-1M Dataset Containing 1,6 Million Newspaper Images Available for Researchers

The visualisation of news through photographs has exploded since the second half of the 20th century (Kester & Kleppe 2015). However, methods that are employed to analyse the (re)use of visual materials are labour-intensive because Humanities researchers tend to analyse their sources manually (Burke 2001). To estimate the increase in the use of pressphotographs in Dutch newspapers, Kester & Kleppe (2015) e.g manually analysed a sample of 385 newspapers and 5.877 press photographs over the period 1870-2013. To find the recurring use of photographs in Dutch history textbooks, Kleppe (2012) followed a same approach by manually analysing over 5.000 photographs in 400 history textbooks, creating the ‘Foto’s in Nederlandse Geschiedenisschoolboeken (FiNGS) (Photos in Dutch History textbooks) dataset (Kleppe 2013b). Even though manually created and annotated datasets such as FiNGS contain rich & well-annotated data, their scope remains limited given its labour-intensive creation and analyses process. Therefor this poster presents the KBK-1M dataset, that was created specifically for (Digital) Humanities researchers. This dataset contains a collection of 1.603.395 captioned images extracted from Dutch digitised newspapers stored in the Dutch National Library (KB) Newspaper archive of the period 1922-1994. On our poster, we will describe how we obtained the images, what types of research questions it could tailor and how researchers can obtain the dataset for their research purposes.

Full abstract (in pdf) here. See poster below.KBK-1M Poster A1When you are at DHBenelux and if you would like to meet our colleagues, please feel free to approach them during the meeting or at one of the sessions Steven (Digital Textual Analysis II) and Martijn (Digital Art & Culture I) are chairing.

Finally, we are also very happy that during DHBenelux we will launch the call for our Researcher in Residence program 2017 that allows young Digital Humanities researchers to come and work with us in 2017. All details on the call are now online at http://blog.kbresearch.nl/2016/06/08/call-for-proposals-kb-researcher-in-residence-2017/ 

 

 

]]>
http://blog.kbresearch.nl/2016/06/07/kb-at-dhbenelux-2016/feed/ 0
Valid, but not accessible EPUB: crazy fixed layouts http://blog.kbresearch.nl/2016/04/04/valid-but-not-accessible-epub-crazy-fixed-layouts/ http://blog.kbresearch.nl/2016/04/04/valid-but-not-accessible-epub-crazy-fixed-layouts/#respond Mon, 04 Apr 2016 10:13:18 +0000 http://blog.kbresearch.nl/?p=1672 EpubCheck is an invaluable tool for assessing the quality of EPUB files. Still, it is possible that EPUBs that are valid according to the format specification (and thus EpubCheck) are nevertheless inaccessible to some users. Some weeks ago a colleague sent me an EPUB 2 file that produced some really strange behaviour across a number of viewer applications. For a start, the text wouldn’t reflow properly after re-sizing the viewer window, and increasing the font size resulted in garbled text. Running the file through EpubCheck did return some validation errors, but none of these were related to the behaviour I was getting. Closer inspection revealed some very peculiar stylesheet and HTML use.

Crazy Fixed Layout

As I cannot share the original file for rights reasons, I fired up the Sigil e-book editor and made a handcrafted EPUB that reproduces its behaviour. You can download the file here. If you open it in an e-book viewer, it will probably look perfectly normal at first sight. For example, here’s a screenshot I made using the Calibre viewer:

calibre_normal

Next I reduced the width of the viewer window. One would expect the text to re-flow to the new width. Instead this happened:

calibre_resized_screen

After increasing the font size, I ended up with this:

calibre_largefont

I got similar results in Chome’s Readium extension. On my e-Ink reader, a Sony PRS-T2, the book rendered as follows:

sony_fixedlayout

However, I wasn’t able to change the font size.

Analysis

The file passes validation in EpubCheck 4.0.1 without errors. However, the output does contain a series of warnings about the use of absolute positions in a stylesheet:

CSS-017, WARN, [CSS selector specifies absolute position.], OEBPS/Styles/styles.css (6-2)
CSS-017, WARN, [CSS selector specifies absolute position.], OEBPS/Styles/styles.css (24-1)
CSS-017, WARN, [CSS selector specifies absolute position.], OEBPS/Styles/styles.css (43-1)
::

To really understand what causes the problem, we need to look inside the file’s HTML and CSS resources. Here’s some of the HTML that underlies the text:

<p id="p01" class="para">This is an <em>EPUB</em> 2 file that uses a fixed layout.</p>
<p id="p02" class="para">This is achieved by placing each line inside a</p>
<p id="p03" class="para"><em>paragraph</em> element. Each <em>paragraph</em> element</p>
<p id="p04" class="para">is placed at a fixed position on the page. Even</p>
<p id="p05" class="para">though this file is valid <em>EPUB</em>, this is a pretty</p>
<p id="p06" class="para"> terrible idea, because in most readers the text</p>
<p id="p07" class="para">will not reflow after resizing the viewer window.</p>

So, every line is wrapped inside a paragraph element, each of which has a unique id selector. These refer to style definitions in the EPUB‘s stylesheet. Here are the definitions for the first two lines:

#p01
{
position:absolute;
left:40px;
top:80px;
letter-spacing:0.42px;
word-spacing:0.1em;
}
#p02
{
position:absolute;
left:40px;
top:120px;
letter-spacing:0.42px;
word-spacing:0.1em;

Each style definition specifies a line’s position on the canvas (left, top); moreover, these co-ordinates are defined as absolute positions. This means that each line is placed at a fixed position, regardless of whether this makes any sense given the actual dimensions of the viewer window (or device), or the user’s preferred font size. It seems that the intention of the producer of the original EPUB (from which I derived my example) was to create some sort of “fixed layout” document. However, this doesn’t make much sense for books with simple, text-only layouts (as in this case). Worse, depending on the viewing device and the user the file may be effectively inaccessible. For example, someone with a visual impairment may only be able to read an EPUB using very large font sizes, which in this case results in garbled text.

Crazy Columns

Things can even get worse. I once came across an EPUB that used similar tricks to achieve a two-column layout. Again I’m not able to share the original file, so I created another EPUB that mimicks its behavour. In the Calibre viewer it looks like this:

calibre_columns

As with the first example, the text doesn’t reflow after resizing the viewer window, and increasing the font resulted in this:

calibre_columns_largefont

This is what I got when I opened the file in my Sony e-Ink reader:

sony_crazycolumns1

After I increased the font size this happened:

sony_crazycolumns2

Similarly, when I tried to copy the text in the file to the clipboard, and then pasted it in a text editor, I ended up with this:

This is an EPUB filepage. Even though thisthat uses a two-columnfile is valid EPUB, there’slayout. For each column,no way to establish theevery line is placed atlogical reading order ofa fixed position on thethe text.

Ouch!

Analysis

Again, throwing this file at EpubCheck 4 doesn’t result in any validation errors, although just like the previous file there are some warnings about the use of absolute positions in the stylesheet:

CSS-017, WARN, [CSS selector specifies absolute position.], OEBPS/Styles/styles.css (13-1)

A peek inside the HTML reveals the true horrors of this EPUB. This is how the text is encoded:

<div class="pos" style="left: 40px; top: 100px;">This is an <em>EPUB</em> file<div>
<div class="pos" style="left: 260px; top: 100px;">page. Even though this</div>
<div class="pos" style="left: 40px; top: 140px;">that uses a two-column</div>
<div class="pos" style="left: 260px; top: 140px;">file is valid <em>EPUB</em>, there's</div>

So, every line of each column is wrapped in a division element that has a fixed position. The class pos in the stylesheet defines the general layout of each division element. In this case, it specifies that all positions are (again) absolute:

.pos {position:absolute;
} 

Technically this is pretty similar to the first example. Note that the above HTML doesn’t contain any semantic information on the fact that there are two separate columns. Worse, the order of the text in the HTML doesn’t even follow the actual reading order! This also explains the results after copying and pasting. Screen reader applications will not be able to handle this either, which makes books like these inaccessible to many visually impaired users. All of this could have been avoided if the book’s producer had followed the W3C multi-column layout specification.

Conclusion

I don’t know how common (or rare) EPUBs like the above are. They may just be weird edge cases. Nevertheless, their existence indicates that checking for validity alone may not be sufficient to ensure accessibility for all users (in particular those with a visual impairment). In any case, files like these can be identified relatively easily by checking EpubCheck‘s output for the presence of a CSS-017 warning (“CSS selector specifies absolute position”)1. These examples also underline the importance of guidelines and best practices. Several good resources for making accessible EPUB are available from the EPUB 3 Accessibility Guidelines, including a useful Accessibility QA Checklist. I would also be interested in hearing other people’s experiences with “weird” EPUBs like these.

Postscript

Alberto Pettarin pointed me to his blog post (Current) Fixed Layout eBooks Considered Harmful. Written in 2015, it addresses the problems with current implementations of fixed layouts in EPUB, and if you found this blog post interesting, I would suggest to check out Alberto’s blog as well.

Alberto’s Twitter feed also drew my attention to an interesting EPUB with the program of the recent EPUB Summit in Bordeaux. You can download it here (you need to unzip it first!). The file is interesting because:

  1. It does not pass validation by EpubCheck (the mimetype file entry is not the first file resource in the archive)
  2. It uses a fixed, multi-column layout that doesn’t scale in either Readium or Calibre‘s viewer (changing the font size has no effect), and I’m wondering if it is usable at all on any handheld devices!

There’s some irony in that this file was published by EDRLab, an organisation that describes itself as “the European headquarter for IDPF and Readium Foundation”, and which mentions “support for people who have print disabilities” as a “key part”of its mission. Oh well …

The EPUBs used for this blog post are part of the EPUB KB policy testing repository. This is an annotated set of openly licensed EPUB files that were specifically created for testing purposes.


  1. Note that EpubCheck 3 (now outdated) does not report this warning, so always use EpubCheck 4.
]]>
http://blog.kbresearch.nl/2016/04/04/valid-but-not-accessible-epub-crazy-fixed-layouts/feed/ 0
The future of EPUB? A first look at the EPUB 3.1 Editor’s draft http://blog.kbresearch.nl/2016/03/10/the-future-of-epub-a-first-look-at-the-epub-3-1-editors-draft/ http://blog.kbresearch.nl/2016/03/10/the-future-of-epub-a-first-look-at-the-epub-3-1-editors-draft/#comments Thu, 10 Mar 2016 16:51:15 +0000 http://blog.kbresearch.nl/?p=1666  

About a month ago the International Digital Publishing Forum, the standards body behind the EPUB format, published an Editor’s Draft of EPUB 3.1. This is meant to be the successor of the current 3.0.1 version. IDPC has set up a community review, which allows interested parties to comment on the draft. The proposed changes relative to EPUB 3.0.1 are summarised in this document. A note at the top states (emphasis added by me):

The EPUB working group has opted for a radical change approach to the addition and deletion of features in the 3.1 revision to move the standard aggressively forward with the overarching goals of alignment with the Open Web Platform and simplification of the core specifications.

As Gary McGath pointed out earlier, this is a pretty bold statement for what is essentially a minor version. The authors of the draft also mention that they expect it “will provoke strong reactions both for and against”, and that changes that raise “strong negative reactions” from the community “will be reviewed for future drafts”.

This blog post is an attempt to identify the main implications of the current draft for libraries and archives: to what degree would the proposed changes affect (long-term) accessibility? Since the current draft is particularly notable for its aggressive removal of various existing EPUB features, I will focus on these. These observations are all based on the 30 January 2016 draft of the changes document.

Removed support for EPUBCFI for linking

The EPUB Canonical Fragment Identifier (EPUBCFI) “defines a standardized method for referencing arbitrary content within an EPUB Publication”. Until EPUB 3.0.1, Reading Systems were required to support EPUBCFI for hyperlinking within and between documents. This requirement is dropped in EPUB 3.1 (although it would still be possible to use EPUBCFI for annotations and bookmarks).

In principle this change could result in problems if an EPUB that uses CFI for hyperlinks is opened in a 3.1 reading system: in that case the hyperlinks would not work. However, according to EPUB editor Matt Garrish, authors simply do not use CFI for hyperlinking. He also mentions a check by Google on their corpus of millions of books, which only turned up a few instances of CFI use. One of these was a link in an EPUB best practices book, while the remaining ones were all part of the EPUB test suite documents. If these results are representative of all EPUBs “in the wild”, the implications of the change would be negligible.

Reduced set of metadata elements in Package Document

EPUB 3.1 imposes restrictions on the metadata elements that can be embedded in the Package Document. Up to version 3.0.1, the full Dublin Core Metadata Element Set was supported, whereas in 3.1 only the dc:identifier, dc:title, dc:language, dc:creator, dc:publisher and dc:type elements are allowed. Additional metadata can be included, but they need to be defined in a separate resource (file), which is referenced from the metadata element using the link element. Below is an example that uses a MARC file:

<link rel="record"
  href="meta/9780000000001.xml" 
  media-type="application/marc"/>

Complicating things further, the EPUB 3.1 Packages draft says:

Linked resources that are not Publication Resources are not subject to Core Media Type requirements [EPUB31] and may be located inside or outside [EPUB31] the EPUB Container. Retrieval of Remote Resources is optional.

So, linked metadata resources can have any possible format, and they may not even be included in the EPUB container. Even though these changes would have no direct consequences for long-term accessibility, they would seriously complicate document processing (e.g. ingest) workflows that rely on the metadata in the Package Document. It would also affect end users who rely on these metadata fields to sort and find their ebooks.

Note: the discussion thread on this topic in the issue tracker is worth checking out, as it contains some excellent additional observations.

Removal of the NCX

EPUB 2 documents contain the NCX file (“Navigation Control file for XML”), which provides a mechanism to navigate a publication. It is essentially a hierarchical table of contents. The NCX was superseded by the Navigation Document in EPUB 3.0.1. However, the NCX was allowed in EPUB 3.01 publications, which was useful for keeping EPUB 3 publications compatible with older (EPUB 2-based) reading systems1. The 3.1 draft forbids the NCX altogether, which means that such “hybrid” EPUBs are not possible without breaking the specification.

The main consequence of this is that it would make EPUB 3.1 files incompatible with older reading systems. More specifically, basic navigation functionality such as direct access to a chapter from the table of contents would not work.

To get an approximate idea of the impact of this, I had a look at the EPUB 3 support grid, which gives detailed information about the support of specific EPUB 3 features for commonly used devices, apps, and reading systems. This link shows support of the toc nav element, which defines the primary navigational hierarchy in the Navigation Document. Only 55% (34 out of 62) of all tested reading systems fully support the toc nav element, with 37% (23 out of 62) not supporting it at all2. This may not be a big deal for users of software-based reading systems (which make up the majority of the support grid), but users of (older) E-ink readers often don’t have the option to upgrade their devices. A good example is this (now discontinued) Sony e-Ink hardware reader. Unfortunately, E-ink devices appear to be underrepresented in the support grid. For example, it contains no information whatsoever on any of the popular Kobo readers.

The proposal to remove the NCX provoked strong reactions in the community review, with one respondent stating it would lead to “dropping support for millions of eInk reading systems”. It would also contradict this statement from the EPUB 3.0.1 specification (emphasis added by me):

The NCX feature defined in [OPF2] is superseded by the EPUB Navigation Document [ContentDocs301]. EPUB 3 Publications may include an NCX (as defined in OPF 2.0.1) for EPUB 2 Reading System forwards compatibility purposes, but EPUB 3 Reading Systems must ignore the NCX.

The explicit reference to EPUB 3 Publications (not EPUB 3.0.1 Publications!!) implies that the statement applies to EPUB 3 in general. Removing the NCX in another EPUB 3 release would be at odds with this.

Removal of the guide Element

The guide element was an optional data structure in EPUB 2 that provided “convenient access” to structural components of a publication. It was deprecated in EPUB 3.0.1. Without any data on the actual usage of this feature, it is difficult to say much about the impact of its complete removal (this was also pointed out by one respondent to the community review).

Removal of the bindings Element

In EPUB 3.0.1 the bindings element could be used to define fallbacks for foreign resources. According to EPUB editor Matt Garrish “this feature is not widely used or supported”, and the impact on accessibility appears to be negligible.

Removal of the switch Element

The switch element in EPUB 3.0.1 allows one to define alternative representations of XML fragments. Here’s an example:

<epub:switch id="cmlSwitch">
   
   <epub:case required-namespace="http://www.xml-cml.org/schema">
      <cml xmlns="http://www.xml-cml.org/schema">
         <molecule id="sulfuric-acid">
            <formula id="f1" concise="H 2 S 1 O 4"/>
         </molecule>
      </cml>
   </epub:case>
   
   <epub:default>
      <p>H<sub>2</sub>SO<sub>4</sub></p>
   </epub:default>
   
</epub:switch>

Here, we have a chemical formula in ChemML format and in standard HTML. ChemML is not natively supported in EPUB, so by default a reader will display the HTML version. However, wrapping both in a switch element would allow a ChemML-capable reader to render that representation instead.

I asked EPUB editor Matt Garrish how an EPUB 3.1-compliant reader would render content that is wrapped in a switch element. He replied that by default all of the switch content would be rendered. So for the example above, a reader would try to render both the HTML and the ChemML versions (with the latter failing on most reading systems). Matt stressed the significance of the switch element, adding that people have been using it, “if not extensively”.

Removal of the trigger Element

The trigger element in EPUB 3.0.1 is used to define simple user interfaces for multimedia content. Since this can be done natively in HTML 5, it is dropped from EPUB 3.1. Here editor Matt Garrish explains that the feature is both “sparsely used” (referring to a survey of publishers) and “poorly supported”.

Miscellaneous changes

Apart from the changes above (which all remove features from the existing specification), the EPUB 3.1 draft also adds a number of new features, and clarifies some existing ones. I won’t go over them in detail, but here’s a brief overview:

Finally, the draft contains clarifications on Foreign Resource Fallbacks and Scripting Support.

EPUB 3.1 or EPUB 4.0?

By now it should be clear that the aggressive removal of features in EPUB 3.1 would have some far-reaching consequences. This is particularly true for the removal of the NCX, which would make EPUB 3.1 files incompatible with many existing E-ink readers. It would do this by ruling out the option to make backward-compatible “hybrid” files. As Gary McGath pointed out earlier, introducing “radical changes” in what is essentially a minor version is pretty unusual practice for any standard. Nowadays, most software and file formats use some variation of semantic versioning, with version numbers that follow the general form MAJOR.MINOR.PATCH. Here, each component of the version number has a well-defined meaning:

  1. MAJOR version is increased in case of incompatible API changes,
  2. MINOR version is increased when functionality is added in a backwards-compatible manner, and
  3. PATCH version is increased in case of backwards-compatible bug fixes.

Since the current draft includes multiple backward-incompatible changes, this makes me wonder why the editors didn’t name it EPUB 4.0 instead! Kovid Goyal, lead developer of the popular Calibre software, made the following comment on this:

[I]f you want to make backwards incompatible changes, please, dont do it in a point release. From glancing over your changes document, it seems to me that you want to make several breaking changes. That’s great, EPUB 3 could do with some serious breaking. But name it EPUB 4. I really dont want to have tell my users that calibre supports EPUB 3.1 but not EPUB 3.

I agree with Kovid here. Having multiple sub-versions of EPUB 3, with some of them being backward-compatible with EPUB 2, while this backward compatibility is explicitly ruled out in another sub-version, is bound to create a situation that will be incomprehensible for most e-book buyers. Worse, it could even undermine overall confidence in the format. For memory institutions it would also make the management of EPUB 3 publications unnecessarily complicated. Not only would some EPUB 3.1 files not render correctly in an EPUB 3.0.1 reader, the opposite would be true as well.

Flashback

In my 2012 report on EPUB for archival preservation I already mentioned the stability of the EPUB format as a concern:

EPUB 3 shows quite major changes relative to version 2, which raises concerns about the format’s stability over time. These concerns are reinforced by the fact that EPUB 3 is heavily dependent on (X)HTML5 and CSS3, both of which are unfinished “works in progress”, which may undergo various changes before being finalised.

These concerns are once more confirmed by the current EPUB 3.1 draft. However, it remains to be seen how many of these changes will make it to the final version. The community review process is ongoing at this moment, so if you’re getting a little uneasy after reading this blog post, there’s still time to get involved and make your voice heard!

Acknowledgement

Thanks to Matt Garrish for his prompt replies to my questions on Github.


  1. See here how O’Reilly’s keeps their EPUB 3 books compatible with EPUB 2 readers
  2. This figure includes reading systems for which support is unknown
  3. See the HTML5 Reference for a discussion of the differences between both syntaxes

 

]]>
http://blog.kbresearch.nl/2016/03/10/the-future-of-epub-a-first-look-at-the-epub-3-1-editors-draft/feed/ 2
Jpylyzer 2015 round-up http://blog.kbresearch.nl/2015/12/08/jpylyzer-2015-round-up/ http://blog.kbresearch.nl/2015/12/08/jpylyzer-2015-round-up/#respond Tue, 08 Dec 2015 14:49:11 +0000 http://blog.kbresearch.nl/?p=1547

Yesterday (7 December) we released version 1.16.0 of the jpylyzer tool, which is this year’s third release of the software (excluding bugfix releases). This blog post gives a brief overview of the main jpylyzer improvements that have been implemented over this year.

Changes in XML output

The 1.14 release introduced two output improvements. Most importantly, an XML Schema Definition (XSD) was created. The schema formally defines the output format, and it also makes it possible to validate output files. In addition, a namespace declaration was added. These changes make the post-processing of jpylyzer‘s output more straightforward.

The 1.16 release added the statusInfo element, which tells you whether the validation completed without any internal errors. It contains the following sub-elements:

  • success: a Boolean flag that indicates whether the validation attempt
    completed normally (“True”) or not (“False”). A value of “False” indicates
    an internal error that prevented jpylyzer from validating the file.
  • failureMessage: if the validation attempt failed (value of success
    equals “False”), this field gives further details about the reason of the failure.

This means that the general structure of the output now looks like this:

outputStructure

Recursive traversal of directory trees

Another feature that was introduced with the 1.14 release is the --recurse option. This allows one to recursively traverse a directory tree. The code for this feature was created by Adam Retter, Jaishree Davey and Laura Damian of The National Archives (UK).

Memory mapping

The 1.15 release introduced the use of memory mapping for reading input images. This results in better performance when processing (very) large files. Images that would cause a memory error in previous versions are now handled without any problem. Also, the processing of very large files can be significantly faster than in earlier releases, and is less prone to freezing other processes that are simultaneously running on the machine. This improvement was suggested by Stefan Weil of Mannheim University Library, and the changes are based on a patch he submitted.

Two examples illustrate the benefits of this change:

  • This 2 GB image
    resulted in a memory error with jpylyzer 1.14.2 on a Windows machine with 4 GB RAM. The latest versions process the file without problems.
  • On a Linux Mint machine with 8 GB RAM, this 6.7 GB image
    also resulted in a memory error. Again, the current version handles the file without any problem.

This doesn’t mean that memory errors are now a thing of the past entirely; they may still occur under some circumstances. For instance, a test with the 6.7 GB image failed on a Linux Mint machine with 4 GB RAM. So it seems prudent to make sure that the amount of available RAM always exceeds the maximum image size by a fairly wide safety margin. Also, chip architecture and operating system may put further constraints on the amount of memory than can be mapped at a time.

Improved exception handling

Prior to release 1.16.0, an exception during the processing of an image could cause jpylyzer to crash. For example, an extremely large image can result in an internal memory error, and this would grind jpylyzer to a halt. This is particularly problematic when using the new --recurse option: in this case a single jpylyzer invocation may involve the processing of thousands of images at a time. One single (e.g. extremely large) image could then result in unusable output; moreover, it would be difficult to identify which image caused the crash in the first place! Release 1.16.0 introduces improved exception handling that allows jpylyzer to handle such situations more gracefully.

Robustness

The combined effect of the exception handling, memory mapping and status output should make jpylyzer releases from 1.16.0 onwards significantly more robust than previous versions. As an example, here’s some (simplified) output for a 6.5 GB JP2 that caused a memory error:

<?xml version='1.0' encoding='UTF-8'?>
<jpylyzer>
    <toolInfo>
        <toolName>jpylyzer.py</toolName>
        <toolVersion>1.16.0</toolVersion>
    </toolInfo>
    <fileInfo>
        <fileName>AS16-P-4102.jp2</fileName>
        <filePath>/home/johan/testJpylyzer/AS16-P-4102.jp2</filePath>
        <fileSizeInBytes>6745365021</fileSizeInBytes>
        <fileLastModified>Wed Dec  2 20:05:29 2015</fileLastModified>
    </fileInfo>
    <statusInfo>
        <success>False</success>
        <failureMessage>memory error (file size too large)</failureMessage>
    </statusInfo>
    <isValidJP2>False</isValidJP2>
    <tests/>
    <properties/>
</jpylyzer>

Previous versions would simply crash in this situation. Now, automated workflows can simply check for the value of the success field to verify the status of the validation. More importantly, if the jpylyzer invocation involved multiple input files (e.g. through the --recurse option), errors like these will not stop the processing of the remaining files.

64-bit Windows binaries

Finally, from version 1.15.1 onwards we are now providing 64 bit Windows binaries of jpylyzer (previously only 32-bit binaries were available).

Links

Jpylyzer website

]]>
http://blog.kbresearch.nl/2015/12/08/jpylyzer-2015-round-up/feed/ 0