Digital Humanities – KB Research http://blog.kbresearch.nl Research at the National Library of the Netherlands Fri, 24 Aug 2018 13:17:55 +0000 en-US hourly 1 https://wordpress.org/?v=4.4.2 Bridging the gap between quantitative and qualitative research in digital newspaper archives http://blog.kbresearch.nl/2017/01/25/bridging-the-gap/ http://blog.kbresearch.nl/2017/01/25/bridging-the-gap/#comments Wed, 25 Jan 2017 08:26:10 +0000 http://blog.kbresearch.nl/?p=2061 This blog post is written by Thomas Smits, KB Researcher-in-residence from May 2017

One of the central and most far-reaching promises of the so-called Digital Humanities has been the possibility to analyse large datasets of cultural production, such as books, periodicals, and newspapers, in a quantitative way. Since the early 2000s, humanities 3.0, as Rens Bod has called it, was posited as being able to discover new patterns, mostly over long periods of time, that were overlooked by traditional qualitative approaches.[1] In the last couple of weeks a study by a team of academics led by Professor Nello Christianini of the University of Bristol made headlines: “This AI found trends hidden in British history for more than 150 years” (Wired) and “What did Big Data find when it analysed 150 years of British history? (Phys.org). Did Big Data and Humanities 3.0 finally deliver on its promise? And could the KB’s collection of digitised newspapers be used for similar research?

The study, “Content analysis of 150 years of British periodicals”, is based on a corpus of 28.6 billion words, contained in 35.9 million articles of 120 regional, or local British newspapers from the period 1800-1950. [2] Focussing on six spheres – values and beliefs, UK politics, technology, economy, social change, and popular culture – the study is mostly based on the ‘use frequency’ of n-grams: the number of times a (combination of) word(s) appears in relation to all the words of the corpus in a specific year. For example, if a corpus consists of 100 words and the 2-gram, a combination of two words, “Digital Humanities” appears three times, the use frequency of this 2-gram is 0,03. In short: by applying n-grams the researchers were able to measure the relative importance of certain words, or combinations of words.

By using this method, the researchers were able to pinpoint specific historic events in their corpus, such as coronations, the election of a new pope, and outbreaks of several contagious diseases. More importantly, they used their method to test the validity of certain long-held notions about the nineteenth century. For example, the study suggests a very clear timeline in the emergence of the concept of “Britishness” in the popular imagination. While recent studies have posited that ‘national identity’ has deep historical roots, predating the nineteenth century, Christianini and his colleagues found that “British” overtook “English” only in the late-nineteenth century, supporting the close connection between the production of national identity, modernity, and the rise of mass media.[3]

Bias

While the researchers carefully composed their corpus, further contextualisation could make future research less biased. First of all, while the 120 newspaper titles studied in this research represent roughly 14% of all published titles, it remains unclear to what part of the press landscape these titles belonged. For example, the study neglects to discuss the fact that newspapers predominantly aimed to reach middle class readers. This bias is further enhanced by two factors: publications directed at lower classes, such as those of the chartist movement or the so-called penny press, were often deemed to be unworthy of archiving.[4] This process is enhanced by digitisation: well-known nineteenth century titles are more likely to be digitised than lesser-known, but not necessarily less influential, cheaper and/or radical publications.[5] Furthermore, by focussing on regional newspapers, the study aspires to mitigate the London-centric bias of research based on newspaper coverage. However, it neglects to account for the fact that in the first half of the nineteenth century many regional newspapers copied articles from London-based newspapers on a large scale, while syndication of content achieved, to some extent, the same result in the final decades of the century.[6]

Most crucially, the researchers seem to equate attention given in newspapers to historical significance. By doing so, they run the risk of failing to acknowledge the most important aspect of the medial form of the newspaper: its focus on ‘newsworthy’ events. The importance of certain long-term developments, which were never perceived as being radically new by contemporary commentators, can only be recognised with the benefit of hindsight. This leads to a somewhat paradoxical situation: digital newspaper archives are used to discover long-term trends, while newspaper discourse is mostly centred on short-term developments.

Delpher

What can researchers using the Delpher corpus learn from this study? The research department of the KB has already made important steps in the large-scale analysis of digitsed newspapers. Its open access n-gram viewer, developed by the University of Amsterdam, enables any user to replicate important parts of the British research. For example, a search for ‘nieuwe paus [new pope]’, yields the same results as the British study. In addition, recent projects of the KB’s fellows and researchers-in-residence use Dutch digitised newspapers in innovative ways. I hope that my own project, which applies two computer vision techniques to images in Dutch newspapers, can continue this tradition.

One could even ask the question why British digitised newspapers are not used more frequently for research similar to that of Christianini. Two private companies, Gale and FindMyPast, provide access to the collection of digitised historical newspapers, originally archived by the British Library. In a recent article, I compare these companies to Trove, the digital collection of Australian newspapers maintained by the government.[7] While public digital collections, such as Trove and Delpher, encourage users to tweak the archive, providing them with access to ‘raw’ data and API’s, private companies, such as FindMyPast, focus on a specific kind of use, in this case amateur historians studying their family history. This results in the fact that access to raw data is expensive, which makes it relatively hard, especially for junior researchers, to use it. In my opinion, we should take a critical look at the network of actors involved in the digitisation of newspapers, books, and other sources. Private companies, such as Google and FindMyPast, increasingly shape our access to the past and our use of historical sources. As I argue in my article, we should continue to discuss how this influences the ways that both researchers and the general public are able to interpret the past and relate to it in ways that are meaningful to them.

Bridging the gap

(Media) historians using qualitative methods would be wise to take note of the results of this study. It opens up a new world of possible research and shows how quantitative analysis can be used to substantiate existing theories. More importantly, the article raises the question if the strict separation between qualitative and quantitative research, or distant and close reading, is useful in distinguishing ‘traditional’ methods from their ‘digital’ counterparts. As the study amply shows, insights from traditional research are essential in defining the questions and contextualising both the corpus and the results of this kind of data-driven research. The project points to the importance of interdisciplinary research teams and, hopefully, will further undermine the trenches into which practitioners of the ‘digital’ and the ‘traditional’ humanities have grouped themselves.

[1] R. Bod, “Who is afraid of patterns? The Particular versus the Universal and the Meaning of Humanities 3.0” BMGN 128, no. 4 (2013): 171-80,  175.

[2] Landall-Welfare et al, “Content analysis of 150 years of British periodicals” PNAS (published ahead of print January 9, 2017). doi:10.1073/pnas.1606380114.

[3] This body of scholarship is mostly connected to Benedict Anderson’s concept of the imagined community. B. Anderson, Imagined Communities. London: Verso, 1983.

[4] For an explanation of what French press historian Jean-Pierre Bacot has called the ‘downward spiral of popularity’ of nineteenth-century newspapers and periodicals see Andrew King’s work on the London Journal: J.P. Bacot, La presse illustrée au XIXe siècle: une histoire oublié. Limoges: PULIM, 2005, 75 : A. King, The London Journal 1845-83: Periodicals, Production, and Gender. Aldershot: Ashgate, 2004, 16.

[5] A. Hobbs, “The Deleterious Dominance of The Times in Nineteenth-Century Scholarship” Journal of Victorian Culture 18, no. 4 (2013): 472-497.

[6] M. Beals, “Musings on a Multimodal Analysis of Scissors-and-Paste Journalism (Part 1),” accessed November 22, 2016, http://mhbeals.com/musings-on-a-multimodal-analysis-of-scissors-and-paste-journalism; B. Nicholson, “‘You Kick the Bucket; We Do the Rest!’: Jokes and the Culture of Reprinting in the Transatlantic Press” Media History 17, no. 3 (2012): 277-278.

[7] T. Smits, “Making the News National: Using Digitized Newspapers to Study the Distribution of the Queen’s Speech by W. H. Smith & Son, 1846–1858” Victorian Periodicals Review 49, no. 4 (2016): 598-625. DOI: 10.1353/vpr.2016.0041

]]>
http://blog.kbresearch.nl/2017/01/25/bridging-the-gap/feed/ 1
Farewell; Work on Discerning Journalistic Styles continues! http://blog.kbresearch.nl/2017/01/17/farewell-work-on-discerning-journalistic-styles-continues/ http://blog.kbresearch.nl/2017/01/17/farewell-work-on-discerning-journalistic-styles-continues/#respond Tue, 17 Jan 2017 11:42:40 +0000 http://blog.kbresearch.nl/?p=2055 At the end of December our current researcher-in-residence dr. Frank Harbers of Groningen University ended his project ‘Discerning Journalistic Styles’. In this blogpost he describes the outcomes and plans for the future.

It is January 2017, meaning my period as researcher-in-residence at the KB has come to an end. It also means that my project Discerning Journalistic Styles (DJS) has come to an end. It was a really nice and valuable experience and a fruitful project in which we (I couldn’t have done it without the expertise of KB programmer Juliette Lonij) have managed to create a classification tool that automatically determines the genre of news articles. You can try the tool yourself at: http://www.kbresearch.nl/genre. Just paste a Dutch news article in the text box, press the button below and the result will appear on the right side; simple as that!

genre-classifier

Currently the tool predicts the correct genre in 65% of the cases. This might not seem that high at face value. However, we need to take into account 1) that genres are ideal types that never manifest themselves in their pure form and boundaries between different genres are fluid; 2) that genres are dynamic concepts that change over time, and 3) that genre is a typical example of a ‘latent content’ category, meaning that determining a genre involves a considerable amount of interpretation. It is therefore unsurprising that classifying genres manually is also difficult and that human coders also regularly disagree on what the correct genre of a text is. In fact, it is not unusual that 20 to 30% of the time, human coders disagree on what the right genre of a (historical) news article is. With that in mind, 65% is a solid result – which is not to say that it doesn’t need to be improved.

In that sense the research has only just begun. Not only because in the coming period we will keep concern ourselves with presenting the results on conferences and in academic articles, but also because we are developing research projects that follow-up on DJS. I therefore hope this won’t be the last time I visit The Hague to delve into the historical newspaper collection. If you are curious about the tool, please keep a close eye on the new Lab website of the KB that will be launched soon. On that website we aim to give much more detailed information on the tool and its background.

]]>
http://blog.kbresearch.nl/2017/01/17/farewell-work-on-discerning-journalistic-styles-continues/feed/ 0
DH Clinics – librarians unite! http://blog.kbresearch.nl/2016/11/11/dh-clinics-librarians-unite/ http://blog.kbresearch.nl/2016/11/11/dh-clinics-librarians-unite/#comments Fri, 11 Nov 2016 09:52:27 +0000 http://blog.kbresearch.nl/?p=2005 You might have heard someone from @KBNLResearch mention DH Clinics, or a colleague at the libraries of the Vrije Universiteit or Universiteit Leiden, but what are they, why do we need them and who are they for?

The DH Clinics are our attempt of spreading the DH-word amongst our Dutch colleagues. We wanted to set up a community of librarians who were involved in DH, in order to learn from each other and discuss new methods and initiatives. However, we soon learned that a lot of academic libraries in the Netherlands were still thinking about DH and how to implement it in their organisations. We’re speaking early 2015 now and luckily, a lot has happened since, but we believe a small impulse is needed to speed everything along.

And that is why we are now organising DH Clinics for Dutch academic librarians (and possibily also some archivists). The idea of the clinics is that we tackle several major themes of DH over six full-day sessions. The mornings are dedicated to lectures about the theme (think of, for example, text and data mining) and are open to a public that is interested in DH. In the afternoon, we’ll have hands-on workshops with specific applications. This part of the session is meant for people who are actually working with DH researchers, or want to do so.

We’re not setting out to re-train librarians into programmers or data crunchers, but we do want to provide them with the basics of DH with which they should increase their knowledge level to such an extent that they are able to follow the (online) discussions in the field, give tips to beginning researchers, perhaps even use some of the tools for their own work and ideally to engage with the very rich online content to learn more, such as The Programming Historian or Library Carpentry (both of which are used as inspiration for our clinics).

We’re developing the clinics with the Working Out Loud-principles, which means we will regularly share what we are doing and invite you to comment on it. Since we’re working together with the VU and UBL, this blog is not the only one you should keep an eye on, but we’ll announce everything via Twitter as well, so if you’re not already following us, now is the perfect time to start! @KBNLresearch

]]>
http://blog.kbresearch.nl/2016/11/11/dh-clinics-librarians-unite/feed/ 2
Tackling problems and making progress http://blog.kbresearch.nl/2016/11/03/tackling-problems-and-making-progress/ http://blog.kbresearch.nl/2016/11/03/tackling-problems-and-making-progress/#respond Thu, 03 Nov 2016 15:57:11 +0000 http://blog.kbresearch.nl/?p=1994 Our current Researcher-in-Residence, Frank Harbers, is well under way with his project “Discerning Journalistic Styles. Exploring Automated Analysis of Journalism’s Modes of Expression”. In this blogpost he gives an update on his project and its progress.

Frank Harbers

It has been several months since I wrote the first blog about my work as researcher-in-residence and the research project is in full swing by now. The first phase of the project , connecting the metadata from my own database to the historical newspaper data (and metadata) in Delpher is finished and we are fully enveloped in the main part of the project: training a classifier to automatically determine the genre of historical newspaper articles.

The first phase was not as successful as we hoped, but we have managed to create a – modest – dataset to train the classifier. Initially, we hoped to be able to connect the metadata about approximately 33.000 Dutch newspaper articles to the data in Delpher. A crucial factor in the success of this attempt was the extent to which the segmentation of newspaper articles in Delpher matched the way the newspaper articles were segmented for the content analysis that resulted in the set of metadata about the historical newspaper articles. Unfortunately, it was far from a perfect match. For that reason the newspapers before 1945 could not be included – basically half of the metadata. Furthermore, De Volkskrant after the Second World War has not been digitized by the KB. In addition, the segmentation of De Telegraaf in the postwar period was so different that we couldn’t include that either. In the end, this meant that we could only use the data of Algemeen Handelsblad/NRC Handelsblad in the postwar period. So quickly we saw our dataset shrink from the potential 33.000 articles to a modest 2000 articles. A bit of a setback, but fortunately we can still use this smaller dataset to train a genre classifier. This experience does make clear how crucial segmentation is for the creation of datasets that can be fruitfully used for digital humanities research into journalism history.

At the moment, we are working on the second phase of the project. We have identified several genres that we would like to classify. These genres, such as news reports, reportages, interviews, opinion articles, reviews, news analyses, can shed light on the way journalism developed from a reflective, opinion-oriented way of doing journalism to a more event-centered and fact-oriented journalism practice. At the core of this part is the translation of the genre definitions to clear linguistic markers that can be identified automatically. Take for instance the news report, a genre that is defined by the use of the inverted pyramid (a story structure in which typical journalistic questions, like Who, What, Where and When, are answered in the first paragraph. Moreover, it often contains direct quotes from sources and is generally a fairly concise article written in a depersonalized, objective style. Question is how you can recognize these features automatically in the text. In this case, the quotes can be recognized by the presence of quotation marks (for which a high quality OCR is crucial) and we will attempt to identify the inverted pyramid structure by using named identity recognition to see whether questions concerning who was involved and where and when it happened are answered. We hope the depersonalized style can be captured by looking at the lack of a first person perspective (the use of the pronoun ‘I’ or ‘We’) and the lack of adjectives that create a colorful and subjective account.

Juliette Lonij, programmer on this project, is currently developing the Python software to extract the features on which the classifier will run. She looked into different natural language processing software packages to pre-process the article texts and chose to use FROG for tokenization and Part-of-Speech tagging, which facilitates our research needs quite well (other packages might be added in the future). And today, we have just run a first exploratory test with the classifier, which showed promising results. In the coming weeks we will keep on testing and refining the classifier. So wish us luck!

 

]]>
http://blog.kbresearch.nl/2016/11/03/tackling-problems-and-making-progress/feed/ 0
Abstracts applications for Researcher-in-residence 2017 http://blog.kbresearch.nl/2016/10/25/abstracts-applications-for-researcher-in-residence-2017/ http://blog.kbresearch.nl/2016/10/25/abstracts-applications-for-researcher-in-residence-2017/#respond Tue, 25 Oct 2016 08:52:57 +0000 http://blog.kbresearch.nl/?p=1932 Below you will find the abstracts that were submitted and unfortunately not accepted for the 2017 run of the Researcher-in-residence programme. The abstracts are in alphabetical order. If your abstract is published here and you would like to have your name posted with it, please contact us and let us know. The accepted projects and their abstracts can be found here.

We want to thank all researchers for their interesting proposals, wish them all the best for 2017 and hope to see them again in a following year!

Deep learning OCR post-correction – dr. Janneke van der Zwaan

Humanities research makes extensive use of digital archives. Most of these archives, including the KB newspaper data, consist of digitized text. One of the major challenges of using these collections for research is the fact that Optical Character Recognition (OCR) on scanned historical documents is far from perfect. Although it is hard to quantify the impact of OCR mistakes on humanities research (Traub et al., 2015), it is known that these mistakes have a negative impact on basic text processing techniques such as sentence boundary detection, tokenization, and part-of-speech tagging (Lopresti, 2009). As these basic techniques are often used prior to performing more advanced techniques and most advanced techniques use words as features, it is likely that OCR mistakes have a negative impact on more advanced text mining tasks humanities researchers are interested in, such as named entity recognition, topic modeling, and sentiment analysis.

The goal of the proposed research is to bring the digitized text closer to the original newspaper articles by applying post-correction. Post-correction involves improving digitized text quality by manipulating the textual output of the OCR process directly. The idea is that better quality data boosts eHumantities research. Although the quality of the KB newspaper data would definitely benefit from improving the OCR process itself (i.e., improved image recognition), post-correction will still be necessary, because the quality of historical newspapers is suboptimal for OCR (e.g., due to poor paper and print quality) (Arlitsch & Herbert, 2004).

Existing approaches for OCR post-correction generally make use of extensive dictionaries to replace words in the OCRed text that do not occur in the dictionary with words that do (see e.g., Alex et al., 2012, Strange et al., 2014, Volk et al. 2011). Based on the assumption that a number of characters in every word will be identified correctly, words not in the dictionary are replaced with alternatives that are as similar as possible to the text recognized, possibly taking into account word frequencies to solve ties. The main problem with these existing approaches is that they do not take into account the context in which words occur.

Deep learning techniques provide an opportunity to take this context into account. I propose to learn a character based language model of Dutch newspaper articles. This is a model of the character sequences occurring in the text of a corpus (see Karpathy (2015) for examples). OCR mistakes can be viewed as deviations from this model. Mistakes can be fixed by intervening when text deviates too much from the model.

Back to top ↑


Faith in Old Age. A biographical and micro-historical study into the relation between religious beliefs and social-cultural perceptions, experiences and practices of ageing (c. 1800-1950)

In 2041 4,7 million inhabitants of the Netherlands (26,5%) will be 65 years or older. One third will be above eighty and likely in need of care. Parallel to this development we are moving away from the welfare state towards a participation society in which ‘a strengths-based approach that encourages citizens to be more in control of their own lives, of their own communities, and eventually of society as a whole’ is needed. The outcomes of this biographical and microhistorical study will contribute on a fundamental level to that need.

In this study I will research how and to what extent small scale life histories reflect the relation between personal religious beliefs and social-cultural perceptions, experiences and practices of ageing and caring for the aged and if and how these religious convictions reflected and shaped the urban social-cultural ageing practice in the past (c. 1800-1950).

Although the past does not provide answers for the future, this study will open up longer perspectives on ideas and practices of ageing and ageing care and thus facilitate the construction of broader imaginative and critical perceptions of ourselves in our society and the way we (can) act today. By working together with other academic disciplines and professionals in social development this project will fuel a necessary scientific, political and public dialogue on ageing and what will motivate people in the near future to ‘reconquer the initiative’ to care for the aged from the welfare state.

In the long run it will also contribute to the understanding of what religious societies are and how they function, which has relevance for other academic disciplines as well as for society as a whole.

Back to top ↑


Mapping the Early Modern Dutch News(papers) (1618-1795)

The first Dutch newspaper was published in 1618. It marked the beginning of the rise of newspapers in the Dutch Republic. It was the recent launch of the newspaper database Delpher in 2013 which caused an increase in the research on early modern newspapers (Van Groesen, 2013; Van Groesen, 2015; Der Weduwen, 2015). Despite the digital disclosure of these historical newspapers, most research is done manually and on a small scale. It still remains difficult to determine the larger picture of the news provision in the early modern period. The big questions about the precise content, extent and origin of this news, are yet to be answered.

The aim of this project is to use Delpher as an instrument to visualize the origin and spread (both in location and tempo) of the early modern news in the Dutch Republic. By enriching metadata it becomes possible to show where the news came from, and how long the news was on the way from a certain region. By using the origin and the date of the news, it is possible to make a digital map based on early modern newspaper data.

Early modern newspapers consisted mainly of a single sheet of paper filled with foreign news.

The newspaper was clearly divided into blocks per country (with the name of a country as a caption above). Utilizing this recognizable and clear layout as a filter, it is possible to show where the news originated and the percentage of space of the delimited text. This project focuses on developing a deep learning program based on computer vision techniques which can automatically determine and extract the news items from digitized newspapers.

The blocks with news (sorted by country) contain separate items which are clearly recognizable with indented paragraphs. This standard format is important for software enhanced filtering. Each item begins with the city (or region) of origin, and date. This project will develop a method which adds metadata (country, city and date) to each news item. While great progress is made on fully searchable newspapers via OCR, this project adds a valuable dimension which allows a view from above.

Ultimately questions will be answered such as: in what period did Swedish news get more attention (in frequency and coverage)? And: from which countries originated news during the Nine Years’ War? Furthermore, the rise and scale of transnational news circulation can be mapped. With this, for the first time, it can be clearly established where news came from. The dating of news is also very important. Combined with dates of publication of the newspaper, dating each news item results in answering the important question of how long news was underway. On the basis of the origin and date, a digital map with a timeline will be created. Another version of this map, a more experimental one will be created which uses time instead of distance as a measure between locations.

Back to top ↑


Segmentation and Categorization of Advertisements in Delpher’s Newspapers: An Eighteenth Century Feasibility Case Study

This project aims to arrive at a better level of segmentation and categorization for individual advertisements in Delpher’s newspapers, with a particular focus on the eighteenth century collection. In this way, this invaluable historical resource of advertisements can be studied in much greater detail than is currently the case. This is extremely important for historical researchers, because the entire market economy passes by in these advertisements. They contain essential information about the flow of many goods and services in the Dutch Republic and subsequently the Kingdom of the Netherlands, which cannot yet be accessed in detail on such a big scale.

Currently, search queries for any product or service will give results that are not obviously relevant, because the nature of an advertisement is not disclosed in the search results. The smallest entities in Delpher are sections of advertisements, with metadata that have little to say about individual advertisements. Sections of advertisements can hold between 1 and 25+ advertisements, and the diversity of material within these sections 25+ is massive. Therefore, users still have to check the relevance of their results manually, because a query can occur in any kind of advertisement.

With a better level of segmentation and categorization, individual advertisements are recognizable as separate entities. In this project, both shape and content of the advertisements are used to arrive at a better level of segmentation. By making better use of markers for relevance on a page, such as indentations or capitalized words, it is possible to split segments of advertisements into smaller entities, that have meaning on their own. Once technical feasibility has been fully established, then advertisements can be studied, clustered, enriched and re-used in superior ways. For the purpose of categorization, a library of predefined categories of advertisements will be created, which allows users to narrow down to a baseline of relevant advertisements much quicker.

As a test case for this approach, a specific yet recognizable category of advertisements will be used: eighteenth century advertisements for auctions of drug components. These advertisements contain essential information about the early modern drug trade, but their contents overlap with other categories of advertisements: it is hard to find solid search queries to isolate this category of advertisements from others. Labelling these advertisements on the basis of a predefined category makes it possible to analyze them in greater detail, and to arrive at a valid thesis about the drug trade. This is of crucial importance to understand the mechanisms of the premodern medical

marketplace: many aspects that receive substantial attention from scholars (clinical testing, prescriptive procedures, preparation of remedies and so on) require understanding of the import and availability of raw materials.

Thus, this project will clarify the commercial dimension of early modern medicines, as a test case for developments of the market economy as a whole.

Back to top ↑


Sound Patterns of Golden Age Theatrical Emotions in KB/DBNL’s digised theatre plays

Sound Patterns of Golden Age Theatrical Emotions: the development of a tool to reveal the correlation between phonological patterns and emotions in KB/DBNL’s digitised Theatre Plays, to unravel the aural elements of historical texts.

What did emotions in the Dutch theatre sound like in the Golden Age? Did comedies sound different to tragedies? Did angry men on stage sound different to angry women? How did queens in love sound in relation to servants in love?

My PhD project researches the role that phonological patterns play in the expression of emotions in early modern Dutch theatre plays, and the way analysis of phonological patterns can contribute to new methods of author identification.

With a quantitative approach to modelling sounds and emotions, the project includes 200 digitised Dutch theatre plays provided by the Digital Library for Dutch Literature (DBNL), covering the entire early modern period in the Netherlands from 1570 to 1800. The project will result in a historical sound pattern timeline, which fits in with the results of the Historic Embodied Emotion Model, HEEM (Leemans e.a. 2015, Leemans e.a. forthcoming 2016), product of the emotion mining project, conducted by the Amsterdam Centre for Cross-disciplinary Emotion and Sensory Studies (ACCESS). The selection of plays involved in my research corresponds to the corpus of the ACCESS research group. With HEEM, ACCESS has created a new technique of sentiment mining. STAGE, (“Sounds in Theatre plays featuring Golden Age Emotions”) adds a new element by associating emotions with phonological patterns, and bringing together the fields of the history of emotions and those of (historical) phonology, musicology, history of theatre and computational linguistics.

Counting and analysis of phonological patterns will help reveal how the history of emotional expressions on stage has evolved. In addition the data this tool generates could open other opportunities for research on sound patterns in texts. Furthermore, as the tool has modular construction, this enables its application to related projects in the field of computational linguistics, (historical) phonology and ‘distant reading’ in a wide range of texts.

The supervisors are Prof. Inger Leemans (Vrije Universiteit Amsterdam) and Prof. Karina van Dalen-Oskam (Universiteit van Amsterdam, Huygens-ING).

Keywords: History of Emotions, History of (Dutch) Theatre, Historical Phonological Patterns, Machine Learning, Open Access Tool, Digital Humanities Research Question: How did the expression of emotions on stage evolve in the Netherlands during the Golden Age? How do phonological patterns relate to historical emotional expressions on stage? My PhD project aims to reveal the role phonological patterns play in the expression of emotions in early modern Dutch theatre.

Back to top ↑


Understanding petitioning behavior in the Batavian Republic (1795-1801) through enhanced access to serial government sources

For scholars working on the last decades of the eighteenth century, the Early Dutch Books Online (EDBO) dataset is an invaluable corpus of source material.

To literary works, political writings and other conventional texts in this dataset, adequate access is provided through Delpher and Nederlab. There is, however, at least one important source type for which the present search options of these tools are not optimal. As EBDO contains virtually the entire printed output of the revolutionary Batavian Republic, it also includes the many multi-volume proceedings of local, provincial, and national representative bodies that were printed  in order to ensure a maximally transparent government. For historians, the Dagverhaal der handelingen van de Nationaale Vergadering and other such serial government publications are immensely rich sources that are also notoriously tough to work with. It is my conviction that the accessibility of this source type could be greatly improved by an approach more comparable – but not identical – to that already applied to other datasets, such as KB Kranten en Staten-Generaal Digitaal.

As a researcher-in-residence I therefore intend to build on my experience in working with this source type to create, in close collaboration with the KB digital humanities team, customized search options in Delpher. I propose a multifaceted approach with a primary focus on the use of automatic segmentation to separate the daily or weekly instalments in which these serial sources were published, the sessions of the representative bodies, and the deliberative elements of which each session was made up. If users can be enabled to search only the deliberative elements that are relevant to them in the proceedings of multiple governmental bodies at the same time and if their search results can be sorted by date or session, this opens up a whole new realm of research opportunities.

As for my own research, I want to deploy the enhanced searchability that should result from this project to address a set of research questions concerning the petitioning behavior of Dutch citizens during the Batavian Republic. Between 1795 and 1801 citizens petitioned all levels of government, as they had done in the old regime Dutch Republic but on a much larger scale and often with a more overtly political agenda. The description and discussion of petitions in the proceedings of various government bodies, the inventorying of which will be made manageable by this project, provides insight in how citizens related to local and supra-local contexts and how they came to terms with the great ideological and institutional transformations of their day. These questions are at the heart of my current research project The primacy of local belonging. Private papers, petitioning, and periodical press, 1747-1848.

In the long run, I consider my contribution to meeting the objectives set out in this proposal an investment in new research and teaching initiatives.

Moreover, the benefits of this project could become greater in the future as the knowledge and skills gained from it might eventually also be applied to other serial sources in the EDBO dataset.

Back to top ↑


Unlocking the STCN

The aim of the project is to use the Short-Title Catalogue, Netherlands (STCN) as an instrument for an easy-to-use tool to visualize trends in the history of the book in the sixteenth and seventeenth century based on a SPARQL generator.

The STCN is the national bibliography of the Dutch printed book up to the year 1800. It is a catalogue with over 204.000 titles. But the STCN is more than a catalogue. It is an overview of the printed book in the sixteenth and seventeenth century in the Netherlands. An overview in which all sorts of data lies hidden which can give us insight in how the book changed in these centuries. The STCN is available online, but would benefit from additional ways to consult its underlying data to gain insight in larger trends. For example, it can be used to show changes in the book in a specific genre, a specific decade, observing typographical developments, or the most active printer or author based on location or year. It is possible to use the STCN as a research tool via the Advanced Search function, but this often comes down to tallying.

In the proposed project, the STCN will be used to create an accessible (RDF) dataset and a tool (SPARQL generator) for researchers, students and other interested parties. This tool will be an easy accessible platform in which the user can request information about trends in the printed book. Recently it has become possible to answer questions based on the STCN with the help of the RDF query language SPARQL. However, this is a difficult language to master for occasional users. Last year, the Koninklijke Bibliotheek (KB) offered a STCN SPARQL-workshop for researchers. Despite the high turnout of interested researchers, working with SPARQL proved to be too difficult for most participants. The proposed tool will ease dealing with SPARQL language.

With the help of a SPARQL generator, the tool provides the users the option to combine selected variables to answer their questions in just a few clicks.

The output can be used to display information visually, in charts or plots.

This digital humanities project will increase the potential use of the STCN.

Trends and changes in the book which now has to be dug out the STCN, will soon be just a few clicks away which allows for a better understanding of the emergence of the printed book. The project will be more than a plaything for data mining the STCN hosted by the KB. This experimental SPARQL generator can be used for other (KB) projects with RDF data or the Semantic Web.

]]>
http://blog.kbresearch.nl/2016/10/25/abstracts-applications-for-researcher-in-residence-2017/feed/ 0
Our Researchers-in-residence 2017 will be…. http://blog.kbresearch.nl/2016/10/03/our-researchers-in-residence-2017-will-be/ http://blog.kbresearch.nl/2016/10/03/our-researchers-in-residence-2017-will-be/#respond Mon, 03 Oct 2016 14:31:49 +0000 http://blog.kbresearch.nl/?p=1901 Earlier this year we sent out our Call for Proposals for our Researcher-in-Residence Program 2017. This program offers a chance to early career researchers to work in the library with the Digital Humanities team and KB data. In return, we learn how researchers use the data of the KB. Together we will address their research question in a 6 month project using the digital collections of the KB and computational techniques. The output of the project will be incorporated in the KB Research Lab and is ideally beneficial for a larger (scholarly) community.

This year, we received nine proposals that focused on a wide range of datasets and techniques. Last week, a group of seven leading Dutch Digital Humanities professors met at the KB to discuss each proposal thoroughly. Today we are excited to announce the names and projects of our two Researchers-in-Residence 2017!

20160923_122908

20160923_122943 

The first researcher is Melvin Wevers of Utrecht University who will be focusing on advertisements in our Digital Newspapers, please find his abstract below.

melvin-wevers

Combining Textual Content and Non-Textual Features of Digitized Newspaper Advertisements to Study Historical Developments in the Dutch Consumer Society

The KB’s digitized newspaper collection provides an important and exciting set of advertisements. After all, newspapers played a major role in the dissemination of advertisements. This project aims to analyze how advertisements in digitized newspapers can be used to study historical changes in the Dutch consumer society. Roland Marchand argues that advertisements provide an insight into the ideals and aspirations of past realities. Advertisements show the state of technology, the social functions of products, and provide information on the society in which a product was sold (Marchand, 1985). Academic work on advertisements often focuses on a specific symbolic connotation, such as gender or consumerism. This project aims to build on this scholarship.

In this project, I will develop computational methods to identify trends and breakpoints in newspaper advertisements that represent the Dutch consumer society. Schreurs contends that advertisements changed markedly in the Netherlands during the twentieth century (Schreurs, 2001). He claims that advertisements became more visual and gained prominence in media. I intend to use the KB’s researcher-in-residence fellowship to test this hypothesis in a quantifiable manner. In collaboration with the KB, I develop computational methods to analyze three aspects of the advertisements in newspapers between 1850 and 1950. First, the position, size, and frequency of newspaper ads. These metadata are indicators of the prominence and cultural impact of advertisements. For instance, a large advertisement on the front page in a national newspaper has more impact than a small ad on page 8 in a local newspaper. After aggregating these metadata, we can analyze their temporal dynamics. Do changes over time in these aspects reveal characteristics of the Dutch advertising landscape? For instance, did advertisements increase in size and/or move to the front pages?

Secondly, I propose to develop methods to cluster ads by brand or product group using text mining techniques on the textual content of advertisements—available as OCR-ed text. This clustering can help to understand whether the trends found in the metadata are product-specific. The information derived from specific product groups can be used to test existing hypotheses posed in corporate histories or histories of the advertising industry. For instance, did the position and size of cigarette ads changed over time?

Thirdly, I focus on the visual aspect of advertisements. For this aspect, I would like to examine whether computer vision techniques can be applied to advertisements. The precision of computer vision techniques is far from perfect, and therefore this last step would be mostly exploratory (Snoek et al. 2015). A large part of the meaning in advertisements was expressed in images. The extraction of advertisements from the corpus allows for the analysis of visual information in advertisements. Can computer vision be used to identify logos, objects, and people in ads? The dataset’s richness and size offers the KB the possibility to set important future steps in the field of computer vision.

Our second researcher-in-residence is Thomas Smits of Radboud University. Thomas will also focus on newspapers but will address the illustrations.  Please find his abstract below.

thomas-smits

Illustrations to Photographs: using computer vision to analyse news pictures in Dutch newspapers, 1860-1940

Most digital humanities projects are based on the analysis of text. However, in our increasingly visually orientated world, it has become clear that we should also devise ways to analyse visual material. In the last couple of years, the KB has made important steps in this emergent field: Delpher provides users with the opportunity to search for ‘images with caption’ in its database of digitized newspapers and the KBK-1M database, which holds all the images published in the KB’s digitized newspapers between 1923 and 1995, provides researchers with the opportunity to analyse the visual material of this collection in a viable way.

The proposed research will apply two computer vision techniques to sort the images of the KBK-1M database according to the way in which they were reproduced (engraving/half-tone) and shed a new light on an important transitional phase in the history of the visual culture of the news. Several media historians suggest that around 1900 both illustrations and photographs were considered to be objective visual representations of the news. However, relying on case studies, they have been unable to pinpoint this period. By introducing a digital humanities approach to this question, the proposed project will describe and analyse this period. It consists of two phases, connected to two digital humanities components.

First of all, building on the PhoCon project of Elliott & Kleppe (2016), the KBK-1M database will be expanded to include the period 1860-1923. Using the power of the SURFSara’s Cartesius supercomputer, the first phase will apply the technique of a recent project of Fyfe & Ge (2016) to the images in the expanded database. Fyfe and Ge analysed images in three Victorian illustrated newspapers by measuring their pixel ratio and the entropy level. By juxtaposing these two so-called low-level features, images could be sorted according to the technique used for their reproduction (engraving/half-tone).

The second phase will explore how (a combination of) two open source applications (OpenCV/Caffe) can be used to fine-tune the recognition of engravings and photographs. Both programs can create so-called cascade classifiers that are able to detect faces, objects, or patterns on images in a large dataset by comparing them to a manually created training set. Based on the results of the application of Fyfe & Ge’s method, several cascade classifiers can be created are able to detect specific patterns of different reproduction techniques.

The project will provide researchers with a new way to sort, discover patterns, and make sense of the visual material contained in the KB’s collection of digitized newspapers. Second, by applying computer vision techniques to study an important development in the history of the visual culture of the news, it introduces a digital humanities approach to the relatively theoretical field of nineteenth-century visual culture studies.

Both projects will take place in 2017 and we will keep you updated on their progress on this blog. If you are curious about the work of our current researcher-in-residence Frank Harbers, please see his latest blog post.

If you have any questions about the projects or programme, feel free to contact us via dh@kb.nl. All other submitted abstracts will be posted in a separate blog later this week.

]]>
http://blog.kbresearch.nl/2016/10/03/our-researchers-in-residence-2017-will-be/feed/ 0
Lets get to work! http://blog.kbresearch.nl/2016/07/21/lets-get-to-work/ http://blog.kbresearch.nl/2016/07/21/lets-get-to-work/#respond Thu, 21 Jul 2016 07:04:29 +0000 http://blog.kbresearch.nl/?p=1855 Since 1 July, our new researcher-in-residence dr. Frank Harbers joined our Research Department to work on his project ‘Discerning Journalistic Styles. Exploring Automated Analysis of Journalism’s Modes of Expression’. He will share his experiences through regular blogposts and we’re happy to share his first below. If you would like to be our researcher-in-residence in 2017, please see the Call for Proposals which is currently open.

Frank Harbers

Lets get to work!

Two weeks ago I took the train in Groningen at 7.16 AM and arrived around 10 AM in The Hague to start my fellowship as researcher-in-residence. My first day mainly consisted of tasks of practical and organizational nature (login data, an access pass, printer codes, etc.). Martijn Kleppe gave me a tour of the building with all its corners and corridors. I hadn’t seen more than the general and special collections reading room, where I spent quite some time perusing the original historical newspaper material during my PhD research into the development of the press from the 19th century onwards.

During my PhD research (‘Between Personal Experience and Detached Information. The Development of Reporting and the Reportage in Great Britain, the Netherlands and France, 1885-2005’) I studied the development of reporting by manually coding a sample of historical newspapers for characteristics such as topic, genre, sourcing practices, images; a long and arduous endeavor. To make this easier, in my current project at the National Library of the Netherlands (KB) ‘Discerning Journalistic Styles. Exploring Automated Analysis of Journalism’s Modes of Expression’ I will attempt to automate the classification of genre of newspaper articles. The large database of metadata of newspaper articles I compiled during my PhD provides the necessary already-coded test material. Fortunately, I don’t have to do this all by myself, but I am lucky that Juliette Lonij, who knows so much more about the technical side of this type of research, will collaborate with me on this project.

And now that all practicalities have been arranged, we could ‘really’ get started this week. The first problem we will try to solve is one of more practical nature: my database with metadata about the historical newspapers has to be linked to the actual digitized newspaper material that is found in Delpher. It is the first necessary step to be able to compile the datasets we will use to explore and experiment with the different approaches and tools to automate the classification of genre.

What makes this project so challenging is the fact that genres, such as reportages, news reports, background analyses or interviews, cannot be recognized by the topical content (such as sports for example), but only by its stylistic and formal characteristics. A reportage, for instance, is typified by the many depictions of the atmosphere, but such depictions can relate to the aggressive atmosphere in a football stadium, an impression of the natural beauty of the Amazon, or to the tension that can be felt during a police arrest. You have to focus on different features, such as the use of adjectives for example.

Manually classifying the genre of articles is time consuming, which means that you can only code a limited amount of newspaper material. This limitation makes generalizing statements about the development of journalism and reporting within a particular cultural context, like the Netherlands, problematic. It would therefore be an important step forward if the classification of genre can be automated. That way much larger amounts of material can be examined, making the historical analyses more robust. So, lets get to work!

]]>
http://blog.kbresearch.nl/2016/07/21/lets-get-to-work/feed/ 0
KB Digital Humanities Team at DH2016 http://blog.kbresearch.nl/2016/07/08/kb-digital-humanities-team-at-dh2016/ http://blog.kbresearch.nl/2016/07/08/kb-digital-humanities-team-at-dh2016/#respond Fri, 08 Jul 2016 06:51:31 +0000 http://blog.kbresearch.nl/?p=1839 This year we will again be at the Digital Humanities Conference. After visiting the conference in Nebraska, Lausanne and Sydney we very much look forward to meeting international Digital Humanities scholars this year in Cracow, Poland. Three members of our Digital Humanities team will be attending the conference: Juliette Lonij, Steven Claeyssens and Martijn Kleppe.

Juliette Lonij will present a paper together with our former researcher-in-residence Pim Huijnen on his KB project ‘From keyword search to discourse mining – the meaning of scientific management in Dutch vocabulary, 1900-1940’ during the session ‘Extracting textual content 6’, Friday 15 July 2.30-4pm in room MSWB. During the poster session on Wednesday afternoon, Martijn Kleppe will present our new KBK-1M dataset at booth 066: ‘1 Million Dutch Newspaper Images available for researchers: the KBK-1M dataset’. We published both short abstracts below. Martijn is also one of the co-organizers of the AVinDH workshop and will chair a session on ‘Images and Art’ Friday 15 July 11.30 am-1pm in room MADB.

We look forward to the conference and are also eager to get in touch with researchers who are interested in the call for our (fully paid!) Researcher-in-Residence program 2017 which is currently open. If you would like to hear more on our program and possibilities please do not hesitate to approach Steven, Martijn or Juliette if you see them at one of the sessions or breaks. If you want to be sure to meet them you can also send them an email at dh@kb.nl or send a tweet to our @KBNLResearch account.

Schermafbeelding 2016-07-08 om 08.17.12

From Keyword Search To Discourse Mining – The Meaning Of Scientific Management In Dutch Vocabulary, 1900-1940

Pim Huijnen (Utrecht University), Juliette Lonij (National Library of the Netherlands)

In this paper we present a technique to enable the historical study of ideas instead of words. It aimed at assisting humanities scholars in overcoming the limitations of traditional keyword searching by making use of context-specific dictionaries. The aim of the project in the context of which this technique was developed, was twofold: first, to create a method for dictionary extraction from a representative text corpus, based on existing methods and algorithms. Second, to find a way of executing dictionary searches in the KB’s digitized newspaper archive and visualizing the results. Both components of the project were tested and evaluated by means of a case study on the impact of American scientific management theories in the Dutch public sphere during the first half of the 20th Century. Using the approach described here, we were able to discover and analyze shifts in the way the modernization of Dutch business was discussed.

 

1 Million Dutch Newspaper Images available for researchers: The KBK-1M Dataset

Martijn Kleppe (National Library of the Netherlands), Desmond Elliott (University of Amsterdam)

The visualisation of news through photographs has exploded since the second half of the 20th century (Kester & Kleppe 2015). However, methods that are employed to analyse the (re)use of visual materials are labour-intensive because Humanities researchers tend to analyse their sources manually (Burke 2001). To estimate the increase in the use of pressphotographs in Dutch newspapers, Kester & Kleppe (2015) e.g manually analysed a sample of 385 newspapers and 5.877 press photographs over the period 1870-2013. To find the recurring use of photographs in Dutch history textbooks, Kleppe (2012) followed a same approach by manually analysing over 5.000 photographs in 400 history textbooks, creating the ‘Foto’s in Nederlandse Geschiedenisschoolboeken (FiNGS) (Photos in Dutch History textbooks) dataset (Kleppe 2013b). Even though manually created and annotated datasets such as FiNGS contain rich & well-annotated data, their scope remains limited given its labour-intensive creation and analyses process. Therefor this poster presents the KBK-1M dataset, that was created specifically for (Digital) Humanities researchers. This dataset contains a collection of 1.603.395 captioned images extracted from Dutch digitised newspapers stored in the Dutch National Library (KB) Newspaper archive of the period 1922-1994. On our poster, we will describe how we obtained the images, what types of research questions it could tailor and how researchers can obtain the dataset for their research purposes.

Please see the poster below.

KBK-1M Poster A1 DEF

]]>
http://blog.kbresearch.nl/2016/07/08/kb-digital-humanities-team-at-dh2016/feed/ 0
Call for proposals KB Researcher-in-residence 2017 http://blog.kbresearch.nl/2016/06/08/call-for-proposals-kb-researcher-in-residence-2017/ http://blog.kbresearch.nl/2016/06/08/call-for-proposals-kb-researcher-in-residence-2017/#comments Wed, 08 Jun 2016 09:49:00 +0000 http://blog.kbresearch.nl/?p=1798 The Koninklijke Bibliotheek (KB), National Library of the Netherlands is seeking proposals for its Researcher-in-residence program to start in 2017. This program offers a chance to early career researchers to work in the library with the Digital Humanities team and KB data. In return, we learn how researchers use the data of the KB. Together we will address your research question in a 6 month project using the digital collections of the KB and computational techniques. The output of the project will be incorporated in the KB Research Lab and is ideally beneficial for a larger (scholarly) community.

The KB and digitisation

The Koninklijke Bibliotheek (KB), National Library of the Netherlands  is a research library with a broad collection in the fields of Dutch history, culture and society, and as a national library collects and stores all (digital) publications that appear in the Netherlands, as well as a part of the international publications about the Netherlands. The KB has planned to have digitised and OCRed its entire collection of books, periodicals and newspapers from 1470 onward by the year 2030. Already in 2016, about 15% of this enormous task was completed, either from the KB itself or via public-private partnerships as Google Books and ProQuest. Over 20 million book-, newspaper- and magazine papers are currently available via the search portal www.delpher.nl. The project will be carried out in the Research Department of the KB and there will be two consecutive placements in 2017.

Who are we looking for?

Early career researchers who are:

  • PhD-students that are in their final stages of their PhD project or researchers that have obtained their PhD between 2011 and 2016
  • Employed at a university or research institute in the EU,
  • Interested in using one (or more) of the digital collections of the KB,
  • Available for 0.5 fte over a period of 6 months (Jan – Jun 2017 or Jul – Dec 2017) and able to spend at least 1 day a week at the KB.

What can we offer you?

  • A secondment with the KB for 0,5 fte for a period of 6 months based on your current salary
  • Access to all data sets of the KB,
  • An office space,
  • Travel costs within the Netherlands,
  • Support from a programmer, collection and data specialists.

Which collections do we have?

You can use any digital collection of the KB and even combine it with an external collection, if copyright allows. Several of our digitised collections are described in more detail on our website, such as the parliamentary papers and the medieval illuminated manuscripts.

You can also browse through our collection of more than 1 million newspapers, magazines, radio bulletins and books on Delpher.nl.

What kind of projects are we looking for?

We’re open to all kinds of projects that use our data and benefit your research and other users of the KB and/or the KB Research Lab. The KB Research Department currently focuses on research projects that improve, enrich, connect and analyse our data by using techniques and methods from the domains of Information Retrieval (IR), Natural Language Processing (NLP) and Machine Learning (ML). We encourage you to define your project by:

  1. formulating a fundamental research question that stems from your field of expertise and that can be linked to the applied techniques at the KB Research Department,
  2. formulating a project that is different from the previous executed Researcher in Residence projects that can be found on our blog.

For more inspiration also take a look at the previously submitted proposals on our blog: here, here, here and here.

How do I apply?

Fill out this form before 31 August 2016 to submit your project, after having read carefully our terms and conditions. The form contains the following elements: details, project description (including research question, theoretical background and applied methods and techniques), outcomes, work plan, personal background, your availability in 2017 and a checkbox on our terms and conditions.

Before you start working on your proposal, we encourage you take a look at the form so you will be able to fill it out in the most efficient manner.

Don’t forget to read the terms and conditions of this call and agree to them.

All proposals will first be reviewed by an internal KB committee and then forwarded to an external committee of representative experts from several Dutch universities and institutions that consists of:

  • prof. dr. Franciska de Jong, Erasmus University Rotterdam & Clarin
  • prof. dr. Sally Wyatt, eHumanities & Maastricht University
  • prof. dr. Karina van Dalen-Oskam, Huygens ING & University of Amsterdam
  • prof. dr. Joris van Eijnatten, Utrecht University
  • prof. dr. Maarten de Rijke, University of Amsterdam
  • prof. dr. Marcel Broersma, University of Groningen
  • Prof. dr. Emiel Krahmer, University of Tilburg
  • Prof. dr. Hilde de Weerdt, University of Leiden
  • Prof. dr. Arjen de Vries, Radboud University

All entries will be judged on:

  • Originality and quality
  • Link with techniques and methods currently applied at the KB Research Department (Information Retrieval, Natural Language Processing and Machine Learning)
  • Feasibility (technically, legally and practically)
  • How the KB data will be showcased and used
  • Whether the end results are of use for a wider community

You will be notified of the outcome of this call in October 2016.

For answer to more questions, read our FAQ. Please also read the terms of this call and placement.

Respondents are strongly advised to contact dh@kb.nl in advance of proposal submission to discuss eligibility, project details, prerequisites, and KB support with the Digital Humanities team, consisting of Lotte Wilms, Steven Claeyssens, Martijn Kleppe, Juliette Lonij and Willem Jan Faber.

]]>
http://blog.kbresearch.nl/2016/06/08/call-for-proposals-kb-researcher-in-residence-2017/feed/ 2
FAQ Call for Proposals Researcher-in-Residence http://blog.kbresearch.nl/2016/06/08/faq-call-for-proposals-researcher-in-residence-2017/ http://blog.kbresearch.nl/2016/06/08/faq-call-for-proposals-researcher-in-residence-2017/#respond Wed, 08 Jun 2016 09:48:50 +0000 http://blog.kbresearch.nl/?p=1813 Updated 04 June 2018

I don’t live or work in the Netherlands. Can I apply? 
Probably! Contact us at dh@kb.nl and we’ll discuss your options.

I want to use my own dataset. Is that possible?
Sure! As long as you also use one of the datasets of the KB and it doesn’t limit the publication of the project end results.

I don’t know how to code, is that a problem?
Not at all. We have skilled programmers who can help you with your project or we will try to find a match for you if you prefer someone else. This would mean submitting as a team and will cut the budget in half. Reach out to us to discuss the options.

I don’t speak Dutch. Is your content still interesting to me?
That depends on your research question :) It might not be so appealing to linguists, but could offer an novel collection for computer scientists. Contact us to see which collections we have and we can discuss what might be the most interesting set for you.

Why will you publish my abstract?
We want to show others what types of proposals we have received to offer future researchers an insight into the selection process and to prevent them from entering a similar project.

Can I submit a project I’ve submitted previously (at another institution)?
We’d like you to submit an original idea. It can be one you have had lying around for some time, but we’d appreciate projects that haven’t been done before. Projects that have been previously entered into a similar program should be changed significantly before resubmitting.

Can I also work fulltime on my project for a period of 3 months?
We prefer you to work part-time so you can spend a total of 6 months with us. This also allows you to continue your research or teaching obligations at your university.

Will you be able to reimburse any housing or hotel costs?
Unfortunately, when you come from outside the Netherlands, we are not able to find and fund your housing or pay for your travel expenses to the KB. However, we do fund travel costs within the Netherlands allowing you to come to and work in the KB, the Hague wherever you are based in the Netherlands.

I want to use my own programmer, can I?
Yes, you can. We even encourage you to bring in extra people when you want to address a subject we’re not experts in (such as multimedia). However, the budget remains the same, so it will have to be split between you. We do ask that the whole team is available in the KB for at least one day a week. If you want to know whether we can help you or if you should bring someone in, please contact us at dh@kb.nl.

I don’t know if my idea is what you’re looking for. What can I do?
You are welcome to contact us at dh@kb.nl to discuss your ideas and the possibilities.

Can I submit more than one project?
Please focus your efforts on one great project.

Who will be judging the entries?
The entries will be judged by an internal committee and then forwarded to an external committee of representative experts from several Dutch universities and institutions.

What will you judge my project on?
We will judge the entries on criteria such as feasibility (technically, legally and practically), how the KB data will be showcased and used and whether the end results are of use for a wider community. Next to this, we will also look at the originality and quality of the proposal and the amount of support needed (and in this case, more is not necessarily worse!).

What happens if you submit a plagiarized project?
When we notice your your project is plagiarized, we will not consider your application for placement. You are responsible for the originality and authenticity of the project, but we will keep our eyes open.

What happens to any software I write for my project?
All software in the projects, whether you or we write it, will be made available on the KB Lab and Github page under an open source license.

What happens to the data I collect/produce in my project?
At the KB Lab we try to be as open as possible. All data produced in the programme is to be made available for research purposes, either through the KB Lab, KB Data Services or DANS, and where possible will receive a CC-license.

Can I publish any papers about the project?
Yes, we even encourage you to do so. If necessary, we’re happy to help.

]]>
http://blog.kbresearch.nl/2016/06/08/faq-call-for-proposals-researcher-in-residence-2017/feed/ 0