New E-Books from Cambridge Core

A further 253 e-books across most subject areas have been added to DiscoverEdCambridge Core hosts books from a variety of publishers including Edinburgh University Press, Cambridge University Press, Boydell & Brewer.  A list of the new titles and subject areas can be found in the spreadsheet here.

Posted in New e-resources, Updates | Tagged , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , | Comments Off on New E-Books from Cambridge Core

International Journal of Sport Management and Marketing – new e-journal

We have a new e-journal subscription – International Journal of Sport Management and Marketing 

International Journal of Sport Management and Marketing (IJSMM) is a multidisciplinary journal which aims to provide a unique focus on a wide range of sport management and sport technology topics. It covers advances in theory, new concepts, methods and applications and case studies. Each issue disseminates quality sport-related research relevant to sport technology and sport management, examining both hard and soft perspectives in managing sporting organisations in the public and private sectors.

Posted in New e-resources | Tagged , , , , | Comments Off on International Journal of Sport Management and Marketing – new e-journal

Automated item data extraction from old documents

Overview

The Problem

We have a collection of historic papers from the Scottish Court of Session. These are collected into cases and bound together in large volumes, with no catalogue or item data other than a shelfmark. If you wish to find a particular case within the collection, you are restricted to a manual, physical search of likely volumes (if you’re lucky you might get an index at the start!).

Volumes of Session Papers in the Signet Library, Edinburgh

Volumes of Session Papers in the Signet Library, Edinburgh

The Aim

I am hoping to use computer vision techniques, OCR, and intelligent text analysis to automatically extract and parse case-level data in order to create an indexed, searchable digital resource for these items. The Digital Imaging Unit have digitised a small selection of the papers, which we will use as a pilot to assess the viability of the above aim.

Stage One – Image preparation

Using Python and OpenCV to extract text blocks

I am indebted to Dan Vanderkam‘s work in this area, especially his blog post ‘Finding blocks of text in an image using Python, OpenCV and numpy’ upon which this work is largely based.

The items in the Scottish Session Papers collection differ from the images that Dan was processing, being images of older works, which were printed with a letterpress rather than being typewritten.

The Session Papers images are lacking a delineating border, backing paper, and other features that were used to ease the image processing. In addition, the amount, density and layout of text items is incredibly varied across the corpus, further complicating the task.

The initial task is to find a crop of the image to pass to the OCR engine. We want to give it as much text as possible in as few pixels as possible!

Due to the nature of the images, there is often a small amount of text from the opposite page visible (John’s blog explains why) and so to save some hassle later, we’re going to start by cropping 50px from each horizontal side of the image, hopefully eliminating these bits of page overspill.

A cropped version of the page

A cropped version of the page

Now that we have the base image to work on, I’ve started with the simple steps of converting it to grayscale, and then applying an inverted binary threshold, turning everything above ~75% gray to white, and everything else to black. The inversion is to ease visual understanding of the process. You can view full size versions by clicking each image.

A grayscale version of the page

Grayscale

75% Threshold

75% Threshold

The ideal outcome is that we eliminate smudges and speckles, leaving only the clear printed letters. This entailed some experimenting with the threshold level, as you can see in the image above, a lot of speckling remains. Dropping the threshold to only leave pixels above ~60% gray was a large improvement, and to ~45% even more so:

60% Threshold

60% Threshold

45% Threshold

45% Threshold

At a threshold of 45%, some of the letters are also beginning to fade, but this should not be an issue, as we have successfully eliminated almost all the noise, which was the aim here.

We’re still left with a large block at the top, which was the black backing behind the edge of the original image. To eliminate this, I experimented with several approaches:

  • Also crop 50px from the top and bottom of the images – unfortunately this had too much “collateral damage” as a large amount of the images have text within this region.
  • Dynamic cropping based on removing any segments touching the top and bottom of the image – this was a more effective approach but the logic for determining the crop became a bit convoluted.
  • Using Dan’s technique of  applying Canny edge detection and then use a rank filter to remove ~1px edges – this was the most successful approach, although it still had some issues when the text had a non-standard layout.

I settled on the Canny/Rank filter approach to produce these results:

Result of Canny edge finder

Result of Canny edge finder

With rank filter

With rank filter

Next up, we want to find a set of masks that covers the remaining white pixels on the page. This is achieved by repeatedly dilating the image, until only a few connected components remain:

You can see here that the “faded” letters from the thresholding above have enough presence to be captured by the dilation process. These white blocks now give us a pretty good record of where the text is on the page, so we now move onto cropping the image.

Dan’s blog has a good explanation of solving the Subset Sum problem for a dilated image, so I will apply his technique (start with the largest white block, and add more if they improve the amount of white pixels at a favourable increase in total area size, with some tweaking to the exact ratio):

With final bounding

With final bounding

So finally, we apply this crop to the original image:

Final cropped version

Final cropped version

As you can see, we’ve now managed to accurately crop out the text from the image, helping to significantly reduce the work of the OCR engine.

My final modified version of Dan’s code can be found here: https://github.com/mbennett-uoe/sp-experiments/blob/master/sp_crop.py

In my next blog post, I’ll start to look at some OCR approaches and also go through some of the outliers and problem images and how I will look to tackle this.

Comments and questions are more than welcome 🙂

Mike Bennett – Digital Scholarship Developer

 

Posted in Uncategorized | Tagged , , , , , , | Comments Off on Automated item data extraction from old documents

Web of Science Upgrade to version 5.25 – Maintenance Sunday 25 June 2017

Web of Science is undergoing scheduled maintenance on Sunday June 25, from 13.00 (BST) to Monday June 26, at 01.00 (BST).

During this time, access to the service may be intermittent and the service should be considered at risk during the maintenance period.

Clarivate Analytics apologise for any inconvenience as a result.

During the maintenance period there will also be an upgrade of Web of Science to version 5.25.

For information on the new features of the WoS upgrade (including screenshots) please see the WoS Release notes at:

http://wok.mimas.ac.uk/support/documentation/WoS525-external-release-notes.pdf

Features include:

* Modernised discovery workfows.

* Enriched analytics workfows.

* Expanded and updated data.

* Improved Product Quality.

Posted in Access issues | Tagged , , | Comments Off on Web of Science Upgrade to version 5.25 – Maintenance Sunday 25 June 2017

Conservation Student- Claire Hutchison

In this week’s blog, Claire, our first placement student of the summer describes her experience of working with us…

My name is Claire and I am a conservation student from Northumbria University, specialising in paper conservation. This is the first of several voluntary work placements I must carry out as part of my Master’s degree. Working at the conservation studio the past two weeks has challenged me to work with new materials outside of my specialism.

Claire working in the studio

I have been working on artworks that will be displayed as part of the University’s new exhibition, ‘Highlands to Hindustan’ which goes on display at the end of July. My role was to conserve the works where needed, but to also improve the storage and display of the objects. This meant a great deal of multi-tasking and time management. Read More

Posted in Student Placement | Tagged , , , | Comments Off on Conservation Student- Claire Hutchison

Unexpected delights at the UCF

There is a large area of grass growing at the side of the University Collections Facility (née Library Annexe) which is usually just grass and moss with the odd daisy daring to poke its head above the grassy parapets. However, a lunchtime stroll resulted in a delightful find of a Common Spotted Orchid (Dactylorhiza fuchsii) growing amongst the buttercups (Ranunculus acris) and grasses at South Gyle.

Brown honey bees looked conspicuous busily collecting nectar from the white clover (Trifolium repens) which is in full flower. The tiny purple flowers of selfheal (Prunella vulgaris) looked especially lovely as a complementary colour to the shining yellow buttercups. A daisy (Bellis perennis) or two is also flowering. A few thistles (Cirsium vulgare) are nearly flowering and small sheep sorrel plants (Rumex acetosella) are appearing. The type of grassland here (neutral to alkaline) is typical of one which has not had chemicals or artificial fertilisers put on it and is typical of how grassland looked before ploughing and fertilising became common practices.

Sandi Phillips, Collections Management Assistant

Posted in Uncategorized | Tagged , , , , , , | Comments Off on Unexpected delights at the UCF

New Law E-Books Available

We have purchased the 2017 Hart Collection from Bloomsbury.  We have already added 61 titles to DiscoverEd and the remaining titles will be added as they are published.  A full list of the 125 titles are listed here and the status column K will show the available titles already added to DiscoverEd.

Posted in New e-resources | Tagged , , , | Comments Off on New Law E-Books Available

Digital library of pre-modern Japanese works for open access

The National Institute of Japanese Literature (Kokubunken) has made their new新日本古典籍総合データベース/Database of Pre-Modern Japanese Works (tentative edition) freely available at: https://kotenseki.nijl.ac.jp/ (Japanese interface) and  https://kotenseki.nijl.ac.jp/?ln=en (English interface).

The database, built out of Kokusho Sōmokuroku, a Japanese reference book published by Iwanami Shoten, and the largest of its kind, contains about 300,000 of 500,000 entries listed in the original book, along with digitized versions of materials referred to by the entries. It also allows one action search across repositories of multiple institutions.

For more information, please see the Project to Build an International Collaborative Research Network for Pre-modern Japanese Texts (NIJL-NW project) websites at:

http://www.nijl.ac.jp/pages/cijproject/

http://kotenseki.nijl.ac.jp/page/about.html#DB

Posted in Japanese Studies | Tagged | Comments Off on Digital library of pre-modern Japanese works for open access

Edinburgh’s International Inventions Exhibitions

Did you know we had International Inventions Exhibitions in Edinburgh in 1886 and 1890? The most famous international exhibition was held in London in 1851 at the Crystal Palace, but there were many others in the major European Capitals. The first one in Edinburgh was held on the Meadows, about where Jawbone Walk is now. The second was out west at Shandon, between the colonies and the canal. Music was important in both, with brass bands, organ recitals and piping contests. In 1890, you could stand in a phone booth in Edinburgh and listed to orchestras playing in Galashiels. Phones were exciting new technologies then, and their modern equivalents still are!

 

Posted in Uncategorized | Comments Off on Edinburgh’s International Inventions Exhibitions

The Perils of Technology: Aviation During the First World War

The First World War was the first in which air warfare played a significant part. While aircraft were ultimately to change the face of warfare, the demands of the war provided a rapid boost to this very new technology.

At the outbreak of the war effective powered flight was a technology not much more than ten years old, but each of the participating countries already had an armed air service of some sort. The allies had 208 aeroplanes between them, and Germany 180.  There were also airships, which initially seemed better-tried and more practical, although their importance diminished as the war progressed.  The British aircraft were split between the Royal Flying Corps (RFC), under the command of the army, and the Royal Naval Air Service (RNAS), under the command of the Navy.  These were merged to form the Royal Air Force (RAF) in 1918.

The types of aircraft in use were very varied, because each country had several of their own manufacturers and models of aircraft, and almost any machine which was available might be pressed into service. Aircraft models evolved rapidly, as technology was improved, and the needs of the war changed.

Initially, aeroplanes were regarded as primarily useful for reconaissance, and stable two-seaters, such as the British B.E.2, were preferred.  These were not very manoeuvreable and were poor at defending themselves or evading enemy anti-aircraft guns, which led to the development of fast, single-seater fighters, such as the French S.P.A.D.

The Germans made parallel developments; their early reconnaisance aircraft including monoplanes with distinctive swept-back, birdlike wings, such as the Rumpler Taube.  By the end of the war bigger, heavier aircraft designed for bombing had been developed.

From the beginning of the war both allied and German air forces had to establish, for the first time, how to make their aircraft recognisable, both to other airmen and to those on the ground. National markings were rapidly adopted, and by the beginning of 1915 both French and British authorities had produced posters showing silhouettes of enemy aircraft, entitled ‘Fire on these’.

In late 1914 or early 1915 the French produced the very first book of aircraft silhouettes for recognition purposes, Silhouettes D’Avions, Diagrams of Aeroplanes. These were produced with text in French and English and distributed to both troops and airmen.  An alternative version was produced on cards, for better durability.  Updated editions and supplements were issued to reflect new developments.  We hold two versions of this in our collections – the very first edition, showing French, British and German aircraft of late 1914, and a supplement of French aircraft from September 1915

.

Examination of the two shows just why these publications were of limited success in preventing both ground troops and airmen from attacking the wrong aircraft. The pictures are not very high quality, and do not show up the main features of the individual aircraft particularly clearly, especially to an untrained eye.   Take the picture of the B.E.2 – it is quite difficult to make out that it is trying to indicate that the upper wings are longer than the lower ones.

The illustrations for the two editions are printed from different artwork, which shows up ambiguities in the lines.  The Parasol Morane is included in both.  In one illustration there are lines which might be either substantial structural struts, or nearly invisible cable, but all of which are omitted from the other picture entirely.

The S.P.A.D. was a very common aeroplane type, but it takes some effort to work out from its picture that the propeller was located in the middle of its fuselage.

There is even one illustration we have not been able to identify, the Avion de Chasse Morane, which does not seem to entirely correspond to any aircraft made by the Morane company and which was used during the war, that we can find a modern record of!

These manuals were all produced in a hurry; they illustrate only a selection of models, and do not show variants or all the latest developments of equipment. Even more confusing, not to say downright unhelpful, is the earlier edition, which summarises all the other British models of aircraft as being similar to the ones illustrated.

Towards the end of the First World War the problems of distinguishing aircraft, while not solved, were somewhat reduced. In 1917 the RFC ordered 1000 Bristol Fighter aircraft, so that although there were still many different models of aeroplane in the skies, there began to be some standardisation.  More successfully, in early 1918 the French set up a ‘Flying circus’ which toured examples of the different models of their aircraft around the British airfields.  This had the happy result of not only familiarising the RFC with the appearance of the different models, but the opportunity to compare performance and develop some camerarderie with the French aviators.

Despite the shortcomings of these identification manuals, this approach continues to be used today, better pictures and combining it with other methods of teaching improving its effectiveness, although today it is more likely to be used by enthusiasts for civilian aircraft, and combined with a mobile phone app for detecting and tracking aircraft.

 

Posted in Uncategorized | Tagged , , , , | Comments Off on The Perils of Technology: Aviation During the First World War

Follow @EdUniLibraries on Twitter

Collections

Default utility Image Hill and Adamson Collection: an insight into Edinburgh’s past My name is Phoebe Kirkland, I am an MSc East Asian Studies student, and for...
Default utility Image Cataloguing the private papers of Archibald Hunter Campbell: A Journey Through Correspondence My name is Pauline Vincent, I am a student in my last year of a...

Projects

Default utility Image Cataloguing the private papers of Archibald Hunter Campbell: A Journey Through Correspondence My name is Pauline Vincent, I am a student in my last year of a...
Default utility Image Archival Provenance Research Project: Lishan’s Experience Presentation My name is Lishan Zou, I am a fourth year History and Politics student....

Archives

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.