Keith Munro, new Research Data Support Assistant

Hello, my name is Keith Munro and on March 4th 2024 I began my new role as a Research Data Support Assistant. Immediately prior to joining the Research Data Service (RDS), I studied for a PhD in Computer and Information Science at the University of Strathclyde. My thesis studied the information behaviour of hikers on the West Highland Way, see below for a photo of me during data gathering, with a particular focus on embodied information that walkers encountered, the classification of information behaviour in situ and well-being benefits resulting from the activity. I was lucky to present at the Information Seeking In Context conference in Berlin in 2022 and I am still working on getting a number of the findings from my thesis published in the months ahead. I passed my viva on Feb 2nd, so the timing of starting this job has been excellent.

Before my PhD, I studied for a MSc in Information and Library Studies, also from the University of Strathclyde, so there was always a plan to work in the library and information sector, but as my Masters degree was finishing during the outbreak of the Covid-19 pandemic in Spring/Summer 2020, I decided to take an interesting diversion, the scenic route, if you will, with the PhD! My Masters thesis was on the information behaviour of DJ’s, motivated by my own, lucky to do it but not exactly high-profile, experience as a DJ. From this, I was very fortunate to win the International Association of Music Librarians (UK & Ireland branch) E.T. Bryant Memorial Prize, awarded for a significant contribution to the literature in the field of music information. Subsequently, findings from this have been published in the Journal of Documentation and Brio.

Since starting my new role I have been greatly impressed by the team I have joined, who all bring a wealth of experience from across the academic spectrum and have also been very warm in welcoming me and in sharing knowledge. I hope I can bring my study and research experience to complement what the RDS team is doing and I am excited to be learning more about research data management. The size of the University of Edinburgh can be daunting and learning all the acronyms will take some time I suspect, but the range of research I have already encountered in reviewing submissions to DataShare has been fascinating, including Martian rock impacts and horse knees, something I’m sure will continue to be the case!

Digital Curation Interviews project with DCC

In this guest blog post, Clara Lines Diaz reports on last year’s Digital Curation Interviews with University of Edinburgh researchers, conducted by the Digital Curation Centre (DCC) on behalf of Digital Research Services.

The project was initiated by staff in the Research Data Service to gain an overview of the research data and software management practices and challenges across the University through in-depth interviews with researchers. The DCC was selected as best placed to conduct the interviews, given its expertise on the subject matter and location at University of Edinburgh. The information was collected through semi-structured interviews during Spring 2023.

2 women talking at table


Image by WOCinTech Chat, Flickr

The motivation to collect this information was to help ensure that researchers are supported in their specific needs and to contribute to shaping the research data management (RDM) services. The choice of in-depth interviews as a method was also expected to help build deeper relationships between service providers and users.

For the semi-structured interviews we had some topic blocks as below, and some prepared questions within each of those blocks. This was used more as a check list for us and we gave the interviewees space to focus on or bring up anything they considered relevant.

Topic blocks:

  • A: Introduction: Research line, projects, collaborations
  • B: Data provenance, types and reuse
  • C: Data/Software management practices
  • D: Influences on data/software management practices
  • E: Data management challenges and sources of assistance

This type of interview works well for exploratory studies like this because it allows common and maybe unexpected patterns to emerge, but also has some caveats around comparability, as not all interviews cover exactly the same topics in the same level of detail. This means that in the results we were able to indicate, for example, how many people mentioned using a particular service, but we could not infer that the others don’t use it, just because they did not mention it.

To select the participants, we contacted research support staff in the three colleges and asked them to suggest participants or send the invitation around. It felt like there was a high interest to discuss these topics and make the challenges they encounter heard, especially among researchers in the College of Science and Engineering (CSE). The interviewees were all involved in data intensive research, with a mix of senior and early career researchers. The interviews were planned to last around 45 minutes but there was some variation in the duration.

From the 14 interviews, four were with staff from the College of Medicine and Veterinary Medicine (CMVM), eight with staff from CSE and two with staff from the College of Arts, Humanities and Social Sciences (CAHSS). The oversampling of CSE interviews was intended as the service team was particularly interested to hear about their practices, which are less well known to them.

Once we had all the interview notes, we extracted the comments and classified them by themes. This was the basis for the final project report, which included a selection of the themes and possible points for action for the Research Data Service and the Research Computing Service in five key areas:

  • Data sharing and reuse was common practice, but there were challenges and areas where further support would be beneficial.
  • Code sharing and collaborative development was widespread and growing, but support and services were perceived as being less mature than that provided for data.
  • External collaboration with university-hosted services could be challenging.
  • Awareness of FAIR and open science was variable.
  • There was an appetite for more training, both for students and staff.

Sharing and reuse had a special focus in the interviews and the first two points are connected to that. Most interviewees had a lot to say about challenges related to sharing and reusing data, especially those working with sensitive data. Some extra advice to help people with those challenges would help. Most interviewees also discussed storing and, in some cases, sharing their code. GitHub is in general preferred for that. Sharing code is in general considered very time consuming.

A briefing was given to the Digital Research Services in August, 2023 and the Research Data Support team was given the transcripts and full results to inform service development.

Clara Lines Diaz
Research Data Specialist
Digital Curation Centre

Large scale data transport service launched

Research Services, IT Infrastructure Division, are pleased to report that a project that allows researchers to transfer terabytes of data between the University of Edinburgh and external collaborators has been completed. The service uses a transport mechanism known as Globus to set up multiple connections between host and client to transfer data instead of relying on a single point-to-point connection. This results in very large data being transferred between sites in parallel, allowing faster transfer.

The service is integrated with the University’s research data platform, DataStore, allowing researchers to specify specific folders that can be used as “endpoints” to the transfer. Many users have already taken advantage of the service, but it is key to note that this will not improve data transfer speeds within the University itself, rather that bottlenecks in the wider Internet can be mitigated.

For more information, University of Edinburgh users may view the RSS Wiki.

Mike Wallis
IS ITI

New feature: Sharing DataVault data with an external user

By popular demand, the Research Data Service is pleased to announce the arrival of a brand new feature: the DataVault Outward Staging Area (DOSA), a free-of-charge benefit to DataVault depositors.

Engraving depicting a stagecoach with people in front of a building

What is a staging area? Somewhere your data can be held temporarily, on the way to somewhere else. Just like a traditional staging post for stagecoaches, as shown in this engraving.

Imagine: your multi-terabyte dataset is safe-and-sound in your vault, you’ve cited it in a paper you’ve just published, and an external researcher has asked you for a copy. What will you do?

Simple: send a request to IS Helpline (or data-support@ed.ac.uk) asking us to create a DOSA folder for your data.

We’ll then use DOSA to give temporary (two months) external access to a copy of your deposit, using a Globus FTP endpoint. We’ll retrieve a copy of your data to the folder. And we’ll provide you with the Globus endpoint, which you send to the researcher. They may need to install some software to get the data. Alternatively, for datasets under 500 GB, we suggest a DataSync link will be more suitable. We set that up and provide it to you in the same way as the Globus endpoint. The difference for the end user is they can use the DataSync link (+ password) from their browser. Let us know if you have a preference for a Globus endpoint or a DataSync link (otherwise we’ll decide automatically based on the size).

 

Workflow diagram showing data moving from DataVault into DOSA, and from DOSA out to DataSync or globus

Workflow: We retrieve your deposit to your DOSA folder. We provide you with either a Globus endpoint or a DataSync link, to provide to your external person who made the request.

The DOSA is part of our networked active data storage, DataStore, but separate from the other staging area we provide for users making a new deposit (‘the DataVault staging area’), for the inward route.

Since 2016 researchers have been archiving data in Edinburgh DataVault. The DOSA is available for any DataVault deposit, old or new.

DataVault Outward Staging Area (DOSA): Sharing data with an external user

Not sure you’ll remember the name of the service? Worry not! I have a mnemonic device for you: just remember that a ‘dosa’ is an Indian savoury pancake. What’s not to like?

Photo of a folded dosa pancake on a tray with dishes of savoury sauces.

Pauline Ward
Data Repository Operations Officer
University of Edinburgh