Large scale data transport service launched

Research Services, IT Infrastructure Division, are pleased to report that a project that allows researchers to transfer terabytes of data between the University of Edinburgh and external collaborators has been completed. The service uses a transport mechanism known as Globus to set up multiple connections between host and client to transfer data instead of relying on a single point-to-point connection. This results in very large data being transferred between sites in parallel, allowing faster transfer.

The service is integrated with the University’s research data platform, DataStore, allowing researchers to specify specific folders that can be used as “endpoints” to the transfer. Many users have already taken advantage of the service, but it is key to note that this will not improve data transfer speeds within the University itself, rather that bottlenecks in the wider Internet can be mitigated.

For more information, University of Edinburgh users may view the RSS Wiki.

Mike Wallis
IS ITI

End of an era – 2017-2020 RDM Roadmap Review (part 1)

Looking back on three years that went into completing our RDM Roadmap in this period of global pandemic and working from home, feels a bit anti-climactic. Nevertheless, the previous three years have been an outstanding period of development for the University’s Research Data Service, and research culture has changed considerably toward openness, with a clearer focus on research integrity. Synergies between ourselves as service providers and researchers seeking RDM support have never been stronger, laying a foundation for potential partnerships in future.

thumbnail image of poster

FAIR Roadmap Review Poster

A complete review was written for the service steering group in October last year (available on the RDM wiki to University members). This was followed by a poster and lightning talk prepared for the FAIR Symposium in December where the aspects of the Roadmap that contributed to FAIR principles of research data (findable, accessible, interoperable, reusable) were highlighted.

The Roadmap addressed not only FAIR principles but other high level goals such as interoperability, data protection and information security (both related to GDPR), long-term digital preservation, and research integrity and responsibility. The review examined where we had achieved SMART-style objectives and where we fell short, pointing to gaps either in provision or take-up.

Highlights from the Roadmap Review

The 32 high level objectives, each of which could have more than one deliverable, were categorised into five categories. In terms of Unification of the Service there were a number of early wins, including a professionally produced short video introducing the service to new users; a well-designed brochure serving the same purpose; case study interviews with our researchers also in video format – a product of a local Innovation Grant project; and having our service components well represented in the holistic presentation of the Digital Research Services website.

Gaps include the continuing confusion about service components starting with the name ‘Data’___ [Store, Sync, Share, Vault]; the delay of an overarching service level definition covering all components; and the ten-year old Research Data Policy. (The policy is currently being refreshed for consultation – watch this space.)

A number of Data Management Planning goals were in the Roadmap, from increasing uptake, to building capacity for rapid support, to increasing the number of fully costed plans, and ensuring templates in DMPOnline were well tended. This was a mixed success category. Certainly the number of people seeking feedback on plans increased over time and we were able to satisfy all requests and update the University template in DMPOnline. The message on cost recovery in data management plans was amplified by others such as the Research Office and school-based IT support teams, however many research projects are still not passing on RDM costs to the funders as needed.

Not many schools or centres created DMP templates tailored to their own communities yet, with the Roslin Institute being an impressive exception; the large majority of schools still do not mandate a DMP with PhD research proposals, though GeoSciences and the Business School have taken this very seriously. The DMP training our team developed and gave as part of scheduled sessions (now virtually) were well taken up, more by research students than staff. We managed to get software code management into the overall message, as well as the need for data protection impact assessments (DPIAs) for research involving human subjects, though a hurdle is the perceived burden of having to conduct both a DPIA and a DMP for a single research project. A university-wide ethics working group has helped to make linkages to both through approval mechanisms, whilst streamlining approvals with a new tool.

In the category of Working with Active Data, both routine and extraordinary achievements were made, with fewer gaps on stated goals. Infrastructure refreshment has taken place on DataStore, for which cost recovery models have worked well. In some cases institutes have organised hardware purchases through the central service, providing economies of scale. DataSync (OwnCloud) was upgraded. Gitlab was introduced to eventually replace Subversion for code versioning and other aspects of code management. This fit well with Data and Software Carpentry training offered by colleagues within the University to modernise ways of doing coding and cleaning data.

A number of incremental steps toward uptake of electronic notebooks were taken, with RSpace completing its 2-year trial and enterprise subscriptions useful for research groups (not just Labs) being managed by Software Services. Another enterprise tool, protocols.io, was introduced and extended as a trial. EDINA’s Noteable service for Jupyter Notebooks is also showcased.

By far and away the most momentous achievement in this category was bringing into service the University Data Safe Haven to fulfil the innocuous sounding goal of “Provide secure setting for sensitive data and set up controls that meet ISO 27001 compliance and user needs.” An enormous effort from a very small team brought the trusted secure environment for research data to a soft launch at our annual Dealing with Data event in November 2018, with full ISO 27001 standard certification achieved by December 2019. The facility has been approved by a number of external data providers, including NHS bodies. Flexibility has been seen as a primary advantage, with individual builds for each research project, and the ability for projects to define their own ‘gatekeeping’ procedures, depending on their requirements. Achieving complete sustainability on income from research grants however has not proven possible, given the expense and levels of expertise required to run this type of facility. Whether the University is prepared to continue to invest in this facility will likely depend on other options opening up to local researchers such as the new DataLoch, which got its start from government funding in the Edinburgh and South East Scotland region ‘city deal’.

As for gaps in the Working with Data category, there were some expressions of dissatisfaction with pricing models for services offered under cost recovery although our own investigation found them to be competitively priced. We found that researchers working with external partners, especially in countries with different data protection legislation, continue to find it hard work to find easy ways to collaborate with data. Centralised support for databases was never agreed on by the colleges because some already have good local support. Encryption is something that could benefit from a University key management system but researchers are only offered advice and left to their own mechanisms not to lose the keys to their research treasures; the pilot project that colleagues ran in this area was unfortunately not taken forward.

In part 2 of this blog post we will look at the remaining Roadmap categories of Data Stewardship and Research Data Support.

Robin Rice
Data Librarian and Head of Research Data Support
Library and University Collections

Highlights from the RDM Programme Progress Report: May to July 2016

The following key results were highlighted in the RDM Programme Progress Report:

  • There were 42 new users and 69 data management plans created with DMPOnline.
  • An additional 1.5PB has been procured for DataStore’s general capacity expansions.
  • The Roslin Institute has deposited 16 datasets into Data Vault.
  • DataShare upload release (2.1) went live on 23 May 2016.
  • There are now 334 dataset records in PURE, an increase of 124 records from the last reporting period (February to April 2016).
  • 54 datasets have been deposited into DataShare.
  • The University of Edinburgh was recommended as a preferred supplier on the Framework for the Research Data Management Shared Services for Jisc Services Ltd (JSL) for the following Lots:
  • Lot 2: Repository Interfaces
  • Lot 3: Data Exchange Interface
  • Lot 6: Research Data Preservation Tools Development
  • Lot 8: User Experience Enhancements
  • A total of 390 staff and postgraduates attended RDM courses and workshops during this quarter.
  • A total of 3,649 learners enrolled for the 5-week RDMS MOOC rolling course from March through July, 2016 and a total of 461 people completed the course in the same time frame.
  • There were 5,198 MANTRA sessions recorded from May to July with 58 to 60 percent identified as new users.
  • Set up an RDM Forum in collaboration with College of Arts, Humanities and Social Sciences (CAHSS) Research Officer and Research Outputs Co-ordinator. The first RDM forum is scheduled for Wednesday, 7 September 2016.

Data Management Planning highlights

We currently hold sample data management plans for grant applications submitted to the Arts and Humanities Research Council (AHRC) the Economic and Social Research Council (ESRC) and the Medical Research Council (MRC).

 Active Data Infrastructure highlights

DataStore

An additional 1.5PB has been procured for general capacity expansions. This capacity will primarily be deployed to the College of Medicine & Veterinary Medicine (CMVM) and the College of Science & Engineering (CSE).

MRC Institute of Genetics & Molecular Medicine (IGMM) has purchased an additional 1.2PB of capacity, and this is now deployed in their dedicated file system.

Data Stewardship highlights

DataShare

The large data sharing investigation was completed for DataShare and reported previously. Upload release (2.1) went live on 23 May 2016. Download release planned following ‘embargo release’ and ShareGeo spatial data migration.

Data Vault

There was a soft release of Data Vault in February 2016, with the Roslin Institute depositing 16 datasets during this quarter.

PURE

There are now 334 dataset records in PURE, an increase of 124 records from the last reporting period (February to April 2016).

Research Data Discovery Service (RDDS)

Two PhD interns are working on School engagement activities (dataset records into PURE / datasets into DataShare) for Divinity & Division of Infection and Pathway Medicine; contracts end 16 September 2016. One PhD intern retrospectively added DataShare metadata to PURE for data deposits prior to PURE Data Catalogue functionality; contract to end 16 September 2016. A fourth PhD intern (to work with School of Informatics) is awaiting for approval.

Data Management Support highlights

A total of 390 staff and postgraduates attended RDM courses and workshops during this quarter.

Other related research data management support activities to highlight

  • Working with sensitive data in research’ guide was written for research staff and students in social sciences.
  • Another guide is being written on ‘Sharing and retaining data’ for research staff and students in social sciences.
  • Set up an RDM Forum in collaboration with College of Arts, Humanities and Social Sciences (CAHSS) Research Officer and Research Outputs Co-ordinator. The first RDM forum is scheduled for Wednesday, 7 September 2016.

Other activities to highlight

The outcome of Jisc RDM Shared Services bid that was submitted in March 2016

The Procurement Panel has recommended University of Edinburgh as a preferred supplier on the Framework for the Research Data Management Shared Services for Jisc Services Ltd (JSL) for the following Lots:

  • Lot 2: Repository Interfaces
  • Lot 3: Data Exchange Interface
  • Lot 6: Research Data Preservation Tools Development
  • Lot 8: User Experience Enhancements

Unfortunately, the Procurement Panel has decided not to recommend University of Edinburgh for the following Lots:

  • Lot 1: Research Data Repository
  • Lot 4: Research Information and Administration Systems Integrations

National and International Engagement Activities

From May to June

Çuna Ekmekcioglu gave a talk on ”Understanding and overcoming challenges to sharing personal and sensitive dataat the Recon: Research Communication & Data Visualisation Conference, 24th June 2016, The Edinburgh Centre for Carbon Innovation (ECCI).

Stuart MacDonald and Rocio von Jungenfeld ran three workshops for the IS Innovation Fund project, Data-X: Pioneering Research Data Exhibition, with PhD students from across the University. Introduction to Data-X: Pioneering Research Data Exhibition.

In June

Stuart MacDonald presented peer-reviewed presentation to IASSIST conference, Bergen: Supporting the development of a national Research Data Discovery Service – a Pilot Project.

Robin Rice presented a poster at Open Repositories 2016, Dublin: Data Curation Lifecycle Management at the University of Edinburgh.

Pauline Ward presented a lightning talk at Open Repositories 2016, Dublin:  Growing Open Data: Making the sharing of XXL-sized research data files online a reality, using Edinburgh DataShare.

Stuart MacDonald was an invited speaker at NFAIS (National Federation of Abstracting and Information Services) Fostering Open Science Virtual Seminar: NFAIS Fostering Open Science Virtual Seminar.

In July

Robin Rice gave two presentations (invited and peer-reviewed) at LIBER 2016, Helsinki: University of Edinburgh RDM Training: MANTRA & beyond; Designing and delivering an international MOOC on Research Data Management and Sharing.

Robin Rice filled in for Stuart Lewis as invited speaker for JISC-CNI 2016, London: Managing active research in the University of Edinburgh.

This is the last quarterly report as the Research Data Management (RDM) Roadmap Project (August 2012 to July 2016) came to a close on 31 July 2016.

There will be discussions with the RDM Steering Group to decide how future reporting will be conducted. These reports will be released on the Research Data Blog as well.

Tony Mathys
Research Data Management Service Co-ordinator