bookmark

Wednesday, 19 November 2008

Consorcio Madrono - research data seminar

Stuart Macdonald and Luis Martinez were invited to speak at a research data seminar (http://www.consorciomadrono.es/noticias_eventos/evento11.html organised by Consorcio Madrono - a consortium of 7 Madrid university libraries. The audience consisted of primarily of library managers and directors but also included researchers, IT and e-research/e-science specialists. With the aid of translation professionals Alicia Lopez Medina (UNED) gave an interesting overview of current initiatives in Spain including those relating to research data management, e-research and repositories. She indicated that a concerted effort is required in Spain to address issues surrounding the 'Data Deluge' and data management in academic settings as currently there are no platforms in place to cater for research data in an Institutional Repository environment. Luis introduced the concept of research data with typologies and examples, explained data management and data curation with their associated benefits and challenges, finishing with the findings of the study currently taking place at the University of Oxford - Scoping digital repository services for research data management. Stuart introduced the concept of data libraries (as they exist currently in UK tertiary education) and made mention of DISC-UK. He then discussed the DataShare project showcasing deliverables and anticipated outcomes, and the Data Audit framework. He finished by talking about Web 2.0 data visualisation tools that can be employed independently of or potentially integrated into a repository environment. Celia Russell (ESDS International) then detailed why research data is worth preserving and provided an overview of instistutional, national, european and global level data infrastrutures with particular reference to e-science and national/international grid environments and initiatives. The day ended with an extended and lively panel discussion which proved as fruitful to panelist as it did to the enthusiastic audience.
What became evident during the discussions was the generic use of the word 'data' to describe various different kinds of digital research output – highly curated collections, large scale dataset gathering, lab-based data generation, secondary social-science data (researcher-created data) and research data products and summary data. These all have different patterns of generation and use with variable lifetimes and life cycles. Distinguishing between finished research data packages and on-going data production and analysis also needs some thinking out. There are scale, storage and format issues. There is perhaps some scope to delineate between pre-publication and post-publication data - the latter being more likely to be repository (and data librarian) friendly, the former the domain of a new breed of 'data scientists' conversant with subject but with less interest in metadata, discovery and preservation. It may well be time to discuss each of the patterns of generation and use individually with a view to establishing where they differ and where there are commonalities, in addition to articulating curatorial roles, responsibilities and relationships for each of said patterns of generation and use!

Plenty of food for thought!

Presentations from the seminar will be posted here soon.

Stuart Macdonald
DISC-UK DataShare

Wednesday, 12 November 2008

DataShare deliverables over last 6 months

Having just submitted our October progress report, it seems we've accomplished quite a lot over the last 6 months. Not bad, considering it included summertime!

Summary:
The project has changed some of its deliverables following the change in LSE’s status to an associate partner due to staffing shortages; this includes a greater emphasis on data audits at each of the remaining partners to reach out to users earlier in the data lifecycle and to better meet their needs for support in data management. Edinburgh participated in the Data Audit Framework Development project led by HATII/DCC at University of Glasgow and conducted its own DAF Implementation project, both funded by JISC, and led by Robin Rice, with a new team member, Cuna Ekmekcioglu.

The project members have continued to engage with contacts at UK and international institutions – especially in the US and Australia - who are building services for data sharing. The project team has participated in professional development activities, and disseminated deliverables at conferences, in articles, and through their website and blog. A briefing paper on geo-spatial Web 2.0 visualisation tools was written. The findings from Oxford’s Scoping digital repository services for research data management project were disseminated. Several project team members participated in the Edinburgh Repository Fringe, along with peer projects amongst the partners, e.g. Kultur and EdShare (Southampton), ShareGeo and the Depot
(EDINA), and the CRIG International Roadshow (Oxford). An article in Online by Luis Martinez Uribe and Stuart Macdonald and an interview in CILIPS Update brought attention to the profession of data librarians, which was further amplified by the recent JISC-commissioned report by Key Perspectives, The Skills, Role and Career Structure of Data Scientists.

The partners have created and received peer review on a Dublin Core based metadata schema for datasets in DSpace and EPrints, worked on procedures for storing and preserving databases, and have developed a content model for a database of sound files in Fedora. The Edinburgh DataShare repository was soft-launched, with an option for depositors to append the open data license developed by the Open Data Commons.

The progress report also includes specific progress made at each partner institution and a new evaluation plan, to be carried out by Sheila Anderson at Kings College London.

Saturday, 25 October 2008

My Faves for Friday, October 24, 2008

Ben O'Steen's blog post describing the DISC-UK DataShare project's approach at Oxford to incorporating a dataset type (a phonetics database) into a Fedora repository.

[tags: Oxford, blogs, data curation, formats, metadata]

See the rest of my Faves at Faves

Wednesday, 15 October 2008

Update from LSE - Associate Partner in DISC-UK DataShare

LSE Library has been developing a new approach to the management and development of Data Library Services, following difficulties in recruiting to a single Data Library Manager post. We have identified the various strands of the service offered by the Library, and we are developing a team approach which will enable the flexible delivery and development of LSE Library services, along with the ability to actively participate in research and development activities in the field.

The Library has recently made the first appointment in this new team - with Dave Puplett taking up the post of Data Librarian on 6 October. Dave has worked in the Library since 2007, on a variety of projects focussing on improving access to electronic resources and exploiting new
web technologies. He has recently completed work on the JISC funded VIF project dealing with version identification of academic research in digital repositories.

Dave ('d dot puplett at lse dot ac dot uk') will be the main contact for the Data Library service.

Over the coming months we will continue to develop the Library's data services team and our commitment to the Data Library service remains as strong as ever.

Nicola Wright
Information Services Manager
London School of Economics and Political Science

My Faves for Tuesday, October 14, 2008

"Publishing data long-term with an accompanying journal publication can produce a large number of citations which will be good for your career in the Research Excellence Framework" -- said Michael Wilson, a researcher associated with the MRC Psycholinguistic Database, who presented this finding at the e-Science All-Hands Meeting in Edinburgh last month.

Edinburgh DataShare is an institutional data repository set up by the Data Library to help researchers do just this. It uses the same DSpace software as the Edinburgh Research Archive, ERA, and operates in tandem with the Publications Repository to provide a place for researchers to share the datasets on which their published papers are based.

[University Newsletter article announcing the soft launch of Edinburgh DataShare, a project deliverable and pilot service.]

[tags: edinburgh, repository, service, data sharing, data publishing]

See the rest of my Faves at Faves

Tuesday, 30 September 2008

My Faves for Monday, September 29, 2008

Gail Steinhart, co-chair of the working group, forwarded me a link to this paper during the summer, and I’m very pleased to have read it. The group, formed in 2006, has been investigating issues, current activities, and opportunities for the Library to get involved in “digital research data curation.” Thus, it serves as a very useful US equivalent to our DISC-UK State of the Art Review, but also hones in on the specific issues within a given institution, which is what I’d like to help the Information Services do within the University of Edinburgh.

The white paper begins with an environmental scan beyond Cornell, before turning to the strengths and potential areas of collaboration within the University. It looks at the actual and potential role of the academic research library, international organisations such as CODATA, activities in the UK including the importance of Liz Lyon’s 2007 report on roles and responsibilities, the EU DRIVER project, The Australian National Data Service and the activities at Monash University (“noteworthy in terms of utilizing institutional repositories for research data”), and developments in the US including the formation of the federal Interagency Working Group on Digital Data and the DataNet initiative funded by the NSF, as well as recent commercial activities by Sun, Google, and Microsoft. Institutions within the US mentioned for moving forward the state of the art include the San Diego Supercomputer Centre (for SRB, iRODS, and Data Central), Purdue University (for its Distributed Data Curation Centre, D2C2), University of Washington and Johns Hopkins University.

Four US universities are named as pursuing educational opportunities in data curation – Indiana University’s School of Informatics, University of Illinois at Urbana-Champaign, University of North Carolina at Chapel Hill, and Syracuse University.

A section on data curation issues covers financial sustainability, appraisal and selection, digital preservation, intellectual property, confidentiality and privacy, and participation by data owners. The recommendations made by the group include the need to seek out and cultivate partnerships, and the need to develop new services for Cornell researchers.

[tags: report, USA, libraries, policy, data curation, data management, repositories, training]

See the rest of my Faves at Faves

Saturday, 27 September 2008

My Faves for Friday, September 26, 2008

Neil introduces the report as a whole. The part about research data is extracted below:

The activities of the Alliance Initiative are directed to three areas: First, the partners wish to formulate a common data policy in order to promote both the need for action and to demonstrate the usefulness of primary data infrastructures for scientists and scholars.

Secondly, the partners wish to foster cooperation between scientists and information specialists and to offer funding for pilot projects. Such projects should develop subject-specific standards and methods of data curation and archiving; they should also define the division of labour required in the process.

These steps have the overall goal of establishing a reliable system of digital archives for primary research data, and to ensure that these remain accessible internationally and their data reusable in various interdisciplinary contexts.
Finally, the third and ultimate aim is to establish a system of discipline specific, internationally networked data repositories for primary research data. However,
this task can and should only be tackled when sufficient experience has been acquired from the funding and evaluation of pilot projects. This is to ensure that
the new structures respond to the requirements of the individual subject disciplines and are embraced by them.

[tags: blogs, report, data curation, Germany]

See the rest of my Faves at Faves