Wednesday, 19 November 2008
Consorcio Madrono - research data seminar
What became evident during the discussions was the generic use of the word 'data' to describe various different kinds of digital research output – highly curated collections, large scale dataset gathering, lab-based data generation, secondary social-science data (researcher-created data) and research data products and summary data. These all have different patterns of generation and use with variable lifetimes and life cycles. Distinguishing between finished research data packages and on-going data production and analysis also needs some thinking out. There are scale, storage and format issues. There is perhaps some scope to delineate between pre-publication and post-publication data - the latter being more likely to be repository (and data librarian) friendly, the former the domain of a new breed of 'data scientists' conversant with subject but with less interest in metadata, discovery and preservation. It may well be time to discuss each of the patterns of generation and use individually with a view to establishing where they differ and where there are commonalities, in addition to articulating curatorial roles, responsibilities and relationships for each of said patterns of generation and use!
Plenty of food for thought!
Presentations from the seminar will be posted here soon.
Stuart Macdonald
DISC-UK DataShare
Wednesday, 12 November 2008
DataShare deliverables over last 6 months
Summary:
The project has changed some of its deliverables following the change in LSE’s status to an associate partner due to staffing shortages; this includes a greater emphasis on data audits at each of the remaining partners to reach out to users earlier in the data lifecycle and to better meet their needs for support in data management. Edinburgh participated in the Data Audit Framework Development project led by HATII/DCC at University of Glasgow and conducted its own DAF Implementation project, both funded by JISC, and led by Robin Rice, with a new team member, Cuna Ekmekcioglu.
The project members have continued to engage with contacts at UK and international institutions – especially in the US and Australia - who are building services for data sharing. The project team has participated in professional development activities, and disseminated deliverables at conferences, in articles, and through their website and blog. A briefing paper on geo-spatial Web 2.0 visualisation tools was written. The findings from Oxford’s Scoping digital repository services for research data management project were disseminated. Several project team members participated in the Edinburgh Repository Fringe, along with peer projects amongst the partners, e.g. Kultur and EdShare (Southampton), ShareGeo and the Depot
(EDINA), and the CRIG International Roadshow (Oxford). An article in Online by Luis Martinez Uribe and Stuart Macdonald and an interview in CILIPS Update brought attention to the profession of data librarians, which was further amplified by the recent JISC-commissioned report by Key Perspectives, The Skills, Role and Career Structure of Data Scientists.
The partners have created and received peer review on a Dublin Core based metadata schema for datasets in DSpace and EPrints, worked on procedures for storing and preserving databases, and have developed a content model for a database of sound files in Fedora. The Edinburgh DataShare repository was soft-launched, with an option for depositors to append the open data license developed by the Open Data Commons.
The progress report also includes specific progress made at each partner institution and a new evaluation plan, to be carried out by Sheila Anderson at Kings College London.
Saturday, 25 October 2008
My Faves for Friday, October 24, 2008
Ben O'Steen's blog post describing the DISC-UK DataShare project's approach at Oxford to incorporating a dataset type (a phonetics database) into a Fedora repository.
[tags: Oxford, blogs, data curation, formats, metadata]
Wednesday, 15 October 2008
Update from LSE - Associate Partner in DISC-UK DataShare
The Library has recently made the first appointment in this new team - with Dave Puplett taking up the post of Data Librarian on 6 October. Dave has worked in the Library since 2007, on a variety of projects focussing on improving access to electronic resources and exploiting new
web technologies. He has recently completed work on the JISC funded VIF project dealing with version identification of academic research in digital repositories.
Dave ('d dot puplett at lse dot ac dot uk') will be the main contact for the Data Library service.
Over the coming months we will continue to develop the Library's data services team and our commitment to the Data Library service remains as strong as ever.
Nicola Wright
Information Services Manager
London School of Economics and Political Science
My Faves for Tuesday, October 14, 2008
"Publishing data long-term with an accompanying journal publication can produce a large number of citations which will be good for your career in the Research Excellence Framework" -- said Michael Wilson, a researcher associated with the MRC Psycholinguistic Database, who presented this finding at the e-Science All-Hands Meeting in Edinburgh last month.
Edinburgh DataShare is an institutional data repository set up by the Data Library to help researchers do just this. It uses the same DSpace software as the Edinburgh Research Archive, ERA, and operates in tandem with the Publications Repository to provide a place for researchers to share the datasets on which their published papers are based.
[University Newsletter article announcing the soft launch of Edinburgh DataShare, a project deliverable and pilot service.]
[tags: edinburgh, repository, service, data sharing, data publishing]
Tuesday, 30 September 2008
My Faves for Monday, September 29, 2008
Gail Steinhart, co-chair of the working group, forwarded me a link to this paper during the summer, and I’m very pleased to have read it. The group, formed in 2006, has been investigating issues, current activities, and opportunities for the Library to get involved in “digital research data curation.” Thus, it serves as a very useful US equivalent to our DISC-UK State of the Art Review, but also hones in on the specific issues within a given institution, which is what I’d like to help the Information Services do within the University of Edinburgh.
The white paper begins with an environmental scan beyond Cornell, before turning to the strengths and potential areas of collaboration within the University. It looks at the actual and potential role of the academic research library, international organisations such as CODATA, activities in the UK including the importance of Liz Lyon’s 2007 report on roles and responsibilities, the EU DRIVER project, The Australian National Data Service and the activities at Monash University (“noteworthy in terms of utilizing institutional repositories for research data”), and developments in the US including the formation of the federal Interagency Working Group on Digital Data and the DataNet initiative funded by the NSF, as well as recent commercial activities by Sun, Google, and Microsoft. Institutions within the US mentioned for moving forward the state of the art include the San Diego Supercomputer Centre (for SRB, iRODS, and Data Central), Purdue University (for its Distributed Data Curation Centre, D2C2), University of Washington and Johns Hopkins University.
Four US universities are named as pursuing educational opportunities in data curation – Indiana University’s School of Informatics, University of Illinois at Urbana-Champaign, University of North Carolina at Chapel Hill, and Syracuse University.
A section on data curation issues covers financial sustainability, appraisal and selection, digital preservation, intellectual property, confidentiality and privacy, and participation by data owners. The recommendations made by the group include the need to seek out and cultivate partnerships, and the need to develop new services for Cornell researchers.
[tags: report, USA, libraries, policy, data curation, data management, repositories, training]
Saturday, 27 September 2008
My Faves for Friday, September 26, 2008
Neil introduces the report as a whole. The part about research data is extracted below:
The activities of the Alliance Initiative are directed to three areas: First, the partners wish to formulate a common data policy in order to promote both the need for action and to demonstrate the usefulness of primary data infrastructures for scientists and scholars.
Secondly, the partners wish to foster cooperation between scientists and information specialists and to offer funding for pilot projects. Such projects should develop subject-specific standards and methods of data curation and archiving; they should also define the division of labour required in the process.
These steps have the overall goal of establishing a reliable system of digital archives for primary research data, and to ensure that these remain accessible internationally and their data reusable in various interdisciplinary contexts.
Finally, the third and ultimate aim is to establish a system of discipline specific, internationally networked data repositories for primary research data. However,
this task can and should only be tackled when sufficient experience has been acquired from the funding and evaluation of pilot projects. This is to ensure that
the new structures respond to the requirements of the individual subject disciplines and are embraced by them.
[tags: blogs, report, data curation, Germany]