bookmark

Showing posts with label data sharing. Show all posts
Showing posts with label data sharing. Show all posts

Sunday, 8 February 2009

Data Walkabout 6: Brisbane, University of Queensland


From the city centre, it is a pleasant and fast ferry ride up the Brisbane River to the University of Queensland. This next Data Walkabout stop gave me the chance to chat with the dynamic Belinda Weaver (at yet another outdoor campus cafe). Although she's on secondment and not currently working on the institutional repository, my impression is that she's accomplished so much already she could be allowed to take a break.

I inquired about the institutional survey which she initiated and Margaret Henty expanded to other universities, Investigating Data Management Practices in Australian Universities. The outcomes provide a baseline of evidence at each participating institution but, like so much else in
Australia, they don't stop there, but take action to foster change.

The status quo for data management amongst researchers was - perhaps depressingly - found to be much the same as that in the UK (through SToRE, DAF and other surveys): often a junior researcher is put in charge, there is no standard practice, there aren't rewards for doing things well, and few consequences for doing it poorly. Often the problem arises only after something goes wrong and data are lost.

Problematically, universities have not seen data management as a responsibility nor something for which they need to provide services. As a direct result of this survey University of Queensland has put data management/loss into their overall risk strategy. Belinda believes a risk management approach is a powerful way to influence institutional senior management to support proper data management.

As we were sipping our coffee, Christiaan Kortekaas was walking by, and Belinda waved him over. Christiaan is the inventor of the Fez open source interface to the Fedora repository software for the University of Queensland Library, which is a competent rival to proprietary solutions. It's quite flexible, and can offer different metadata schemas (e.g. MODS, Dublin Core, etc.) and a variety of classification schemes.

As for libraries, Belinda saw multiple roles (her secondment replacements are pursuing this now). Although data management support is often in no one's job description, it is commonly repository managers who fill this void - perhaps due to the rallying encouragement of APSR, the Australian Partnership for Sustainable Repositories. Specifically, librarians could provide support in describing data structures (metadata & documentation), providing training and templates for data management; writing data rescue case studies; and exit plans for data producers leaving university. Whereas research offices tend to focus on new grants and fostering collaboration, and IT services on servers and cost recovery, libraries are in a unique position to help researchers in finding relevant tools and technology (Web 2.0, etc.) to enhance their research - just as they help them find publications literature. She even thinks that librarians should be based with faculty rather than all in the library building itself, so they can be part of the team. This has worked well, for example in the hospital, where librarians work alongside clinical researchers.

Belinda emphasised that researchers are not necessarily aware that librarians 'know stuff' about tools and technologies, so advocacy is needed. Her publicity poster urges staff and students to "join the growing number of UQ academics and researchers" who are preserving their digital research material with UQ eSpace. Smiling faces of people provide an 'imagine' scenario about materials they can deposit and the implicit benefits of doing so, making sharing research output seem the most natural thing in the world.

Here's a wee gem from Belinda: because repositories are a new service, people don't realise the huge potential. If researchers think they don't need libraries, then adding value to the research chain is vital. So: libraries should be re-purposing themselves around repositories.

Friday, 30 January 2009

Data Walkabout: Wellington


The second stop on my data walkabout was New Zealand's capital. I spent the morning with the Information Management team at Statistics New Zealand learning how they take initiative on documenting and archiving legacy datasets for long-term preservation. I'd heard Euan Cochrane's clever presentation at last year's IASSIST conference, and so I knew Stats NZ is unusual as a national statistical agency for adopting the XML-based DDI standard (Data Documentation Iniative).

I had previously only heard of DDI being used as a dissemination tool before, as within the software invented by the national data archive community, Nesstar, which allows the user to select cases and variables and do basic online analysis before downloading the entire dataset. So I was surprised to hear that while the team marks up datasets in DDI (ver 2), using an XML editor such as Stylus Studio, they don't disseminate them that way, but simply store them, basically in a dark archive which a handful of people have access to, for posterity.

As for dissemination, survey tables and other aggregate datasets are published on the website. For individual-level microdata, there are three ways to obtain them: a personal visit to the secure Data Lab, by requesting and obtaining a "CURF" - Confidentialised Unit Record File, or via Remote Access through ATOM (Access to Microdata). Access is restricted, reviewed on a per request basis, and all involve a cost recovery charge. Individual data on New Zealanders, it is felt, must be carefully guarded since the population is so small and people have unique attributes to which they could be identified.

A newly formed team,led by Hamish James, who along with Euan ensured my visit was hospitable and informative, has a mission of maintaining an enduring national statistical resource. Passage of time has proven that a) data are meaningless without metadata, and b) that there is reluctance from business units to part with data even to an organisational data archive. So the team is working hard to build trust with statisticians who collect and analyse data through effective preservation of legacy datasets. Eventually, workflows adopted by statisticians will ensure that newer data are properly documented and cared for from the start, hopefully making the archiving process easier.

The team collaborates with other preservation organisations in the city, Archives New Zealand and the National Library, who all meet regularly to exchange best practice. They use tools such as JHOVE (to produce checksums for checking data integrity), DROID, which provides a PRONOM identifier that gives a full description of the file format, and the National Library of NZ Metadata Harvester, which produces an XML file from which an XSLT stylesheet is produced. A local script then helps to fill in a PREMIS preservation metadata record.

In the afternoon I had the pleasure of meeting with Isabella Cawthorne and Julia Watson from the Ministry of Research, Science and Technology (MoRST) over coffee near the Beehive Parliament building (pictured). Isabella, as a policy-maker for research funding, is concerned about incentivising researchers to manage and share data to avoid having to fund projects that "reinventing the wheel". She says New Zealand needs coordination to get the best value out of environmental research. Julia is working on the e-Research front: the high speed Karen network has been set up in New Zealand, but applications and middleware still needs to be developed. They both believe BESTGrid is a good "bottom-up" example that could be an exemplar for further collaboration and development.

Their ideal scenario for environmental data sharing is a federated approach (rather than a central archive), but with authenticated access, based on levels of quality assured data. (New Zealand is considering joining the Australian Access Federation, which would offer a Shibboleth-based approach to authentication.) They shared a discussion paper commissioned by MORST called Environment Data 2.0: building the digital platform for a sustainable future, which sets out this vision.

I found this substantial food for thought: what can policy makers and funders put in place to best encourage data sharing in research?

Monday, 7 January 2008

Edinburgh DataShare takes shape

As part of the collaborative DISC-UK DataShare project, Edinburgh University Data Library has now launched a pilot data repository for members of the university to use for sharing their data, much in the way that Edinburgh University Library has been offering an open access document repository for a few years already (called ERA). The new DSpace repository is called Edinburgh DataShare, reflecting its part in the wider project as well as its primary purpose.