bookmark

Thursday, 30 July 2009

Repository Fringe 2009 : Pecha Kucha

What is Pecha Kucha? A formal presentation style, which gives rise to some very inspiring talks. In essence, you have 20 slides in your talk (not 19, not 21), and each slide is displayed for 20 seconds. The whole presentation therefore lasts for 6 minutes and 40 seconds.

This morning's sessions are in three bundles of three speakers each and we'll be voting them (using golden nuggets and cowboy hats to indicate our preference!).

Group A

James Toon, University of Edinburgh
: ERIScotland
  • ERIS is: Enhancing Repository Infrastructure Scotland
  • IRIS Scotland was the predecessor to this project and it looked to see if central rather than local repositories would be more useful. The project was considered a success and we wanted to build on the momentum
  • Grant Funding Call 12/08 - Strand A5 - Repository Enhancement
  • IRIS took a top down view of the issue, ERIS is taking a bottom up view of the issue working with Research pools
  • Reasearch Pools have issues too and we need to get a better understanding of their needs.
  • The community developing repositories haven't engaged as much as they should have done with research users and that leads to problems
  • They are working and talking to researchers and repository managers together.
  • We need to find a unified way to curate and encode repositories on a national and global scale.
  • We also want to make recommendations and suggestions for training and policy around repositories
  • Not only the functionality but also the tech is being considered as part of this work.
  • We also want to deliver enhancements already proposed - these will be full implementations not just demonstrators.
  • BUT we are not developing just because we can. The development teams are working closely with the user engagement work
  • Most of the goals are based on work already started
  • We don't know what the research pools will actually want. We know local and aggregation factors will both be important.
  • We only have a vague idea of what is needed at this stage but we have to be more confident about this if we are making high level recommendations.
  • Planning, making a business case and policy are all crucial to having solid long term impact.
  • One big assessment of user needs and demands summarizes must of the work of ERIS.
  • We think this work is achievable but we have to realistic about long term preservation and access.
  • We need unified communities, be trusted and build a sustainable repository.

Les Carr, Southampton: "Repository Challenges"
  • When you set up a repository you must fit many many targets and multiple agendas.
  • There is value in the ability to share freely but there is a catch...
  • We have to adapt to the web as researchers. We are not used to doing our science in public.
  • Preservation and saving our material is a huge problem. From a data point of view it's an enormous bogeyman problem but we have to do what we can in the realistic way.
  • e-Learning - the ability for repositories to not only preserve articles but also data and materials for e-learning. There are all sorts of ways of researching that are expensive and very non textual and the only account is journal articles and notes.
  • Business is part of academia and funding and business cases are important.
  • Repositories must be efficient and effective. It's for the repository to provide these services.
  • We have set up repositories to be like a box of lego - so you can put data together as lots of modular components
  • We need to Pimp our research ride.
  • The cloud: there is so much power and capability both in the cloud and on our computers. Fitting into day to day working practices and activities is crucial.
  • A repository is like gears on a bike. It mediates between the components of the research world
  • We don't need to be chemically enhanced to do well but repositories DO need to be helped to be as useful as they promise to be.
  • No technology on it's own is the answer. The combination and the blend of people with technology that makes for so much more powerful a future.

Guy McGarva, EDINA: ShareGeo - Discovering and Sharing Geospatial Data
  • This is an overview of ShareGeo. A deposit tool that forms part of Digimap.
  • Digimap provides access to licensed geo data sources.
  • A lot of data exists out there and it's hard to find and use - especially if derived from licensed data.
  • The repository uses DSpace and allows stats and other functionality to track use and connect to services.
  • The data in ShareGeo can be either open access data or that derived or tied to licensed data.
  • Example data includes land use, grids, research generated metadata etc.
  • We are trying to formalize the process of sharing and reuse.
  • ShareGeo is based upon the work of the Grade project which found a need for this type of specialist sharing of licensed data.
  • We currently have a fairly high numbers of logins, users and downloads but little upload of data.
  • We use a map based search and spatial queries of data sets are enabled making the finding of data easier for users.
  • The footprints of datasets are shown on maps which is useful but quite powerful when combined with spatial queries.
  • ShareGeo take a single zipped file for deposit of material which means one or many files can make up a deposit. We have a size limit of 1Gb at the moment but this is just to make managing the service manageable.
  • The automatic ingest maps the data for ShareGeo and geospatial metadata is automatically added to the item in ShareGeo.
  • Licenses are part of the deposit process so that you always know what type of data and usage you are dealing with.
  • Issues regarding take up include the difficulties in the closed nature of the site. There are also commercial sites that also provide data sharing facilities.
  • Future improvements include looking at sourcing and adding more open data, creating a sister open access version of ShareGeo etc.
  • The main issue for us right now is how to get more data deposited and how to build up our community of users.

Q&A for Group A

Q (Balvier, JISC): How are you doing the aggregation for ERIS?

A (James Toon): The NLS leads the aggregation. Very standard aggregation in use right now...

Follow up (Balvier, JISC):
Have you spoken to Paul Walk at UKOLN as they are working on something in this area.

A (James Toon):
We've chatted. Right now normalizing data is really what we want to be able to do to provide a good API.


Q: Regarding ShareGeo: How hard was it to get PostgreSQL to do the Geo searching.

A: Not too bad but we do it in quite a basic way from lists. We're looking at a more geospatial extension that might allow a more sophisticated solution.

Follow up comment from the floor: You might look at LocalSOLR.


Group B

Richard Jones
, Sympletic: Symplectic Repository Tools
  • Richard will be talking about some of the repository tools that his company Symplectic produce.
  • Richard works on repository integration tools to go with the repository systems they make
  • The aim is how do we provide a deposit tool to make things easy and efficient.
  • We're starting with an image of Researcher publication lists which form part of their repositories.
  • And a full text tab - you can see what file you have uploaded, permissions and the ROMEO publisher policies.
  • Publications pull in data from lots of sources. They connect the repository with SWORD and AtomPub as the method.
  • Why not just sword? Well it's only designed for creating not updating/removing/changing items.
  • AtomPub exists in a RESTful environment so extra functionality can be added in.
  • Some real complications though. Repositories are designed to be static but Symplectic is a more dynamic environment in a constant state of flux.
  • Repository workflows are a complication - there are three stages really: working copy; review; archiving.
  • So if you blend static and dynamic repository systems what do you get? A really complex slide - but Richard assures us that we can find out more in the tutorial tomorrow!
  • Benefits for the researcher - you can update and correct your data if/as needed which can also mean better metadata creation.
  • Where next? Lots of bells and whistles with repository tools linked to publication tools as well as helping to make standards grow.
  • Richard will be talking more about this tomorrow.
Julian Cheal, UKOLN: “Repository Deposit Using Adobe Air
  • What is it that we're trying to capture?
  • Academics write things in their notebooks, it's not easily absorbed into repositories.
  • You could make researchers work on computers...
  • There's a quote from Bruce Chatwin: "Losing my passport was the least of my worries. Losing my notebook was a disaster" - data is important to researchers and they need to know it's safely archived.
  • But we need to make it more straightforward for depositing materials.
  • Adobe Air is a runtime environment and combines the world of the web with the desktop. It's cross platform. It's a rich internet application.
  • So who's made an AIR? Various twitter clients, The BBC and various advertising companies for a start.
  • Academics want stats and relationships to funders etc.
  • Julian has thus made a prototype that looks - deliberately - very much like Flickr uploader.
  • The finished product is able to drag and drop, easy to use, and pretty to look at.
  • It's a small application - Julian is showing us all the files involved and it's only a few.
  • Academics want easy repositories so drag and drop functionality on their desktop is perfect. It uses as SQLite database so can synchronize offline data as soon your machine is re-connected to the internet.
  • The application talks to SWORD, looks up ROMEO, the name project etc. to catch automatic metadata.
  • Screen shots indicate that you drag and drop, add metadata as you want. You can add lots or less metadata as appropriate. Auto-complete makes this easy.
  • JISC has offered to have a deposit event to combine all the deposit apps. This will take place in October.
Hannah Payne & Antony Corfield, The Welsh Repository Network: "The Welsh Repository Network: A tasty bit on the side!"
  • URNIP is the JISC repository enhancement project for the WRN.
  • Wales has a diverse HE landscape - very varied size institutions with vary different needs.
  • We have face to face and video conferences with institutions and we're doing site visits.
  • Each summer there is a library and IT development event and we are using this to communicate with our project partners
  • We share support calls and share work via Google code.
  • We are working with 4 partner institutions on deposit to see whether deposit increases with changes to the deposit process. But we are only at the pilot stage at the moment.
  • We will be reporting the best models, policies and possibilities as part of this project.
  • e-Theses and dissertations: the National Library of Wales already collects all paper copies but we want to see if they can be a hub for electronic deposit too.
  • The e-Thesis project will connect preservation and metadata functionality.
  • Auto-complete is a hit so we are looking at the work of a previous project Deposit Plait which looked at harvesting and checking data via web services.
  • Users wanted import and export of metadata to link to other educational databases and services.
  • Embedded players and multimedia deposit were the highest user priority. Holograms, art works and film are all key research outputs for various institutions in Wales.
  • We are not a standalone project but fit into the wider repository landscape. We want a cross-project forum so we can establish a set of services and support across repositories in the UK.
  • Diolch yn fawr am gwrando! (Thank you for listening).

Q & A for Group B

Q (Les Carr): The scottish ERIS project is from shared research, is the welsh one more based on library collaboration?

A (Hannah): Yes. ERIS is very research focused. We are perhaps a step back looking at collaboration and development at this point.

Q (Hugh Glaser): For the adobe air application, where do you get data for auto-complete functionality?

A (Julian): I use SWORD APIs where possible but some data you have to grab and work around as not everyone has a suitable API.

Q: You don't send data to the national centre for text mining for instance to find keywords?

A (Julian):
If they have an API that can be used then I'd be very happy to tie that in to the tool.





Group C

Joyce Lewis, Southampton: "Marketing and Repositories - Tell me a Story"
  • Joyce is talking about the importance of stories and how repositories can help us to tell stories.
  • "People don't care about cold facts. They care about pictures and stories" - Nancye Green.
  • Back when Joyce started at the University they did news releases and the broadcast media weren't really targeted and only a few releases were picked up by the print media. Once published the story was also lost.
  • The university environment has changed now. Lots of universities, lots of research and lots of enthusiasm to show.
  • Quote being shown here is along the lines of the fact that universities do a poor job of telling investors what they get for their money.
  • At the moment Joyce tells the stories about the university through text with links and a picture but there is SO much more on the web that could be used. We don't want people to get tangled up in lots of unlinked resources on the web though.
  • Impact is key to the RAE and REF and this has to be thought about when we think about what we record and promote and how.
  • Project called Tell Tale is about telling the story of research. It's not funded yet but fingers crossed...
  • The project would catalogue adaptations to a repository necessary to capture the research story.
  • It would involve enhancing the content by putting it together into a story.
  • We want to create narratives automatically with story templates and narrative generation software that links around to items.
  • Then what is left is to demonstrate success through these stories and the usage of the content they talk about.
  • The bottom line is that we want to tell a better story.
William Nixon & Gordan Allan, Glasgow: "Enrich - Research System and Repository Integration"
  • William and Gordan are talking about the JISC funded project Enrich.
  • The project aims to bring disconnected research elements together.
  • Research systems are miles from repository systems...
  • What is a research project? It's an idea. It may or may not have funding or licenses or artifacts associated with it.
  • There is a sense of research alchemy. And some research MUST publish, others may not have to.
  • The research lifecycle includes a short burst of publishing but a lot of unpublished work.
  • University of Glasgow's Research System which tracks funding and licensing and we've started connecting that to the repository.
  • Repositories need to relate better with the research systems. Records can, when set up this way, now be pushed out via RSS and Twitter for instance. Enlighten is a service whose use has been growing more and more. They are at about 40% full text right now but a requirement to deposit material at the university should get the repository nearer the 100%.
  • In the old day repositories and research systems were separate silos.
  • Junction boxes are the future - we're about services. Turning the repository as a junction box to other resources. For example data from repositories is used to generate publications list on staff pages at the university. To add your publications you have to deposit them.
  • Most searches are not native - people come in via Google and other search engines.
  • We have freely available global open access.
  • But key to success are good relationships, easy clear systems and processes and the university policies really help us to be successful.
  • Enrich will bring together lots more data to tie into the hybrid repository/research Enlighten service.
Jo Walsh, EDINA: "Geoparsing text"
  • Jo wants to introduce herself to this community as she is just starting to move into a new role with EDINA and to engage with the repositories community.
  • Geoparser - developed in this very building - which is based on a grammar based named entity recognition technique that allows geo tags to be added to text automatically.
  • The recognition links to the Gazetteer service. They work well together and the more places you find, the more accurate the look up will be. Text context creates an idea of geo context.
  • GeoCrossWalk has been around for a while and it has a service status now. It uses Ordnance Survey to identify places. It's an enormous and useful service but it has been limited to Digimap users and licensed users. This confuses users.
  • This year we will also be expanding the service with an open access Gazetteer using geonames.org and the same type of system as the licensed version. The results will be variable BUT geonames has a wiki style ability to edit so errors can be identified and fixed.
  • Geoparser webservice will be a simple RESTful API for document placename extraction and markup. You can use OS data OR the OpenDate Gazetteer.
  • Jo is looking for a sense of user requirements and how this tool fits in specifically with repository needs.
  • Linking items across the repository seems to be one useful case.
  • You might use techniques to bootstrap geographic metadata for archives of textual components.
  • Spatially searching archives and nearby related material would be another use case.
  • Please contact Jo with comments, feedback and use cases.

Q & A for Group C

Q: What kind of licence do you pick for the open access geo stuff?

A (Jo): Actually it's from other sources and inherited.


Q: How do you feel about where institutional repositories are going?

A (William and Gordan): We feel more like an 8 year overnight success right about now. We've visited every department in the university. Everyone asks How not Why deposit these days. They used to ask why they should. We've seeped into the research process. We've been very supported by our Vice Principal for research. We're really started to realise the potential of all the data we have been gatehring. And we are in a post RAE, pre REF place so we're looking at how to repurpose the repository to suit that change best.

Repository Fringe 2009: Welcome and Opening Keynote

Today and tomorrow this blog will be covering 2009: Beyond the Repository Fringe and we're just about to get started with an introduction from Simon Bains, Head of Digital Library at Edinburgh University.

Simon is introducing us to our lovely venue for the day - The Informatics Forum - and the fact that we are due to have a fire alarm this morning so there may be a very short gap in blogging. Simon is also introducing our official Welcome from Sheila Cannell, Head of Edinburgh University Library Services.

Sheila is warmly welcoming us with a magnificent image of the Udderbelly in Bristo Square which is one of the most visible temporary Edinburgh Festival Fringe highlights in the area this month. Sheila urges us to think about giant purple cows as a way to think outside the box and lays out some goals for the next few days from her role as a manager with a huge interest in repositories for increasing open access.

There are huge changes around us and the current financial climate will act as a catalyst for changes to the methods of scholarly communication. We've been talking about these ideas for years but the edge of the financial crisis may be the trigger for real change.

We have traditional scholarly communications in journals and we have open access. will finances change that balance. Have we normalized the processes of repos in our day to day practice. How is the funding and staffing set up - are resources permanent or short term. Do we need to normalize their place in the institution?

There is a three fold role for repositories:
  • curation
  • promotion
  • marketing
But does that lead to difficulties if we don't have a clear message about what our repositories are about.

You will talk about multiple repositories and depositing to multiple repositories but how do we deposit to them in a simple unified way, it's good for preservation to have many copies but how many times are we having to deposit material.

We need to think and understand whether what we do with repositories is important to the researcher and whether they still see the traditional scholarly communications as the main aim.

And disciplinary difference: biologists work and think differently to physicists or mathematicians or social sciences in terms of scholarly communication and in terms of scholarly practice.

All of the topics Sheila has talked about would warrant a paper and she hopes that some of them will be progressed through discussion and consideration in the course of the next few days.

Sheila closes by hoping that the Repository Fringe 2009 is a success and that everyone has a chance to see and enjoy Edinburgh as part of their visit.

Simon Bains returns to introduce our keynote speakers Ben O'Steen and Sally Rumsey.

Opening Keynote:
Ben O’Steen and Sally Rumsey (Oxford) – “A sneak preview at the A-list stars of future repositories: blockbuster technical developments and the cultural drivers behind them”

Sally opens by explaining that she and Ben will be handing back and forth with Sally looking at the more library view of repositories whilst Ben will be talking about the more technical whizzy end of affairs.

Sir Thomas Bodley set up the library in Oxford and Sally is taking us through the history of the library including a lovely quote from Francis Bacon that the Bodley "is an arc to save knowledge". We're are also looking at search, 1620 style: a paper list.

The original library building fast ran out of space and the Radcliffe Camera, the Radcliffe science library and the new Bodlien library were all built. By 1914 the library received a million items a Year. It continues to grow and grow and Sally shows us a preview of the storage facility in Swindon which will be helping the Bodley deal with the volume of material by 2010.


There is a usage agreement for the library - you must not kindle any fire or flame for instance - based on the traditional one that users must still sign today. And the sign on the library states that it is a "republic for lettered men". This is an interesting phrase for thinking about repositories. Are we building the digital equivalent of the Bodly? The growth curve for repositories so far is encouraging but we cannot grow such resources over night.

Realizations as a catalyst for change - the realization can be as important as change itself.
Repositories can be treated as a concept and Sally uses the term in its plural on purpose. The single repository as a thing is on it's way out. The repository as a box is no more. They may even be invisible as they are built into other services. One factor that's moving us away from the stand alone is integration with other technical systems as well as changes in soft academic systems. Repository staff can also act as catalysts for networking and collaborative working especially in the light of the REF (Reearch Excellence Framework).

As we move forward we are beginning to achieve Clifford Lynch's idea of repositories as a set of services.

Over to Ben who starts by saying that the internet is the most successful repository in the world. We are separating the service and the storage and that's what should be occuring. They should be distributed across a number of notes. There should be multiple ways to search and access content. Any service or storage can disappear or be added or upgraded without effecting the other systems unduly. If you lose your index of a repository you can be lost, you want this internet method of multple acces spoints fixing this. You want to make your repositories like the web.

"The future is here. it's just not evenly distributing yet"
- William Gibson, NPR talk of the nation 1999.

Ben is explaining that he knew the quote but it took a while to find the citation. He eventually found it on someone else website and connections and this is an example of how people look for information (find out more here: http://bit.ly/89AtD). The citation was found using Google to find where the quote was from. The web is the world. The web is usage. Following the trail of usage through searching and contacts and then through to a recording. But NPR have changed their website since the citation was originally found and they changed the metadata trail for this quote. The new URL has a single ID to that single recording now.

People search for things. It's incidental that they can find the documents not the things. Search is about full text.

There are some issues. All things have names of some sort. But there are things we need to fix. Dates and events don't always work on the web. We can do that though, this is the challenge for repositories. We can provide documents that directly relate to a thing. we can provide URLs. this will help people find things in repositories better than they currently are ale. The key is knowing HOW the document relates to a thing - critique, reference etc.

A second realization - we've been giving ourselves names on the web for a while. Ben is showing his various web IDs (on Twitter, Facebook, etc). How about we do this more evenly so that there are pages for projects and researchers that link to existing names.

Power comes from the power of relating names. This is unbelievable powerful.

And there has been a social seachange in how people appear on the web. Rather than "do you have a profile on.." to "are you on.." - people are themselves online, not some random profile but all linked to the real person. And where are we going with this...

Linked data and http names means names for things and connections between things.

Library of Congress are publishing their authority lists as linked data in RDF. That's a great way of making LCSH more useful on the web (See: http://id.loc.gov). Yahoo and Google index RDF embedded in HTML pages (as RDFa) and that's hugely useful for linking and connecting and making search more useful and items more useful and visible. You need clear ideas about what you are doing as a repository.

Back to Sally: In some areas policies are in place and well developed and they should drive everything. OpenDOAR provides a tool to help with this, but the Preserv project at Southamption found a real gap here. DISC UK DataShare are also looking at issues of repositories management (have a special session this afternoon at RepoFringe2009).

There has been a huge expansion of items and types of items that you expect to see in repositories since 2000. Originally they were for refereed published literature. But now a much much wider range of materials are being handled (eg JORUM). Some of the most successful repositories have been been single subject repositories. Academics like them and use them and continue to deposit in them. We need to be able to deposit in multiple places at once and policy must drive that. We need that to get buy in from a lot of data creators

Back to Ben: We have reinvented too many wheels already. Don't fight it, work with it. Use the standards that people are using now. Defacto standards are important so don't feel tied to the ISO standards if they are not what is actually in use. Already out there are:

Transfer
  • Files - http
  • Lists - atom, rss

Create update etc
  • HTTP POST, etc.

Names
  • URIs - these are already used by people connecting to Wikipedia for instance.

Lookups
  • DNS resolvers

Using what is in use means instant communities. They have techniques and tried and tested software to access our materials already. You don't have to write things from scratch. You can experiment quickly and usefully. How do these fit into researchers workflow? If they don't you need to ditch it and move on. Tools and techniques may not be perfect but might do the job.

If you genuinely are doing something new you need a community to assist you, if no one else is interested than should you be doing it that way? There are some projects who go ahead on their own when good alternatives are already available and in use.

We don't have Defacto standards for:
  • real time event notifications through the browser.
  • Simultaneous collaborative document editing (Google Wave may have some relevance here).
  • Data qualified and ranked by evidence. Search engines do very poorly at this as: how can they know and use what YOU trust.

So, audience participation time: Ben asks us to name some repositories:

  • flickr
  • youtube
  • kfupm

And adds a long list including:
  • slideshare
  • facebook
  • google docs

People use all these things. They are useful and links between different versions of the same documents are useful for qualifying all connected items.

There are no common standard or APIs for these repositories but they all contain a set of things. And these are useful things.

If you want to get stuff from these repositories into yours you will not get a sip. You won't get a focused package of what you want. There are mechanisms being developed but nothing so far and many repositories don't have these abilities. What's now? What's current? Use that!

Realisation: object transfer is still in a divergent state. For the moment we just have to cope with lots of containers and folders. No negotiation for the format of a SIP: you deal with what you are given. And sometimes you have to harvest what you can (e.g. Pubmed - you have to grab and keep your copy).

But we can cope because there are is a Normal Archival Process already in place (it's for physical objects but the digital issues are similar):

  • Accept delivery of boxed of stuff and record roughly what was received. Things get permanent IDs now. That means you have an audit trail and provenance for all items.
  • Triage the contents within a stable environment: deal with fragile things first, things that will deteriorate; sort out issues that arise with rights holders, depositors - this is always a dialogue; some things may stay in the box for a LONG time and just be on a shelf waiting but as long as they have been triaged the box can sit until it needs handling
  • Identify actions that need to be taken to ensure future access.
  • Characterize and catalogue the contents using relevant tools. For instance Oxford have an ephemera collection that just doesn't work with MARC so a new schema was needed. This action will sometimes be called for.
  • Update archival records so that people can find the content (if they are allowed to).

So we already do this. This is our accession process.

The media may be different in the digital deposit/accession version of events but the process need not be and/or can evolve as necessary.

Not all storage is the same:
  • The absolute biggest benefit to any repository is to separate out the concerns of storage and services . It will make your life so much easier.
  • Oxford have a "bitbucket" - a huge safe storage machine where things can sit until they can be dealt with.

Hardware, software, people and storage will come and go. You content is constant. And we need to respect what scholars deposit because of that.

Back to Sally now for experience of scholars and repositories:

"When it's one click deposit I'll do it"

A diagram of what researchers should do in the deposit and publication process explains the confusion and barrier to deposit. So we need a way to make things clear and easy:
  • Deposit by stealth and through other easy solutions
  • Multiple repository deposit regime (MuRDer!)
  • Answer related problems that worry people such as the issue of multiple versions
  • Automation. automation, automation
Nature publishing have recently offered to deposit items into repositories and maybe even Institutional Repositories (IRs). Will other publishers follow suit?

Copyright: wouldn't it be nice if things were uniform between publishers? Some are becoming more open but it's a very long way from consistent.

A recent BL report highlighted that "restriction thretened to lock away digital content in a way we would never countenance for printed material. ". Researchers are used to a more open way of working with print and with other (non repository) items online.

Legal deposit as a parallal to repository mandates and their role for archiving and access. In 1610 Bodley did a deal with the Stationers and Newspaper makers company to be able to request a copy of anything published. If copies ran/sold out the Bodley could be used to find/replicate an item. Strong parallels with repositories and deposit in them. Perhaps we should be aiming for universal scope, independence and size (as mentioned in a 1910 Bodley document) with our repositories?

Preservation aims towards preserving access. Assured secure storage and permanent access needs to be well managed. And aided by intra-library agreements and funding moves.

Shared and distributed expertise. Example being mentioned here is the LC putting collections on Flickr - the metadata isn't going to be perfect but you get some metadata created quickly and some will be good.

Recent RIN report: Creating Catalogues: Bibliographic records in a networked world looked towards making material available and findable in repositories but will it really come true? It would be great if it did.

And back to Ben: we are looking at/for disproportionate feedback loop
  • The perception that a small effort leads to a very great benefit.
  • This leads to the idea that more little efforts have bigger results.
Ben is showing a duck hunt screen capture. High scores are technically trivial but psychologically important in gaming. Are usage stats for scholarly items any less useful in this way? Reusage stats (trackbacks, tweets and references) are incredibly important. Vanity stats can really drive deposit. Another screen capture of the new Ghostbusters game is on screen now - 6 buttons to hit for huge feedback. How many boxes and buttons do we have on deposit forms? We need much better feedback if we actually want people to deposit their work.

Back to Sally: peer review is super important. Knowing how lab books, data etc fits in is also important. A journal article is just a summary of research and things could change very differently if other outputs become available or take over.

There are new forms of dissemination and publishing - a semantically marked up article which has been commented and colour coded and linked back to other items is being shown. The meaning of the word is highlighted and links out, the article links to figures, data and other items so that they are all available as part of the article. It's not an article it's a huge resource.

Aren't more people going to want to do this? Once authors find out it's possible they will want to do it.

Open Access
We start with the example of Sally's ancestors. They marched on land but needed a permit to access this and though the landowner under used the space he blocked access so there was a mass trespass to prove the point. Some were imprisoned for their actions but the march resulted in a law change - the introduction of the Right to Roam (updated in 2000). What many people had failed to realize was that the land had been free to access before and should be again.

There is a parallel here to Open Access. The change in legislation that the trespassers got should be a positive indicator.

There is a perception that Free isn't good and we have to change that. So many complex open access options of authors to deal with. Unfortunately they will probably hover around for some time but hopefully things will get more manageable.

And finally a preview...
  • We think repositories are really moving. It's going to be long slow incremental change.

But we are still waiting on
  • Easy multiple deposit.
  • Collaboration between publishers and IRs etc.
  • Simplifying of everything.
Final word from Ben: Print on Demand is going to be big. People will take what's useful to them and mix it up a bit. What does a book mean when it's £2 and you create it in minutes? You can have a printing machine available to open up access and mixing options.

  • You can print off a set of articles into a book on a librries book printer
  • your colleagues comments tweets and reviews are interleaved with the test
  • Your colleagues were found from your Professional networks
  • YOU can do all this already!

You can create a bookmark list of plates from 18th century books online which you believe to be the work of one anonymous artist - this list is research in itself

Permanent books, temporary magazines? Is this true? How about facsimilies etc?

We've been talking about preservation and access so some demos to close the presentation:

  • Ben shows us a book printed on demand, another from a facsimile, and a traditionally published one. They all look the same.
  • We can't preserve access to all research. Research on a computer game has to be emulated. you can't preserve the actual researched activity any other way at the moment.
  • We don't have all the media we want yet. People continue to create new media - Ben shows a video (included on this blog somewhere: http://blog.karagos.com/) BUT you can pan around the video: this is a new form of video. This could be a way of broadcasting. This could be an archive of a choreographed piece. Research is not just text. It can be all sorts of formats.
  • Ben takes a picture on his phone. And he's got a £25 mobile printer that print out stickers. You can send them from your phone and have a printed image from a wallet sized printer in a few seconds.
Laptops aren't what you carry all the time. But your phone is and the mobile printer lets you grab a paper print of a map or similar on the move. There's lots of this stuff out there already. What people say they want is NOT what they actually want. Print on Demand is going to be good and useful for research. It's what people will be doing soon!

And on that Ben and Sally conclude their session. And Simon speaks for us all in saying that that was tremendously interesting!


Q & A

Q (Ian Stuart): People have identities. Personal and Professional identities are separate for many at the moment. Can people mix identities in this sort of linked landscape?

A (Ben): The power of linking is huge. If you link identities that is a double edged sword. It's useful but can get you in trouble too. Some use consistent nicknames online for their personal presences but some are just getting better at managing their online presence in general. Some universities are leading the way to get a professional presence online but we need to encourage that and link to materials and qualify work appropriately.


Q (Les Carr): We're almost 10 years since the first meeting of the Open Archiving Initiative. We are now even less able to define what a repository is. And yet our institutions are set up to deliver applications that are well defined and look like databases. How do we deliver services that are relevant and business critical but also open and flexible.

A (Ben): You don't have to have all the data about an item in the same place as the items, you just need names and connections. You can keep some domain separation but it can be political.

A (Sally): Sometimes you need to just do something, you can't get it perfect, you have to demonstrate what is possible. Demonstrating what is possible may be what's needed to move forwards.

Wednesday, 29 July 2009

Preview of Forthcoming Attractions: Repository Fringe 2009


My name is Nicola Osborne and I am the Social Media Officer for Edina. For the next two days I will be using the DataShare blog to cover the Repository Fringe 2009 event which is being held at the Informatics Forum of the University of Edinburgh on Thursday 30th and Friday 31st July 2009.

I'll be sitting in on most of the sessions and posting up notes and summaries here as well as Tweeting the highlights of the Fringe. If you are interested in joining in from your own desk you will be able to see streaming video of some sessions, view the Tweet stream and take part by commenting, looking at images, taking part in polls, etc. via the Repository Fringe 2009 website and our CoverItLive stream. Or just keep an eye on this blog all day Thursday and Friday.

If you have any questions or comments please leave a comment below or email me (nicola.osborne@ed.ac.uk). If you are attending Repository Fringe then please say hello and let me know how I'm doing with the live blogging. If you are blogging, tweeting or uploading films or images please use one of the hashtags #RepoFringe09 or #RF09.

Thursday, 16 July 2009

My Faves for Wednesday, July 15, 2009

Blog post from Chris Keene on the "The Data Imperative: Libraries and Research Data" event held at Oxford in June.

Interesting summaries and comments on speakers - Paul Jeffries, Luis Martinez, Sally Rumsey, Alma Swan, Simon Hodson, Martin Lewis.

[tags: research data, data management, Oxford, blogs, data librarians, data repositories]

See the rest of my Faves at Faves

Wednesday, 17 June 2009

My Faves for Tuesday, June 16, 2009

MIXED is a project of DANS, Data Archiving and Networked Services. MIXED is to contribute to digital preservation, by dealing with the problem of file formats. Over time, file formats become obsolete. When that happens, the information in such file types is no longer accessible. MIXED follows the strategy of converting files to XML as soon as possible, preferably when data is ingested into the archive. MIXED also converts these XML files to formats of choice by the archive user.

[tags: Data Curation, Data Preservation]

See the rest of my Faves at Faves

Tuesday, 9 June 2009

DataShare final deliverables

The official project period has ended, though follow-up activity continues. It may be worth noting some recent deliverables:

May, 2009: DISC-UK members along with Ann Green (Yale) and Gail Steinhart (Cornell) gave a half-day training workshop - Data Requirements and Digital Repositories - to twenty-one data professionals at the IASSIST/IFDO 2009 conference in Tampere, Finland, based on the recently published Guide (see below).

May, 2009: Policy-making for Research Data in Repositories: A Guide is available for download (Adobe PDF). The guide is intended to be used as a decision-making and planning tool for institutions with digital repositories in existence or in development that are considering adding research data to their digital collections.

May, 2009: DataShare Final Report now available in JISC Repository. Executive Summary also available separately [PDF].

May, 2009: Two papers were contributed to the most recent edition of IASSIST Quarterly: the project manager summarised the work of the DataShare Project; Luis Martinez-Uribe reported on the requirements gathering exercise on researchers' needs at Oxford.

April, 2009: The project manager gave an invited paper - Lessons Learned from the DISC-UK DataShare and Data Audit Framework Implementation Projects at the Digital Curation Practice, Promise and Prospects (DIGCCUR) conference in Chapel Hill, North Carolina, 1-3 April.

Many thanks to everyone who got in touch with us during the project. It's been real!