Showing posts with label metadata. Show all posts
Showing posts with label metadata. Show all posts

28 Jun 2021

Remote Cataloguing Projects During Lockdown at the Library of TCD and the 1872 Printed Catalogue Conversion Project

The LAI Cataloguing and Metadata Group's AGM (2020) and networking event took place on 10th March, 2021 hosted virtually by TCD Library courtesy of the LAI Zoom account. This event featured a number of presentations on remote Cataloguing Projects during Lockdown at the Library of Trinity College Dublin and the 1872 Printed Catalogue Conversion Project.

Christoph Supprian-Schmidt (Acting Keeper, Collection Management), in his opening remarks outlined the situation for cataloguers during lockdown and explained how Library of Trinity College Dublin used the COVID-19 lockdown period to work on a range of remote cataloguing projects. He noted that at the one-year mark of the original lockdown, these projects will have added well over 200,000 records to Trinity's main online catalogue, Stella Search with most records are coming from the 1872 Printed Catalogue Conversion Project, - the focus of the presentations of the evening. 

While the loading e-book records, cataloguing digital collections and the remote cataloguing of new books (from scanned title and key pages and other cataloguing data), continued, TCD Library also used this time to work on the 1872 printed catalogue conversion project. This work was facilitated by access to scanned pages (and OCR data) of legacy catalogues.


 

The TCD Library Printed Catalogue 1835-1887 by Trevor Peare, former Keeper (Readers’ Services)

Trevor Peare presented the historical background to the 1872 Printed Catalogue and its 40-year conversion project. James Henthorn Todd (1805-1869) was primarily responsible for the first TCD Library printed catalogue. Todd entered Trinity 1820, graduating with an honours degree in Science in 1824 followed by a Fellowship and ordination 1831. He was appointed Assistant Librarian 1834 and finding the existing catalogues inadequate, he began work on a new library catalogue in 1835. By 1846, the entire library had been re-catalogued., Todd was appointed Librarian in 1852 and the first volume of the printed catalogue was published in 1864. Todd died in June 1869 and Henry Hutton and Jan Hessels were appointed as editors 1872. The final volume of the nine-volume set was published in 1887. The edition of 250 copies included 48 Presentation Copies.

John Gabriel Byrne entered TCD in 1952 and graduated top of the class in engineering in 1956. He also studied French, Latin and Greek. He completed his PhD 1957 –1961 and began lecturing in 1963. He was appointed the first Chair of Computer Science in 1973 and became interested in the printed catalogue in 1985. He arranged to begin scanning of the original text in 1990 and the first database and search system 1993 was available in-house in TCD in 1993 and on the internet by 2005.  The 5121 pages of one set of the eight volumes were separated in 1987 in order to make a microfiche copy and these pages, which were provided by Dr. Charles Benson, Keeper of Early Printed Books, were used to develop this on-line system. There are about 250,000 entries in the catalogue (including 'see references').The catalogue contains entries in at least eighteen languages. English and Latin occur most frequently and other languages in the Roman alphabet include French, Italian, Spanish, Portuguese, German, Dutch, Icelandic, Danish, Norwegian, Swedish, Welsh and Irish. 

Dirty secrets of OCR, or, how to wrangle a big set of bibliographic legacy data by Joe Nankivell, Junior Bibliographer (Early Printed Books and Special Collections) 

Joe Nankivell described the process of transforming the raw data from John Byrne’s OCR project into MARC records, with a particular focus on the data-cleaning side. The data had been shared on a memory stick that contained all Professor Byrne’s project files, including the records that formed the basis of his searchable online version of the printed catalogue. These records were distributed across over 5,000 files, one for each printed page, each containing between 30 and 50 records. The first task was to merge all these records into a single file where they could be manipulated in bulk, to impose consistency across the dataset prior to line-by-line proofreading by the wider team.

The OCR records had a superficial resemblance to rudimentary MARC records, with clearly identifiable bibliographic elements – a main heading (usually the author), the title, an imprint statement with place and date of publication, and a shelfmark. They could not be simply transformed and loaded, however, for three main reasons: Mapping difficulties due to inconsistent data structure, and more complex records that needed nuanced approach. OCR errors, as the original scans were low resolution. Missing data information not captured by OCR, or lacking in the original record.

The talk focused mostly on the first of these problems. One of the largest issues was how the 1872 catalogue handled multiple editions of the same title. These needed to be represented in MARC with an individual record for each edition, but in the printed catalogue they are filed under a single uniform title, usually reflecting the earliest edition held by TCD. This in turn appeared in the OCR data as a single record. Joe described the process of separating these out into new records using OpenRefine data-cleaning software, which proved to be the ideal tool for working with such a large and complex dataset. 


Some of the OCR errors could also be identified and cleaned at the batch-edit stage, as they followed predictable patterns. And the data was further enriched at this stage by separating the imprint out into fresh fields for place and date of publication, as well as printer, series, date range, language, and other information that was available in some of the more detailed records. This allowed the creation of more technically precise MARC records, populating fields for country, language and date. 


With the batch edits complete by the end of April, the dataset was shared among colleagues from across TCD Library, who painstakingly compared each line of the data with all 5,121 pages of the printed catalogue. This work went on over the rest of the year, and was finally complete just before Christmas 2020. In the final phase, the proofread data was given its final integrity checks and further augmented by Niamh Harte, the project manager who converted the records into MARC format and loaded them into TCD’s live online catalogue one volume at a time. All the presenters paid tribute to the work of their TCD colleagues Niamh Harte, Barbara McDonald and John Byrne on this project.

 *Special thanks to Joe Nankivell for his help in summarising his work on data-wrangling and OCR for this blogpost. 

 Patricia Moloney is secretary of the LAICMG and works as a cataloguing librarian on the Dónal Ó Súilleabháin Collection in Special Collections, Glucksman Library, University of Limerick.

25 Jun 2021

Controlled vocabularies should no longer be created and used because they are biased.

Blogpost by Alison Kindegran. MLIS student at UCD. Graduating in 2022.

"LCSH" by Travelin' Librarian / CC BY-NC 2.0.

Introduction: 

Controlled vocabularies are organised and arranged word and phrases used to retrieve items via navigation and searches. The purpose of this essay is to form an opinion on the narrative:

“Controlled vocabularies should no longer be created and used because they are biased.”

To form a well-rounded opinion extensive research was conducted. While it was easy to find many articles and journals agreeing that controlled vocabularies are biased and therefore should no longer be used, the argument to continue the use of controlled vocabularies was under represented. There were by far more challenges noted than benefits. However, the writer of this essay did take into consideration all points and did not allow the majority of those in agreement with the statement to sway their opinion from the outset and used self-debate for both sides with supporting articles for each side to form the understanding of both sides and ultimately to form a conclusion.

Benefits of Controlled Vocabularies:
The main areas that are beneficial are:

•    Users with limited knowledge of a topic:
It is believed that even if a user has limited knowledge on a topic they wish to perform a search, the main benefit of controlled vocabularies is that once you have the heading then all classifications or variants will be found.  

•    Vocabulary Deficits:
The use of controlled vocabularies is beneficial for vocabulary deficits; this is where one user may have limited terms on a subject. The benefit is, if they have one term it will still retrieve their search, showing the additional terms available thus adding new terminologies to the user’s vocabulary which would close the gap of deficit in the future and an additional benefit to the user.

•    Topics/Concepts Covered:
Controlled vocabularies guarantee the topic and concepts are covered in the article if the subject heading is listed this assures the topic is covered in the article. Controlled vocabularies can make a search more specific. The hierarchy structure will go from broad to specific as the user goes through the headings where they are narrowed down.

Challenges of Controlled Vocabularies:
The main challenges of controlled vocabularies:

•    Synonymous Concepts:
There are a number of popular used examples when researching the synonymous challenges and the most common examples for this are around soda, pop, soda pop and coke. Soda pop and coke are examples of words that often represent the same idea, or thing. However, those are used differently in different regions and some regional dialects use different terms altogether. The author would not use these words at all. As the author is from Ireland, the words used would generally be fizzy drink, mineral. Also here in Ireland, there is no use of a particular brand or drink type used interchangeably to refer to a number of drinks just the drink noted. For example, coke is used in the above example as it referring to any fizzy which seems to be acceptable in the USA. However, Coca-Cola also known as Coke is to refer to this drink only in Ireland as Coke would be considered completely different to Fanta Orange (alternative brand and drink type) this would not be used interchangeably.

•    Word Form:
Word form is also a challenge. An example would be the word “Online”. The writer would use online all in one and not hyphenated. However, the word online is also acceptable in other formats such as: on line and on-line.

•    Homographs:
Words that look that same but have different meaning, these may or may not be pronounced the same. The pronunciation is not an issue but rather the same spelling of different meanings which would result in the incorrect result. For example, if the user wanted to search bat referring to the animal their result would also include bat referring to the sports equipment. One could argue that the use of qualifiers would remove any issues with this for example: bat (mammal). Similar issues arise with homophones, words that sound the same but are spelled different such as fowl and foul.

The above is not an exhaustive list but an example of the main benefits and challenges with controlled vocabularies. So where is the discussion or argument for bias?

The writer would argue that the bias begins with ignoring that it exists. In the research carried out the challenges that were noted by those who overall would advocate for the use of controlled vocabularies did not address bias at all. This is a concern as not addressing these concerns from those who believe that controlled vocabularies are bias is not helpful in their efforts to portray the benefits and advocate for its use. By and large, those who advocated for the use of controlled vocabularies where libraries, universities or individuals associated with these libraries or universities where controlled vocabularies are used and/or a taught subject. The writer of this essay believes that as it is used and/or taught they do not wish to portray it in any negative light. However, the writer believes that anyone can advocate for something while still addressing any issues it has such as bias.

Ways in which controlled vocabularies are bias:
•    Outdated Terminology:
This is particularly the case in racial categories. There are not reflective of current terms and inclusive of all racial groups.

For those who identify as non-binary, the LCSH term “gender non-conforming people” is an exact match for “non-binary people”. Being gender non-conforming is not the same as being non-binary, although some will identify with both terms. People do not have to be non-binary in order to be gender non-conforming.

•    Too simple in terms and non-representative:
An example of this would be that all First Nations groups are not part of the LCSH or Library of Congress Name Authority File.

•    Non-Proactive Approach:
Most would accept that historic terms for certain groups of people could be described as historical, inaccurate, non-representative and offensive so would expect changes to be made. However, a major issue is the non-proactive approach in addressing these required changes. LCSH as an example have been slow to update the changes. There are a number of resources which would help them make these changes including representing bodies or groups if they were not fully educated on the correct terms. They are freely available.

•    Language Vocabularies:  All languages are not included.

With a number of clear bias found as set out above consideration was given to what the alternative would be if controlled vocabulary was no longer created or used. The result being keywords being used as an alternative. Using keywords, non-controlled vocabulary or natural language as it is also known has a set of advantages and disadvantages of its own which are outside the scope of this essay. In summary, these searches would likely result in a broader search which may include non-related topics.

It is difficult to portray in general terms the impact that these biases have on those they impact. In order to highlight that impact the writer will include an individual focus piece as part of this essay. The piece will focus on one particular individual and their experiences.

Individual Focus: Safiya Umoja Noble

Safiya Umoja Noble is an Associate Professor at UCLA in the Departments of Information Studies and African American Studies. She is also an author of the book Algorithms of Oppression: How Search Engines Reinforce Racism (NYU Press). She has also done a number of Tedx Talks including one entitled “How biased are our algorithms”.

Safiya talks about her mixed-race background and how she had grown up during a time of cultural revolution. The 1970’s was when civil rights movement was highly active and the Black Power Movement was “creating change for how African Americans saw themselves” (Noble, 2014).

Her white mother was aware of a previous study that was done in the 1940’s where black children were given a black doll and a while doll and asked which doll was the best, the prettiest, which one did they like the most? The black children would all pick the white doll. Her mother tried to instil her daughters’ pride in her own black heritage and did not want her daughter to feel negative thoughts and feelings towards her own race.

This experience led Noble to research this further in 2009. One of the first searches she conducted in 2011 was a simple entry into Google “black girls”. The results were of a pornographic nature where they sexualised black girls. In contrast a search for “white girls” returned blonde haired blue eyed general pictures of girls with no implied sexualisation for this general search term in the results. She worked to change this and only months later the algorithm changed and this result in a search for “black girls” is no longer the case.

She is advocating for the pursuit of socially responsible information and technology. She talks about the exploitation of children in the Democratic Republic of Congo where they work in deplorable conditions to source parts required for technology. She wants everyone to consider the part they play in all the above.

In her book, Algorithms of Oppression: How Search Engines Reinforce Racism (NYU Press) she talks about how the public generally views what they find on places like Google to be credible and fair. However, her research found that it was largely misrepresentative. Taking marginalised and oppressed people and making them further marginalised and oppressed with their racist and sexist algorithm bias. She talks about how it is designed to make some heard and others silenced. She finishes the book by seeking alternatives. She argues that there needs to be searches that are for public interest and not driven by marketing and money-making schemes.

The writer of this essay would agree that they also would have considered Google or any other search engine to be credible and fair. While the writer would not consider it the best resource for research, it is the most used way to source information in the modern digital age. The term “Google it” is often used as a response to a query raised by others rather than suggesting an alternative search engine or other resources. It has become an easy and acceptable form of mass use for information. The writer has used it for personal use and to find that the results were most likely skewed in a way that further marginalised people is alarming and seriously problematic. It has truly opened the writers mind to the impact of this on those it affects and the public at large.

Conclusion

There is no doubt that bias exists within controlled vocabularies. The bias has been formed due to historical bias, intentional bias, unintentional bias and unconscious bias. The individual focus piece opened up the problems further. However, it is the writer’s belief that the use of controlled vocabularies remains beneficial. The writer believes that with a number of changes the bias can be removed and the benefits of controlled vocabularies will remain. The end of controlled vocabularies does not address the biases but rather shifts it somewhere else. That is why the recommendations will focus on removing bias and making better more inclusive controlled vocabulary options.  The writer offers a number of recommendations to address and remove forms of bias within the controlled vocabularies. The benefits can be advocated for while making changes to the issues. It does not have to be one or the other.

Recommendations:

Education: While some bias will have been intentional the writer would like to believe that these are historical and that no truly inclusive database owners, creators and users would like to continue with these bias terms. It is fully acceptable that the desire to create change is there but the knowledge on terms may not. This is where the education around the terms will be required. There are a number of international and local groups that can assist with correct terminology who should be sought out and collaborated with to bring about these changes.

Support the change:  Change is a difficult transition for most people. Change is usually met with resistance. In order to support those affected by this bias, first of all, support change. This is an opportunity for those affected by bias to be supported and to support those addressing the change with guidance and education.

Evolving terminology: Understand that this change will be ongoing. A term that was acceptable decades ago can be considered outdated or offensive today. Understand that the same will be true for words used right now. They become obsolete and new terms will inevitably develop over time.  
 
Understanding: Just because a term does not offend or affect one individual, it must be understood that it may not be appropriate for those who are part of a particular community or group. Don’t use outdated terminology and don’t accept it within your controlled vocabulary or indeed anywhere else. Advocate for your family, friends, colleagues and even strangers who are part of these marginalised communities.

Accept Feedback: Provide an option for users to give feedback on existing controlled vocabularies. Those who are affected are best placed to assist with change.

Show Progress: All words which are under review for removal or areas where terms are to be created could be highlighted. The changes will take time and overall will be an ongoing challenge. Let everyone know that the bias is being challenged and addressed from within.

Proactive: This is one of the most important steps. Without being proactive and creating the changes required the argument turns in favour of no longer using controlled vocabularies due to bias.

References:

21 May 2019

Beyond records storage… Institutional repository Digital CSIC as service for open science

Digital CSIC is the institutional repository (IR) of the Spanish National Research Council (CSIC). CSIC’s network of libraries and archives is in charge of the leadership and management of this IR.

A few preliminary notes: the institution and its libraries

CSIC is the largest public, research institution in Spain and the third largest in Europe. Its researchers generate approximately the 20% of all scientific production in the country. Its mission is to foster, coordinate, develop and promote scientific and technological research of a multidisciplinary nature in order to contribute to the advancement of knowledge and economic, social and cultural development, as well as training staff and advise public and private organizations on these matters.

CSIC’s research scope involves the following fields:
  • Biology and biomedicine.
  • Humanities and social science.
  • Natural resources.
  • Agricultural sciences.
  • Physical science and technologies.
  • Materials science and technology.
  • Food science and technology.
  • Chemical science and technology.
There are research centres all over the country that belong to CSIC. In a number of them there is a library and/or an archive.

No few services are managed thanks to a well-conceived network of libraries and archives:
  • A discovery tool that provides access to all information resources (papers, books, digitalised collections, databases, software licenses, etc.) kept, subscribed and managed by CSIC’s libraries.
  • Remote access to those resources, despite not being physically in the institution.
  • Traditional services, such as loan, interlibrary loan, user/library orientation, reading room and reproduction of documents.
  • A digital reference service.
  • An institutional repository in which research outcomes are archived: Digital CSIC. All members of the research community of CSIC can upload metadata-enriched files to it.
  • The Digital.CSIC Direct Archiving Service by which research community can delegate the archiving of its research outcomes to librarians so as to ensure higher-quality metadata and a faster uploading.
  • The service GRANADO aimed at improving the management of libraries space as well as ensuring the conservation of its collections regardless of its format.
  • 100% Digital plan, which is offered by CSIC’s network of libraries and archives to CSIC institutes without libraries. It includes a number of library services.

A new librarianship context: from open access to open science

According to the Open Access (OA) libguide of the University of Pittsburgh library system (2019), Open Access refers to:
  • “A family of copyright licensing policies under which authors and copyright owners make their works publicly available
  • A movement in higher education to increase access to scholarly research and communication, not limiting it solely to subscribers or purchasers of works
  • A response to the current crisis in scholarly communication”.
Although providing free online access to journal articles began many years before the term "open access" was formally coined, computer scientists had been sharing anonymous archives through FTP since the 1970s and physicists had been self-archiving on arxiv since the 1990s (History of open access).

The concept Open Access was not formally established until the 2000s due to these statements: the Budapest Open Access Initiative (2002), the Bethesda Statement on Open Access Publishing (2003), and the Berlin Declaration on Open Access to Knowledge in the Sciences and Humanities (2003).

Two ways to accomplish Open Access statements emerged: green (research outcomes published on IRs) and gold roads (papers published on OA on their respective journals). However, the high costs of article processing charges (ACP) (Khoo, 2019) for pursuing gold road have resulted in that IRs are sometimes the only possible way for OA.

Not many years ago, the scope and sense of openness were widened by The Royal Society (2012) through its thought-provoking book Science as an open enterprise. The transcendence of Open Science (Anglada & Abadal, 2018) have come to the European Union. Indeed, European Commission (2019) has taken it into account on its new policies regarding research across Europe.

The FOSTER Plus (Fostering the practical implementation of Open Science in Horizon 2020 and beyond) project designed this taxonomy that organises all the related concepts:

Open Science Taxonomy. Source: Foster Open Science.
Open Science have brought out several relevant issues on how research is carried out, its outcomes and benefits for society, and the agents involved:
  • The fourth scientific revolution concerns big data, data mining and software.
  • Speaking of outcomes, openness does not only refer to papers published on journals or proceedings, but also the research data regardless of its format, e.g. databases, photographs, presentations, web sites and pages, videos, didactic materials, datasets, software and code. Moreover, research outcomes does not only belong to publishers and/or researchers, but also to society.
  • Data must be FAIR: Findable, Accessible, Interoperable and Reusable (Wilkinson et al, 2016. GO FAIR, 2019).
  • Ethics counts: data ownership, intellectual property rights, research integrity (SPARC Europe, 2019), privacy, security and safety.
  • There is much more need for investing in scientific literacy, science communication and open education than ever.
  • Now, a more variety of partnerships between research agents and society is feasible.
  • Evaluation of science and its metrics must change, as the current cites-based system is not enough to foster open science among scientists.

Digital CSIC as service for open science and researchers community

Looking at this new data-information-and-knowledge environment we will undoubtedly have to face, librarians must wonder how to adapt ourselves, our libraries and profession to address the issue.

Specifically, as for institutional repositories, the following are the actions taking up by the Digital CSIC IR to go beyond any digital library and play a service role for open science and research community.

Digital preservation

It goes without saying that the archiving of research papers on IRs contributes to their digital preservation. If we keep in mind the Levels of Digital Preservation established by the Digital Library Federation (2018), an Open Access Repository in and of itself can be a “tool” to cope with the five functional areas: storage and geographic location, file fixity and data integrity, information security, metadata and file formats.

Digital preservation must be planned. Although IRs can be a great deal of help, they must be tools that are integrated into a well-conceived digital preservation planification.

Digital CSIC (2019) currently offers the following digital-preservation-oriented actions:
  • Backups.
  • Storage of magnetic tapes.
  • Conversion of formats to more secure ones.
  • Periodic checks of the files integrity to prevent their corruption.
  • Monitoring of the technology environment to foresee possible migrations of obsolete formats or software.
  • Metadata for digital preservation.
  • Recommendation for file formats.

The archive of science

Digital CSIC pursue the archiving of all the research outcomes of its institution. As I said before, according to open science view, outcomes involves a wide range of resources: papers published on journals or proceedings, databases, photographs, presentations, web sites and pages, videos, didactic materials, datasets and software. That is precisely the mission of archives: the archiving and preservation of all the records generated as a result of the activity of the institution in which it is integrated and depends on. So, in a sense, an OA IR is the archive of science produced on its institution. In case copyright and intellectual issues do not allow to publish some resources on Open Access, it does not mean that those cannot have an embargo or be in closed access in order to preserve them.

Digital CSIC, which is built upon the software Dspace, has one community per field of knowledge in which CSIC researchers research. I listed those fields in the first epigraph of this post, all of them are accessible via https://digital.csic.es/community-list. There is a sub-community per each research institute devoted to a determined field of knowledge. Then, there are as many collections inside each sub-community as different types of information-or-data resources resulting from the research carried out by that research institute. The principle of provenance is present, thus the archive of science.

FAIR data

Taking FAIR Principles (GO FAIR, 2019) into consideration, I show how the IR Digital CSIC accomplishes them as followed:

Findable

F1. It uses the handle system to assign an URI to each digital object.

F2. The IR publishes intrinsic metadata, librarians ensure the contextual metadata and librarians along with researchers are committed in the description process to ensure rich metadata, such as the measurement devices used, the units of the captured data, the species involved, the genes/proteins/other that are the focus of the study, the physical parameter space of observed or simulated astronomical data sets, questions and concepts linked to longitudinal data, calculation of the properties of materials, or any other details about the experiment.

F3. Digital CSIC does it through dc.identifier.uri.

F4. They are, as Digital CSIC is indexed by the Spanish national aggregator RECOLECTA, OpenAIRE, share.osf.io, core.ac.uk, base-search.net, Google Scholar as well as being registered on re3data.org.

Accesible

A1.1 and A1.2. It uses OAI-PMH.

Interoperable

I1. It supports MARC, Dublin Core, RDF, ORE, MODS, METS and DIDL.

I2 and I3. It does.

Reusable

R1.1. dc.rights and dc.rights.license are used.

R1.2. dc.date.accessioned, dc.date.available and dc.description.provenance are used.

R1.3. It is partially accomplished. Digital CSIC tends to use dc.description as last resource.

Open Peer Review Module

Digital CSIC have integrated the first Open Peer Review Module (OPRM) for open access repositories that allows to make reviews and comments on already archived digital objects.

Open Peer Review Module. Source: Digital CSIC.
This tool is especially useful for receiving feedback that is bound to facilitate the improvement of research outcomes.

Impact, (alt)metrics and statistics of research

How can Digital CSIC measure the impact of its archived files?

First all of all, in the web page about general statistics, we can see them in terms of:
  • Number of research institutes per community (field of knowledge).
  • Number of items per community.
  • Number of items per research institute (top 20).
  • Number of research institutes by geographical distribution.
  • Number of items by geographical distribution.
  • Types of items (articles, conference paper, etc.).
  • Types of archived items per research institute.
  • Open Access: the percentage of OA items by type, year of deposit and community.
We can delimit them by date (year and/or month).

It also shows the number of archived objects by communities (field of knowledge), sub-communities (research centres), collections (types of documents per research centres) and authors. By research groups and research projects are being tested.
Source: Digital CSIC.
Source: Digital CSIC.
Source: Digital CSIC.
It is also possible to view statistics of any of the communities in terms of count of views, sub-community view count, collection view count, item count view and item download count. Besides, we can examine those by region/country/city in a geo map (thanks to Google Maps API) and along time.

As for single archived digital objects, we can see its views and downloads by region and along time. There is also information about altmetrics.
Source: Digital CSIC.
Source: Digital CSIC.

Web pages for researchers

Digital CSIC provides the possibility to generate web pages for researcher. They consist of:
  • An URI.
  • A personal statement with a nice picture.
  • Integration of profiles of other networks and IDs.
  • Statistics.
  • Concentration and organization of all their scientific production.

Automated archiving

Digital CSIC and some publishers came to an agreement so that they are archiving all the journal papers on this IR as long as their filiation contains CSIC.

Consultancy

Digital CSIC and its librarians give advice to the research community regarding a number of topics:
  • Technical use of Digital CSIC and guidance in the metadata description according to its policies.
  • Open Access mandates.
  • Profiles for researchers and research groups.
  • Intellectual property, copyright and licensing.
  • FAIR data.
  • Data management plans.

Final thoughts

The increasingly consciousness regarding the importance of open access, which we can even measure (Dubinsky, 2019), is undoubtedly good news. However, promoting open access is not enough. Institutional repositories seem need to evolve from merely digital libraries for storage of items. Librarians of research institutions must change their mindset to a service-oriented one. Service, here, has to do with open science and the researchers community. I have presented current developments of Digital CSIC, I hope it would be inspiring for other librarians. 

References

Anglada, L.; Abadal, E. (2018). ¿Qué es la ciencia abierta?. Anuario ThinkEPI, 12, 292-298.
doi: 10.3145/thinkepi.2018.43

Berlin Declaration on Open Access to Knowledge in the Sciences and Humanities (2003). Retrieved from https://openaccess.mpg.de/67605/berlin_declaration_engl.pdf

Bethesda Statement on Open Access Publishing (2003). Retrieved from https://legacy.earlham.edu/~peters/fos/bethesda.htm

Budapest Open Access Initiative (2002). Retrieved from https://www.budapestopenaccessinitiative.org/read

Digital CSIC (2019). Digital preservation policy. Retrieved from http://digital.csic.es/dc/politicas/#politica8

Digital Library Federation (2018). Levels of Digital Preservation. Retrieved from https://ndsa.org/activities/levels-of-digital-preservation/

Dubinsky, E. (2019). Does open access make cents? Return on investment in the institutional repository. College & Research Libraries News, 80(5). doi: 10.5860/crln.80.5.281.

European Commission (2019). Open Science. Retrieved from https://ec.europa.eu/research/openscience/index.cfm

Foster Open Science. Open Science Taxonomy. Retrieved from https://www.fosteropenscience.eu/themes/fosterstrap/images/taxonomies/os_taxonomy.png

GO FAIR. FAIR Principles. Retrieved from https://www.go-fair.org/fair-principles/

History of open access. Retrieved from https://en.wikipedia.org/wiki/History_of_open_access

Khoo, S. Y.-S. (2019). Article Processing Charge Hyperinflation and Price Insensitivity: An Open Access Sequel to the Serials Crisis. LIBER Quarterly, 29(1), 1–18. doi: 10.18352/lq.10280

SPARC Europe (2019). Research Integrity through Open Science and FAIR Data. Retrieved from https://sparceurope.org/wp-content/uploads/dlm_uploads/2019/03/SPARCEurope_ResearchIntegrityBrief.pdf

The Royal Society (2012). Science as an open enterprise. Retrieved from https://royalsociety.org/~/media/royal_society_content/policy/projects/sape/2012-06-20-saoe.pdf

University of Pittsburgh library system (2019). Open Access @ Pitt: All About OA. Retrieved from https://pitt.libguides.com/openaccess

Wilkinson, M. D. et al (2016). The FAIR Guiding Principles for scientific data management and stewardship. Scientific data, 3:160018. doi: 10.1038/sdata.2016.18.

3 May 2016

Open access to research data : a pilot project in Sweden

Guest post by Ulf-Göran Nilsson and Stefan Carlstein librarians at Jönköping University Library.


Jönköping University Library and two research groups at Jönköping University, CHILD (Children, Health, Intervention, Learning and Development) and Computer Science and Informatics, are currently participating in a national pilot study of open research data.


In recent decades, rapid technological development has brought new opportunities and possibilities to collect and make information available. In doing so, it has established new ways of conducting research. A wide range of funders, the EU and Vetenskapsrådet, the Swedish Research Council, considers that it is of great importance economically to make data available in general and research in particular.


To archive research data was previously a responsibility of each institution in Sweden. But according to the Swedish Research Council's proposed guidelines for the coming years “Proposal for National Guidelines for Open Access to Scientific Information", will universities now have to increasingly take responsibility both to preserve research data on long   term and to make it available where it is possible.


Jönköping University is one of five universities that are part of a pilot study in arrangement of Swedish National Data Service, SND. The Swedish Research Council has appointed SND as a national resource for the coordination of existing and newly established databases within the social sciences, humanities and health sciences. SND offers support to Swedish research by facilitating researchers access to data within and outside of Sweden as well as offer support for research during the whole research process. SND presents Swedish research outside of Sweden. The project started in April 2016 and is expected to continue throughout the year.


The pilot project aims to develop a model for dealing with the challenges that Swedish universities and the Swedish National Data Service will meet in connection with the increasing demands for the long time preservation of digital research and the availability to the research data. One important part of the project will be transfer of competence from the Swedish National Data Service to a Research Support Unit, RSU, at the university to manage research data, metadata enrichment, ensuring the format for long term storage and making the data available when it is possible. RSU is a concept from SND and it is defined in this project to be manned by the university library and the archive with support from the IT department and the university lawyer. The archive is part of the same department as the university library at Jönköping University and it is an advantage with the close connection between library and archive. The Research Support Unit will in the next step train and support researchers in two ways. The first training will be conducted together with SND with the two selected research groups and in the final training for a broader group of researchers conducted by the RSU on its own.


The pilot project is focusing on the most important parts in research data management: Working with various types of research data and file formats, development of forms for metadata management, metadata profiles and metadata standards, valuation of the metadata required for data to be useful for secondary research, design of data management plans, analyzing tools and methods to document data, development of procedures and responsibilities for the archiving of research documents and finally reviewing the basic legal aspects that affect the handling of research data.




1 Jul 2015

Developing Curation Ready Projects Workshop

Last week I attended a very interesting one-day workshop that covered the topical theme of digital curation and adjacent considerations. What I actually liked most about this session was the presence of a diverse audience: the room was not just filled with curious librarians, but also included artists concerned with how to best preserve their creative outputs in the digital realm (one facet of the digital curation process), as well as researchers interested in the holistic nature of curation within the digital humanities context.

The day started off with Debra F. Laefer who spoke about the higher-order rationale of open access within the digital realm (presentation entitled The euros and cents of open access data). This may include the need to fulfil institutional and research sponsorship requirements, but also public expectations, which essentially relate to the establishment of equitable access to taxpayer sponsored research outputs. Laefer also noted the immediate researcher benefits of openly sharing research data and adjoining journal publications. They include, for example, ownership validation, attraction of potential collaborators for joint publications, provision of new insights and their timely usage in a variety of academic and non-academic contexts.

Kalpana Shankar provided a brief introduction to digital curation by demarcating its boundaries (what it is and isn't). In a nutshell, digital curation represents a set of activities including the selection, preservation, maintenance, collection and archiving of digital assets. It covers the entire lifecycle of the digital item and not just an isolated activity, such as digitisation for example. Kalpana pointed out the various threats to digital material (from software rot to sociopolitical issues) and offered cogent motivations for carrying out digtial curation efforts. Pertinent examples were provided to this end and the digital curation lifecyle model was discussed. The OAIS framework was also introduced for the purpose of illustrating the considerations behind the design of archival systems.

Amber Cushing focused in her presentation on the specific activities involved in the creation of a digital preservation plan: 1) Identify your interest or activity, 2) Describe the interest or activity, 3) Select a subset using appraisal, 4) Select and plan a preservation strategy. Appraisal refers to the tricky process of determining material significance and enduring value. Depending on context and scope, this can (is) cognitively taxing: what represents value to me might not represent value to you. This task is especially difficult when arguing for project funding support...

Jenny O'Neill reiterated the importance of digital file preservation and then drilled into the conceptual and practical considerations around metadata that attach to digital objects. Rights management was discussed (Creative Commons) as well as various DRI guides.

From a librarian (my) perspective, digital curation is of critical importance. Considering Chris Alen Sula's conceptual model very much highlights the intricacies involved here.

 
 
Finally, I'd like to point to Charles W. Bailey's recently published Research Data Curation Bibliography, which might be of interest to some of you. This selective bibliography includes over 350 English-language articles, books, and technical reports that are useful in understanding the curation of digital research data in academic and other research institutions.

14 May 2015

Open access & research data management: Horizon 2020 and beyond (UCC, 14th-15th April 2015)

Guest post by Maura Flynn, Breeda Herlihy and Ronan Madden, all UCC Library

Speakers and Organisers day 1. Picture courtesy of Richard Bradfield

This two-day training event was held in UCC in April, and was jointly hosted by UCC Library, UCC Research Support Services, Teagasc and the Repository Network of Ireland (RNI). The first of its kind to be held in Ireland, the event introduced attendees to the concepts of open research and research data management within the context of Horizon 2020. With speakers from the U.K. and Ireland sharing best practice, the event was an invaluable learning experience, and timely in the context of Horizon 2020’s Open Data Pilot.

To stage the event, the project team was successful in securing funding from the FP7 funded FOSTER project. FOSTER (Facilitate Open Science Training for European Research) is a two-year EU funded project which aims to promote & ‘foster’ open science in research, and to optimise research visibility and impact and the adoption of EU open access policies.

Research data management (RDM) generally refers to the processes of organising, structuring, storing, and preserving the data used or generated during a research project. Numerous factors are now influencing the drive for open data, but chief among them is the influence of funders seeking transparency and a demonstration of the wider impact of the research they are financing. In Horizon 2020 a limited pilot on open access to research data is being implemented, with participating projects required to develop a Data Management Plan (DMP). There is an expectation that this trend of research funding programmes requiring data management plans is set to continue, as has been happening in the U.K.

In addition to compliance, RDM benefits researchers and institutions through the potential for re-use of data, and the opportunity to demonstrate research excellence. Many institutions are taking a lead by establishing research data policies and seeking to coordinate cross-campus approaches to gathering and maintaining data. This often involves research support services, IT teams, libraries and researchers working together. However, RDM has been described by Cox et al. (2014) as a ‘wicked problem’, complex and difficult to define, requiring solutions that are flexible and pragmatic. RDM is still in the early stages at many Irish institutions, and this event offered an opportunity to learn from others and to draw on the expertise of those who are further down the road. It was a chance also for making connections within and across institutions in Ireland and the U.K.

David O'Connell opening Day 1. Picture courtesy of Richard Bradfield

Day 1: ‘Open research in H2020: how to increase your chances of success’

The first day was targeted to researchers and small and medium enterprises interested in developing Horizon 2020 proposals. David O’Connell, Director of Research Support Services, UCC, provided the opening remarks, mentioning that as former chief editor of ‘Nature Reviews Microbiology’ he has had a long interest in open access publishing, and a strong interest now in the application of open access to research data.

The project team were lucky to have support and guidance from Martin Donnelly from the Digital Curation Centre (DCC). The DCC is a UK-based world-leading centre of expertise in digital information curation, providing expert advice to higher education. Martin played an invaluable advisory role in the run-up to the event. Although he was unable to attend due to unavoidable reasons, he provided four recorded presentations for the event at short notice. Day 1 began with his first presentation: an overview of Open Science and Open Data in Horizon 2020. He started by providing a background to open access and RDM, looking back at open access in FP7, before looking at open science in Horizon 2020, and the specifics of the open data pilot.

Joe Doyle, Intellectual Property Manager, Enterprise Ireland, provided a background to how intellectual property relates to both innovation and collaboration, describing IP as a bridge between the creative and the commercial. Open access can generate greater collaboration, but it is important to acknowledge that what is free to access is not necessarily free to use without limits. Open access and patents can work hand-in-hand, as patents are about disclosing data. While they can’t be copied, much can be learned from previous innovations.

Jonathan Tedds, Senior Research Fellow, Department of Health Sciences, University of Leicester, spoke of RDM from the perspective of researchers, giving examples of projects he has been involved in, and issues encountered. He originally became convinced of the benefits of data sharing through his work as an astronomer, when he would ‘stitch together’ data he had generated for re-use. He cited the Royal Society (2012) report ‘Science as Open Enterprise’ suggesting that publishing articles without making data available is a form of scientific malpractice, and he noted that the number of papers based upon reuse of archived observations now exceeds those based on the use described in the original proposal. However researchers in many fields, especially those involved in smaller projects, need help to comply with funder requirements. He emphasised the iterative nature of research and data management planning, and the challenge of sustaining research software, not just the underlying data. The HALOGEN project was a good example of combining different kinds of data from different fields, achieved by creating a central scalable database infrastructure to support the project. The BRISSkit project involved developing software to link applications to create a data warehouse of anonymised (consented) patient data. It brings bed-side patient data to university researchers to be used for new biomedical research.

Group shot. Picture courtesy of Richard Bradfield


In the afternoon, Martin Donnelly’s second presentation focussed on data management plans (DMPs), providing an overview of these and their benefits. He went on to outline various data related policies and requirements in Europe and elsewhere, plus the supports and resources that are available to those writing DMPs, including those provided by the DCC. He demonstrated the DMPonline tool, which was created by the DCC, and can be customised by institutions. It can be used by researchers at the point of application and throughout the research project, and can be used for sharing and co-writing plans.

Brian Clayton, Research Cloud Service Manager, UCC, spoke of RDM as a work-in-progress at UCC. He described current UCC research cloud paid services which include data storage and compute services. The service has expanded to offer elements of data management, and a draft RDM policy is currently awaiting University committee approval. The aspiration is that RDM services can be provided at zero-cost to the researcher. Many outstanding issues will need to be explored, particularly in regard to data sharing, metadata, and who will carry out the various roles within the University.

Peter Mooney, Environmental Research Scientist, Environmental Protection Agency, looked back at over a decade of RDM at the EPA. As far back as 2004 the EPA made a commitment to researchers they were funding, that they would preserve data free of charge, and be responsible for long term management and infrastructure. The SAFER data archive was launched in 2006, linking data to papers and reports. Collaboration with researchers has been key to its success and development. Data reporting is now an essential element of the reporting process on EPA funded projects. He outlined some of the lessons learned, and suggested that open data is often misunderstood by researchers, and metadata is often a mystery, or seen as a burden. Modelling data correctly at the start of a project increases usability, and researchers would benefit from understanding the basics of relational databases. As an example, he cited over-reliance on Excel rather than using databases. He also cautioned against long embargo periods which only serve to make data lose relevance.

The final speaker of the day was Evelyn Flanagan, Data Manager at the UCC Clinical Research Facility, who spoke of her role as a data manager in clinical trials. She discussed how core principles of data management are a fundamental element of good clinical practice (GCP), before providing a thorough description of the ‘data sequence’ from protocol design right through to the report writing stage. She examined each stage of the process, including database design for case report forms (CRF), the importance of good metadata, data collection and data entry procedures. Like the previous speakers she stressed the value of DMPs at the early stages of a project, and how they underpin good practice at each stage of the data sequence.


John Fitzgerald opening Day 2. Picture courtesy of Richard Bradfield
Day 2: ‘Research data management – institutional needs, targets and training’

The second day of the event was aimed at institutional support staff who can provide support to researchers engaging with RDM. Many speakers came from the UK where the policies of the UK research funders (RCUK) require researchers to engage with RDM. In Ireland, the open data pilot in Horizon 2020 is the first signal that research performing institutions here will have to address RDM in the coming years.

John FitzGerald, University Librarian and Head of Information Services, UCC, gave the opening remarks mentioning how RDM will ‘challenge us as professionals with broadly curatorial problems’ as we seek to ‘manage the ecosystem in which data exists’. The first invited speaker of the day, Martin Donnelly of the DCC, provided a clear overview into RDM for support staff. Although unable to attend the event in person, Martin provided a recorded presentation which was very well received by those present.

Stuart McDonald, Research Data Management Service Co-ordinator, University of Edinburgh, spoke about their comprehensive approach to RDM services which began early in 2008 with a JISC funded pilot project. There were some audible gasps when he outlined the resourcing and staffing of the RDM programme at Edinburgh where £1.2 million has been allocated via internal funding. Aside from the resourcing, it was also illuminating to see how Edinburgh approaches data management before, during and after research. They are now investigating how to ensure that systems used for data management do not duplicate effort required by researchers which they will undoubtedly be happy to hear about.

David McElroy, Research Services Librarian at the University of East London, demonstrated how they have used Eprints, their existing institutional repository software, to create a new data repository, Data.uel. Publications archived in their open access publications repository are then linked to underlying data archived in their data repository. This of course ensures traceability and reproducibility of research. It was really useful to see the repository development path taken from decision making and planning to functional and metadata specifications and right through to mock ups and branding.

The third speaker, Jonathan Greer highlighted how Queens University Belfast is taking an ‘incremental approach’ to RDM services as they seek to align the plans and policies of the institution with the practice of their researchers. He offered some consolation to those uninitiated in RDM services by relaying how challenging it can be to roll out a service in such a complex area.

In the afternoon, Gareth Cole, Research Data Manager at Loughborough University and formerly of the University of Exeter, outlined how both university libraries approached the delivery of training and support. This was very useful as it became clear throughout the day that there is no one size fits all approach to RDM services.

Julia Barrett, Research Services Manager, UCD Library, summarised how she has shaped their research services to facilitate effective data management and sharing in UCD. It was encouraging to see the potential for a range of services which the library can offer and Julia has categorised these into ‘Discover’; ‘Create / Analyse’; ‘Manage’ and ‘Disseminate / Publish’ services.

Louise Farragher, Information Specialist, Health Research Board, introduced the PASTEUR4OA project which seeks to align open access policies across Europe. While an earlier question from the audience queried the effectiveness of lots of policy, Louise was quick to reinforce the message that policy is a good starting point for open access adoption.

Finally Dermot Frost, Research IT services at Trinity College Dublin, gave an engaging account of his experiences of developing the technical infrastructure for the Digital Repository of Ireland (DRI). The DRI is a ‘green-field repository’, to be launched publicly in June 2015 at DPASSH and is Ireland’s trusted repository for humanities and social sciences data. The DRI has had a large inter-disciplinary project team and Dermot stressed that while the language barrier (tech vs. non-tech) can be challenging, it was very useful for the exchange of ideas to have different people on board.

Q&A featuring Day 2 speakers. Picture courtesy of Richard Bradfield

Overall take home messages

1. Challenging: developing RDM services can be challenging due to the complexity and variety of research data. However, it is possible to learn from established services at other institutions. All speakers were very open to sharing their own experiences, tools and resources for late adopters of RDM services. All highlighted the well-established services available in the UK e.g. Digital Curation Centre as well as various online tools and resources which can be reused.

2. Planning: it is essential to plan out a roadmap after first establishing an understanding of the needs of the stakeholders. Stuart McDonald discussed the Data Audit Framework used at Edinburgh to identify research data assets and their management before developing an RDM policy and service. Dermot Frost mentioned that a ‘repository needs data to justify its existence’ and so the DRI has a stakeholder advisory group to ensure that depositors were involved from the planning stages.

3. Cross campus collaboration required: due to the complexity of RDM, the different types of stakeholders involved and emerging funder requirements, coordination across the institution is essential for an effective approach to service development.

4. Planning at the research project level: the importance of DMPs in the early stages of projects was emphasised by all Day 1 speakers. They ensure good data management practice at each stage of the research process.

References

Cox, A.M., Pinfield. S., & Smith, J. (2014). Moving a brick building: UK libraries coping with research data management as a ‘wicked’ problem. Journal of Librarianship and Information Science, 46 (4), 299-316. doi:10.1177/0961000614533717

Royal Society. (2012). Science as Open Enterprise. Retrieved from https://royalsociety.org/policy/projects/science-public-enterprise/Report/

8 May 2015

Digital Preservation: Not Just Clouds & Unicorns (ANLTC – NLI 29th April – 1st May 2015 – Report)

Guest post by Elaine Harrington, Special Collections Librarian, UCC Library

I had previously attended a one-day seminar run by the DPC on Getting Started in Digital Preservation. This three-day course run by the DPTP is an intermediate course for practitioners of digital preservation. Over the course Ed Pinsent, a digital archivist and Steph Taylor, a senior consultant, both with University of London Computer Centre (ULCC) showed us tools, methods and strategies for engaging with digital preservation. We viewed practical examples, examined case studies and challenging and complex objects, and participated in group exercises to better understand what digital preservation is.

Over the three days there were moments when I thought I was in a different universe where acronyms ruled (AIP, SIP, DIP,  JHOVE, PLATO, SCAPE, SCOUT, METS, MODS, TDR) or on a Star Wars’ set (constant references to DROID) or looking at antique cars (parallel situation of finding parts to replace wear and tear in cars or older technologies). By the end of third day I was beginning to return to Earth.

The course was broken into modules each of which lasted approximately 45 minutes. Although the course was intensive there was plenty of opportunity to ask questions and Ed and Steph included plenty of examples. At certain points for example when we were discussing ‘migration’ in methods of digital preservation we noted that ‘migration’ would also feature in file formats and as part of a ‘Migration Strategy Exercise.’

Due to sheer volume of concepts and information covered over the three days it is impossible to write about all the modules.

What is Digital Preservation?
According to the National Archives at Kew a digital preservation policy is the mandate for an archive to support the preservation of digital records through a structured and managed digital preservation strategy.” In practical terms the following are needed for digital preservation:
a database to manage preservation and store metadata
tools to perform ingest functions
a place to store digital objects
an access or delivery platform
rules, workflows, policies
an IT infrastructure
people and skills

OAIS Model
Ed and Steph used the OAIS Model and its terminology throughout the course to illuminate the digital preservation process.

Courtesy of University of London Computer Centre



Day 1
On the first day we examined what the OAIS Model is and some of the implications in using it. This was useful as it would be used in some of the group exercises over the next three days and we would have the appropriate terminology to use. This section was followed by modules on methods of digital preservation and exercise; significant properties and the Performance Model; file format: their structure and treatment; and metadata for practitioners. Significant properties varied depending on the file type: 16 significant properties for moving images compared to 6 for audio.

It was clear from the exercise on digital preservation methods that while we understood what was being said to us it was another matter entirely to be given a method and to discuss the pros and cons of that method. Approaches included: migration, emulation and technology preservation. The group I was in was given the bit-level only approach which focuses on maintaining only the 1s and 0s of code.

Courtesy of Elaine Harrington


It was a little bit worrying when Ed said that someone (not on the course!) thought a way to preserve technology was to dip a laptop in Perspex and then chip it off in 20 years! If the computer specs were known and 3D printers still exist in 20 years’ time perhaps it would be possible to 3D print any parts that would be needed to fix a physical piece of technology.

Real world examples were used to explain each module. For example DIOSCURI was used to show how emulation works. The National Library & Archives of the Netherlands use DIOSCURI to run old operating systems such as DOS and WordPerfect 5.1.

Courtesy of University of London Computer Centre


Ed and Steph also mentioned Atari systems and Pac-Man. The Centre for Computer History in Cambridge was established to tell the story of the Information Age through exploring the historical, social and cultural impact of developments in personal computing.

Courtesy of Centre for Cambridge History of Computing.


Day 2
On the second day we covered XML for digital preservation; tools for ingest; how to do migration including an exercise; METS; PREMIS and an exercise; making a business case and an exercise; assessment, audit and TDR.

XML is Extensible Markup Language. Like HTML, XML uses tags but whereas HTML describes presentation XML describes content. There may be:
XML Schema which has specification for elements and tags you will use. In a digital preservation plan the schema being used must be declared.
XML Stylesheet which displays the underlying XML and renders text in useful way for readers
XML Document which is the document you are authoring and which describes the object.

XML is a preservable text format in that it is open, documented and is not tied to vendor or platform. It is both good for storing and conveying metadata. There are different types of metadata: Descriptive, technical, rights, structural and preservation, and XML can be used to describe them all. The Library of Congress uses XML to represent their metadata records in MARC, MODS and METS. XML can enclose a digital object and be used to build and AIP (see OAIS Model). XML allows for interoperability.

Courtesy of University of London Computer Centre


XML & Migration
It is not enough simply to have digital preservation for the file but also for the metadata. The metadata may be stored separately to the object, within a database or metadata is embedded in the files requiring preservation. Metadata can be used for the source file format and when migrating to the new target file format for example word moving to pdf. Migration exercises are like Fight Club: There will always be losses. We have to decide when migrating what could be lost, what is an acceptable loss, what should not be lost such as significant properties and what are the choice that need to be made so that only acceptable losses happen. Ed and Steph suggested doing very detailed use cases before migration.

Day 3
On the third day we looked at metadata exercise; email preservation; social media: communicating with the user community; social media: user community and engagement; understanding legal issues for preservation and access; preservation of databases; and managed storage.

Metadata Exercise
On paper we were shown a painting and its museum cataloguing record. The painting had been digitised and metadata was present. There were gaps in the metadata which had to be identified and what preservation data was also required. This exercise highlighted that no matter the source of the metadata some metadata will not be present.

Courtesy of Elaine Harrington


Cloud Storage
Cloud storage providers should meet ISO standards and should care about auditing standards. Discussion during an exercise showed that institutions who have cloud storage should limit the holdings to within the EU at the very least. If material is held on the cloud and moved to American then it is subject to different copyright laws, different data protection. Copyright law has not yet caught up with digital content. Indeed if a project for digital preservation has EU funding then the storage, cloud or otherwise, may need to remain in the EU. Cloud companies don’t mention how long the objects will be stored for and considering how fast technology changes (who remembers VHS or Betamax?) will the objects require a new digital preservation storage facility in a very short span of time? Of equal concern was cost: it may require little money to insert an object into cloud storage but it could take a long time and much more money to extract an object from cloud storage. If an object is requested will it pass through multiple countries’ before reaching its destination? This may happen as cloud providers move data on a regular basis.  We were advised to always read the fine print!

Conclusion
There is much more to digital preservation than placing objects in cloud storage and all the processes and details are real and not imaginary like unicorns. A good deal of discussion is required no matter which method of digital preservation is chosen, no matter which method of storage for digital preservation is chosen and no matter which tools are used during the process. It was clear that we should all be engaging in digital preservation and that we should be engaging right now.

Thanks to Ed Pinsent and Steph Taylor who shared their experiences and expertise so freely. The slides are available through Creative Commons and UCLC. Thanks also to ANLTC and NLI for organising and hosting the event.