Showing posts with label digital preservation. Show all posts
Showing posts with label digital preservation. Show all posts

Apr 15, 2014

Survey of digital curation and curators

I am conducting a survey of digital curation and digital curators.

If you are involved in taking care of digital materials of any type, form and purpose and are interested in the advancement of digital curation as a professional field, feel free to take the survey and share it with colleagues. The survey takes about 20 min to complete and can be found at http://bit.ly/1osgTQ7

Jun 13, 2013

Sustainable software

An interesting post from the LOC Digital Preservation blog about digitization of historical documents (Profiles in Science project) and choices that need to be made to make such a project working over time. Oddly enough, I used Profiles in Science quite a lot while working on my dissertation because I needed materials by and about Arthur Kornberg, a biochemist who was said to have created "life in a test tube."

To summarize, any system grows over time and some components get upgraded or replaced, while new ones are developed. As the technological landscape keeps changing, we want to keep up with it and make sure that digital collections are available and usable in ways that correspond to current views on access and use. In other words, you can't just "build it and leave it alone".

Forces that threaten stability and sustainability of software and data:

  • Software bugs
  • Loss of needed functions
  • Hardware or operating system incompatibility
  • New policy requirements
  • Security flaws
  • Product support / licencing cost
  • Loss of backward compatibility
  • Product abandonment

Factors that encourage sustainability of the system include:

  • ability to view and modify the code
  • wider use and testing
  • active development and support
  • standards awareness rather than focus on proprietary formats
  • non-restrictive licencing
  • ability to import and export data and code
  • compatibility with multiple platforms
  • backward compatibility
  • minimal customization

I'd probably add "community support for the project" to the list of factors that encourage sustainability. Once the funding and interest from initial researchers and developers are gone, the community of users can help maintain and upgrade the system. Product abandonment is a serious threat for many digital projects. Just yesterday I saw a site with good scholarly content that had a link to gopher resources...

Apr 25, 2013

Strategy for Civil Earth Observations - Data Management for Societal Benefit

The US National Science and Technology Council recently released a National Strategy for Civil Earth Observations. The goal of this strategy is to provide a framework for developing a more detailed plan that would enable "stable, continuous, and coordinated global Earth observation capabilities for the benefit of society."

The strategy establishes a way to evaluate Earth-observing systems and their information products around 12 societal benefit areas: agriculture and forestry, biodiversity, climate, disasters, ecosystems (terrestrial and freshwater), energy and mineral resources, human health, ocean and coastal resources, space weather, transportation, water resources, weather, and reference measurements. The production and dissemination of information products should be based on the following principles:

  • Full and open access
  • Timeliness
  • Non-discrimination
  • Minimum cost
  • Preservation
  • Information quality
  • Ease of use
Data management in federal agencies that are responsible for earth science data is described based on the three components of the data life cycle: planning and production, data management, and usage. The latter two components are the main focus of the data management strategy. The suggested activities for those are:

  • Data management
    • Data collection and processing - initial steps to store data and create usable data records.
    • Quality control - follow the principles of the “Quality Assurance Framework for Earth Observation” (QA4EO)
    • Documentation - basic information about the sensor systems, location and time available at the moment of data collection, etc.
    • Dissemination - data should be offered in formats that are known to work with a broad range of scientific or decision-support tools. Common vocabularies, semantics, and data models should be employed.
    • Cataloging - establishing formal standards-based catalog services, building thematic or agency-specific portals, enabling commercial search engines to index data holdings, and implementing emerging techniques such as feeds, self-advertising data, and casting.
    • Preservation and stewardship - guarantee the authenticity and quality of digital holdings over time.
    • Usage tracking - measuring whether the data are actually being used; to enable better usage tracking, data should be made available through application programming interfaces (APIs).
    • Final disposition - not all data and derived products must be archived, derived products that most users have access to may adequately replace raw data and processing algorithms.
  • Usage activities
    • Discovery - enabled by dissemination, cataloging and documentation activities.
    • Analysis - includes quick evaluaionts to assess the usefulness of a data set and an actual scientific analysis.
    • Product generation - creating new products by averaging, combining, differencing, interpolating, or assimilating data.
    • User feedback - mechanisms to provide feedback to improve usability and resolve data-related issues.
    • Citation - different data products, e.g., classifications, model runs, data subsets, etc., need to be citable.
    • Tagging - identify a data set as relevant to some event, phenomenon, purpose, program, or agency without needing to modify the original metadata.
    • Gap analysis - the determination by users that more data are needed, which influences the requirements-gathering for new data life cycles.

Each activity raises a lot of questions and challenges. The activities of cataloging, usage tracking, final disposition, tagging and gap analysis are particularly interesting. They raise questions that are rarely addressed in the data management literature. Does anybody use data that are being shared? Do all the data need to be preserved? How can we avoid duplicates and unnecessary modifications of metadata if data are being re-used? To what extent do we need to serve immediate user interests versus the future possibilities for research?

Nov 16, 2012

Workflows for Digital Preservation and Curation

Notes from today workshop on workflows for digital preservation and curation.

Curation means adding value to the data during its life cycle, i.e., making sure that has meaning and context and can be re-used outside the creator's environment. Curation infrastructure includes repositories, access procedures, policies, processes and institutional support. The unit of curation and preservation is a file. It's important to maintain files integrity by ensuring their fixity, duplicate storage and format validation.
To preserve files, they need to be in proper formats, i.e., durable (transparent, documented, used widely, renderable) and supported with standards (syntactic and semantic). Syntactic standards don't have context (e.g., in CSV we don't know how columns were created and what they mean). Semantic standards are better.

In preservation a good practice is to use a master file for preservation (highest quality and fidelity) and derivative files for active use and delivery. For example, high-resolution TIFF for preservation and lossy JPEG for viewing.

To implement curation and preservation and practices mentioned above, the following activities are often part of the workflow: ongoing verification (file integrity and object integrity), metadata management, management of obsolescence (hardware, software, formats, documentation).
Workflows systems simplify repetitive and mundane activities, facilitate best practices and coordination and outreach. Various systems that support scientific, research, and software development workflows exist, e.g., Kepler , Triana, Taverna , Ptolemy II and BPEL.

Trident is an open source software package from Microsoft with many components and functions. Data to Insight Center has been working on developing components to support data ingest and curation activities.
The following components has been developed so far: fixity (MD5 checksum), data integrity (JHOVE for format verification and validation), metadata creation (MIX and METS data generator and validator), format normalization and generation (PPT/DOCX to PDF, XLSX to CSV, TIFF to JPEG and some more), persistent identification (DOI generator), repository integration (ingest to Dspace via Sword, DOI
generator).

workflow
Example of Single Object Workflow

Digital Curation Activities in Trident
It's been great to play with these components and see how repetitive and boring tasks can be automated. The biggest question, of course, is whether the tool is ready for wider adoption. It's relatively easy to work with existing components, but it requires coding experience to modify them and create new ones.

Sep 6, 2012

Digital preservation issues

Notes from Digital preservation, archival science and methodological foundations for digital libraries (S. Ross, 2012, New Review of Information Networking, 17:1, 43-68, doi).

At the beginning the article makes an important observation - there is more to preserving digital objects than saving the content. Approaches to preservation should also include a) retaining the environment and context of creation and use and b) reproducing the experience of use.

The middle of the article is of lesser interest. Many points, such as a lack of systematic practices, policies or research strategies in preservation, can be skipped.

Suggestions for research agenda in this paper come primarily from the DigitalPreservationEurope project (DPE). There are nine important areas of work in digital preservation:

  1. Restoration - restoring damaged digital objects, including content, context and experience and verifying their completeness.
  2. Conservation - saving digital objects before they are damaged and making sure they cannot be damaged or destroyed in the future.
  3. Collection management - making decisions about what goes in and out, etc.
  4. Risk management - determining and quantifying uncertainties and minimizing various threats.
  5. Interpretability and functionality - making sure digital objects remain meaningful, authentic, and usable.
  6. Cohesion and interoperability - maintaining connections and transitions across systems, time, and repositories.
  7. Automation - developing tools for handling big quantities of information.
  8. Preserving the context - retaining information about how the object was created and used.
  9. Storage - developing infrastructure for storing digital objects.

The article concludes with the statement that there is an urgent need for a theory of digital preservation and curation. To me it seems that we have enough theories to rely on. Once structures (technological and social) that support digital preservation become adopted and used, we can start observing existing practices and then decide whether we need a new theory. Otherwise, there is a danger of coming up with something trivial and calling it a new theory.