Showing posts with label data curation. Show all posts
Showing posts with label data curation. Show all posts

Jun 6, 2016

The Net Data directory

The Berkman Center for Internet & Society announced the launch of the Net Data Directory - a free, publicly available database of data about the Internet that covers topics such as cyber-security, civil and human rights, social media and many more. The directory currently contains about 150 data source records and includes many types of sources, including website rankings, opinion surveys, maps of activities and so on.

The press release says that records are maintained by researchers at the Berkman Center, which means that keeping the directory current, relevant and error-free will be a challenge. As the number of sources grows, it will also be harder to navigate the directory through search and browse, without more sophisticated tools of filtering, recommendations, and visualizations.

Jun 29, 2015

Increasing sharing, expanding user base, and estimating impact of yourresearch data using service tools and social media - a use case study

It has become increasingly important to communicate and share your research with users and estimate the potential impact of your research data, both during and after its development. Dr. Ge Peng, a research scholar at Cooperative Institute for Climate and Satellites in North Carolina (CICS-NC), and Thomas Maycock, Science Public Information Officer, share their thoughts on different ways of sharing your data, expanding your user base, and broadening the impact of your research. They compare several communication platforms and service tools in this set of slides, outlining some advantages and disadvantages based on their own experiences.

Below, they describe how tools and services have been used with one of the main research products: the scientific data stewardship maturity assessment model.

The main research product featured in this presentation is the scientific data stewardship maturity assessment model, in the form of a matrix, which is jointly developed by scientists and data managers from CICS-NC and NOAA’s National Centers for Environmental Information (NCEI). The matrix provides a unified framework for assessing stewardship practices applied to individual digital Earth Science datasets and is published by the Data Science Journal (http://dx.doi.org/10.2481/dsj.14-049).

Prior to baselining and publishing the matrix in the peer-reviewed journal, the vehicle of slideshare.com was utilized to allow public viewing and, by invitation only, downloads of the beta versions of the matrix along with a set of background slides. The slides were used in communicating, either directly or via e-mail lists, to people at data-management-oriented conferences, groups, and organizations. This communication has proven to be very beneficial in improving the consistency of content in the matrix by obtaining feedback from a much wider pool of experts in the field.

The capability of creating shorter and meaningful URLs by tinyurl.com was useful for customizing URL links for use in tweets, e-mails, and presentations.

The vehicle of figshare.com was used to issue a persistent digital object identifier (DOI) in a timely fashion. This DOI is included in the matrix journal publication to provide users with sustained and trackable access to the latest version of the matrix.

Dr. Peng indicated that keeping both sets of slides (matrix and high-level background) at slideshare.com after the publication of the matrix allows her to continue to reach out to users in domains and countries beyond her original expectations. The analytics provided by slideshare.com provide a good indicator of potential interest and impact.

The web stories and social media are coordinated efforts between CICS-NC and NCEI by the communication teams from both organizations. As indicated in slide 6, they bring noticeable traffic to the site. Again, the online presence of those web stories, tweets and Facebook has a long-lasting effect.

Submitted by Ge Peng, Cooperative Institute for Climate and Satellites in North Carolina (CICS-NC)

Feb 20, 2015

Repository features to motivate more data sharing

One of the challenges of creating data stewardship infrastructure is engaging the users and meeting and prioritizing their needs, particularly the needs of long-tail science research. "What would motivate researchers to make their data available?" is a question we continuously grapple with. A recent study "Potential contributor perspectives on desirable characteristics of an online data environment for spatially referenced data" published in First Monday asked a very similar question in the context of geographic data. The researchers hypothesized that potential data contributors of small scale, local spatial data would be more willing to share their data if a repository included a simple, clear licensing mechanism, a simple process for attaching descriptions to the data, and a simple post-publication peer evaluation/commenting mechanism.

The paper draws on 10 qualitative interviews and 110 responses to an online questionnaire. The qualitative interview responses were mixed; they don't seem to reveal any patterns or unusual concerns. Some of the quantitative results were also mixed, but some provide good numbers to support the hypotheses:

  • 90% of respondents said attribution (licensing) is important
    • 62% think that non-commercial attribution is important
    • 54% think that restricting re-use is important, i.e., others may use the data but not modify it in any way
  • 93% said ability to attach keywords or other descriptions to data is important
  • 78% said that commenting capability is important
  • 85% said that stability and long-term maintenance of the repositories matters

Conclusion:

This research, subject to the caveats listed below, suggests that it would be desirable from the perspective of potential contributors of data to provide infrastructure capability that would:

  • allow users to attach conditions to the use of their data;
  • provide basic information that could be translated into standards based metadata; and,
  • receive comments and feedback from users.

Feb 18, 2015

Research Data Alliance/US Call for Fellows

I'm a co-PI on a project that provides a great opportunity to the early career researchers and professionals to engage with the Research Data Alliance and help to improve data practices and make data management and data sharing easier and more transparent. Below are the details from the call for fellows:
The Research Data Alliance (RDA) invites applications for its newly redesigned fellowship program. The program’s goal is to engage early career researchers in the US in Research Data Alliance (RDA), a dynamic and young global organization that seeks to eliminate the technical and social barriers to research data sharing.

The successful Fellow will engage in the RDA through a 12-18 month project under the guidance of a mentor from the RDA community. The project is carried out within the context of an RDA Working Group (WG), Interest Group (IG), or Coordination Group (i.e., Technical Advisory Board), and is expected to have mutual benefit to both Fellow and the group’s goals.

Fellows receive a stipend and travel support and must be currently employed or appointed at a US institution.

Fellows have a chance to work on real-world challenges of high importance to RDA, for instance:
  • Engage with social sciences experts to study the human and organizational barriers to technology sharing
  • Apply a WG product to a need in the Fellow’s discipline
  • Develop plan and disseminate RDA research data sharing practices
  • Develop and test adoption strategies
  • Study and recommend strategies to facilitate adoption of outputs from WGs into the broader RDA membership and other organizations
  • Engage with potential adopting organizations and study their practices and needs
  • Develop outreach materials to disseminate information about RDA and its products
  • Adapt and transfer outputs from WGs into the broader RDA membership and other organizations
The program involves one or two summer internships and travel to RDA plenaries during the duration of the fellowship (international and domestic travel). Fellows will receive a $5000 stipend for each summer of the fellowship. Fellows will be paired with a mentor from the RDA community.

Through the RDA Data Share program, fellows will participate in a cohort building orientation workshop offering training in RDA and data sciences. This workshop is held at the beginning of the fellowship. RDA Data Share program coordinators will work with Fellows and mentors to clarify roles and responsibilities at the start of the fellowship.

Criteria for selection: The Fellows engaging in the RDA Data Share program are sought from a variety of backgrounds: communications, social, natural and physical sciences, business, informatics, and computer science. The RDA Data Share program will look for a T-shaped skill set, where early signs of cross discipline competency are combined with evidence of teamwork and communication skills, and a deep competency in one discipline.

Additional criteria include: interest in and commitment to data sharing and open access; demonstrated ability to work in teams and within a limited time framework; and benefit to the applicant’s career trajectory.

Eligibility: Graduate students and postdoctoral researchers at institutions of higher education in the United States, and early career researchers at U.S.-based research institutions who graduated with a relevant master’s or PhD and are no more than three years beyond receipt of their degree. Applications from traditionally underserved populations are strongly encouraged to apply.

To apply: Interested candidates are invited to submit their resume/curriculum vitae and a 300-500 word statement that briefly describes their education, interests in data issues, and career goals to datashare-inquiry-l@list.indiana.edu. Candidates are encouraged to browse the RDA website https://rd-alliance.org/ and pages of interest and working groups to identify relevant topics and mutual interests.

Important dates:
April 16, 2015 – Fellowship applications are due
May 1, 2015 – Award notifications
June 18-19, 2015 – Fellowship begins with the orientation workshop in Bloomington, IN

RDA Data Share, funded by the Alfred P. Sloan Foundation under award G-2014-13746, engages students and early career researchers in the Research Data Alliance. This engagement builds on foundational infrastructure funded by the National Science Foundation grant # ACI-1349002.

Feb 11, 2015

Institutional analysis of data practices

A short summary of a paper published in JASIST recently: Mayernik, M. S. (2015), Research data and metadata curation as institutional issues. J Assn Inf Sci Tec. doi:10.1002/asi.23425.

The paper begins by noticing a mismatch between the findings of two studies on the data practices in climate science. One of them (a report commissioned by the UK Research Information Network RIN) described the level of data sharing in climate science as low and the other (the book by Edwards "A vast machine...") argued that data sharing was a strong and common norm in climate science. Which one is true? Or, could it be that both studies are correct and climate science includes both the high and the low data sharing levels?

Data practices are institutionalized within a number of social systems, including formal organizations (such as universities and research centers), rules and sanctions (such as funding agency requirements and professional guidelines), and the norms of modern Western science, so the case study analysis in this paper is grounded in the institutional framework that has five characteristics: (a) norms and symbols, (b) intermediaries, (c) routines, (d) standards, and (e) material objects. Norms are largely associated with the norms of science (Merton and later work), symbols are logos and other visible signs of collective identity, but also terminological choices and metaphors. Intermediaries are individuals or collectives who connect resources and facilitate relationships. Routines are frequently repeated patterns of action and interaction, for example, meal or socializing routines. Standards are rules and specifications that define how information can be presented, organized, and transferred. Material objects are ... material objects.

The case studies are comparisons between data practices at the Center for Embedded Networked Sensing (CENS) and the Long Term Ecological Research (LTER) network and between the University Corporation for Atmospheric Research (UCAR)and the National Center for Atmospheric Research (NCAR).

Although there are some interesting observations in these case studies, it seemed that the first, conceptual part of the paper was stronger than the second. The five characteristics of the institutional framework were applied rather narrowly, without revealing many interconnections and directionality. For example, the standards section focuses on metadata standards and their choice. Are there any other standards relevant to data practices? How does the choice of standards affect norms and what is the role of intermediaries in establishing routines and other aspects of data practices? Another much more important question is: Once we describe the variability of data practices within and across disciplines, what's next? What exactly is the role of each institutional carrier in data practices?

Jan 13, 2014

Identifier Test-bed Activities Report (ESIP Federation)

Below is a brief summary from a recent report to ESIP Federation's Data Stewardship Committee that evaluated identifier schemes for Earth system science data and information(see also executive summary and links). The report seems to be a hands-on continuation of the paper published in 2011 "On the utility of identification schemes for digital earth science data: an assessment and recommendations" by Ruth Duerr and others(link).
The paper introduced four uses cases and three assessment criteria:
Use cases:
  • unique identification (identify a piece of data, no matter which copy)
  • unique location (locate an authoritative copy)
  • citable location (identify cited data)
  • scientifically unique identification (to tell whether two data instances have the same info even if the formats are different)
Assessment criteria:
  • Technical value (e.g., scalability, interoperability, security, compatibility, technological viability)
  • User value (e.g., publishers' commitment, transparency)
  • Archive value (e.g., maintenance, cost, versatility)
The report took those use cases, expanded assessment criteria and used all of it to test the implementation of 8 identification schemes, DOI, ARK, UUID, XRI, OID, Handles, PURL, LSID, and URI/URN/URL, using two datasets: the Glacier Photo Collection from the National Snow and Ice Data Center (JPEG and TIFF images) and a numerical data set from the NASA's Moderate Resolution Imaging Spectroradiometer (MODIS) sensor.
Report recommendations:
  • UUID are most appropriate as unique identifiers, any other use requires effort.
  • DOI, ARK and Handles are the most suitable as unique locators, DOI and ARK also support citable locators. Handles need a local dedicated server. ARKs are cheaper than others, but DOIs are accepted by publishers.
  • PURL has no means for creating opaque identifiers and the API support for batch operations is poor.
  • The rest of the ID schemes are less suitable.

It seems that the overall conclusion is that DOI and ARK are generally better, but there is a need for support of multiple ID schemes in a system. From the report I didn't quite get whether any of the ID schemes can support the fourth use case - scientifically unique identification. The paper argued that "none of the identifier schemes assessed here even minimally address this use case".

Oct 24, 2013

Faculty engagement for librarians and curators

The Council on Library and Information Resources (CLIR) postdoctoral fellowships in data curation encourage connections between library, technology and research. CLIR fellows are hosted by a variety of institutions and have widely varying skills and responsibilities. Every month they (we) get together to discuss current issues and challenges related to digital curation. The most recent session focused on faculty engagement and featured two guests: Gabrielle Dean, the curator of literary rare books and manuscripts at Johns Hopkins University and Kelly Miller, the director of teaching and learning services and head of the college library at UCLA. Two current CLIR fellows, John Kratz and Bridget Whearty, led the session.

The notes below are my attempt to synthesize many useful pieces of advice and information shared during that session.

Engagement in the context of librarians/curators working with faculty is a relatively new term. Previously, the word "outreach" was used more often. "Outreach" has a sense of a method or certain approach, e.g., providing guidelines or distributing best practices. "Engagement" has a sense of participation, shared goals and activities. As any other type of engagement, faculty engagement is rather difficult. So here are some tips:

  • Set small goals and gradually extend your network, because engagement is an incremental activity that involves trust, relationship building and a lot of trials and errors.
  • Be positive (or even nice and cheerful if you can), your positive attitude toward your and other’s work will pass to others and invite them to be more open and interactive.
  • Be modest. Not that many people may be interested in the library and its services, automatically assuming that it’s valuable to others may backfire.
  • Be curious. Engagement is an opportunity to learn about many interesting things and people. Don’t be afraid to ask questions or even be naive sometimes, it will pay off in more knowledge and more connections.
  • Always say "yes" (within reasonable limits). This may be important at the beginning of engagement initiatives. Willingness to do the work communicates good will and stimulates interest. Get involved in possibly tangential projects, attend informal gatherings, create connections and then follow up and maintain connections via phone calls, emails, newsletters, etc.
  • Learn "the language". Engaging other audiences sometimes means getting into conversations without adequate background knowledge or expertise. Learning about faculty research in advance can help with terminology and ability to ask intelligent questions.
  • Talk, don't just listen. Listening is useful, but it’s also important to talk. Conversing, (i.e., listening, asking questions, and encouraging others to ask questions) helps to find shared points or issues to address.
  • Look for creative opportunities for engagement, e.g., shared learning, teaching, activist and interest groups, and so on. Engagement doesn't have to be limited to faculty. Engaging other interested groups, including undergraduate students, local schools, or private organizations can be useful and quite rewarding.
  • Seek effective ways of gathering information. Some faculty may not be responsive to emails or phone calls, graduate students may be more responsive to certain requests, surveys are not effective due to low response rate. Sometimes it depends on the institution and the nature of curated content. Think through priorities, audiences, and context in order to get best results.
  • Mistakes happen. False starts and even failures in faculty engagement are common for everyone. Rather than dwelling on mistakes, try to learn from them and do better next time.
  • Avoid political entanglements and personal battles. These things happen in many if not all organizations. To maintain a good working environment, try to stay positive, focus on the goals, look for opportunities to be creative and don’t take matters personally.

I wish there was more literature (both formal and informal) on this topic that could answer questions and help in practice. For example, what should someone know before starting a job that involves faculty engagement? Are there certain skill that might be helpful? What is the nature of the relationship between curators and faculty - is it a peer-to-peer or a nurse-doctor relationship? Can/should it be changed? What are the best ways to gather information about your targeted audiences and their needs? So on and so forth.

To stimulate a discussion on this topic and, perhaps, encourage more writing, an engagement interest group has been created within the Research Data Alliance (RDA). The group has a narrower focus, because it emphasizes the engagement of researchers and other stakeholders in research data sharing and re-use. Nevertheless, it may be a good platform for continuing a conversation on engagement and building a knowledge base/wiki.

Oct 2, 2013

About research objects

Notes from the article by Bechhofer, Buchan, De Roure, Missier, Ainsworth et al. "Why linked data is not enough", Future Generation Computer Systems, 2011, (pdf).

Scientific research is increasingly digital and collaborative, therefore a new framework is needed that would facilitate the reuse and exchange of digital knowledge. Simply publishing data fails to reflect the research methodology and respect the rights and reputation of the researcher.

The concept of Research Objects (ROs) as semantically rich aggregations of resources can serve as a cornerstone of such new framework. ROs would include research questions, hypotheses, abstracts, organisational context (e.g., ethical and governance approvals, investigators, etc.), study design, methods (workflows, scripts, services, software packages. etc.), data, results, answers (e.g., publications, slides, DOIs), etc. The authors argue that this approach is better than linked data, but later they acknowledge that linked data works fine, it just needs to be revised and extended.

Important assumptions in the paper:

  • ROs work well in the context of e-Laboratories - environments that are mostly based on automated management systems and execution of in silico experiments
  • Reproducible research is ultimately possible in any domain and always desirable.
  • All elements of scientific research can be made explicit and encoded in a machine-readable way, if not now, then in the future.

Terms that refer to different ways of reusability:

  • Reusable - reuse as a whole or single entity.
  • Repurposeable - reuse as parts, e.g., taking an RO and substituting alternative services or data for those used in the study.
  • Repeatable - repeat the study, perhaps years later.
  • Reproducible - reproduce or replicate a result (start with the same inputs and methods and see if a prior result can be confirmed).
  • Replayable - automated studies can be replayed rather than executed again.
  • Referenceable - citataions for ROs.
  • Revealable - audit the steps performed in the research in order to be convinced of the validity of results.
  • Respectful - credit and attribution.

The authors describe several environments that try to implement aggregation of resources into ROs approach.

  • myExperiment Virtual Research Environment relies on the notion of "packs", collections of items that can be shared as a single entity.
  • Systems Biology of Microorganisms (SysMO) project has a web plaform SysMO-DB and a catalog SysmoSEEK. It relies on a JERM (Just Enough Results Model), which is based on the ISA (Investigation/Study/Assay) format. Another approach to support ROs within the systems biology community is SBRML (Systems Biology Results Markup Language). Most of the experiments in this domain are wet lab experiments, so traceability and referenceability are more relevant than repeatability and replayability.
  • MethodBox is part of the Obesity e-Lab, that allows researchers to "shop for variables" from studies related to obesity in the UK. The paper doesn't describe what method is used to support RO aggregations.

Packs in myExperiment is the most advanced implementation of the idea of ROs and, ironically, it's based on linked data: "Work in myExperiment makes use of the OAI-ORE vocabulary and model in order to deliver ROs in a Linked Data friendly way" (p. 10).

OAI-ORE defines standards for the description and exchange of aggregations of Web resources. It is agnostic to relationship types, so it needs to be extended. The authors propose the following extensions: the Research Objects Upper Model (ROUM) and the Research Object Domain Schemas (RODS). ROUM provides basic vocabulary to describe general properties of RO, such as the basic lifecycle states. RODS provide domain specific vocabulary. Not much details are provided about these two extensions.

Rather than arguing that linked data is not enough, it seems that the paper argues that current implementations of linked data in packaging scientific results needs to be revised to explicitly include the structure of aggregations. The purpose of articulating structure in a machine-readable way is to create an environment where every component of research (including hypotheses, methods, data and results) can be re-enacted. A more obvious and important conclusion from the discussion about ROs is that a) we need to keep encouraging exchange and sharing of research in ways that are more transparent; b) there is still a shortage of platforms to do that. MyExperiment is a nice example, but it's still domain and platform-specific.

The approach described in this paper is quite forward-looking. It is a call for rather radical changes in scientific practices. I wonder how many labs have automated experiment management environments where all datasets, workflows, scripts and results can be connected and reconstructed without much back-channeling. Another question is how much effort it takes to create ROs in a way that would make science fully "re-enactable". We probably won't be able to do that with legacy data.

Aug 5, 2013

Research Data Alliance (RDA) Plenary

Spreading the word about an interesting event/initiative I'm involved with:

Second Plenary of the Research Data Alliance, National Academy of Sciences, Washington, DC, September 16–18, 2013

The Research Data Alliance invites participants from all walks of the research/data world to join us at the National Academy of Sciences for our Second Plenary!

  • Great keynote speakers — including John Wilbanks of Sage Bionetworks and Carole Palmer, Univ. of Illinois
  • Representatives from supporting agencies like the US National Science Foundation, the US National Institute of Standards and Technology, the European Commission, and the Department of Innovation through the Australian National Data Service (ANDS)
  • People from organizations like ESIP, CODATA, Microsoft Research, and W3C
  • Opportunity to participate in a poster session
  • Working sessions, breakouts, working and interest groups and much more

More info at https://www.rd-alliance.org/future-events.

Jul 15, 2013

The Practice of Data Curation - Archive Journal issue

A little bit of self-promotion.

Archive Journal in its latest 360° section focuses on the practice of data curation.

As research and teaching produce ever-increasing amounts of data in analog and digital forms, what we do with that data is a question that librarians, archivists, scholars, teachers, and students must address. The four contributors discuss what “data curation” is and might become. We invite you to read through the responses by author or by question.

I was one of the contributing authors. It was a great pleasure and a challenge to write my responses.