Showing posts with label data sharing. Show all posts
Showing posts with label data sharing. Show all posts

Sep 16, 2016

Data for humanitarian purposes

Unni Karunakara, a former president of "Doctors without borders", gave a talk at International Data Week 2016 on September 13 about the role of data in humanitarian organizations. The talk was very powerful in its simplicity and urgent need for better data and its management and dissemination. It was a story of human suffering, but also a story of care and integrity in using data to alleviate it.

Humanitarian action can be defined as moral activity grounded in the ethics of assistance to those in need. Four principles guide humanitarian action:
  • humanity (respect for the human)
  • impartiality (provide assistance because of person's need, not politics or religion
  • neutrality (tell the truth regardless of interests)
  • independence (work independently from governments, businesses, or other agencies)
These principles affect how to collect and use data and how to ensure that data helps. Data collected for humanitarian action is evidence that can be used for direct medical action and for bearing witness, which is a very important activity of humanitarian organizations:  
“We are not sure that words can always save lives, but we know that silence can certainly kill." (quoted from another MSF president)
Awareness of serious consequences of data for humanitarian action makes "Doctors without borders" work only with data they collect themselves and use stories they witnessed firsthand. Restraint and integrity in data collection is crucial in maintaining credibility of the organization.

Lack of data or lack of mechanisms to deliver necessary data hurts people. Thus, in Ebola outbreak it took the World Health Organization about 8 months to declare emergency and 3000 people died because data was not available in time or in the right form. The Infectious Diseases Data Observatory (IDDO) was created to help with tracking and researching infectious diseases by sharing data, but many ethical, legal, etc. issues still need to be solved.

Humanitarian organizations often do not have trustworthy data available, either because of competing definitions or lack of data collection systems. For example, because of the differences in defining "civilian casualty" numbers of civilians killed in drone strikes range from a hundred to thousands. Or, in developing countries or conflict zones where census activities are absent or dangerous, counting graves or tents becomes a proxy of mortality, mobility rates and other important indicators. Crude estimates then are the only available evidence.

"Doctors without borders" (MSF) does a lot to share and disseminate its information. It has an open data / access policy and aspires to share data, while placing high value on security and well-being of people it helps.


Apr 18, 2016

Dataset on Parkinson's disease

In March 2016 Sage Bionetworks released a dataset that captures the everyday experiences of over 9,500 people with Parkinson's disease (press release). The data described in the data paper "The mPower study, Parkinson disease mobile data collected using ResearchKit" was collected via the mPower iPhone app, where participants were presented with tasks (referred to as ‘memory’, ‘tapping’, ‘voice’, and ‘walking’ activities) and asked to fill out surveys.

Not everybody agreed to share their data broadly with the research community. Out of 14,684 verified participants 9,520 (65%) agreed to share broadly, the rest split between withdrawing from the study and agreeing to share narrowly with the team only:

Study cohort description
Figure 1: mPower study cohort description. From http://www.nature.com/articles/sdata201611#methods

To provide proper safeguards and to balance sharing and privacy, the research team established a data governance structure. Access is granted to qualified researchers who agree to specific conditions for use, including the following:

  • participants cannot be re-identified
  • the data may not be redistributed
  • findings need to be published in open access venues
  • both participants and research team need to be acknowledged as data contributors
This effort is another example of the newly forming data sharing culture. And it uses Synapse that seems to make sharing easier from both technical and policy perspectives.

Jan 11, 2016

Pantheon 1.0: A manually verified dataset of globally famous biographies



Scientific Data has published a description of an interesting dataset: "Pantheon 1.0, a manually verified dataset of globally famous biographies". This data collection effort contributes to quantitative data for studying historical information, especially, the information about famous people and events.


Data collection workflow
Workflow diagram (Image from the paper)
The authors retrieved over 2 mln records about famous ("globally known") individuals from Google's Freebase, narrowed down the dataset to individuals who have metadata in English Wikipedia and then reduced it further to people who have records in more than 25 different languages in Wikipedia.

Manual cleaning and verification includes a controlled vocabulary for occupations, popularity metrics (defined as a number of Wikipedia edits adjusted by age and pageviews).




The dataset is available for download at Harvard Dataverse http://dx.doi.org/10.7910/DVN/28201. Another entertaining part is a visualization interface at http://pantheon.media.mit.edu that allows to explore the data and answer questions like "Where were globally known individuals in Math born?" (21% in France) or "Who are the globally known people born within present day by country?". Turns out that Russia produced a lot of politicians and writers, while the US gave us many actors, singers and musicians. 

Globally known people born in the US (from http://pantheon.media.mit.edu/treemap/country_exports/US/all/-4000/2010/H15/pantheon)


Sep 17, 2015

Valuable lessons from sharing and non-sharing of data

A vivid story from Buzzfeed "Scientists Are Hoarding Data And It’s Ruining Medical Research" describes two related cases - one where researchers voluntarily shared their entire dataset and how the re-analysis  found errors and miscalculations and another one where the data or any results from a largest drug trial were not released for 7 years because the researchers feared criticism and continued double-checking their data. The details from each of the cases are worth following up, but the author of the story comes to an important conclusion that we need to accept that science works through checks and corrections, stop unfair criticisms and doubts in researchers' credibility, and start sharing data for better science, better knowledge, and ultimately, better informed decisions that impact our lives:

And here is where I think the threads come together. The press releases on the reanalysis of the Miguel and Kremer deworming trial in Kenya will go live this week. Somewhere, I’m sure, people will attack or mock them for their errors. One way or another, I can’t believe they won’t feel bruised by the reanalysis. And that is where we have gone wrong. It’s not just naive to expect that all research will be perfectly free from errors, it’s actively harmful.


There is a replication crisis throughout research. When subjected to independent scrutiny, results routinely fail to stand up. We are starting to accept that there will always be glitches and flaws. Slowly, as a consequence, the culture of science is shifting beneath everyone’s feet to recognise this reality, work with it, and create structural changes or funding models to improve it.

 

Sep 15, 2015

Dataset: Roads and cities of 18th century France

An interesting dataset has been described in the Scientific Data journal and shared via the Harvard Dataverse repository - "Roads and cities of 18th century France".
The database presented here represents the road network at the french national level described in the historical map of Cassini in the 18th century. The digitization of this historical map is based on a collaborative methodology that we describe in detail. This dataset can be used for a variety of interdisciplinary studies, covering multiple spatial resolutions and ranging from history, geography, urban economics to network science.

The repository page showed 268 downloads on Sept 15, 2015, so hopefully, some examples of data re-use will follow this publication.

Aug 31, 2015

Lessons from replication of research in psychology

Science magazine has published an article “Estimating the reproducibility of psychological science”, which reports the first findings from 100 replications completed by 270 contributing authors. A quasi-random sample was drawn from three psychology journals: Psychological Science (PSCI), Journal of Personality and Social Psychology (JPSP), and Journal of Experimental Psychology: Learning, Memory, and Cognition (JEP:LMC). The replications were performed by teams and then independently reviewed by other researchers and reproduced by another analyst. The reproducibility was evaluated using significance and P values, effect sizes, subjective assessments of replication teams, and meta-analyses of effect sizes. Some highlights from the results:

  • 35 studies in the replications showed positive effect of p < 0.05 compared to 97 original studies

  • 82 studies showed a stronger effect size in the original study than in the replication

  • Effect size comparisons showed a 47.4% replication success rate

  • 39 studies were subjectively rated as successfully replicated


While some news about this publication reported failures in the test (e.g., Nature’s "Over half of psychology studies fail reproducibility test"), the Science article emphasized the challenges of reproducibility itself and care with which interpretations of successes and failures need to be made. The authors of the study pointed out that while replications produced weaker evidence for the original findings,
“It is too easy to conclude that successful replication means that the theoretical understanding of the original finding is correct. Direct replication mainly provides evidence for the reliability of a result. If there are alternative explanations for the original finding, those alternatives could likewise account for the replication. Understanding is achieved through multiple, diverse investigations that provide converging support for a theoretical interpretation and rule out alternative explanations.

It is also too easy to conclude that a failure to replicate a result means that the original evidence was a false positive. Replications can fail if the replication methodology differs from the original in ways that interfere with observing the effect. We conducted replications designed to minimize a priori reasons to expect a different result by using original materials, engaging original authors for review of the designs, and conducting internal reviews. Nonetheless, unanticipated factors in the sample, setting, or procedure could still have altered the observed effect magnitudes.”

Jul 20, 2015

Study: Biomedical data sharing and reuse

A recent publication in PLoS ONE "Biomedical Data Sharing and Reuse: Attitudes and Practices of Clinical and Scientific Research Staff" surveyed the Intramural Research Program at the US National Institutes of Health (NIH) with regard to data management, data sharing, and data re-use.  The authors received 190 responses and analyzed 135 (scientific and clinical staff). Below are the highlights from their findings:

  • ~60% of respondents rated relevance of data re-use as high, while ~15% rated it as low

  • ~25% rated their expertise in re-using data as high, while ~45% rated it as low

  • ~61% reported that they had never uploaded a dataset into a repository, while ~71% said they had shared data directly with another researcher

  • `30%  indicated that it took them more than 10 hours to prepare data for sharing

  • Only 20 respondents provided reasons for not sharing data and their reasons were pretty scattered (see image below):


Image from the study (t016), responses to "...the reason(s) for not sharing your data"

The data from this study is available on Figshare, but it is not the full survey dataset, it's a subset to support the results of this publication. And for some reason it doesn't contain free text responses to the non-sharing question. It's always informative to see what people say beyond the provided standard categories (that is usually the most interesting story in my mind). Perhaps, there were no free text responses.

Citation: Federer LM, Lu Y-L, Joubert DJ, Welsh J, Brandys B (2015) Biomedical Data Sharing and Reuse: Attitudes and Practices of Clinical and Scientific Research Staff. PLoS ONE 10(6): e0129506. doi:10.1371/journal.pone.0129506

Jul 15, 2015

Archaeology meets modern scanning technology for preservation and re-use

Submitted by Annemiek van der Kuil, edited by Inna Kouper

Image from "The strange case of 60 frothy beads: puzzling Early Iron Age glass beads from the Netherlands" conference paper by D.J. Huisman et al.
Dr. Dominique Ngan-Tillard, a professor at the Faculty of Civil Engineering and Geosciences at Delft University of Technology, the Netherlands, has deposited a dataset into the 3TU.Datacentrum repository that contains tomography scans of early Iron Age glass beads found during the archaeological excavations in the Netherlands.

The dataset supports a conference publication by Dr. Ngan-Tillard and others “The strange case of 60 frothy beads: puzzling Early Iron Age glass beads from the Netherlands”. The micro-CT scans helped to identify gas bubbles and mineral and metal inclusions in the glass beads, which allowed the researchers to conclude that “the Zutphen glass beads are the result of local, inexpert, reworking of imported glass objects” (p. 231, conference paper).

In addition to the in-depth analysis of the beads’ structure, the scans serve as a form of virtual preservation of the ornaments. Stored in a data repository and made publicly available, they can help other archaeologists, as well as material scientists and museums in their research and educational activities. In the future 3D prints of the ornaments can be produced for a better understanding of the art of making glass and jewels.

According to Dr. Ngan-Tillard’ comment on the 3TU.Datacentrum website, storing digital collections of archaeological remains together with their meta-data and interpretation will help advance both arts and research and create more challenges for our knowledge.

Watch a short video about frothy beads or see the full story at http://datacentrum.3tu.nl/en/researchers-about-3tudatacentrum/showcase-ngan-tillard/

Jul 10, 2015

Withholding data - questionable science or scientific misconduct

Nicole Janz writes in the LSE Impact of Social Sciences blog  that not sharing one’s research data should be considered a scientific misconduct. This will help to fight data secrecy and establish better research practices. A few key points from the post:

  • Many researchers don’t share data even if they promise to do so - see, for example, Krawczyk and Reuben’s 2012 study “(Un)available upon request: field experiment on researchers' willingness to share supplementary materials [see also “How and why researchers share data (and why they don’t)”]

  • Scientific misconduct definitions usually includes fabrication, falsification or plagiarism. Sharing research data provides evidence that there was no fabrication or falsification involved, hence it’s crucial in avoiding misconduct allegations and demonstrating proper conduct.

  • A broader definition of scientific misconduct includes departure from accepted standards and practices of a research community. As many research communities strive to be open with regard to the evaluation of their knowledge claims, obligations to share data can be seen as part of the research standards and practices. Hence, data secrecy can be considered a questionable research practice or a misconduct.

The continuum of research practices described by Janz ranges from the gold standards of open data, open code, pre-registration and version control to questionable research practices of p-hacking, sloppy statistical methods and other manipulations to withholding data to misconduct with its fabrication, falsification, and plagiarism.

[caption id="" align="aligncenter" width="433"]From Janz' LSE Impact of Social Sciences blog post: Research practices continuum From Janz' LSE Impact of Social Sciences blog post: Research practices continuum[/caption]

Jul 7, 2015

The International Polar Year 2007-2008 (IPY-4) and the importance ofdata management

The International Polar Year is an international collaboration that focuses on the Arctic and the Antarctic, or polar regions. The polar regions have many unique phenomena, but the cold harsh environment makes them expensive to visit and study. It takes a large multi-country collaborative effort to put together expeditions, install equipment, and collect data. The first three IPYs occurred in 1882–1883,1932–1933, and 1957–58 respectively. The fourth IPY took place between March 2007 and March 2009.

The fourth IPY was dramatically different from the previous efforts (Mokraine and Parsons, 2013). A $1.2 bln effort with participants from more than 60 countries, it had an ambitious vision to enable international sharing and reuse of multidisciplinary datasets and keep the data discoverable, open, linked, useful, and safe (Parsons, Godoy et al., 2011). The enormous efforts to initiate, coordinate, improve, and sustain IPY data stewardship have seen both successes and failures, with some components of the IPY infrastructure struggling to exist and be useful (Lessons and legacies..., 2012).

A fair amount of IPY data is available via such online portals as the IPY data page at the National Snow and Ice Data Center (NSIDC) in the US, the NASA Global Change Master Directory (GCMD) IPY portal, or the Global Cryosphere Watch portal. Some of it, such as a global IPY Data and Information System (IPYDIS)  or the Discovery, Access, and Delivery of Data for IPY (DADDI) are broken. Most importantly though, missing is a way to track and access all the IPY data via a federated or centralized catalog. There is no good consistent way of international polar data to “function locally and reach globally”, to use Mokraine and Parsons’ words.

The challenges of making heterogeneous data and metadata work together were exacerbated by the lack of focused international funding for planning data archiving post-IPY, differences in data policies and researchers “hoarding” data (Lessons and legacies..., 2012; Carlson, 2011). Despite many IPY projects adopting a free and open data-sharing policy, compliance with it and, ultimately, sharing was rather low. Additionally, the researchers in IPY-4 didn’t have access to data from the first IPY projects, some of the data were not available in the digital form, while others were scattered or lost. The data centers (WDCs) that were supposed to support the increasing IPY data streams, lacked mechanisms of working with heterogeneous data, e.g., they couldn’t support social and ecological data.

Despite the difficulties, the IPY data management experience is crucial to the advancement of global data services and the norms of data sharing and re-use. As Mark Parsons, Secretary General of the  Research Data Alliance and former Senior Associate Scientist and the Lead Project Manager at NSIDC put it,
“We were perhaps rather naive going in to IPY. Many of the organizers came from the geoscience background of earlier the IPYs and assumed data systems would exist that could handle IPY data. We weren’t prepared for the incredible diversity of IPY4 with data ranging from Indigenous knowledge to satellite remote sensing to genomic sequencing to cosmology. Although it is unclear what percentage of IPY data are available and much is surely lost, new data services were created and sustained, international coordination continues in sustained organisations, and we learned a lot about different disciplinary cultures and their attitudes to data sharing. The IPY Data Policy was aggressive and not fully honored, but it did drive changes in national policies towards more timely and open release of data. Most critically we saw a change in the conversation within polar science from whether to share to when to share and now how to share. We have a long way to go, but polar data are significantly more accessible than they were prior to IPY.”

Mark’s and others’ publications, some of which are listed below, are a good source of all the lessons learned from IPY data stewardship efforts, one important lesson being that “[e]xperts in data management are critical members of any team attempting internationally coordinated science ...” (Lessons and legacies..., 2012).

Resources

Carlson D. 2011. A lesson in sharing. Nature 469 (293).

Lessons and legacies of the International Polar Year 2007-2008. 2012.

Mokrane M and MA Parsons. 2014. Learning from the international polar year to build the future of polar data management. Data Science Journal 13.

Parsons MA, Ø Godøy, E LeDrew, TF de Bruin, B Danis, S Tomlinson, and D Carlson. 2011. A conceptual framework for managing very diverse data for complex interdisciplinary science. Journal of Information Science 37 (6): 555-569.

Parsons MA, T de Bruin, S Tomlinson, H Campbell, Ø Godøy, J LeClert, and IPY Data Policy and Management SubCommittee. 2011. The state of polar data—the IPY experience. In Understanding Earth’s Polar Challenges: International Polar Year 2007-2008. Ed. Krupnik I, I Allison, R Bell, P Cutler, D Hik, J López-Martínez, V Rachold, E Sarukhanian, and C Summerhayes. Edmonton, Canada: CCI Press.

Jun 29, 2015

Increasing sharing, expanding user base, and estimating impact of yourresearch data using service tools and social media - a use case study

It has become increasingly important to communicate and share your research with users and estimate the potential impact of your research data, both during and after its development. Dr. Ge Peng, a research scholar at Cooperative Institute for Climate and Satellites in North Carolina (CICS-NC), and Thomas Maycock, Science Public Information Officer, share their thoughts on different ways of sharing your data, expanding your user base, and broadening the impact of your research. They compare several communication platforms and service tools in this set of slides, outlining some advantages and disadvantages based on their own experiences.

Below, they describe how tools and services have been used with one of the main research products: the scientific data stewardship maturity assessment model.

The main research product featured in this presentation is the scientific data stewardship maturity assessment model, in the form of a matrix, which is jointly developed by scientists and data managers from CICS-NC and NOAA’s National Centers for Environmental Information (NCEI). The matrix provides a unified framework for assessing stewardship practices applied to individual digital Earth Science datasets and is published by the Data Science Journal (http://dx.doi.org/10.2481/dsj.14-049).

Prior to baselining and publishing the matrix in the peer-reviewed journal, the vehicle of slideshare.com was utilized to allow public viewing and, by invitation only, downloads of the beta versions of the matrix along with a set of background slides. The slides were used in communicating, either directly or via e-mail lists, to people at data-management-oriented conferences, groups, and organizations. This communication has proven to be very beneficial in improving the consistency of content in the matrix by obtaining feedback from a much wider pool of experts in the field.

The capability of creating shorter and meaningful URLs by tinyurl.com was useful for customizing URL links for use in tweets, e-mails, and presentations.

The vehicle of figshare.com was used to issue a persistent digital object identifier (DOI) in a timely fashion. This DOI is included in the matrix journal publication to provide users with sustained and trackable access to the latest version of the matrix.

Dr. Peng indicated that keeping both sets of slides (matrix and high-level background) at slideshare.com after the publication of the matrix allows her to continue to reach out to users in domains and countries beyond her original expectations. The analytics provided by slideshare.com provide a good indicator of potential interest and impact.

The web stories and social media are coordinated efforts between CICS-NC and NCEI by the communication teams from both organizations. As indicated in slide 6, they bring noticeable traffic to the site. Again, the online presence of those web stories, tweets and Facebook has a long-lasting effect.

Submitted by Ge Peng, Cooperative Institute for Climate and Satellites in North Carolina (CICS-NC)

Jun 5, 2015

Data Sharing in Human Paleogenetics

A study published in PLOS ONE “When Data Sharing Gets Close to 100%: What Human Paleogenetics Can Teach the Open Science Movement”  examined patterns of sharing data related to polymorphisms in ancient mitochondrial DNA and Y and X chromosomes. The authors focused on data that can be considered derivative, i.e., they’re derived from the processing of raw data obtained via such methods as DNA purification, Polymerase Chain Reaction, and so on. 162 PubMed papers on ancient human DNA containing a total of 207 datasets were retrieved.

The authors classified types of sharing (in the text, files for download, supplementary materials, online database) and sent out a survey to the papers’ authors about their choices in sharing human DNA data.

202 datasets out of 207 (97.6%) were fully available and reusable. Five datasets that were initially withheld, have been published later. At the same time more than half of the datasets (57.7%) were shared in the body of the published article, rather than via a database or in separate files.

Among the 33 researchers who responded to the survey, most of them acknowledged the importance of making their studies open to scientific inquiry. Many (97%) also agreed that data sharing should be a common practice in science. The authors of the PLOS paper make a conclusion that a) awareness of the importance of openness in science may help achieve a high data sharing rate; b) modality of sharing (i.e., how data is shared) plays an important role in sharing behavior, and c) openness to the scientific scrutiny of data in human paleogenetics coupled with the adoption of rigorous standards and cross-laboratory validation has been crucial in establishing the field and its scientific rigor and data reliability.

May 5, 2015

The mystery of a missing dataset

An interesting blog post from computer scientist and engineer David Rosenthal details his stepdaughter's quest to locate a dataset which she had downloaded in 2011, but which is now unavailable.

Although this was an important study in her field (sustainability and life cycle analysis) the original link to download the data is broken  and there appear to be no archived versions of the dataset.

Submitted by Isabel Chadwick, Open University

Apr 15, 2015

Selective sharing of usability research data in an organization

Toni Rosati is a data curator and a usability researcher at the National Snow and Ice Data Center (NSIDC). Toni is involved in several projects at NSIDC, including the Advanced Cooperative Arctic Data and Information Service (ACADIS ) and the Science-Driven Cyberinfrastructure: Integrating Permafrost Data, Services, and Research Applications (PermaData).

Toni and I met at the 5th Research Data Alliance plenary and sat down during one of the coffee breaks to talk about qualitative data and the challenges of sharing them. As a usability researcher and a member of the ACADIS team, Toni conducts tests of the Arctic Data Explorer (ADE ), a federated search tool for interdisciplinary Arctic science data. Her research results in recommendations to ADE managers and software developers to improve appearance, functionality, and quality of search results of the tool. Diving deep into the metadata and the code that make the ADE possible, Toni is also, ultimately, looking to improve data management and data sharing practices.

The usability tests generate a wealth of qualitative data including interview recordings and written transcripts. But, can the raw be shared? Not at this point, says Toni, for the following reasons:

  • Toni has been working closely with their institutional review board (IRB), an ethical committee that reviews and approves research involving humans in the U.S. to come to this conclusion. To ensure privacy, raw identifiable data must be anonymized.

  • The raw data are collected in the context of an organization and are most valuable for the organization itself rather than for an outside scholarly community; however, the methods and results are extremely valuable to the outside community and will be shared.


Does anything need to or can be shared in this situation?

In short, yes. As Toni pointed out, the most valuable sharing of qualitative data in an organizational context is the sharing of data that has undergone expert interpretation. “I’m collecting a lot of qualitative data, but to be most useful to my teammates, they have to be distilled into quick actionable items,” says Toni. Qualitative data are sometimes hard to communicate, and building trust in its validity requires time.

Toni and ADE Principal Investigator Lynn Yarmey are writing a paper outlining the user experience / usability research methods and processes they undertook, and the results and lessons learned. Ultimately, usability research is intended to create software with an end user focus that is intuitive, complete, and pleasurable to use. Ms. Rosati is passionate about such research and welcomes your questions.

Mar 25, 2015

Sharing in Paleontology leads to re-use

Submitted by Isabel Chadwick, Open University

This opinion piece on the PLOS Integrative Paleontology blog "And this is why we should always provide our data..." is about a  paleontologist who published data about dinosaur teeth in PLOS One, which was reused, leading to a new discovery about the evolution of small theropod dinosaurs.

Source: Andrew Farke. And this is why we should always provide our data. The Integrative PAleontologists (blog). January 25, 2013. Retrieved from http://blogs.plos.org/paleo/2013/01/25/and-this-is-why-we-should-always-provide-our-data/ 

(licensed under the CCBY Creative Commons Attribution 3.0 Unported License.)

Mar 11, 2015

Depositing in Netherlands repository adds value, researchers say

3TU.Datacentrum offers researchers within the technical and engineering sciences in the Netherlands several research data services. One of these services is a certified data repository.

A brochure describes good practices on some of the data sets in 3TU.Datacentrum. The brochure also contains a few good examples (stories), where researchers explain how depositing their research data to 3TU.Datacentrum, and making it openly available, has added value to their research.

Contributed by Annemiek van der Kuil, 3TU.Datacentrum

3TU.DC_goodpractices_cover

Mar 9, 2015

The Availability of Research Data Declines Rapidly with Article Age

The Availability of Research Data Declines Rapidly with Article Age - article from Current Biology (2014)

Highlights:

  • The availability of data from 516 studies between 2 and 22 years old was studied

  • The odds of a data set being reported as extant fell by 17% per year

  • Broken e-mails and obsolete storage devices were the main obstacles to data sharing

  • Policies mandating data archiving at publication are clearly needed


Submitted by Isabel Chadwick (Open University)

Source

Vines et al. The Availability of Research Data Declines Rapidly with Article Age Current Biology Volume 24, Issue 1, 6 January 2014, Pages 94–97 doi:10.1016/j.cub.2013.11.014

http://www.sciencedirect.com/science/article/pii/S0960982213014000

Feb 20, 2015

Repository features to motivate more data sharing

One of the challenges of creating data stewardship infrastructure is engaging the users and meeting and prioritizing their needs, particularly the needs of long-tail science research. "What would motivate researchers to make their data available?" is a question we continuously grapple with. A recent study "Potential contributor perspectives on desirable characteristics of an online data environment for spatially referenced data" published in First Monday asked a very similar question in the context of geographic data. The researchers hypothesized that potential data contributors of small scale, local spatial data would be more willing to share their data if a repository included a simple, clear licensing mechanism, a simple process for attaching descriptions to the data, and a simple post-publication peer evaluation/commenting mechanism.

The paper draws on 10 qualitative interviews and 110 responses to an online questionnaire. The qualitative interview responses were mixed; they don't seem to reveal any patterns or unusual concerns. Some of the quantitative results were also mixed, but some provide good numbers to support the hypotheses:

  • 90% of respondents said attribution (licensing) is important
    • 62% think that non-commercial attribution is important
    • 54% think that restricting re-use is important, i.e., others may use the data but not modify it in any way
  • 93% said ability to attach keywords or other descriptions to data is important
  • 78% said that commenting capability is important
  • 85% said that stability and long-term maintenance of the repositories matters

Conclusion:

This research, subject to the caveats listed below, suggests that it would be desirable from the perspective of potential contributors of data to provide infrastructure capability that would:

  • allow users to attach conditions to the use of their data;
  • provide basic information that could be translated into standards based metadata; and,
  • receive comments and feedback from users.

Feb 18, 2015

Research Data Alliance/US Call for Fellows

I'm a co-PI on a project that provides a great opportunity to the early career researchers and professionals to engage with the Research Data Alliance and help to improve data practices and make data management and data sharing easier and more transparent. Below are the details from the call for fellows:
The Research Data Alliance (RDA) invites applications for its newly redesigned fellowship program. The program’s goal is to engage early career researchers in the US in Research Data Alliance (RDA), a dynamic and young global organization that seeks to eliminate the technical and social barriers to research data sharing.

The successful Fellow will engage in the RDA through a 12-18 month project under the guidance of a mentor from the RDA community. The project is carried out within the context of an RDA Working Group (WG), Interest Group (IG), or Coordination Group (i.e., Technical Advisory Board), and is expected to have mutual benefit to both Fellow and the group’s goals.

Fellows receive a stipend and travel support and must be currently employed or appointed at a US institution.

Fellows have a chance to work on real-world challenges of high importance to RDA, for instance:
  • Engage with social sciences experts to study the human and organizational barriers to technology sharing
  • Apply a WG product to a need in the Fellow’s discipline
  • Develop plan and disseminate RDA research data sharing practices
  • Develop and test adoption strategies
  • Study and recommend strategies to facilitate adoption of outputs from WGs into the broader RDA membership and other organizations
  • Engage with potential adopting organizations and study their practices and needs
  • Develop outreach materials to disseminate information about RDA and its products
  • Adapt and transfer outputs from WGs into the broader RDA membership and other organizations
The program involves one or two summer internships and travel to RDA plenaries during the duration of the fellowship (international and domestic travel). Fellows will receive a $5000 stipend for each summer of the fellowship. Fellows will be paired with a mentor from the RDA community.

Through the RDA Data Share program, fellows will participate in a cohort building orientation workshop offering training in RDA and data sciences. This workshop is held at the beginning of the fellowship. RDA Data Share program coordinators will work with Fellows and mentors to clarify roles and responsibilities at the start of the fellowship.

Criteria for selection: The Fellows engaging in the RDA Data Share program are sought from a variety of backgrounds: communications, social, natural and physical sciences, business, informatics, and computer science. The RDA Data Share program will look for a T-shaped skill set, where early signs of cross discipline competency are combined with evidence of teamwork and communication skills, and a deep competency in one discipline.

Additional criteria include: interest in and commitment to data sharing and open access; demonstrated ability to work in teams and within a limited time framework; and benefit to the applicant’s career trajectory.

Eligibility: Graduate students and postdoctoral researchers at institutions of higher education in the United States, and early career researchers at U.S.-based research institutions who graduated with a relevant master’s or PhD and are no more than three years beyond receipt of their degree. Applications from traditionally underserved populations are strongly encouraged to apply.

To apply: Interested candidates are invited to submit their resume/curriculum vitae and a 300-500 word statement that briefly describes their education, interests in data issues, and career goals to datashare-inquiry-l@list.indiana.edu. Candidates are encouraged to browse the RDA website https://rd-alliance.org/ and pages of interest and working groups to identify relevant topics and mutual interests.

Important dates:
April 16, 2015 – Fellowship applications are due
May 1, 2015 – Award notifications
June 18-19, 2015 – Fellowship begins with the orientation workshop in Bloomington, IN

RDA Data Share, funded by the Alfred P. Sloan Foundation under award G-2014-13746, engages students and early career researchers in the Research Data Alliance. This engagement builds on foundational infrastructure funded by the National Science Foundation grant # ACI-1349002.

Feb 11, 2015

Institutional analysis of data practices

A short summary of a paper published in JASIST recently: Mayernik, M. S. (2015), Research data and metadata curation as institutional issues. J Assn Inf Sci Tec. doi:10.1002/asi.23425.

The paper begins by noticing a mismatch between the findings of two studies on the data practices in climate science. One of them (a report commissioned by the UK Research Information Network RIN) described the level of data sharing in climate science as low and the other (the book by Edwards "A vast machine...") argued that data sharing was a strong and common norm in climate science. Which one is true? Or, could it be that both studies are correct and climate science includes both the high and the low data sharing levels?

Data practices are institutionalized within a number of social systems, including formal organizations (such as universities and research centers), rules and sanctions (such as funding agency requirements and professional guidelines), and the norms of modern Western science, so the case study analysis in this paper is grounded in the institutional framework that has five characteristics: (a) norms and symbols, (b) intermediaries, (c) routines, (d) standards, and (e) material objects. Norms are largely associated with the norms of science (Merton and later work), symbols are logos and other visible signs of collective identity, but also terminological choices and metaphors. Intermediaries are individuals or collectives who connect resources and facilitate relationships. Routines are frequently repeated patterns of action and interaction, for example, meal or socializing routines. Standards are rules and specifications that define how information can be presented, organized, and transferred. Material objects are ... material objects.

The case studies are comparisons between data practices at the Center for Embedded Networked Sensing (CENS) and the Long Term Ecological Research (LTER) network and between the University Corporation for Atmospheric Research (UCAR)and the National Center for Atmospheric Research (NCAR).

Although there are some interesting observations in these case studies, it seemed that the first, conceptual part of the paper was stronger than the second. The five characteristics of the institutional framework were applied rather narrowly, without revealing many interconnections and directionality. For example, the standards section focuses on metadata standards and their choice. Are there any other standards relevant to data practices? How does the choice of standards affect norms and what is the role of intermediaries in establishing routines and other aspects of data practices? Another much more important question is: Once we describe the variability of data practices within and across disciplines, what's next? What exactly is the role of each institutional carrier in data practices?