Showing posts with label big data. Show all posts
Showing posts with label big data. Show all posts

Feb 13, 2017

Data quality - a short overview

love your data image

A short overview of data quality definitions and challenges in support of the Love Your Data week #lyd17 (February 13 - 17, 2017). The main theme is "Data Quality" and I was part of preparing daily content. Many of the aspects discussed below are elaborated through stories and resources for each day on the LYD website: Defining Data QualityDocumenting, Describing, DefiningGood Data ExamplesFinding the Right DataRescuing Unloved Data

Data quality is the degree to which data meets the purposes and requirements of its use. Good data, therefore, is the data that can be used for the task at hand even if it has some issues (e.g., missing data, poor metadata, value inconsistencies, etc.) Data that has errors, is hard to retrieve or understand, or has no context or traces of where it came from is generally considered bad.

Numerous attempts to define data quality over the last few decades relied on diverse methodologies and identified multiple dimensions of data or information quality (Price and Shanks, 2005). The importance of quality of data is recognized in business and commercial data warehousing (Fan and Geertz, 2012; Redman, 2001), in government operations (Information Quality Act, 2001) and by international agencies involved in data-intensive activities (IMF, 2001). Many research domains have also developed frameworks to evaluate quality of information, including decision, measurement, test, and estimation theories (Altman, 2012).

Attempts to develop discipline-independent frameworks resulted in several models, including models that define quality as data-related versus system-related (Wand and Wang, 1996), as product and service quality (Khan, Strong and Wang, 2002), as syntactic, semantic and pragmatic dimensions (Price and Shanks, 2005), and as user-oriented and contextual quality (Dedeke, 2000). Despite these many attempts to define discipline-independent data quality frameworks, they have not been widely adopted and more frameworks continue to appear. Several systematic syntheses compared many existing frameworks to only point out the complexity and multidimensionality of data quality (Knight and Burn, 2005; Battini et al, 2009).

Data / information quality research grapples with the following fundamental questions (Ge and Helfert, 2007):

  • how to assess quality
  • how to manage quality
  • what impact quality has on organization
The multitude of definitions, frameworks, and contexts in which data quality is used demonstrate that making data quality a useful paradigm is a persisting challenge that can benefit from establishing a dynamic network of researchers and practitioners in the area of data quality and from developing a framework that would be general and yet flexible enough to accommodate highly specific attributes and measurements from particular domains.

data quality attributes
3 Reasons Why Data Quality Should Be Your Top Priority This Year
Each dimension of data quality, such as completeness, accuracy, timeliness, or consistency creates challenges for data quality.

Completeness, for example, is the extent to which data is not missing or is of sufficient breadth and depth for the task at hand (Khan, Strong and Wang, 2002). If a dataset has missing values due to non-response or errors in processing, there is a danger that representativeness of the sample is reduced and thus inferences about the population are distorted. If the dataset contains inaccurate or outdated values, problems with modeling and inference arise.

As data goes through many stages during the research lifecycle, from its collection / acquisition to transformation and modeling to publication, each of the stages creates additional challenges for maintaining integrity and quality of data. In one of the most recent attempts to discredit climate change studies, for example, the authors of the study were blamed for not following the NOAA Climate Data Record policies that maintain standards for documentation, software processing, and access and preservation (Letzter, 2017). This brings out possibilities for further studies:
  • How does non-compliance with policies undermine the quality of data?
  • What role does scientific community consensus play in establishing the quality of data?
  • Should quality management efforts focus on improving the quality of data at every stage or the quality of procedures so that possibilities of errors are minimized? 
Another aspect of data quality that complicates formalized treatment of initial dimensions is that data is often heterogeneous and can be applied in varied contexts. As has been pointed above, data quality frameworks and approaches are being developed in business, government, and research contexts and quality solutions have to consider structured, semi-structured, and unstructured data and their combinations. Most of the previous data quality research focused on structured or semi-structured data. Additionally, spatial, temporal, and volume dimensions of data contribute to quality assessment and management.

Madnick et al. (2009) identify three approaches to possible solutions to data quality: technical or database approach, computer science / information technology (IT) approach, and digital curation approach. Technical solutions include data integration and warehousing, conceptual modeling and architecture, monitoring and cleaning, provenance tracking and probabilistic modeling. Computer / IT solutions include assessments of data quality, organizational studies, studies of data networks and flows, establishment of protocols and standards, and others. Digital curation includes paying attention to metadata, long-term preservation, and provenance.

Most likely, some combination of the above is the best approach. Quality depends on how data was collected as well as on how it was subsequently stored, curated, and made available to others. Data quality is a responsibility that is shared between data providers, data curators and data consumers. While data providers can ensure the quality of their individual datasets, curators help with consistency, coverage and metadata. Maintaining current and consistent metadata across copies and systems also benefits contributions from those who intend to re-use the data. Data and software documentation is another aspect of data quality that cannot be solved technically and needs a combination of organizational / information science solutions.

References and further reading:

Sep 2, 2016

Workshop: Data Quality in Era of Big Data

The center where I work organizes a workshop of possible interest to many who work with data. Scholarships are available.

Data Quality in Era of Big Data
Bloomington, Indiana
28-29 September 2016


Throughout the history of modern scholarship, the exchange of scholarly data was undertaken through personal interactions among scholars or through highly curated data archives. In either case, implicit or explicit provenance mechanisms gave a relatively high degree of insurance of the quality of the data. However, the ubiquity of the web and mobile digital culture has produced disruptive new forms of data. We need to ask ourselves what we know about the data and what we can trust. Failure to answer these questions endangers the integrity of the science produced from these data.

The workshop will examine questions of quality:
·        Citizen science data
·        Health records
·        Integrity
·        Completeness; boundary conditions
·        Instrument quality
·        Data trustworthiness
·        Data provenance
·        Trust in data publishing

The 2 day workshop begins with a half day of tutorials.  The main workshop begins early afternoon on 28 September and continuing to noon on the 29 September.  With sufficient interest, there may be another training session following noon conclusion of the main workshop on 29 September.

Early Career Travel Funds:
Travel funds are available for early career researchers, scholars, and practitioners http://d2i.indiana.edu/mbdh/#scholarships

Important Dates:
·        Workshop:  Sep 28-29, 2016
·        Deadline for requesting early career travel funds:  Sep 9, 2016 midnight EDT
·        Notification of travel funding:  Sep 13, 2016
·        Registration deadline:  Sep 19, 2016
 ​
Organizing Committee:
General Chairs:  Beth Plale, Indiana University

Program Committee
Carl Lagoze, University of Michigan, chair
Devan Donaldson, Indiana University
H.V. Jagadish, University of Michigan
Xiaozhong Liu, Indiana University
Jill Minor, Indiana University
Val Pentchev, Indiana University
Hridesh Rajan, Iowa State University

Early Career Chairs
Devan Donaldson, Indiana University 
Xiaozhong Liu, Indiana University

Local Arrangements Chair
Jill Minor, Indiana University

Apr 10, 2016

Big data analytics overview

The paper Beyond the hype: Big data concepts, methods, and analytics (2015, International Journal of Information Management, Vol. 35, N 2, pp. 137–144) reviews definitions and analytics techniques of big data and discusses some future developments. The article begins with a chart showing an explosion of publications in the Proquest database, which is quite similar to the chart in our JASIST publication "Big data, bigger dilemmas". Both charts show that 2013 was the year when the term "big data" gained popularity:
"Beyond the hype ..."
"Big data, bigger dilemmas..."
The paper cites Diebold's paper "A personal perspective on the origin(s) and development of “big data”: The phenomenon, the term, and the discipline" to describe the origin of the term "big data":
"... the term “big data … probably originated in lunch-table conversations at Silicon Graphics Inc. (SGI) in the mid-1990s, in which John Mashey figured prominently".
After summarizing aspects of big data that were discussed many times elsewhere (volume, velocity, variety, veracity, etc.), the article provides a useful summary of the types of analytics that are common in big data research:
  1. Text analytics
    • Information extraction
      • Entity recognition
      • Relation extraction
  2. Text summarization
    • Extractive (location and frequency of text units)
    • Abstractive (semantic information)
  3. Question answering
  4. Audio (speech) analytics
    • Transcript-based approach (large-vocabulary continuous speech recognition, LVCSR)
    • Phonetic-based approach
  5. Video analytics
  6. Social media analytics
    • Content-based analytics
    • Structure-based analytics
      • Community detection
      • Social influence analysis
      • Link prediction
  7. Predictive analytics
In conclusion the paper argues for new techniques that would address such issues as the irrelevance of statistical significance, heterogeneity and computational efficiency in big data.

Jan 5, 2015

Data politics and the dark sides of data

A bit lengthy post about the dark sides of data discusses whether data and its vast amounts and ubiquitous collection mechanisms help to "tell the truth to power", i.e., to change the world for the better. Will picking up the traces and revealing wrongdoing fix the world? Most likely not, because it's not clear whether we will care or do anything because of data. Here is a great quote:

... lawyers cannot fix human rights abuses, scientists cannot fix global warming, whistle-blowers cannot fix secret services and activists cannot fix politics, and nobody really knows how global finances work - regardless of the data they have at hand...

The framework of countering power and problems with data need to be revised. It's not about quantity or even quality of data, but about using whatever little we know to address not only our understandings (i.e., our rational capacities), but also our feelings and beliefs. Here is what gets in the way (those dark sides that everyone should think about):

  • corporate infrastructure for data and its cultures that creates an illusion of free and neutral services (i.e., services that have no monetary and no political cost)
  • non-transparency of most digital data and lack of control over it, which prevents us from copying, deleting, or processing our own data
  • de-politicizing of digital data or constructing data as fuel for innovation and services rather than a ground for moral, ethical, and political decisions

The post is rather pessimistic, but changes do not happen at once, so we should probably keep trying.

May 9, 2014

Big data report from the White House

Another big data review, this time from the White House - "Big Data: Seizing Opportunities, Preserving Values" (pdf). The report explains what big data is (large, diverse, complex, longitudinal, distributed, making possible unexpected discoveries and creating an asymmetry of power between those who hold the data and those who intentionally or inadvertently supply it) and describes implications of big data for public and private sectors. In addition to many known and less known examples of how big data can be good or bad, the report provides initial thoughts on recommendations for big data governance. It divided its approach to policy framework into four overlapping core areas:

1. Big data and citizens - improve public services while preventing the government from accruing unlimited power by using increased surveillance, algorithmic profiling, and metadata tracking.

2. Big data and consumers - reduce cost of commercial services and personalize them while mitigating security breaches and risks of discrimination based on consumer profiles and lack of consumer awareness and data transparency.

3. Big data and discrimination - do less harm and prevent discriminatory uses of identification and re-identification techniques.

4. Big data and privacy - get used to less privacy while reconsidering the notice and consent framework.

In the concluding section the report had the following recommendations:

  • Advance the Consumer Privacy Bill of Rights.
  • Pass National Data Breach Legislation.
  • Extend Privacy Protections to non-U.S. Persons.
  • Ensure Data Collected on Students in School is Used for Educational Purposes.
  • Expand Technical Expertise to Stop Discrimination.
  • Amend the Electronic Communications Privacy Act.

It's a thorough report and is definitely worth a read, but similarly to my and my colleagues big data review (pre-print), it's just the beginning of studying implications and governance of big data.