Paper published in a book (Scientific congresses, symposiums and conference proceedings)
DataPrism: Exposing Disconnect between Data and Systems
Galhotra, Sainyam; Fariha, Anna; DE PAULA LOURENCO, Raoni et al.
2022In SIGMOD 2022 - Proceedings of the 2022 International Conference on Management of Data
Peer reviewed
 

Files


Full Text
3514221.3517864.pdf
Author postprint (2.56 MB)
Download

All documents in ORBilu are protected by a user license.

Send to



Details



Keywords :
causal testing; data profiles; debugging; root-cause identification; Causal testing; Central component; Data driven; Data profiles; Debugging; Driven system; Health monitoring system; Property; Root cause; Root cause identification; Software; Information Systems
Abstract :
[en] As data is a central component of many modern systems, the cause of a system malfunction may reside in the data, and, specifically, particular properties of data. E.g., a health-monitoring system that is designed under the assumption that weight is reported in lbs will malfunction when encountering weight reported in kilograms. Like software debugging, which aims to find bugs in the source code or runtime conditions, our goal is to debug data to identify potential sources of disconnect between the assumptions about some data and systems that operate on that data. We propose DataPrism, a framework to identify data properties (profiles) that are the root causes of performance degradation or failure of a data-driven system. Such identification is necessary to repair data and resolve the disconnect between data and systems. Our technique is based on causal reasoning through interventions: when a system malfunctions for a dataset, DataPrism alters the data profiles and observes changes in the system's behavior due to the alteration. Unlike statistical observational analysis that reports mere correlations, DataPrism reports causally verified root causes-in terms of data profiles-of the system malfunction. We empirically evaluate DataPrism on seven real-world and several synthetic data-driven systems that fail on certain datasets due to a diverse set of reasons. In all cases, DataPrism identifies the root causes precisely while requiring orders of magnitude fewer interventions than prior techniques.
Disciplines :
Computer science
Author, co-author :
Galhotra, Sainyam;  University of Chicago, Chicago, United States
Fariha, Anna;  Microsoft, Seattle, United States
DE PAULA LOURENCO, Raoni  ;  University of Luxembourg > Interdisciplinary Centre for Security, Reliability and Trust (SNT) > SerVal ; NYU - New York University [US-NY]
Freire, Juliana;  New York University, New York, United States
Meliou, Alexandra;  University of Massachusetts Amherst, Amherst, United States
Srivastava, Divesh;  At&t Chief Data Office, Bedminster, United States
External co-authors :
yes
Language :
English
Title :
DataPrism: Exposing Disconnect between Data and Systems
Publication date :
10 June 2022
Event name :
Proceedings of the 2022 International Conference on Management of Data
Event place :
Philladelphia, Usa
Event date :
12-06-2022 => 17-06-2022
Main work title :
SIGMOD 2022 - Proceedings of the 2022 International Conference on Management of Data
Publisher :
Association for Computing Machinery
ISBN/EAN :
978-1-4503-9249-5
Peer reviewed :
Peer reviewed
Funders :
ACM SIGMOD
Available on ORBilu :
since 22 November 2023

Statistics


Number of views
53 (1 by Unilu)
Number of downloads
99 (0 by Unilu)

Scopus citations®
 
9
Scopus citations®
without self-citations
6
OpenCitations
 
2
OpenAlex citations
 
8

Bibliography


Similar publications



Contact ORBilu