Intervening at the Source: A Data-Centric Approach to Trustworthy AI

Babak Salimi -
Data Science

Date: -
Location: Eurecom

Abstract: AI systems now make decisions in health care, hiring, and public services, and when they fail the cause is usually in the data: biased samples, wrong labels, missing coverage, or content a model absorbed during training and later reproduced. Data is also the one input we can deliberately change and then measure the effect of. This talk presents a data-centric approach to trustworthy AI organized around three questions: which data to keep, how far to trust it, and what a model does with it. I describe a model-agnostic method for valuing training data at scale, which identifies helpful and harmful examples without retraining; a framework rooted in possible-worlds semantics that learns from uncertain or incomplete data and certifies which predictions can be trusted; and targeted interventions for large language models that curb memorized training data, prevent forgetting during fine-tuning, and make agent ensembles more reliable. The common thread is that reliability improves when the intervention targets the actual point of failure in the data rather than treating symptoms in the model. Bio: Babak Salimi is an Assistant Professor at the Halicioglu Data Science Institute and the Department of Computer Science and Engineering at UC San Diego. His research bridges data management and machine learning, focusing on data-centric foundations for trustworthy AI: data valuation and cleaning, learning under data uncertainty, and the reliability and safety of large language models and agentic systems. His work has been recognized with the Best Demonstration Paper Award at VLDB 2018, the Best Paper Award at SIGMOD 2019, the Research Highlight Award at SIGMOD 2020, and an NSF CAREER Award in 2024.