publications
publications by categories in reversed chronological order. generated by jekyll-scholar.
2026
- An Economic Framework for Generative Engines: Advertising or Subscription?Luyang Zhang , Cathy Jiao , Beibei Li , and Chenyan XiongPreprint, 2026
Generative Engines (GEs) such as ChatGPT and Google’s AI Overviews are rapidly reshaping search economics by delivering synthesized responses that allow users to bypass third-party websites, cutting those sites’ advertising revenue. Yet this shift also leaves GEs facing their own monetization problem: whether to insert ads into synthesized responses or keep them ad-free to drive subscription conversions. In this paper, we introduce a dynamic framework to study this problem, which captures how query-level design choices shape user engagement, retention, and subscription conversion over time. Using this framework, we show that the optimal policy follows a cutoff rule: ads should only be shown to users only when the immediate ad payoff exceeds the long-term value of providing ad-free responses. This cutoff shifts toward with-ad responses when i) ad revenue is high or ii) users are less sensitive to ads, and toward ad-free responses when iii) subscription conversion becomes relatively more valuable. In addition, the presence of rival GEs shifts the optimal policy further toward ad-free responses, as ad-heavy monetization becomes less sustainable when users can freely switch to alternatives. Our findings reveal incentives for real-life generative engine providers to adopt designs that enhance user experience and long-term sustainability.
- Rigorous Interpretation Is a Form of EvaluationIsabelle Lee , Emmy Liu , Cathy Jiao , Brihi Joshi , Dani Yogatama , Fazl Barez , and Michael SaxonIn Proceedings of the Workshop on Evaluating Evaluations (EvalEval), Jul 2026
Current machine learning models are evaluated through behavioral snapshots, with benchmark accuracies, win rates and outcome-based metrics. Model explanations and evaluations, however, are fundamentally intertwined: understanding why a model produces a behavior can be as important as measuring what it produces. If we trusted interpretability, we argue that it can serve not merely as diagnostics but as a richer and more principled form of model evaluation beyond surface-level performance metrics. We explore three ways interpretability can function evaluatively: (1) fixing problems by identifying the root causes of unwanted behavior, (2) detecting subtly faulty mechanisms that invalidate model outputs, and (3) predicting potential issues before they arise by fully understanding the model’s weaknesses. To fulfill its evaluative potential, we argue that interpretability methods must generate claims that are falsifiable, reproducible, and predictive—that is, interpretability must meet scientific standards.
- Effective Synthetic Data Curation Requires Group-Level SignalsCathy Jiao and Chenyan XiongarXiv preprint, Jul 2026
Synthetic data now is essential to LLM training, used to strengthen advanced capabilities such as autonomous and long-horizon task execution. Yet recent work shows that training on it at scale can degrade model generation, making it important to decide what synthetic data is worth training on. While current data curation practices do so with individual-level signals (i.e., estimates of each data sample’s training utility in isolation), across pre-training and post-training settings we show that this is insufficient for synthetic data, and that group-level signals (i.e., estimates of utility that account for interactions among data samples) are necessary for effective data curation. First, we show that individual-level signals are blind to how samples jointly affect training: synthetic datasets with different compositions can be indistinguishable under individual-level influence yet differ sharply under group-level influence, and curating by the latter yields better downstream performance, particularly in generative capability. Second, we find that group-level signals matter more as training pipelines become increasingly synthetic: among widely used data curation methods, only those incorporating them improve over baseline, with gains increasing when weights capturing relations among samples are amplified. Finally, we translate these findings into practice – for model developers under a compute budget, we offer a cheap diagnostic that prioritizes which groups of synthetic data most need group-level estimation, recovering much of the benefit of full group-level scoring at a fraction of the compute cost.
- Efficient Dataset Selection for Continual Adaptation of Generative RecommendersCathy Jiao , Juan Elenter , Praveen Ravichandran , Bernd Huber , Joseph Cauteruccio , Todd Wasson , Timothy Heath , Chenyan Xiong , Mounia Lalmas , and Paul BennettIn ICLR CAO Workshop, Jul 2026
Selected for an oral presentation at the ICLR 2026 CAO Workshop (6 of 76 papers, top 8%).
Recommendation systems must continuously adapt to evolving user behavior, yet the volume of data generated in large-scale streaming environments makes frequent full retraining impractical. This work investigates how targeted data selection can mitigate performance degradation caused by temporal distributional drift while maintaining scalability. We evaluate a range of representation choices and sampling strategies for curating small but informative subsets of user interaction data. Our results demonstrate that gradient-based representations, coupled with distribution-matching, improve downstream model performance, achieving training efficiency gains while preserving robustness to drift. These findings highlight data curation as a practical mechanism for scalable monitoring and adaptive model updates in production-scale recommendation systems.
2025
- A survey of data attribution: Methods, applications, and evaluation in the era of generative aiJunwei Deng , Yuzheng Hu , Pingbang Hu , Ting-Wei Li , Shixuan Liu , Jiachen T Wang , Dan Ley , Qirun Dai , Benhao Huang , Jin Huang , Cathy Jiao , Hoang Anh Just , Yijun Pan , Jingyan Shen , Yiwen Tu , Weiyi Wang , Xinhe Wang , Shichang Zhang , Shiyuan Zhang , Ruoxi Jia , Himabindu Lakkaraju , Hao Peng , Weijing Tang , Chenyan Xiong , Jieyu Zhao , Hanghang Tong , Han Zhao , and Jiaqi W MaPreprint, Jul 2025
Training data is the fuel of modern artificial intelligence (AI), fundamentally shaping the capabilities, limitations, and biases of AI systems. The emergence of large-scale generative models has elevated the importance of understanding how data influences their behaviors, bringing the field of data attribution to the forefront. This survey provides a comprehensive overview of data attribution, covering its methods, applications, and evaluation protocols, with a particular emphasis on the challenges and opportunities arising in the era of generative AI. We start by introducing a conceptual framework for attribution centered on three core questions: what to attribute (model behaviors), attribute to what (training entities), and how to attribute (influence measures). Within this framework, we systematically review major attribution approaches, including those based on influence functions, weighted marginal contributions, training dynamics, and simulators. We then examine key applications of data attribution, such as data selection, fact tracing, adversarial attacks and defenses, and the emerging data economy. Finally, we critically assess common evaluation criteria, including the quality of counterfactual predictions, utility in downstream tasks, and computational efficiency. We conclude with a forward-looking perspective on the future of data attribution, highlighting key open challenges and promising directions for future research.
- DATE-LM: Benchmarking Data Attribution Evaluation for Large Language ModelsCathy Jiao* , Yijun Pan* , Emily Xiao* , Daisy Sheng , Niket Jain , Hanzhang Zhao , Ishita Dasgupta , Jiaqi W. Ma , and Chenyan XiongNeurIPS, Jul 2025
Data attribution methods quantify the influence of training data on model outputs and are becoming increasingly relevant for a wide range of LLM research and applications, including dataset curation, model interpretability, data valuation. However, there remain critical gaps in systematic LLM-centric evaluation of data attribution methods. To this end, we introduce DATE-LM (Data Attribution Evaluation in Language Models), a unified benchmark for evaluating data attribution methods through real-world LLM applications. DATE-LM measures attribution quality through three key tasks – training data selection, toxicity/bias filtering, and factual attribution. Our benchmark is designed for ease of use, enabling researchers to configure and run large-scale evaluations across diverse tasks and LLM architectures. Furthermore, we use DATE-LM to conduct a large-scale evaluation of existing data attribution methods. Our findings show that no single method dominates across all tasks, data attribution methods have trade-offs with simpler baselines, and method performance is sensitive to task-specific evaluation design. Finally, we release a public leaderboard for quick comparison of methods and to facilitate community engagement. We hope DATE-LM serves as a foundation for future data attribution research in LLMs.
- Fairshare Data Pricing via Data Valuation for Large Language ModelsLuyang Zhang* , Cathy Jiao* , Beibei Li , and Chenyan XiongNeurIPS, Jul 2025
Training data is a pivotal resource for building large language models (LLMs), but unfair pricing in data markets poses a serious challenge for both data buyers (e.g., LLM builders) and sellers (e.g., human annotators), which discourages market participation, reducing data quantity and quality. In this paper, we propose a fairshare pricing framework that sets training data prices using data valuation methods to quantify their contribution to LLMs. In our framework, buyers make purchasing decisions using data valuation and sellers set prices to maximize their profits based on the anticipated buyer purchases. We theoretically show that pricing derived from our framework is tightly linked to data valuation and buyers’ budget, optimal for both buyers and sellers. Through market simulations using current LLMs and datasets (math problems, medical diagnosis, and physical reasoning), we show that our framework is fairshare for buyers by ensuring their purchased data is reflective of model training value, leading to higher LLM task performances per-dollar spent on data, and fairshare for sellers by ensuring they sell their data at optimal prices. Our framework lays the foundation for future research on equitable and sustainable data markets for large-scale AI.
- On the Feasibility of In-Context Probing for Data AttributionCathy Jiao , Gary Gao , Aditi Raghunathan , and Chenyan XiongNAACL, Jul 2025
Data attribution methods are used to measure the contribution of training data towards model outputs, and have several important applications in areas such as dataset curation and model interpretability. However, many standard data attribution methods, such as influence functions, utilize model gradients and are computationally expensive. In our paper, we show in-context probing (ICP) – prompting a LLM – can serve as a fast proxy for gradient-based data attribution for data selection under conditions contingent on data similarity. We study this connection empirically on standard NLP tasks, and show that ICP and gradient-based data attribution are well-correlated in identifying influential training data for tasks that share similar task type and content as the training data. Additionally, fine-tuning models on influential data selected by both methods achieves comparable downstream performance, further emphasizing their similarities. We also examine the connection between ICP and gradient-based data attribution using synthetic data on linear regression tasks. Our synthetic data experiments show similar results with those from NLP tasks, suggesting that this connection can be isolated in simpler settings, which offers a pathway to bridging their differences.
2024
- Examining Prosody in Spoken Navigation Instructions for People with DisabilitiesCathy Jiao , Aaron Steinfeld , and Maxine EskenaziIn Proceedings of the Third Workshop on Bridging Human–Computer Interaction and Natural Language Processing, Jun 2024
The introduction of conversational systems have made synthesized speech technologies common tools for daily activities. However, not all synthetic speech systems are designed with the needs of people with disabilities in mind. This paper describes a study in which 198 people – 80 participants with self-reported disabilities and 118 participants without – were recruited to listen to navigation instructions from a spoken dialogue system with different prosodic features. Results showed that slowing down speech rate aids in participants’ number recall, but not in noun recall. From our results, we provide suggestions for developers for building accessible synthetic speech systems.
2023
- Understanding the Effectiveness of Very Large Language Models on Dialog EvaluationJessica Huynh , Cathy Jiao , Prakhar Gupta , Shikib Mehri , Payal Bajaj , Vishrav Chaudhary , and Maxine EskenaziIWSDS, Jun 2023
Language models have steadily increased in size over the past few years. They achieve a high level of performance on various natural language processing (NLP) tasks such as question answering and summarization. Large language models (LLMs) have been used for generation and can now output human-like text. Due to this, there are other downstream tasks in the realm of dialog that can now harness the LLMs’ language understanding capabilities. Dialog evaluation is one task that this paper will explore. It concentrates on prompting with LLMs: BLOOM, OPT, GPT-3, Flan-T5, InstructDial and TNLGv2. The paper shows that the choice of datasets used for training a model contributes to how well it performs on a task as well as on how the prompt should be structured. Specifically, the more diverse and relevant the group of datasets that a model is trained on, the better dialog evaluation performs. This paper also investigates how the number of examples in the prompt and the type of example selection used affect the model’s performance.
2022
- The DialPort toolsJessica Huynh , Shikib Mehri , Cathy Jiao , and Maxine EskenaziIn Proceedings of the 23rd Annual Meeting of the Special Interest Group on Discourse and Dialogue, Sep 2022
The DialPort project (\urlhttp://dialport.org/), funded by the National Science Foundation (NSF), covers a group of tools and services that aim at fulfilling the needs of the dialog research community. Over the course of six years, several offerings have been created, including the DialPort Portal and DialCrowd. This paper describes these contributions, which will be demoed at SIGDIAL, including implementation, prior studies, corresponding discoveries, and the locations at which the tools will remain freely available to the community going forward.
- Improving Zero and Few-shot Generalization in Dialogue through Instruction TuningPrakhar Gupta , Cathy Jiao , Yi-Ting Yeh , Shikib Mehri , Maxine Eskenazi , and Jeffrey P. BighamEMNLP, Sep 2022
Instruction tuning is an emergent paradigm in NLP wherein natural language instructions are leveraged with language models to induce zero-shot performance on unseen tasks. Instructions have been shown to enable good performance on unseen tasks and datasets in both large and small language models. Dialogue is an especially interesting area to explore instruction tuning because dialogue systems perform multiple kinds of tasks related to language (e.g., natural language understanding and generation, domain-specific interaction), yet instruction tuning has not been systematically explored for dialogue-related tasks. We introduce InstructDial, an instruction tuning framework for dialogue, which consists of a repository of 48 diverse dialogue tasks in a unified text-to-text format created from 59 openly available dialogue datasets. Next, we explore cross-task generalization ability on models tuned on InstructDial across diverse dialogue tasks. Our analysis reveals that InstructDial enables good zero-shot performance on unseen datasets and tasks such as dialogue evaluation and intent detection, and even better performance in a few-shot setting. To ensure that models adhere to instructions, we introduce novel meta-tasks. We establish benchmark zero-shot and few-shot performance of models trained using the proposed framework on multiple dialogue tasks.
- Improving compositional generalization for multi-step quantitative reasoning in question answeringArmineh Nourbakhsh , Cathy Jiao , Sameena Shah , and Carolyn RoséEMNLP, Sep 2022
Quantitative reasoning is an important aspect of question answering, especially when numeric and verbal cues interact to indicate sophisticated, multi-step programs. In this paper, we demonstrate how modeling the compositional nature of quantitative text can enhance the performance and robustness of QA models, allowing them to capture arithmetic logic that is expressed verbally. Borrowing from the literature on semantic parsing, we propose a method that encourages the QA models to adjust their attention patterns and capture input/output alignments that are meaningful to the reasoning task. We show how this strategy improves program accuracy and renders the models more robust against overfitting as the number of reasoning steps grows. Our approach is designed as a standalone module which can be prepended to many existing models and trained in an end-to-end fashion without the need for additional supervisory signal. As part of this exercise, we also create a unified dataset building on four previously released numerical QA datasets over tabular data.