A Decision Theoretic Framework for Measuring AI Reliance
Humans frequently make decisions with the aid of artificially intelligent (AI) systems. A common pattern is for the AI to recommend an action to the human who retains control over the final decision. Researchers have identified ensuring that a human has appropriate reliance on an AI as a critical component of achieving complementary performance. We argue that the current definition of appropriate reliance used in such research lacks formal statistical grounding and can lead to contradictions. We propose a formal definition of reliance, based on statistical decision theory, which separates the concepts of reliance as the probability the decision-maker follows the AI’s recommendation from challenges a human may face in differentiating the signals and forming accurate beliefs about the situation. Our definition gives rise to a framework that can be used to guide the design and interpretation of studies on human-AI complementarity and reliance. Using recent AI-advised decision-making studies from literature, we demonstrate how our framework can be used to separate the loss due to mis-reliance from the loss due to not accurately differentiating the signals. We evaluate these losses by comparing a baseline and a benchmark for complementary performance defined by the expected payoff achieved by a rational decision-maker facing the same decision task as the behavioral decision-makers.
Validating LLM simulations as behavioral evidence
A growing literature uses large language models (LLMs) as synthetic participants to generate cost-effective and nearly instantaneous responses in social science experiments. However, there is limited guidance on when such simulations support valid inference about human behavior. We contrast two strategies for obtaining valid estimates of causal effects and clarify the assumptions under which each is suitable for exploratory versus confirmatory research. Heuristic approaches seek to establish that simulated and observed human behavior are interchangeable through prompt engineering, model fine-tuning, and other repair strategies designed to reduce LLM-induced inaccuracies. While useful for many exploratory tasks, heuristic approaches lack the formal statistical guarantees typically required for confirmatory research. In contrast, statistical calibration combines auxiliary human data with statistical adjustments to account for discrepancies between observed and simulated responses. Under explicit assumptions, statistical calibration preserves validity and provides more precise estimates of causal effects at lower cost than experiments that rely solely on human participants. Yet the potential of both approaches depends on how well LLMs approximate the relevant populations. We consider what opportunities are overlooked when researchers focus myopically on substituting LLMs for human participants in a study.
Explaining and Improving Information Complementarities in Multi-Agent Decision-making
Multiple agents are increasingly combined to make decisions with the expectation of achieving complementary performance, where the decisions they make together outperform those made individually. However, knowing how to improve the performance of collaborating agents requires knowing what information and strategies each agent employs. With a focus on human-AI pairings, we contribute a decision-theoretic framework for characterizing the value of information. By defining complementary information, our approach identifies opportunities for agents to better exploit available information in AI-assisted decision workflows. We present a novel explanation technique (ILIV-SHAP) that adapts Shapley value explanations to highlight human-complementing information. We validate the effectiveness of our framework and ILIV-SHAP through a study of human-AI decision-making and demonstrate the framework on examples from chest X-ray diagnosis and deepfake detection. We find that presenting ILIV-SHAP with AI predictions leads to reliably greater reductions in human decision-maker’s error over non-AI assisted decisions more than vanilla SHAP.
Hullman, J., D. Broska, H. Sun, and A. Shaw. 2026. This human study did not involve human subjects: Validating LLM simulations as behavioral evidence. arXiv:2602.15785.
Guo, Z., B. Ustun, and J. Hullman. 2026. Explanations are a means to an end: Decision theoretic explanation evaluation. Proceedings of the 43rd International Conference on Machine Learning (ICML).
Guo, Z., Y. Wu, J. Hartline, and J. Hullman. 2026. Explaining and improving information complementarities in multi-agent decision-making. Proceedings of the Fourteenth International Conference on Learning Representations (ICLR).
Musslick, S., L. Bartlett, S. Chandramouli, M. Dubova, F. Gobet, T. Griffiths, J. Hullman, et al. 2025. Automating the practice of science: Opportunities, challenges, and implications. Proceedings of the National Academy of Sciences 122(5): e2401238121.
Hullman, J., A. Kale, and J. Hartline. 2025. Underspecified human decision experiments considered harmful. Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, 1–15.
Guo, Z., Y. Wu, J. Hartline, and J. Hullman. 2024. A decision theoretic framework for measuring AI reliance. Proceedings of the ACM Conference on Fairness, Accountability, and Transparency in Artificial Intelligence 221–36.
Subramonyam, H. and J. Hullman. 2024. Are we closing the loop yet? Gaps in the generalizability of VIS4ML research. IEEE Transactions on Visualization and Computer Graphics 30(1): 1–11.
Gelman, A., J. Hullman, and L. Kennedy. 2024. Casual quartets: Different ways to achieve the same average treatment effect. The American Statistician 78(3): 267–72.
Zhang, D., J. Hartline, and J. Hullman. 2024. Designing shared information displays for agents of varying strategic sophistication. Proceedings of the ACM on Human-Computer Interaction 8(CSCW1): 1–34.
Zhang, D., A. Chatzimparmpas, N. Kamali, and J. Hullman. 2024. Evaluating the utility of conformal prediction sets for AI-advised image labeling. Proceedings of the ACM Conference on Computer-Human Interaction (CHI) 1–19.
Gelman, A., J. Hullman, and L. Kennedy. 2023. Casual quartets: Different ways to achieve the same average treatment effect. The American Statistician 78(3): 267–72.
Wu, Y., Z. Guo, J. Hartline, and J. Hullman. 2023. The rational agent benchmark for data visualization. IEEE Transactions on Visualization and Computer Graphics 30(1): 338–47.
Nanayakkara, P., J. Bater, X. Hu, J. Hullman, and J. Rogers. 2022. Visualizing privacy-utility trade-offs in differentially private data releases. Proceedings of Privacy Enhancing Technologies (2): 601–18.
Hullman, J. and A. Gelman. 2021. Designing for interactive exploratory data analysis requires theories of graphical inference. Harvard Data Science Review 3(3).