Publications
Lab authors are bolded below.
2026
Empirical Evidence and Analysis of a Critical Pitfall in Reward Learning from Human Feedback
Taha Shaheen, S. West, and Yu Zhang
Reward learning via human feedback is a crucial capability for beneficial AI. Current methods are built on decision-making theories that assume a matched dynamics model between the learning agent and the feedback provider. However, humans often form imperfect internal dynamics models, and their feedback reflects these misconceptions. While this relationship has long been hypothesised, its manifestation in sequential decision-making remains largely an assumption. Our work provides the first comprehensive empirical investigation of this relationship through a randomized controlled trial (N=211). We followed a two-stage design where we first initialized the participants' understanding of the dynamics in a grid-world navigation domain and then manipulated it using text-based instructions. Causal mediation analysis revealed that humans' internal models play a mediating role in feedback behaviour. We show that this relationship is invariant across visual contexts and is robust to three common feedback types: pairwise preferences, trajectory corrections, and off-switch interventions. These findings confirm a critical limitation of current reward learning methods and establish the missing psychological foundation for approaches that incorporate dynamics understanding.
Safe Explicable Policy Search
Akkamahadevi Hanni, J. Montano, and Yu Zhang
Assigning Multi-Robot Tasks to Multitasking Robots
Winston Smith and Yu Zhang
2025
Provably Efficient Reward Transfer in Reinforcement Learning with Discrete Markov Decision Processes
Kevin Jatin Vora and Yu Zhang
In this paper, we propose a new solution to reward adaptation (RA) in reinforcement learning, where the agent adapts to a target reward function based on one or more existing source behaviors learned a priori under the same domain dynamics but different reward functions. While learning the target behavior from scratch is possible, it is often inefficient given the available source behaviors. Our work introduces a new approach to RA through the manipulation of Q-functions. Assuming the target reward function is a known function of the source reward functions, we compute bounds on the Q-function and present an iterative process (akin to value iteration) to tighten these bounds. Such bounds enable action pruning in the target domain before learning even starts. We refer to this method as "Q-Manipulation" (Q-M). The iteration process assumes access to a lite-model, which is easy to provide or learn. We formally prove that Q-M, under discrete domains, does not affect the optimality of the returned policy and show that it is provably efficient in terms of sample complexity in a probabilistic sense. Q-M is evaluated in a variety of synthetic and simulation domains to demonstrate its effectiveness, generalizability, and practicality.
ACoL: From Abstractions to Grounded Languages for Robust Coordination of Task Planning Robots
Yu Zhang
2024
Safe Explicable Planning
Akkamahadevi Hanni, Andrew Boateng, and Yu Zhang
2023
Implicit Projection: Improving Team Situation Awareness for Tacit Human-Robot Interaction via Virtual Shadows
Andrew Boateng, W. Zhang, and Yu Zhang
Max Markov Chain
Yu Zhang and Mitchell Bucklew
From Abstractions to Grounded Languages for Robust Coordination of Task Planning Robots
Yu Zhang
2022
Explicable Policy Search
Ze Gong and Yu Zhang
2021
Generating Active Explicable Plans in Human-Robot Teaming
Akkamahadevi Hanni and Yu Zhang
Order Matters: Generating Progressive Explanations for Planning Tasks in Human-Robot Teaming
Mehrdad Zaker Shahrak, S. Marpally, Akshay Sharma, Ze Gong, and Yu Zhang
Achieving Multi-Tasking Robots in Multi-Robot Tasks
Winston Smith and Yu Zhang
Active Explicable Planning for Human-Robot Teaming
Akkamahadevi Hanni and Yu Zhang
Virtual Shadow Rendering for Maintaining Situation Awareness in Proximal Human-Robot Teaming
Andrew Boateng and Yu Zhang
Explicable Policy Search via Preference-Based Learning under Human Biases
Ze Gong and Yu Zhang
2020
Online Explanation Generation for Human-Robot Teaming
Mehrdad Zaker Shahrak, Ze Gong, Nikhillesh Sadassivam, and Yu Zhang
As AI becomes an integral part of our lives, the development of explainable AI, embodied in the decision-making process of an AI or robotic agent, becomes imperative. For a robotic teammate, the ability to generate explanations to justify its behavior is one of the key requirements of explainable agency. Prior work on explanation generation has been focused on supporting the rationale behind the robot's decision or behavior. These approaches, however, fail to consider the mental demand for understanding the received explanation. In other words, the human teammate is expected to understand an explanation no matter how much information is presented. In this work, we argue that explanations, especially those of a complex nature, should be made in an online fashion during the execution, which helps spread out the information to be explained and thus reduce the mental workload of humans in highly cognitive demanding tasks. However, a challenge here is that the different parts of an explanation may be dependent on each other, which must be taken into account when generating online explanations. To this end, a general formulation of online explanation generation is presented with three variations satisfying different "online" properties. The new explanation generation methods are based on a model reconciliation setting introduced in our prior work. We evaluated our methods both with human subjects in a simulated rover domain, using NASA Task Load Index (TLX), and synthetically with ten different problems across two standard IPC domains. Results strongly suggest that our methods generate explanations that are perceived as less cognitively demanding and much preferred over the baselines and are computationally efficient.
What Is It You Really Want of Me? Generalized Reward Learning with Biased Beliefs about Domain Dynamics
Ze Gong and Yu Zhang
2019
Explicability as Minimizing Distance from Expected Behavior
Anagha Kulkarni, Y. Zha, Tathagata Chakraborti, S. Vadlamudi, Yu Zhang, and Subbarao Kambhampati
In order to have effective human-AI collaboration, it is necessary to address how the AI agent's behavior is being perceived by the humans-in-the-loop. When the agent's task plans are generated without such considerations, they may often demonstrate inexplicable behavior from the human's point of view. This problem may arise due to the human's partial or inaccurate understanding of the agent's planning model. This may have serious implications from increased cognitive load to more serious concerns of safety around a physical agent. In this paper, we address this issue by modeling plan explicability as a function of the distance between a plan that agent makes and the plan that human expects it to make. We learn a regression model for mapping the plan distances to explicability scores of plans and develop an anytime search algorithm that can use this model as a heuristic to come up with progressively explicable plans. We evaluate the effectiveness of our approach in a simulated autonomous car domain and a physical robot domain.
Online Explanation Generation for Human-Robot Teaming
Mehrdad Zaker Shahrak, Ze Gong, Nikhillesh Sadassivam, Akkamahadevi Hanni, and Yu Zhang
As AI becomes an integral part of our lives, the development of explainable AI, embodied in the decision-making process of an AI or robotic agent, becomes imperative. For a robotic teammate, the ability to generate explanations to justify its behavior is one of the key requirements of explainable agency. Prior work on explanation generation has been focused on supporting the rationale behind the robot's decision or behavior. These approaches, however, fail to consider the mental demand for understanding the received explanation. In this work, we argue that explanations, especially those of a complex nature, should be made in an online fashion during the execution, which helps spread out the information to be explained and thus reduce the mental workload of humans in highly cognitive demanding tasks. To this end, a general formulation of online explanation generation is presented with three variations satisfying different "online" properties, based on a model reconciliation setting introduced in our prior work.
2018
Interactive Plan Explicability in Human-Robot Teaming
Mehrdad Zaker Shahrak, A. Sonawane, Ze Gong, and Yu Zhang
Behavior Explanation as Intention Signaling in Human-Robot Teaming
Ze Gong and Yu Zhang
Temporal Spatial Inverse Semantics for Robots Communicating with HumansFinalist for the Best Paper Award in Cognitive Robotics
Ze Gong and Yu Zhang
Explain by Goal Augmentation: Explanation Generation as Inverse Planning
Z. Chen and Yu Zhang
Interactive Plan Explicability in Human-Robot Teaming
Mehrdad Zaker Shahrak and Yu Zhang
Explicability as Minimizing Distance from Expected Behavior
Anagha Kulkarni, Y. Zha, Tathagata Chakraborti, S. Vadlamudi, Yu Zhang, and Subbarao Kambhampati
In order to have effective human-AI collaboration, it is necessary to address how the AI agent's behavior is being perceived by the humans-in-the-loop. When the agent's task plans are generated without such considerations, they may often demonstrate inexplicable behavior from the human's point of view. In this paper, we address this issue by modeling plan explicability as a function of the distance between a plan that agent makes and the plan that human expects it to make. We learn a regression model for mapping the plan distances to explicability scores of plans and develop an anytime search algorithm that can use this model as a heuristic to come up with progressively explicable plans. We evaluate the effectiveness of our approach in a simulated autonomous car domain and a physical robot domain.
Robot Signaling its Intentions in Human-Robot Teaming
Ze Gong and Yu Zhang
2017
Plan Explanations as Model Reconciliation: Moving Beyond Explanation as Soliloquy
Tathagata Chakraborti, Sarath Sreedharan, Yu Zhang, and Subbarao Kambhampati
When AI systems interact with humans in the loop, they are often called on to provide explanations for their plans and behavior. Past work on plan explanations primarily involved the AI system explaining the correctness of its plan and the rationale for its decision in terms of its own model. Such soliloquy is wholly inadequate in most realistic scenarios where the humans have domain and task models that differ significantly from that used by the AI system. We posit that the explanations are best studied in light of these differing models. In particular, we show how explanation can be seen as a "model reconciliation problem" (MRP), where the AI system in effect suggests changes to the human's model, so as to make its plan be optimal with respect to that changed human model. We will study the properties of such explanations, present algorithms for automatically computing them, and evaluate the performance of the algorithms.
Plan Explicability and Predictability for Robot Task Planning
Yu Zhang, Sarath Sreedharan, Anagha Kulkarni, Tathagata Chakraborti, Hankz Hankui Zhuo, and Subbarao Kambhampati
Simultaneous Feature and Body-Part Learning for Real-Time Robot Awareness of Human Behaviors
F. Han, X. Yang, C. Reardon, Yu Zhang, and H. Zhang
Sequence-based Multimodal Apprenticeship Learning For Robot Perception and Decision Making
F. Han, X. Yang, Yu Zhang, and H. Zhang
2016
A Formal Analysis of Required Cooperation in Multi-agent Planning
Yu Zhang, Sarath Sreedharan, and Subbarao Kambhampati
It is well understood that, through cooperation, multiple agents can achieve tasks that are unachievable by a single agent. However, there had been no formal characterization of situations where cooperation is required to achieve a goal, thus warranting the use of multiple agents. We provide such a formal characterization for multi-agent planning problems with sequential action execution. We first show that determining whether there is required cooperation is, in general, intractable even in this limited setting, so we start our analysis with a subset of more restrictive problems where agents are homogeneous. For such problems, we identify two conditions that can cause required cooperation: when neither holds, the problem is single-agent solvable, and otherwise we provide upper bounds on the minimum number of agents required. For the remaining problems with heterogeneous agents, we further divide them into two subsets, and for one of these we propose the concept of a transformer agent to reduce the number of agents that need to be considered, which is used to improve planning performance.
Planning with Resource Conflicts in Human-Robot Cohabitation
Tathagata Chakraborti, Yu Zhang, and Subbarao Kambhampati
Plan Explicability for Robot Task Planning
Yu Zhang, Sarath Sreedharan, Anagha Kulkarni, Tathagata Chakraborti, Hankz Hankui Zhuo, and Subbarao Kambhampati
A Formal Framework for Studying Interaction in Human-Robot Societies
Tathagata Chakraborti, Kartik Talamadupula, Yu Zhang, and Subbarao Kambhampati
2015
A Human Factors Analysis of Proactive Support in Human-robot Teaming
Yu Zhang, Vignesh Narayanan, Tathagata Chakraborti, and Subbarao Kambhampati
It has long been assumed that for effective human-robot teaming, it is desirable for assistive robots to infer the goals and intents of humans and take proactive actions to help them achieve those goals. However, there had not been a systematic evaluation of the accuracy of this claim. On the face of it, there are several ways a proactive robot assistant can in fact reduce the effectiveness of teaming: it can increase the cognitive load of the human teammate by performing actions that are unanticipated, and misinterpretations or delays in goal and intent recognition due to partial observations and limited communication can also reduce performance. In this project, we perform an analysis of human factors on the effectiveness of proactive support in human-robot teaming, evaluated in a simulated urban search and rescue task in which the efficacy of teaming depends not only on individual performance but also on how teammates interact with each other. In this task, the human teammate remotely controls a robot while working with an intelligent robot teammate.
Planning for Serendipity
Tathagata Chakraborti, G. Briggs, Kartik Talamadupula, Yu Zhang, M. Scheutz, D. Smith, and Subbarao Kambhampati
There has been a lot of focus on human-robot cohabitation issues that are often orthogonal to many aspects of human-robot teaming, such as producing socially acceptable robot behaviors and de-conflicting plans of robots and humans in shared environments. An interesting offshoot of these settings that has largely been overlooked is the problem of planning for serendipity: planning for stigmergic collaboration without explicit commitments between agents in cohabitation. In this project, we formalize this notion of planning for serendipity for the first time and provide an integer-programming-based solution. We illustrate the different modes of this planning technique on a typical urban search and rescue scenario, and show a real-life implementation of the ideas on a Nao robot interacting with a human colleague.
DisCoF+: Asynchronous DisCoF with Flexible Decoupling for Cooperative Pathfinding in Distributed Systems
K. Kim, J. Campbell, W. Duong, Yu Zhang, and G. Fainekos
In our prior work, we outlined an approach, named DisCoF, for cooperative pathfinding in distributed systems with limited sensing and communication range. Contrasting to prior works on cooperative pathfinding with completeness guarantees, which often assume the access to global information, DisCoF does not make this assumption. The implication is that at any given time in DisCoF, the robots may not all be aware of each other, which is often the case in distributed systems. As a result, DisCoF represents an inherently online approach since coordination can only be realized in an opportunistic manner between robots that are within each other's sensing and communication range. However, there are a few assumptions made in DisCoF to facilitate a formal analysis, which must be removed to work with distributed multi-robot platforms. In this paper, we present DisCoF+, which extends DisCoF by enabling an asynchronous solution, as well as providing flexible decoupling between robots for performance improvement. We also extend the formal results of DisCoF to DisCoF+. Furthermore, we evaluate our implementation of DisCoF+ and demonstrate a simulation of it running in a distributed multi-robot environment. Finally, we compare DisCoF+ with DisCoF in terms of plan quality and planning performance.
Capability Models and Their Applications in Planning
Yu Zhang, Sarath Sreedharan, and Subbarao Kambhampati
One important challenge for a set of agents to achieve more efficient collaboration is for these agents to maintain proper models of each other. An important aspect of these models is that they are often not provided, and hence must be learned from plan execution traces. As a result, these models of other agents are inherently partial and incomplete. Most existing agent models are based on action modeling and do not naturally allow for incompleteness. We introduce a modeling approach based on the representation of capabilities, which has several unique advantages. First, we show that the structures of capability models can be learned or easily specified, and both model structure and parameter learning are robust to high degrees of incompleteness in plan traces (for example, with only start and end states partially observed). Furthermore, parameter learning can be performed efficiently online via Bayesian learning. As a result, capability models are useful in applications where traditional models are difficult to obtain, or where models must be learned from incomplete plan traces, such as robots learning human models from observations and interactions.
Automated Planning for Peer-to-peer Teaming and its Evaluation in Remote Human-Robot Interaction
Vignesh Narayanan, Yu Zhang, N. Mendoza, and Subbarao Kambhampati
2014
DisCoF: Cooperative Pathfinding in Distributed Systems with Limited Sensing and Communication Range
Yu Zhang, K. Kim, and G. Fainekos
This project addresses the multi-agent pathfinding problem in distributed systems that are subject to limited sensing and communication range. Cooperative pathfinding is typically addressed in one of two ways in the literature: fully coupled approaches consider all robots together and construct plans simultaneously, while decoupled approaches construct plans for only a subset of robots at a time. Decoupled approaches can be much faster, but are often suboptimal and incomplete, and the few decoupled approaches that do achieve completeness typically assume access to global information, which may not be available in distributed robotic systems. We provide a window-based approach to cooperative pathfinding with limited sensing and communication range, called DisCoF. Robots are assumed to be fully decoupled initially, and may gradually increase their level of coupling online and in a distributed fashion; in cases where global information is needed to solve a problem instance, DisCoF eventually couples all robots together. A completeness analysis of DisCoF is provided.
Coalition Coordination for Tightly Coupled Multirobot Tasks with Sensor Constraints
Yu Zhang, L. E. Parker, and Subbarao Kambhampati
We propose a coordination mechanism to address coalition execution in tightly coupled multirobot tasks. It provides a flexible method to reason about synergies with overlapping coalitions, thus enabling multitasking robots in multi-robot tasks, which not only improves efficiency but also reduces resource requirements during task execution. This means the mechanism enables tasks that could not be easily handled before, especially when critical resources are rare but commonly required. This coordination mechanism is based on the concept of sensor constraint, introduced by information sharing between robots. We have proven that this mechanism is sound and complete in finding a coordination solution given a few assumptions.
A Formal Analysis of Required Cooperation in Multi-agent Planning
Yu Zhang and Subbarao Kambhampati
Research on multi-agent planning has been popular in recent years. While previous research has been motivated by the understanding that, through cooperation, multi-agent systems can achieve tasks that are unachievable by single-agent systems, there are no formal characterizations of situations where cooperation is required to achieve a goal, thus warranting the application of multi-agent systems. In this paper, we provide such a formal discussion from the planning aspect. We first show that determining whether there is required cooperation (RC) is intractable in general. Then, by dividing the problems that require cooperation into two classes, problems with heterogeneous and homogeneous agents, we aim to identify all the conditions that can cause RC in these two classes. We establish that when none of these identified conditions hold, the problem is single-agent solvable. Furthermore, with a few assumptions, we provide an upper bound on the minimum number of agents required for RC problems with homogeneous agents.
2013
IQ-ASyMTRe: Forming Executable Coalitions for Tightly Coupled Multirobot Tasks
Yu Zhang and L. E. Parker
While most previous research on forming coalitions concentrates mainly on loosely coupled multirobot tasks, a more challenging problem is to address tightly coupled multirobot tasks that involve close robot coordination, which often requires capability sharing. General methods for autonomous capability sharing have been shown to greatly improve the flexibility of distributed systems. However, in addition to the interaction constraints between the robots and the environment required by the tasks, these methods may introduce additional interaction constraints between robots based on how the capabilities are shared. The satisfiability of these constraints in the current situation determines the feasibility of potential coalitions. To achieve system autonomy, the ability to identify potential coalitions that are feasible for task execution is critical. We introduce a general approach that incorporates this capability, extending the ASyMTRe architecture into IQ-ASyMTRe, which is able to find coalitions in which these required constraints are satisfied. When used to form coalitions, IQ-ASyMTRe sets up only feasible coalitions, enabling tasks to be executed autonomously. We have formally proven that IQ-ASyMTRe is sound and complete for forming executable coalitions.
Considering Inter-Task Resource Constraints in Task Allocation
Yu Zhang and L. E. Parker
Task allocation with single-task robots, multi-robot tasks, and instantaneous assignment has been shown to be strongly NP-hard. Although this problem has been studied extensively, few efficient approximation algorithms have been provided given its inherent complexity. We provide discussion and analysis of two natural greedy heuristics for solving this problem, then introduce a new greedy heuristic that considers inter-task resource constraints to approximate the influence between different assignments. Instead of only looking at the utility of an assignment, our approach computes the expected loss of utility, due to the assigned robots and task, as an offset, and uses the offset utility for making the greedy choice. A formal analysis of the new heuristic shows that solution quality is bounded by two different factors, and we provide a new algorithm to approximate the heuristic for improved performance.
Multi-Robot Task Scheduling
Yu Zhang and L. E. Parker
2012
Coalition Formation and Execution in Multi-robot Tasks
Yu Zhang
Task Allocation with Executable Coalitions in Multirobot Tasks
Yu Zhang and L. E. Parker
2011
Solution Space Reasoning to Improve IQ-ASyMTRe in Tightly-Coupled Multirobot Tasks
Yu Zhang and L. E. Parker
In FLOW, when an information flow is interrupted, or when flow quality no longer satisfies the task's requirements, the task robot can initiate a flow relaxation process for the affected coalition. Since this process can update the set of sensor constraints, the coordination solution also needs to be recreated. This only needs to be performed on the initiating coalition and any coalitions set up after it in the previous coordination process, unless a new coordination solution cannot be found with these coalitions after relaxation. This paper improves on that process by reasoning about the solution space, providing a more robust and flexible flow relaxation process.
2010
IQ-ASyMTRe: Synthesizing Coalition Formation and Execution for Tightly-Coupled Multirobot Tasks
Yu Zhang and L. E. Parker
In IQ-ASyMTRe, robot capabilities are built as schemas based on schema theory, where each schema represents a motor, sensory, computational, or communication capability of the agent. Given a task, we must determine how to connect the different schemas of different robots to satisfy the task's requirements. To create the solution space of potential connection solutions for a task, the reasoning algorithm first checks all components that can output the required information instances for the task, then checks recursively for the inputs of those components until each path either ends in a source component or in a conflict with the referent instantiation constraint. In a second phase, the robots temporarily activate their capabilities to dynamically instantiate the information flows from sources to sinks.
A General Information Quality Based Approach for Satisfying Sensor Constraints in Multirobot Tasks
Yu Zhang and L. E. Parker
As coalitions are formed in FLOW, sensor constraints among robots are also established. How to keep these constraints satisfied throughout execution, from initial configuration to task completion, remains an open issue, and environmental factors, both static and dynamic, can influence whether the constraints continue to hold. This paper proposes a general method to address these issues across applications with different sensors. The method combines sensor models, environment sampling, and a measure of information quality with a sampled motion model. Local information-quality measures are then combined systematically to compute an overall flow quality.