Yu Chen 陈禹

Ph.D. student · IIIS, Tsinghua University

Hi! My name is Yu Chen.

I am a third-year Ph.D. student at the Institute for Interdisciplinary Information Sciences (IIIS), Tsinghua University, where I am very fortunate to be advised by Prof. Longbo Huang. My research lies at the intersection of AI and decision-making: I develop the theory of sequential decision-making under uncertainty, from bandits to reinforcement learning, and study how AI systems such as LLM agents make decisions.

Starting in November 2026, I will be visiting Caltech as a visiting student, under the guidance of Prof. Adam Wierman. Previously, I was a research intern at the MSR Asia Theory Center from February to August 2024, under the guidance of Dr. Wei Chen.

Prior to my Ph.D., I received my Bachelor of Science in Mathematics from Tsinghua University.

A common thread in my work is decision-making in settings where standard practice or standard assumptions fall short: LLM agents and LLM training that rely on heuristics without a well-defined objective, feedback that is heavy-tailed rather than bounded, and objectives that weigh risk rather than only the expected return. In each setting, I look for formulations that lead to principled algorithms and, where possible, provable guarantees. My current work is organized around three directions.

Decision-making in LLMs and LLM agents. Large language models, and the agents built on them, increasingly make decisions on their own: which skills or tools to use, what to keep in a limited context, and how to act on uncertain feedback. I study what makes these decisions good, and how ideas from sequential decision-making and optimization can make them more reliable and efficient. Recent work includes skill selection for LLM agents, objectives for fine-tuning, and robust on-policy distillation.

Heavy-tailed online learning. In many real problems, rewards and losses are heavy-tailed: rare, extreme outcomes dominate the average and break the estimators that standard algorithms rely on. I study how to learn and decide robustly under such feedback, ideally without knowing in advance how heavy the tails are or whether the environment is stochastic or adversarial. Recent work includes best-of-both-worlds algorithms for heavy-tailed bandits and Markov decision processes.

Risk-sensitive reinforcement learning. In safety-critical applications, a good policy must care not only about the expected return but also about the risk of bad outcomes. I study when risk-aware objectives, such as conditional value-at-risk and other risk measures over the return distribution, can be learned as efficiently as the standard objective, including in large state spaces that require function approximation. Work in this direction includes provably efficient algorithms for risk-sensitive RL with iterated CVaR objectives, Lipschitz risk measures, and partial observability.