Cosmos 3: Omnimodal World Models for Physical AI
NVIDIA
arXiv, 2026
Improving Reinforcement Learning from Human Feedback with Efficient Reward Model Ensemble
Shun Zhang, Zhenfang Chen, Sunli Chen, Yikang Shen, Zhiqing Sun, and Chuang Gan
arXiv, 2024
We introduce efficient reward model ensemble approaches for reinforcement learning from human feedback (RLHF), achieving better alignment with human values under computational constraints.
Adaptive Online Replanning with Diffusion Models
Siyuan Zhou, Yilun Du, Shun Zhang, Mengdi Xu, Yikang Shen, Wei Xiao, Dit-Yan Yeung, and Chuang Gan
Conference on Neural Information Processing Systems (NeurIPS), 2023
When used for planning, diffusion models generate complete plans at once without incorporating new observations during the execution. Our algorithm determines when it is necessary to replan with new observations, and replans efficiently.
Planning with Large Language Models for Code Generation
Shun Zhang, Zhenfang Chen, Yikang Shen, Mingyu Ding, Joshua B. Tenenbaum, and Chuang Gan
International Conference on Learning Representations (ICLR), 2023
Our algorithm combines Monte-Carlo tree search with the Transformer beam search algorithm for code generation. It's more sample-efficient than a well-accepted sampling + filtering baseline.
Prompting Decision Transformer for Few-shot Policy Generalization
Mengdi Xu, Yikang Shen, Shun Zhang, Yuchen Lu, Ding Zhao, Joshua B. Tenenbaum, and Chuang Gan
International Conference on Machine Learning (ICML), 2022
Using trajectory segments of different tasks as prompts for a Decision Transformer to achieve meta-offline reinforcement learning.
Efficiently Finding Approximately-Optimal Queries for Improving Policies and Guaranteeing Safety
Shun Zhang
Ph.D. Dissertation, 2020
Querying to Find a Safe Policy Under Uncertain Safety Constraints in Markov Decision Processes
Shun Zhang, Edmund H. Durfee, and Satinder Singh
AAAI Conference on Artificial Intelligence (AAAI), 2020
An agent is uncertain about which policies are safe. It either finds a safe policy to accomplish a task or proves that no safe policies exist using a minimum number of queries.
Minimax-Regret Querying on Side Effects for Safe Optimality in Factored Markov Decision Processes
Shun Zhang, Edmund H. Durfee, and Satinder Singh
International Joint Conference on Artificial Intelligence (IJCAI), 2018
A novel formulation and a query selection algorithm for the avoiding negative side effects problem in safe reinforcement learning.
Approximately-Optimal Queries for Planning in Reward-Uncertain Markov Decision Processes
Shun Zhang, Edmund H. Durfee, and Satinder Singh
International Conference on Automated Planning and Scheduling (ICAPS), 2017
A provably-optimal query selection algorithm to resolve reward uncertainty for better planning in reward-uncertain Markov decision processes.
Autonomous Intersection Management for Semi-Autonomous Vehicles
Tsz-Chiu Au, Shun Zhang, and Peter Stone
Handbook of Transportation, 2015