PhD candidate at The Hong Kong Polytechnic University (joint training with EIT, Ningbo), working on reinforcement learning for large language models and embodied agents. Models are trainable; the environments that train them are not. I want to make the environment trainable, the way models are.
Homepage 路 Email 路 Google Scholar 路 ORCID 路 Hong Kong
- C3: exact per-decision credit for cooperative LLM agents by transcript replay, with a method-agnostic audit of credit quality. [paper]
- AccuracyParadox-RLHF: reproducible RLHF training and evaluation pipelines and reference reward models. [paper, EMNLP 2024]

