About
I am a Researcher at Ant Group, currently working on Agentic RL and Coding Agents. My research focuses on enabling LLM-based agents to learn from interaction, diagnose their own failures, and continuously improve their capabilities in real-world environments. Previously, I worked on foundation model pre-training and large language model alignment, including post-training and preference alignment. I received my PhD from The Chinese University of Hong Kong, Shenzhen in 2026, advised by Dr. Xiang Wan and Prof. Tsung-Hui Chang.
I am always open to research collaborations and discussions! Please drop me with an email.News
Selected updates
One paper accepted at EMNLP 2026
One paper accepted at ICML 2026Oral
Three papers accepted at ACL 2026
One paper accepted at ICLR 2026
Earlier updates
One paper accepted at AAAI 2026
One paper accepted at NeurIPS 2025
Two papers accepted at EMNLP 2025
One paper accepted at TMLR
One paper accepted at ACM Multimedia 2025Oral
One paper accepted at TMLR
One paper accepted at ACL 2025Oral
One paper accepted at IEEE TCSVT
One paper accepted at NAACL 2025
Featured Publications
* Equal contribution. † Tech lead / Corresponding author
-
VisCache: Visual KV Cache Pruning for Efficient Vision Large Language Model Inference
-
Mitigating Reward Hacking in RLHF via Bayesian Non-negative Reward Modeling
-
Eliminating Inductive Bias in Reward Models with Information-Theoretic Guidance
-
APLOT: Robust Reward Modeling via Adaptive Preference Learning with Optimal Transport
-
RLHF in an SFT Way: From Optimal Solution to Reward-Weighted Alignment
All Publications
Grouped by research theme
LLM Alignment, Reward Modeling & Data Curation
-
Mitigating Reward Hacking in RLHF via Bayesian Non-negative Reward Modeling
-
Eliminating Inductive Bias in Reward Models with Information-Theoretic Guidance
-
APLOT: Robust Reward Modeling via Adaptive Preference Learning with Optimal Transport
-
RLHF in an SFT Way: From Optimal Solution to Reward-Weighted Alignment
Prompting, Safety & Task-Specific NLP
-
Safeguarding LLM Fine-tuning via Push-Pull Distributional Alignment
-
Atoxia: Red-teaming Large Language Models with Target Toxic Answers
-
Self-Instructed Derived Prompt Generation Meets In-Context Learning
-
A Simple yet Effective Subsequence-Enhanced Approach for Cross-Domain NER
-
Graph Enhanced Contrastive Learning for Radiology Findings Summarization
Benchmarks & Healthcare-Centric LLMs
Learning with Limited or Imbalanced Data
-
Synthesizing Minority Samples for Long-tailed Classification via Distribution Matching
-
Prototype-oriented Clean Subset Extraction for Noisy Long-tailed Classification
-
Learning to Re-weight Examples with Optimal Transport for Imbalanced Classification
-
Diverse Condensed Data Generation via Class Preserving Distribution Matching
Other Publications
-
Rotary Position Embedding-Based Transformer Hawkes Process for Event-Type Big Data
-
Intermediate Domain Alignment and Morphology Analogy for Patent-Product Image Retrieval
-
Toward Explainable and Fine-Grained 3D Grounding through Referring Textual Phrases
-
Cognitive-Visual Fusion for Target Classification in Rapid SAR Image Series Visual Presentation
Experience
-
2026 — Present
Researcher · Ant Group
Agentic RL and Coding Agent
-
Jul 2025 — Feb 2026
Research Intern · Qwen, Alibaba
Reward hacking and RLHF · Mentor: Pengyu Cheng
-
Apr — Jun 2025
Research Intern · Noah's Ark Lab, Huawei
Data mixture for MLLM pre-training · Mentor: Lu Hou
-
Dec 2024 — Mar 2025
Research Intern · LIGHT STUDIO, Tencent
LLM-based machine translation · Mentor: Xiaoqi Jiao