I am a PhD student at Northeastern University, advised by Prof. Weiyan Shi. I also work closely with Prof. David Bau. I was a scholar in the MATS program, working on reasoning at Prof. Dawn Song’s Lab. I obtained my Master’s degree in Computer Science at UMass Amherst where I was lucky to be advised by Prof. Andrew McCallum and Prof. Hong Yu. I finished my undergraduate study in Computer Science at HKUST.
I am generally interested in understanding the mechanisms of AI models to improve and control them. I am currently working on post-training, especially on understanding and mitigating emergent behaviors. I am broadly interested in continual learning, reasoning and interp. Feel free to email me if you would like to collaborate.
News
- Gave a talk at Meta on how chat-template tokens shape AI safety, covering The Piggyback Hypothesis and LLMs Encode Harmfulness and Refusal Separately.
- The Piggyback Hypothesis of Generalization was accepted to NeurIPS 2026.
- Our True Thinking Score was added to the Tinker cookbook by Thinking Machines.
- Can Aha Moments Be Fake? was accepted to EMNLP 2026 Findings.
- Quanta Magazine covered our work on true vs. decorative thinking steps in chain-of-thought.
- New preprint: The Piggyback Hypothesis, explaining and mitigating emergent misalignment.
Older news
- New preprint: Can Aha Moments Be Fake? Identifying true and decorative thinking steps in CoT.
- LLMs Encode Harmfulness and Refusal Separately was accepted to NeurIPS 2025.
- Joined the MATS program for the summer, working on reasoning with Prof. Dawn Song’s lab.
- Started my PhD at Northeastern University, advised by Prof. Weiyan Shi.
- LLMs are In-context Teachers for Knowledge Reasoning was accepted to EMNLP 2024 Findings.
- Two papers accepted: one at ICML 2024, one at ACL 2024.
Publications
2026
- NeurIPS 2026
- EMNLP 2026 FindingsCan Aha Moments Be Fake? Identifying True and Decorative Thinking Steps in Chain-of-ThoughtCovered by Quanta Magazine; added to the Tinker cookbook.
2025
- NeurIPS 2025
2024
- EMNLP 2024 FindingsLarge Language Models are In-context Teachers for Knowledge ReasoningPDF
TL;DR
We propose the encoding specificity hypothesis, inspired by humans’ memory retrieval, to understand prompting LLMs. Effective prompts should match LLMs’ own training distribution.
- ICML 2024
- ACL 2024
- IEEE JBHIAdaptive Fusion of Deep Learning with Statistical Anatomical Knowledge for Robust Patella Segmentation from CT ImagesIEEE Journal of Biomedical and Health Informatics
2023
- NeurIPS 2023 WorkshopSELF-EXPLAIN: Teaching Large Language Models to Reason Complex Questions by ThemselvesWorkshop on Robustness of Zero/Few-shot Learning in Foundation Models
- NeurIPS 2023 WorkshopStudent as an Inherent Denoiser of Noisy Teacher3rd Workshop on Efficient Natural Language and Speech ProcessingarXiv
TL;DR
We find the model converges to clean labels faster during knowledge distillation, so we leverage early checkpoints to denoise teacher labels.
- ICML 2023 WorkshopIn-Context Exemplars as Clues to Retrieving from Large Associative MemoryNeural Conversational AI @ ICML 2023 · Associative Memory & Hopfield Networks @ NeurIPS 2023
2022
- arXiv
* Equal contribution
Academic service
Reviewer for ICML, NeurIPS, CoNLL, AAAI, ICLR
