I am a PhD student at Northeastern University, advised by Prof. Weiyan Shi. I also work closely with Prof. David Bau. I was a scholar in the MATS program, working on reasoning at Prof. Dawn Song’s Lab. I obtained my Master’s degree in Computer Science at UMass Amherst where I was lucky to be advised by Prof. Andrew McCallum and Prof. Hong Yu. I finished my undergraduate study in Computer Science at HKUST.

I am generally interested in understanding the mechanisms of AI models to improve and control them. I am currently working on post-training, especially on understanding and mitigating emergent behaviors. I am broadly interested in continual learning, reasoning and interp. Feel free to email me if you would like to collaborate.

News

Older news
  • New preprint: Can Aha Moments Be Fake? Identifying true and decorative thinking steps in CoT.
  • LLMs Encode Harmfulness and Refusal Separately was accepted to NeurIPS 2025.
  • Joined the MATS program for the summer, working on reasoning with Prof. Dawn Song’s lab.
  • Started my PhD at Northeastern University, advised by Prof. Weiyan Shi.
  • LLMs are In-context Teachers for Knowledge Reasoning was accepted to EMNLP 2024 Findings.
  • Two papers accepted: one at ICML 2024, one at ACL 2024.

Publications

2026

  1. NeurIPS 2026
    The Piggyback Hypothesis of Generalization: Explaining and Mitigating Emergent Misalignment
    Jiachen Zhao, Zhengxuan Wu, Aryaman Arora, Yiyou Sun, David Bau, Weiyan Shi
  2. EMNLP 2026 Findings
    Can Aha Moments Be Fake? Identifying True and Decorative Thinking Steps in Chain-of-Thought
    Jiachen Zhao*, Yiyou Sun*, Weiyan Shi, Dawn Song
    Covered by Quanta Magazine; added to the Tinker cookbook.

2025

  1. NeurIPS 2025
    LLMs Encode Harmfulness and Refusal Separately
    Jiachen Zhao, Jing Huang, Zhengxuan Wu, David Bau, Weiyan Shi

2024

  1. EMNLP 2024 Findings
    Large Language Models are In-context Teachers for Knowledge Reasoning
    Jiachen Zhao, Zonghai Yao, Zhichao Yang, Hong Yu
  2. ICML 2024
    Learning and Forgetting Unsafe Examples in Large Language Models
    Jiachen Zhao, Zhun Deng, David Madras, James Zou, Mengye Ren
  3. ACL 2024
    Multistage Collaborative Knowledge Distillation from a Large Language Model for Semi-Supervised Sequence Generation
    Jiachen Zhao, Wenlong Zhao*, Andrew Drozdov*, Benjamin Rozonoyer, Md Arafat Sultan, Jay-Yoon Lee, Mohit Iyyer, Andrew McCallum
  4. IEEE JBHI
    Adaptive Fusion of Deep Learning with Statistical Anatomical Knowledge for Robust Patella Segmentation from CT Images
    Jiachen Zhao, Tianshu Jiang, Yi Lin, Justin Chan, Ping-Keung Lewis Chan, Chunyi Wen, Hao Chen
    IEEE Journal of Biomedical and Health Informatics

2023

  1. NeurIPS 2023 Workshop
    SELF-EXPLAIN: Teaching Large Language Models to Reason Complex Questions by Themselves
    Jiachen Zhao, Zonghai Yao, Zhichao Yang, Hong Yu
    Workshop on Robustness of Zero/Few-shot Learning in Foundation Models
  2. NeurIPS 2023 Workshop
    Student as an Inherent Denoiser of Noisy Teacher
    Jiachen Zhao
    3rd Workshop on Efficient Natural Language and Speech Processing
  3. ICML 2023 Workshop
    In-Context Exemplars as Clues to Retrieving from Large Associative Memory
    Jiachen Zhao
    Neural Conversational AI @ ICML 2023 · Associative Memory & Hopfield Networks @ NeurIPS 2023

2022

  1. arXiv

* Equal contribution

Academic service

Reviewer for ICML, NeurIPS, CoNLL, AAAI, ICLR