About Me

Hi! I'm Jianshuo Dong, a PhD student at Tsinghua University, advised by Prof. Han Qiu.

Research Interests

  • Safety & Security of Large Model Systems: Convey trustworthy large model services to downstream users.
  • Autonomous Agentic AI: Developing strong, safe, and efficient autonomous agents.
  • Explainable AI: Reverse-engineer and closely monitor AI models.

Email: dongjs23(at)mails(dot)tsinghua(dot)edu(dot)cn

News

2026-09-05 Congratulations to Yutong! Our paper INTENT-AS-A-TOOL was accepted to EMNLP 2026 Main!
2026-08-01 Our new paper on IPI probing is now available on arXiv!
2026-05-01 One paper (SafeSearch) was accepted to ICML 2026 as a regular paper!
2026-04-04 One paper (IFEval++) got accepted to ACL 2026 main (oral)!
2025-10-17 I became a PhD candidate at Tsinghua University!
2025-08-23 One paper (Leakage-intent probing) got accepted to EMNLP 2025 main (oral)!
2025-08-10 I created this website for sharing my research and career!

Selected Papers

Agentic misalignment running example

INTENT-AS-A-TOOL Makes it Easy to Track Agentic Misalignment

EMNLP 2026 Main
Yutong Zhang, Jianshuo Dong, Peng Xu, Long Wang, Jie Zhang, Tianwei Zhang, Xiaoping Zhang, Han Qiu
TL;DR: Tracking agentic misalignment through fine-grained intent-tool call probabilities.
Explanation of hidden prompt-injection exposure signals

Your Agentic LLMs Secretly Encode Indirect Prompt-Injection Exposure in Hidden States

arXiv:2608.02657
Jianshuo Dong, Yiming Liu, Maosen Zhang, Nan Deng, Peng Xu, Xiaoping Zhang, Tianwei Zhang, Jie Zhang, Han Qiu
TL;DR: Detecting, defending against, and explaining indirect prompt-injection exposure through latent signals in agentic LLMs.
SafeSearch motivating search-agent incident

SafeSearch: Automated Red-Teaming of LLM-Based Search Agents

ICML 2026 Regular
Jianshuo Dong, Sheng Guo, Hao Wang, Xun Chen, Zhuotao Liu, Tianwei Zhang, Ke Xu, Minlie Huang, Han Qiu
TL;DR: To proactively discover the potential vulnerabilities of LLM-based search agents.
Reliability gap in instruction-following models

Revisiting the Reliability of Language Models in Instruction-Following

ACL 2026 Oral
Jianshuo Dong, Yutong Zhang, Yan Liu, Zhenyu Zhong, Tao Wei, Chao Zhang, Han Qiu
TL;DR: Investigating nuance-oriented reliability of LLMs in instruction-following with varied phrasings.
Pipeline for measuring cognitive habits of reasoning models

Towards Understanding the Cognitive Habits of Large Reasoning Models

MIR 2026
Jianshuo Dong, Yujia Fu, Chuanrui Hu, Chao Zhang, Han Qiu
TL;DR: Studying whether large reasoning models exhibit human-like cognitive habits.
Intention-based prompt-leakage example

"I've Decided to Leak": Probing Internals Behind Prompt Leakage Intents

EMNLP 2025 Oral
Jianshuo Dong, Yutong Zhang, Yan Liu, Zhenyu Zhong, Tao Wei, Ke Xu, Minlie Huang, Chao Zhang, Han Qiu
TL;DR: Diving into the internals to understand LLMs' prompt leakage intents.
Cellular specification refinement tasks

Can Large Language Models Automate the Refinement of Cellular Network Specifications?

arXiv:2507.04214
Jianshuo Dong, Yuanjie Li, Jun Liu, Hewu Li, Han Qiu
TL;DR: Evaluate and improve the performance of LLMs in refining cellular network specifications concerning security issues.
Engorgio output-length comparison

An Engorgio Prompt Makes Large Language Model Babble on

ICLR 2025 Poster
Jianshuo Dong, Ziyuan Zhang, Qingjie Zhang, Tianwei Zhang, Hao Wang, Hewu Li, Qi Li, Chao Zhang, Ke Xu, Han Qiu
TL;DR: An inference cost attack targeting modern auto-regressive LLMs.
One-bit flip attack optimization illustration

One-bit Flip is All You Need: When Bit-flip Attack Meets Model Training

ICCV 2023 Poster
Jianshuo Dong, Han Qiu, Yiming Li, Tianwei Zhang, Yuanjie Li, Zeqi Lai, Chao Zhang, Shu-Tao Xia
TL;DR: Insert unactivated backdoor in the model training process and make it activated by bit-flip attack.

Experiences

Education

2023 - 2028 (expected) Tsinghua University, Beijing, China — Ph.D. Student
2019 - 2023 Wuhan University, Wuhan, China — B.E.

Internships

2026 Jun - Aug ByteDance · Mamoda Team — Research Intern, GUI Agent Reinforcement Learning

Academic Services

Teaching

2025 Spring Trustworthy Machine Learning, Tsinghua University — Teaching Assistant

Reviewing

AAAI 2027
ARR 2027 Aug
ICLR 2026; 2025 (Notable)
ICML 2026 (Gold)
NeurIPS 2025 (Top)