Qingyu Zhang

I build LLM agents that hold reliable multi-turn conversations in real-world settings, and I work on making large models more efficient.

I’m a third-year master’s student at the Chinese Information Processing Laboratory, Institute of Software, Chinese Academy of Sciences. As an algorithm intern at ByteDance, I’m building User Agents: simulated users that make the evaluation of sales and customer-service agents more realistic and enable multi-turn reinforcement learning. Before that, I built AI-Salesman, an RL-driven sales-dialogue system deployed in production at Meituan (AAAI 2026).

On model efficiency, my latest first-author work, ShortOPD (arXiv 2026), recovers the free-form generation ability of pruned LLMs through short-to-long on-policy distillation, at a fraction of the training cost. Earlier, at Baichuan Intelligence, I co-developed the layer-pruning method ShortGPT (ACL Findings 2025). I also contribute to the open-source toolkits AutoAlign and ShortX.

张清宇 Portrait of Qingyu Zhang

Research interests

  • LLM agents
  • Long context
  • Model compression & efficiency
  • Post-training

Education

  • M.S. in Computer Science and Technology Institute of Software, Chinese Academy of Sciences 2024 – Present
  • B.S. in Computer Science and Technology Fuzhou University 2020 – 2024

News

  • Jul 2026 ShortOPD released on arXiv (first author).
  • Dec 2025 AI-Salesman accepted to AAAI 2026 (first author).
  • Nov 2025 Open-sourced ShortX, a unified pruning toolkit for AI models.
  • Jun 2025 Contributed to AutoAlign, an open-source toolkit for automated LLM alignment.
  • May 2025 ShortV accepted to ICCV 2025.
  • May 2025 ShortGPT accepted to ACL Findings 2025 (co-first author).

Experience

ByteDance

Present

Algorithm Intern

  • Led the R&D of a User Agent framework for multi-turn evaluation across business lines; over 80% of the evaluation data it generates is directly usable by the business.
  • Exploring how User Agents can power multi-turn RL for sales agents.

Meituan

Algorithm Intern

  • Led the R&D of an RL-based dialogue optimization system for LLMs, covering the full training, inference, and evaluation pipeline.
  • Deployed in live business, lifting the core conversion rate by 10–20%.
  • Published as first author: AI-Salesman (AAAI 2026).

Baichuan Intelligence

Foundation Model Intern

  • Studied layer redundancy in Transformers and proposed a layer-pruning method (ShortGPT, ACL Findings 2025).
  • Studied the lower bound of the RoPE base for long context (Base of RoPE Bounds Context Length, NeurIPS 2024).
  • Proposed a variant of the Needle-in-a-Haystack evaluation (patent granted).

Institute of Software, Chinese Academy of Sciences

Research Intern

  • Adapted and optimized SFT/DPO for the Megatron framework (AutoAlign, ACL Demo 2025).
  • Ran large-scale distributed training on Ascend 910B with the ModelLink framework.

Publications

Projects

Inference Acceleration for the 70B LLaMA-2 Large Language Model

As a coaching assistant for the ASC24 Student Supercomputer Challenge, I helped the team optimize the inference performance of the LLaMA-2-70B model using efficient frameworks like vLLM and designed data parallelism strategies to significantly reduce latency.

Training the 10-Billion Parameter Yuan-1.0 LLM

As a key member of the ASC23 team, I trained a 10B-level large language model using the DeepSpeed-Megatron framework, combining tensor, pipeline, and data parallelism. Our work won the First Prize.

Awards

  • Jun 2024 Outstanding Graduate, Fuzhou University.
  • May 2023 First Prize, 10th ASC Student Supercomputer Challenge.
  • Nov 2022 First Prize, 13th National College Student Mathematics Competition.