Master Student of Computer Science and Technology

Qingyu Zhang

I’m Qingyu Zhang, a second-year master’s student at Chinese Information Processing Laboratory in the Institute of Software Chinese Academy of Sciences. My current research focuses on AI sales and customer-service agents, especially reliable multi-turn interaction and evaluation in real-world business scenarios. I am also actively exploring User Agents for realistic evaluation and multi-turn interaction.

My internship experience spans foundation-model pretraining at Baichuan Intelligence, post-training and dialogue optimization at Meituan, and building user agents for evaluation and multi-turn training at ByteDance. Along the way, I have worked on AI-Salesman for reliable LLM-driven telemarketing, ShortGPT for layer pruning and model efficiency, and open-source toolkits such as AutoAlign and ShortX.

Research interests

  • Agent
  • LLM Long Context
  • LLM Compression & Efficiency
  • LLM Post-training

Education

  • M.S. in Computer Science and TechnologyInstitute of Software, CAS
    2024 - Present
  • B.S. in Computer Science and TechnologyCollege of Computer and Data Science, Fuzhou University
    2020 - 2024

News

  • Jul, 2026 One paper “ShortOPD” is released on arXiv as first author.
  • Dec, 2025 One paper “AI-Salesman” is accepted by AAAI 2026 as first author.
  • Nov, 2025 Open-sourced ShortX project, a unified pruning toolkit for AI models.
  • Jun, 2025 Contributed to AutoAlign project, an open-source toolkit for automated LLM alignment.
  • May, 2025 One paper “ShortV” is accepted by ICCV 2025.
  • May, 2025 One paper “ShortGPT” is accepted by ACL Findings 2025.

Experience

Algorithm Intern · ByteDance Jan 2026 – Present · Beijing, China
  • Led the R&D of a User Agent framework supporting multi-turn evaluation needs across business lines, with over 80% of the generated evaluation data being business-usable.
  • Exploring viable paradigms for applying the User Agent to multi-turn RL for Sales Agents.
Algorithm Intern · Meituan Dec 2024 – Jan 2026 · Beijing, China
  • Led the R&D of an RL-based dialogue optimization system for large models, building the full pipeline of training, inference, and evaluation.
  • Deployed in a live business environment, increasing core business conversion rate by 10%~20%.
  • Published as first author (AI-Salesman, AAAI, 2026).
Foundation Model Intern · Baichuan Intelligence Jan 2024 – Oct 2024 · Beijing, China
  • Investigated Transformer redundancy and proposed a layer-based pruning method (ShortGPT, ACL Findings, 2025).
  • Researched the lower bounds of RoPE Base (Base of RoPE Bounds Context Length, NeurIPS, 2024).
  • Proposed a variant of the “Needle in a Haystack” evaluation method (Patent Granted).
Research Intern · Institute of Software, Chinese Academy of Sciences Oct 2023 – Sep 2024 · Beijing, China
  • Adapted and optimized SFT/DPO algorithms for the Megatron framework (ACL Demo, 2025).
  • Implemented large-scale distributed training on Ascend 910b using the ModelLink framework.

Publications

Projects

Inference Acceleration for the 70B LLaMA-2 Large Language Model

Inference Acceleration for the 70B LLaMA-2 Large Language Model

As a coaching assistant for the ASC24 Student Supercomputer Challenge, I helped the team optimize the inference performance of the LLaMA-2-70B model using efficient frameworks like vLLM and designed data parallelism strategies to significantly reduce latency.
Training the 10-Billion Parameter Yuan-1.0 LLM

Training the 10-Billion Parameter Yuan-1.0 LLM

As a key member of the ASC23 team, I trained a 10B-level large language model using the DeepSpeed-Megatron framework, combining tensor, pipeline, and data parallelism. Our work won the First Prize.

Awards

  • Jun, 2024 Honored as an Outstanding Graduate at Fuzhou University.
  • May, 2023 Won the First Prize in the 10th ASC Student Supercomputer Challenge.
  • Nov, 2022 Won the First Prize in the 13th National College Student Mathematics Competition.