Skip to content CV
Education
- M.S. in Computer Technology, University of Science and Technology of China, 2024.09-2027.06 (expected)
- B.S. in Software Engineering, Zhengzhou University, 2020.09-2024.06
Project Experience
- 2025.06-2026.01: Process- and Verification-Aware Reward Modeling (PRM/VRM) for Scientific Reasoning Alignment
- First Author. Tech stack: RLVR, PPO, DeepSpeed, vLLM.
- End-to-end alignment pipeline: targeted hallucination in scientific reasoning on LLaMA-3 / Qwen3 backbones, leading the design of the full pipeline from data construction and SFT to RLVR.
- Data synthesis and quality control: built the QuantumQA dataset (92k+ samples) using Self-Instruct with rejection sampling for high-quality CoT trajectories, and a multi-agent automated cleaning and verification framework to ensure data quality.
- Reward model innovation: proposed a Verification-aware Reward Model (VRM) for long-chain reasoning, with an Adaptive Reward Fusion (ARF) mechanism that dynamically balances sparse code-execution outcome rewards against dense semantic process rewards, mitigating reward hacking and credit-assignment problems.
- Distributed RLVR training: built a PPO training loop on 8×H200 with DeepSpeed ZeRO-3; designed VRM-based fine-grained advantage estimation with token-level KL penalty for policy-shift control. Achieved a +8.1% relative gain in Pass@1 on the QuantumQA test set over the SFT baseline while keeping training stable.
- 2025.07-2025.09: Multi-Agent Automated Cleaning and Verification Framework for Scientific Reasoning Data
- First Author. Tech stack: LangGraph, Code Interpreter, Data Filtering, Self-Correction.
- Multi-agent collaborative cleaning architecture: tackled logical hallucinations in synthetic data with a multi-agent pipeline — Planner decomposes verification goals, Solver translates natural language into executable code, and Reviewer compares results, enabling automated grading and filtering of data quality.
- Code-execution-grounded ground-truth verification: integrated Python as an external tool, executing code cells in reasoning trajectories and comparing outputs to automatically detect and filter samples with logical fallacies or computational errors, ensuring training-data reliability.
- 2023.09-2024.2: Stranger Detection System for Laboratory
- Person in Charge
- Developed the front-end interface using Vue.js, implementing real-time monitoring, alert push notifications, and historical record querying.
- Constructed the back-end service based on the Spring Boot framework to handle front-end requests, including video stream processing, face recognition result comparison, and alert generation and push.
- Responsible for server selection, configuration, and deployment, using Docker containerization for rapid service deployment and horizontal scaling.
- Built and trained a face recognition model using PyTorch. The system runs stably with a recognition accuracy of over 98%.
- 2022.09-2023.4: Driver Abnormal Behavior Detection Project
- Core Member
- The team designed and implemented a driver abnormal behavior detection model based on Yolo v5, used TensorRT to accelerate model inference, compressed the model using a knowledge distillation algorithm based on instance conditions, and deployed the model on edge terminal devices.
- The trained model achieved an accuracy of 94.78% on the validation set.
- Responsible for model building, the entire front-end and back-end development process, and project deployment.
- The project won the national third prize in the Computer Design Competition.
- 2021.05-2024.06: Application of Decision Tree and Feedforward Neural Network Algorithms in Coordinating Multi-Agent Players
- Person in Charge
- Evaluated the positions of our players using a feedforward neural network algorithm to obtain corresponding scores.
- Combined with information such as the degree of noise affecting each player on the field and the player’s status, a decision tree algorithm was used to determine the best playing strategy for the players.
- Without considering the position changes of players in different formations, the team’s winning rate was improved by 8.7%.
- This project won the first prize in the 3D simulation group of the China Robot Competition and ROBOCUP Robot World Cup China Tournament.
Skills
- LLM Alignment (SFT, RLHF/RLVR, PPO, Reward Modeling)
- Distributed Training (DeepSpeed ZeRO, multi-GPU)
- LLM Inference & Deployment (vLLM)
- Multi-Agent Systems (LangGraph, tool use, code interpreters)
- LLM Finetune
- Javaweb Development
- Linux System Operation
- Deep Learning
- Image Processing
Honors and Awards
- First Prize in the 3D simulation group of the 2022 China Robot Competition and ROBOCUP Robot World Cup China Tournament
- National Third Prize in the 2023 Computer Design Competition
- University-level completion of the 2021 College Student Innovation and Entrepreneurship Challenge
- Second and Third-class Student Scholarships and National Encouragement Scholarship in 2020 and 2021
- Rated as a ‘Triple-A’ Student at Zhengzhou University in 2022
Service and Leadership
- Served as the person in charge of the Zhengzhou University Simulation Laboratory, leading the team to complete the RoboCup football robot-simulation 3D project.
- Served as the deputy class monitor, actively communicating with teachers and cooperating with the class monitor to complete class work.