Ph.D. Student @ CAS
Ming Ma
Internship Experience
Alibaba Tongyi Lab 2026.02 - Present
Research Intern
Qwen-Intelligence Team: Agentic Post train
Ant Group 2025.10 - 2026.01
Research Intern
Ling Team: Pre-training Quality
Microsoft Research Asia (MSRA) 2025.07 - 2025.10
Research Intern
Multi-Agent & Debugging
Papers
- A closed-loop intervention–validation framework that auto-debugs LLM multi-agent systems beyond passive failure-log analysis.
- A co-evolving dual-graph architecture (outline + knowledge) for open-ended deep-research agents that detects gaps and steers retrieval.
- Reveals that in-context learning relies on distributed local task vectors carried by label words rather than a single global encoding.
- A bidirectional question–answer coherence filter for selecting high-value synthetic code-instruction data.
- A training-free intermediate-layer decoding method enabling accurate early exit on single-token LLM tasks.
- Uses environment checks to assign progress credit to intermediate actions in long-horizon agentic reinforcement learning without an additional reward model.
- Explores the planning capabilities of multimodal models and evaluates their performance on agentic tasks.
- An abductive RL framework using drift-diffusion models to switch between deadlock and exploration in long-horizon sparse-reward tasks.
- OmniMemBench: Towards Scalable Evaluation of Long-Term Omni-Modal Agent Memory (NeurIPS 2026)Evaluates long-term omni-modal agent memory with text-only filtering to ensure questions require multimodal evidence.
Education
Chinese Academy of Sciences (CAS)
2022.09 - Present
Shandong University
2019.03 - 2022.06
Beijing Institute of Technology
2019.09 - 2020.06
Naval Aviation University
2017.08 - 2019.03
Links
Life & Photography
Writing & Media
Gaming
游戏研究社 - “阎王的恋爱小曲”背后,是人们挣脱引力的翅膀