CV

I lead DreamX, a 100+ member product-facing AI team at Alibaba AMAP building spatial intelligence models and systems that understand and predict, generate and simulate, plan and act in the real world. My research traces an arc from neural architecture search through Vision Transformer design and multimodal foundation models to LLM reasoning, world models, agent systems, and large-scale AMAP products serving 300M+ users every day. I have authored 120+ research papers and preprints, including 70+ papers at top-tier AI conferences, with 15,000+ citations (6,000+ from first-authored works) across open-source projects.

15,000+ Citations
6,000+ First-Author Citations
120+ Publications
100+ Team Members

Awards & Recognition


Professional Experience

Alibaba Group — AMAP Mar 2024 – Present
Senior Director & Head of DreamX

Leading AMAP’s 100+ member spatial-intelligence team across understanding and prediction, generation and simulation, planning and action, and shared model, data, infrastructure, and evaluation foundations.

  • Led DreamX research resulting in 50+ team papers at top venues (ICLR, CVPR, ICML, ACL, KDD, ICCV, ECCV, NeurIPS, AAAI, EMNLP, SIGGRAPH) and 40+ open-source projects hosted through the AMAP-ML GitHub organization
  • Key first-author works: GPG (ICLR 2026, adopted by ByteDance’s VERL framework), USP (ICCV 2025); key team works: DreamX-Phi, DreamX-World, LongHorizon-Harness, SkillClaw, Tree-GRPO, FASA, CoEvolve
  • Multimodal technology supports AMAP’s Saojie Bang (扫街榜) pipeline; large-scale industrial Agent work contributes to AI Companion (AI 伴行) — alongside AMAP products serving 300M+ users every day
Meituan — Visual Intelligence Department May 2020 – Mar 2024
Senior Technical Manager

Built the Visual Intelligence team from scratch. Directed research in Vision Transformers, multimodal large models, and industrial AI systems.

  • Created Twins (NeurIPS 2021), which systematized layer-wise local-global attention interleaving in a general-purpose vision backbone; CPVT (ICLR 2023); and VisionLLaMA (ECCV 2024), which introduced auto-scaling 2D RoPE for LLaMA-style vision backbones
  • Built MobileVLM, a compact VLM designed for real-time on-device deployment; reproduced LLaMA 7B from scratch
  • Open-sourced YOLOv6, a widely used industrial object detection framework; developed QARepVGG to address quantization challenges in RepVGG-style deployment
  • Shipped 3D perception for autonomous delivery vehicles and drones, reducing annotation and serving costs
Xiaomi — Artificial Intelligence Department Mar 2017 – May 2020
Senior Technical Manager

Founded Xiaomi’s AutoML team and produced a series of influential neural architecture search works.

  • FairNAS (ICCV 2021), FairDARTS (ECCV 2020), DARTS- (ICLR 2021), FALSR — establishing new standards for fair and robust architecture search
  • Won 2nd place in Xiaomi’s first “Million Dollar Prize” (Automated Neural Network Design)
  • FALSR super-resolution algorithm personally endorsed by CEO Lei Jun
Beijing KingStar System Control Co., Ltd. Jun 2013 – Mar 2017
Deputy Director
  • Core contributor to “Complex Power Grid Autonomy — Collaborative Automatic Voltage Control” project
  • Contributed 20 invention patents; awarded National Science and Technology Progress First Prize (2018)
IBM Research China (CRL) Jul 2012 – May 2013
Research Scientist
  • Large-scale data analytics and machine learning solutions at IBM China Research Lab

Selected Publications

LLM Reasoning

  • GPG: Simple & Strong RL for Reasoning — ICLR 2026 · 1st Author · 100+ citations
  • Tree-GRPO: Tree Search for Agent RL — ICLR 2026
  • CoEvolve: Agent-Data Co-Evolution — ACL 2026
  • MathForge: Difficulty-Aware GRPO — ICLR 2026
  • AutoDrive-R2: Reasoning VLA for Driving — ICLR 2026

Generative AI & World Models

  • DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation · 2026 Technical Report · project
  • USP: Unified Pretraining for Gen & Understanding — ICCV 2025 · 1st Author
  • DCW: SNR-t Bias of Diffusion Models — CVPR 2026
  • S2-Guidance: Training-Free Diffusion Enhancement — ICLR 2026
  • EPG: End-to-End Pixel Generation without VAE — ICLR 2026
  • DreamX-World: Interactive World Model · 2026 Technical Report · code

AI Agents & Intelligent Mobility

Foundation Architectures

Vision-Language & Detection

  • MobileVLM: Real-Time Mobile Vision-Language Model · 1st Author
  • YOLOv6: Industrial Object Detection
  • SpatialGenEval: Spatial Intelligence Benchmark — ICLR 2026
  • PromptDet: Open-Vocabulary Detection — ECCV 2022

AutoML & Neural Architecture Search

Full publication list (120+ papers)


Professional Service

Area Chair: ICLR Area Chair: NeurIPS SPC: AAAI SPC: IJCAI

Education


Patents