Portrait of Fanxu Meng

REDstar@Dots Studio · Xiaohongshu Inc.

Fanxu Meng 孟繁续

I work on efficient architectures for long-context large language models and parameter-efficient fine-tuning.

I am currently with Dots Studio at Xiaohongshu Inc. My research focuses on improving the efficiency and scalability of long-context LLMs through architectural innovations, as well as developing effective parameter-efficient fine-tuning methods. I have served as a reviewer for NeurIPS, ICML, ICLR, CVPR, TPAMI, etc.

I received my Ph.D. from the Institute for Artificial Intelligence at Peking University, advised by Prof. Muhan Zhang, and my master’s degree from Harbin Institute of Technology, Shenzhen, advised by Prof. Guangming Lu.

During my master’s and Ph.D. studies, I spent a total of four years at Tencent YouTu as both an intern and a full-time researcher, collaborating with Xing Sun, Hao Cheng, and Ke Li.

  • Efficient architectures for long-context LLMs
  • Parameter-efficient fine-tuning

Research

Selected publications

View all on Google Scholar
HISA hierarchical block-to-token indexing workflow

COLM 2026

HISA: Efficient Hierarchical Indexing for Fine-Grained Sparse Attention

A training-free, coarse-to-fine indexer that accelerates token-level sparse attention while preserving fine-grained selection.

Yufei Xu*, Fanxu Meng*, Fan Jiang, Yuxuan Wang, Ruijie Zhou, Zhaohui Wang, Jiexi Wu, Zhixin Pan, Xiaojuan Tang, Wenjie Pei, Tongxuan Liu, Di Yin, Xing Sun, Muhan Zhang

Diagram illustrating Tensor Parallel Latent Attention

ASPLOS 2026 · Summer cycle

TPLA: Tensor Parallel Latent Attention

An attention mechanism designed for tensor parallelism and prefill–decode disaggregation.

Xiaojuan Tang*, Fanxu Meng*, Pingzhi Tang, Yuxuan Wang, Di Yin, Xing Sun, Muhan Zhang

TransMLA architecture illustration

NeurIPS 2025 · Spotlight · Top 3.19%

TransMLA: Multi-Head Latent Attention Is All You Need

A theoretical and practical study of MLA, including conversions of Llama and Qwen models to the DeepSeek-style architecture.

Fanxu Meng*, Pingzhi Tang*, Xiaojuan Tang, Zengwei Yao, Xing Sun, Muhan Zhang

HD-PiSSA method illustration

EMNLP 2025 · Oral · Top 3.98%

HD-PiSSA: High-Rank Distributed Orthogonal Adaptation

High-rank parameter updates for data-parallel fine-tuning.

Yiding Wang*, Fanxu Meng*, Xuefeng Zhang, Fan Jiang, Pingzhi Tang, Muhan Zhang

CLOVER method illustration

ICML 2025

CLOVER: Cross-Layer Orthogonal Vectors Pruning and Fine-Tuning

An absorb–decompose approach to model pruning and fine-tuning.

Fanxu Meng, Pingzhi Tang, Fan Jiang, Muhan Zhang

Stripe-wise pruning illustration

NeurIPS 2020

Pruning Filter in Filter

A stripe-wise structured pruning method for convolutional neural networks.

Fanxu Meng*, Hao Cheng*, Ke Li, Huixiang Luo, Xiaowei Guo, Guangming Lu, Xing Sun

Filter grafting training illustration

CVPR 2020

Filter Grafting for Deep Neural Networks

A training strategy that reactivates invalid filters to improve representation capacity.

Fanxu Meng*, Hao Cheng*, Ke Li, Zhixin Xu, Rongrong Ji, Xing Sun, Guangming Lu

* Equal contribution.

Contact

Let’s talk research.

I welcome conversations about foundation model architectures, model adaptation, and efficient inference.