VLM
How visual tokens get selected and compressed under a fixed context budget. Long video is where this hurts most — and where evaluation is easiest to fool.
Research Homepage
I am an undergraduate student at Zhejiang University. I work at the boundary between perception and reasoning — systems that take in images, video, or documents and have to decide what to do with them, not just describe them.
My research interests lie in VLM, Agent, WM, and Generative Models.
01 / Research
How visual tokens get selected and compressed under a fixed context budget. Long video is where this hurts most — and where evaluation is easiest to fool.
Multi-step tool use that stays grounded: planning, retrieval, and verification loops that catch their own mistakes instead of confidently continuing.
Learned dynamics as a substrate for planning — what a model has to represent before imagining the next state becomes useful rather than decorative.
Diffusion and autoregressive generation, with an eye on controllability and how generated data feeds back into training.
02 / Work
54 one-page cards covering the LeetCode Hot 100 and the 代码随想录 roadmap — one card per algorithmic pattern, each with the definition, recognition signals, a worked diagram, C++ / Python templates, and a mnemonic. Ships with two ready-to-present decks and the full generation scripts.
The Zhejiang University Beamer template, packaged as a skill an AI agent can invoke end to end: copy the resources, write the .tex, compile with XeLaTeX, open the PDF. Official ZJU blue, 16:9, CJK typesetting already wired up.