Publications
arXiv · 2026
KiToke: Kernel-based Interval-aware Token Compression for Video Large Language Models
Training-free video token compression that removes global redundancy while preserving temporal coherence, including at token retention ratios as low as 1%.
IEEE TPAMI · 2026
Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM
Context-rich object representations and explicit identifiers unify 3D grounding, captioning, and spatial reasoning without task-specific fine-tuning.
CVPR · 2025
RoboGround: Robotic Manipulation with Grounded Vision-Language Priors
A large-scale simulated manipulation dataset and a grounding-aware robot policy that uses object masks to turn vision-language priors into actions.
NeurIPS · 2024
Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers
Object identifiers connect point clouds and multi-view images to language for 3D grounding, captioning, and question answering. Ranked first on ScanRefer and Scan2Cap in September 2024.
NAACL · 2025
Data-Efficiently Learn Large Language Model for Universal 3D Scene Perception
Chat-3D · published version
More publications 12
NeurIPS Datasets and Benchmarks · 2024