科研成果

(2025). Spiking point transformer for point cloud classification. In AAAI.

PDF Code

(2025). Semantic-Aware Late-Stage Supervised Contrastive Learning for Fine-Grained Action Recognition. In IEEE Transactions on Circuits and Systems for Video Technology.

PDF Dataset

(2025). MMSupcon: An image fusion-based multi-modal supervised contrastive method for brain tumor diagnosis. In Artificial Intelligence in Medicine.

PDF Code

(2025). MeDKCoOp: Dual Knowledge-guided Graph Prompt Learning for Biomedical Vision-Language Models. In Proceedings of the 33rd ACM International Conference on Multimedia.

PDF Code

(2025). Hybrid Vision Transformer and Convolutional Neural Network for Super-Resolution Image Quality Assessment. In ICCV.

PDF

(2025). Hierarchical Task-aware Temporal Modeling and Matching for few-shot action recognition. In Neurocomputing.

PDF

(2025). Flow-anchored consistency models. In arXiv preprint arXiv:2507.03738.

PDF Code

(2025). Event-Enhanced Blurry Video Super-Resolution. In AAAI.

PDF Code

(2025). Entropy-Adapter-Based Deep Image Compression for User-Generated Content with Knowledge Distillation. In DCC.

PDF

(2025). Enhancing zero-shot brain tumor subtype classification via fine-grained patch-text alignment. In Expert Systems with Applications.

PDF

(2025). Enhancing Visual Tracking by Leveraging High-frequency Information within Event Signals.

PDF

(2025). Enhancing Visual Question Answering Via Clustered In-Context Sequence Configuration. In ICIP.

PDF

(2025). Egvd: Event-guided video deraining. In IEEE Transactions on Neural Networks and Learning Systems.

PDF Code

(2025). Efficient Spiking Point Mamba for Point Cloud Analysis. In arXiv preprint arXiv:2504.14371.

PDF

(2025). Efficient Event-Based Semantic Segmentation via Exploiting Frame-Event Fusion: A Hybrid Neural Network Approach. In AAAI.

PDF Code

(2025). Dome-DETR: DETR with density-oriented feature-query manipulation for efficient tiny object detection. In Proceedings of the 33rd ACM International Conference on Multimedia.

PDF Code

(2025). Create Anything Anywhere: Layout-Controllable Personalized Diffusion Model for Multiple Subjects. In arXiv preprint arXiv:2505.20909.

PDF

(2024). Wavelet-like Transform with Subbands Fusion in Decoupled Structure for Deep Image Compression. In PCS.

PDF

(2024). Visual perception by large language model’s weights. Advances in Neural Information Processing Systems. In NeurIPS.

PDF Code

(2024). Understanding of facial features in face perception: insights from deep convolutional neural networks. In Frontiers in Computational Neuroscience.

PDF Dataset

(2024). Tmformer: Token merging transformer for brain tumor segmentation with missing modalities. In AAAI.

PDF

(2024). Task navigator: Decomposing complex tasks for multimodal large language models. In CVPR.

PDF

(2024). Semi-Supervised Medical Image Segmentation Via Dynamic Pseudo-Label Refinement. In IEEE International Symposium on Biomedical Imaging.

PDF Code

(2024). Semantic-Enhanced Point-Box Joint Prompting for Video Object Segmentation. In ICIP.

PDF

(2024). Scene adaptive sparse transformer for event-based object detection. In CVPR.

PDF Code

(2024). Perceptual Image Compression With Conditional Diffusion Transformers. In VCIP.

PDF

(2024). Optimized Decoupled Structure with Non-Local Attention for Deep Image Compression. In ICIP.

PDF

(2024). Multi-modal generative embedding model.

PDF

(2024). Multi-modal Diffusion Network with Controllable Variability for Medical Image Segmentation. In IEEE International Conference on Bioinformatics and Biomedicine.

PDF

(2024). Image captioning with multi-context synthetic data. In AAAI.

PDF

(2024). GRACE: GRadient-based Active Learning with Curriculum Enhancement for Multimodal Sentiment Analysis. In Proceedings of the 32nd ACM International Conference on Multimedia.

PDF

(2024). Feature Compression With 3D Sparse Convolution. In VCIP.

PDF

(2024). EvTexture: event-driven texture enhancement for video super-resolution.

PDF Code

(2024). Event-based stereo depth estimation by temporal-spatial context learning. In IEEE Signal Processing Letters.

PDF

(2024). Event-based head pose estimation: Benchmark and method. In ECCV.

PDF Code

(2024). Event-assisted low-light video object segmentation. In CVPR.

PDF Code

(2024). Estme: Event-driven spatio-temporal motion enhancement for micro-expression recognition. In ICME.

PDF

(2024). Ee-mllm: A data-efficient and compute-efficient multimodal large language model.

PDF

(2024). Deep multi-threshold spiking-UNet for image processing. In Neurocomputing.

PDF Code

(2024). D-FINE: Redefine regression task in DETRs as fine-grained distribution refinement..

PDF Code

(2024). Anatomical consistency distillation and inconsistency synthesis for brain tumor segmentation with missing modalities.

PDF

(2024). A micro-expression recognition system with event cameras. In IEEE International Conference on Multimedia and Expo Workshops.

PDF

(2023). Video super-resolution via event-driven temporal alignment. In ICIP.

PDF Code

(2023). Text-Only Image Captioning with Multi-Context Data Generation. In CoRR.

(2023). Learned rate-distortion cost prediction for ultrafast screen content intra coding. In IEEE Transactions on Circuits and Systems for Video Technology.

PDF

(2023). Get: Group event transformer for event-based vision. In ICCV.

PDF Code

(2023). Eoformer: Edge-oriented transformer for brain tumor segmentation. In International Conference on Medical Image Computing and Computer-Assisted Intervention.

PDF Code

(2023). Electron microscopy image registration using correlation volume. In International Symposium on Biomedical Imaging.

PDF Code

(2023). Deep spiking-unet for image processing. In CoRR.

PDF Code

(2023). Better and faster: Adaptive event conversion for event-based object detection. In AAAI.

PDF

(2023). Attention-guided contrastive masked image modeling for transformer-based self-supervised learning. In ICIP.

PDF Code

(2021). Training spiking neural networks with accumulated spiking flow. In AAAI.

PDF Code