论文清单
论文与学术产出
下列清单覆盖视频理解、具身感知、世界模型、数字人与视频编码等方向。
- 一、视频理解与多模态内容分析
- 二、具身感知、三维场景理解与自主系统
- 三、时空运动建模与世界模型支撑技术
- 四、数字人、动作生成与视频内容生成
- 五、视频编码、压缩与国际标准
- 感知层(方向一、方向二) → 给世界模型提供观测编码和三维环境表征
- 预测层(方向三) → 给世界模型提供状态演化预测能力
- 生成层(方向四) → 给具身智能提供交互模拟与训练数据
- 编码层(方向五) → 给整个系统提供底层效率保障
汇总:截至 2026 年 6 月,创始团队及合作网络在视觉智能、多模态与具身相关方向发表的论文,约 200+
| 论文标题 | 作者 | 会议/期刊 | 年份 |
|---|---|---|---|
| Adversarial image detection based on spatial and frequency information | Yucheng Li, Shuai Wang, Guozhi Li, Jiangyao Shi, Wenyi Wang, Xiangdong Han, Lu Pan | Displays | 2026 |
| Audio-Visual Synergy for High-Fidelity Portrait Animation: A Two-Stage Framework with Identity-Motion Disentanglement | Yuzhi Lu, Yuanzong Mei, Wenyi Wang, Xi Chen, Xiaowen Chen, Zhi Wang, Feng Xu, Jianwen Chen | ISCAS | 2026 |
| Cross-Modal Probabilistic Representation Learning for Conversational Emotion Recognition | Zhou Z, Xu F, Wang W, Chen J | IEEE Transactions on Multimedia | 2026 |
| EdgeFER: A Structure-Guided and Hardware-Friendly Framework for Dynamic Facial Expression Recognition | Zhou Z, Chen A, Teka N A, Chen J | ISCAS | 2026 |
| RIA-NET: Realistic Image Animation through Semantic-aware Feature Learning Networks | Teka N A, Kumie G A, Chen A, Chen J | Journal of King Saud University Computer and Information Sciences | 2026 |
| Towards Generalizable Deepfake Detection: Mitigating Training Bias via Generalization Bound Tightening | Yuzhi Lu, Wenyi Wang, Xi Chen, Xiaowen Chen, Jianwen Chen, Ce Zhu | ACM MM | 2026 |
| AMT-Net: Adversarial Motion Transfer Network with Disentangled Shape and Pose for Realistic Image Animation | Teka N A, Alemu K G, Assefa M, et al | IEEE Access | 2025 |
| Multi-Scale Spatial-Frequency Features Representation and Learnable Cross Modal Feature Fusion in DeepFake Detection | Yuzhi Lu, Wenyi Wang, Xiaowen Chen, Fengyu Wang, Shuai Wang, Jianwen Chen | ICIP | 2025 |
| TransMask-Anim: Transformer-Driven Mask Perturbation and Keypoint Correspondence for Latent Space Image Animation | Teka N A, Zhou Z, Chen J | ICCWAMTIP | 2025 |
| Deepfake face discrimination based on self-attention mechanism | Shuai Wang, Donghui Zhu, Jian Chen, Jiangbo Bi, Wenyi Wang | Pattern Recognition Letters | 2024 |
| Face Animation Based on Multiple Sources and Perspective Alignment | Mei Y, Wang W, Liu X, et al | Virtual Reality & Intelligent Hardware | 2024 |
| Locational Detection of False Data Injection Attacks in Smart Grids: A Graph Convolutional Attention Network Approach | Wei Xia, Demin He, Lisha Yu | IEEE Internet of Things Journal | 2024 |
| Multi-visual Modality Micro Drone-Based Structural Damage Detection | Agyemang I O, Zeng L, Chen J, Adjei-Mensah I, Acheampong D | Engineering Applications of Artificial Intelligence | 2024 |
| Perception-oriented video frame interpolation via asymmetric blending | Guangyang Wu, Xin Tao, Changlin Li, Wenyi Wang, Xiaohong Liu, Qingqing Zheng | CVPR | 2024 |
| Accflow: Backward accumulation for long-range optical flow | Guangyang Wu, Xiaohong Liu, Kunming Luo, et al | ICCV | 2023 |
| Cascaded Temporal and Spatial Attention Network for Solar Adaptive Optics Image Restoration | Chi Zhang, Shuai Wang, Libo Zhong, Qingqing Chen, Changhui Rao | Astronomy & Astrophysics | 2023 |
| Cheap-fake detection with llm using prompt engineering | Guangyang Wu, Weijie Wu, Xiaohong Liu, Kele Xu, Tianjiao Wan, Wenyi Wang | ICMEW | 2023 |
| Distributed Collaborative Tensor Beamforming via Gaussian Entropy over Array Networks | Wei Xia, Lisha Yu, Guoqing Xia | IEEE Transactions on Vehicular Technology | 2023 |
| Dynamic UAV Swarm Confrontation: An Imitation Based on Mobile Adaptive Networks | Wei Xia, Zhuoyang Zhou, Wanyue Jiang, Yuhan Zhang | IEEE Transactions on Aerospace and Electronic Systems | 2023 |
| Fastllve: Real-time low-light video enhancement with intensity-aware look-up table | Wenhao Li, Guangyang Wu, Wenyi Wang, Peiran Ren, Xiaohong Liu | ACM MM | 2023 |
| Improved First-Order Motion Model of Image Animation with Enhanced Dense Motion and Repair Ability | Xu Y, Xu F, Liu Q, Chen J | Applied Science | 2023 |
| LiDAR Point Cloud Compression, Processing and Learning for Autonomous Driving | Abbasi R, Bashir A K, Alyamani H J, et al | IEEE Transactions on Intelligent Transportation Systems | 2023 |
| On the PMU Placement Optimization for the Detection of False Data Injection Attacks | Wei Xia, Demin He, Junbin Chen | IEEE Systems Journal | 2023 |
| Resilient Distributed Estimation Against FDI Attacks: A Correntropy-based Approach | Wei Xia, Yuhan Zhang | Information Sciences | 2023 |
| Spatial-Temporal Consistency Refinement Network for Dynamic Point Cloud Frame Interpolation | Ren L, Zhao L, Sun Z, Zhang Z, Chen J | IEEE ICMEW | 2023 |
| Widely Linear Null Broadening GNSS Anti-interference for Polarization Sensitive Arrays | Wei Xia, Yuhan Zhang, Hao Yu | IEEE Transactions on Aerospace and Electronic Systems | 2023 |
| 二阶广义总变分约束的太阳图像多帧盲解卷积 | 王帅、何春元、荣会钦等 | 光电工程 | 2023 |
| 低秩先验的相位差法太阳图像重建 | 王帅、鲍华、何春元等 | 光电工程 | 2023 |
| 深度学习点云质量增强方法综述 | 陈建文、赵丽丽、任蓝草、孙卓群、张新峰、马思伟 | 中国图象图形学报 | 2023 |
| 针对人脸识别卷积神经网络的局部背景区域对抗攻击 | 张晨晨、王帅、王文一等 | 光电工程 | 2023 |
| Adversarial Examples Detection Based on Error Level Analysis and Space Mapping | Sizhao Huang, Shuai Wang, Jian Chen, Guozhi Li, Wenyi Wang | ICASSP | 2022 |
| Enhanced Surveillance Video Compression with Dual Reference Frames Generation | Lei Zhao, Shiqi Wang, Shanshe Wang, Yan Ye, Siwei Ma, Wen Gao | IEEE Transactions on Circuits and Systems for Video Technology | 2022 |
| Evolution of AVS Video Coding Standards: Twenty Years of Innovation and Development | Siwei Ma, Li Zhang, Shiqi Wang, Chuanmin Jia, Shanshe Wang, Tiejun Huang, Feng Wu, Wen Gao | Science China Information Sciences | 2022 |
| FPX-NIC: An FPGA-Accelerated 4K Ultra-High-Definition Neural Video Coding System | Chuanmin Jia, Xinyu Hang, Shanshe Wang, Yaqiang Wu, Siwei Ma, Wen Gao | IEEE Transactions on Circuits and Systems for Video Technology | 2022 |
| Joint Local and Nonlocal Progressive Prediction for Versatile Video Coding | Meng Lei, Falei Luo, Xinfeng Zhang, Shanshe Wang, Siwei Ma | IEEE Transactions on Image Processing | 2022 |
| Multiframe Correction Blind Deconvolution for Solar Image Restoration | Shuai Wang, Huiqin Rong, Chunyuan He, Libo Zhong, Changhui Rao | PASP | 2022 |
| MVFI-Net: Motion-aware Video Frame Interpolation Network | Lin X, Zhao L, Liu X, Chen J | ACCV | 2022 |
| P-STMO: Pre-trained Spatial Temporal Many-to-One Model for 3D Human Pose Estimation | Wenkang Shan, Zhenhua Liu, Xinfeng Zhang, Shanshe Wang, Siwei Ma, Wen Gao | ECCV | 2022 |
| RangeINet: Fast LiDAR Point Cloud Temporal Interpolation | Zhao L, Lin X, Wang W, Ma K K, Chen J | ICASSP | 2022 |
| Real-time LiDAR Point Cloud Compression Using Bi-directional Prediction and Range-adaptive Floating-point Coding | Zhao L, Ma K K, Lin X, Wang W, Chen J | IEEE Transactions on Broadcasting | 2022 |
| Real-Time Scene-Aware LiDAR Point Cloud Compression Using Semantic Prior Representation | Zhao L, Ma K K, Liu Z, Yin Q, Chen J | IEEE TCSVT | 2022 |
| Scalable Intra Coding Optimization for Video Coding | Jiaqi Zhang, Meng Wang, Chuanmin Jia, Shanshe Wang, Siwei Ma, Wen Gao | IEEE Transactions on Circuits and Systems for Video Technology | 2022 |
| An Unsupervised Optical Flow Estimation for LiDAR Image Sequences | Guo X, Lin X, Zhao L, Zhu Z, Chen J | ICIP | 2021 |
| Deep Inter Prediction via Reference Frame Interpolation for Blurry Video Coding | Zhu Z, Zhao L, Lin X, Guo X, Chen J | VCIP | 2021 |
| Lossless Point Cloud Attribute Compression with Normal-based Intra Prediction | Yin Q, Ren Q, Zhao L, Chen J, Wang W | BMSB | 2021 |
| RAI-Net: Range-Adaptive LiDAR Point Cloud Frame Interpolation Network | Zhao L, Zhu Z, Lin X, et al | BMSB | 2021 |
| Regularized Intermediate Layers Attack: Adversarial Examples with High Transferability | Xiaorui Li, Weiyu Cui, Jiawei Huang, Wenyi Wang, Jianwen Chen | ICIP | 2021 |
| AIM 2020 Challenge on Efficient Super-Resolution: Methods and Results | K. Zhang, et al | ECCV Workshops | 2020 |
| Author Classification Using Transfer Learning and Predicting Stars in Co-Author Networks | Rashid Abbasi, Ali Kashif Bashir, Jianwen Chen, et al | Software—Practice & Experience | 2020 |
| RDH Based Dynamic Weighted Histogram Equalization for Secure Transmission and Cancer Prediction | Rashid Abbasi, Jianwen Chen, Yasser Al-Otaibi, et al | Multimedia Systems | 2020 |
| Reconstruction of Natural Visual Scenes from Neural Spikes with Deep Neural Networks | Yichen Zhang, Shanshan Jia, Yajing Zheng, Zhaofei Yu, Yonghong Tian, Siwei Ma, Tiejun Huang, Jian K. Liu | Neural Networks | 2020 |
| Robust Prior-Based Single Image Super Resolution under Multiple Gaussian Degradations | W. Wang, G. Wu, W. Cai, L. Zeng, J. Chen | IEEE Access | 2020 |
| Substitute Model Generation for Black-box Adversarial Attack Based on Knowledge Distillation | W. Cui, X. Li, J. Huang, W. Wang, S. Wang | ICIP | 2020 |
| 3D Facial Expression Recognition Using Deep Feature Fusion CNN | K. Tian, L. Zeng, S. McGrath, Q. Yin, W. Wang | ISSC | 2019 |
| A Compatible Framework for RGB-D SLAM in Dynamic Scenes | Zhao L, Liu Z, Chen J, Cai W, Wang W, Zeng L | IEEE Access | 2019 |
| Efficient Screen Content Coding Based on Convolutional Neural Network Guided by a Large-Scale Database, ICIP, 2019 | Zhao L, Wei Z, Cai W, Wang W, Zeng L, Chen J | — | 2019 |
| Feature Level MRI Fusion Based on 3D Dual Tree Compactly Supported Shearlet Transform | Chang Duan, Shuai Wang, Qihong Huang, et al | Journal of Visual Communication and Image Representation | 2019 |
| Khan, et al | Abubakar Hassan Sani, Jianwen Chen, Muhammad T | Design and Simulation of 3.9 GHz Microstrip Parallel Coupled Line Band Pass Filter. ICCWAMTIP | 2019 |
| Neural Network Based Image and Video Coding Technologies | Chuanmin Jia, Zhenghui Zhao, Shanshe Wang, Siwei Ma, Wen Gao | Telecommunications Science | 2019 |
| PRED: A Parallel Network for Handling Multiple Degradations via Single Model in Single Image Super-Resolution | G. Wu, L. Zhao, W. Wang, L. Zeng, J. Chen | ICIP | 2019 |
| Residual MultiSmoothlets | Shuai Wang, Dianmin Hu, Shuai Li, Ce Zhu, Changhui Rao | BMSB | 2019 |
| A Multi-scale Deconvolution Semantic Segmentation Network for Joint Detection and Segmentation | Feng N, Dong L, Zhang Q, et al | JCRAI | 2018 |
| Adaptive Illumination Normalization via Adaptive Illumination Preprocessing and Modified Weber-Face | Jianwen Chen, Zhen Zeng, Rumin Zhang, et al | Applied Intelligence | 2018 |
| Adaptive Shadow Removal Algorithm for Face Images | Z. Zeng, R. Zhang, J. Chen, L. Zeng, W. Wang, S. McMrath | ICST | 2018 |
| An Algorithm for Obstacle Detection Based on YOLO and Light Field Camera | Chen J, Zhang R, Yang Y, Wang W, Zeng L | ICST | 2018 |
| Extended Smoothlets: An Efficient Multi-Resolution Adaptive Transform | Shuai Wang, Chunmei Wang, Qian Zhang, et al | Journal of Visual Communication and Image Representation | 2018 |
| High Efficient VR Video Coding Based on Auto Projection Selection Using Transferable Features, VCIP, 2018 | Chen J, Zhao L, Zhang M, et al | — | 2018 |
| Multiple Low-Ranks Plus Sparsity Based Tensor Reconstruction for Dynamic MRI | Shan Wu, Yipeng Liu, Tengteng Liu, et al | DSP | 2018 |
| Robust Multi-Frame Super-Resolution Based on Spatially Weighted Half-Quadratic Estimation and Adaptive BTV Regularization | X. Liu, L. Chen, W. Wang, J. Zhao | IEEE TIP | 2018 |
| Using Hybrid Sensing Method for Motion Capture in the Standardized Volleyball Techniques Training | Lin J, Zhang R, Chen J, et al | ICST | 2018 |
| A Bayesian Approach for Camouflaged Moving Object Detection | Xiang Zhang, Ce Zhu, Shuai Wang, Yipeng Liu, Mao Ye | IEEE TCSVT | 2017 |
| A Real-Time Obstacle Detection Algorithm for the Visually Impaired Using Binocular Camera | Rumin Zhang, Wenyi Wang, Liaoyuan Zeng, Jianwen Chen, et al | CSPS | 2017 |
| A Robust and Real-Time Full 3D Reconstruction Method Based on Multiple Kinect | Peng X, Zeng L, Wang W, et al | CSPS | 2017 |
| Adaptive Progressive Motion Vector Resolution Selection Based on Rate-Distortion Optimization | Zhao Wang, Shiqi Wang, Jian Zhang, Siwei Ma | IEEE Transactions on Image Processing | 2017 |
| Complex Wavelet-Domain Image Watermarking Algorithm Using L1-Norm Function-Based Quantization | Jinhua Liu, Yuanyuan Xu, Shuai Wang, Ce Zhu | Circuits Systems & Signal Processing | 2017 |
| Enhancing Wedgelet-based Depth Modeling in 3D-HEVC | Chang Duan, Yuhuan Shen, Yingying Zhang, Shuai Wang, Ce Zhu, Meng Yang | APSIPA | 2017 |
| Entropy of Primitive: From Sparse Representation to Visual Information Evaluation | Siwei Ma, Xinfeng Zhang, Jian Zhang, Chuanmin Jia, Shiqi Wang, Wen Gao | IEEE Transactions on Circuits and Systems for Video Technology | 2017 |
| Hybrid Laplace Distribution-Based Low Complexity Rate-Distortion Optimized Quantization | Jing Cui, Shanshe Wang, Shiqi Wang, Xinfeng Zhang, Siwei Ma, Wen Gao | IEEE Transactions on Image Processing | 2017 |
| Inter-View Dependency-Based Rate Control for 3D-HEVC | Songchao Tan, Siwei Ma, Shanshe Wang, Shiqi Wang, Wen Gao | IEEE Transactions on Circuits and Systems for Video Technology | 2017 |
| Video Compressive Sensing Reconstruction via Reweighted Residual Sparsity | Chen Zhao, Siwei Ma, Jian Zhang, Ruiqin Xiong, Wen Gao | IEEE Transactions on Circuits and Systems for Video Technology | 2017 |
| A Novel Perception Oriented Image Color Representation | W. Wang, Y. Luo, J. Hu, J. Zhao | I2MTC | 2016 |
| Hiding Depth Information in Compressed 2D Image/Video Using Reversible Watermarking | W. Wang, J. Zhao | Multimedia Tools and Applications | 2016 |
| Low-Rank Decomposition-Based Restoration of Compressed Images via Adaptive Noise Estimation | Xiang Zhang, Weisi Lin, Ruiqin Xiong, Xianming Liu, Siwei Ma, Wen Gao | IEEE Transactions on Image Processing | 2016 |
| Nonlocal In-Loop Filter: The Way Toward Next-Generation Video Coding? | Siwei Ma, Xinfeng Zhang, Jian Zhang, Chuanmin Jia, Shiqi Wang, Wen Gao | IEEE MultiMedia | 2016 |
| Real-time Video Chroma Keying: A Parallel Approach Based on Local Texture and Global Color Distribution | L. Yin, W. Wang, J. Zhao | IET Image Processing | 2016 |
| Structure Prior Effects in Bayesian Approaches of Quantitative Susceptibility Mapping | Shuai Wang, Weiwei Chen, Chunmei Wang, et al | BioMed Research International | 2016 |
| Video Chroma Keying via Global Sampling and Trimap Propagation | C. Hao, W. Wang, J. Zhao | Multimedia Systems | 2016 |
| AVS2: Making Video Coding Smarter | Siwei Ma, Tiejun Huang, Cliff Reader, Wen Gao | IEEE Signal Processing Magazine | 2015 |
| Research Progress of Quantitative Susceptibility Mapping in MRI | Wang Shuai, Duan Chang, Zhang Ping, et al | Journal of Biomedical Engineering | 2015 |
| Robust Image Chroma-Keying: A Quadmap Approach Based on Global Sampling and Local Affinity | W. Wang, J. Zhao | IEEE Transactions on Broadcasting | 2015 |
| Stereoscopic Chroma Key Matting Using Statistical Analysis in CIECAM02 Color Space | L. Yin, W. Wang, J. Zhao | HAVE | 2015 |
| Gauthier, Ajay Gupta, et al | Weiwei Chen, Susan A | Quantitative Susceptibility Mapping of Multiple Sclerosis Lesions at Various Ages. Radiology | 2014 |
| Intracranial Calcifications and Hemorrhages: Characterization with Quantitative Susceptibility Mapping | Weiwei Chen, Wenzhen Zhu, Iihami Kovanlikaya, et al | Radiology | 2014 |
| Inverse Problem in Quantitative Susceptibility Mapping | Jae Kyu Choi, Hyoung Suk Park, Shuai Wang, Yi Wang, Jin Keun Seo | SIAM Journal on Imaging Sciences | 2014 |
| Low Complexity Adaptive View Synthesis Optimization in HEVC Based 3D Video Coding | Siwei Ma, Shiqi Wang, Wen Gao | IEEE Transactions on Multimedia | 2014 |
| Remote Image Fusion Based on PCA and Dual Tree Compactly Supported Shearlet Transform | Chang Duan, Qi Hong Huang, Xue Gang Wang, Shuai Wang | Journal of Information Hiding and Multimedia Signal Processing | 2014 |
| Categorization of Various Methods for Quantitative Susceptibility Mapping (QSM) and Their Noise Properties | Shuai Wang, Tian Liu, Weiwei Chen, et al | ISMRM | 2013 |
| Color Range Determination and Alpha Matting for Color Images | Z. Luo, W. Wang, J. Zhao, L. Yu | IST | 2013 |
| High-resolution Remote Sensing Image Segmentation Based on Improved RIU-LBP and SRM | Jian Cheng, Lan Li, Shuai Wang, Haijun Liu | EURASIP Journal on Wireless Communications and Networking | 2013 |
| Magnetic Susceptibility Anisotropy | Cynthia Wisnieff, Tian Liu, Pascal Spincemaille, Shuai Wang, et al | NeuroImage | 2013 |
| Mode Dependent Coding Tools for Video Coding | Siwei Ma, Shiqi Wang, Qin Yu, Junjun Si, Wen Gao | IEEE Journal of Selected Topics in Signal Processing | 2013 |
| Parallel Fast Inter Mode Decision for H.264/AVC Encoding | Chen J, Villasenor J, Luo G, He Y | Journal of Visual Communication and Image Representation | 2013 |
| A Pixel-Wise Directional Intra Prediction Method | Wang Y, Chen J, He Y | Journal of Visual Communication and Image Representation | 2012 |
| Adaptive Frequency Weighting for High Performance Video Coding | Chen J, Zheng J, Xu F, Villasenor J | IEEE TCSVT | 2012 |
| Efficient Video Coding Using Legacy Algorithmic Approaches | Chen J, Xu F, He Y, et al | IEEE Transactions on Multimedia | 2012 |
| UMHexagonS Algorithm for MPEG Type-1 Video Encoder | Li S, Chen J, He Y | Journal of Harbin Engineering University | 2011 |