孙晓艳

孙晓艳

导师

孙晓艳,中国科学技术大学信息学院讲席教授、博士生导师、类脑智能技术及应用国家工程实验室副主任,国家级特聘专家。 1997年获哈尔滨工业大学计算机科学与工程学士学位,1999年获哈尔滨工业大学计算机科学与工程硕士学位,2003获哈尔滨工业大学计算机科学与工程博士学位。2019年加入中科大之前,历任微软亚洲研究院研究员、主管研究员、资深研究员。 长期专注于图像、视频处理及计算机视觉方向的基础与应用研究,研究成果以论文形式在多个相关领域顶级国际会议和刊物上发表100余篇。担任相关领域十余个著名国际会议和顶级国际期刊的技术委员、编委、评审等职务,包括IEEE JETCAS高级编委和Signal Processing: Image Communication期刊的领域编辑,并担任IEEE MSA(Multimedia Systems & Applications)技术委员会委员 (2014-2018)。获国家技术发明奖二等奖“高效数字视频编解码技术及其在国际标准与国家标准中的应用”;获北京市科学技术奖一等奖“视频编码关键技术研究及其对MPEG-4国际标准的贡献”;获IEEE TCSVT 最佳期刊论文奖(大陆首篇);获VCIP 最佳学生论文奖。

地址: 安徽省合肥市蜀山区黄山路443号科技实验楼西901(230027)

邮箱: sunxiaoyan@ustc.edu.cn

Interests
  • 视频图像处理
  • 计算机视觉
  • 人工智能
  • 类脑计算
Publications
  1. Token-Wise Attention-Guided Semantic Quality Assessment for Compressed Visual Features 2026
  2. Holo-World: Unified Camera, Object and Weather Control for Video World Model 2026
  3. Salient Diagnostic Value Perception For Preoperative Posterior Fossa Tumor Diagnosis 2026
  4. SSCM: A Spatial-Semantic Consistent Model for Multi-Contrast MRI Super-Resolution 2026
  5. Facm: Flow-anchored consistency models 2026
  6. ReactID: Synchronizing Realistic Actions and Identity in Personalized Video Generation 2026
  7. Seeing the unseen: Zooming in the dark with event cameras 2026
  8. RiO-DETR: DETR for Real-time Oriented Object Detection 2026
  9. EvTexture++: Event-Driven Texture Enhancement for Video Super-Resolution 2026
  10. Dome-DETR: DETR with density-oriented feature-query manipulation for efficient tiny object detection 2025
  11. DT-UFC: Universal large model feature coding via peaky-to-balanced distribution transformation 2025
  12. Enhancing zero-shot brain tumor subtype classification via fine-grained patch-text alignment 2025
  13. MeDKCoOp: Dual Knowledge-guided Graph Prompt Learning for Biomedical Vision-Language Models 2025
  14. Dash: 4d hash encoding with self-supervised decomposition for real-time dynamic scene rendering 2025
  15. Efficient spiking point mamba for point cloud analysis 2025
  16. Hybrid Vision Transformer and Convolutional Neural Network for Super-Resolution Image Quality Assessment 2025
  17. Vquala 2025 challenge on image super-resolution generated content quality assessment: Methods and results 2025
  18. Enhancing Visual Question Answering Via Clustered In-Context Sequence Configuration 2025
  19. MMSupcon: An image fusion-based multi-modal supervised contrastive method for brain tumor diagnosis 2025
  20. Semamil: Semantic reordering with retrieval-guided state space modeling for whole slide image classification 2025
  21. Create anything anywhere: Layout-controllable personalized diffusion model for multiple subjects 2025
  22. Incomplete multi-modal brain tumor segmentation via learnable sorting state space model 2025
  23. D-FINE: Redefine regression task of DETRs as fine-grained distribution refinement 2025
  24. Efficient event-based semantic segmentation via exploiting frame-event fusion: A hybrid neural network approach 2025
  25. Event-enhanced blurry video super-resolution 2025
  26. Spiking point transformer for point cloud classification 2025
  27. Hierarchical Task-aware Temporal Modeling and Matching for few-shot action recognition 2025
  28. Semantic-aware late-stage supervised contrastive learning for fine-grained action recognition 2025
  29. Enhancing Visual Tracking by Leveraging High-frequency Information within Event Signals 2025
  30. Visual perception by large language model’s weights 2024
  31. Feature compression with 3d sparse convolution 2024
  32. Perceptual image compression with conditional diffusion transformers 2024
  33. Multi-modal diffusion network with controllable variability for medical image segmentation 2024
  34. Asymmetric event-guided video super-resolution 2024
  35. Optimized decoupled structure with non-local attention for deep image compression 2024
  36. Semantic-enhanced point-box joint prompting for video object segmentation 2024
  37. Event-adapted video super-resolution 2024
  38. Event-based head pose estimation: Benchmark and method 2024
  39. Ee-mllm: A data-efficient and compute-efficient multimodal large language model 2024
  40. A micro-expression recognition system with event cameras 2024
  41. Estme: Event-driven spatio-temporal motion enhancement for micro-expression recognition 2024
  42. Advancing presurgical non-invasive molecular subgroup prediction in medulloblastoma using artificial intelligence and MRI signatures 2024
  43. Task navigator: Decomposing complex tasks for multimodal large language models 2024
  44. Event-assisted low-light video object segmentation 2024
  45. Microcinema: A divide-and-conquer approach for text-to-video generation 2024
  46. Scene adaptive sparse transformer for event-based object detection 2024
  47. Deep multi-threshold spiking-UNet for image processing 2024
  48. Multi-modal generative embedding model 2024
  49. Semi-supervised medical image segmentation via dynamic pseudo-label refinement 2024
  50. Understanding of facial features in face perception: insights from deep convolutional neural networks 2024
  51. Image captioning with multi-context synthetic data 2024
  52. Tmformer: Token merging transformer for brain tumor segmentation with missing modalities 2024
  53. Anatomical consistency distillation and inconsistency synthesis for brain tumor segmentation with missing modalities 2024
  54. Panacea: Panoramic and controllable video generation for autonomous driving 2024
  55. Attention-guided contrastive masked image modeling for transformer-based self-supervised learning 2023
  56. Video super-resolution via event-driven temporal alignment 2023
  57. Eoformer: Edge-oriented transformer for brain tumor segmentation 2023
  58. Get: Group event transformer for event-based vision 2023
  59. Learned rate-distortion cost prediction for ultrafast screen content intra coding 2023
  60. Multimodal sentiment analysis with preferential fusion and distance-aware contrastive learning 2023
  61. Better and faster: Adaptive event conversion for event-based object detection 2023
  62. Deep spiking-unet for image processing 2023
  63. Text-Only Image Captioning with Multi-Context Data Generation. 2023
  64. Dual progressive prototype network for generalized zero-shot learning 2021
  65. Task-independent knowledge makes for transferable representations for generalized zero-shot learning 2021
  66. Uncertainty-aware label rectification for domain adaptive mitochondria segmentation 2021
  67. VAE^ 2: Preventing Posterior Collapse of Variational Video Predictions in the Wild 2021
  68. Enriching optical flow with appearance information for action recognition 2020
  69. Posterior-guided neural architecture search 2020
  70. Spatiotemporal fusion in 3D CNNs: A probabilistic view 2020