<?xml version="1.0" encoding="utf-8" standalone="yes" ?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Transformer | ViLab</title>
    <link>https://vilab.team/tag/transformer/</link>
      <atom:link href="https://vilab.team/tag/transformer/index.xml" rel="self" type="application/rss+xml" />
    <description>Transformer</description>
    <generator>Hugo Blox Builder (https://hugoblox.com)</generator><language>en-us</language><lastBuildDate>Fri, 11 Apr 2025 00:00:00 +0000</lastBuildDate>
    <image>
      <url>https://vilab.team/media/icon_hu2896232876136423579.png</url>
      <title>Transformer</title>
      <link>https://vilab.team/tag/transformer/</link>
    </image>
    
    <item>
      <title>Spiking point transformer for point cloud classification</title>
      <link>https://vilab.team/publication/spiking-point-transformer-for-point-cloud-classification/</link>
      <pubDate>Fri, 11 Apr 2025 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/publication/spiking-point-transformer-for-point-cloud-classification/</guid>
      <description>&lt;p&gt;本文提出Spiking Point Transformer（SPT），首个基于Transformer的脉冲神经网络框架，用于三维点云分类。SPT设计队列驱动采样直接编码，在降低计算成本的同时保留关键支撑点；并引入混合动力学积分发放神经元（HD-IF），模拟选择性神经元激活，减少对特定人工神经元的过度依赖。在多个真实与合成点云基准上取得领先结果，理论能耗较ANN对应模型降低至少6.4倍。&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>祝贺实验室 3 项科研成果发表在 AAAI 2025！</title>
      <link>https://vilab.team/event/publication-news-021641c63fc12e96/</link>
      <pubDate>Fri, 11 Apr 2025 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/event/publication-news-021641c63fc12e96/</guid>
      <description>&lt;p&gt;热烈祝贺开大纯同学、李和倍同学、吴沛熹同学！近期，实验室共有 3 项科研成果正式发表。&lt;/p&gt;
&lt;h2 id=&#34;event-enhanced-blurry-video-super-resolution&#34;&gt;Event-enhanced blurry video super-resolution&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;祝贺开大纯同学！该论文已发表在 Proceedings of the AAAI Conference on Artificial Intelligence 39 (4), 4175-4183。&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;作者：&lt;/strong&gt; Dachun Kai、Yueyi Zhang、Jin Wang、Zeyu Xiao、Zhiwei Xiong、Xiaoyan Sun&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;发表载体：&lt;/strong&gt; Proceedings of the AAAI Conference on Artificial Intelligence 39 (4), 4175-4183&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;发表时间：&lt;/strong&gt; 2025年4月11日&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;相关链接：&lt;/strong&gt; &lt;a href=&#34;https://ojs.aaai.org/index.php/AAAI/article/view/32438&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;论文链接&lt;/a&gt; · &lt;a href=&#34;https://ojs.aaai.org/index.php/AAAI/article/download/32438/34593&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;PDF&lt;/a&gt; · &lt;a href=&#34;https://github.com/DachunKai/Ev-DeblurVSR&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;代码&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&#34;论文介绍&#34;&gt;论文介绍&lt;/h3&gt;
&lt;p&gt;In this paper, we tackle the task of blurry video super-resolution (BVSR), aiming to generate high-resolution (HR) videos from low-resolution (LR) and blurry inputs. Current BVSR methods often fail to restore sharp details at high resolutions, resulting in noticeable artifacts and jitter due to insufficient motion information for deconvolution and the lack of high-frequency details in LR frames. To address these challenges, we introduce event signals into BVSR and propose a novel event-enhanced network, Ev-DeblurVSR. To effectively fuse information from frames and events for feature deblurring, we introduce a reciprocal feature deblurring module that leverages motion information from intra-frame events to deblur frame features while reciprocally using global scene context from the frames to enhance event features. Furthermore, to enhance temporal consistency, we propose a hybrid deformable alignment module that fully exploits the complementary motion information from inter-frame events and optical flow to improve motion estimation in the deformable alignment process. Extensive evaluations demonstrate that Ev-DeblurVSR establishes a new state-of-the-art performance on both synthetic and real-world datasets. Notably, on real data, our method is 2.59 dB more accurate and 7.28× faster than the recent best BVSR baseline FMA-Net.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&#34;efficient-event-based-semantic-segmentation-via-exploiting-frame-event-fusion-a-hybrid-neural-network-approach&#34;&gt;Efficient event-based semantic segmentation via exploiting frame-event fusion: A hybrid neural network approach&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;祝贺李和倍同学！该论文已发表在 &lt;em&gt;AAAI&lt;/em&gt; 39(17)。&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;















&lt;figure  &gt;
  &lt;div class=&#34;d-flex justify-content-center&#34;&gt;
    &lt;div class=&#34;w-100&#34; &gt;&lt;img alt=&#34;Efficient event-based semantic segmentation via exploiting frame-event fusion: A hybrid neural network approach&#34; srcset=&#34;
               /event/publication-news-021641c63fc12e96/images/paper-02_hu7998433290983327509.webp 400w,
               /event/publication-news-021641c63fc12e96/images/paper-02_hu5742706483688984605.webp 760w,
               /event/publication-news-021641c63fc12e96/images/paper-02_hu5093957591575336652.webp 1200w&#34;
               src=&#34;https://vilab.team/event/publication-news-021641c63fc12e96/images/paper-02_hu7998433290983327509.webp&#34;
               width=&#34;760&#34;
               height=&#34;237&#34;
               loading=&#34;lazy&#34; data-zoomable /&gt;&lt;/div&gt;
  &lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;作者：&lt;/strong&gt; Hebei Li、Yansong Peng、Jiahui Yuan、Peixi Wu、Jin Wang、Yueyi Zhang、Xiaoyan Sun&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;发表载体：&lt;/strong&gt; &lt;em&gt;AAAI&lt;/em&gt; 39(17)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;发表时间：&lt;/strong&gt; 2025年4月11日&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;相关链接：&lt;/strong&gt; &lt;a href=&#34;https://ojs.aaai.org/index.php/AAAI/article/view/34013&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;论文链接&lt;/a&gt; · &lt;a href=&#34;https://ojs.aaai.org/index.php/AAAI/article/download/34013/36168&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;PDF&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&#34;论文介绍-1&#34;&gt;论文介绍&lt;/h3&gt;
&lt;p&gt;本文提出一种高效的混合神经网络框架，用于事件相机语义分割。该框架包含处理事件流的脉冲神经网络（SNN）分支和处理帧图像的人工神经网络（ANN）分支，并设计了自适应时间加权（ATW）注入器、事件驱动稀疏（EDS）注入器和通道选择融合（CSF）模块，以充分融合帧与事件的互补时空信息。在DDD17-Seg、DSEC-Semantic和M3ED-Semantic数据集上取得了最先进精度，并在DSEC-Semantic上降低63%能耗。&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&#34;spiking-point-transformer-for-point-cloud-classification&#34;&gt;Spiking point transformer for point cloud classification&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;祝贺吴沛熹同学！该论文已发表在 &lt;em&gt;AAAI&lt;/em&gt; 39(20)。&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;















&lt;figure  &gt;
  &lt;div class=&#34;d-flex justify-content-center&#34;&gt;
    &lt;div class=&#34;w-100&#34; &gt;&lt;img alt=&#34;Spiking point transformer for point cloud classification&#34; srcset=&#34;
               /event/publication-news-021641c63fc12e96/images/paper-03_hu7240971347687367392.webp 400w,
               /event/publication-news-021641c63fc12e96/images/paper-03_hu15235344109779677276.webp 760w,
               /event/publication-news-021641c63fc12e96/images/paper-03_hu7343380604242636640.webp 1200w&#34;
               src=&#34;https://vilab.team/event/publication-news-021641c63fc12e96/images/paper-03_hu7240971347687367392.webp&#34;
               width=&#34;760&#34;
               height=&#34;418&#34;
               loading=&#34;lazy&#34; data-zoomable /&gt;&lt;/div&gt;
  &lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;作者：&lt;/strong&gt; Peixi Wu、Bosong Chai、Hebei Li、Menghua Zheng、Yansong Peng、Zeyu Wang、Xuan Nie、Yueyi Zhang、Xiaoyan Sun&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;发表载体：&lt;/strong&gt; &lt;em&gt;AAAI&lt;/em&gt; 39(20)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;发表时间：&lt;/strong&gt; 2025年4月11日&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;相关链接：&lt;/strong&gt; &lt;a href=&#34;https://ojs.aaai.org/index.php/AAAI/article/view/35459&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;论文链接&lt;/a&gt; · &lt;a href=&#34;https://ojs.aaai.org/index.php/AAAI/article/view/35459/37614&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;PDF&lt;/a&gt; · &lt;a href=&#34;https://github.com/PeppaWu/SPT&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;代码&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&#34;论文介绍-2&#34;&gt;论文介绍&lt;/h3&gt;
&lt;p&gt;本文提出Spiking Point Transformer（SPT），首个基于Transformer的脉冲神经网络框架，用于三维点云分类。SPT设计队列驱动采样直接编码，在降低计算成本的同时保留关键支撑点；并引入混合动力学积分发放神经元（HD-IF），模拟选择性神经元激活，减少对特定人工神经元的过度依赖。在多个真实与合成点云基准上取得领先结果，理论能耗较ANN对应模型降低至少6.4倍。&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Perceptual image compression with conditional diffusion transformers</title>
      <link>https://vilab.team/publication/perceptual-image-compression-with-conditional-diffusion-tran/</link>
      <pubDate>Sun, 08 Dec 2024 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/publication/perceptual-image-compression-with-conditional-diffusion-tran/</guid>
      <description>&lt;p&gt;本文提出一种基于条件扩散模型的感知图像压缩方法，旨在解决现有生成式压缩方法性能提升有限和模型复杂度高的问题。方法采用扩散Transformer作为解码器，并利用Swin Transformer实现高效架构，以增强生成能力；同时引入多尺度特征融合模块，为解码器提供更丰富的信息特征。实验结果表明该方法在感知图像压缩任务上取得了优越的性能。&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Tmformer: Token merging transformer for brain tumor segmentation with missing modalities</title>
      <link>https://vilab.team/publication/tmformer-token-merging-transformer-for-brain-tumor-segmentat/</link>
      <pubDate>Sun, 24 Mar 2024 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/publication/tmformer-token-merging-transformer-for-brain-tumor-segmentat/</guid>
      <description>&lt;p&gt;本文提出 TMFormer，一种用于缺失模态脑肿瘤分割的 Token 合并 Transformer。该方法通过提取并合并可用模态为更紧凑的 token 序列，解决现有方法以零图填充缺失模态带来的特征偏差与冗余计算问题。其核心包括单模态 Token 合并块（UMB）和多模态 Token 合并块（MMB），分别增强单模态表示并缓解多模态融合偏差。在 BraTS 2018 和 2020 数据集上的实验表明，TMFormer 在缺失模态场景下优于现有方法。&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Attention-guided contrastive masked image modeling for transformer-based self-supervised learning</title>
      <link>https://vilab.team/publication/attention-guided-contrastive-masked-image-modeling-for-trans/</link>
      <pubDate>Sun, 08 Oct 2023 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/publication/attention-guided-contrastive-masked-image-modeling-for-trans/</guid>
      <description>&lt;p&gt;本文提出注意力引导的对比掩码图像建模方法（ACoMIM），融合对比学习与掩码图像建模两种自监督范式，并利用视觉Transformer的注意力机制提升表征能力。该方法包含两个预训练任务：一是根据注意力引导预测掩码区域的特征，二是比较掩码图像与未掩码图像的全局特征。两个任务相互补充，有效缓解了图像信息稀疏与分布不均的问题，在多种下游任务上验证了方法的有效性。&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Eoformer: Edge-oriented transformer for brain tumor segmentation</title>
      <link>https://vilab.team/publication/eoformer-edge-oriented-transformer-for-brain-tumor-segmentat/</link>
      <pubDate>Sun, 01 Oct 2023 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/publication/eoformer-edge-oriented-transformer-for-brain-tumor-segmentat/</guid>
      <description>&lt;p&gt;本文提出边缘导向Transformer（EoFormer），用于脑肿瘤MRI图像分割。该方法采用CNN-Transformer混合编码器，CNN提取局部低级特征，Transformer建模长距离依赖以生成全局高级特征；解码器集成边缘导向Sobel与Laplacian锐化模块，增强边缘信息。同时引入高效注意力与重参数化技术，提升特征表示能力与分割精度。&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Get: Group event transformer for event-based vision</title>
      <link>https://vilab.team/publication/get-group-event-transformer-for-event-based-vision/</link>
      <pubDate>Sun, 01 Oct 2023 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/publication/get-group-event-transformer-for-event-based-vision/</guid>
      <description>&lt;p&gt;本文提出一种基于分组的事件视觉Transformer骨干网络GET，用于事件相机视觉任务。GET将事件按时间戳和极性分组为Group Token，并在特征提取过程中解耦时空信息与极性信息。通过事件双自注意力模块和分组Token聚合模块，实现空间与时间-极性信息的有效通信与整合，充分利用事件数据特性，提升事件视觉任务性能。&lt;/p&gt;
</description>
    </item>
    
  </channel>
</rss>
