<?xml version="1.0" encoding="utf-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.3 20210610//EN" "JATS-journalpublishing-oasis-article1-3.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="research-article" dtd-version="1.3" xml:lang="en">
<front>
<journal-meta>
<journal-id journal-id-type="nlm-ta">Intell. Robot.</journal-id>
 <journal-id journal-id-type="publisher-id">ir</journal-id>
 <journal-title-group>
<journal-title>Intelligence &#38; Robotics</journal-title>
<abbrev-journal-title>IR</abbrev-journal-title>
</journal-title-group>
<issn pub-type="epub">2770-3541</issn>
 <issn pub-type="ppub">2770-3541</issn>
 <publisher>
<publisher-name>Intelligence &#38; Robotics</publisher-name>

</publisher>
 </journal-meta>
 <article-meta>
 <article-id pub-id-type="publisher-id">IR-2026-21</article-id>
 <article-id pub-id-type="doi">10.20517/ir.2026.21</article-id>

<article-categories>
<subj-group subj-group-type="heading">
<subject>Research Article</subject>
</subj-group>
</article-categories>
 <title-group>
 <article-title>ConCast: a CBAM-enhanced SimVP with temporal consistency regularization for precipitation nowcasting</article-title>
      </title-group> 

 <contrib-group>
			  <contrib contrib-type="author">

<name>
 <surname>Zhao</surname>
 <given-names>Peihan</given-names>
 </name>

 <xref ref-type="aff" rid="aff1">1</xref>
<xref ref-type="aff" rid="I#">#</xref>
 </contrib>
 <contrib contrib-type="author">

<name>
 <surname>Wang</surname>
 <given-names>Rong</given-names>
 </name>

 <xref ref-type="aff" rid="aff2">2</xref>
<xref ref-type="aff" rid="I#">#</xref>
 </contrib>
 <contrib contrib-type="author">

<name>
 <surname>Zhang</surname>
 <given-names>Yiye</given-names>
 </name>

 <xref ref-type="aff" rid="aff2">2</xref>
 </contrib>
 <contrib contrib-type="author" corresp="yes"><contrib-id contrib-id-type="orcid">http://orcid.org/0000-0003-2458-6774</contrib-id>

<name>
 <surname>Yang</surname>
 <given-names>Xiaofei</given-names>
 </name>
<email>xiaofeiyang@gzhu.edu.cn</email>
 <xref ref-type="aff" rid="aff3">3</xref>
 <xref ref-type="corresp" rid="cor1">&#42;</xref>
 </contrib>
<contrib contrib-type="author" corresp="yes"><contrib-id contrib-id-type="orcid">http://orcid.org/0000-0002-8419-0109</contrib-id>

<name>
 <surname>Li</surname>
 <given-names>Chunshan</given-names>
 </name>
<email>lics@hit.edu.cn</email>
 <xref ref-type="aff" rid="aff2">2</xref>
 <xref ref-type="corresp" rid="cor2">&#42;</xref>
 </contrib>
      </contrib-group>

 <aff id="aff1">
<label><sup>1</sup></label>
<addr-line>Digitalization Division, Department of Science and Information Technology, China Energy Investment Corporation Co., Ltd., Beijing 100011, China.</addr-line>
</aff>
<aff id="aff2">
<label><sup>2</sup></label>
<addr-line>School of Computer Science and Technology, Harbin Institute of Technology, Weihai 264209, Shandong, China.</addr-line>
</aff>
<aff id="aff3">
<label><sup>3</sup></label>
<addr-line>School of Electronic and Communication Engineering, Guangzhou University, Guangzhou 510182, Guangdong, China.</addr-line>
</aff>
<aff id="I#">
<label><sup>#</sup></label>
<addr-line>Authors contributed equally to this work and share first authorship.</addr-line>
</aff>

 <author-notes>

 <corresp id="cor1">Correspondence to: Prof. Xiaofei Yang, School of Electronic and Communication Engineering, Guangzhou University, Guangzhou 510182, Guangdong, China. E-mail: <email>xiaofeiyang@gzhu.edu.cn</email>; Prof. Chunshan Li, School of Computer Science and Technology, Harbin Institute of Technology, Weihai 264209, Shandong, China. E-mail: <email>lics@hit.edu.cn</email></corresp>
<fn fn-type="other"><p><bold>Received:</bold> 18 Apr 2026 | <bold>First Decision:</bold> 8 Jun 2026 | <bold>Revised:</bold> 19 Jun 2026 | <bold>Accepted:</bold> 15 Jul 2026 | <bold>Published:</bold> 31 Jul 2026</p>
</fn>
<fn fn-type="other"><p><bold>Academic Editor:</bold> Simon Yang | <bold>Copy Editor:</bold> Pei-Yun Wang | <bold>Production Editor:</bold> Pei-Yun Wang</p>
</fn>


</author-notes>
      <pub-date pub-type="ppub">
        <year>2026</year>
      </pub-date>
      <pub-date pub-type="epub">
        <day>31</day>
        <month>7</month>
        <year>2026</year>
      </pub-date>
      <volume>6</volume>
 <issue>3</issue>
 <fpage>427</fpage>
 <lpage>43</lpage>
      <permissions>
        <copyright-statement>© The Author(s) 2026.</copyright-statement>
        <license xlink:href="https://creativecommons.org/licenses/by/4.0/">
          <license-p>© The Author(s) 2026. <bold>Open Access</bold> This article is licensed under a Creative Commons Attribution 4.0 International License (<uri xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</uri>), which permits unrestricted use, sharing, adaptation, distribution and reproduction in any medium or format, for any purpose, even commercially, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons license, and indicate if changes were made.</license-p>
        </license>
      </permissions>


<abstract><p>Precipitation nowcasting is essential for weather warning and rapid-response decision-making, yet existing deep spatiotemporal models often struggle to emphasize meteorologically salient echo structures and to keep their predictions coherent across future frames. We address these issues with ConCast, an enhanced SimVP for radar-based precipitation nowcasting that couples convolutional block attention module (CBAM) refinement with temporal consistency regularization (TCR). The CBAM stage strengthens channel-wise and spatial feature selection so that the network attends to informative precipitation patterns, whereas an auxiliary cosine consistency term constrains the relationship between successive predictions during training. On the Shanghai-Radar and SEVIR datasets, ConCast improves forecasting quality over representative baselines and produces sharper and temporally steadier echo fields. Ablation studies show that the two components act on different aspects of the problem, with attention sharpening spatial discrimination and the consistency term improving sequence coherence, and that their gains do not overlap. Because ConCast adds only marginal overhead to SimVP, attention refinement combined with TCR is a practical option for nowcasting.</p></abstract>

<kwd-group kwd-group-type="author-created">
		<kwd>Precipitation nowcasting</kwd>
		<kwd>very short-term precipitation forecasting</kwd>
		<kwd>SimVP</kwd>
		<kwd>radar echo data</kwd>
		<kwd>temporal consistency regularization</kwd>
		<kwd>attention mechanisms</kwd>
      </kwd-group>

</article-meta>
</front>

<body>

<sec id="s1">
<label>1</label><title>1. INTRODUCTION</title>
<p>Very short-term precipitation forecasting, commonly referred to as precipitation nowcasting, aims to predict the evolution of rainfall over the next few minutes to several hours. Reliable nowcasting is particularly important in densely populated regions, where sudden convective events can rapidly trigger traffic disruption, urban flooding, and operational risks in transportation and energy systems. Traditional numerical weather prediction (NWP)<sup>[<xref ref-type="bibr" rid="b1">1</xref>]</sup> provides physically grounded forecasts, while radar echo extrapolation methods, such as tracking radar echoes by correlation (TREC)<sup>[<xref ref-type="bibr" rid="b2">2</xref>]</sup> and optical-flow-based nowcasting<sup>[<xref ref-type="bibr" rid="b3">3</xref>]</sup>, remain important tools for real-time operational forecasting. Seamless or blending systems further combine NWP guidance with extrapolation to exploit their complementary strengths across lead times<sup>[<xref ref-type="bibr" rid="b4">4</xref>]</sup>. In contrast, deep learning methods can learn precipitation dynamics directly from historical observations, enabling fast inference and flexible deployment.</p>

<p>Deep learning has made substantial progress in spatiotemporal sequence forecasting. Recurrent neural network (RNN)-based models, such as ConvLSTM<sup>[<xref ref-type="bibr" rid="b5">5</xref>]</sup>, TrajGRU<sup>[<xref ref-type="bibr" rid="b6">6</xref>]</sup>, PredRNN<sup>[<xref ref-type="bibr" rid="b7">7</xref>]</sup>, PredRNN++<sup>[<xref ref-type="bibr" rid="b8">8</xref>]</sup>, MIM<sup>[<xref ref-type="bibr" rid="b9">9</xref>]</sup>, SwinLSTM<sup>[<xref ref-type="bibr" rid="b10">10</xref>]</sup>, PhyDNet<sup>[<xref ref-type="bibr" rid="b11">11</xref>]</sup>, and MAU<sup>[<xref ref-type="bibr" rid="b12">12</xref>]</sup>, process images sequentially and preserve temporal continuity through hidden-state transitions. These models are effective for local motion modeling, but their recurrent nature can weaken long-range dependency learning and increase optimization difficulty over extended forecasting horizons. Related video prediction studies, including action-conditioned and stochastic formulations<sup>[<xref ref-type="bibr" rid="b13">13</xref>-<xref ref-type="bibr" rid="b16">16</xref>]</sup>, likewise emphasize the importance of motion realism and multi-step stability.</p>

<p>More recent direct prediction architectures, such as SimVP<sup>[<xref ref-type="bibr" rid="b17">17</xref>]</sup>, Tau<sup>[<xref ref-type="bibr" rid="b18">18</xref>]</sup>, and Earthformer<sup>[<xref ref-type="bibr" rid="b19">19</xref>]</sup>, jointly model multiple historical frames and directly decode future sequences. In the broader weather forecasting literature, large-context and foundation-style models such as MetNet<sup>[<xref ref-type="bibr" rid="b20">20</xref>]</sup>, its 12-hour forecasting variant<sup>[<xref ref-type="bibr" rid="b21">21</xref>]</sup>, FourCastNet<sup>[<xref ref-type="bibr" rid="b22">22</xref>]</sup>, Pangu-Weather<sup>[<xref ref-type="bibr" rid="b23">23</xref>]</sup>, GraphCast<sup>[<xref ref-type="bibr" rid="b24">24</xref>]</sup>, and ClimaX<sup>[<xref ref-type="bibr" rid="b25">25</xref>]</sup> have further demonstrated the scalability of data-driven forecasting across longer horizons and larger spatial domains. Such models are computationally attractive and easier to parallelize, yet they may still produce overly smooth predictions or lose sharp local structures, particularly when the model fails to emphasize the most relevant echo regions. SimVP is selected as the baseline in this work because of its simple encoder-mid-decoder design, strong efficiency, and competitive performance without relying on recurrent connections or transformer self-attention.</p>

<p>Generative and probabilistic forecasting strategies have also been introduced to improve the realism of the output. Methods such as CasCast<sup>[<xref ref-type="bibr" rid="b26">26</xref>]</sup>, SV2P<sup>[<xref ref-type="bibr" rid="b27">27</xref>]</sup>, STRPM<sup>[<xref ref-type="bibr" rid="b28">28</xref>]</sup>, DGMR<sup>[<xref ref-type="bibr" rid="b29">29</xref>]</sup>, NowcastNet<sup>[<xref ref-type="bibr" rid="b30">30</xref>]</sup>, DiffCast<sup>[<xref ref-type="bibr" rid="b31">31</xref>]</sup>, and ExtremeCast<sup>[<xref ref-type="bibr" rid="b32">32</xref>]</sup> improve sharpness or uncertainty modeling by redesigning architectures and training objectives. Nevertheless, two challenges remain central in radar-based nowcasting: identifying the most informative regions in cluttered echo scenes and maintaining stable temporal transitions across successive predicted frames.</p>

<p>To address these challenges, we propose ConCast, an enhanced SimVP model for precipitation nowcasting with two complementary components:</p>

<p>&#9679; <bold>Convolutional block attention module (CBAM):</bold> CBAM<sup>[<xref ref-type="bibr" rid="b33">33</xref>]</sup> guides the model to focus on informative channel-wise and spatial features while suppressing irrelevant information.</p>

<p>&#9679; <bold>Temporal consistency regularization (TCR):</bold> A cosine consistency loss is applied across successive predictions to provide additional temporal regularization during training.</p>

<p>ConCast therefore aims to sharpen spatial discrimination and prediction coherence simultaneously while maintaining the low computational footprint of SimVP. The remainder of this paper reviews related work, describes the proposed method, reports the experimental results, and discusses their implications.</p>

<p>The main contributions of this work are summarized as follows:</p>

<p>&#9679; We propose ConCast, a lightweight precipitation nowcasting framework that enhances SimVP with CBAM-based feature refinement and TCR.</p>

<p>&#9679; We validate the proposed framework on two heterogeneous benchmarks, Shanghai-Radar and SEVIR, covering both regional radar forecasting and broader cross-scene evaluation.</p>

<p>&#9679; We provide quantitative and ablation evidence showing that the two components address different aspects of the task and together improve both forecasting quality and temporal consistency.</p>

</sec>


<sec id="s2">
<label>2</label><title>2. RELATED WORK</title>

<sec id="s2-1">
<label>2.1</label><title>2.1. Operational and learning-based precipitation nowcasting</title>
<p>Precipitation nowcasting has attracted considerable attention for its role in real-time weather prediction. Operational systems have traditionally relied on two major sources of guidance. NWP models provide physically grounded atmospheric evolution but can be less responsive at very short lead times because of initialization, spin-up, and computational constraints<sup>[<xref ref-type="bibr" rid="b1">1</xref>]</sup>. Radar echo extrapolation methods, including TREC-based motion estimation<sup>[<xref ref-type="bibr" rid="b2">2</xref>]</sup> and enhanced optical-flow techniques<sup>[<xref ref-type="bibr" rid="b3">3</xref>]</sup>, directly infer echo displacement from recent observations and are therefore widely used for immediate forecasts. In practice, seamless nowcasting systems often blend extrapolation with NWP so that observation-driven forecasts dominate early lead times and model guidance contributes increasingly at longer lead times<sup>[<xref ref-type="bibr" rid="b4">4</xref>]</sup>.</p>

<p>Learning-based methods offer an alternative by training neural models to infer precipitation evolution from historical radar sequences. In recent years, data-driven nowcasting has evolved from early sequence modeling approaches toward more expressive architectures that aim to better capture precipitation motion, local structure, and uncertainty.</p>

<p>Existing nowcasting models mainly fall into three categories. Recurrent architectures, such as ConvLSTM<sup>[<xref ref-type="bibr" rid="b5">5</xref>]</sup>, TrajGRU<sup>[<xref ref-type="bibr" rid="b6">6</xref>]</sup>, PredRNN<sup>[<xref ref-type="bibr" rid="b7">7</xref>]</sup>, PredRNN++<sup>[<xref ref-type="bibr" rid="b8">8</xref>]</sup>, and MIM<sup>[<xref ref-type="bibr" rid="b9">9</xref>]</sup>, model temporal evolution through hidden-state transitions and are effective for motion-aware prediction. Encoder-decoder approaches, including U-Net<sup>[<xref ref-type="bibr" rid="b34">34</xref>]</sup>, RainNet<sup>[<xref ref-type="bibr" rid="b35">35</xref>]</sup>, and related radar forecasting models<sup>[<xref ref-type="bibr" rid="b36">36</xref>]</sup>, preserve spatial detail through multi-scale feature extraction and skip connections. More recent large-context and generative methods, such as MetNet<sup>[<xref ref-type="bibr" rid="b20">20</xref>]</sup>, DGMR<sup>[<xref ref-type="bibr" rid="b29">29</xref>]</sup>, NowcastNet<sup>[<xref ref-type="bibr" rid="b30">30</xref>]</sup>, and DiffCast<sup>[<xref ref-type="bibr" rid="b31">31</xref>]</sup>, further improve realism and forecasting quality by enlarging the receptive context or modeling uncertainty explicitly.</p>

<p>Despite this progress, accurately identifying salient echo regions and maintaining coherent multi-frame evolution remain challenging, especially for efficient radar-based nowcasting systems. These issues are particularly relevant when a model must preserve sharp local precipitation structures without sacrificing temporal stability across future predictions.</p>

</sec>


<sec id="s2-2">
<label>2.2</label><title>2.2. Attention mechanisms in weather prediction</title>
<p>Attention mechanisms have become an important tool for improving spatiotemporal prediction models by emphasizing informative regions and suppressing irrelevant responses. In precipitation nowcasting, this capability is particularly valuable because key echo structures often occupy only a limited portion of the scene, while surrounding areas may contain clutter or weak background patterns. Attention-based designs have been explored in both recurrent and non-recurrent forecasting models. For example, SwinLSTM<sup>[<xref ref-type="bibr" rid="b10">10</xref>]</sup> incorporates transformer-style attention into recurrent prediction, Tau<sup>[<xref ref-type="bibr" rid="b18">18</xref>]</sup> introduces temporal attention for efficient spatiotemporal predictive learning, and Earthformer<sup>[<xref ref-type="bibr" rid="b19">19</xref>]</sup> further demonstrates the effectiveness of space-time attention mechanisms in Earth system forecasting.</p>

<p>Beyond global or transformer-style attention, lightweight plug-in modules have also been widely studied for efficient feature refinement. CBAM<sup>[<xref ref-type="bibr" rid="b33">33</xref>]</sup> sequentially applies channel attention and spatial attention, while ECA-Net<sup>[<xref ref-type="bibr" rid="b37">37</xref>]</sup> shows that efficient channel attention alone can provide meaningful gains with minimal additional complexity. Compared with heavier attention mechanisms, such modules are attractive for radar nowcasting because they can enhance localized precipitation features, storm boundaries, and high-value echo regions without substantially increasing computation.</p>

</sec>


<sec id="s2-3">
<label>2.3</label><title>2.3. Temporal regularization for forecasting</title>
<p>Similarity-based objectives from representation learning encode structure by preserving similarity among related samples rather than relying solely on reconstruction. In time-series settings, Zhang <italic>et al.</italic><sup>[<xref ref-type="bibr" rid="b38">38</xref>]</sup> examined what makes such objectives effective for forecasting, and SimTS<sup>[<xref ref-type="bibr" rid="b39">39</xref>]</sup> showed that simple similarity-based objectives improve forecasting-oriented sequence representations. These studies indicate that similarity constraints can help sequential models learn smoother and more discriminative temporal dynamics.</p>

<p>Temporal consistency has also been used as an explicit regularizer in video and sequence modeling. Lai <italic>et al.</italic><sup>[<xref ref-type="bibr" rid="b40">40</xref>]</sup> enforce frame-to-frame consistency to stabilize per-frame processed video; temporal cycle-consistency learning<sup>[<xref ref-type="bibr" rid="b41">41</xref>]</sup> aligns embeddings across time through a differentiable consistency objective, and TCR has been applied to improve stability in domain-adaptive video segmentation<sup>[<xref ref-type="bibr" rid="b42">42</xref>]</sup>. A common theme across these works is that constraining the relationship between neighboring frames encourages temporally coherent outputs, which is particularly relevant for radar echo forecasting, where small inconsistencies between adjacent predictions may accumulate and lead to unrealistic precipitation evolution.</p>

</sec>

</sec>


<sec id="s3">
<label>3</label><title>3. METHOD</title>

<sec id="s3-1">
<label>3.1</label><title>3.1. SimVP architecture</title>
<p>The proposed method is built upon SimVP, which learns spatial and temporal dependencies from sequential radar echo images. As shown in <xref ref-type="fig" rid="Figure1">Figure 1</xref>, the model takes a sequence of historical radar frames as input and predicts multiple future precipitation frames. The framework contains three main components: an encoder for hierarchical spatial feature extraction, a spatiotemporal module for latent dynamics modeling, and a decoder for reconstructing the future sequence. CBAM is inserted into the feature transformation pipeline to refine informative channel-wise and spatial responses. During training, the predicted sequence is supervised by a pixel-wise reconstruction loss and an auxiliary temporal consistency loss between adjacent predictions.</p>

<fig id="Figure1">
    <label>Figure 1</label>
    <caption style="columns:2;">
        <p>Overall framework of the proposed method. Historical radar frames are first encoded into hierarchical spatial features, then processed by a spatio-temporal module with CBAM refinement, and finally decoded into future precipitation frames. SimVP predicts five frames per forward pass, and four autoregressive rollouts generate the complete 20-frame forecast. Only representative frames are shown for clarity. During training, the predictions are jointly supervised by the MSE loss and the temporal consistency loss. CBAM: Convolutional block attention module; MSE: mean squared error.</p>
    </caption>
    <graphic xlink:href="ir6021.fig.1.jpg"></graphic>
</fig>



<p>More specifically, let the input radar sequence be <inline-formula><tex-math id="M1">$$ X \in \mathbb{R}^{B \times T \times C \times H \times W} $$</tex-math></inline-formula>, where <inline-formula><tex-math id="M2">$$ B $$</tex-math></inline-formula> is the batch size, <inline-formula><tex-math id="M3">$$ T $$</tex-math></inline-formula> is the number of historical frames, and <inline-formula><tex-math id="M4">$$ C $$</tex-math></inline-formula>, <inline-formula><tex-math id="M5">$$ H $$</tex-math></inline-formula>, and <inline-formula><tex-math id="M6">$$ W $$</tex-math></inline-formula> denote the channel number, height, and width, respectively. The input is first reshaped into <inline-formula><tex-math id="M7">$$ BT $$</tex-math></inline-formula> individual frames and processed by a spatial encoder. The encoder consists of <inline-formula><tex-math id="M8">$$ N_S $$</tex-math></inline-formula> stacked convolutional blocks with alternating strides generated by the pattern <inline-formula><tex-math id="M9">$$ [1, 2, 1, 2, \ldots] $$</tex-math></inline-formula>, so that spatial downsampling is introduced in every second block. Each block uses a <inline-formula><tex-math id="M10">$$ 3 \times 3 $$</tex-math></inline-formula> convolution, followed by Group Normalization and a LeakyReLU activation. This design keeps the backbone lightweight while still expanding the receptive field through progressive spatial compression.</p>

<p>Formally, after reshaping <inline-formula><tex-math id="M11">$$ X $$</tex-math></inline-formula> into <inline-formula><tex-math id="M12">$$ \tilde{X} \in \mathbb{R}^{BT \times C \times H \times W} $$</tex-math></inline-formula>, the spatial encoder produces</p>

<p><disp-formula> <label>(1)</label> <tex-math id="E1"> $$ \begin{equation} E^{(l)} = \phi_l(E^{(l-1)}), \quad l=1, \ldots, N_S, \end{equation} $$ </tex-math></disp-formula></p>

<p>where <inline-formula><tex-math id="M13">$$ E^{(0)}=\tilde{X} $$</tex-math></inline-formula> and <inline-formula><tex-math id="M14">$$ \phi_l(\cdot) $$</tex-math></inline-formula> denotes the <inline-formula><tex-math id="M15">$$ l $$</tex-math></inline-formula>-th spatial encoding block. The final latent feature is <inline-formula><tex-math id="M16">$$ E^{(N_S)} \in \mathbb{R}^{BT \times C' \times H' \times W'} $$</tex-math></inline-formula>, and the first-stage feature <inline-formula><tex-math id="M17">$$ E^{(1)} $$</tex-math></inline-formula> is retained as a shallow skip representation for decoding.</p>

<p>After spatial encoding, the latent tensor is rearranged back to the sequence form and passed to a middle spatiotemporal translator. The latent sequence is reshaped from <inline-formula><tex-math id="M18">$$ \mathbb{R}^{B \times T \times C' \times H' \times W'} $$</tex-math></inline-formula> to <inline-formula><tex-math id="M19">$$ \mathbb{R}^{B \times (TC') \times H' \times W'} $$</tex-math></inline-formula> so that temporal interaction can be modeled in a channel-compressed latent space. Each translation block first applies a <inline-formula><tex-math id="M20">$$ 1 \times 1 $$</tex-math></inline-formula> convolution to mix channel information and then uses multiple grouped convolution branches with kernel sizes <inline-formula><tex-math id="M21">$$ \{3, 5, 7, 11\} $$</tex-math></inline-formula> to capture multi-scale dynamics. The outputs of these branches are summed to obtain a richer latent representation. A U-shaped latent encoder-decoder with skip connections across depth is then used to aggregate both local and large-scale motion patterns without relying on recurrent units or self-attention.</p>

<p>Let</p>

<p><disp-formula> <label>(2)</label> <tex-math id="E2"> $$ \begin{equation} Z^{(0)} = \text{Reshape}\left(E^{(N_S)}\right) \in \mathbb{R}^{B \times (TC') \times H' \times W'}.  \end{equation} $$ </tex-math></disp-formula></p>

<p>For each multi-scale translation block, the hidden feature is computed as</p>

<p><disp-formula> <label>(3)</label> <tex-math id="E3"> $$ \begin{equation} Z^{(m)} = \sum\limits_{k \in \{3, 5, 7, 11\}} g_k\left(\psi\left(Z^{(m-1)}\right)\right), \end{equation} $$ </tex-math></disp-formula></p>

<p>where <inline-formula><tex-math id="M22">$$ \psi(\cdot) $$</tex-math></inline-formula> denotes the <inline-formula><tex-math id="M23">$$ 1 \times 1 $$</tex-math></inline-formula> channel mixing convolution and <inline-formula><tex-math id="M24">$$ g_k(\cdot) $$</tex-math></inline-formula> denotes a grouped convolution with kernel size <inline-formula><tex-math id="M25">$$ k $$</tex-math></inline-formula>. The latent encoder-decoder structure further introduces skip connections, so the decoder-side latent update can be written as</p>

<p><disp-formula> <label>(4)</label> <tex-math id="E4"> $$ \begin{equation} \hat{Z}^{(m)} = h_m\left([\hat{Z}^{(m-1)}; Z^{(M-m)}]\right), \end{equation} $$ </tex-math></disp-formula></p>

<p>where <inline-formula><tex-math id="M26">$$ [\cdot;\cdot] $$</tex-math></inline-formula> denotes channel concatenation and <inline-formula><tex-math id="M27">$$ h_m(\cdot) $$</tex-math></inline-formula> is the <inline-formula><tex-math id="M28">$$ m $$</tex-math></inline-formula>-th decoder-side multi-scale block.</p>

<p>Finally, the decoder mirrors the encoder by using the reversed stride schedule to progressively recover spatial resolution. The first-stage encoder feature is preserved as a shallow skip connection and is concatenated with the last decoder feature before the final reconstruction block. A <inline-formula><tex-math id="M29">$$ 1 \times 1 $$</tex-math></inline-formula> readout convolution then maps the hidden representation back to the predicted radar frames. Overall, the backbone comprises a spatial encoder, a multi-scale latent translator, and a skip-connected spatial decoder, providing an efficient basis for the subsequent attention and temporal-regularization enhancements.</p>

<p>The decoder output is therefore obtained as</p>

<p><disp-formula> <label>(5)</label> <tex-math id="E5"> $$ \begin{equation} D^{(l)} = \varphi_l(D^{(l-1)}), \quad l=1, \ldots, N_S-1, \end{equation} $$ </tex-math></disp-formula></p>

<p><disp-formula> <label>(6)</label> <tex-math id="E6"> $$ \begin{equation} \hat{Y} = \rho \left( \varphi_{N_S}\left([D^{(N_S-1)}; E^{(1)}]\right) \right), \end{equation} $$ </tex-math></disp-formula></p>

<p>where <inline-formula><tex-math id="M30">$$ D^{(0)}=\text{Reshape}^{-1}(\hat{Z}^{(M)}) $$</tex-math></inline-formula>, <inline-formula><tex-math id="M31">$$ \varphi_l(\cdot) $$</tex-math></inline-formula> denotes the <inline-formula><tex-math id="M32">$$ l $$</tex-math></inline-formula>-th decoder block, and <inline-formula><tex-math id="M33">$$ \rho(\cdot) $$</tex-math></inline-formula> is the final <inline-formula><tex-math id="M34">$$ 1 \times 1 $$</tex-math></inline-formula> readout convolution. The model output satisfies <inline-formula><tex-math id="M35">$$ \hat{Y} \in \mathbb{R}^{B \times T \times C \times H \times W} $$</tex-math></inline-formula>.</p>

</sec>


<sec id="s3-2">
<label>3.2</label><title>3.2. Encoder-decoder with CBAM</title>
<p>The encoder and decoder follow the general structure of an encoder-decoder network, with CBAM embedded to enhance attention to informative features. The encoder gradually downsamples the input sequence and extracts hierarchical representations, while the decoder upsamples the encoded features to generate future radar frames. This design preserves the computational simplicity of SimVP while strengthening the network's sensitivity to salient echo intensity patterns, storm boundaries, and localized precipitation cores.</p>

<p>The encoder uses repeated convolutional blocks, where stride-1 layers preserve the current spatial resolution, and stride-2 layers perform downsampling. The decoder follows the reverse process with upsampling blocks to progressively restore the original resolution. This alternating downsampling and upsampling strategy reduces computation while retaining a direct shallow skip path from the first encoder stage to the final reconstruction stage. Such a design helps recover fine precipitation structures that might otherwise be lost in the latent transformation process.</p>

<p>CBAM is applied sequentially with a channel-attention stage followed by a spatial-attention stage. Given an input feature map <inline-formula><tex-math id="M36">$$ x \in \mathbb{R}^{C \times H \times W} $$</tex-math></inline-formula>, the refined feature can be written as</p>

<p><disp-formula> <label>(7)</label> <tex-math id="E7"> $$ \begin{equation} x' = \mathcal{M}_{\text{channel}}(x) \otimes x, \end{equation} $$ </tex-math></disp-formula></p>

<p><disp-formula> <label>(8)</label> <tex-math id="E8"> $$ \begin{equation} x'' = \mathcal{M}_{\text{spatial}}(x') \otimes x', \end{equation} $$ </tex-math></disp-formula></p>

<p>where <inline-formula><tex-math id="M37">$$ \mathcal{M}_{\text{channel}}(\cdot) $$</tex-math></inline-formula> denotes channel attention, <inline-formula><tex-math id="M38">$$ \mathcal{M}_{\text{spatial}}(\cdot) $$</tex-math></inline-formula> denotes spatial attention, and <inline-formula><tex-math id="M39">$$ \otimes $$</tex-math></inline-formula> denotes element-wise multiplication.</p>

<p>For channel attention, global average pooling and global max pooling are first applied to <inline-formula><tex-math id="M40">$$ x $$</tex-math></inline-formula>, and the pooled descriptors are passed through a shared two-layer bottleneck with a ReLU activation:</p>

<p><disp-formula> <label>(9)</label> <tex-math id="E9"> $$ \begin{equation} \mathcal{M}_{\text{channel}}(x) = \sigma \left( \text{MLP}(\text{AvgPool}(x)) + \text{MLP}(\text{MaxPool}(x)) \right), \end{equation} $$ </tex-math></disp-formula></p>

<p>where <inline-formula><tex-math id="M41">$$ \sigma(\cdot) $$</tex-math></inline-formula> is the sigmoid function. In this process, the pooled channel descriptors are transformed by a shared bottleneck and then broadcast back to the original spatial resolution.</p>

<p>For spatial attention, channel-wise average pooling and max-pooling are computed on the channel-refined feature <inline-formula><tex-math id="M42">$$ x' $$</tex-math></inline-formula>, concatenated along the channel dimension, and processed by a convolution layer:</p>

<p><disp-formula> <label>(10)</label> <tex-math id="E10"> $$ \begin{equation} \mathcal{M}_{\text{spatial}}(x') = \sigma \left( f^{k \times k} \left( [\text{AvgPool}_{c}(x'); \text{MaxPool}_{c}(x')] \right) \right), \end{equation} $$ </tex-math></disp-formula></p>

<p>where <inline-formula><tex-math id="M43">$$ [\cdot ; \cdot] $$</tex-math></inline-formula> denotes channel concatenation and <inline-formula><tex-math id="M44">$$ f^{k \times k} $$</tex-math></inline-formula> is a convolution with an odd kernel size <inline-formula><tex-math id="M45">$$ k $$</tex-math></inline-formula>. This stage uses a convolution to generate the spatial attention map, which is then expanded and multiplied with the input feature to obtain the final CBAM output.</p>

</sec>


<sec id="s3-3">
<label>3.3</label><title>3.3. TCR</title>
<p>To provide additional temporal regularization during training, a cosine consistency loss is applied across multiple predicted frames. The loss is computed between adjacent outputs so that neighboring predictions remain close in representation space while the mean squared error (MSE) term still anchors each frame to the ground truth. This regularization is particularly useful for multi-step nowcasting, where small inconsistencies introduced in early predictions may accumulate and lead to unrealistic echo evolution later in the sequence.</p>

<p>The temporal consistency loss is defined as</p>

<p><disp-formula> <label>(11)</label> <tex-math id="E11"> $$ \begin{equation} \mathcal{L}_{\text{tc}} = \sum\limits_{i=1}^{T-1} \text{CosineEmbeddingLoss}(\hat{Y}_i, \hat{Y}_{i+1}), \end{equation} $$ </tex-math></disp-formula></p>

<p>where <inline-formula><tex-math id="M46">$$ \hat{Y}_i $$</tex-math></inline-formula> and <inline-formula><tex-math id="M47">$$ \hat{Y}_{i+1} $$</tex-math></inline-formula> are predictions at time steps <inline-formula><tex-math id="M48">$$ i $$</tex-math></inline-formula> and <inline-formula><tex-math id="M49">$$ i+1 $$</tex-math></inline-formula>, respectively, and <inline-formula><tex-math id="M50">$$ T $$</tex-math></inline-formula> is the total number of prediction steps. The cosine embedding loss can be written as</p>

<p><disp-formula> <label>(12)</label> <tex-math id="E12"> $$ \begin{equation} \text{CosineEmbeddingLoss}(x, y) = 1 - \frac{x \cdot y}{\|x\| \|y\|}, \end{equation} $$ </tex-math></disp-formula></p>

<p>where <inline-formula><tex-math id="M51">$$ x $$</tex-math></inline-formula> and <inline-formula><tex-math id="M52">$$ y $$</tex-math></inline-formula> are flattened feature vectors from two adjacent predictions.</p>

</sec>


<sec id="s3-4">
<label>3.4</label><title>3.4. Overall loss function</title>
<p>The overall training objective combines the MSE loss with the temporal consistency loss:</p>

<p><disp-formula> <label>(13)</label> <tex-math id="E13"> $$ \begin{equation} \mathcal{L}_{\text{total}} = \mathcal{L}_{\text{MSE}} + \lambda \cdot \mathcal{L}_{\text{tc}}, \end{equation} $$ </tex-math></disp-formula></p>

<p>where <inline-formula><tex-math id="M53">$$ \lambda $$</tex-math></inline-formula> controls the relative contribution of TCR. In all main experiments, <inline-formula><tex-math id="M54">$$ \lambda $$</tex-math></inline-formula> is set to 0.1 unless otherwise specified.</p>

<p>The reconstruction term preserves pixel-level fidelity to the target radar field, while the temporal consistency term regularizes the trajectory of the predicted sequence in feature space. As a result, the model is encouraged not only to minimize frame-wise errors, but also to generate temporally coherent future evolution.</p>

</sec>

</sec>


<sec id="s4">
<label>4</label><title>4. EXPERIMENTS</title>

<sec id="s4-1">
<label>4.1</label><title>4.1. Experimental settings</title>
<p>Following prior work, the continuous radar sequences are divided into multiple events. Each event contains 25 frames, and the first 5 frames are used to predict the subsequent 20 frames. For SEVIR, this setting corresponds to forecasting the next 100 min from the previous 25 min of observations because the temporal interval is 5 min per frame. For Shanghai-Radar, where the frame interval is 6 min, the same protocol corresponds to forecasting the next 120 min from the previous 30 min of observations. All frames are resized to <inline-formula><tex-math id="M55">$$ 128 \times 128 $$</tex-math></inline-formula> pixels before training and evaluation.</p>

<p>The model is implemented in PyTorch. All models are trained on a server equipped with an Intel(R) Xeon(R) Gold 6448H @ 2.40 GHz CPU, eight NVIDIA A40 GPUs, and 1 TB of RAM. We train each model for 200 epochs using the AdamW optimizer with an initial learning rate of <inline-formula><tex-math id="M56">$$ 1\times10^{-3} $$</tex-math></inline-formula>, a weight decay of 0.01, a cosine learning-rate decay schedule, and a total batch size of <inline-formula><tex-math id="M57">$$ 4\times8=32 $$</tex-math></inline-formula>. The temporal consistency weight is set to <inline-formula><tex-math id="M58">$$ \lambda=0.1 $$</tex-math></inline-formula> for the main results. All compared methods follow the same input-output protocol to ensure a fair comparison.</p>

</sec>


<sec id="s4-2">
<label>4.2</label><title>4.2. Evaluation metrics</title>
<p>The models are assessed using CSI, HSS, bias score (BS), LPIPS, SSIM, temporal gradient error (TGE), and MSE. For threshold-based metrics, TP, FP, TN, and FN are computed after binarizing the predicted and ground-truth radar fields with the same intensity threshold. For Shanghai-Radar, we use reflectivity thresholds of 20, 30, 35, and 40 dBZ, with 35 dBZ as the primary strong-convection threshold. For SEVIR, we follow common VIL evaluation practice and use thresholds of 16, 74, 133, 160, 181, and 219. Unless otherwise specified, the main categorical analysis reports the mean score across the corresponding threshold set. Among these metrics, the pooled CSI variants (CSI-POOL4 and CSI-POOL16) are additionally reported in the main quantitative comparison, whereas BS and TGE are used in the ablation analysis to examine forecast bias and temporal coherence, respectively.</p>

<p><disp-formula> <label>(14)</label> <tex-math id="E14"> $$ \begin{equation} \text{CSI} = \frac{TP}{TP + FN + FP}, \end{equation} $$ </tex-math></disp-formula></p>

<p>where CSI is also known as the threat score (TS),</p>

<p><disp-formula> <label>(15)</label> <tex-math id="E15"> $$ \begin{equation} \text{HSS} = \frac{2(TP \cdot TN - FN \cdot FP)}{(TP + FN)(FN + TN) + (TP + FP)(FP + TN)}, \end{equation} $$ </tex-math></disp-formula></p>

<p><disp-formula> <label>(16)</label> <tex-math id="E16"> $$ \begin{equation} \text{BS} = \frac{TP + FP}{TP + FN}, \end{equation} $$ </tex-math></disp-formula></p>

<p>where BS measures whether a model over-forecasts or under-forecasts event frequency. A value closer to 1 indicates a better balance between predicted and observed events. CSI-POOL4 and CSI-POOL16 are pooled CSI variants computed after applying spatial max-pooling with kernel sizes 4 and 16, respectively. These pooled scores relax small spatial displacement errors and evaluate whether the predicted precipitation event occurs in the neighboring region.</p>

<p><disp-formula> <label>(17)</label> <tex-math id="E17"> $$ \begin{equation} \text{TGE} = \frac{1}{T-1}\sum\limits_{t=1}^{T-1}\left\|(\hat{Y}_{t+1}-\hat{Y}_{t})-(Y_{t+1}-Y_t)\right\|_1, \end{equation} $$ </tex-math></disp-formula></p>

<p>where TGE directly measures whether the temporal change between adjacent predicted frames matches the observed temporal change. Lower TGE indicates better temporal coherence relative to the ground truth.</p>

<p><disp-formula> <label>(18)</label> <tex-math id="E18"> $$ \begin{equation} \text{LPIPS}(I_p, I_t) = \sum\limits_l w_l \left\| \phi_l(I_p) - \phi_l(I_t) \right\|_2^2, \end{equation} $$ </tex-math></disp-formula></p>

<p>where <inline-formula><tex-math id="M59">$$ \phi_l $$</tex-math></inline-formula> denotes the features extracted from the <inline-formula><tex-math id="M60">$$ l $$</tex-math></inline-formula>-th layer of a deep model,</p>

<p><disp-formula> <label>(19)</label> <tex-math id="E19"> $$ \begin{equation} \text{SSIM}(x, y) = \frac{(2\mu_x\mu_y + C_1)(2\sigma_{xy} + C_2)}{(\mu_x^2 + \mu_y^2 + C_1)(\sigma_x^2 + \sigma_y^2 + C_2)}, \end{equation} $$ </tex-math></disp-formula></p>

<p>and</p>

<p><disp-formula> <label>(20)</label> <tex-math id="E20"> $$ \begin{equation} \text{MSE} = \frac{1}{N} \sum\limits_{i=1}^{N} (x_i - y_i)^2.  \end{equation} $$ </tex-math></disp-formula></p>

</sec>


<sec id="s4-3">
<label>4.3</label><title>4.3. Datasets</title>

<sec id="s4-3-1">
<label>4.3.1</label><title>4.3.1. Shanghai-Radar dataset</title>
<p>The Shanghai-Radar dataset<sup>[<xref ref-type="bibr" rid="b43">43</xref>]</sup> was collected from the dual-polarization weather monitoring radar located in Pudong, Shanghai. It contains continuous radar echo frames acquired between October 2015 and July 2018. Each frame covers approximately <inline-formula><tex-math id="M61">$$ 501\; \text{km} \times 501\; \text{km} $$</tex-math></inline-formula> and is sampled every 6 min. For training, validation, and testing, the dataset is filtered into 25-frame precipitation events, with reflectivity values normalized to the range [0, 70] dBZ. We use 1, 534 events for training, 526 for validation, and 526 for testing. After preprocessing, every frame is resized to <inline-formula><tex-math id="M62">$$ 128 \times 128 $$</tex-math></inline-formula> pixels for model input and output.</p>

</sec>


<sec id="s4-3-2">
<label>4.3.2</label><title>4.3.2. SEVIR dataset</title>
<p>The SEVIR dataset<sup>[<xref ref-type="bibr" rid="b44">44</xref>]</sup> includes satellite imagery, NEXRAD radar mosaics, and lightning event data. In this work, we use the VIL modality for precipitation forecasting. The dataset contains more than 20, 000 severe weather events, each lasting 4 h and covering an area of <inline-formula><tex-math id="M63">$$ 384\; \text{km} \times 384\; \text{km} $$</tex-math></inline-formula>. The native VIL frames are provided on a <inline-formula><tex-math id="M64">$$ 384 \times 384 $$</tex-math></inline-formula> grid with a 5-minute temporal interval. Following the same protocol as Shanghai-Radar, we extract 25 consecutive frames from each event, use the first 5 frames as input, predict the next 20 frames, and resize all frames to <inline-formula><tex-math id="M65">$$ 128 \times 128 $$</tex-math></inline-formula> pixels.</p>

</sec>

</sec>


<sec id="s4-4">
<label>4.4</label><title>4.4. Quantitative results</title>
<p>The results on the Shanghai-Radar and SEVIR datasets are summarized in <xref ref-type="table" rid="Table1">Tables 1</xref> and <xref ref-type="table" rid="Table2">2</xref>. Across most evaluation metrics, the proposed model outperforms the baselines, indicating superior forecasting skill and greater structural fidelity.</p>

<table-wrap id="Table1">
    <label>Table 1</label>
    <caption style="columns:2;">
        <p>Performance comparison of different methods on the Shanghai-Radar dataset</p>
    </caption>

    <table>
  <thead>
    <tr>
        <td style="class:table_top_border" align="left"><bold>Method</bold></td>
        <td style="class:table_top_border" align="center"><bold>CSI <inline-formula><tex-math id="M66">$$ \uparrow $$</tex-math></inline-formula></bold></td>
        <td style="class:table_top_border" align="center"><bold>CSI-POOL4 <inline-formula><tex-math id="M67">$$ \uparrow $$</tex-math></inline-formula></bold></td>
        <td style="class:table_top_border" align="center"><bold>CSI-POOL16 <inline-formula><tex-math id="M68">$$ \uparrow $$</tex-math></inline-formula></bold></td>
        <td style="class:table_top_border" align="center"><bold>HSS <inline-formula><tex-math id="M69">$$ \uparrow $$</tex-math></inline-formula></bold></td>
        <td style="class:table_top_border" align="center"><bold>LPIPS <inline-formula><tex-math id="M70">$$ \downarrow $$</tex-math></inline-formula></bold></td>
        <td style="class:table_top_border" align="center"><bold>SSIM <inline-formula><tex-math id="M71">$$ \uparrow $$</tex-math></inline-formula></bold></td>
        <td style="class:table_top_border" align="center"><bold>MSE <inline-formula><tex-math id="M72">$$ \downarrow $$</tex-math></inline-formula></bold></td>
    </tr>
  </thead>
    <tfoot>
				<tr><td align="left" colspan="8">Higher values are better for CSI, CSI-POOL4, CSI-POOL16, HSS, and SSIM. Lower values are better for LPIPS and MSE. Bold and underlined values indicate the best and second-best results for each metric, respectively. MSE: Mean squared error.</td></tr>
			</tfoot>

  <tbody>
    <tr>
        <td style="class:table_top_border2" align="left">SimVP</td>
        <td style="class:table_top_border2" align="center">0.3614</td>
        <td style="class:table_top_border2" align="center">0.3803</td>
        <td style="class:table_top_border2" align="center">0.4257</td>
        <td style="class:table_top_border2" align="center"><underline>0.5083</underline></td>
        <td style="class:table_top_border2" align="center">0.3691</td>
        <td style="class:table_top_border2" align="center">0.7477</td>
        <td style="class:table_top_border2" align="center"><underline>29.6566</underline></td>
    </tr>
    <tr>
        <td align="left">PhyDNet</td>
        <td align="center"><underline>0.3729</underline></td>
        <td align="center"><underline>0.3875</underline></td>
        <td align="center"><underline>0.4596</underline></td>
        <td align="center">0.5054</td>
        <td align="center"><underline>0.3330</underline></td>
        <td align="center"><underline>0.7740</underline></td>
        <td align="center">32.9153</td>
    </tr>
    <tr>
        <td align="left">Tau</td>
        <td align="center">0.3652</td>
        <td align="center">0.3679</td>
        <td align="center">0.4186</td>
        <td align="center">0.5004</td>
        <td align="center">0.3730</td>
        <td align="center">0.7593</td>
        <td align="center">31.2825</td>
    </tr>
    <tr>
        <td align="left">MAU</td>
        <td align="center">0.3582</td>
        <td align="center">0.3776</td>
        <td align="center">0.4485</td>
        <td align="center">0.4923</td>
        <td align="center">0.3483</td>
        <td align="center">0.7451</td>
        <td align="center">33.1546</td>
    </tr>
    <tr>
        <td align="left">Earthformer</td>
        <td align="center">0.2222</td>
        <td align="center">0.2089</td>
        <td align="center">0.2353</td>
        <td align="center">0.3198</td>
        <td align="center">0.3726</td>
        <td align="center">0.6928</td>
        <td align="center">35.8055</td>
    </tr>
    <tr>
        <td align="left">PredRNN</td>
        <td align="center">0.1993</td>
        <td align="center">0.2158</td>
        <td align="center">0.2686</td>
        <td align="center">0.2779</td>
        <td align="center">0.4280</td>
        <td align="center">0.6532</td>
        <td align="center">54.0558</td>
    </tr>
    <tr>
        <td align="left">PredRNN++</td>
        <td align="center">0.1855</td>
        <td align="center">0.2198</td>
        <td align="center">0.3325</td>
        <td align="center">0.2983</td>
        <td align="center">0.3593</td>
        <td align="center">0.6792</td>
        <td align="center">57.6103</td>
    </tr>
    <tr>
        <td style="class:table_bottom_border" align="left">Ours</td>
        <td style="class:table_bottom_border" align="center"><bold>0.4038</bold></td>
        <td style="class:table_bottom_border" align="center"><bold>0.4384</bold></td>
        <td style="class:table_bottom_border" align="center"><bold>0.5207</bold></td>
        <td style="class:table_bottom_border" align="center"><bold>0.5419</bold></td>
        <td style="class:table_bottom_border" align="center"><bold>0.2893</bold></td>
        <td style="class:table_bottom_border" align="center"><bold>0.7824</bold></td>
        <td style="class:table_bottom_border" align="center"><bold>29.4490</bold></td>
    </tr>
  </tbody>
</table>

</table-wrap>



<table-wrap id="Table2">
    <label>Table 2</label>
    <caption style="columns:2;">
        <p>Performance comparison of different methods on the SEVIR dataset</p>
    </caption>

    <table>
  <thead>
    <tr>
        <td style="class:table_top_border" align="left"><bold>Method</bold></td>
        <td style="class:table_top_border" align="center"><bold>CSI <inline-formula><tex-math id="M73">$$ \uparrow $$</tex-math></inline-formula></bold></td>
        <td style="class:table_top_border" align="center"><bold>CSI-POOL4 <inline-formula><tex-math id="M74">$$ \uparrow $$</tex-math></inline-formula></bold></td>
        <td style="class:table_top_border" align="center"><bold>CSI-POOL16 <inline-formula><tex-math id="M75">$$ \uparrow $$</tex-math></inline-formula></bold></td>
        <td style="class:table_top_border" align="center"><bold>HSS <inline-formula><tex-math id="M76">$$ \uparrow $$</tex-math></inline-formula></bold></td>
        <td style="class:table_top_border" align="center"><bold>LPIPS <inline-formula><tex-math id="M77">$$ \downarrow $$</tex-math></inline-formula></bold></td>
        <td style="class:table_top_border" align="center"><bold>SSIM <inline-formula><tex-math id="M78">$$ \uparrow $$</tex-math></inline-formula></bold></td>
        <td style="class:table_top_border" align="center"><bold>MSE <inline-formula><tex-math id="M79">$$ \downarrow $$</tex-math></inline-formula></bold></td>
    </tr>
  </thead>
    <tfoot>
				<tr><td align="left" colspan="8">Higher values are better for CSI, CSI-POOL4, CSI-POOL16, HSS, and SSIM. Lower values are better for LPIPS and MSE. Bold and underlined values indicate the best and second-best results for each metric, respectively. MSE: Mean squared error.</td></tr>
			</tfoot>
  <tbody>
    <tr>
        <td style="class:table_top_border2" align="left">SimVP</td>
        <td style="class:table_top_border2" align="center">0.2644</td>
        <td style="class:table_top_border2" align="center"><underline>0.2797</underline></td>
        <td style="class:table_top_border2" align="center"><underline>0.3341</underline></td>
        <td style="class:table_top_border2" align="center">0.3411</td>
        <td style="class:table_top_border2" align="center">0.4032</td>
        <td style="class:table_top_border2" align="center">0.6123</td>
        <td style="class:table_top_border2" align="center">492.88</td>
    </tr>
    <tr>
        <td align="left">PhyDNet</td>
        <td align="center">0.2631</td>
        <td align="center">0.2706</td>
        <td align="center">0.3095</td>
        <td align="center">0.3388</td>
        <td align="center">0.3852</td>
        <td align="center">0.6122</td>
        <td align="center">457.98</td>
    </tr>
    <tr>
        <td align="left">Tau</td>
        <td align="center">0.2578</td>
        <td align="center">0.2651</td>
        <td align="center">0.3072</td>
        <td align="center">0.3322</td>
        <td align="center">0.4205</td>
        <td align="center">0.5095</td>
        <td align="center">472.98</td>
    </tr>
    <tr>
        <td align="left">MAU</td>
        <td align="center">0.2638</td>
        <td align="center">0.2736</td>
        <td align="center">0.3245</td>
        <td align="center">0.3402</td>
        <td align="center">0.3631</td>
        <td align="center">0.6244</td>
        <td align="center">461.52</td>
    </tr>
    <tr>
        <td align="left">Earthformer</td>
        <td align="center">0.1882</td>
        <td align="center">0.1982</td>
        <td align="center">0.2238</td>
        <td align="center">0.2341</td>
        <td align="center"><bold>0.3458</bold></td>
        <td align="center">0.5076</td>
        <td align="center">555.36</td>
    </tr>
    <tr>
        <td align="left">PredRNN</td>
        <td align="center">0.1622</td>
        <td align="center">0.1724</td>
        <td align="center">0.2156</td>
        <td align="center">0.2092</td>
        <td align="center">0.4116</td>
        <td align="center">0.5535</td>
        <td align="center">739.14</td>
    </tr>
    <tr>
        <td align="left">PredRNN++</td>
        <td align="center"><underline>0.2663</underline></td>
        <td align="center">0.2768</td>
        <td align="center">0.3258</td>
        <td align="center"><underline>0.3430</underline></td>
        <td align="center">0.3598</td>
        <td align="center"><underline>0.6297</underline></td>
        <td align="center"><underline>454.46</underline></td>
    </tr>
    <tr>
        <td style="class:table_bottom_border" align="left">Ours</td>
        <td style="class:table_bottom_border" align="center"><bold>0.2765</bold></td>
        <td style="class:table_bottom_border" align="center"><bold>0.2961</bold></td>
        <td style="class:table_bottom_border" align="center"><bold>0.3489</bold></td>
        <td style="class:table_bottom_border" align="center"><bold>0.3622</bold></td>
        <td style="class:table_bottom_border" align="center"><underline>0.3494</underline></td>
        <td style="class:table_bottom_border" align="center"><bold>0.6317</bold></td>
        <td style="class:table_bottom_border" align="center"><bold>445.03</bold></td>
    </tr>
  </tbody>
</table>

</table-wrap>






<p>On the Shanghai-Radar dataset, the proposed model achieves the best results on all reported metrics. Relative to the SimVP baseline, it raises CSI from 0.3614 to 0.4038 and HSS from 0.5083 to 0.5419, while lowering LPIPS from 0.3691 to 0.2893 and MSE from 29.6566 to 29.4490. The pooled scores improve in parallel (CSI-POOL4 from 0.3803 to 0.4384 and CSI-POOL16 from 0.4257 to 0.5207), and the gains over the strongest competing baselines on CSI and SSIM further indicate that the proposed design better preserves local echo structures and fine-grained precipitation patterns rather than simply smoothing the output.</p>

<p>On the SEVIR dataset, the proposed model again achieves the best performance on CSI, CSI-POOL4, CSI-POOL16, HSS, SSIM, and MSE, while obtaining the second-best LPIPS. It reaches a CSI of 0.2765 and an HSS of 0.3622, exceeding both the SimVP baseline (0.2644 and 0.3411) and the strongest competing baseline, PredRNN++ (0.2663 and 0.3430), and attains the lowest MSE of 445.03. Its LPIPS of 0.3494 trails only Earthformer (0.3458), whose categorical skill is far lower (CSI 0.1882), so the small LPIPS gap does not reflect better forecasting ability. Because SEVIR contains more diverse storm morphologies and broader scene variability than Shanghai-Radar, these results indicate that the proposed framework generalizes effectively beyond a single regional radar distribution.</p>

</sec>


<sec id="s4-5">
<label>4.5</label><title>4.5. Dataset characteristics</title>
<p>The representative examples in <xref ref-type="fig" rid="Figure2">Figures 2</xref> and <xref ref-type="fig" rid="Figure3">3</xref> illustrate the different visual characteristics of the two datasets. The Shanghai-Radar samples exhibit relatively concentrated regional echo structures, whereas SEVIR contains more diverse large-scale storm morphologies. This contrast highlights the value of evaluating the model on both a regional radar dataset and a more heterogeneous benchmark.</p>

<fig id="Figure2">
    <label>Figure 2</label>
    <caption style="columns:2;">
        <p>Visual examples from the Shanghai-Radar dataset. The samples exhibit relatively concentrated regional echo structures and localized precipitation evolution in the Shanghai area.</p>
    </caption>
    <graphic xlink:href="ir6021.fig.2.jpg"></graphic>
</fig>



<fig id="Figure3">
    <label>Figure 3</label>
    <caption style="columns:2;">
        <p>Visual examples from the SEVIR dataset using the VIL modality. The samples show more diverse storm structures and broader scene variability than Shanghai-Radar.</p>
    </caption>
    <graphic xlink:href="ir6021.fig.3.jpg"></graphic>
</fig>







</sec>


<sec id="s4-6">
<label>4.6</label><title>4.6. Ablation studies</title>
<p>To isolate the contribution of each component, we compare the SimVP baseline, SimVP with CBAM, SimVP with TCR, and the full ConCast model that combines both, as reported in <xref ref-type="table" rid="Table3">Tables 3</xref> and <xref ref-type="table" rid="Table4">4</xref>. This component-isolation protocol directly tests whether the attention module, the temporal regularization term, and their combination each contribute to final performance.</p>

<table-wrap id="Table3">
    <label>Table 3</label>
    <caption style="columns:2;">
        <p>Component-isolation ablation on the Shanghai-Radar dataset</p>
    </caption>

    <table>
  <thead>
    <tr>
        <td style="class:table_top_border" align="left"><bold>Method</bold></td>
        <td style="class:table_top_border" align="center"><bold>CSI <inline-formula><tex-math id="M80">$$ \uparrow $$</tex-math></inline-formula></bold></td>
        <td style="class:table_top_border" align="center"><bold>HSS <inline-formula><tex-math id="M81">$$ \uparrow $$</tex-math></inline-formula></bold></td>
        <td style="class:table_top_border" align="center"><bold>BS <inline-formula><tex-math id="M82">$$ \rightarrow  1 $$</tex-math></inline-formula></bold></td>
        <td style="class:table_top_border" align="center"><bold>TGE <inline-formula><tex-math id="M83">$$ \downarrow $$</tex-math></inline-formula></bold></td>
        <td style="class:table_top_border" align="center"><bold>LPIPS <inline-formula><tex-math id="M84">$$ \downarrow $$</tex-math></inline-formula></bold></td>
        <td style="class:table_top_border" align="center"><bold>SSIM <inline-formula><tex-math id="M85">$$ \uparrow $$</tex-math></inline-formula></bold></td>
        <td style="class:table_top_border" align="center"><bold>MSE <inline-formula><tex-math id="M86">$$ \downarrow $$</tex-math></inline-formula></bold></td>
    </tr>
  </thead>
    <tfoot>
				<tr><td align="left" colspan="8">Bold values indicate the best result for each metric. BS: Bias score; TGE: temporal gradient error; MSE: mean squared error; CBAM: convolutional block attention module; TCR: temporal consistency regularization.</td></tr>
			</tfoot>
  <tbody>
    <tr>
        <td style="class:table_top_border2" align="left">SimVP</td>
        <td style="class:table_top_border2" align="center">0.3614</td>
        <td style="class:table_top_border2" align="center">0.5083</td>
        <td style="class:table_top_border2" align="center">1.12</td>
        <td style="class:table_top_border2" align="center">4.86</td>
        <td style="class:table_top_border2" align="center">0.3691</td>
        <td style="class:table_top_border2" align="center">0.7477</td>
        <td style="class:table_top_border2" align="center">29.6566</td>
    </tr>
    <tr>
        <td align="left">SimVP+CBAM</td>
        <td align="center">0.3876</td>
        <td align="center">0.5298</td>
        <td align="center">1.06</td>
        <td align="center">4.68</td>
        <td align="center">0.3157</td>
        <td align="center">0.7758</td>
        <td align="center">29.5842</td>
    </tr>
    <tr>
        <td align="left">SimVP+TCR</td>
        <td align="center">0.3789</td>
        <td align="center">0.5236</td>
        <td align="center">1.04</td>
        <td align="center">4.31</td>
        <td align="center">0.3418</td>
        <td align="center">0.7645</td>
        <td align="center">29.5317</td>
    </tr>
    <tr>
        <td style="class:table_bottom_border" align="left">Ours</td>
        <td style="class:table_bottom_border" align="center"><bold>0.4038</bold></td>
        <td style="class:table_bottom_border" align="center"><bold>0.5419</bold></td>
        <td style="class:table_bottom_border" align="center"><bold>1.01</bold></td>
        <td style="class:table_bottom_border" align="center"><bold>4.18</bold></td>
        <td style="class:table_bottom_border" align="center"><bold>0.2893</bold></td>
        <td style="class:table_bottom_border" align="center"><bold>0.7824</bold></td>
        <td style="class:table_bottom_border" align="center"><bold>29.4490</bold></td>
    </tr>
  </tbody>
</table>

</table-wrap>



<table-wrap id="Table4">
    <label>Table 4</label>
    <caption style="columns:2;">
        <p>Component-isolation ablation on the SEVIR dataset</p>
    </caption>

    <table>
  <thead>
    <tr>
        <td style="class:table_top_border" align="left"><bold>Method</bold></td>
        <td style="class:table_top_border" align="center"><bold>CSI <inline-formula><tex-math id="M87">$$ \uparrow $$</tex-math></inline-formula></bold></td>
        <td style="class:table_top_border" align="center"><bold>HSS <inline-formula><tex-math id="M88">$$ \uparrow $$</tex-math></inline-formula></bold></td>
        <td style="class:table_top_border" align="center"><bold>BS <inline-formula><tex-math id="M89">$$ \rightarrow  1 $$</tex-math></inline-formula></bold></td>
        <td style="class:table_top_border" align="center"><bold>TGE <inline-formula><tex-math id="M90">$$ \downarrow $$</tex-math></inline-formula></bold></td>
        <td style="class:table_top_border" align="center"><bold>LPIPS <inline-formula><tex-math id="M91">$$ \downarrow $$</tex-math></inline-formula></bold></td>
        <td style="class:table_top_border" align="center"><bold>SSIM <inline-formula><tex-math id="M92">$$ \uparrow $$</tex-math></inline-formula></bold></td>
        <td style="class:table_top_border" align="center"><bold>MSE <inline-formula><tex-math id="M93">$$ \downarrow $$</tex-math></inline-formula></bold></td>
    </tr>
  </thead>
    <tfoot>
				<tr><td align="left" colspan="8">Bold values indicate the best result for each metric. BS: Bias score; TGE: temporal gradient error; MSE: mean squared error; CBAM: convolutional block attention module; TCR: temporal consistency regularization.</td></tr>
			</tfoot>
  <tbody>
    <tr>
        <td style="class:table_top_border2" align="left">SimVP</td>
        <td style="class:table_top_border2" align="center">0.2644</td>
        <td style="class:table_top_border2" align="center">0.3411</td>
        <td style="class:table_top_border2" align="center">1.15</td>
        <td style="class:table_top_border2" align="center">18.92</td>
        <td style="class:table_top_border2" align="center">0.4032</td>
        <td style="class:table_top_border2" align="center">0.6123</td>
        <td style="class:table_top_border2" align="center">492.88</td>
    </tr>
    <tr>
        <td align="left">SimVP+CBAM</td>
        <td align="center">0.2718</td>
        <td align="center">0.3547</td>
        <td align="center">1.09</td>
        <td align="center">18.21</td>
        <td align="center">0.3668</td>
        <td align="center">0.6264</td>
        <td align="center">462.37</td>
    </tr>
    <tr>
        <td align="left">SimVP+TCR</td>
        <td align="center">0.2696</td>
        <td align="center">0.3509</td>
        <td align="center">1.07</td>
        <td align="center">17.35</td>
        <td align="center">0.3821</td>
        <td align="center">0.6218</td>
        <td align="center">468.54</td>
    </tr>
    <tr>
        <td style="class:table_bottom_border" align="left">Ours</td>
        <td style="class:table_bottom_border" align="center"><bold>0.2765</bold></td>
        <td style="class:table_bottom_border" align="center"><bold>0.3622</bold></td>
        <td style="class:table_bottom_border" align="center"><bold>1.03</bold></td>
        <td style="class:table_bottom_border" align="center"><bold>16.88</bold></td>
        <td style="class:table_bottom_border" align="center"><bold>0.3494</bold></td>
        <td style="class:table_bottom_border" align="center"><bold>0.6317</bold></td>
        <td style="class:table_bottom_border" align="center"><bold>445.03</bold></td>
    </tr>
  </tbody>
</table>

</table-wrap>













<p>The component-isolation results in <xref ref-type="table" rid="Table3">Tables 3</xref> and <xref ref-type="table" rid="Table4">4</xref> let us attribute the gains to each module separately. Relative to the SimVP baseline, adding CBAM yields a larger improvement in spatial skill and structural fidelity: on Shanghai-Radar, it lifts CSI from 0.3614 to 0.3876 and cuts LPIPS from 0.3691 to 0.3157, clearly exceeding the CSI of 0.3789 obtained by adding TCR alone. Adding TCR instead yields a pronounced reduction in TGE, from 4.86 to 4.31 <italic>vs.</italic> 4.68 for CBAM on Shanghai-Radar and from 18.92 to 17.35 <italic>vs.</italic> 18.21 on SEVIR, and brings the BS closer to one. The two effects are therefore largely complementary rather than redundant: CBAM sharpens where echoes are placed, whereas TCR constrains how predictions evolve between adjacent frames. The full model combines both effects and attains the best balanced performance on both datasets, achieving the highest CSI and HSS, the lowest LPIPS and TGE, and a BS closest to one. The SimVP+TCR variant, in particular, improves TGE over both SimVP and SimVP+CBAM, which isolates the contribution of the temporal regularization term independently of the attention module.</p>

</sec>


<sec id="s4-7">
<label>4.7</label><title>4.7. Lead-time-wise evaluation</title>
<p>Because forecast skill in nowcasting degrades with longer lead times, we report the lead-time-wise performance across all 20 forecast steps in <xref ref-type="fig" rid="Figure4">Figure 4</xref>. All metrics deteriorate monotonically as the horizon extends, confirming that the later steps are intrinsically harder. On Shanghai-Radar, the proposed model holds a clear and consistent advantage over SimVP in both CSI and HSS throughout the sequence, and the gap widens at the longer lead times where the baseline degrades fastest, while the two MSE curves remain close. On SEVIR, the CSI curves of the two models nearly overlap, but the proposed model retains a small edge in HSS and, more notably, yields a visibly lower MSE that grows with the lead time, indicating that the temporal regularization mainly suppresses the accumulation of pixel-level error at the far horizon.</p>

<fig id="Figure4">
    <label>Figure 4</label>
    <caption style="columns:2;">
        <p>Lead-time-wise performance across all 20 forecast steps on Shanghai-Radar and SEVIR. Higher CSI and HSS are better, while lower MSE is better. MSE: Mean squared error.</p>

    </caption>
    <graphic xlink:href="ir6021.fig.4.jpg"></graphic>
</fig>




</sec>


<sec id="s4-8">
<label>4.8</label><title>4.8. Sensitivity to the temporal consistency weight</title>
<p>The temporal consistency weight <inline-formula><tex-math id="M94">$$ \lambda $$</tex-math></inline-formula> in Equation (13) balances the reconstruction term against the temporal regularization term. We analyze its effect on the Shanghai-Radar dataset in <xref ref-type="table" rid="Table5">Table 5</xref>. Skill scores improve as <inline-formula><tex-math id="M95">$$ \lambda $$</tex-math></inline-formula> increases from zero and peak around <inline-formula><tex-math id="M96">$$ \lambda=0.1 $$</tex-math></inline-formula>, the value used in the main results. Beyond this point, the BS keeps decreasing below one, and TGE starts to rise again while CSI, HSS, and SSIM degrade, consistent with the concern that an overly strong consistency penalty over-smooths adjacent predictions and freezes the predicted echo evolution. A moderate value of <inline-formula><tex-math id="M97">$$ \lambda=0.1 $$</tex-math></inline-formula> therefore provides the best trade-off between forecast skill and temporal coherence.</p>

<table-wrap id="Table5">
    <label>Table 5</label>
    <caption style="columns:2;">
        <p>Sensitivity analysis of <inline-formula><tex-math id="M98">$$ \lambda $$</tex-math></inline-formula> on the Shanghai-Radar dataset</p>
    </caption>

    <table>
  <thead>
    <tr>
        <td style="class:table_top_border" align="center"><bold><inline-formula><tex-math id="M99">$$ \lambda $$</tex-math></inline-formula></bold></td>
        <td style="class:table_top_border" align="center"><bold>CSI <inline-formula><tex-math id="M100">$$ \uparrow $$</tex-math></inline-formula></bold></td>
        <td style="class:table_top_border" align="center"><bold>HSS <inline-formula><tex-math id="M101">$$ \uparrow $$</tex-math></inline-formula></bold></td>
        <td style="class:table_top_border" align="center"><bold>BS <inline-formula><tex-math id="M102">$$ \rightarrow  1 $$</tex-math></inline-formula></bold></td>
        <td style="class:table_top_border" align="center"><bold>TGE <inline-formula><tex-math id="M103">$$ \downarrow $$</tex-math></inline-formula></bold></td>
        <td style="class:table_top_border" align="center"><bold>SSIM <inline-formula><tex-math id="M104">$$ \uparrow $$</tex-math></inline-formula></bold></td>
        <td style="class:table_top_border" align="center"><bold>MSE <inline-formula><tex-math id="M105">$$ \downarrow $$</tex-math></inline-formula></bold></td>
    </tr>
  </thead>
    <tfoot>
				<tr><td align="left" colspan="7">Moderate temporal consistency improves coherence, while overly large values over-smooth adjacent predictions. Bold values indicate the best result for each metric. BS: Bias score; TGE: temporal gradient error; MSE: mean squared error.</td></tr>
			</tfoot>
  <tbody>
    <tr>
        <td style="class:table_top_border2" align="center">0</td>
        <td style="class:table_top_border2" align="center">0.3876</td>
        <td style="class:table_top_border2" align="center">0.5298</td>
        <td style="class:table_top_border2" align="center">1.06</td>
        <td style="class:table_top_border2" align="center">4.68</td>
        <td style="class:table_top_border2" align="center">0.7758</td>
        <td style="class:table_top_border2" align="center">29.5842</td>
    </tr>
    <tr>
        <td align="center">0.01</td>
        <td align="center">0.3941</td>
        <td align="center">0.5356</td>
        <td align="center">1.04</td>
        <td align="center">4.43</td>
        <td align="center">0.7786</td>
        <td align="center">29.5119</td>
    </tr>
    <tr>
        <td align="center">0.05</td>
        <td align="center">0.4002</td>
        <td align="center">0.5397</td>
        <td align="center">1.02</td>
        <td align="center">4.25</td>
        <td align="center">0.7813</td>
        <td align="center">29.4725</td>
    </tr>
    <tr>
        <td align="center">0.1</td>
        <td align="center"><bold>0.4038</bold></td>
        <td align="center"><bold>0.5419</bold></td>
        <td align="center"><bold>1.01</bold></td>
        <td align="center"><bold>4.18</bold></td>
        <td align="center"><bold>0.7824</bold></td>
        <td align="center"><bold>29.4490</bold></td>
    </tr>
    <tr>
        <td align="center">0.2</td>
        <td align="center">0.3994</td>
        <td align="center">0.5384</td>
        <td align="center">0.98</td>
        <td align="center">4.22</td>
        <td align="center">0.7807</td>
        <td align="center">29.6038</td>
    </tr>
    <tr>
        <td align="center">0.5</td>
        <td align="center">0.3916</td>
        <td align="center">0.5301</td>
        <td align="center">0.96</td>
        <td align="center">4.47</td>
        <td align="center">0.7742</td>
        <td align="center">30.0846</td>
    </tr>
    <tr>
        <td style="class:table_bottom_border" align="center">1.0</td>
        <td style="class:table_bottom_border" align="center">0.3769</td>
        <td style="class:table_bottom_border" align="center">0.5163</td>
        <td style="class:table_bottom_border" align="center">0.91</td>
        <td style="class:table_bottom_border" align="center">4.93</td>
        <td style="class:table_bottom_border" align="center">0.7628</td>
        <td style="class:table_bottom_border" align="center">31.2154</td>
    </tr>
  </tbody>
</table>

</table-wrap>



</sec>


<sec id="s4-9">
<label>4.9</label><title>4.9. Efficiency comparison</title>
<p>Since real-time nowcasting systems require both accuracy and efficiency, we further compare model complexity and inference cost. The results in <xref ref-type="table" rid="Table6">Table 6</xref> show that the proposed model introduces almost no increase in parameter count, FLOPs, or inference time relative to the SimVP baseline, while increasing GPU memory usage because of the additional attention operations. Even with this overhead, the model stays competitive with most recurrent and transformer-based alternatives, so the added attention does not undermine the deployment advantages of the underlying backbone.</p>

<table-wrap id="Table6">
    <label>Table 6</label>
    <caption style="columns:2;">
        <p>Efficiency comparison of representative methods</p>
    </caption>

    <table>
  <thead>
    <tr>
        <td style="class:table_top_border" align="left"><bold>Method</bold></td>
        <td style="class:table_top_border" align="center"><bold>Params</bold></td>
        <td style="class:table_top_border" align="center"><bold>FLOPs</bold></td>
        <td style="class:table_top_border" align="center"><bold>Inference time (ms/sample) <inline-formula><tex-math id="M106">$$ \downarrow $$</tex-math></inline-formula></bold></td>
        <td style="class:table_top_border" align="center"><bold>FPS <inline-formula><tex-math id="M107">$$ \uparrow $$</tex-math></inline-formula></bold></td>
        <td style="class:table_top_border" align="center"><bold>GPU memory (MB) <inline-formula><tex-math id="M108">$$ \downarrow $$</tex-math></inline-formula></bold></td>
    </tr>
  </thead>
    <tfoot>
				<tr><td align="left" colspan="6">FPS denotes the number of predicted samples per second during inference under the same hardware setting. Lower values are better for Params, FLOPs, inference time, and GPU memory. Higher values are better for FPS. Bold and underlined values indicate the best and second-best results for each efficiency metric, respectively.</td></tr>
			</tfoot>
  <tbody>
    <tr>
        <td style="class:table_top_border2" align="left">PhyDNet</td>
        <td style="class:table_top_border2" align="center">3.09M</td>
        <td style="class:table_top_border2" align="center">286.21G</td>
        <td style="class:table_top_border2" align="center">57.19</td>
        <td style="class:table_top_border2" align="center">17.49</td>
        <td style="class:table_top_border2" align="center"><bold>28.38</bold></td>
    </tr>
    <tr>
        <td align="left">PredRNN</td>
        <td align="center">23.84M</td>
        <td align="center">585.86G</td>
        <td align="center">56.60</td>
        <td align="center">17.67</td>
        <td align="center">129.22</td>
    </tr>
    <tr>
        <td align="left">PredRNN++</td>
        <td align="center">38.58M</td>
        <td align="center">867.72G</td>
        <td align="center">76.41</td>
        <td align="center">13.09</td>
        <td align="center">279.21</td>
    </tr>
    <tr>
        <td align="left">SimVP</td>
        <td align="center">14.87M</td>
        <td align="center">95.94G</td>
        <td align="center"><underline>49.50</underline></td>
        <td align="center"><underline>20.20</underline></td>
        <td align="center">182.03</td>
    </tr>
    <tr>
        <td align="left">MAU</td>
        <td align="center">7.63M</td>
        <td align="center">103.43G</td>
        <td align="center">110.57</td>
        <td align="center">9.04</td>
        <td align="center">230.52</td>
    </tr>
    <tr>
        <td align="left">Tau</td>
        <td align="center">11.74M</td>
        <td align="center">82.89G</td>
        <td align="center"><bold>24.79</bold></td>
        <td align="center"><bold>40.34</bold></td>
        <td align="center"><underline>178.35</underline></td>
    </tr>
    <tr>
        <td align="left">Earthformer</td>
        <td align="center"><bold>912.38K</bold></td>
        <td align="center"><bold>22.69G</bold></td>
        <td align="center">289.91</td>
        <td align="center">3.45</td>
        <td align="center">2904.55</td>
    </tr>
    <tr>
        <td style="class:table_bottom_border" align="left">Ours</td>
        <td style="class:table_bottom_border" align="center">14.87M</td>
        <td style="class:table_bottom_border" align="center">95.94G</td>
        <td style="class:table_bottom_border" align="center">49.84</td>
        <td style="class:table_bottom_border" align="center">20.06</td>
        <td style="class:table_bottom_border" align="center">246.87</td>
    </tr>
  </tbody>
</table>

</table-wrap>



</sec>

</sec>


<sec id="s5">
<label>5</label><title>5. DISCUSSION</title>
<p>The experiments indicate that CBAM-based feature refinement and TCR contribute in different ways. The gains in CSI and HSS point to CBAM helping the model concentrate on meteorologically informative echo regions, whereas the temporal consistency loss mainly improves the coherence of adjacent predictions. The two therefore address separate aspects of the nowcasting problem rather than reinforcing the same one.</p>

<p>The method also remains inexpensive relative to heavier recurrent and transformer-based alternatives. This matters for operational nowcasting, where inference speed and deployment simplicity often weigh as much as forecast quality. Because ConCast extends an efficient backbone instead of redesigning it, it keeps most of the computational advantage of SimVP while still improving accuracy.</p>

<p>The study also has several limitations. First, the method relies solely on radar-derived image sequences and does not incorporate auxiliary meteorological information such as wind fields, topography, or multi-source satellite observations. Second, the temporal consistency term is defined only between adjacent predictions and may therefore be insufficient for modeling longer-range temporal dependencies or uncertainty accumulation. Third, the current experiments focus on fixed-resolution inputs and benchmark settings; further evaluation under operational conditions, such as missing frames, domain shifts, and extreme precipitation events, would provide a more comprehensive assessment of robustness.</p>

<p>Future work may therefore explore multi-modal meteorological fusion, stronger physics-aware regularization, and uncertainty-aware forecasting objectives. It would also be valuable to investigate adaptive temporal regularization strategies and event-focused evaluation protocols for extreme rainfall nowcasting.</p>

</sec>


<sec id="s6">
<label>6</label><title>6. CONCLUSIONS</title>
<p>This study presented ConCast, an enhanced SimVP for precipitation nowcasting that integrates CBAM-based feature refinement with TCR. On the Shanghai-Radar and SEVIR datasets, it improved forecasting skill and structural quality over representative baselines while retaining most of the underlying backbone's efficiency.</p>

<p>These results indicate that pairing attention refinement with temporal regularization is a workable way to raise nowcasting quality without sacrificing efficiency.</p>

</sec>


<sec id="s7">
<title>DECLARATIONS</title>

<sec id="s7-1">
<title>Authors' contributions</title>
<p>Made substantial contributions to the conception and design of the study, data analysis and interpretation, and drafting and revision of the manuscript: Zhao, P.; Wang, R.; Zhang, Y.; Yang, X.; Li, C.</p>

<p>All authors read and approved the final manuscript.</p>

</sec>


<sec id="s7-2">
<title>Availability of data and materials</title>
<p>The data supporting this study are publicly available. The Shanghai-Radar dataset is available from Harvard Dataverse at <ext-link ext-link-type="uri" xlink:href="https://dataverse.harvard.edu/dataset.xhtml?persistentId=doi:10.7910/DVN/2GKMQJ">https://dataverse.harvard.edu/dataset.xhtml?persistentId=doi:10.7910/DVN/2GKMQJ</ext-link> and is also mirrored at <ext-link ext-link-type="uri" xlink:href="https://drive.google.com/file/d/14JB4ElkZKHzqxIGKMFrbnY2P4zcae8RA/view">https://drive.google.com/file/d/14JB4ElkZKHzqxIGKMFrbnY2P4zcae8RA/view</ext-link>. The SEVIR dataset is available at <ext-link ext-link-type="uri" xlink:href="https://github.com/MIT-AI-Accelerator/neurips-2020-sevir">https://github.com/MIT-AI-Accelerator/neurips-2020-sevir</ext-link>, with an example tutorial at <ext-link ext-link-type="uri" xlink:href="https://github.com/MIT-AI-Accelerator/eie-sevir/blob/master/examples/SEVIR_Tutorial.ipynb">https://github.com/MIT-AI-Accelerator/eie-sevir/blob/master/examples/SEVIR_Tutorial.ipynb</ext-link>. SEVIR can also be downloaded from the public AWS S3 bucket according to the official instructions provided by the dataset authors.</p>

</sec>


<sec id="s7-3">
<title>AI and AI-assisted tools statement</title>
<p>During the preparation of this manuscript, the AI tool ChatGPT (GPT-5.5, released 2026-04-24) was used solely for language editing. The tool did not influence the study design, data collection, analysis, interpretation, or the scientific content of the work. All authors take full responsibility for the accuracy, integrity, and final content of the manuscript.</p>

</sec>


<sec id="s7-4">
<title>Financial support and sponsorship</title>
<p>This paper was funded by Natural Science Foundation of Shandong Province (No.ZR2025MS997).</p>

</sec>


<sec id="s7-5">
<title>Conflicts of interest</title>
<p>Zhao, P. is affiliated with China Energy Investment Corporation Co., Ltd., while the other authors have declared that they have no conflicts of interest.</p>

</sec>


<sec id="s7-6">
<title>Ethical approval and consent to participate</title>
<p>Not applicable.</p>

</sec>


<sec id="s7-7">
<title>Consent for publication</title>
<p>Not applicable.</p>

</sec>


<sec id="s7-8">
<title>Copyright</title>
<p>&#169; The Author(s) 2026.</p>

</sec>

</sec>

</body>

<back>

<ref-list>
      <title>References</title>
      <ref id="b1">
      <label>1</label>
        <element-citation publication-type="journal">
        <person-group person-group-type="author">
	   <name>
	    <surname>Sun</surname>
	    <given-names>J.</given-names>
	   </name>
	   <name>
	    <surname>Xue</surname>
	    <given-names>M.</given-names>
	   </name>
	   <name>
	    <surname>Wilson</surname>
	    <given-names>J. W.</given-names>
	   </name>
	  <etal/>
          </person-group>
          <article-title>Use of NWP for nowcasting convective precipitation: recent progress and challenges</article-title>
          <source>Bull. Am. Meteorol. Soc.</source>
          <year>2014</year>
          <volume>95</volume>
          <fpage>409</fpage>
          <lpage>26</lpage>
		<pub-id pub-id-type="doi">10.1175/BAMS-D-11-00263.1</pub-id>
		 <annotation><p>Sun, J.; Xue, M.; Wilson, J. W.; et al. Use of NWP for nowcasting convective precipitation: recent progress and challenges. <italic>Bull. Am. Meteorol. Soc.</italic> <bold>2014</bold>, <italic>95</italic>, 409–26.</p></annotation></element-citation>
     </ref>

      <ref id="b2">
      <label>2</label>
        <element-citation publication-type="journal">
        <person-group person-group-type="author">
	   <name>
	    <surname>Liang</surname>
	    <given-names>Q. Q.</given-names>
	   </name>
	   <name>
	    <surname>Feng</surname>
	    <given-names>Y. R.</given-names>
	   </name>
	   <name>
	    <surname>Deng</surname>
	    <given-names>W. J.</given-names>
	   </name>
	  <etal/>
          </person-group>
          <article-title>A composite approach of radar echo extrapolation based on TREC vectors in combination with model-predicted winds</article-title>
          <source>Adv. Atmos. Sci.</source>
          <year>2010</year>
          <volume>27</volume>
          <fpage>1119</fpage>
          <lpage>30</lpage>
		<pub-id pub-id-type="doi">10.1007/s00376-009-9093-4</pub-id>
		 <annotation><p>Liang, Q. Q.; Feng, Y. R.; Deng, W. J.; et al. A composite approach of radar echo extrapolation based on TREC vectors in combination with model-predicted winds. <italic>Adv. Atmos. Sci.</italic> <bold>2010</bold>, <italic>27</italic>, 1119–30.</p></annotation></element-citation>
     </ref>

      <ref id="b3">
      <label>3</label>
        <element-citation publication-type="journal">
        <person-group person-group-type="author">
	   <name>
	    <surname>Bechini</surname>
	    <given-names>R.</given-names>
	   </name>
	   <name>
	    <surname>Chandrasekar</surname>
	    <given-names>V.</given-names>
	   </name>
          </person-group>
          <article-title>An enhanced optical flow technique for radar nowcasting of precipitation and winds</article-title>
          <source>J. Atmos. Ocean. Technol.</source>
          <year>2017</year>
          <volume>34</volume>
          <fpage>2637</fpage>
          <lpage>58</lpage>
		<pub-id pub-id-type="doi">10.1175/JTECH-D-17-0110.1</pub-id>
		 <annotation><p>Bechini, R.; Chandrasekar, V. An enhanced optical flow technique for radar nowcasting of precipitation and winds. <italic>J. Atmos. Ocean. Technol.</italic> <bold>2017</bold>, <italic>34</italic>, 2637–58.</p></annotation></element-citation>
     </ref>

      <ref id="b4">
      <label>4</label>
        <element-citation publication-type="journal">
        <person-group person-group-type="author">
	   <name>
	    <surname>Chen</surname>
	    <given-names>M. X.</given-names>
	   </name>
	   <name>
	    <surname>Qin</surname>
	    <given-names>R.</given-names>
	   </name>
	   <name>
	    <surname>Song</surname>
	    <given-names>L. Y.</given-names>
	   </name>
	  <etal/>
          </person-group>
          <article-title>SMART2022: a project supporting weather forecasting and services for the Beijing 2022 Olympic and Paralympic Winter Games - a success story of precise observations, accurate forecasting, and meticulous service in the field of meteorology</article-title>
          <source>Bull. Am. Meteorol. Soc.</source>
          <year>2025</year>
          <volume>106</volume>
          <fpage>E1401</fpage>
          <lpage>33</lpage>
		<pub-id pub-id-type="doi">10.1175/BAMS-D-24-0146.1</pub-id>
		 <annotation><p>Chen, M. X.; Qin, R.; Song, L. Y.; et al. SMART2022: a project supporting weather forecasting and services for the Beijing 2022 Olympic and Paralympic Winter Games - a success story of precise observations, accurate forecasting, and meticulous service in the field of meteorology. <italic>Bull. Am. Meteorol. Soc.</italic> <bold>2025</bold>, <italic>106</italic>, E1401–33.</p></annotation></element-citation>
     </ref>

      <ref id="b5">
      <label>5</label>
        <note><p>Shi, X.; Chen, Z.; Wang, H.; Yeung, D. Y.; Wong, W.; Woo, W. Convolutional LSTM network: a machine learning approach for precipitation nowcasting. <italic>arXiv</italic> <bold>2015</bold>, arXiv: 1506.04214. Available online: <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.48550/arXiv.1506.04214">https://doi.org/10.48550/arXiv.1506.04214</ext-link>. (accessed 20 Jul 2026)</p><p content-type="code">10.48550/arXiv.1506.04214</p></note>
     </ref>

      <ref id="b6">
      <label>6</label>
        <note><p>Shi, X.; Gao, Z.; Lausen, L.; et al. Deep learning for precipitation nowcasting: a benchmark and a new model. <italic>arXiv</italic> <bold>2017</bold>, arXiv: 1706.03458. Available online: <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.48550/arXiv.1706.03458">https://doi.org/10.48550/arXiv.1706.03458</ext-link>. (accessed 20 Jul 2026)</p><p content-type="code">10.48550/arXiv.1706.03458</p></note>
     </ref>

      <ref id="b7">
      <label>7</label>
        <note><p>Wang, Y.; Long, M.; Wang, J.; Gao, Z.; Yu, P. S. PredRNN: recurrent neural networks for predictive learning using spatiotemporal LSTMs. In <italic>Proceedings of the 31st International Conference on Neural Information Processing Systems</italic>, 2017. Curran Associates Inc., 2017; pp. 879-88.</p><p content-type="code">10.5555/3294771.3294855</p></note>
     </ref>

      <ref id="b8">
      <label>8</label>
        <note><p>Wang, Y.; Gao, Z.; Long, M.; Wang, J.; Yu, P. S. PredRNN++: towards a resolution of the deep-in-time dilemma in spatiotemporal predictive learning. <italic>arXiv</italic> <bold>2018</bold>, arXiv: 1804.06300. Available online: <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.48550/arXiv.1804.06300">https://doi.org/10.48550/arXiv.1804.06300</ext-link>. (accessed 20 Jul 2026)</p><p content-type="code">10.48550/arXiv.1804.06300</p></note>
     </ref>

      <ref id="b9">
      <label>9</label>
        <note><p>Wang, Y.; Zhang, J.; Zhu, H.; Long, M.; Wang, J.; Yu, P. S. Memory in memory: a predictive neural network for learning higher-order non-stationarity from spatiotemporal dynamics. In <italic>2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</italic>, Long Beach, USA, Jun 15-20, 2019. IEEE; 2019. pp. 9146-54.</p><p content-type="code">10.1109/CVPR.2019.00937</p></note>
     </ref>

      <ref id="b10">
      <label>10</label>
        <note><p>Tang, S.; Li, C.; Zhang, P.; Tang, R. SwinLSTM: improving spatiotemporal prediction accuracy using swin transformer and LSTM. In <italic>2023 IEEE/CVF International Conference on Computer Vision (ICCV)</italic>, Paris, France, Oct 01-06, 2023. IEEE; 2023. pp. 13424-33.</p><p content-type="code">10.1109/ICCV51070.2023.01239</p></note>
     </ref>

      <ref id="b11">
      <label>11</label>
        <note><p>Guen, V. L.; Thome, N. Disentangling physical dynamics from unknown factors for unsupervised video prediction. In <italic>2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</italic>, Seattle, USA, Jun 13-19, 2020. IEEE; 2020. pp. 11471-81.</p><p content-type="code">10.1109/CVPR42600.2020.01149</p></note>
     </ref>

      <ref id="b12">
      <label>12</label>
        <note><p>Chang, Z.; Zhang, X.; Wang, S.; et al. MAU: a motion-aware unit for video prediction and beyond. In <italic>Proceedings of the 35th International Conference on Neural Information Processing Systems</italic>, 2021. Curran Associates Inc., 2021; pp. 26950-62.</p><p content-type="code">10.5555/3540261.3542325</p></note>
     </ref>

      <ref id="b13">
      <label>13</label>
        <note><p>Finn, C.; Goodfellow, I.; Levine, S. Unsupervised learning for physical interaction through video prediction. <italic>arXiv</italic> <bold>2016</bold>, arXiv: 1605.07157. Available online: <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.48550/arXiv.1605.07157">https://doi.org/10.48550/arXiv.1605.07157</ext-link>. (accessed 20 Jul 2026)</p><p content-type="code">10.48550/arXiv.1605.07157</p></note>
     </ref>

      <ref id="b14">
      <label>14</label>
        <note><p>Lotter, W.; Kreiman, G.; Cox, D. Deep predictive coding networks for video prediction and unsupervised learning. <italic>arXiv</italic> <bold>2016</bold>, arXiv: 1605.08104. Available online: <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.48550/arXiv.1605.08104">https://doi.org/10.48550/arXiv.1605.08104</ext-link>. (accessed 20 Jul 2026)</p><p content-type="code">10.48550/arXiv.1605.08104</p></note>
     </ref>

      <ref id="b15">
      <label>15</label>
        <note><p>Villegas, R.; Yang, J.; Hong, S.; Lin, X.; Lee, H. Decomposing motion and content for natural video sequence prediction. <italic>arXiv</italic> <bold>2017</bold>, arXiv: 1706.08033. Available online: <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.48550/arXiv.1706.08033">https://doi.org/10.48550/arXiv.1706.08033</ext-link>. (accessed 20 Jul 2026)</p><p content-type="code">10.48550/arXiv.1706.08033</p></note>
     </ref>

      <ref id="b16">
      <label>16</label>
        <note><p>Denton, E.; Fergus, R. Stochastic video generation with a learned prior. <italic>arXiv</italic> <bold>2018</bold>, arXiv: 1802.07687. Available online: <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.48550/arXiv.1802.07687">https://doi.org/10.48550/arXiv.1802.07687</ext-link>. (accessed 20 Jul 2026)</p><p content-type="code">10.48550/arXiv.1802.07687</p></note>
     </ref>

      <ref id="b17">
      <label>17</label>
        <note><p>Gao, Z.; Tan, C.; Wu, L.; Li, S. Z. SimVP: simpler yet better video prediction. <italic>arXiv</italic> <bold>2022</bold>, arXiv: 2206.05099. Available online: <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.48550/arXiv.2206.05099">https://doi.org/10.48550/arXiv.2206.05099</ext-link>. (accessed 20 Jul 2026)</p><p content-type="code">10.48550/arXiv.2206.05099</p></note>
     </ref>

      <ref id="b18">
      <label>18</label>
        <note><p>Tan, C.; Gao, Z.; Wu, L.; Xu, Y.; Xia, J.; Li, S. Temporal attention unit: towards efficient spatiotemporal predictive learning. In <italic>2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</italic>, Vancouver, Canada, June 17-24, 2023. IEEE; 2023. pp. 18770-82.</p><p content-type="code">10.1109/CVPR52729.2023.01800</p></note>
     </ref>

      <ref id="b19">
      <label>19</label>
        <note><p>Gao, Z.; Shi, X.; Wang, H.; et al. Earthformer: exploring space-time transformers for earth system forecasting. <italic>arXiv</italic> <bold>2022</bold>, arXiv: 2207.05833. Available online: <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.48550/arXiv.2207.05833">https://doi.org/10.48550/arXiv.2207.05833</ext-link>. (accessed 20 Jul 2026)</p><p content-type="code">10.48550/arXiv.2207.05833</p></note>
     </ref>

      <ref id="b20">
      <label>20</label>
        <note><p>Sønderby, C. K.; Espeholt, L.; Heek, J.; et al. MetNet: a neural weather model for precipitation forecasting. <italic>arXiv</italic> <bold>2020</bold>, arXiv: 2003.12140. Available online: <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.48550/arXiv.2003.12140">https://doi.org/10.48550/arXiv.2003.12140</ext-link>. (accessed 20 Jul 2026)</p><p content-type="code">10.48550/arXiv.2003.12140</p></note>
     </ref>

      <ref id="b21">
      <label>21</label>
        <note><p>Espeholt, L.; Agrawal, S.; Sønderby, C. K.; et al. Skillful twelve hour precipitation forecasts using large context neural networks. <italic>arXiv</italic> <bold>2021</bold>, arXiv: 2111.07470. Available online: <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.48550/arXiv.2111.07470">https://doi.org/10.48550/arXiv.2111.07470</ext-link>. (accessed 20 Jul 2026)</p><p content-type="code">10.48550/arXiv.2111.07470</p></note>
     </ref>

      <ref id="b22">
      <label>22</label>
        <note><p>Pathak, J.; Subramanian, S.; Harrington, P.; et al. FourCastNet: a global data-driven high-resolution weather model using adaptive fourier neural operators. <italic>arXiv</italic> <bold>2022</bold>, arXiv: 2202.11214. Available online: <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.48550/arXiv.2202.11214">https://doi.org/10.48550/arXiv.2202.11214</ext-link>. (accessed 20 Jul 2026)</p><p content-type="code">10.48550/arXiv.2202.11214</p></note>
     </ref>

      <ref id="b23">
      <label>23</label>
        <element-citation publication-type="journal">
        <person-group person-group-type="author">
	   <name>
	    <surname>Bi</surname>
	    <given-names>K.</given-names>
	   </name>
	   <name>
	    <surname>Xie</surname>
	    <given-names>L.</given-names>
	   </name>
	   <name>
	    <surname>Zhang</surname>
	    <given-names>H.</given-names>
	   </name>
	   <name>
	    <surname>Chen</surname>
	    <given-names>X.</given-names>
	   </name>
	   <name>
	    <surname>Gu</surname>
	    <given-names>X.</given-names>
	   </name>
	   <name>
	    <surname>Tian</surname>
	    <given-names>Q.</given-names>
	   </name>
          </person-group>
          <article-title>Accurate medium-range global weather forecasting with 3D neural networks</article-title>
          <source>Nature</source>
          <year>2023</year>
          <volume>619</volume>
          <fpage>533</fpage>
          <lpage>38</lpage>
		<pub-id pub-id-type="doi">10.1038/s41586-023-06185-3</pub-id>
		 <annotation><p>Bi, K.; Xie, L.; Zhang, H.; Chen, X.; Gu, X.; Tian, Q. Accurate medium-range global weather forecasting with 3D neural networks. <italic>Nature</italic> <bold>2023</bold>, <italic>619</italic>, 533–38.</p></annotation></element-citation>
     </ref>

      <ref id="b24">
      <label>24</label>
        <note><p>Lam, R.; Sanchez-Gonzalez, A.; Willson, M.; et al. GraphCast: learning skillful medium-range global weather forecasting. <italic>arXiv</italic> <bold>2022</bold>, arXiv: 2212.12794. Available online: <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.48550/arXiv.2212.12794">https://doi.org/10.48550/arXiv.2212.12794</ext-link>. (accessed 20 Jul 2026)</p><p content-type="code">10.48550/arXiv.2212.12794</p></note>
     </ref>

      <ref id="b25">
      <label>25</label>
        <note><p>Nguyen, T.; Brandstetter, J.; Kapoor, A.; Gupta, J. K.; Grover, A. ClimaX: a foundation model for weather and climate. <italic>arXiv</italic> <bold>2023</bold>, arXiv: 2301.10343. Available online: <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.48550/arXiv.2301.10343">https://doi.org/10.48550/arXiv.2301.10343</ext-link>. (accessed 20 Jul 2026)</p><p content-type="code">10.48550/arXiv.2301.10343</p></note>
     </ref>

      <ref id="b26">
      <label>26</label>
        <note><p>Gong, J.; Bai, L.; Ye, P.; et al. CasCast: skillful high-resolution precipitation nowcasting via cascaded modelling. <italic>arXiv</italic> <bold>2024</bold>, arXiv: 2402.04290. Available online: <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.48550/arXiv.2402.04290">https://doi.org/10.48550/arXiv.2402.04290</ext-link>. (accessed 20 Jul 2026)</p><p content-type="code">10.48550/arXiv.2402.04290</p></note>
     </ref>

      <ref id="b27">
      <label>27</label>
        <note><p>Babaeizadeh, M.; Finn, C.; Erhan, D.; Campbell, R. H.; Levine, S. Stochastic variational video prediction. <italic>arXiv</italic> <bold>2017</bold>, arXiv: 1710.11252. Available online: <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.48550/arXiv.1710.11252">https://doi.org/10.48550/arXiv.1710.11252</ext-link>. (accessed 20 Jul 2026)</p><p content-type="code">10.48550/arXiv.1710.11252</p></note>
     </ref>

      <ref id="b28">
      <label>28</label>
        <note><p>Chang, Z.; Zhang, X.; Wang, S.; Ma, S.; Gao, W. STRPM: a spatiotemporal residual predictive model for high-resolution video prediction. In <italic>2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</italic>, New Orleans, USA, Jun 18-24, 2022. IEEE; 2022. pp. 13926-35.</p><p content-type="code">10.1109/CVPR52688.2022.01356</p></note>
     </ref>

      <ref id="b29">
      <label>29</label>
        <element-citation publication-type="journal">
        <person-group person-group-type="author">
	   <name>
	    <surname>Ravuri</surname>
	    <given-names>S.</given-names>
	   </name>
	   <name>
	    <surname>Lenc</surname>
	    <given-names>K.</given-names>
	   </name>
	   <name>
	    <surname>Willson</surname>
	    <given-names>M.</given-names>
	   </name>
	  <etal/>
          </person-group>
          <article-title>Skilful precipitation nowcasting using deep generative models of radar</article-title>
          <source>Nature</source>
          <year>2021</year>
          <volume>597</volume>
          <fpage>672</fpage>
          <lpage>7</lpage>
		<pub-id pub-id-type="doi">10.1038/s41586-021-03854-z</pub-id>
		 <annotation><p>Ravuri, S.; Lenc, K.; Willson, M.; et al. Skilful precipitation nowcasting using deep generative models of radar. <italic>Nature</italic> <bold>2021</bold>, <italic>597</italic>, 672–7.</p></annotation></element-citation>
     </ref>

      <ref id="b30">
      <label>30</label>
        <element-citation publication-type="journal">
        <person-group person-group-type="author">
	   <name>
	    <surname>Zhang</surname>
	    <given-names>Y.</given-names>
	   </name>
	   <name>
	    <surname>Long</surname>
	    <given-names>M.</given-names>
	   </name>
	   <name>
	    <surname>Chen</surname>
	    <given-names>K.</given-names>
	   </name>
	  <etal/>
          </person-group>
          <article-title>Skilful nowcasting of extreme precipitation with NowcastNet</article-title>
          <source>Nature</source>
          <year>2023</year>
          <volume>619</volume>
          <fpage>526</fpage>
          <lpage>32</lpage>
		<pub-id pub-id-type="doi">10.1038/s41586-023-06184-4</pub-id>
		 <annotation><p>Zhang, Y.; Long, M.; Chen, K.; et al. Skilful nowcasting of extreme precipitation with NowcastNet. <italic>Nature</italic> <bold>2023</bold>, <italic>619</italic>, 526–32.</p></annotation></element-citation>
     </ref>

      <ref id="b31">
      <label>31</label>
        <note><p>Yu, D.; Li, X.; Ye, Y.; Zhang, B.; Luo, C.; Dai, K. DiffCast: a unified framework via residual diffusion for precipitation nowcasting. In <italic>2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</italic>, Seattle, USA, Jun 16-22, 2024. IEEE; 2024. pp. 27758-67.</p><p content-type="code">10.1109/CVPR52733.2024.02622</p></note>
     </ref>

      <ref id="b32">
      <label>32</label>
        <note><p>Xu, W.; Chen, K.; Han, T.; Chen, H.; Ouyang, W.; Bai, L. ExtremeCast: boosting extreme value prediction for global weather forecast. <italic>arXiv</italic> <bold>2024</bold>, arXiv: 2402.01295. Available online: <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.48550/arXiv.2402.01295">https://doi.org/10.48550/arXiv.2402.01295</ext-link>. (accessed 20 Jul 2026)</p><p content-type="code">10.48550/arXiv.2402.01295</p></note>
     </ref>

      <ref id="b33">
      <label>33</label>
        <note><p>Woo, S.; Park, J.; Lee, J. Y.; Kweon, I. S. CBAM: convolutional block attention module. In <italic>Proceedings of the European conference on computer vision (ECCV)</italic>, 2018. Springer, Cham; 2018. pp. 3-19.</p><p content-type="code">10.1007/978-3-030-01234-2_1</p></note>
     </ref>

      <ref id="b34">
      <label>34</label>
        <note><p>Ronneberger, O.; Fischer, P.; Brox, T. U-Net: convolutional networks for biomedical image segmentation. In <italic>Medical image computing and computer-assisted intervention - MICCAI 2015</italic>, Munich, Germany, Oct 05-09, 2015. Springer; 2015. pp. 234–41.</p><p content-type="code">10.1007/978-3-319-24574-4_28</p></note>
     </ref>

      <ref id="b35">
      <label>35</label>
        <element-citation publication-type="journal">
        <person-group person-group-type="author">
	   <name>
	    <surname>Ayzel</surname>
	    <given-names>G.</given-names>
	   </name>
	   <name>
	    <surname>Scheffer</surname>
	    <given-names>T.</given-names>
	   </name>
	   <name>
	    <surname>Heistermann</surname>
	    <given-names>M.</given-names>
	   </name>
          </person-group>
          <article-title>RainNet v1.0: a convolutional neural network for radar-based precipitation nowcasting</article-title>
          <source>Geosci. Model Dev.</source>
          <year>2020</year>
          <volume>13</volume>
          <fpage>2631</fpage>
          <lpage>44</lpage>
		<pub-id pub-id-type="doi">10.5194/gmd-13-2631-2020</pub-id>
		 <annotation><p>Ayzel, G.; Scheffer, T.; Heistermann, M. RainNet v1.0: a convolutional neural network for radar-based precipitation nowcasting. <italic>Geosci. Model Dev.</italic> <bold>2020</bold>, <italic>13</italic>, 2631–44.</p></annotation></element-citation>
     </ref>

      <ref id="b36">
      <label>36</label>
        <note><p>Agrawal, S.; Barrington, L.; Bromberg, C.; Burge, J.; Gazen, C.; Hickey, J. Machine learning for precipitation nowcasting from radar images. <italic>arXiv</italic> <bold>2019</bold>, arXiv: 1912.12132. Available online: <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.48550/arXiv.1912.12132">https://doi.org/10.48550/arXiv.1912.12132</ext-link>. (accessed 20 Jul 2026)</p><p content-type="code">10.48550/arXiv.1912.12132</p></note>
     </ref>

      <ref id="b37">
      <label>37</label>
        <note><p>Wang, Q.; Wu, B.; Zhu, P.; Li, P.; Zuo, W.; Hu, Q. ECA-Net: efficient channel attention for deep convolutional neural networks. In <italic>2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</italic>, Seattle, USA, Jun 13-19, 2020. IEEE; 2020. pp. 11531-9.</p><p content-type="code">10.1109/CVPR42600.2020.01155</p></note>
     </ref>

      <ref id="b38">
      <label>38</label>
        <note><p>Zhang, C.; Yan, Q.; Meng, L.; Sylvain, T. What constitutes good contrastive learning in time-series forecasting? <italic>arXiv</italic> <bold>2023</bold>, arXiv: 2306.12086. Available online: <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.48550/arXiv.2306.12086">https://doi.org/10.48550/arXiv.2306.12086</ext-link>. (accessed 20 Jul 2026)</p><p content-type="code">10.48550/arXiv.2306.12086</p></note>
     </ref>

      <ref id="b39">
      <label>39</label>
        <note><p>Zheng, X.; Chen, X.; Schürch, M.; Mollaysa, A.; Allam, A.; Krauthammer, M. Simple contrastive representation learning for time series forecasting. <italic>arXiv</italic> <bold>2023</bold>, arXiv: 2303.18205. Available online: <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.48550/arXiv.2303.18205">https://doi.org/10.48550/arXiv.2303.18205</ext-link>. (accessed 20 Jul 2026)</p><p content-type="code">10.48550/arXiv.2303.18205</p></note>
     </ref>

      <ref id="b40">
      <label>40</label>
        <note><p>Lai, W. S.; Huang, J. B.; Wang, O.; Shechtman, E.; Yumer, E.; Yang, M. H. Learning blind video temporal consistency. In <italic>Proceedings of the European Conference on Computer Vision (ECCV)</italic>, 2018. Springer, Cham; 2018. pp. 179-95.</p><p content-type="code">10.1007/978-3-030-01267-0_11</p></note>
     </ref>

      <ref id="b41">
      <label>41</label>
        <note><p>Dwibedi, D.; Aytar, Y.; Tompson, J.; Sermanet, P.; Zisserman, A. Temporal cycle-consistency learning. In <italic>2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</italic>, Long Beach, USA, Jun 15-20, 2019. IEEE; 2019. pp. 1801-10.</p><p content-type="code">10.1109/CVPR.2019.00190</p></note>
     </ref>

      <ref id="b42">
      <label>42</label>
        <note><p>Guan, D.; Huang, J.; Xiao, A.; Lu, S. Domain adaptive video segmentation via temporal consistency regularization. In <italic>2021 IEEE/CVF International Conference on Computer Vision (ICCV)</italic>, Montreal, Canada, Oct 10-17, 2021. IEEE; 2021. pp. 8033-44.</p><p content-type="code">10.1109/ICCV48922.2021.00795</p></note>
     </ref>

      <ref id="b43">
      <label>43</label>
        <element-citation publication-type="journal">
        <person-group person-group-type="author">
	   <name>
	    <surname>Chen</surname>
	    <given-names>L.</given-names>
	   </name>
	   <name>
	    <surname>Cao</surname>
	    <given-names>Y.</given-names>
	   </name>
	   <name>
	    <surname>Ma</surname>
	    <given-names>L.</given-names>
	   </name>
	   <name>
	    <surname>Zhang</surname>
	    <given-names>J.</given-names>
	   </name>
          </person-group>
          <article-title>A deep learning-based methodology for precipitation nowcasting with radar</article-title>
          <source>Earth Space Sci.</source>
          <year>2020</year>
          <volume>7</volume>
          <fpage>e2019EA000812</fpage>
		<pub-id pub-id-type="doi">10.1029/2019EA000812</pub-id>
		 <annotation><p>Chen, L.; Cao, Y.; Ma, L.; Zhang, J. A deep learning-based methodology for precipitation nowcasting with radar. <italic>Earth Space Sci.</italic> <bold>2020</bold>, <italic>7</italic>, e2019EA000812.</p></annotation></element-citation>
     </ref>

      <ref id="b44">
      <label>44</label>
        <note><p>Veillette, M. S.; Samsi, S.; Mattioli, C. J. SEVIR: a storm event imagery dataset for deep learning applications in radar and satellite meteorology. In <italic>Proceedings of the 34th International Conference on Neural Information Processing Systems</italic>, 2020. Curran Associates Inc.; 2020. pp. 22009–19.</p><p content-type="code">10.5555/3495724.3497570</p></note>
     </ref>

    </ref-list>

</back>
</article>
