﻿<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.0 20120330//EN" "http://jats.nlm.nih.gov/publishing/1.0/JATS-journalpublishing1.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-id journal-id-type="nlm-ta">Intell. Robot.</journal-id>
      <journal-id journal-id-type="publisher-id">IR</journal-id>
      <journal-title-group>
        <journal-title>Intelligence &amp; Robotics</journal-title>
      </journal-title-group>
      <issn pub-type="epub">2770-3541</issn>
      <publisher>
        <publisher-name>OAE Publishing Inc.</publisher-name>
      </publisher>
    </journal-meta>
    <article-meta>
	<article-id>IR-2026-052601</article-id>
      <article-id pub-id-type="doi">10.20517/ir.2026.32</article-id>
      <article-categories>
        <subj-group>
          <subject>Research Article</subject>
        </subj-group>
      </article-categories>
      <title-group>
        <article-title>AD-FIT: industrial anomaly detection via fusion of IoT sensing and network traffic data</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <name>
            <surname>Wang</surname>
            <given-names>Fei</given-names>
          </name>
          <xref ref-type="aff" rid="I1">
            <sup>1</sup>
          </xref>
          <xref ref-type="aff" rid="I2">
            <sup>2</sup>
          </xref>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Wang</surname>
            <given-names>Lei</given-names>
          </name>
          <xref ref-type="aff" rid="I3">
            <sup>3</sup>
          </xref>
        </contrib>
        <contrib contrib-type="author" corresp="yes">
          <name>
            <surname>Lv</surname>
            <given-names>Mingqi</given-names>
          </name>
          <xref ref-type="aff" rid="I3">
            <sup>3</sup>
          </xref>
          <xref ref-type="corresp" rid="cor1" />
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Chen</surname>
            <given-names>Honglong</given-names>
          </name>
          <xref ref-type="aff" rid="I1">
            <sup>1</sup>
          </xref>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Zhang</surname>
            <given-names>Zhaobo</given-names>
          </name>
          <xref ref-type="aff" rid="I2">
            <sup>2</sup>
          </xref>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Tao</surname>
            <given-names>Yongfeng</given-names>
          </name>
          <xref ref-type="aff" rid="I4">
            <sup>4</sup>
          </xref>
        </contrib>
      </contrib-group>
      <aff id="I1">
        <sup>1</sup>College of Control Science and Engineering, China University of Petroleum (East China), Qingdao 266580, Shandong, China.</aff>
		<aff id="I2">
        <sup>2</sup>Peking University Institute of Advanced Information Technology, Hangzhou 310000, Zhejiang, China.</aff>
      <aff id="I3">
        <sup>3</sup>College of Geographic Information, Zhejiang University of Technology, Hangzhou 310023, Zhejiang, China.</aff>
      <aff id="I4">
        <sup>4</sup>Shaoxing Smart City Group Co., LTD., Shaoxing 312000, Zhejiang, China.</aff>
      <author-notes>
        <corresp id="cor1">Correspondence to: Prof. Mingqi Lv, College of Geographic Information, Zhejiang University of Technology, Hangzhou 310023, Zhejiang, China. E-mail: <email>mingqilv@zjut.edu.cn</email></corresp>
        <fn fn-type="other">
          <p>
            <bold>Received:</bold> 26 May 2026 | <bold>First Decision:</bold> 6 Jul 2026 |  <bold>Revised:</bold> 14 Sep 2026 | <bold>Accepted:</bold> 15 Sep 2026 |  <bold>Published:</bold> 30 Sep 2026</p>
        </fn>
        <fn fn-type="other">
          <p>
            <bold>Academic Editor:</bold> Lei Lei |  <bold>Copy Editor:</bold> Pei-Yun Wang |  <bold>Production Editor:</bold> Pei-Yun Wang</p>
        </fn>
      </author-notes>
      <pub-date pub-type="ppub">
        <year>2026</year>
      </pub-date>
      <pub-date pub-type="epub">
        <day>30</day>
        <month>9</month>
        <year>2026</year>
      </pub-date>
      <volume>6</volume>
	  <issue>3</issue>
      <fpage>729</fpage>
	  <lpage>49</lpage>
      <permissions>
        <copyright-statement>© The Author(s) 2026.</copyright-statement>
        <license xlink:href="https://creativecommons.org/licenses/by/4.0/">
          <license-p>© The Author(s) 2026. <bold>Open Access</bold> This article is licensed under a Creative Commons Attribution 4.0 International License (<uri xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</uri>), which permits unrestricted use, sharing, adaptation, distribution and reproduction in any medium or format, for any purpose, even commercially, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons license, and indicate if changes were made.</license-p>
        </license>
      </permissions>
      <abstract>
        <p>Anomaly detection is an important research topic in the Industrial Internet of Things (IIoT). In recent years, deep learning has been exploited to analyze complex IIoT data and build anomaly detection models. Due to the lack of abnormal samples and the difficulty of labeling industrial data, unsupervised deep learning has become the mainstream technique for IIoT anomaly detection, with autoencoders being the most representative approaches. However, the existing autoencoder-based IIoT anomaly detection models predominantly focus on a single data modality, since existing data fusion frameworks are mostly designed for supervised learning tasks, while it is infeasible to simultaneously reconstruct heterogeneous data modalities in a unified autoencoder. To address this limitation, this paper focuses on IoT sensing data and network traffic data and proposes AD-FIT, a novel autoencoder framework for IIoT anomaly detection via the fusion of IoT sensing and network traffic data in an unsupervised manner. Specifically, it creates multiple local autoencoders with different architectures to fit the two data modalities, and then fuses their reconstruction errors through a global autoencoder. We conducted extensive experiments based on a public IIoT dataset. Experimental results show that AD-FIT achieves the best overall anomaly detection performance among the evaluated baseline methods, with F1-score improvements ranging from 7.1% to 38.6%.</p>
      </abstract>
      <kwd-group>
        <kwd>Anomaly detection</kwd>
        <kwd>Industrial Internet of Things</kwd>
        <kwd>autoencoder</kwd>
        <kwd>heterogeneous data fusion</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec1">
      <title>1. INTRODUCTION</title>
      <p>In recent years, the advent of the Industrial Internet of Things (IIoT) has transformed traditional industrial landscapes by integrating advanced connectivity and smart technologies into manufacturing processes. IIoT systems encompass a myriad of interconnected devices and sensors that collectively contribute to the efficient monitoring and control of industrial operations. This paradigm shift has not only ushered in a new era of productivity but has also introduced novel challenges, particularly in the realm of anomaly detection, which is the identification of deviations from expected behavior. An anomaly can indicate a malfunction, an attack, or other types of unexpected events<sup>[<xref ref-type="bibr" rid="B1">1</xref>]</sup>. Early detection of anomalies can prevent equipment failures, improve process efficiency, and enhance overall performance.</p>
      <p>To facilitate information exchange among humans, devices, and systems, an IIoT system connects a large number of monitoring devices and sensors through communication networks and collects data that can reflect the process and status of production. As a result, most state-of-the-art IIoT anomaly detection methods apply data-driven techniques, including traditional machine learning and deep learning. Since the operating processes and collected IIoT data are highly complex, making it difficult to understand the abnormal states, labeling anomalous states in the collected IIoT data requires substantial domain knowledge and is costly. In practice, obtaining labeled abnormal samples from real IIoT systems is often difficult or impossible. Therefore, most previous studies focus on unsupervised learning techniques<sup>[<xref ref-type="bibr" rid="B2">2</xref>,<xref ref-type="bibr" rid="B3">3</xref>]</sup>. Studies using traditional machine learning typically involve two steps. They first extract a large number of features from industrial data, and then discover anomalies by using outlier detection algorithms [e.g., clustering<sup>[<xref ref-type="bibr" rid="B4">4</xref>,<xref ref-type="bibr" rid="B5">5</xref>]</sup>, One-Class support vector machine (SVM)<sup>[<xref ref-type="bibr" rid="B6">6</xref>]</sup>, iForest<sup>[<xref ref-type="bibr" rid="B7">7</xref>]</sup>]. However, feature engineering relies on domain knowledge, and the high-dimensional and dynamic nature of IIoT data makes it difficult to design effective features.</p>
      <p>In recent years, thanks to the ability to automatically learn features from raw data, deep learning has been increasingly applied to IIoT anomaly detection tasks. Many studies focus on multivariate time-series IoT sensing data and use unsupervised reconstruction-based deep learning models<sup>[<xref ref-type="bibr" rid="B8">8</xref>,<xref ref-type="bibr" rid="B9">9</xref>]</sup>. These models learn representations of normal patterns and detect anomalies using reconstruction errors. Due to the temporal nature of IoT sensing data, recurrent neural network (RNN)-based models have also been widely investigated<sup>[<xref ref-type="bibr" rid="B8">8</xref>,<xref ref-type="bibr" rid="B10">10</xref>]</sup>. Several studies have also attempted to detect anomalies from the IIoT network traffic data. Specifically, they extract statistical and temporal features from raw network data packets<sup>[<xref ref-type="bibr" rid="B11">11</xref>]</sup> and build anomaly detection models based on autoencoders<sup>[<xref ref-type="bibr" rid="B12">12</xref>,<xref ref-type="bibr" rid="B13">13</xref>]</sup>. Various backbone networks [e.g., multilayer perceptron (MLP)<sup>[<xref ref-type="bibr" rid="B13">13</xref>]</sup>, long short-term memory (LSTM)<sup>[<xref ref-type="bibr" rid="B14">14</xref>]</sup>] that can adapt to these statistical and temporal features are exploited to implement the autoencoders. Then, these features are reconstructed based on the decoders, and the reconstruction errors are used for network traffic anomaly detection. Recent work has further applied sequence-to-sequence LSTM autoencoders with attention mechanisms to real-time network-based anomaly detection in Modbus/TCP industrial control systems<sup>[<xref ref-type="bibr" rid="B15">15</xref>]</sup>.</p>
      <p>Recent transformer and graph-based methods, including Anomaly Transformer, TranAD, GRELEN, MTGFlow, MEMTO, D3R, SARAD, and graph mixture-of-experts models, further improve multivariate time-series anomaly detection by modeling temporal dependencies, inter-variable relations, and nonstationary behavior<sup>[<xref ref-type="bibr" rid="B16">16</xref>-<xref ref-type="bibr" rid="B23">23</xref>]</sup>. Multimodal sensing representation learning, as exemplified by FOCAL, also shows the value of contrastive alignment across heterogeneous sensing signals<sup>[<xref ref-type="bibr" rid="B24">24</xref>]</sup>. Recent IIoT intrusion-detection research also benefits from realistic benchmarks and attention-based models, including Edge-IIoTset, CICIoT2023, and SACNN-IDS<sup>[<xref ref-type="bibr" rid="B25">25</xref>-<xref ref-type="bibr" rid="B27">27</xref>]</sup>.</p>
      <p>Despite recent advances, existing IIoT anomaly-detection studies mainly focus on a single data modality and stable operating conditions. Although dynamic weighted domain adaptation has been explored for bearing fault diagnosis under time-varying speeds<sup>[<xref ref-type="bibr" rid="B28">28</xref>]</sup>, robust unsupervised fusion of heterogeneous IoT sensing and network traffic data remains insufficiently explored, limiting the ability to detect anomalies spanning both domains. Modern IIoT systems are highly interconnected, and anomalies arising in one domain may affect the entire infrastructure. Therefore, a unified framework that can effectively integrate information from both data sources is needed. However, developing such a framework remains challenging for the following reasons.</p>
      <p>Complex correlations: IoT sensing data are usually collected from many sensors, and network traffic data are usually represented by many features, resulting in high-dimensional data. Different dimensions may have mutual correlations. For example, the data of one sensor could affect the data of another sensor, and a specific state of the system can be jointly influenced by multiple features. Moreover, the IoT sensing and network traffic data have substantially different correlation patterns, which are difficult to learn with a single unified model.</p>
      <p>Heterogeneous data modalities: The IoT sensing data and network traffic data represent fundamentally distinct forms of information. The inherent differences in structure, format, and semantics make it challenging to seamlessly map these diverse data types into a unified feature space. Although previous work has extensively studied heterogeneous data fusion methods, most are designed for supervised learning tasks, and it is infeasible to simultaneously reconstruct heterogeneous data modalities in a unified autoencoder.</p>
      <p>To this end, we propose AD-FIT, a novel framework for IIoT anomaly detection that fuses IoT sensing and network traffic data. AD-FIT adopts the following strategies to address the above-mentioned challenges. First, AD-FIT uses graphs to explicitly model the correlations between different feature dimensions and applies graph neural networks (GNNs) to learn the correlation strengths and incorporate them into the anomaly detection model. Second, AD-FIT fuses the IoT sensing data and network traffic data based on an ensemble framework of autoencoders across diverse data modalities. In summary, this paper makes the following contributions.</p>
      <p>First, we propose an ensemble framework to fuse multiple autoencoders across two data modalities, i.e., IoT sensing data and network traffic data. This framework operates unsupervisedly and does not require anomalous training samples.</p>
      <p>Second, we design the autoencoders using GNNs to explicitly learn complex correlations between sensors and integrate them into the models.</p>
      <p>Third, we conducted extensive experiments on a public multimodal IIoT dataset. The results show that AD-FIT improves the F1-score by approximately 7.1% to 38.6% compared with the baseline methods.</p>
    </sec>
    <sec id="sec2">
      <title>2. RELATED WORK</title>
      <sec id="sec2-1">
        <title>2.1. Industrial anomaly detection</title>
        <p>Industrial data are typically multivariate time series characterized by large volume, high dimensionality, and strong temporal dynamics, and thus the traditional univariate anomaly detection methods<sup>[<xref ref-type="bibr" rid="B2">2</xref>-<xref ref-type="bibr" rid="B6">6</xref>]</sup> may not perform effectively.</p>
        <p>To handle complex multivariate time-series data, most existing studies use learning-based techniques (including machine learning and deep learning), which can be roughly divided into two categories, i.e., supervised learning techniques and unsupervised learning techniques. In terms of supervised learning techniques, Griffin <italic>et al</italic>. proposed an anomaly detection method based on neural networks and decision trees to detect anomalies across multiple industrial processes<sup>[<xref ref-type="bibr" rid="B29">29</xref>]</sup>. Nanduri <italic>et al</italic>. proposed the use of recurrent neural networks to detect abnormal events that may reduce flight safety factors<sup>[<xref ref-type="bibr" rid="B10">10</xref>]</sup>. Janssens <italic>et al</italic>. proposed a convolutional neural network (CNN)-based feature-learning system for detecting fault states in rotating machinery<sup>[<xref ref-type="bibr" rid="B30">30</xref>]</sup>.</p>
        <p>Although supervised learning techniques can better distinguish anomalies by explicitly learning the latent patterns from anomaly samples, their practical deployment is challenging because labeled anomalous training data are severely limited. Because normal industrial data can be easily obtained at scale, unsupervised learning techniques are widely used for industrial multivariate time-series anomaly detection. Earlier studies applied traditional unsupervised machine-learning methods such as clustering, One-Class SVM, and iForest. For example, Amruthnath and Gupta applied multiple clustering algorithms for anomaly detection (e.g., K-Means, fuzzy C-Means)<sup>[<xref ref-type="bibr" rid="B31">31</xref>]</sup>. These algorithms identify behaviors that deviate from normal patterns. Diez-Olivan <italic>et al</italic>. proposed an anomaly detection method based on One-Class SVM to detect anomalies in sensor data by obtaining anomaly scores based on the distance between samples and the separating hyperplane<sup>[<xref ref-type="bibr" rid="B32">32</xref>]</sup>. Joshi <italic>et al</italic>. conducted anomaly detection based on a hidden Markov model (HMM), which constructs a model by extracting features and estimating anomaly probabilities from the generated state sequence<sup>[<xref ref-type="bibr" rid="B33">33</xref>]</sup>. Nguyen and Vien proposed AE-1SVM, which combines autoencoder-based representation learning with a One-Class SVM for unsupervised anomaly detection<sup>[<xref ref-type="bibr" rid="B34">34</xref>]</sup>.</p>
        <p>However, traditional machine-learning techniques rely heavily on feature engineering, which is challenging for high-dimensional and dynamic industrial data. Consequently, researchers have extensively investigated generative encoder-decoder networks for unsupervised anomaly detection in time-series data. For example, Xu <italic>et al</italic>. proposed Donut, a variational autoencoder (VAE)-based unsupervised anomaly detection method that identifies abnormal time points using the reconstruction probability of time-series observations<sup>[<xref ref-type="bibr" rid="B35">35</xref>]</sup>. Lu <italic>et al</italic>. detected anomalies in rotating mechanical components by using a stacked autoencoder<sup>[<xref ref-type="bibr" rid="B36">36</xref>]</sup>. Zhang <italic>et al</italic>. proposed MSCRED, which uses a ConvLSTM autoencoder to learn the multiple levels of system operation patterns characterized by multi-scale signature matrices in different time steps<sup>[<xref ref-type="bibr" rid="B37">37</xref>]</sup>. Yin <italic>et al</italic>. integrated a CNN with a recurrent autoencoder for detecting anomalies in IoT systems<sup>[<xref ref-type="bibr" rid="B38">38</xref>]</sup>. Specifically, they used a two-stage sliding window strategy to design the encoder for better feature extraction. Muneer <italic>et al</italic>. proposed a hybrid model based on a deep autoencoder neural network (DANN) with five layers for detecting anomalies in a real-world gas turbine dataset<sup>[<xref ref-type="bibr" rid="B39">39</xref>]</sup>. Although the encoder-decoder deep neural networks have achieved promising results for industrial anomaly detection, their performance may be further improved by explicitly learning the relationships between industrial data and network traffic data and appropriately balancing the two data sources.</p>
      </sec>
      <sec id="sec2-2">
        <title>2.2. Cross-domain data fusion</title>
        <p>IoT sensing data and network traffic data represent distinct modalities with different representations, distributions, scales, and densities. For example, industrial sensing data typically consist of measurements collected by multiple sensor devices on a production line, which may include data such as device temperature and pressure. Network traffic data, by contrast, come from packet analysis and include attributes such as source and destination addresses. Consequently, fusing data across these modalities is challenging. Data fusion methods can be categorized as stage-based fusion methods, feature-level fusion methods, and semantic-level fusion methods<sup>[<xref ref-type="bibr" rid="B40">40</xref>]</sup>.</p>
        <p>The stage-based methods use different datasets at different stages of a data fusion task. Thus, these datasets are loosely connected, and cross-modal consistency is not explicitly required. For example, Xiao <italic>et al.</italic> utilized spatial trajectories to detect stay points, converting them into feature vectors based on surrounding points of interest (POIs)<sup>[<xref ref-type="bibr" rid="B41">41</xref>]</sup>. These feature vectors are hierarchically clustered to form a tree structure, representing users’ location histories and facilitating similarity measurement between users based on their hierarchical graphs.</p>
        <p>For feature-level fusion, a common approach is to learn a unified feature representation from disparate datasets using deep neural networks. For example, Ngiam <italic>et al.</italic> introduced a deep autoencoder architecture to capture intermediate feature representations across modalities, exploring three learning settings: cross-modality learning, shared representation learning, and multimodal fusion, thereby improving single-modality representations and capturing correlations across multiple modalities<sup>[<xref ref-type="bibr" rid="B42">42</xref>]</sup>. Nagrani <italic>et al.</italic> introduced a transformer-based multimodal architecture that uses attention bottlenecks to constrain cross-modal interactions through a small set of bottleneck latent units, thereby improving fusion performance while reducing computational cost<sup>[<xref ref-type="bibr" rid="B43">43</xref>]</sup>. More recently, IoT-GRAF represents heterogeneous sensing and network observations through graph learning for multimodal IoT anomaly and intrusion detection<sup>[<xref ref-type="bibr" rid="B44">44</xref>]</sup>. Cross-domain representation learning has also been used to jointly capture network behaviors and physical-process states in industrial control systems<sup>[<xref ref-type="bibr" rid="B45">45</xref>]</sup>.</p>
        <p>Feature-based data fusion methods treat features as numerical or categorical values without considering their semantic meaning, whereas semantic-level methods capture the semantic information contained in each dataset, recognize relations between features, and provide interpretable and meaningful fusion by incorporating the significance of each dataset and the interplay among the datasets. Although researchers have investigated cross-domain fusion methods, their scope is often limited to specific data-query or application settings. For example, Sun <italic>et al.</italic> studied k-nearest-neighbor temporal aggregate queries that rank locations by combining spatial distance with temporally aggregated attributes<sup>[<xref ref-type="bibr" rid="B46">46</xref>]</sup>. Decision-level fusion of network and physical anomaly detectors has been shown to improve cyber–physical threat detection compared with relying on either data source alone<sup>[<xref ref-type="bibr" rid="B47">47</xref>]</sup>. More recently, an unsupervised adversarial fusion framework combined sensor and network representations through dual cross-attention to detect subtle cyber–physical anomalies<sup>[<xref ref-type="bibr" rid="B48">48</xref>]</sup>. Despite this progress, robust unsupervised fusion of heterogeneous sensing and network-traffic modalities for IIoT anomaly detection remains insufficiently explored.</p>
      </sec>
    </sec>
    <sec id="sec3">
      <title>3. METHODOLOGY</title>
      <sec id="sec3-1">
        <title>3.1. Preliminary</title>
        <p>In this section, we first define the key concepts and formulate the anomaly detection problem. We then introduce the overall AD-FIT architecture, followed by the modality-specific backbone autoencoders, the unsupervised data-fusion framework, and the final anomaly-detection procedure.</p>
        <p>
          <bold>Definition 1</bold> (<italic>IoT Sensing Sample</italic>): The IoT sensing sample of time slot <italic>t</italic> is denoted as <italic>X<sub>t</sub></italic> ∈ <italic>R<sup>N</sup></italic><sup>×</sup><italic><sup>W</sup></italic>, where <italic>N</italic> is the number of sensors, and <italic>W</italic> is the sampling-window length. Here, <italic>X<sub>t</sub></italic> is a subsequence of the multivariate time series from the <italic>N</italic> sensors during the time period of (<italic>t</italic>-<italic>W</italic>, <italic>t</italic>]. Note that a physical sensor may generate multiple readings at each time slot, e.g., an accelerometer generates three readings at each time slot. For notational convenience, we treat each univariate time series in <italic>X<sub>t</sub></italic> as sampled from an individual virtual sensor.</p>
        <p>
          <bold>Definition 2</bold> (<italic>Network Traffic Sample</italic>): The network traffic sample of time slot <italic>t</italic> (denoted as <italic>Y<sub>t</sub></italic>) is the collection of network data packets captured during the time period of (<italic>t</italic>-<italic>W</italic>, <italic>t</italic>].</p>
        <p>
          <bold>Definition 3</bold> (<italic>IIoT Data Sample</italic>): The IIoT data sample of time slot <italic>t</italic>, denoted as <italic>S<sub>t</sub></italic> = (<italic>X<sub>t</sub></italic>, <italic>Y<sub>t</sub></italic>), is composed of an IoT sensing sample and a network traffic sample during the same time period (<italic>t</italic>-<italic>W</italic>, <italic>t</italic>].</p>
        <p>
          <bold>Problem Definition</bold>: Industrial anomaly detection in this paper is formulated as learning from IIoT data samples to obtain a function <italic>f</italic>, which takes <italic>S<sub>t</sub></italic> as input and produces an output value <italic>y<sub>t</sub></italic> ∈ {0, 1}, where <italic>y<sub>t</sub></italic> denotes whether an anomaly occurs at time slot <italic>t</italic>.</p>
        <p>
          <bold>Architecture of AD-FIT:</bold> <xref ref-type="fig" rid="fig1">Figure 1</xref> shows the architecture of AD-FIT, which comprises three modules, i.e., IoT-AE module, Network-AE module, and Fusion module. The IoT-AE module applies bootstrap sampling to the IoT sensing samples to create multiple training subsets and then trains multiple autoencoders on the IoT sensing data. Each IoT sensing autoencoder consists of three layers: the graph attention network (GAT) layer, the encoder layer, and the decoder layer. The Network-AE module also applies bootstrap sampling to the set of network traffic samples to create multiple training subsets and then trains multiple autoencoders on the network traffic data. Each autoencoder for network traffic data consists of three layers, i.e., the feature extraction layer, the encoder layer, and the decoder layer. The Fusion module can be viewed as a meta-learner on top of the IoT-AE and Network-AE modules. It takes the reconstruction root mean square errors (RMSEs) produced by the local autoencoders as input and trains a global autoencoder to reconstruct these RMSEs.</p>
        <fig id="fig1" position="float">
          <label>Figure 1</label>
          <caption>
            <p>The architecture of AD-FIT. IIoT: Industrial Internet of Things; IoT: Internet of Things; AE: autoencoder; RMSE: root mean square error.</p>
          </caption>
          <graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ir6032.fig.1.jpg" />
        </fig>
        <p>In the real-time anomaly detection phase, given an IIoT data sample <italic>S<sub>t</sub></italic> = (<italic>X<sub>t</sub></italic>, <italic>Y<sub>t</sub></italic>), <italic>X<sub>t</sub></italic> and <italic>Y<sub>t</sub></italic> are fed into the corresponding autoencoders of the IoT-AE and Network-AE modules, respectively, which produce a series of reconstruction RMSEs. The Fusion module then combines these RMSEs to produce a global RMSE for anomaly detection.</p>
        <p>AD-FIT assumes that the input IoT sensing and network-traffic data have undergone basic quality control and are sufficiently reliable. In practical IIoT deployments, sensor drift, missing measurements, and corrupted readings may distort reconstruction errors and propagate to the fusion decision. Although the present framework does not explicitly model reliability and trust, reliability-aware data validation or weighting could be incorporated to improve deployment robustness<sup>[<xref ref-type="bibr" rid="B49">49</xref>]</sup>.</p>
      </sec>
      <sec id="sec3-2">
        <title>3.2. The backbone autoencoders</title>
        <p>AD-FIT is essentially an ensemble framework of autoencoders, built separately for IoT sensing data and network traffic data before fusion. Because the two data modalities have different characteristics, we used different backbone autoencoders for each.</p>
        <sec id="sec3-2-1">
          <title>3.2.1. IoT sensing data-based autoencoder</title>
          <p>The backbone autoencoder for IoT sensing data is composed of three layers, i.e., the GAT layer, the encoder layer, and the decoder layer. The GAT layer learns potential correlations between different sensors. The encoder layer learns spatial and temporal patterns in IoT sensing data. The decoder layer reconstructs the input data.</p>
          <p>(1) The GAT Layer<break/>The IoT sensing data are represented as high-dimensional multivariate time series. First, these dimensions, which represent different sensors, may exhibit potential correlations and cross-effects. For example, detecting some anomalies requires considering data from multiple sensors simultaneously. Second, different sensors may contribute differently to anomaly detection. For example, pressure data in a hydraulic system are usually more important for anomaly detection than the data from other sensors. To address these issues, the GAT layer uses a graph to model the correlations between different dimensions of the IoT sensing data (Step 1), and then the correlation strengths are learned using a GAT subnetwork (Step 2).</p>
          <p>
            <bold>Step 1</bold> (Sensor graph construction): A sensor graph is represented as <italic>G</italic> = (<italic>V</italic>, <italic>E</italic>, <italic>F</italic>), where each node <italic>v<sub>i</sub></italic> ∈ <italic>V</italic> represents a dimension of the IoT sensing data (corresponding to a sensor), each edge <italic>e<sub>ij</sub></italic> ∈ <italic>E</italic> represents a correlation between sensor <italic>v<sub>i</sub></italic> and <italic>v<sub>j</sub></italic>, and <italic>F</italic> represents the set of node features, where each element <italic>f<sub>i</sub></italic> denotes the original feature of node <italic>v<sub>i</sub></italic> (i.e., the univariate time-series data collected from sensor <italic>v<sub>i</sub></italic>). Domain knowledge can determine whether two sensors are correlated when constructing the sensor graph. However, because domain knowledge is often hard to obtain, we design <italic>G</italic> as a fully connected graph when domain knowledge is unavailable.</p>
          <p>
            <bold>Step 2</bold> (Correlation strength learning): Following the graph attention mechanism proposed by Veličković <italic>et al.</italic><sup>[<xref ref-type="bibr" rid="B50">50</xref>]</sup>, we apply GAT to learn the correlation strength between each pair of nodes in <italic>G</italic>. As illustrated in <xref ref-type="fig" rid="fig2">Figure 2</xref>, for each edge <italic>e<sub>ij</sub></italic> in <italic>G</italic>, its correlation strength <italic>w<sub>ij</sub></italic> is calculated according to Equation (1), where <italic>a<sub>ij</sub></italic> is the unnormalized attention score, <italic>k</italic> is the index of a node in the first-order neighborhood of <italic>v<sub>i</sub></italic>, and <italic>L</italic> is the number of nodes in this neighborhood, including <italic>v<sub>i</sub></italic> itself.</p>
			<p><disp-formula> <label>(1)</label> <tex-math id="E1"> $$  w_{i j}=\frac{\exp \left(a_{i j}\right)}{\sum_{k=1}^{L} \exp \left(a_{i k}\right)} $$ </tex-math></disp-formula></p>
          <fig id="fig2" position="float">
            <label>Figure 2</label>
            <caption>
              <p>Correlation-strength learning based on GAT. GAT: Graph attention network.</p>
            </caption>
            <graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ir6032.fig.2.jpg" />
          </fig>
          <p>The unnormalized attention score <italic>a<sub>ij</sub></italic> is further calculated according to Equation (2), where <italic>q</italic> is a learnable parameter vector and <italic>σ</italic>(…) is a nonlinear activation function.</p>
		  <p><disp-formula> <label>(2)</label> <tex-math id="E1"> $$  a_{i j}=\sigma\left(q^{\mathrm{T}} \cdot\left(f_{i} \oplus f_{j}\right)\right) $$ </tex-math></disp-formula></p>
          <p>The normalized correlation strengths are subsequently used to aggregate the neighboring features and update the embedding <italic>g<sub>i</sub></italic> of node <italic>v<sub>i</sub></italic> according to Equation (3), where <italic>w<sub>ik</sub></italic> is the normalized correlation strength from node <italic>v<sub>k</sub></italic> to node <italic>v<sub>i</sub></italic> and <italic>f<sub>k</sub></italic> is the original feature of node <italic>v<sub>k</sub></italic>.</p>
		  <p><disp-formula> <label>(3)</label> <tex-math id="E1"> $$  g_{i}=\sigma\left(\sum_{k=1}^{L} w_{i k} f_{k}\right) $$ </tex-math></disp-formula></p>
          <p>After the two steps, each IoT sensing sample <italic>X<sub>t</sub></italic> is represented as a feature matrix <italic>GX<sub>t</sub></italic> ∈ <italic>R<sup>N</sup></italic><sup>×</sup><italic><sup>W</sup></italic>, where the <italic>i</italic>-th row of <italic>GX<sub>t</sub></italic> represents the embedding vector of node <italic>v<sub>i</sub></italic> (i.e., <italic>g<sub>i</sub></italic>).</p>
          <p>(2) The Encoder Layer<break/>The autoencoder maps the original IoT sensing samples into a lower-dimensional latent feature space using an encoder and then reconstructs the latent features into the original sample space using a decoder. The autoencoder was trained by gradually reducing the errors between original samples and reconstructed samples through backpropagation.</p>
          <p>Because the latent features have lower dimensionality than the original samples, the latent features can capture the dominant patterns of the original samples. For the anomaly detection task, the autoencoder is trained based on normal samples (or mostly normal samples), so the latent features learned by the trained autoencoder can represent the dominant patterns of normal samples. If a reconstructed sample deviates greatly from the original, it means the original sample does not conform to the main patterns of normal samples, indicating an anomaly.</p>
          <p>Considering the temporal characteristics of IoT sensing samples, we employ an LSTM encoder based on the recurrent architecture proposed by Hochreiter and Schmidhuber<sup>[<xref ref-type="bibr" rid="B51">51</xref>]</sup>. As shown in <xref ref-type="fig" rid="fig3">Figure 3</xref>, given an IoT sensing sample <italic>X<sub>t</sub></italic> = [<italic>x<sub>t</sub></italic><sub>-</sub><italic><sub>W+</sub></italic><sub>1</sub>, <italic>x<sub>t</sub></italic><sub>-</sub><italic><sub>W</sub></italic><sub>+2</sub>, …, <italic>x<sub>t</sub></italic>] (<italic>x<sub>t</sub></italic> ∈ <italic>R<sup>N</sup></italic><sup>×1</sup> denotes the sensor-reading vector at time slot <italic>t</italic>), we first input <italic>X<sub>t</sub></italic> into the GAT layer to obtain the updated feature matrix <italic>GX<sub>t</sub></italic> = [<italic>gx<sub>t</sub></italic><sub>-</sub><italic><sub>W+</sub></italic><sub>1</sub>, <italic>gx<sub>t</sub></italic><sub>-</sub><italic><sub>W</sub></italic><sub>+2</sub>, …, <italic>gx<sub>t</sub></italic>]. At each time slot <italic>t</italic>, the encoder updates its hidden state <italic>h<sub>t</sub><sup>ie</sup></italic> from the previous state <italic>h<sub>t</sub></italic><sub>-1</sub><italic><sup>ie</sup></italic> and the current GAT-enhanced feature <italic>gx<sub>t</sub></italic> according to Equation (4), where the superscript “ie” denotes the IoT sensing data encoder.</p>
		  <p><disp-formula> <label>(4)</label> <tex-math id="E1"> $$  h_{t}^{i e}=\operatorname{LSTM}\left(h_{t-1}^{i e}, g x_{t}\right) $$ </tex-math></disp-formula></p>
          <fig id="fig3" position="float" width="450">
            <label>Figure 3</label>
            <caption>
              <p>The architecture of the LSTM autoencoder. LSTM: Long short-term memory; GAT: graph attention network; IoT: Internet of Things.</p>
            </caption>
            <graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ir6032.fig.3.jpg" />
          </fig>
          <p>(3) The Decoder Layer<break/>As shown in <xref ref-type="fig" rid="fig3">Figure 3</xref>, the decoder takes <italic>h<sub>t</sub><sup>ie</sup></italic> as input and also uses an LSTM subnetwork to output <italic>W</italic> hidden state vectors <italic>h<sub>t-W</sub></italic><sub>+1</sub><italic><sup>id</sup></italic>, <italic>h<sub>t-W</sub></italic><sub>+2</sub><italic><sup>id</sup></italic>, …, <italic>h<sub>t</sub><sup>id</sup></italic> based on Equation (5). After that, a fully connected subnetwork is used to reconstruct <italic>h<sub>t-W</sub></italic><sub>+1</sub><italic><sup>id</sup></italic>, <italic>h<sub>t-W</sub></italic><sub>+2</sub><italic><sup>id</sup></italic>, …, <italic>h<sub>t</sub><sup>id</sup></italic> into the same dimensions as the original samples <inline-formula><tex-math id="M1">$$ \tilde{X} $$</tex-math></inline-formula><italic><sub>t</sub></italic> ∈ <italic>R<sup>N</sup></italic><sup>×</sup><italic><sup>W</sup></italic>.</p>
		  <p><disp-formula> <label>(5)</label> <tex-math id="E1"> $$  h_{t}^{i d}=\operatorname{LSTM}\left(h_{t-1}^{i d}, h_{t}^{i e}\right) $$ </tex-math></disp-formula></p>
          <p>Following the reconstruction-error objective commonly used in LSTM-based autoencoders, we train the autoencoder for IoT sensing data by minimizing the RMSE between the original sample <italic>X<sub>t</sub></italic> and its reconstruction <inline-formula><tex-math id="M1">$$ \tilde{X} $$</tex-math></inline-formula><italic><sub>t</sub></italic>, as defined in Equation (6), where <italic>N</italic> is the number of sensors; <italic>W</italic> is the sampling-window length; and <italic>Loss<sub>I</sub></italic> denotes the resulting reconstruction loss for sample <italic>X<sub>t</sub></italic>.</p>
		  <p><disp-formula> <label>(6)</label> <tex-math id="E1"> $$  \operatorname{Loss}_{I}=\sqrt{\frac{1}{N W} \sum_{i=1}^{N} \sum_{j=1}^{W}\left(X_{t}[i, j]-\widetilde{X}_{t}[i, j]\right)^{2}} $$ </tex-math></disp-formula></p>
        </sec>
        <sec id="sec3-2-2">
          <title>3.2.2. Network traffic data-based autoencoder</title>
          <p>The backbone autoencoder for network traffic data is composed of three layers, i.e., the feature extraction layer, the encoder layer, and the decoder layer.</p>
          <p>(1) The Feature Extraction Layer<break/>The feature extraction layer includes two steps, i.e., sample generation and feature extraction.</p>
          <p>
            <bold>Sample generation</bold>: The network data packets can be captured using tools such as tcpdump and Wireshark. Then, we generate network traffic samples by selecting network data packets for a specific source-destination pair (same source IP, destination IP, source port, destination port, and protocol) within a specific time period (<italic>t</italic>-<italic>W</italic>, <italic>t</italic>].</p>
          <p>
            <bold>Feature extraction</bold>: We used CICFlowMeter (<uri xlink:href="https://github.com/ahlashkari/CICFlowMeter">https://github.com/ahlashkari/CICFlowMeter</uri>) to extract a fixed set of statistical features from each network-traffic sample. These features characterize flow duration, packet counts and byte volumes in the forward and backward directions, packet-length statistics, transmission rates, and packet inter-arrival times. Identifier fields and label-related fields are excluded. The remaining numerical features are normalized using statistics calculated from the training set and are assembled into an <italic>M</italic>-dimensional vector <italic>y<sub>t</sub></italic> ∈ <italic>R<sup>M</sup></italic><sup>×1</sup>.</p>
          <p>(2) The Encoder Layer<break/>We used an MLP as the encoder for the network traffic samples. Given a network traffic sample <italic>y<sub>t</sub></italic>, the encoder layer maps <italic>y<sub>t</sub></italic> to a latent feature space. Specifically, a two-layer MLP is applied, as shown in Equation (7), where <italic>W</italic><sup>(</sup><italic><sup>i</sup></italic><sup>)</sup> and <italic>b</italic><sup>(</sup><italic><sup>i</sup></italic><sup>)</sup> denote the learnable weight matrix and bias vector of the i-th hidden layer, respectively.</p>
		  <p><disp-formula> <label>(7)</label> <tex-math id="E1"> $$  h^{n e(1)}=\operatorname{ReLU}\left(W^{(1)} \cdot y_{t}+b^{(1)}\right)\\
h^{n e(2)}=\operatorname{ReLU}\left(W^{(2)} \cdot h^{n e(1)}+b^{(2)}\right) $$ </tex-math></disp-formula></p>
          <p>(3) The Decoder Layer<break/>The decoder layer’s task is to reconstruct the latent feature <italic>h<sup>ne</sup></italic><sup>(2)</sup> back to the input space. We also apply a two-layer MLP as the decoder, which outputs a reconstructed sample <inline-formula><tex-math id="M1">$$ \tilde{y} $$</tex-math></inline-formula><italic><sub>t</sub></italic> ∈ <italic>R<sup>M</sup></italic><sup>×1</sup>. To train the autoencoder, we define the loss function for sample <italic>y<sub>t</sub></italic> as the error between the original input sample <italic>y<sub>t</sub></italic> and the reconstructed sample <inline-formula><tex-math id="M1">$$ \tilde{y} $$</tex-math></inline-formula><italic><sub>t</sub></italic>, as defined in Equation (8).</p>
		  <p><disp-formula> <label>(8)</label> <tex-math id="E1"> $$  \operatorname{Loss}_{N}=\sqrt{\frac{1}{M}\sum_{i=1}^{M}(y_t[i]-\tilde{y}_t[i])^{2}} $$ </tex-math></disp-formula></p>
        </sec>
      </sec>
      <sec id="sec3-3">
        <title>3.3. Data fusion via ensemble of autoencoders</title>
        <sec id="sec3-3-1">
          <title>3.3.1. Unsupervised data fusion framework</title>
          <p>Because IoT sensing data and network traffic data typically have distinct dimensions and characteristics, a single, unified autoencoder cannot process both. Instead, we used autoencoders with different structures to model them. However, IoT sensing data and network traffic data collected from the same IIoT system may exhibit semantic correlations. For example, in a smart home environment, an attacker could remotely control a door lock. Because attack behaviors involve infiltrating the household’s smart home network and lurking on edge devices, clues can be reflected in both network traffic data and device log data. Therefore, a data fusion mechanism is needed to learn semantic correlations across the IoT sensing data and network traffic data.</p>
          <p>However, existing data fusion frameworks are designed for supervised learning tasks. For example, the most popular traditional data fusion frameworks (including Bagging<sup>[<xref ref-type="bibr" rid="B52">52</xref>]</sup>, Boosting<sup>[<xref ref-type="bibr" rid="B53">53</xref>]</sup>, and Stacking<sup>[<xref ref-type="bibr" rid="B54">54</xref>]</sup>) perform data fusion by exploiting classification consistency, classification residual, or classification probability distribution, whereas such information is unavailable in unsupervised settings. Deep learning models usually perform data fusion by integrating different subnetworks in a unified learning objective. Unfortunately, it is infeasible to simultaneously reconstruct heterogeneous data modalities in a unified autoencoder.</p>
          <p>To address this challenge, we propose a data-fusion framework for autoencoders with different data modalities. As shown in <xref ref-type="fig" rid="fig4">Figure 4</xref>, this framework comprises three parts: data sampling, local autoencoder training, and global autoencoder training.</p>
          <fig id="fig4" position="float">
            <label>Figure 4</label>
            <caption>
              <p>The data fusion framework for autoencoders. AE: Autoencoder; RMSE: root mean square error; GAE: global autoencoder.</p>
            </caption>
            <graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ir6032.fig.4.jpg" />
          </fig>
          <p>(1) Data Sampling<break/>Ensembling multiple sub-models is an effective approach for data fusion. Diversity among component models is an important factor affecting ensemble performance<sup>[<xref ref-type="bibr" rid="B55">55</xref>]</sup>. Thus, we used bootstrap sampling to create multiple training subsets for both IoT sensing data and network traffic data. Specifically, given a training dataset <italic>D</italic>, we perform <italic>K</italic> bootstrap iterations. In each iteration, a training subset is generated by randomly selecting <italic>N<sup>S</sup></italic> samples from <italic>D</italic> with replacement, which means that a sample in <italic>D</italic> can be selected multiple times or not at all in the training subset. By applying bootstrap sampling, we can generate diverse training subsets for each data modality.</p>
          <p>Finally, we can obtain <italic>K</italic> training subsets for both IoT sensing data (denoted as <italic>IDS</italic> = {<italic>ID</italic><sub>1</sub>, <italic>ID</italic><sub>2</sub>, …, <italic>ID<sub>K</sub></italic>}) and network traffic data (denoted as <italic>TDS</italic> = {<italic>TD</italic><sub>1</sub>, <italic>TD</italic><sub>2</sub>, …, <italic>TD<sub>K</sub></italic>}).</p>
          <p>(2) Local Autoencoder Training<break/>After obtaining the collections of training subsets <italic>IDS</italic> and <italic>TDS</italic> for both data modalities, we train multiple autoencoders through the following steps. First, for each training subset <italic>ID<sub>k</sub></italic> in <italic>IDS</italic>, we train a backbone autoencoder for IoT sensing data (denoted as <italic>IAE<sub>k</sub></italic>). Second, for each training subset <italic>TD<sub>k</sub></italic> in <italic>TDS</italic>, we train a backbone autoencoder for network traffic data (denoted as <italic>NAE<sub>k</sub></italic>). Finally, we can obtain a set of IoT sensing data-based local autoencoders <italic>IAES</italic> = {<italic>IAE</italic><sub>1</sub>, <italic>IAE</italic><sub>2</sub>, …, <italic>IAE<sub>K</sub></italic>} and a set of network traffic data-based local autoencoders <italic>NAES</italic> = {<italic>NAE</italic><sub>1</sub>, <italic>NAE</italic><sub>2</sub>, …, <italic>NAE<sub>K</sub></italic>}. This ensures that each training subset is used to train a dedicated local autoencoder, and the resulting collection of local autoencoders is capable of capturing diverse patterns within the IIoT data.</p>
          <p>(3) Global Autoencoder Training<break/>To fuse the multiple local autoencoders from <italic>IAES</italic> and <italic>NAES</italic>, we train a global autoencoder to capture their joint semantic correlations and behavioral patterns. The core idea is inspired by stacking, i.e., the global autoencoder is trained on the outputs of the local autoencoders, but in an unsupervised manner. The specific steps are as follows.</p>
          <p>
            <bold>RMSE generation</bold>: Given an IIoT data sample <italic>S<sub>t</sub></italic> = (<italic>X<sub>t</sub></italic>, <italic>Y<sub>t</sub></italic>), we input the IoT sensing sample <italic>X<sub>t</sub></italic> into each local autoencoder in <italic>IAES</italic>. For each local autoencoder <italic>IAE<sub>k</sub></italic>, it outputs a reconstructed sample <inline-formula><tex-math id="M1">$$ \tilde{X} $$</tex-math></inline-formula><italic><sub>t</sub></italic>, and we calculate the RMSE between <italic>X<sub>t</sub></italic> and <inline-formula><tex-math id="M1">$$ \tilde{X} $$</tex-math></inline-formula><italic><sub>t</sub></italic> [denoted as <italic>IRMSE</italic>(<italic>t</italic>)<italic><sub>k</sub></italic>]. Similarly, we represent the network traffic sample <italic>Y<sub>t</sub></italic> as a feature vector <italic>y<sub>t</sub></italic> and input <italic>y<sub>t</sub></italic> into each local autoencoder in <italic>NAES</italic>. Each local autoencoder <italic>NAE<sub>k</sub></italic> outputs a reconstructed sample <inline-formula><tex-math id="M1">$$ \tilde{y} $$</tex-math></inline-formula><italic><sub>t</sub></italic>, and we also calculate the RMSE between <italic>y<sub>t</sub></italic> and <inline-formula><tex-math id="M1">$$ \tilde{y} $$</tex-math></inline-formula><italic><sub>t</sub></italic> [denoted as <italic>NRMSE</italic>(<italic>t</italic>)<italic><sub>k</sub></italic>]. Then, we can obtain an RMSE vector for IIoT data sample <italic>S<sub>t</sub></italic> [denoted as <italic>ES</italic>(<italic>t</italic>) = &lt;<italic>IRMSE</italic>(<italic>t</italic>)<sub>1</sub>, <italic>IRMSE</italic>(<italic>t</italic>)<sub>2</sub>, …, <italic>IRMSE</italic>(<italic>t</italic>)<italic><sub>K</sub></italic>, <italic>NRMSE</italic>(<italic>t</italic>)<sub>1</sub>, <italic>NRMSE</italic>(<italic>t</italic>)<sub>2</sub>, …, <italic>NRMSE</italic>(<italic>t</italic>)<italic><sub>K</sub></italic>&gt;], which represents the distribution of anomaly scores across different semantic spaces. Finally, we can obtain a collection of RMSE vectors for the entire training dataset <italic>D</italic> [denoted as <italic>ES</italic> = {<italic>ES</italic>(1), <italic>ES</italic>(2), …, <italic>ES</italic>(|<italic>D</italic>|)}].</p>
          <p>
            <bold>Global autoencoder training</bold>: The global autoencoder (denoted as <italic>GAE</italic>) is trained to capture the normal patterns of the outputs of all the local autoencoders. Specifically, <italic>GAE</italic> takes an RMSE vector <italic>ES</italic>(<italic>t</italic>) as input, and outputs a reconstructed RMSE vector <italic>E</italic><inline-formula><tex-math id="M1">$$ \tilde{S} $$</tex-math></inline-formula>(<italic>t</italic>) through an encoder and a decoder. We utilize the same encoder and decoder structures as in Section 3.2.2 (i.e., an MLP encoder and an MLP decoder). We train <italic>GAE</italic> on the RMSE vector dataset <italic>ES</italic> by minimizing the loss function in Equation (9).</p>
			<p><disp-formula> <label>(9)</label> <tex-math id="E1"> $$  \operatorname{Loss}_{G}=\frac{1}{|D|} \sum_{t=1}^{|D|} \sqrt{\frac{\sum_{i=1}^{2 K}(E S(t)[i]-E \tilde{S}(t)[i])^{2}}{2 K}} $$ </tex-math></disp-formula></p>
          <p>Since <italic>GAE</italic> is trained on the outputs from all data modalities and all semantic spaces, it is expected to discover latent anomalies that are difficult to detect based on a single autoencoder or a single data modality. For example, some latent and stealthy anomalies might not necessarily cause all local autoencoders to generate high reconstruction RMSEs.</p>
        </sec>
        <sec id="sec3-3-2">
          <title>3.3.2. Anomaly detection</title>
          <p>Given an IIoT data sample <italic>S<sub>t</sub></italic> = (<italic>X<sub>t</sub></italic>, <italic>Y<sub>t</sub></italic>), we first input it to the set of local autoencoders <italic>AES</italic> = {<italic>IAE</italic><sub>1</sub>, <italic>IAE</italic><sub>2</sub>, …, <italic>IAE<sub>K</sub></italic>, <italic>NAE</italic><sub>1</sub>, <italic>NAE</italic><sub>2</sub>, …, <italic>NAE<sub>K</sub></italic>}, which outputs a set of RMSEs <italic>ES</italic>(<italic>t</italic>) = {<italic>IRMSE</italic>(<italic>t</italic>)<sub>1</sub>, <italic>IRMSE</italic>(<italic>t</italic>)<sub>2</sub>, …, <italic>IRMSE</italic>(<italic>t</italic>)<italic><sub>K</sub></italic>, <italic>NRMSE</italic>(<italic>t</italic>)<sub>1</sub>, <italic>NRMSE</italic>(<italic>t</italic>)<sub>2</sub>, …, <italic>NRMSE</italic>(<italic>t</italic>)<italic><sub>K</sub></italic>}. Second, we input <italic>ES</italic>(<italic>t</italic>) into the global autoencoder <italic>GAE</italic>, which outputs a reconstructed RMSE vector <italic>E</italic><inline-formula><tex-math id="M1">$$ \tilde{S} $$</tex-math></inline-formula>(<italic>t</italic>). We calculate the global RMSE between <italic>ES</italic>(<italic>t</italic>) and <italic>E</italic><inline-formula><tex-math id="M1">$$ \tilde{S} $$</tex-math></inline-formula>(<italic>t</italic>) [denoted as <italic>GRMSE</italic>(<italic>t</italic>)]. Finally, <italic>GRMSE</italic>(<italic>t</italic>) is compared with a predefined threshold <italic>δ</italic>. If <italic>GRMSE</italic>(<italic>t</italic>) &gt; <italic>δ</italic>, <italic>S<sub>t</sub></italic> is identified as abnormal, indicating a potential deviation from the expected patterns. Otherwise, <italic>S<sub>t</sub></italic> is identified as normal.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec4">
      <title>4. EXPERIMENTS</title>
      <p>This section evaluates AD-FIT through four sets of experiments. We first introduce the dataset and evaluation strategies. We then compare AD-FIT with representative baseline methods, analyze the effect of the anomaly detection threshold, examine the contribution of different data modalities through ablation experiments, and finally present representative cases to illustrate the effectiveness of the proposed data fusion framework. The experimental results are discussed together with their underlying reasons and implications.</p>
      <sec id="sec4-1">
        <title>4.1. Experiment setup</title>
        <sec id="sec4-1-1">
          <title>4.1.1. Dataset</title>
          <p>We evaluated AD-FIT based on the ToN_IoT dataset<sup>[<xref ref-type="bibr" rid="B56">56</xref>]</sup>, which includes heterogeneous data sources collected from IoT sensors and network traffic. It was collected from a realistic and large-scale testbed network designed by the IoT Lab of UNSW Canberra Cyber. The testbed network emulates a complex and scalable IIoT environment that includes virtual machines, physical systems, hacking platforms, cloud/fog platforms, and IoT sensors. Specifically, the IoT sensor subset of ToN_IoT includes 21 data dimensions sampled from 7 sensors, and the network traffic subset comprises 46 features. ToN_IoT was collected from March 31 to April 27, 2019 (spanning nearly a month), including 167 MB of IoT sensing data and 3.16 GB of network traffic data. The IoT sensing data were sampled at 1 Hz. The complete original dataset has a normal-to-abnormal event ratio of approximately 24:1. However, after preprocessing, temporal alignment, and test-sample construction, anomalous samples account for approximately 40.7% of the test set used in our experiments.</p>
          <p>In the experiments, the window length <italic>W</italic> was set to five sampling intervals. Since the IoT sensing data were sampled at 1 Hz, each IoT sensing sample contained five consecutive observations. Non-overlapping windows with a stride of five sampling intervals were applied to both the IoT sensing data and network traffic data, and the data falling within the same time window were paired to construct an IIoT data sample.</p>
          <p>Unless otherwise specified, <italic>K</italic> was set to 10 in all experiments. Accordingly, 10 local autoencoders were trained for each data modality, and the RMSE vector input to the global autoencoder had a dimension of 2<italic>K</italic>.</p>
        </sec>
        <sec id="sec4-1-2">
          <title>4.1.2. Evaluation strategies</title>
          <p>To comprehensively evaluate the anomaly detection performance, we used Accuracy, Precision, Recall, and F1-score as evaluation metrics. Accuracy alone is insufficient because a model may achieve high accuracy while failing to detect anomalous events. Hence, following the standard definitions provided by Sokolova and Lapalme<sup>[<xref ref-type="bibr" rid="B57">57</xref>]</sup>, we used Precision, Recall, and F1-score, as shown in Equations (10)-(12), where <italic>TP</italic> is the number of abnormal samples that are accurately detected, <italic>FP</italic> is the number of normal samples that are mistakenly identified as abnormal, and <italic>FN</italic> is the number of abnormal samples that are mistakenly identified as normal.</p>
		  <p><disp-formula> <label>(10)</label> <tex-math id="E1"> $$ Precision=\frac{T P}{T P+F P} $$ </tex-math></disp-formula></p>
		  <p><disp-formula> <label>(11)</label> <tex-math id="E1"> $$ Recall=\frac{T P}{T P+F N} $$ </tex-math></disp-formula></p>
		  <p><disp-formula> <label>(12)</label> <tex-math id="E1"> $$  F 1=\frac{2 \times Precision \times Recall }{Precision+Recall} $$ </tex-math></disp-formula></p>
        </sec>
      </sec>
      <sec id="sec4-2">
        <title>4.2. Experiment 1: Comparison Experiment</title>
        <p>To evaluate the comparative performance of AD-FIT, we compared it with the following seven baseline methods. All these baseline methods (a) operate in an unsupervised manner without requiring anomalous training samples; (b) consider both data modalities (i.e., the IoT sensing data and network traffic data); and (c) were configured using their best-performing parameter settings.</p>
        <p>
          <bold>Threshold</bold>: It first calculates the value range for each attribute in the normal samples. Then, given an IIoT data sample, if the value of any of its attributes exceeds the corresponding normal value range, it is identified as abnormal. IoT sensing samples have attributes corresponding to different sensors, while network traffic samples have attributes that are extracted statistical features.</p>
        <p>
          <bold>OC-SVM</bold>: It refers to the One-Class SVM anomaly detection model<sup>[<xref ref-type="bibr" rid="B6">6</xref>]</sup>. Specifically, it first learns a boundary that encapsulates the normal samples in a high-dimensional feature space, and then detects abnormal samples that fall outside this boundary.</p>
        <p>
          <bold>iForest</bold>: It refers to the iForest anomaly detection model, which leverages an ensemble of decision trees<sup>[<xref ref-type="bibr" rid="B7">7</xref>]</sup>. Specifically, it first splits samples by features using multiple decision trees, then considers samples with shorter average path lengths as anomalies.</p>
        <p>
          <bold>MLP-AE</bold>: Autoencoder-based anomaly detection identifies anomalous samples according to reconstruction errors<sup>[<xref ref-type="bibr" rid="B58">58</xref>]</sup>. In our MLP-AE baseline, both the encoder and decoder are implemented using two-layer MLPs.</p>
        <p>
          <bold>LSTM-AE</bold>: It also refers to an autoencoder-based anomaly detection model, which uses a stacked two-layer LSTM as encoder and decoder<sup>[<xref ref-type="bibr" rid="B14">14</xref>]</sup>.</p>
        <p>
          <bold>Hard Voting</bold>: It first trains a collection of 2<italic>K</italic> local autoencoders for both data modalities. Then, these local autoencoders make anomaly detection decisions independently. Finally, given an IIoT data sample, the sample is classified as abnormal if more than <italic>K</italic> local autoencoders classify it as abnormal.</p>
        <p>
          <bold>Soft Voting</bold>: It also first trains a collection of 2<italic>K</italic> local autoencoders for both data modalities. Then, the final anomaly detection decision is based on the local autoencoders’ reconstruction RMSEs. Specifically, for an IIoT data sample, we calculate the average reconstruction RMSE across all local autoencoders and compare it with the predefined threshold <italic>δ</italic>.</p>
        <p>In these baseline methods, <bold>Threshold</bold> is a rule-based method. <bold>OC-SVM</bold>, <bold>iForest</bold>, <bold>MLP-AE</bold>, and <bold>LSTM-AE</bold> are feature-level fusion methods. Specifically, given an IIoT data sample <italic>S<sub>t</sub></italic> = (<italic>X<sub>t</sub></italic>, <italic>Y<sub>t</sub></italic>), they first flatten the multivariate time-series IoT sensing sample <italic>X<sub>t</sub></italic> into a 1-dimensional vector <italic>x<sub>t</sub></italic>, and extract the statistical feature vector <italic>y<sub>t</sub></italic> for <italic>Y<sub>t</sub></italic>. Then, they concatenate <italic>x<sub>t</sub></italic> and <italic>y<sub>t</sub></italic> into a unified vector <italic>z<sub>t</sub></italic>, and train the anomaly detection models based on the unified vectors. <bold>Hard Voting</bold>, <bold>Soft Voting</bold>, and <bold>AD-FIT</bold> are model-level fusion methods. They allow the sub-models to generate outputs independently and fuse these outputs based on a global strategy. The comparison results are shown in <xref ref-type="table" rid="t1">Table 1</xref>, and the following tendencies could be discerned from the results.</p>
        <table-wrap id="t1">
          <label>Table 1</label>
          <caption>
            <p>The comparison with different baseline methods</p>
          </caption>
          <table frame="hsides" rules="groups">
            <thead>
              <tr>
                <td style="border-bottom:1;" />
                <td style="border-bottom:1;">
                  <bold>Accuracy</bold>
                </td>
                <td style="border-bottom:1;">
                  <bold>Precision</bold>
                </td>
                <td style="border-bottom:1;">
                  <bold>Recall</bold>
                </td>
                <td style="border-bottom:1;">
                  <bold>F1</bold>
                </td>
              </tr>
            </thead>
            <tbody>
              <tr>
                <td>
                  <bold>Threshold</bold>
                </td>
                <td>0.726</td>
                <td>0.996</td>
                <td>0.329</td>
                <td>0.495</td>
              </tr>
              <tr>
                <td>
                  <bold>OC-SVM</bold>
                </td>
                <td>0.669</td>
                <td>0.557</td>
                <td>0.911</td>
                <td>0.691</td>
              </tr>
              <tr>
                <td>
                  <bold>iForest</bold>
                </td>
                <td>0.774</td>
                <td>0.822</td>
                <td>0.568</td>
                <td>0.672</td>
              </tr>
              <tr>
                <td>
                  <bold>LSTM-AE</bold>
                </td>
                <td>0.758</td>
                <td>0.641</td>
                <td>0.923</td>
                <td>0.756</td>
              </tr>
              <tr>
                <td>
                  <bold>MLP-AE</bold>
                </td>
                <td>0.678</td>
                <td>0.558</td>
                <td>0.998</td>
                <td>0.716</td>
              </tr>
              <tr>
                <td>
                  <bold>Hard Voting</bold>
                </td>
                <td>0.697</td>
                <td>0.664</td>
                <td>0.517</td>
                <td>0.581</td>
              </tr>
              <tr>
                <td>
                  <bold>Soft Voting</bold>
                </td>
                <td>0.844</td>
                <td>0.801</td>
                <td>0.821</td>
                <td>0.810</td>
              </tr>
              <tr>
                <td>
                  <bold>AD-FIT</bold>
                </td>
                <td>0.903</td>
                <td>0.879</td>
                <td>0.883</td>
                <td>0.881</td>
              </tr>
            </tbody>
          </table>
          <table-wrap-foot>
            <fn>
              <p>OC-SVM: One-Class support vector machine; LSTM: long short-term memory; AE: autoencoder; MLP: multilayer perceptron.</p>
            </fn>
          </table-wrap-foot>
        </table-wrap>
        <p>First, although threshold-based detection is widely used in engineering practice, it performs poorly on this dataset, particularly when detecting gradual pattern deviations without abrupt changes in sensor signals. Based on our experimental results, we found that threshold-based detection can detect only simple anomalies in IoT devices, such as significant temperature fluctuations. However, most anomalies in the dataset are caused by complex events, such as network intrusions. For instance, attackers may disrupt the operation of remote sensing devices through intrusion, causing slight deviations from normal patterns in the detected data.</p>
        <p>Second, deep-learning-based methods (i.e., <bold>MLP-AE</bold> and <bold>LSTM-AE</bold>) perform better than traditional machine-learning-based methods (i.e., <bold>OC-SVM</bold> and <bold>iForest</bold>). This is because the deep-learning-based methods have more powerful learning capability that can capture the latent and semantic features of the heterogeneous data, while traditional machine-learning-based methods capture only relatively shallow feature representations.</p>
        <p>Third, <bold>LSTM-AE</bold> outperforms <bold>MLP-AE</bold>. LSTM is more effective at modeling time-series data as compared to MLP. This indicates that IIoT data samples contain strong temporal patterns, particularly in the IoT sensing time series.</p>
        <p>Fourth, the suboptimal performance of <bold>Hard Voting</bold> is attributed to its difficulty in detecting complex anomalies. For example, anomalies caused by network intrusions may involve subtle deviations from normal behavior. Additionally, because each autoencoder’s decision is weighted equally, the aggregation may be ineffective when an anomaly can be detected by only a small subset of the local autoencoders.</p>
        <p>Fifth, among the model-level fusion methods, <bold>AD-FIT</bold> outperforms <bold>Hard Voting</bold> and <bold>Soft Voting</bold>. It shows that learning a global model on the outputs of the sub-models is a better model fusion strategy than the simple voting mechanisms. Although rich supervised signals (e.g., classification probability distributions, classification residuals) are unavailable for data fusion under unsupervised learning, the reconstruction errors generated by diverse local autoencoders can still provide clues for model-level fusion to make joint decisions. AD-FIT also has the best overall performance on anomaly detection.</p>
      </sec>
      <sec id="sec4-3">
        <title>4.3. Experiment 2: Parameter Tuning Experiment</title>
        <p>The most important parameter in AD-FIT is <italic>δ</italic>, the anomaly-score threshold (Section 3.3.2). Here, we varied <italic>δ</italic> in the range [0.1, 1], and the experimental results are shown in <xref ref-type="fig" rid="fig5">Figure 5</xref>. First, Precision and Recall show opposite trends. As <italic>δ</italic> increases, the decision criterion becomes more conservative for anomaly detection, so fewer anomalies are detected. As a result, Precision shows a stable increasing phase and Recall shows a stable decreasing phase. Second, Accuracy and F1 exhibit a significant upward trend when increasing <italic>δ</italic> from 0.1 to 0.2, followed by a stable downward trend by further increasing <italic>δ</italic>. The best overall performance was achieved at <italic>δ</italic> = 0.2.</p>
        <fig id="fig5" position="float" width="550">
          <label>Figure 5</label>
          <caption>
            <p>The effect of parameter <italic>δ</italic>.</p>
          </caption>
          <graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ir6032.fig.5.jpg" />
        </fig>
      </sec>
      <sec id="sec4-4">
        <title>4.4. Experiment 3: Ablation Experiment</title>
        <p>To investigate the impact of the two data modalities (i.e., IoT sensing data and network traffic data) on anomaly-detection performance, we conducted separate tests on the autoencoder for IoT sensing data (IoT-AE) and the autoencoder for network traffic data (<bold>Network-AE</bold>). Here, IoT-AE was trained as a global autoencoder according to Section 3.3.1 without local autoencoders of the network traffic data. Similarly, Network-AE was trained as a global autoencoder without local autoencoders of the IoT sensing data.</p>
        <p>In the first experiment, we compared IoT-AE with AD-FIT and other baseline methods as follows. These baseline methods consider only IoT sensing data.</p>
        <p>
          <bold>MLP-IAE</bold>: An autoencoder-based anomaly detection model based only on IoT sensing data. It uses a two-layer MLP as the encoder and decoder.</p>
        <p>
          <bold>LSTM-IAE</bold>: It is also an autoencoder-based anomaly detection model based only on IoT sensing data. It uses a two-layer stacked LSTM as the encoder and decoder.</p>
        <p>
          <bold>MTGNN</bold>: This is a prediction-based anomaly detection model. Specifically, it first trains a multivariate time-series prediction model based on a novel GNN proposed in<sup>[<xref ref-type="bibr" rid="B59">59</xref>]</sup>, and then detects anomalies by comparing the predicted values and the true values.</p>
        <p>The experimental results are shown in <xref ref-type="table" rid="t2">Table 2</xref>. First, LSTM-IAE, MTGNN, and IoT-AE achieve better performance than MLP-IAE. This indicates that temporal patterns are crucial for IoT sensing data-based anomaly detection, and therefore should not be ignored. Second, IoT-AE has a slight advantage over LSTM-IAE. This result shows that using a graph to capture spatial correlations across dimensions of multivariate time-series data benefits anomaly detection. Third, MTGNN performs worse than LSTM-IAE. MTGNN detects anomalies by predicting the value at a given time point and comparing it with the observed value at that time point. The prediction-based anomaly detection model evaluates prediction error at a single time point, while the autoencoder-based anomaly detection model evaluates reconstruction error over a time range. This suggests that a large proportion of anomalies in our dataset are subtle and collectively evolving trend anomalies, rather than significant value changes at a given time point.</p>
        <table-wrap id="t2">
          <label>Table 2</label>
          <caption>
            <p>The evaluation of the autoencoder for IoT sensing data</p>
          </caption>
          <table frame="hsides" rules="groups">
            <thead>
              <tr>
                <td style="border-bottom:1;" />
                <td style="border-bottom:1;">
                  <bold>Accuracy</bold>
                </td>
                <td style="border-bottom:1;">
                  <bold>Precision</bold>
                </td>
                <td style="border-bottom:1;">
                  <bold>Recall</bold>
                </td>
                <td style="border-bottom:1;">
                  <bold>F1</bold>
                </td>
              </tr>
            </thead>
            <tbody>
              <tr>
                <td>
                  <bold>MLP-IAE</bold>
                </td>
                <td>0.701</td>
                <td>0.595</td>
                <td>0.837</td>
                <td>0.696</td>
              </tr>
              <tr>
                <td>
                  <bold>LSTM-IAE</bold>
                </td>
                <td>0.855</td>
                <td>0.776</td>
                <td>0.906</td>
                <td>0.836</td>
              </tr>
              <tr>
                <td>
                  <bold>MTGNN</bold>
                </td>
                <td>0.829</td>
                <td>0.795</td>
                <td>0.783</td>
                <td>0.788</td>
              </tr>
              <tr>
                <td>
                  <bold>IoT-AE</bold>
                </td>
                <td>0.865</td>
                <td>0.821</td>
                <td>0.853</td>
                <td>0.837</td>
              </tr>
              <tr>
                <td>
                  <bold>AD-FIT</bold>
                </td>
                <td>0.903</td>
                <td>0.879</td>
                <td>0.883</td>
                <td>0.881</td>
              </tr>
            </tbody>
          </table>
          <table-wrap-foot>
            <fn>
              <p>IoT: Internet of Things; MLP: multilayer perceptron; IAE: IoT sensing autoencoder; LSTM: long short-term memory; AE: autoencoder.</p>
            </fn>
          </table-wrap-foot>
        </table-wrap>
        <p>In the second experiment, we compared Network-AE with AD-FIT and other baseline methods as follows. These baseline methods consider only network traffic data. We extracted statistical features according to Section 3.2.2 as the input to these baseline methods.</p>
        <p>
          <bold>iForest</bold>: The iForest model was trained using only network traffic data.</p>
        <p>
          <bold>MLP-NAE</bold>: An autoencoder-based anomaly detection model trained only on network traffic data. It uses a two-layer MLP as the encoder and decoder.</p>
        <p>The experimental results are shown in <xref ref-type="table" rid="t3">Table 3</xref>. First, deep-learning-based models (i.e., MLP-NAE and Network-AE) outperform traditional machine-learning-based models (i.e., iForest). This finding is consistent with the results of the previous experiment presented in Section 4.2. Second, Network-AE outperforms MLP-NAE. It demonstrates the advantages of model ensemble, which can reduce the risk of overfitting. Finally, AD-FIT achieves the best overall performance among the evaluated baseline methods. It demonstrates the advantages of multi-view data fusion.</p>
        <table-wrap id="t3">
          <label>Table 3</label>
          <caption>
            <p>Evaluation of the autoencoder on network traffic data</p>
          </caption>
          <table frame="hsides" rules="groups">
            <thead>
              <tr>
                <td style="border-bottom:1;" />
                <td style="border-bottom:1;">
                  <bold>Accuracy</bold>
                </td>
                <td style="border-bottom:1;">
                  <bold>Precision</bold>
                </td>
                <td style="border-bottom:1;">
                  <bold>Recall</bold>
                </td>
                <td style="border-bottom:1;">
                  <bold>F1</bold>
                </td>
              </tr>
            </thead>
            <tbody>
              <tr>
                <td>
                  <bold>iForest</bold>
                </td>
                <td>0.751</td>
                <td>0.799</td>
                <td>0.517</td>
                <td>0.628</td>
              </tr>
              <tr>
                <td>
                  <bold>MLP-NAE</bold>
                </td>
                <td>0.762</td>
                <td>0.635</td>
                <td>0.974</td>
                <td>0.769</td>
              </tr>
              <tr>
                <td>
                  <bold>Network-AE</bold>
                </td>
                <td>0.901</td>
                <td>0.967</td>
                <td>0.784</td>
                <td>0.865</td>
              </tr>
              <tr>
                <td>
                  <bold>AD-FIT</bold>
                </td>
                <td>0.903</td>
                <td>0.879</td>
                <td>0.883</td>
                <td>0.881</td>
              </tr>
            </tbody>
          </table>
          <table-wrap-foot>
            <fn>
              <p>MLP: Multilayer perceptron; NAE: network traffic autoencoder; AE: autoencoder.</p>
            </fn>
          </table-wrap-foot>
        </table-wrap>
      </sec>
      <sec id="sec4-5">
        <title>4.5. Experiment 4: Case Study</title>
        <p>In this section, we use three cases to demonstrate the effectiveness of AD-FIT. <xref ref-type="fig" rid="fig6">Figure 6A</xref> and <xref ref-type="fig" rid="fig6">B</xref> present the IoT sensing data and network traffic data for Case 1, respectively; <xref ref-type="fig" rid="fig6">Figure 6C</xref> and <xref ref-type="fig" rid="fig6">D</xref> present the corresponding data for Case 2; and <xref ref-type="fig" rid="fig6">Figure 6E</xref> and <xref ref-type="fig" rid="fig6">F</xref> present the corresponding data for Case 3. For each case, the IoT sensing data comprise five dimensions (latitude, longitude, temperature, pressure, and humidity), and the network traffic data comprise six representative features.</p>
        <fig id="fig6" position="float">
          <label>Figure 6</label>
          <caption>
            <p>Representative cases illustrating the complementary roles of IoT sensing data and network traffic data in AD-FIT. (A) IoT sensing data of Case 1, which show no obvious deviation from normal patterns; (B) network traffic data of Case 1, which show significant deviations in multiple features; (C) IoT sensing data of Case 2, which show pronounced fluctuations; (D) network traffic data of Case 2, which show no obvious deviation from normal patterns; (E) IoT sensing data of Case 3, in which the <italic>humidity</italic> dimension shows a slight fluctuation; and (F) network traffic data of Case 3, in which the <italic>missed_bytes</italic> feature shows a slight deviation. IoT: Internet of Things.</p>
          </caption>
          <graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ir6032.fig.6.jpg" />
        </fig>
        <p>The first case shows an abnormal sample that cannot be detected by considering only the IoT sensing data, since the five displayed IoT sensing dimensions show no obvious deviation from their normal patterns, as shown in <xref ref-type="fig" rid="fig6">Figure 6A</xref>. On the other hand, we can observe a significant deviation in most features of the network traffic data, as shown in <xref ref-type="fig" rid="fig6">Figure 6B</xref>, where “Norm” stands for the average value of a specific feature in all the normal samples and “Real” stands for the real value of the feature in the target sample. Here, <italic>dst_bytes</italic>, <italic>dst_ip_bytes</italic>, <italic>dst_pkts</italic>, <italic>missed_bytes</italic>, <italic>src_ip_bytes</italic>, and <italic>src_pkts</italic> denote bytes of packets received by the destination system, bytes of the IP header of the destination system, number of packets received by the destination system, bytes of missed packets, bytes of the IP header of the source system, and number of packets sent by the source system. As a result, this sample can be detected by an anomaly detector based on network traffic data. The second case shows the opposite scenario, where the network traffic data show almost no deviation across all features, while the IoT sensing data curve has pronounced fluctuations (as shown in <xref ref-type="fig" rid="fig6">Figure 6C</xref> and <xref ref-type="fig" rid="fig6">D</xref>). Therefore, this sample cannot be detected by an anomaly detector based on network traffic data alone but can be detected using IoT sensing data.</p>
        <p>As shown in <xref ref-type="fig" rid="fig6">Figure 6E</xref> and <xref ref-type="fig" rid="fig6">F</xref>, the third case shows a more subtle anomalous sample, where neither IoT sensing data nor network traffic data exhibit significant deviations from normal temporal patterns. Nevertheless, the <italic>humidity</italic> dimension of the IoT sensing data and the <italic>missed_bytes</italic> feature of the network traffic data show slight fluctuations. By considering the two factors simultaneously, AD-FIT can successfully identify this abnormal sample, highlighting the advantage of the data fusion scheme.</p>
      </sec>
    </sec>
    <sec id="sec5">
      <title>5. CONCLUSIONS</title>
      <p>In this paper, we investigate anomaly detection in IIoT systems. We propose AD-FIT, a novel unsupervised deep learning framework to identify anomalies by fusing multiple autoencoders based on the joint consideration of IoT sensing data and network traffic data in IIoT systems. By exploiting and learning the correlation patterns between multiple data modalities, AD-FIT achieves the best performance among the evaluated baseline methods on the ToN_IoT dataset.</p>
      <p>Overall, the experimental results consistently demonstrate AD-FIT’s effectiveness. The comparison experiment shows that learning a global model from local autoencoder reconstruction errors is more effective than direct feature-level fusion or simple voting strategies. The parameter tuning experiment identifies an appropriate trade-off between precision and recall, while the ablation experiment confirms the complementary contributions of different data modalities. Finally, the case study further illustrates that jointly considering IoT sensing data and network traffic data enables AD-FIT to detect subtle anomalies that may be overlooked when either modality is used alone.</p>
      <p>Future work will focus on two directions. First, AD-FIT detects anomalous events but does not classify them. Hence, equipping AD-FIT with anomaly classification capabilities could facilitate more effective responses to detected anomalies. Second, diagnosing detected anomalies and tracing their root causes are also important directions for future research.</p>
    </sec>
  </body>
  <back>
    <sec>
      <title>DECLARATIONS</title>
      <sec>
        <title>Authors’ contributions</title>
        <p>Responsible for the overall study: Wang, F.</p>
        <p>Wrote the manuscript: Lv, M.; Wang, F.</p>
        <p>Contributed to the discussion and revision of the manuscript: Wang, L.; Chen, H.; Zhang, Z.; Tao, Y.</p>
        <p>All authors read and approved the final manuscript.</p>
      </sec>
      <sec>
        <title>Availability of data and materials</title>
        <p>The data used in this paper are publicly available at <uri xlink:href="https://research.unsw.edu.au/projects/toniot-datasets">https://research.unsw.edu.au/projects/toniot-datasets</uri>.</p>
      </sec>
      <sec>
        <title>AI and AI-assisted tools statement</title>
        <p>During revision of this manuscript, the authors used the AI tool ChatGPT (version GPT-5.6 Sol, released 2026-07-09) to edit the language and improve the clarity of the text. All AI-assisted suggestions were critically reviewed and edited by the authors, who take full responsibility for the final content. No AI-assisted tools were used to generate experimental data, conduct experiments, or determine the research conclusions.</p>
      </sec>
      <sec>
        <title>Financial support and sponsorship</title>
        <p>The work is supported by the Shaoxing Science and Technology Plan Project (2025B11004), Quzhou Science and Technology Research Project (2025K132), National Natural Science Foundation of China (62372410), and the Key Technology Research and Development Program of Shandong Province (2024SZD1A11).</p>
      </sec>
      <sec>
        <title>Conflicts of interest</title>
        <p>Tao, Y. is affiliated with Shaoxing Smart City Group Co., LTD., while the other authors declare no conflicts of interest.</p>
      </sec>
      <sec>
        <title>Ethical approval and consent to participate</title>
        <p>Not applicable.</p>
      </sec>
      <sec>
        <title>Consent for publication</title>
        <p>Not applicable.</p>
      </sec>
      <sec>
        <title>Copyright</title>
        <p>© The author(s) 2026.</p>
      </sec>
    </sec>
    <ref-list>
      <ref id="B1">
        <label>1</label>
        <nlm-citation publication-type="journal">
          <person-group person-group-type="author">
            <name>
              <surname>Pang</surname>
              <given-names>G</given-names>
            </name>
            <name>
              <surname>Shen</surname>
              <given-names>C</given-names>
            </name>
            <name>
              <surname>Cao</surname>
              <given-names>L</given-names>
            </name>
            <name>
              <surname>Hengel</surname>
              <given-names>AVD</given-names>
            </name>
          </person-group>
          <article-title>Deep learning for anomaly detection: a review</article-title>
          <source>ACM Comput Surv</source>
          <year>2022</year>
          <volume>54</volume>
          <fpage>1</fpage>
          <lpage>38</lpage>
          <pub-id pub-id-type="doi">10.1145/3439950</pub-id>
        </nlm-citation>
      </ref>
      <ref id="B2">
        <label>2</label>
        <nlm-citation publication-type="journal">
          <person-group person-group-type="author">
            <name>
              <surname>Omar</surname>
              <given-names>S</given-names>
            </name>
            <name>
              <surname>Ngadi</surname>
              <given-names>A</given-names>
            </name>
            <name>
              <surname>Jebur</surname>
              <given-names>HH</given-names>
            </name>
          </person-group>
          <article-title>Machine learning techniques for anomaly detection: an overview</article-title>
          <source>Int J Comput Appl</source>
          <year>2013</year>
          <volume>79</volume>
          <fpage>33</fpage>
          <lpage>41</lpage>
          <pub-id pub-id-type="doi">10.5120/13715-1478</pub-id>
        </nlm-citation>
      </ref>
      <ref id="B3">
        <label>3</label>
        <nlm-citation publication-type="book">
          <person-group person-group-type="author">
            <name>
              <surname>Audibert</surname>
              <given-names>J</given-names>
            </name>
            <name>
              <surname>Michiardi</surname>
              <given-names>P</given-names>
            </name>
            <name>
              <surname>Guyard</surname>
              <given-names>F</given-names>
            </name>
            <name>
              <surname>Marti</surname>
              <given-names>S</given-names>
            </name>
            <name>
              <surname>Zuluaga</surname>
              <given-names>MA</given-names>
            </name>
          </person-group>
          <comment>USAD: unsupervised anomaly detection on multivariate time series. In <italic>Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery &amp; Data Min</italic>. Association for Computing Machinery; 2020. pp. 3395-404.</comment>
          <pub-id pub-id-type="doi">10.1145/3394486.3403392</pub-id>
        </nlm-citation>
      </ref>
      <ref id="B4">
        <label>4</label>
        <nlm-citation publication-type="journal">
          <person-group person-group-type="author">
            <name>
              <surname>Singhal</surname>
              <given-names>A</given-names>
            </name>
            <name>
              <surname>Seborg</surname>
              <given-names>DE</given-names>
            </name>
          </person-group>
          <article-title>Clustering multivariate time‐series data</article-title>
          <source>J Chemometrics</source>
          <year>2005</year>
          <volume>19</volume>
          <fpage>427</fpage>
          <lpage>38</lpage>
          <pub-id pub-id-type="doi">10.1002/cem.945</pub-id>
        </nlm-citation>
      </ref>
      <ref id="B5">
        <label>5</label>
        <nlm-citation publication-type="journal">
          <person-group person-group-type="author">
            <name>
              <surname>Karim</surname>
              <given-names>S</given-names>
            </name>
            <name>
              <surname>Rousanuzzaman</surname>
              <given-names>M</given-names>
            </name>
            <name>
              <surname>Yunus</surname>
              <given-names>PA</given-names>
            </name>
            <name>
              <surname>Khan</surname>
              <given-names>PH</given-names>
            </name>
            <name>
              <surname>Asif</surname>
              <given-names>M</given-names>
            </name>
          </person-group>
          <article-title>Implementation of K-means clustering for intrusion detection</article-title>
          <source>Int J Sci Res Comput Sci Eng Inf Technol</source>
          <year>2019</year>
          <volume>5</volume>
          <fpage>1232</fpage>
          <lpage>41</lpage>
          <pub-id pub-id-type="doi">10.32628/cseit1952332</pub-id>
        </nlm-citation>
      </ref>
      <ref id="B6">
        <label>6</label>
        <nlm-citation publication-type="journal">
          <person-group person-group-type="author">
            <name>
              <surname>Erfani</surname>
              <given-names>SM</given-names>
            </name>
            <name>
              <surname>Rajasegarar</surname>
              <given-names>S</given-names>
            </name>
            <name>
              <surname>Karunasekera</surname>
              <given-names>S</given-names>
            </name>
            <name>
              <surname>Leckie</surname>
              <given-names>C</given-names>
            </name>
          </person-group>
          <article-title>High-dimensional and large-scale anomaly detection using a linear one-class SVM with deep learning</article-title>
          <source>Pattern Recognit</source>
          <year>2016</year>
          <volume>58</volume>
          <fpage>121</fpage>
          <lpage>34</lpage>
          <pub-id pub-id-type="doi">10.1016/j.patcog.2016.03.028</pub-id>
        </nlm-citation>
      </ref>
      <ref id="B7">
        <label>7</label>
        <nlm-citation publication-type="book">
          <person-group person-group-type="author">
            <name>
              <surname>Liu</surname>
              <given-names>FT</given-names>
            </name>
            <name>
              <surname>Ting</surname>
              <given-names>KM</given-names>
            </name>
            <name>
              <surname>Zhou</surname>
              <given-names>ZH</given-names>
            </name>
          </person-group>
          <comment>Isolation forest. In <italic>2008 Eighth IEEE International Conference on Data Mining</italic>, Pisa, Italy. Dec 15-19, 2008. IEEE; 2008. pp. 413-22.</comment>
          <pub-id pub-id-type="doi">10.1109/ICDM.2008.17</pub-id>
        </nlm-citation>
      </ref>
      <ref id="B8">
        <label>8</label>
        <nlm-citation publication-type="book">
          <person-group person-group-type="author">
            <name>
              <surname>Malhotra</surname>
              <given-names>P</given-names>
            </name>
            <name>
              <surname>Ramakrishnan</surname>
              <given-names>A</given-names>
            </name>
            <name>
              <surname>Anand</surname>
              <given-names>G</given-names>
            </name>
            <name>
              <surname>Vig</surname>
              <given-names>L</given-names>
            </name>
            <name>
              <surname>Agarwal</surname>
              <given-names>P</given-names>
            </name>
            <name>
              <surname>Shroff</surname>
              <given-names>G</given-names>
            </name>
          </person-group>
          <comment>LSTM-based encoder-decoder for multi-sensor anomaly detection. <italic>arXiv</italic> <bold>2016</bold>, arXiv:1607.00148. Available online: <uri xlink:href="https://doi.org/10.48550/arXiv.1607.00148">https://doi.org/10.48550/arXiv.1607.00148</uri>. (accessed 2026-09-21)</comment>
        </nlm-citation>
      </ref>
      <ref id="B9">
        <label>9</label>
        <nlm-citation publication-type="journal">
          <person-group person-group-type="author">
            <name>
              <surname>Zeng</surname>
              <given-names>F</given-names>
            </name>
            <name>
              <surname>Chen</surname>
              <given-names>M</given-names>
            </name>
            <name>
              <surname>Qian</surname>
              <given-names>C</given-names>
            </name>
            <name>
              <surname>Wang</surname>
              <given-names>Y</given-names>
            </name>
            <name>
              <surname>Zhou</surname>
              <given-names>Y</given-names>
            </name>
            <name>
              <surname>Tang</surname>
              <given-names>W</given-names>
            </name>
          </person-group>
          <article-title>Multivariate time series anomaly detection with adversarial transformer architecture in the Internet of Things</article-title>
          <source>Future Gener Comput Syst</source>
          <year>2023</year>
          <volume>144</volume>
          <fpage>244</fpage>
          <lpage>55</lpage>
          <pub-id pub-id-type="doi">10.1016/j.future.2023.02.015</pub-id>
        </nlm-citation>
      </ref>
      <ref id="B10">
        <label>10</label>
        <nlm-citation publication-type="book">
          <person-group person-group-type="author">
            <name>
              <surname>Nanduri</surname>
              <given-names>A</given-names>
            </name>
            <name>
              <surname>Sherry</surname>
              <given-names>L</given-names>
            </name>
          </person-group>
          <comment>Anomaly detection in aircraft data using recurrent neural networks (RNN). In 2016 Integrated Communications Navigation and Surveillance (ICNS), Herndon, USA. Apr 19-21, 2016. IEEE; 2016. pp. 5C2-1-8.</comment>
          <pub-id pub-id-type="doi">10.1109/ICNSURV.2016.7486356</pub-id>
        </nlm-citation>
      </ref>
      <ref id="B11">
        <label>11</label>
        <nlm-citation publication-type="journal">
          <person-group person-group-type="author">
            <name>
              <surname>Hwang</surname>
              <given-names>R</given-names>
            </name>
            <name>
              <surname>Peng</surname>
              <given-names>M</given-names>
            </name>
            <name>
              <surname>Huang</surname>
              <given-names>C</given-names>
            </name>
            <name>
              <surname>Lin</surname>
              <given-names>P</given-names>
            </name>
            <name>
              <surname>Nguyen</surname>
              <given-names>V</given-names>
            </name>
          </person-group>
          <article-title>An unsupervised deep learning model for early network traffic anomaly detection</article-title>
          <source>IEEE Access</source>
          <year>2020</year>
          <volume>8</volume>
          <fpage>30387</fpage>
          <lpage>99</lpage>
          <pub-id pub-id-type="doi">10.1109/access.2020.2973023</pub-id>
        </nlm-citation>
      </ref>
      <ref id="B12">
        <label>12</label>
        <nlm-citation publication-type="journal">
          <person-group person-group-type="author">
            <name>
              <surname>Dutta</surname>
              <given-names>V</given-names>
            </name>
            <name>
              <surname>Pawlicki</surname>
              <given-names>M</given-names>
            </name>
            <name>
              <surname>Kozik</surname>
              <given-names>R</given-names>
            </name>
            <name>
              <surname>Choraś</surname>
              <given-names>M</given-names>
            </name>
          </person-group>
          <article-title>Unsupervised network traffic anomaly detection with deep autoencoders</article-title>
          <source>Log J IGPL</source>
          <year>2022</year>
          <volume>30</volume>
          <fpage>912</fpage>
          <lpage>25</lpage>
          <pub-id pub-id-type="doi">10.1093/jigpal/jzac002</pub-id>
        </nlm-citation>
      </ref>
      <ref id="B13">
        <label>13</label>
        <nlm-citation publication-type="journal">
          <person-group person-group-type="author">
            <name>
              <surname>Xu</surname>
              <given-names>W</given-names>
            </name>
            <name>
              <surname>Jang-Jaccard</surname>
              <given-names>J</given-names>
            </name>
            <name>
              <surname>Singh</surname>
              <given-names>A</given-names>
            </name>
            <name>
              <surname>Wei</surname>
              <given-names>Y</given-names>
            </name>
            <name>
              <surname>Sabrina</surname>
              <given-names>F</given-names>
            </name>
          </person-group>
          <article-title>Improving performance of autoencoder-based network anomaly detection on NSL-KDD dataset</article-title>
          <source>IEEE Access</source>
          <year>2021</year>
          <volume>9</volume>
          <fpage>140136</fpage>
          <lpage>46</lpage>
          <pub-id pub-id-type="doi">10.1109/access.2021.3116612</pub-id>
        </nlm-citation>
      </ref>
      <ref id="B14">
        <label>14</label>
        <nlm-citation publication-type="book">
          <comment>Said Elsayed, M.; Le-Khac, N. A.; Dev, S.; Jurcut, A. D. Network anomaly detection using LSTM based autoencoder. In <italic>Proceedings of the 16th ACM Symposium on QoS and Security for Wireless and Mobile Networks</italic>. Association for Computing Machinery; 2020. pp. 37-45.</comment>
          <pub-id pub-id-type="doi">10.1145/3416013.3426457</pub-id>
        </nlm-citation>
      </ref>
      <ref id="B15">
        <label>15</label>
        <nlm-citation publication-type="journal">
          <person-group person-group-type="author">
            <name>
              <surname>Zare</surname>
              <given-names>F</given-names>
            </name>
            <name>
              <surname>Mahmoudi-Nasr</surname>
              <given-names>P</given-names>
            </name>
            <name>
              <surname>Yousefpour</surname>
              <given-names>R</given-names>
            </name>
          </person-group>
          <article-title>A real-time network based anomaly detection in industrial control systems</article-title>
          <source>Int J Crit Infrastruct Prot</source>
          <year>2024</year>
          <volume>45</volume>
          <fpage>100676</fpage>
          <pub-id pub-id-type="doi">10.1016/j.ijcip.2024.100676</pub-id>
        </nlm-citation>
      </ref>
      <ref id="B16">
        <label>16</label>
        <nlm-citation publication-type="book">
          <person-group person-group-type="author">
            <name>
              <surname>Xu</surname>
              <given-names>J</given-names>
            </name>
            <name>
              <surname>Wu</surname>
              <given-names>H</given-names>
            </name>
            <name>
              <surname>Wang</surname>
              <given-names>J</given-names>
            </name>
            <name>
              <surname>Long</surname>
              <given-names>M</given-names>
            </name>
          </person-group>
          <comment>Anomaly transformer: time series anomaly detection with association discrepancy. In <italic>ICLR 2022 Conference</italic>. 2022. <uri xlink:href="https://openreview.net/forum?id=LzQQ89U1qm_">https://openreview.net/forum?id=LzQQ89U1qm_</uri>. (accessed 2026-09-21)</comment>
        </nlm-citation>
      </ref>
      <ref id="B17">
        <label>17</label>
        <nlm-citation publication-type="journal">
          <person-group person-group-type="author">
            <name>
              <surname>Tuli</surname>
              <given-names>S</given-names>
            </name>
            <name>
              <surname>Casale</surname>
              <given-names>G</given-names>
            </name>
            <name>
              <surname>Jennings</surname>
              <given-names>NR</given-names>
            </name>
          </person-group>
          <article-title>TranAD: deep transformer networks for anomaly detection in multivariate time series data</article-title>
          <source>Proc VLDB Endow</source>
          <year>2022</year>
          <volume>15</volume>
          <fpage>1201</fpage>
          <lpage>14</lpage>
          <pub-id pub-id-type="doi">10.14778/3514061.3514067</pub-id>
        </nlm-citation>
      </ref>
      <ref id="B18">
        <label>18</label>
        <nlm-citation publication-type="book">
          <person-group person-group-type="author">
            <name>
              <surname>Zhang</surname>
              <given-names>W</given-names>
            </name>
            <name>
              <surname>Zhang</surname>
              <given-names>C</given-names>
            </name>
            <name>
              <surname>Tsung</surname>
              <given-names>F</given-names>
            </name>
          </person-group>
          <comment>GRELEN: multivariate time series anomaly detection from the perspective of graph relational learning. In <italic>Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence (IJCAI-22)</italic>. 2022. pp. 2390-7.</comment>
          <pub-id pub-id-type="doi">10.24963/ijcai.2022/332</pub-id>
        </nlm-citation>
      </ref>
      <ref id="B19">
        <label>19</label>
        <nlm-citation publication-type="journal">
          <person-group person-group-type="author">
            <name>
              <surname>Zhou</surname>
              <given-names>Q</given-names>
            </name>
            <name>
              <surname>Chen</surname>
              <given-names>J</given-names>
            </name>
            <name>
              <surname>Liu</surname>
              <given-names>H</given-names>
            </name>
            <name>
              <surname>He</surname>
              <given-names>S</given-names>
            </name>
            <name>
              <surname>Meng</surname>
              <given-names>W</given-names>
            </name>
          </person-group>
          <article-title>Detecting multivariate time series anomalies with zero known label</article-title>
          <source>Proc AAAI Conf Artif Intell</source>
          <year>2023</year>
          <volume>37</volume>
          <fpage>4963</fpage>
          <lpage>71</lpage>
          <pub-id pub-id-type="doi">10.1609/aaai.v37i4.25623</pub-id>
        </nlm-citation>
      </ref>
      <ref id="B20">
        <label>20</label>
        <nlm-citation publication-type="book">
          <person-group person-group-type="author">
            <name>
              <surname>Song</surname>
              <given-names>J</given-names>
            </name>
            <name>
              <surname>Kim</surname>
              <given-names>K</given-names>
            </name>
            <name>
              <surname>Oh</surname>
              <given-names>J</given-names>
            </name>
            <name>
              <surname>Cho</surname>
              <given-names>S</given-names>
            </name>
          </person-group>
          <comment>MEMTO: memory-guided transformer for multivariate time series anomaly detection. In <italic>Proceedings of the 37th International Conference on Neural Information Processing Systems</italic>. Curran Associates Inc.; 2023. pp. 57947-63.</comment>
          <pub-id pub-id-type="doi">10.5555/3666122.3668647</pub-id>
        </nlm-citation>
      </ref>
      <ref id="B21">
        <label>21</label>
        <nlm-citation publication-type="book">
          <person-group person-group-type="author">
            <name>
              <surname>Wang</surname>
              <given-names>C</given-names>
            </name>
            <name>
              <surname>Zhuang</surname>
              <given-names>Z</given-names>
            </name>
            <name>
              <surname>Qi</surname>
              <given-names>Q</given-names>
            </name>
            <etal />
          </person-group>
          <comment>Drift doesn’t matter: dynamic decomposition with diffusion reconstruction for unstable multivariate time series anomaly detection. In <italic>Proceedings of the 37th International Conference on Neural Information Processing Systems</italic>. Curran Associates Inc.; 2023. pp. 10758-74.</comment>
          <pub-id pub-id-type="doi">10.5555/3666122.3666595</pub-id>
        </nlm-citation>
      </ref>
      <ref id="B22">
        <label>22</label>
        <nlm-citation publication-type="book">
          <person-group person-group-type="author">
            <name>
              <surname>Dai</surname>
              <given-names>Z</given-names>
            </name>
            <name>
              <surname>He</surname>
              <given-names>L</given-names>
            </name>
            <name>
              <surname>Yang</surname>
              <given-names>SH</given-names>
            </name>
            <name>
              <surname>Leeke</surname>
              <given-names>M</given-names>
            </name>
          </person-group>
          <comment>SARAD: spatial association-aware anomaly detection and diagnosis for multivariate time series. In <italic>Proceedings of the 38th International Conference on Neural Information Processing Systems</italic>. Curran Associates Inc.; 2024. pp. 48371-410.</comment>
          <pub-id pub-id-type="doi">10.5555/3737916.3739449</pub-id>
        </nlm-citation>
      </ref>
      <ref id="B23">
        <label>23</label>
        <nlm-citation publication-type="journal">
          <person-group person-group-type="author">
            <name>
              <surname>Huang</surname>
              <given-names>X</given-names>
            </name>
            <name>
              <surname>Chen</surname>
              <given-names>W</given-names>
            </name>
            <name>
              <surname>Hu</surname>
              <given-names>B</given-names>
            </name>
            <name>
              <surname>Mao</surname>
              <given-names>Z</given-names>
            </name>
          </person-group>
          <article-title>Graph mixture of experts and memory-augmented routers for multivariate time series anomaly detection</article-title>
          <source>Proc AAAI Conf Artif Intell</source>
          <year>2025</year>
          <volume>39</volume>
          <fpage>17476</fpage>
          <lpage>84</lpage>
          <pub-id pub-id-type="doi">10.1609/aaai.v39i16.33921</pub-id>
        </nlm-citation>
      </ref>
      <ref id="B24">
        <label>24</label>
        <nlm-citation publication-type="book">
          <person-group person-group-type="author">
            <name>
              <surname>Liu</surname>
              <given-names>S</given-names>
            </name>
            <name>
              <surname>Kimura</surname>
              <given-names>T</given-names>
            </name>
            <name>
              <surname>Liu</surname>
              <given-names>D</given-names>
            </name>
            <etal />
          </person-group>
          <comment>FOCAL: contrastive learning for multimodal time-series sensing signals in factorized orthogonal latent space. In <italic>Proceedings of the 37th International Conference on Neural Information Processing Systems</italic>. Curran Associates Inc.; 2023. pp. 47309-38.</comment>
          <pub-id pub-id-type="doi">10.5555/3666122.3668171</pub-id>
        </nlm-citation>
      </ref>
      <ref id="B25">
        <label>25</label>
        <nlm-citation publication-type="journal">
          <person-group person-group-type="author">
            <name>
              <surname>Ferrag</surname>
              <given-names>MA</given-names>
            </name>
            <name>
              <surname>Friha</surname>
              <given-names>O</given-names>
            </name>
            <name>
              <surname>Hamouda</surname>
              <given-names>D</given-names>
            </name>
            <name>
              <surname>Maglaras</surname>
              <given-names>L</given-names>
            </name>
            <name>
              <surname>Janicke</surname>
              <given-names>H</given-names>
            </name>
          </person-group>
          <article-title>Edge-IIoTset: a new comprehensive realistic cyber security dataset of IoT and IIoT applications for centralized and federated learning</article-title>
          <source>IEEE Access</source>
          <year>2022</year>
          <volume>10</volume>
          <fpage>40281</fpage>
          <lpage>306</lpage>
          <pub-id pub-id-type="doi">10.1109/access.2022.3165809</pub-id>
        </nlm-citation>
      </ref>
      <ref id="B26">
        <label>26</label>
        <nlm-citation publication-type="journal">
          <person-group person-group-type="author">
            <name>
              <surname>Neto</surname>
              <given-names>ECP</given-names>
            </name>
            <name>
              <surname>Dadkhah</surname>
              <given-names>S</given-names>
            </name>
            <name>
              <surname>Ferreira</surname>
              <given-names>R</given-names>
            </name>
            <name>
              <surname>Zohourian</surname>
              <given-names>A</given-names>
            </name>
            <name>
              <surname>Lu</surname>
              <given-names>R</given-names>
            </name>
            <name>
              <surname>Ghorbani</surname>
              <given-names>AA</given-names>
            </name>
          </person-group>
          <article-title>CICIoT2023: a real-time dataset and benchmark for large-scale attacks in IoT environment</article-title>
          <source>Sensors</source>
          <year>2023</year>
          <volume>23</volume>
          <fpage>5941</fpage>
          <pub-id pub-id-type="doi">10.3390/s23135941</pub-id>
          <pub-id pub-id-type="pmid">37447792</pub-id>
          <pub-id pub-id-type="pmcid">PMC10346235</pub-id>
        </nlm-citation>
      </ref>
      <ref id="B27">
        <label>27</label>
        <nlm-citation publication-type="journal">
          <person-group person-group-type="author">
            <name>
              <surname>Qathrady</surname>
              <given-names>MA</given-names>
            </name>
            <name>
              <surname>Ullah</surname>
              <given-names>S</given-names>
            </name>
            <name>
              <surname>Alshehri</surname>
              <given-names>MS</given-names>
            </name>
            <etal />
          </person-group>
          <article-title>SACNN‐IDS: a self‐attention convolutional neural network for intrusion detection in industrial Internet of Things</article-title>
          <source>CAAI Trans Intell Technol</source>
          <year>2024</year>
          <volume>9</volume>
          <fpage>1398</fpage>
          <lpage>411</lpage>
          <pub-id pub-id-type="doi">10.1049/cit2.12352</pub-id>
        </nlm-citation>
      </ref>
      <ref id="B28">
        <label>28</label>
        <nlm-citation publication-type="journal">
          <person-group person-group-type="author">
            <name>
              <surname>Li</surname>
              <given-names>X</given-names>
            </name>
            <name>
              <surname>Li</surname>
              <given-names>S</given-names>
            </name>
            <name>
              <surname>Zhang</surname>
              <given-names>G</given-names>
            </name>
            <name>
              <surname>Xie</surname>
              <given-names>Y</given-names>
            </name>
            <name>
              <surname>Wang</surname>
              <given-names>T</given-names>
            </name>
            <name>
              <surname>Chu</surname>
              <given-names>F</given-names>
            </name>
          </person-group>
          <article-title>A novel interpretable dynamic weighted domain adaptation network for cross-domain fault diagnosis of bearings under time-varying speeds</article-title>
          <source>Eng Appl Artif Intell</source>
          <year>2026</year>
          <volume>170</volume>
          <fpage>114173</fpage>
          <pub-id pub-id-type="doi">10.1016/j.engappai.2026.114173</pub-id>
        </nlm-citation>
      </ref>
      <ref id="B29">
        <label>29</label>
        <nlm-citation publication-type="journal">
          <person-group person-group-type="author">
            <name>
              <surname>Griffin</surname>
              <given-names>JM</given-names>
            </name>
            <name>
              <surname>Doberti</surname>
              <given-names>AJ</given-names>
            </name>
            <name>
              <surname>Hernández</surname>
              <given-names>V</given-names>
            </name>
            <name>
              <surname>Miranda</surname>
              <given-names>NA</given-names>
            </name>
            <name>
              <surname>Vélez</surname>
              <given-names>MA</given-names>
            </name>
          </person-group>
          <article-title>Multiple classification of the force and acceleration signals extracted during multiple machine processes: part 1 intelligent classification from an anomaly perspective</article-title>
          <source>Int J Adv Manuf Technol</source>
          <year>2017</year>
          <volume>93</volume>
          <fpage>811</fpage>
          <lpage>23</lpage>
          <pub-id pub-id-type="doi">10.1007/s00170-017-0320-3</pub-id>
        </nlm-citation>
      </ref>
      <ref id="B30">
        <label>30</label>
        <nlm-citation publication-type="journal">
          <person-group person-group-type="author">
            <name>
              <surname>Janssens</surname>
              <given-names>O</given-names>
            </name>
            <name>
              <surname>Slavkovikj</surname>
              <given-names>V</given-names>
            </name>
            <name>
              <surname>Vervisch</surname>
              <given-names>B</given-names>
            </name>
            <etal />
          </person-group>
          <article-title>Convolutional neural network based fault detection for rotating machinery</article-title>
          <source>J Sound Vib</source>
          <year>2016</year>
          <volume>377</volume>
          <fpage>331</fpage>
          <lpage>45</lpage>
          <pub-id pub-id-type="doi">10.1016/j.jsv.2016.05.027</pub-id>
        </nlm-citation>
      </ref>
      <ref id="B31">
        <label>31</label>
        <nlm-citation publication-type="book">
          <person-group person-group-type="author">
            <name>
              <surname>Amruthnath</surname>
              <given-names>N</given-names>
            </name>
            <name>
              <surname>Gupta</surname>
              <given-names>T</given-names>
            </name>
          </person-group>
          <comment>A research study on unsupervised machine learning algorithms for early fault detection in predictive maintenance. In <italic>2018 5th International Conference on Industrial Engineering and Applications (ICIEA)</italic>, Singapore. Apr 26-28, 2018. IEEE; 2018. pp. 355-61.</comment>
          <pub-id pub-id-type="doi">10.1109/IEA.2018.8387124</pub-id>
        </nlm-citation>
      </ref>
      <ref id="B32">
        <label>32</label>
        <nlm-citation publication-type="journal">
          <person-group person-group-type="author">
            <name>
              <surname>Diez-Olivan</surname>
              <given-names>A</given-names>
            </name>
            <name>
              <surname>Pagan</surname>
              <given-names>JA</given-names>
            </name>
            <name>
              <surname>Khoa</surname>
              <given-names>NLD</given-names>
            </name>
            <name>
              <surname>Sanz</surname>
              <given-names>R</given-names>
            </name>
            <name>
              <surname>Sierra</surname>
              <given-names>B</given-names>
            </name>
          </person-group>
          <article-title>Kernel-based support vector machines for automated health status assessment in monitoring sensor data</article-title>
          <source>Int J Adv Manuf Technol</source>
          <year>2018</year>
          <volume>95</volume>
          <fpage>327</fpage>
          <lpage>40</lpage>
          <pub-id pub-id-type="doi">10.1007/s00170-017-1204-2</pub-id>
        </nlm-citation>
      </ref>
      <ref id="B33">
        <label>33</label>
        <nlm-citation publication-type="book">
          <person-group person-group-type="author">
            <name>
              <surname>Joshi</surname>
              <given-names>SS</given-names>
            </name>
            <name>
              <surname>Phoha</surname>
              <given-names>VV</given-names>
            </name>
          </person-group>
          <comment>Investigating hidden Markov models capabilities in anomaly detection. In <italic>Proceedings of the 43rd Annual ACM Southeast Conference</italic>. Association for Computing Machinery; 2005. pp. 98-103.</comment>
          <pub-id pub-id-type="doi">10.1145/1167350.1167387</pub-id>
        </nlm-citation>
      </ref>
      <ref id="B34">
        <label>34</label>
        <nlm-citation publication-type="book">
          <person-group person-group-type="author">
            <name>
              <surname>Nguyen</surname>
              <given-names>MN</given-names>
            </name>
            <name>
              <surname>Vien</surname>
              <given-names>NA</given-names>
            </name>
          </person-group>
          <comment>Scalable and interpretable one-class SVMs with deep learning and random fourier features. In Berlingerio, M.; Bonchi, F.; Gärtner, T.; Hurley, N.; Ifrim, G.; eds. <italic>Machine learning and knowledge discovery in databases</italic>. Cham: Springer International Publishing; 2019. pp. 157-72.</comment>
          <pub-id pub-id-type="doi">10.1007/978-3-030-10925-7_10</pub-id>
        </nlm-citation>
      </ref>
      <ref id="B35">
        <label>35</label>
        <nlm-citation publication-type="book">
          <person-group person-group-type="author">
            <name>
              <surname>Xu</surname>
              <given-names>H</given-names>
            </name>
            <name>
              <surname>Chen</surname>
              <given-names>W</given-names>
            </name>
            <name>
              <surname>Zhao</surname>
              <given-names>N</given-names>
            </name>
            <etal />
          </person-group>
          <comment>Unsupervised anomaly detection via variational auto-encoder for seasonal KPIs in web applications. In <italic>Proceedings of the 2018 World Wide Web Conference</italic>. International World Wide Web Conferences Steering Committee; 2018. pp. 187-96.</comment>
          <pub-id pub-id-type="doi">10.1145/3178876.3185996</pub-id>
        </nlm-citation>
      </ref>
      <ref id="B36">
        <label>36</label>
        <nlm-citation publication-type="journal">
          <person-group person-group-type="author">
            <name>
              <surname>Lu</surname>
              <given-names>C</given-names>
            </name>
            <name>
              <surname>Wang</surname>
              <given-names>Z</given-names>
            </name>
            <name>
              <surname>Qin</surname>
              <given-names>W</given-names>
            </name>
            <name>
              <surname>Ma</surname>
              <given-names>J</given-names>
            </name>
          </person-group>
          <article-title>Fault diagnosis of rotary machinery components using a stacked denoising autoencoder-based health state identification</article-title>
          <source>Signal Process</source>
          <year>2017</year>
          <volume>130</volume>
          <fpage>377</fpage>
          <lpage>88</lpage>
          <pub-id pub-id-type="doi">10.1016/j.sigpro.2016.07.028</pub-id>
        </nlm-citation>
      </ref>
      <ref id="B37">
        <label>37</label>
        <nlm-citation publication-type="journal">
          <person-group person-group-type="author">
            <name>
              <surname>Zhang</surname>
              <given-names>C</given-names>
            </name>
            <name>
              <surname>Song</surname>
              <given-names>D</given-names>
            </name>
            <name>
              <surname>Chen</surname>
              <given-names>Y</given-names>
            </name>
            <etal />
          </person-group>
          <article-title>A deep neural network for unsupervised anomaly detection and diagnosis in multivariate time series data</article-title>
          <source>Proc AAAI Conf Artif Intell</source>
          <year>2019</year>
          <volume>33</volume>
          <fpage>1409</fpage>
          <lpage>16</lpage>
          <pub-id pub-id-type="doi">10.1609/aaai.v33i01.33011409</pub-id>
        </nlm-citation>
      </ref>
      <ref id="B38">
        <label>38</label>
        <nlm-citation publication-type="journal">
          <person-group person-group-type="author">
            <name>
              <surname>Yin</surname>
              <given-names>C</given-names>
            </name>
            <name>
              <surname>Zhang</surname>
              <given-names>S</given-names>
            </name>
            <name>
              <surname>Wang</surname>
              <given-names>J</given-names>
            </name>
            <name>
              <surname>Xiong</surname>
              <given-names>NN</given-names>
            </name>
          </person-group>
          <article-title>Anomaly detection based on convolutional recurrent autoencoder for IoT time series</article-title>
          <source>IEEE Trans Syst Man Cybern Syst</source>
          <year>2022</year>
          <volume>52</volume>
          <fpage>112</fpage>
          <lpage>22</lpage>
          <pub-id pub-id-type="doi">10.1109/tsmc.2020.2968516</pub-id>
        </nlm-citation>
      </ref>
      <ref id="B39">
        <label>39</label>
        <nlm-citation publication-type="journal">
          <person-group person-group-type="author">
            <name>
              <surname>Muneer</surname>
              <given-names>A</given-names>
            </name>
            <name>
              <surname>Mohd Taib</surname>
              <given-names>S</given-names>
            </name>
            <name>
              <surname>Mohamed Fati</surname>
              <given-names>S</given-names>
            </name>
            <name>
              <surname>Balogun</surname>
              <given-names>AO</given-names>
            </name>
            <name>
              <surname>Abdul Aziz</surname>
              <given-names>I</given-names>
            </name>
          </person-group>
          <article-title>A hybrid deep learning-based unsupervised anomaly detection in high dimensional data</article-title>
          <source>Comput Mater Contin</source>
          <year>2022</year>
          <volume>70</volume>
          <fpage>5363</fpage>
          <lpage>81</lpage>
          <pub-id pub-id-type="doi">10.32604/cmc.2022.021113</pub-id>
        </nlm-citation>
      </ref>
      <ref id="B40">
        <label>40</label>
        <nlm-citation publication-type="journal">
          <person-group person-group-type="author">
            <name>
              <surname>Zheng</surname>
              <given-names>Y</given-names>
            </name>
          </person-group>
          <article-title>Methodologies for cross-domain data fusion: an overview</article-title>
          <source>IEEE Trans Big Data</source>
          <year>2015</year>
          <volume>1</volume>
          <fpage>16</fpage>
          <lpage>34</lpage>
          <pub-id pub-id-type="doi">10.1109/tbdata.2015.2465959</pub-id>
        </nlm-citation>
      </ref>
      <ref id="B41">
        <label>41</label>
        <nlm-citation publication-type="journal">
          <person-group person-group-type="author">
            <name>
              <surname>Xiao</surname>
              <given-names>X</given-names>
            </name>
            <name>
              <surname>Zheng</surname>
              <given-names>Y</given-names>
            </name>
            <name>
              <surname>Luo</surname>
              <given-names>Q</given-names>
            </name>
            <name>
              <surname>Xie</surname>
              <given-names>X</given-names>
            </name>
          </person-group>
          <article-title>Inferring social ties between users with human location history</article-title>
          <source>J Ambient Intell Human Comput</source>
          <year>2014</year>
          <volume>5</volume>
          <fpage>3</fpage>
          <lpage>19</lpage>
          <pub-id pub-id-type="doi">10.1007/s12652-012-0117-z</pub-id>
        </nlm-citation>
      </ref>
      <ref id="B42">
        <label>42</label>
        <nlm-citation publication-type="book">
          <person-group person-group-type="author">
            <name>
              <surname>Ngiam</surname>
              <given-names>J</given-names>
            </name>
            <name>
              <surname>Khosla</surname>
              <given-names>A</given-names>
            </name>
            <name>
              <surname>Kim</surname>
              <given-names>M</given-names>
            </name>
            <name>
              <surname>Nam</surname>
              <given-names>J</given-names>
            </name>
            <name>
              <surname>Lee</surname>
              <given-names>H</given-names>
            </name>
            <name>
              <surname>Ng</surname>
              <given-names>AY</given-names>
            </name>
          </person-group>
          <comment>Multimodal deep learning. In <italic>Proceedings of the 28th International Conference on Machine Learning</italic>, Bellevue, USA. 2011. pp. 689-96. <uri xlink:href="https://ai.stanford.edu/~jngiam/papers/NgiamKhoslaKimNamLeeNg2011.pdf">https://ai.stanford.edu/~jngiam/papers/NgiamKhoslaKimNamLeeNg2011.pdf</uri>. (accessed 2026-09-21)</comment>
        </nlm-citation>
      </ref>
      <ref id="B43">
        <label>43</label>
        <nlm-citation publication-type="book">
          <person-group person-group-type="author">
            <name>
              <surname>Nagrani</surname>
              <given-names>A</given-names>
            </name>
            <name>
              <surname>Yang</surname>
              <given-names>S</given-names>
            </name>
            <name>
              <surname>Arnab</surname>
              <given-names>A</given-names>
            </name>
            <name>
              <surname>Jansen</surname>
              <given-names>A</given-names>
            </name>
            <name>
              <surname>Schmid</surname>
              <given-names>C</given-names>
            </name>
            <name>
              <surname>Sun</surname>
              <given-names>C</given-names>
            </name>
          </person-group>
          <comment>Attention bottlenecks for multimodal fusion. <italic>arXiv</italic> <bold>2021</bold>, arXiv:2107.00135. Available online: <uri xlink:href="https://doi.org/10.48550/arXiv.2107.00135">https://doi.org/10.48550/arXiv.2107.00135</uri>. (accessed 2026-09-21)</comment>
        </nlm-citation>
      </ref>
      <ref id="B44">
        <label>44</label>
        <nlm-citation publication-type="book">
          <person-group person-group-type="author">
            <name>
              <surname>Yasaei</surname>
              <given-names>R</given-names>
            </name>
            <name>
              <surname>Moghaddas</surname>
              <given-names>Y</given-names>
            </name>
            <name>
              <surname>Al Faruque</surname>
              <given-names>MA</given-names>
            </name>
          </person-group>
          <comment>IoT-GRAF: IoT graph learning-based anomaly and intrusion detection through multi-modal data fusion. In <italic>2024 Design, Automation &amp; Test in Europe Conference &amp; Exhibition (DATE)</italic>, Valencia, Spain. Mar 25-27, 2024. IEEE; 2024. p. 1-6.</comment>
          <pub-id pub-id-type="doi">10.23919/DATE58400.2024.10546572</pub-id>
        </nlm-citation>
      </ref>
      <ref id="B45">
        <label>45</label>
        <nlm-citation publication-type="journal">
          <person-group person-group-type="author">
            <name>
              <surname>Zhan</surname>
              <given-names>D</given-names>
            </name>
            <name>
              <surname>Zhang</surname>
              <given-names>W</given-names>
            </name>
            <name>
              <surname>Ye</surname>
              <given-names>L</given-names>
            </name>
            <name>
              <surname>Yu</surname>
              <given-names>X</given-names>
            </name>
            <name>
              <surname>Zhang</surname>
              <given-names>H</given-names>
            </name>
            <name>
              <surname>He</surname>
              <given-names>Z</given-names>
            </name>
          </person-group>
          <article-title>Anomaly detection in industrial control systems based on cross-domain representation learning</article-title>
          <source>IEEE Trans Dependable Secure Comput</source>
          <year>2025</year>
          <volume>22</volume>
          <fpage>2505</fpage>
          <lpage>18</lpage>
          <pub-id pub-id-type="doi">10.1109/tdsc.2024.3520155</pub-id>
        </nlm-citation>
      </ref>
      <ref id="B46">
        <label>46</label>
        <nlm-citation publication-type="book">
          <person-group person-group-type="author">
            <name>
              <surname>Sun</surname>
              <given-names>Y</given-names>
            </name>
            <name>
              <surname>Qi</surname>
              <given-names>J</given-names>
            </name>
            <name>
              <surname>Zheng</surname>
              <given-names>Y</given-names>
            </name>
            <name>
              <surname>Zhang</surname>
              <given-names>R</given-names>
            </name>
          </person-group>
          <comment>K-nearest neighbor temporal aggregate queries. In <italic>Proceedings of the 18th International Conference on Extending Database Technology</italic>. 2015. pp. 493-504. <uri xlink:href="https://www.microsoft.com/en-us/research/publication/k-nearest-neighbor-temporal-aggregate-queries/">https://www.microsoft.com/en-us/research/publication/k-nearest-neighbor-temporal-aggregate-queries/</uri>. (accessed 2026-09-21)</comment>
        </nlm-citation>
      </ref>
      <ref id="B47">
        <label>47</label>
        <nlm-citation publication-type="journal">
          <person-group person-group-type="author">
            <name>
              <surname>Canonico</surname>
              <given-names>R</given-names>
            </name>
            <name>
              <surname>Esposito</surname>
              <given-names>G</given-names>
            </name>
            <name>
              <surname>Navarro</surname>
              <given-names>A</given-names>
            </name>
            <name>
              <surname>Romano</surname>
              <given-names>SP</given-names>
            </name>
            <name>
              <surname>Sperlì</surname>
              <given-names>G</given-names>
            </name>
            <name>
              <surname>Vignali</surname>
              <given-names>A</given-names>
            </name>
          </person-group>
          <article-title>Empowered cyber–physical systems security using both network and physical data</article-title>
          <source>Comput Secur</source>
          <year>2025</year>
          <volume>152</volume>
          <fpage>104382</fpage>
          <pub-id pub-id-type="doi">10.1016/j.cose.2025.104382</pub-id>
        </nlm-citation>
      </ref>
      <ref id="B48">
        <label>48</label>
        <nlm-citation publication-type="journal">
          <person-group person-group-type="author">
            <name>
              <surname>Pinto</surname>
              <given-names>A</given-names>
            </name>
            <name>
              <surname>Herrera</surname>
              <given-names>L</given-names>
            </name>
            <name>
              <surname>Donoso</surname>
              <given-names>Y</given-names>
            </name>
            <name>
              <surname>Gutierrez</surname>
              <given-names>JA</given-names>
            </name>
          </person-group>
          <article-title>Cyber-physical anomaly detection a deep adversarial fusion of sensor and network data</article-title>
          <source>Discov Comput</source>
          <year>2026</year>
          <volume>29</volume>
          <fpage>10064</fpage>
          <pub-id pub-id-type="doi">10.1007/s10791-026-10064-6</pub-id>
        </nlm-citation>
      </ref>
      <ref id="B49">
        <label>49</label>
        <nlm-citation publication-type="journal">
          <person-group person-group-type="author">
            <name>
              <surname>Shafin</surname>
              <given-names>SS</given-names>
            </name>
            <name>
              <surname>Karmakar</surname>
              <given-names>G</given-names>
            </name>
            <name>
              <surname>Mareels</surname>
              <given-names>I</given-names>
            </name>
            <name>
              <surname>Balasubramanian</surname>
              <given-names>V</given-names>
            </name>
            <name>
              <surname>Kolluri</surname>
              <given-names>RR</given-names>
            </name>
          </person-group>
          <article-title>Sensor self-declaration of numeric data reliability in Internet of Things</article-title>
          <source>IEEE Trans Reliab</source>
          <year>2025</year>
          <volume>74</volume>
          <fpage>2751</fpage>
          <lpage>65</lpage>
          <pub-id pub-id-type="doi">10.1109/tr.2024.3416967</pub-id>
        </nlm-citation>
      </ref>
      <ref id="B50">
        <label>50</label>
        <nlm-citation publication-type="web">
          <person-group person-group-type="author">
            <name>
              <surname>Veličković</surname>
              <given-names>P</given-names>
            </name>
            <name>
              <surname>Cucurull</surname>
              <given-names>G</given-names>
            </name>
            <name>
              <surname>Casanova</surname>
              <given-names>A</given-names>
            </name>
            <name>
              <surname>Romero</surname>
              <given-names>A</given-names>
            </name>
            <name>
              <surname>Liò</surname>
              <given-names>P</given-names>
            </name>
            <name>
              <surname>Bengio</surname>
              <given-names>Y</given-names>
            </name>
          </person-group>
          <comment>Graph attention networks. <italic>arXiv</italic> <bold>2017</bold>, arXiv:1710.10903. Available online: <uri xlink:href="https://doi.org/10.48550/arXiv.1710.10903">https://doi.org/10.48550/arXiv.1710.10903</uri>. (accessed 2026-09-21)</comment>
        </nlm-citation>
      </ref>
      <ref id="B51">
        <label>51</label>
        <nlm-citation publication-type="journal">
          <person-group person-group-type="author">
            <name>
              <surname>Hochreiter</surname>
              <given-names>S</given-names>
            </name>
            <name>
              <surname>Schmidhuber</surname>
              <given-names>J</given-names>
            </name>
          </person-group>
          <article-title>Long short-term memory</article-title>
          <source>Neural Comput</source>
          <year>1997</year>
          <volume>9</volume>
          <fpage>1735</fpage>
          <lpage>80</lpage>
          <pub-id pub-id-type="doi">10.1162/neco.1997.9.8.1735</pub-id>
          <pub-id pub-id-type="pmid">9377276</pub-id>
        </nlm-citation>
      </ref>
      <ref id="B52">
        <label>52</label>
        <nlm-citation publication-type="journal">
          <person-group person-group-type="author">
            <name>
              <surname>Ngo</surname>
              <given-names>G</given-names>
            </name>
            <name>
              <surname>Beard</surname>
              <given-names>R</given-names>
            </name>
            <name>
              <surname>Chandra</surname>
              <given-names>R</given-names>
            </name>
          </person-group>
          <article-title>Evolutionary bagging for ensemble learning</article-title>
          <source>Neurocomputing</source>
          <year>2022</year>
          <volume>510</volume>
          <fpage>1</fpage>
          <lpage>14</lpage>
          <pub-id pub-id-type="doi">10.1016/j.neucom.2022.08.055</pub-id>
        </nlm-citation>
      </ref>
      <ref id="B53">
        <label>53</label>
        <nlm-citation publication-type="journal">
          <person-group person-group-type="author">
            <name>
              <surname>Kadkhodaei</surname>
              <given-names>HR</given-names>
            </name>
            <name>
              <surname>Moghadam</surname>
              <given-names>AME</given-names>
            </name>
            <name>
              <surname>Dehghan</surname>
              <given-names>M</given-names>
            </name>
          </person-group>
          <article-title>HBoost: a heterogeneous ensemble classifier based on the boosting method and entropy measurement</article-title>
          <source>Expert Syst Appl</source>
          <year>2020</year>
          <volume>157</volume>
          <fpage>113482</fpage>
          <pub-id pub-id-type="doi">10.1016/j.eswa.2020.113482</pub-id>
        </nlm-citation>
      </ref>
      <ref id="B54">
        <label>54</label>
        <nlm-citation publication-type="journal">
          <person-group person-group-type="author">
            <name>
              <surname>Zhang</surname>
              <given-names>H</given-names>
            </name>
            <name>
              <surname>Li</surname>
              <given-names>J</given-names>
            </name>
            <name>
              <surname>Liu</surname>
              <given-names>X</given-names>
            </name>
            <name>
              <surname>Dong</surname>
              <given-names>C</given-names>
            </name>
          </person-group>
          <article-title>Multi-dimensional feature fusion and stacking ensemble mechanism for network intrusion detection</article-title>
          <source>Future Gener Comput Syst</source>
          <year>2021</year>
          <volume>122</volume>
          <fpage>130</fpage>
          <lpage>43</lpage>
          <pub-id pub-id-type="doi">10.1016/j.future.2021.03.024</pub-id>
        </nlm-citation>
      </ref>
      <ref id="B55">
        <label>55</label>
        <nlm-citation publication-type="journal">
          <person-group person-group-type="author">
            <name>
              <surname>Kuncheva</surname>
              <given-names>LI</given-names>
            </name>
            <name>
              <surname>Whitaker</surname>
              <given-names>CJ</given-names>
            </name>
          </person-group>
          <article-title>Measures of diversity in classifier ensembles and their relationship with the ensemble accuracy</article-title>
          <source>Mach Learn</source>
          <year>2003</year>
          <volume>51</volume>
          <fpage>181</fpage>
          <lpage>207</lpage>
          <pub-id pub-id-type="doi">10.1023/a:1022859003006</pub-id>
        </nlm-citation>
      </ref>
      <ref id="B56">
        <label>56</label>
        <nlm-citation publication-type="journal">
          <person-group person-group-type="author">
            <name>
              <surname>Booij</surname>
              <given-names>TM</given-names>
            </name>
            <name>
              <surname>Chiscop</surname>
              <given-names>I</given-names>
            </name>
            <name>
              <surname>Meeuwissen</surname>
              <given-names>E</given-names>
            </name>
            <name>
              <surname>Moustafa</surname>
              <given-names>N</given-names>
            </name>
            <name>
              <surname>Hartog</surname>
              <given-names>FTHD</given-names>
            </name>
          </person-group>
          <article-title>ToN_IoT: the role of heterogeneity and the need for standardization of features and attack types in IoT network intrusion data sets</article-title>
          <source>IEEE Internet Things J</source>
          <year>2022</year>
          <volume>9</volume>
          <fpage>485</fpage>
          <lpage>96</lpage>
          <pub-id pub-id-type="doi">10.1109/jiot.2021.3085194</pub-id>
        </nlm-citation>
      </ref>
      <ref id="B57">
        <label>57</label>
        <nlm-citation publication-type="journal">
          <person-group person-group-type="author">
            <name>
              <surname>Sokolova</surname>
              <given-names>M</given-names>
            </name>
            <name>
              <surname>Lapalme</surname>
              <given-names>G</given-names>
            </name>
          </person-group>
          <article-title>A systematic analysis of performance measures for classification tasks</article-title>
          <source>Inf Process Manag</source>
          <year>2009</year>
          <volume>45</volume>
          <fpage>427</fpage>
          <lpage>37</lpage>
          <pub-id pub-id-type="doi">10.1016/j.ipm.2009.03.002</pub-id>
        </nlm-citation>
      </ref>
      <ref id="B58">
        <label>58</label>
        <nlm-citation publication-type="journal">
          <person-group person-group-type="author">
            <name>
              <surname>Li</surname>
              <given-names>R</given-names>
            </name>
            <name>
              <surname>Li</surname>
              <given-names>Y</given-names>
            </name>
            <name>
              <surname>He</surname>
              <given-names>W</given-names>
            </name>
            <name>
              <surname>Chen</surname>
              <given-names>L</given-names>
            </name>
            <name>
              <surname>Luo</surname>
              <given-names>J</given-names>
            </name>
          </person-group>
          <article-title>Multi-layer reconstruction errors autoencoding and density estimate for network anomaly detection</article-title>
          <source>Comput Model Eng Sci</source>
          <year>2021</year>
          <volume>128</volume>
          <fpage>381</fpage>
          <lpage>97</lpage>
          <pub-id pub-id-type="doi">10.32604/cmes.2021.016264</pub-id>
        </nlm-citation>
      </ref>
      <ref id="B59">
        <label>59</label>
        <nlm-citation publication-type="book">
          <person-group person-group-type="author">
            <name>
              <surname>Wu</surname>
              <given-names>Z</given-names>
            </name>
            <name>
              <surname>Pan</surname>
              <given-names>S</given-names>
            </name>
            <name>
              <surname>Long</surname>
              <given-names>G</given-names>
            </name>
            <name>
              <surname>Jiang</surname>
              <given-names>J</given-names>
            </name>
            <name>
              <surname>Chang</surname>
              <given-names>X</given-names>
            </name>
            <name>
              <surname>Zhang</surname>
              <given-names>C</given-names>
            </name>
          </person-group>
          <comment>Connecting the dots: multivariate time series forecasting with graph neural networks. In <italic>Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery &amp; Data Mining</italic>. Association for Computing Machinery; 2020. pp. 753-63.</comment>
          <pub-id pub-id-type="doi">10.1145/3394486.3403118</pub-id>
        </nlm-citation>
      </ref>
    </ref-list>
  </back>
</article>