<!DOCTYPE article
PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD with MathML3 v1.3 20210610//EN" "JATS-archivearticle1-3-mathml3.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="1.3" xml:lang="en" article-type="research-article"><?properties manuscript?><processing-meta base-tagset="archiving" mathml-version="3.0" table-model="xhtml" tagset-family="jats"><restricted-by>pmc</restricted-by></processing-meta><front><journal-meta><journal-id journal-id-type="nlm-journal-id">101766564</journal-id><journal-id journal-id-type="pubmed-jr-id">49543</journal-id><journal-id journal-id-type="nlm-ta">Mathematics (Basel)</journal-id><journal-id journal-id-type="iso-abbrev">Mathematics (Basel)</journal-id><journal-title-group><journal-title>Mathematics (Basel, Switzerland)</journal-title></journal-title-group><issn pub-type="epub">2227-7390</issn></journal-meta><article-meta><article-id pub-id-type="pmid">39959340</article-id><article-id pub-id-type="pmc">11827645</article-id><article-id pub-id-type="doi">10.3390/math13020180</article-id><article-id pub-id-type="manuscript">HHSPA2053985</article-id><article-categories><subj-group subj-group-type="heading"><subject>Article</subject></subj-group></article-categories><title-group><article-title>Estimating the Relative Risks of Spatial Clusters Using a Predictor-Corrector Method</article-title></title-group><contrib-group><contrib contrib-type="author"><name><surname>Bani-Yaghoub</surname><given-names>Majid</given-names></name><xref rid="A1" ref-type="aff">1</xref></contrib><contrib contrib-type="author"><name><surname>Rekab</surname><given-names>Kamel</given-names></name><xref rid="A1" ref-type="aff">1</xref><xref rid="CR1" ref-type="corresp">*</xref></contrib><contrib contrib-type="author"><name><surname>Pluta</surname><given-names>Julia</given-names></name><xref rid="A1" ref-type="aff">1</xref></contrib><contrib contrib-type="author"><name><surname>Tabharit</surname><given-names>Said</given-names></name><xref rid="A1" ref-type="aff">1</xref></contrib></contrib-group><aff id="A1"><label>1</label>Division of Computing, Analytics and Mathematics, School of Science and Engineering, University of Missouri-Kansas City, Kansas City, MO 64110, USA.</aff><author-notes><corresp id="CR1"><label>*</label><bold>Corresponding Author:</bold>
<email>rekabk@umkc.edu</email></corresp></author-notes><pub-date pub-type="nihms-submitted"><day>6</day><month>2</month><year>2025</year></pub-date><pub-date pub-type="ppub"><year>2025</year></pub-date><pub-date pub-type="pmc-release"><day>14</day><month>2</month><year>2025</year></pub-date><volume>13</volume><issue>2</issue><elocation-id>10.3390/math13020180</elocation-id><abstract id="ABS1"><p id="P1">Spatial, temporal, and space-time scan statistics can be used for geographical surveillance, identifying temporal and spatial patterns, and detecting outliers. While statistical cluster analysis is a valuable tool for identifying patterns, optimizing resource allocation, and supporting decision-making, accurately predicting future spatial clusters remains a significant challenge.</p><p id="P2">Given the known relative risks of spatial clusters over the past <inline-formula><mml:math id="M1" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula> time intervals, the main objective of the present study is to predict the relative risks for the subsequent interval, <inline-formula><mml:math id="M2" display="inline"><mml:mi>k</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula>. Building on our prior research, we propose a predictive Markov chain model with an embedded corrector component. This corrector utilizes either multiple linear regression or exponential smoothing method, selecting the one that minimizes the relative distance between observed and predicted values in the <inline-formula><mml:math id="M3" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula>-th interval.</p><p id="P3">To test the proposed method, we first calculated the relative risks of statistically significant spatial clusters of COVID-19 mortality in the U.S. over seven time intervals from May 2020 to March 2023. Then, for each time interval, we selected the top 25 clusters with the highest relative risks and iteratively predicted the relative risks of clusters from intervals three to seven. The predictive accuracies ranged from moderate to high, indicating the potential applicability of this method for predictive disease analytics and future pandemic preparedness.</p></abstract><kwd-group><kwd>COVID-19</kwd><kwd>Mortality</kwd><kwd>Relative Risk</kwd><kwd>Cluster Analysis</kwd><kwd>Exponential Smoothing</kwd><kwd>Regression</kwd></kwd-group></article-meta></front><body><sec id="S1"><label>1</label><title>Introduction</title><p id="P4">The emergence of COVID-19 highlighted the critical need for precise spatial and temporal surveillance of infectious diseases, especially in densely populated or vulnerable regions. The spatial scan statistic has proven essential for identifying disease clusters in public health data and helping authorities detect, respond to, and mitigate outbreaks. It was first proposed by Kulldorff [<xref rid="R31" ref-type="bibr">31</xref>] and subsequently developed in various studies [<xref rid="R33" ref-type="bibr">33</xref>, <xref rid="R26" ref-type="bibr">26</xref>]. This method, using circular and, later, elliptic scanning windows [<xref rid="R32" ref-type="bibr">32</xref>], efficiently identifies clusters without prior information on the disease&#x02019;s distribution, providing a robust framework for various applications from infectious diseases to cancer epidemiology [<xref rid="R8" ref-type="bibr">8</xref>, <xref rid="R5" ref-type="bibr">5</xref>]. Kulldorff&#x02019;s elliptic spatial scan statistics address the limitations of circular windows, particularly in urban environments where population density and movement patterns are uneven [<xref rid="R32" ref-type="bibr">32</xref>].</p><p id="P5">Additional advancements include the weighted normal spatial scan statistic, which accommodates heterogeneous population data, enabling more precise detection in demographically diverse regions [<xref rid="R26" ref-type="bibr">26</xref>, <xref rid="R45" ref-type="bibr">45</xref>]. Studies on COVID-19, for instance, have applied these techniques to uncover spatial variations and emerging hot spots within urban settings, providing critical insights into the pandemic&#x02019;s dynamics [<xref rid="R2" ref-type="bibr">2</xref>, <xref rid="R3" ref-type="bibr">3</xref>, <xref rid="R4" ref-type="bibr">4</xref>, <xref rid="R6" ref-type="bibr">6</xref>, <xref rid="R34" ref-type="bibr">34</xref>]. Another study, [<xref rid="R16" ref-type="bibr">16</xref>] explored how these spatial methods facilitated response coordination across countries, reflecting the adaptability of spatial scan statistics in different geographic contexts and disease typologies. Additionally, the multivariate scan statistic extends traditional models to analyze multiple data types simultaneously, thereby improving surveillance accuracy in multifactorial diseases [<xref rid="R33" ref-type="bibr">33</xref>]. The utility of Monte Carlo hypothesis testing, as demonstrated in [<xref rid="R14" ref-type="bibr">14</xref>] and further refined for power analyses by [<xref rid="R41" ref-type="bibr">41</xref>], is integral to validating spatial clusters, adding another layer of reliability to these models.</p><p id="P6">These spatial methodologies have also shown efficacy in tracking non-COVID infectious diseases. For example, the space-time models applied airborne diseases like H7N9 influenza [<xref rid="R43" ref-type="bibr">43</xref>, <xref rid="R9" ref-type="bibr">9</xref>, <xref rid="R50" ref-type="bibr">50</xref>] demonstrated their utility in understanding disease spread and preparing targeted interventions. Furthermore, Souris et al. [<xref rid="R42" ref-type="bibr">42</xref>] illustrated the application of these tools in poultry farms, identifying avian influenza re-emergence risks across human and animal health sectors (see also [<xref rid="R12" ref-type="bibr">12</xref>, <xref rid="R21" ref-type="bibr">21</xref>]).</p><p id="P7">In healthcare-associated infection settings, these models have helped detect and control outbreaks, as shown in studies on methicillin-resistant <italic toggle="yes">Methicillin-resistant Staphylococcus aureus</italic> [<xref rid="R27" ref-type="bibr">27</xref>] and carbapenemase-producing <italic toggle="yes">Klebsiella pneumoniae</italic> [<xref rid="R1" ref-type="bibr">1</xref>]. By incorporating real-time data, these systems provide timely alerts, allowing healthcare providers to implement interventions more rapidly [<xref rid="R36" ref-type="bibr">36</xref>]. The syndromic surveillance adaptations, which emerged from the anthrax preparedness studies of [<xref rid="R30" ref-type="bibr">30</xref>] and subsequent research by [<xref rid="R25" ref-type="bibr">25</xref>], underscore the proactive role these systems play in modern epidemiological surveillance.</p><p id="P8">The spatial scan statistic and its derivatives have been instrumental across public health, responding to infectious diseases, environmental exposures, and even syndromic surveillance scenarios. These tools can become even more powerful when applied for prediction, as predictive applications extend the utility of spatial cluster analysis beyond retrospective assessments. The main question is whether predictive methods can be merged with these tools to identify regions at heightened risk of future disease spread, anticipate healthcare resource requirements, and guide targeted interventions in real time, enabling decision-makers to allocate resources more effectively. In this paper, we aim to test the applicability of a predictor-corrector method to the relative risks resulting from the spatial cluster analysis of <inline-formula><mml:math id="M9" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula> time intervals for predicting the relative risks in the <inline-formula><mml:math id="M10" display="inline"><mml:mi>k</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:mi>t</mml:mi><mml:mi>h</mml:mi></mml:math></inline-formula> time interval. In particular, assuming that the relative risks of spatial clusters are known for the past <inline-formula><mml:math id="M11" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula> time intervals, we are interested in predicting the relative risks of spatial clusters in the time interval <inline-formula><mml:math id="M12" display="inline"><mml:mi>k</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula>. Based on our previous work [<xref rid="R46" ref-type="bibr">46</xref>], we propose a Markov chain model with an embedded corrector. The corrector is either a multiple linear regression or an exponential smoothing, whichever with shorter relative distance between the observed and predicted values in interval <inline-formula><mml:math id="M13" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula>.</p><p id="P9">Predicting relative risks in spatial clusters could serve as an early warning system for areas at increased mortality risk, which is especially valuable for addressing resource allocation and intervention planning in critical periods. We use the U.S. COVID-19 mortality data [<xref rid="R44" ref-type="bibr">44</xref>] to evaluate this approach and assess whether this method can accurately forecast relative risks in clusters of significant mortality.</p><p id="P10">The remainder of this paper is organized as follows. <xref rid="S2" ref-type="sec">Section 2</xref> provides details on the data, the SaTScan method [<xref rid="R15" ref-type="bibr">15</xref>], and the results of our spatial cluster analysis. <xref rid="S7" ref-type="sec">Section 3</xref> introduces the predictor-corrector method, while <xref rid="S13" ref-type="sec">Section 4</xref> applies this method to the U.S. COVID-19 mortality data over six time intervals, with the goal of predicting relative risks of the seventh time interval. Finally, <xref rid="S16" ref-type="sec">Section 5</xref> discusses the results and limitations of the present work, as well as potential directions for future research.</p></sec><sec id="S2"><label>2</label><title>Retrospective Analysis of Spatial Clusters</title><sec id="S3"><label>2.1</label><title>Overview</title><p id="P11">We employed a retrospective spatial analysis of COVID-19 clusters across the United States by examining how case distributions correlate with population density. Utilizing a Poisson-based spatial scanning model, we identified clusters by employing a cylindrical scanning window to capture areas of heightened risk relative to surrounding regions. The analysis applies likelihood ratio testing to detect significant spatial clusters, with primary and secondary clusters reported based on statistical significance. Monte Carlo simulations were used to validate the results, offering insights into the geographic spread and demographic impact of COVID-19 across different communities in the U.S.</p></sec><sec id="S4"><label>2.2</label><title>US COVID-19 Mortality Data</title><p id="P12">We obtained COVID-19 mortality data for the United States from <italic toggle="yes">The New York Times</italic>, which provides daily records of COVID-19 deaths from February 2020 onward. This data, publicly available at [<xref rid="R44" ref-type="bibr">44</xref>].</p><p id="P13">The dataset reports daily COVID-19 deaths across all U.S. states and territories. In many states, health departments reconcile these records with death certificates, omitting deaths where COVID-19 was not the cause of death. Following this protocol, non-COVID-19 deaths were removed from the dataset when there was a confirmation that COVID-19 was not the cause, specifically in cases of homicide, suicide, car accidents, or drug overdoses. Details of these adjustments can be found in [<xref rid="R44" ref-type="bibr">44</xref>].</p><p id="P14">To examine significant variations in COVID-19 mortality trends, we focused on data spanning from May 24, 2020, to March 12, 2023. We divided this period into seven time intervals to analyze trends within consistent segments of time: May 24, 2020 &#x02212; September 13, 2020; September 13, 2020 &#x02212; March 14, 2021; March 14, 2021 &#x02212; June 13, 2021; June 13, 2021 &#x02212; October 31, 2021; October 31, 2021 &#x02212; March 13, 2022; March 13, 2022 &#x02212; October 16, 2022; and October 16, 2022 &#x02212; March 12, 2023. Each interval represents a critical phase in the pandemic, allowing us to study temporal fluctuations in COVID-19-related mortality rates (see <xref rid="T1" ref-type="table">Table 1</xref> for a summary). Specifically, these intervals were selected based on the emergence of new <italic toggle="yes">SARS-CoV-2</italic> variants, which played a pivotal role in shaping the trajectory of the pandemic. This approach aimed to capture virological and epidemiological shifts corresponding to significant phases of the outbreak. As shown in <xref rid="F2" ref-type="fig">Figure 2 (a)</xref>, each interval, except the third, includes at least one peak in COVID-19 mortality within the United States. The third interval corresponds to the emergence of the Alpha variant, which did not result in a mortality peak during this period but contributed to a significant peak in the subsequent interval. While this method highlights variant-driven dynamics, we acknowledge that alternative approaches, such as structural change detection methods [<xref rid="R22" ref-type="bibr">22</xref>], could potentially offer a data-driven determination of breakpoints. Incorporating such methods in future studies could further refine the interval selection process and enhance the robustness of predictive modeling.</p></sec><sec id="S5"><label>2.3</label><title>Spatial Statistical Scan</title><p id="P15">To analyze the geographical distribution of COVID-19 cases adjusted for population demographics, we applied a Poisson model. Our spatial analysis utilized a cylindrical scanning approach, with a moving circular base representing the spatial dimension, and set a maximum cluster size of 25% of the population at risk to prevent excessively large clusters. This approach created numerous overlapping windows of varying sizes, allowing full coverage of the study area. Each window was considered as a potential cluster. Under the null hypothesis <inline-formula><mml:math id="M14" display="inline"><mml:msub><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>, the COVID-19 risk is uniform inside and outside the scanning window, while the alternative hypothesis <inline-formula><mml:math id="M15" display="inline"><mml:msub><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mrow><mml:mi>a</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> suggests a higher risk within the window. Each cylinder expanded up to a preset upper limit on cluster size. To identify clusters, we applied a likelihood ratio test.</p><p id="P16">The likelihood ratio <inline-formula><mml:math id="M16" display="inline"><mml:mfrac><mml:mrow><mml:mi>L</mml:mi><mml:mo>(</mml:mo><mml:mi>C</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:math></inline-formula> is calculated by
<disp-formula id="FD1">
<label>(1)</label>
<mml:math id="M17" display="block"><mml:mfrac><mml:mrow><mml:mi>L</mml:mi><mml:mo>(</mml:mo><mml:mi>C</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msup><mml:mrow><mml:mfenced separators="|"><mml:mrow><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>&#x003bc;</mml:mi><mml:mo>(</mml:mo><mml:mi>c</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mfrac></mml:mrow></mml:mfenced></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msup><mml:msup><mml:mrow><mml:mfenced separators="|"><mml:mrow><mml:mfrac><mml:mrow><mml:mi>N</mml:mi><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>N</mml:mi><mml:mo>-</mml:mo><mml:mi>&#x003bc;</mml:mi><mml:mo>(</mml:mo><mml:mi>c</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mfrac></mml:mrow></mml:mfenced></mml:mrow><mml:mrow><mml:mi>N</mml:mi><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mfenced separators="|"><mml:mrow><mml:mfrac><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003bc;</mml:mi><mml:mo>(</mml:mo><mml:mi>T</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mfrac></mml:mrow></mml:mfenced></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:mfrac><mml:mo>,</mml:mo></mml:math>
</disp-formula>
where <inline-formula><mml:math id="M18" display="inline"><mml:mi>L</mml:mi><mml:mo>(</mml:mo><mml:mi>C</mml:mi><mml:mo>)</mml:mo></mml:math></inline-formula> is the maximum likelihood for the cylinder based on the observed cases <inline-formula><mml:math id="M19" display="inline"><mml:msub><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and expected cases <inline-formula><mml:math id="M20" display="inline"><mml:mi>&#x003bc;</mml:mi><mml:mo>(</mml:mo><mml:mi>c</mml:mi><mml:mo>)</mml:mo></mml:math></inline-formula> within the window. Here, <inline-formula><mml:math id="M21" display="inline"><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> is the likelihood under the null hypothesis, <inline-formula><mml:math id="M22" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula> is the total cases observed in the U.S. for each time interval, and <inline-formula><mml:math id="M23" display="inline"><mml:mi>&#x003bc;</mml:mi><mml:mo>(</mml:mo><mml:mi>T</mml:mi><mml:mo>)</mml:mo></mml:math></inline-formula> represents the total expected cases. A likelihood ratio greater than one indicated that observed cases exceeded expectations. The cluster with the highest likelihood was labeled the primary cluster, while other clusters were reported if they achieved significance at <inline-formula><mml:math id="M24" display="inline"><mml:mi>p</mml:mi><mml:mo>&#x0003c;</mml:mo><mml:mn>0.05</mml:mn></mml:math></inline-formula>. P-values for space-time clusters were determined through Monte Carlo simulations, using 999 randomized iterations of the data to assess statistical significance [<xref rid="R15" ref-type="bibr">15</xref>].</p></sec><sec id="S6"><label>2.4</label><title>Identified Significant Clusters</title><p id="P17"><xref rid="F1" ref-type="fig">Figure 1</xref> illustrates the significant spatial clusters of COVID-19 mortality across the United States. High-risk clusters, characterized by a relative risk greater than one, are represented by red circles, while low-risk clusters, with a relative risk less than one, are indicated by blue circles. Panel (a) shows the significant spatial clusters for the entire study period; however, it may not adequately reflect the temporal changes across the defined time intervals from 1 to 7. Panels (b) through (h) correspond to the spatial data associated with intervals 1 through 7, respectively. Panel (b) shows that significant high-risk clusters were initially concentrated in the southern regions of the U.S. As shown in panel (c), these high-risk clusters extend across much of the country, with the notable exceptions of certain areas in the East and Northwest, which may be attributed to lower vaccination rates. Panel (d) illustrates the impact of extensive vaccination efforts, which have proven to be substantially effective. The emergence of the Delta variant is highlighted in panel (e), demonstrating its influence on the overall mortality trends. The effectiveness of vaccination updates and booster shots is evident in panel (f) when it is compared with panel (e). Panel (g) indicates that the significance of the Omicron variant has become pronounced in the central regions of the U.S. Finally, panel (h) reveals a marked reduction in the prevalence of significant high-risk clusters, underscoring the impact of vaccination and public health interventions on COVID-19 mortality rates.</p><p id="P18"><xref rid="F2" ref-type="fig">Figure 2</xref> shows the dynamics of COVID-19 mortality trends and the geographical distribution of risk clusters in the U.S., highlighting key patterns in mortality fluctuations, spatial risk persistence, and transitions between risk classifications. Panel (a) illustrates the time series of COVID-19-related mortality in the U.S. from May 24, 2020, to March 12, 2023, segmented into seven distinct periods. Panel (b) shows the percentage of the U.S. area covered by significant high-risk clusters (i.e., relative risk greater than 1) and low-risk clusters (i.e., relative risk less than 1) for each of the seven intervals. It suggests that both high-risk and low-risk clusters may gradually reach an equilibrium below 20% of the U.S. area, indicating a stable distribution of risk across the country. Panel (c) examines the persistence of high-risk and low-risk clusters from time interval <inline-formula><mml:math id="M25" display="inline"><mml:mi>i</mml:mi></mml:math></inline-formula> to interval <inline-formula><mml:math id="M26" display="inline"><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula>. It demonstrates that approximately 45% of areas classified as high-risk and about 40% of areas classified as low-risk maintain their status, particularly from time interval 5 onward, suggesting a notable consistency in spatial risk distribution over time. Panel (d) illustrates the percent transition of clusters from high-risk to low-risk and vice versa between consecutive intervals. This transition appears to stabilize around 15%, indicating a moderate but consistent change in the risk classification of certain areas over the specified periods.</p><p id="P19"><xref rid="T2" ref-type="table">Table 2</xref> presents a detailed summary of spatial high-risk and low-risk clusters of COVID-19 mortality across intervals 1 through 7. The second and third columns list the number of high-risk and low-risk clusters identified in each interval. Columns 4 through 7 provide the percentage of the total U.S. area covered by these clusters. Notably, in interval 4, high-risk clusters encompassed more than 50% of the U.S. area.</p><p id="P20"><xref rid="T3" ref-type="table">Table 3</xref> details the transitions between high-risk and low-risk COVID-19 mortality clusters, as well as the persistence of clusters across intervals from <inline-formula><mml:math id="M27" display="inline"><mml:mi>i</mml:mi></mml:math></inline-formula> to <inline-formula><mml:math id="M28" display="inline"><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula>. Two percentages are provided: the first represents the overlap of clusters from interval <inline-formula><mml:math id="M29" display="inline"><mml:mi>i</mml:mi></mml:math></inline-formula> with those in interval <inline-formula><mml:math id="M30" display="inline"><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula>, and the second indicates the overlap of clusters in interval <inline-formula><mml:math id="M31" display="inline"><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula> with interval <inline-formula><mml:math id="M32" display="inline"><mml:mi>i</mml:mi></mml:math></inline-formula>. Notably, the values approach steady-state levels over time.</p><p id="P21"><xref rid="F3" ref-type="fig">Figure 3</xref> represents the accuracy of identified spatial clusters. Panel (a) shows the observed-to-expected mortality counts plotted against the estimated relative risk for intervals 1&#x02013;7; panel (b) provides box plots of residuals, indicating low residuals; panel (c) illustrates the mortality data within each identified cluster, giving a localized perspective on mortality distribution; and panel (d) displays the estimated relative risks associated with spatial clusters for each interval, showing the changes in risk levels over time. The observed-to-expected mortality counts and the relative risks of significant spatial clusters are available in the <xref rid="SD1" ref-type="supplementary-material">supplementary document Tables 4</xref> and <xref rid="SD1" ref-type="supplementary-material">5</xref>, respectively.</p></sec></sec><sec id="S7"><label>3</label><title>Modeling Mortality Risk Prediction</title><sec id="S8"><label>3.1</label><title>Markov Chain Model</title><p id="P22">In tackling the challenge of integrating the entire cluster search and prediction process into a manageable model grounded in rigorous statistical analysis, we propose that analyzing trends in transition probabilities across observed clusters over <inline-formula><mml:math id="M33" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula> intervals could enhance the accuracy of our predictions for interval <inline-formula><mml:math id="M34" display="inline"><mml:mi>k</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula>. Given that <inline-formula><mml:math id="M35" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula> is relatively small (between 5 and 10), we employ a predictor-corrector method. The predictor utilizes a Markov chain model [<xref rid="R29" ref-type="bibr">29</xref>, <xref rid="R47" ref-type="bibr">47</xref>], while the corrector optimally chooses between multiple linear regression and exponential smoothing methods based on the distance between transition matrices, aiming for an effective prediction of clusters in interval <inline-formula><mml:math id="M36" display="inline"><mml:mi>k</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula>.</p><p id="P23">Our approach starts with the establishment of initial datasets, labeled <inline-formula><mml:math id="M37" display="inline"><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>, each representing a specific time interval. As the pandemic progresses, additional datasets <inline-formula><mml:math id="M38" display="inline"><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mn>4</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>, and so on are created, forming a sequence that captures different phases of the pandemic. The collection <inline-formula><mml:math id="M39" display="inline"><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> denotes the next set of real data gathered at a subsequent interval. Our goal is to predict the parameters of this new collection <inline-formula><mml:math id="M40" display="inline"><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> using insights from the preceding datasets <inline-formula><mml:math id="M41" display="inline"><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x02026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>.</p><p id="P24">Traditionally, predictions for mortality rates have relied solely on the most recent dataset, <inline-formula><mml:math id="M42" display="inline"><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. However, we recognize that the characteristics of the upcoming dataset <inline-formula><mml:math id="M43" display="inline"><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> may differ from those represented by <inline-formula><mml:math id="M44" display="inline"><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. To improve predictive accuracy, we advocate for a broader analysis of the trends in transition probabilities observed across the series of datasets <inline-formula><mml:math id="M45" display="inline"><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x02026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>.</p><p id="P25">Understanding the patterns and shifts in transition probabilities over time can provide valuable insights into the evolving dynamics of the pandemic. By examining how these probabilities change, we aim to capture underlying trends that may enhance the accuracy of predicting the parameters for the next dataset, <inline-formula><mml:math id="M46" display="inline"><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>. This approach contrasts with conventional methods, suggesting that a more nuanced analysis of historical data could lead to a more robust and reliable predictive model for mortality rates in the ongoing fight against COVID-19.</p></sec><sec id="S9"><label>3.2</label><title>Estimating the parameters of <inline-formula><mml:math id="M47" display="inline"><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula></title><p id="P26">Let <inline-formula><mml:math id="M48" display="inline"><mml:msubsup><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula> represent the prediction for <inline-formula><mml:math id="M49" display="inline"><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>. Predicting <inline-formula><mml:math id="M50" display="inline"><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> is impractical because there is no initial data available before <inline-formula><mml:math id="M51" display="inline"><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>. Similarly, predicting <inline-formula><mml:math id="M52" display="inline"><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> is challenging due to the presence of only one preceding data set, which does not provide sufficient patterns or trends. However, <inline-formula><mml:math id="M53" display="inline"><mml:msubsup><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula> can be predicted for <inline-formula><mml:math id="M54" display="inline"><mml:mi>k</mml:mi><mml:mo>&#x0003e;</mml:mo><mml:mn>2</mml:mn></mml:math></inline-formula> using the data sets <inline-formula><mml:math id="M55" display="inline"><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x02026;</mml:mo><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>. This means that historical data can be maintained as prediction chains throughout the data collection process. More importantly, since predictions are made for each dataset as it becomes available, any inaccurate predictions can be discarded, offering valuable feedback on the quality of the predictive process.</p><p id="P27">Our proposed method corrects the Markov chain prediction of <inline-formula><mml:math id="M56" display="inline"><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> utilizing either multiple linear regression or exponential smoothing at each iteration <inline-formula><mml:math id="M57" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula>. The predictor, denoted as <inline-formula><mml:math id="M58" display="inline"><mml:msubsup><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula>, is formulated as follows:
<disp-formula id="FD2">
<label>(2)</label>
<mml:math id="M59" display="block"><mml:msubsup><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mi>&#x003b1;</mml:mi></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msup><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mfenced separators="|"><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:msup><mml:mrow><mml:mi>&#x003b1;</mml:mi></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:mfenced><mml:msubsup><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msubsup><mml:mo>,</mml:mo></mml:math>
</disp-formula>
which is a weighted sum between a &#x0201c;good corrector&#x0201d; <inline-formula><mml:math id="M60" display="inline"><mml:msub><mml:mrow><mml:msup><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and data <inline-formula><mml:math id="M61" display="inline"><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. The &#x0201c;good corrector&#x0201d; <inline-formula><mml:math id="M62" display="inline"><mml:msub><mml:mrow><mml:msup><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> of <inline-formula><mml:math id="M63" display="inline"><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> based on transition matrices from previous datasets <inline-formula><mml:math id="M64" display="inline"><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x02026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>, whereas the predictor is given by <inline-formula><mml:math id="M65" display="inline"><mml:msup><mml:mrow><mml:mi>&#x003b1;</mml:mi></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msup><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>.</p><p id="P28">Specifically, let <inline-formula><mml:math id="M66" display="inline"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mi>&#x002c6;</mml:mi></mml:mover></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> be a predictor of <inline-formula><mml:math id="M67" display="inline"><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> determined through multiple linear regression, and let <inline-formula><mml:math id="M68" display="inline"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mo>&#x002dc;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> be a predictor of <inline-formula><mml:math id="M69" display="inline"><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> determined via exponential smoothing. <inline-formula><mml:math id="M70" display="inline"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mo>&#x002dc;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> depends on <inline-formula><mml:math id="M71" display="inline"><mml:msup><mml:mrow><mml:mi>&#x003b1;</mml:mi></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msup></mml:math></inline-formula>, an optimal value satisfying <inline-formula><mml:math id="M72" display="inline"><mml:mn>0</mml:mn><mml:mo>&#x0003c;</mml:mo><mml:msup><mml:mrow><mml:mi>&#x003b1;</mml:mi></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msup><mml:mo>&#x0003c;</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula>, and is determined by
<disp-formula id="FD3">
<label>(3)</label>
<mml:math id="M73" display="block"><mml:mrow><mml:munderover><mml:mstyle><mml:mo>&#x02211;</mml:mo></mml:mstyle><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>3</mml:mn></mml:mrow><mml:mi>k</mml:mi></mml:munderover><mml:mi>d</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>T</mml:mi><mml:mo>&#x002dc;</mml:mo></mml:mover></mml:mrow><mml:mi>i</mml:mi></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msup><mml:mi>&#x003b1;</mml:mi><mml:mo>*</mml:mo></mml:msup></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:munder><mml:mrow><mml:mtext>min</mml:mtext></mml:mrow><mml:mrow><mml:mn>0</mml:mn><mml:mo>&#x0003c;</mml:mo><mml:mi>a</mml:mi><mml:mo>&#x0003c;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:munder><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:munderover><mml:mstyle><mml:mo>&#x02211;</mml:mo></mml:mstyle><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>3</mml:mn></mml:mrow><mml:mi>k</mml:mi></mml:munderover><mml:mi>d</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>T</mml:mi><mml:mo>&#x002dc;</mml:mo></mml:mover></mml:mrow><mml:mi>i</mml:mi></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>&#x003b1;</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>}</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mrow></mml:math>
</disp-formula>
<inline-formula><mml:math id="M74" display="inline"><mml:msubsup><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula> is chosen such that
<disp-formula id="FD4">
<label>(4)</label>
<mml:math id="M75" display="block"><mml:mi>d</mml:mi><mml:mfenced separators="|"><mml:mrow><mml:msubsup><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfenced><mml:mo>=</mml:mo><mml:mtext>min</mml:mtext><mml:mfenced open="{" close="}" separators="|"><mml:mrow><mml:mi>d</mml:mi><mml:mfenced separators="|"><mml:mrow><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mi>&#x002c6;</mml:mi></mml:mover></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfenced><mml:mo>,</mml:mo><mml:mi>d</mml:mi><mml:mfenced separators="|"><mml:mrow><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mo>&#x002dc;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfenced></mml:mrow></mml:mfenced><mml:mo>.</mml:mo></mml:math>
</disp-formula>
where <inline-formula><mml:math id="M76" display="inline"><mml:mi>d</mml:mi><mml:mfenced separators="|"><mml:mrow><mml:msubsup><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfenced></mml:math></inline-formula> denotes a distance from transition matrix <inline-formula><mml:math id="M77" display="inline"><mml:msubsup><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula> to matrix <inline-formula><mml:math id="M78" display="inline"><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. We will use a distance measure of the form
<disp-formula id="FD5">
<label>(5)</label>
<mml:math id="M79" display="block"><mml:mi>d</mml:mi><mml:mfenced separators="|"><mml:mrow><mml:msubsup><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfenced><mml:mo>=</mml:mo><mml:mrow><mml:munder><mml:mo stretchy="true">&#x02211;</mml:mo><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:munder></mml:mrow><mml:mrow><mml:munder><mml:mo stretchy="true">&#x02211;</mml:mo><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:munder></mml:mrow><mml:mfrac><mml:mrow><mml:mfenced open="|" close="|" separators="|"><mml:mrow><mml:mfenced separators="|"><mml:mrow><mml:msubsup><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msubsup><mml:mo>(</mml:mo><mml:mi>m</mml:mi><mml:mo>,</mml:mo><mml:mi>n</mml:mi><mml:mo>)</mml:mo><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>(</mml:mo><mml:mi>m</mml:mi><mml:mo>,</mml:mo><mml:mi>n</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mfenced></mml:mrow></mml:mfenced></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>(</mml:mo><mml:mi>m</mml:mi><mml:mo>,</mml:mo><mml:mi>n</mml:mi><mml:mo>)</mml:mo><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mfrac></mml:math>
</disp-formula>
Here, <inline-formula><mml:math id="M80" display="inline"><mml:mi>m</mml:mi></mml:math></inline-formula> represents the matrix row index, while <inline-formula><mml:math id="M81" display="inline"><mml:mi>n</mml:mi></mml:math></inline-formula> denotes the matrix column index. This distance metric accounts for relative differences rather than absolute differences. See Ref. [<xref rid="R38" ref-type="bibr">38</xref>] for more details.</p><p id="P31">As a result, we derive the parameters of <inline-formula><mml:math id="M82" display="inline"><mml:msubsup><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula> by utilizing the previous dataset <inline-formula><mml:math id="M83" display="inline"><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, which is established with real data, and <inline-formula><mml:math id="M84" display="inline"><mml:msubsup><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula>, a predictor for <inline-formula><mml:math id="M85" display="inline"><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> that encapsulates historical information from the datasets <inline-formula><mml:math id="M86" display="inline"><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x02026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>. Throughout this process, we eventually obtain the actual <inline-formula><mml:math id="M87" display="inline"><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> and compare it to our estimate <inline-formula><mml:math id="M88" display="inline"><mml:msubsup><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula> using the aforementioned distance measure. This method allows us to assess the accuracy of our predictions with each successive data collection. Importantly, the distance measure can reveal whether certain datasets are outliers for various reasons, enabling us to exclude those data points from our analysis.</p><p id="P32">The ultimate goal is to determine <inline-formula><mml:math id="M89" display="inline"><mml:msubsup><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula>, for which no corresponding <inline-formula><mml:math id="M90" display="inline"><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> will be available, especially when <inline-formula><mml:math id="M91" display="inline"><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> represents the final data collection before making a prediction. Thus, <inline-formula><mml:math id="M92" display="inline"><mml:msubsup><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula> serves as our prediction for practical use, representing our estimate in the absence of actual observed data for <inline-formula><mml:math id="M93" display="inline"><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>.</p></sec><sec id="S10"><label>3.3</label><title>Predictors of the transition matrix <inline-formula><mml:math id="M94" display="inline"><mml:msubsup><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula></title><p id="P33">We employ one of two methods to predict <inline-formula><mml:math id="M95" display="inline"><mml:msubsup><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula> based on the transition matrices <inline-formula><mml:math id="M96" display="inline"><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x02026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. The value <inline-formula><mml:math id="M97" display="inline"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mo>&#x002dc;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is calculated using exponential smoothing, while <inline-formula><mml:math id="M98" display="inline"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mi>&#x002c6;</mml:mi></mml:mover></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is obtained through multiple linear regression. The method that results in the smaller distance measure, as defined earlier, is chosen for the prediction.</p><sec id="S11"><label>3.3.1</label><title>Derivation of <inline-formula><mml:math id="M99" display="inline"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi mathvariant="bold-italic">T</mml:mi></mml:mrow><mml:mo>&#x002dc;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi mathvariant="bold-italic">k</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> (exponential smoothing):</title><p id="P34"><inline-formula><mml:math id="M100" display="inline"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mo>&#x002dc;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is determined through the recursion
<disp-formula id="FD6">
<label>(6)</label>
<mml:math id="M101" display="block"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi>T</mml:mi><mml:mo>&#x002dc;</mml:mo></mml:mover><mml:mi>k</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mi>&#x003b1;</mml:mi><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x02212;</mml:mo><mml:mi>&#x003b1;</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:msub><mml:mover accent="true"><mml:mi>T</mml:mi><mml:mo>&#x002dc;</mml:mo></mml:mover><mml:mrow><mml:mi>k</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mspace linebreak="newline"/><mml:msub><mml:mover accent="true"><mml:mi>T</mml:mi><mml:mo>&#x002dc;</mml:mo></mml:mover><mml:mrow><mml:mi>k</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>&#x003b1;</mml:mi><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x02212;</mml:mo><mml:mi>&#x003b1;</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:msub><mml:mover accent="true"><mml:mi>T</mml:mi><mml:mo>&#x002dc;</mml:mo></mml:mover><mml:mrow><mml:mi>k</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mspace linebreak="newline"/><mml:mtable columnalign="left"><mml:mtr columnalign="left"><mml:mtd columnalign="left"><mml:mo>&#x022c5;</mml:mo></mml:mtd><mml:mtd columnalign="center"><mml:mo>&#x022c5;</mml:mo></mml:mtd><mml:mtd columnalign="center"><mml:mo>&#x022c5;</mml:mo></mml:mtd></mml:mtr><mml:mtr columnalign="left"><mml:mtd columnalign="left"><mml:mo>&#x022c5;</mml:mo></mml:mtd><mml:mtd columnalign="center"><mml:mo>&#x022c5;</mml:mo></mml:mtd><mml:mtd columnalign="center"><mml:mo>&#x022c5;</mml:mo></mml:mtd></mml:mtr><mml:mtr columnalign="left"><mml:mtd columnalign="left"><mml:mo>&#x022c5;</mml:mo></mml:mtd><mml:mtd columnalign="center"><mml:mo>&#x022c5;</mml:mo></mml:mtd><mml:mtd columnalign="center"><mml:mo>&#x022c5;</mml:mo></mml:mtd></mml:mtr><mml:mtr columnalign="center"><mml:mtd columnalign="center"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi>T</mml:mi><mml:mo>&#x002dc;</mml:mo></mml:mover><mml:mn>3</mml:mn></mml:msub></mml:mrow></mml:mtd><mml:mtd columnalign="center"><mml:mrow><mml:mo>=</mml:mo><mml:mi>&#x003b1;</mml:mi><mml:msub><mml:mi>T</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow></mml:mtd><mml:mtd columnalign="center"><mml:mrow><mml:mo>+</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x02212;</mml:mo><mml:mi>&#x003b1;</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:msub><mml:mover accent="true"><mml:mi>T</mml:mi><mml:mo>^</mml:mo></mml:mover><mml:mn>2</mml:mn></mml:msub></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:math>
</disp-formula>
where <inline-formula><mml:math id="M102" display="inline"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mi>&#x002c6;</mml:mi></mml:mover></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> is an estimate of <inline-formula><mml:math id="M103" display="inline"><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> determined by simple linear regression.</p></sec><sec id="S12"><label>3.3.2</label><title>Derivation of <inline-formula><mml:math id="M104" display="inline"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi mathvariant="bold-italic">T</mml:mi></mml:mrow><mml:mi mathvariant="bold-italic">&#x002c6;</mml:mi></mml:mover></mml:mrow><mml:mrow><mml:mi mathvariant="bold-italic">k</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> (multiple linear regression):</title><p id="P35">The general form of <inline-formula><mml:math id="M105" display="inline"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mi>&#x002c6;</mml:mi></mml:mover></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is
<disp-formula id="FD7">
<label>(7)</label>
<mml:math id="M106" display="block"><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>T</mml:mi><mml:mo>&#x002c6;</mml:mo></mml:mover></mml:mrow><mml:mi>k</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mn>0</mml:mn></mml:msub><mml:mo>+</mml:mo><mml:munderover><mml:mstyle><mml:mo>&#x02211;</mml:mo></mml:mstyle><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:munderover><mml:msub><mml:mi>a</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:msub><mml:mi>T</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:munder><mml:mstyle><mml:mo>&#x02211;</mml:mo></mml:mstyle><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x0003c;</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:munder><mml:munderover><mml:mstyle><mml:mo>&#x02211;</mml:mo></mml:mstyle><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>2</mml:mn></mml:mrow><mml:mi>k</mml:mi></mml:munderover><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi>T</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:msub><mml:mi>T</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>,</mml:mo></mml:mrow></mml:math>
</disp-formula>
where <inline-formula><mml:math id="M107" display="inline"><mml:msub><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> for <inline-formula><mml:math id="M108" display="inline"><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x02026;</mml:mo><mml:mo>,</mml:mo><mml:mi>k</mml:mi><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula> and <inline-formula><mml:math id="M109" display="inline"><mml:msub><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> for <inline-formula><mml:math id="M110" display="inline"><mml:mi>i</mml:mi><mml:mo>&#x0003c;</mml:mo><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x02026;</mml:mo><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:math></inline-formula> are determined through the least square method. In certain cases, the value of the distance is not increased significantly by including the interaction terms <inline-formula><mml:math id="M111" display="inline"><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. In these cases, a simple multiple linear regression model of the form
<disp-formula id="FD8">
<label>(8)</label>
<mml:math id="M112" display="block"><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>T</mml:mi><mml:mo>&#x002c6;</mml:mo></mml:mover></mml:mrow><mml:mi>k</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mn>0</mml:mn></mml:msub><mml:mo>+</mml:mo><mml:munderover><mml:mstyle><mml:mo>&#x02211;</mml:mo></mml:mstyle><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:munderover><mml:msub><mml:mi>a</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:msub><mml:mi>T</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math>
</disp-formula>
may be used.</p></sec></sec></sec><sec id="S13"><label>4</label><title>Predicting Relative Risks of Clusters</title><p id="P36">In this section, we use the relative risk data obtained from the statistical spatial scan (see <xref rid="S2" ref-type="sec">Section 2</xref>) to test the accuracy of our proposed model. In this process, we aim to use the first six datasets of relative risks to predict the seventh set of relative risks associated with interval seven. Note that we only apply our method to the highest 25 relative risks of each interval. Since the seventh dataset is known, we can evaluate the accuracy of our predictions. Given <inline-formula><mml:math id="M113" display="inline"><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x02026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mn>6</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> as the six datasets corresponding to the observations, we seek to determine <inline-formula><mml:math id="M114" display="inline"><mml:msubsup><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mn>7</mml:mn></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula> as an appropriate predictor for <inline-formula><mml:math id="M115" display="inline"><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mn>7</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>. In addition, all coefficients included in the forecasting model were tested for statistical significance, and variables with non-significant coefficients were excluded to ensure the robustness and reliability of the model.</p><sec id="S14"><label>4.1.</label><title>Exponential Smoothing:</title><p id="P37">As described in <xref rid="S7" ref-type="sec">Section 3</xref>, the predictor of chain <inline-formula><mml:math id="M116" display="inline"><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mn>6</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> has the form
<disp-formula id="FD9">
<mml:math id="M117" display="block"><mml:mrow><mml:mtable><mml:mtr><mml:mtd><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi>T</mml:mi><mml:mo>&#x002dc;</mml:mo></mml:mover><mml:mn>6</mml:mn></mml:msub></mml:mrow></mml:mtd><mml:mtd><mml:mo>=</mml:mo></mml:mtd><mml:mtd><mml:mrow><mml:mi>&#x003b1;</mml:mi><mml:msub><mml:mi>T</mml:mi><mml:mn>5</mml:mn></mml:msub></mml:mrow></mml:mtd><mml:mtd><mml:mo>+</mml:mo></mml:mtd><mml:mtd><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x02212;</mml:mo><mml:mi>&#x003b1;</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:msub><mml:mover accent="true"><mml:mi>T</mml:mi><mml:mo>&#x002dc;</mml:mo></mml:mover><mml:mn>5</mml:mn></mml:msub></mml:mrow></mml:mtd></mml:mtr><mml:mtr columnalign="left"><mml:mtd columnalign="left"><mml:mo>&#x022c5;</mml:mo></mml:mtd><mml:mtd/><mml:mtd columnalign="center"><mml:mo>&#x022c5;</mml:mo></mml:mtd><mml:mtd/><mml:mtd columnalign="center"><mml:mo>&#x022c5;</mml:mo></mml:mtd></mml:mtr><mml:mtr columnalign="left"><mml:mtd columnalign="left"><mml:mo>&#x022c5;</mml:mo></mml:mtd><mml:mtd/><mml:mtd columnalign="center"><mml:mo>&#x022c5;</mml:mo></mml:mtd><mml:mtd/><mml:mtd columnalign="center"><mml:mo>&#x022c5;</mml:mo></mml:mtd></mml:mtr><mml:mtr columnalign="left"><mml:mtd columnalign="left"><mml:mo>&#x022c5;</mml:mo></mml:mtd><mml:mtd/><mml:mtd columnalign="center"><mml:mo>&#x022c5;</mml:mo></mml:mtd><mml:mtd/><mml:mtd columnalign="center"><mml:mo>&#x022c5;</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi>T</mml:mi><mml:mo>&#x002dc;</mml:mo></mml:mover><mml:mn>3</mml:mn></mml:msub></mml:mrow></mml:mtd><mml:mtd><mml:mo>=</mml:mo></mml:mtd><mml:mtd><mml:mrow><mml:mi>&#x003b1;</mml:mi><mml:msub><mml:mi>T</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow></mml:mtd><mml:mtd><mml:mo>+</mml:mo></mml:mtd><mml:mtd><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x02212;</mml:mo><mml:mi>&#x003b1;</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:msub><mml:mover accent="true"><mml:mi>T</mml:mi><mml:mo>^</mml:mo></mml:mover><mml:mn>2</mml:mn></mml:msub></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi>T</mml:mi><mml:mo>^</mml:mo></mml:mover><mml:mn>2</mml:mn></mml:msub></mml:mrow></mml:mtd><mml:mtd><mml:mo>=</mml:mo></mml:mtd><mml:mtd><mml:mrow><mml:mn>1.568</mml:mn></mml:mrow></mml:mtd><mml:mtd><mml:mo>+</mml:mo></mml:mtd><mml:mtd><mml:mrow><mml:mn>0.131</mml:mn><mml:msub><mml:mi>T</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:math>
</disp-formula>
where <inline-formula><mml:math id="M118" display="inline"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mi>&#x002c6;</mml:mi></mml:mover></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> is a predictor of <inline-formula><mml:math id="M119" display="inline"><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>, given <inline-formula><mml:math id="M120" display="inline"><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>, determined through simple linear regression. <inline-formula><mml:math id="M121" display="inline"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mo>&#x002dc;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mn>6</mml:mn></mml:mrow></mml:msub><mml:mfenced separators="|"><mml:mrow><mml:msup><mml:mrow><mml:mi>&#x003b1;</mml:mi></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:mfenced></mml:math></inline-formula> is chosen so that
<disp-formula id="FD10">
<mml:math id="M122" display="block"><mml:mrow><mml:msubsup><mml:mo stretchy="false">&#x02211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>3</mml:mn></mml:mrow><mml:mrow><mml:mn>6</mml:mn></mml:mrow></mml:msubsup><mml:mrow><mml:mi>d</mml:mi><mml:mfenced close="" separators="|"><mml:mrow><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mo>&#x002dc;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mfenced separators="|"><mml:mrow><mml:msup><mml:mrow><mml:mi>&#x003b1;</mml:mi></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:mfenced><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mtext>min</mml:mtext></mml:mrow><mml:mrow><mml:mn>0</mml:mn><mml:mo>&#x0003c;</mml:mo><mml:mi>a</mml:mi><mml:mo>&#x0003c;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:msubsup><mml:mo stretchy="false">&#x02211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>3</mml:mn></mml:mrow><mml:mrow><mml:mn>6</mml:mn></mml:mrow></mml:msubsup><mml:mrow><mml:mfenced open="{" close="}" separators="|"><mml:mrow><mml:mi>d</mml:mi><mml:mfenced separators="|"><mml:mrow><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mo>&#x002dc;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>(</mml:mo><mml:mi>&#x003b1;</mml:mi><mml:mo>)</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfenced></mml:mrow></mml:mfenced></mml:mrow></mml:mrow></mml:mrow></mml:mfenced></mml:mrow></mml:mrow><mml:mo>,</mml:mo></mml:math>
</disp-formula>
where the distance <inline-formula><mml:math id="M123" display="inline"><mml:mi>d</mml:mi></mml:math></inline-formula> is defined as in <xref rid="S7" ref-type="sec">Section 3</xref>. represents the values of <inline-formula><mml:math id="M124" display="inline"><mml:mi>D</mml:mi><mml:mo>(</mml:mo><mml:mi>&#x003b1;</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mrow><mml:munderover><mml:mo stretchy="true">&#x02211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>3</mml:mn></mml:mrow><mml:mrow><mml:mn>6</mml:mn></mml:mrow></mml:munderover></mml:mrow><mml:mi>d</mml:mi><mml:mfenced separators="|"><mml:mrow><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mo>&#x002dc;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>(</mml:mo><mml:mi>&#x003b1;</mml:mi><mml:mo>)</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfenced></mml:math></inline-formula>. The minimum occurs for <inline-formula><mml:math id="M125" display="inline"><mml:msup><mml:mrow><mml:mi>&#x003b1;</mml:mi></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mn>0.02</mml:mn></mml:math></inline-formula> giving
<disp-formula id="FD11">
<mml:math id="M126" display="block"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mo>&#x002dc;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mn>6</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>0.02</mml:mn><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mn>5</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mn>0.98</mml:mn><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mi>&#x002c6;</mml:mi></mml:mover></mml:mrow><mml:mrow><mml:mn>5</mml:mn></mml:mrow></mml:msub></mml:math>
</disp-formula>
Also, note that
<disp-formula id="FD12">
<mml:math id="M127" display="block"><mml:mi>d</mml:mi><mml:mfenced separators="|"><mml:mrow><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mo>&#x002dc;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mn>6</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mn>6</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mfenced><mml:mo>=</mml:mo><mml:mrow><mml:msubsup><mml:mo stretchy="false">&#x02211;</mml:mo><mml:mrow><mml:mi>m</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>6</mml:mn></mml:mrow></mml:msubsup><mml:mrow><mml:mrow><mml:msubsup><mml:mo stretchy="false">&#x02211;</mml:mo><mml:mrow><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>6</mml:mn></mml:mrow></mml:msubsup><mml:mrow><mml:mfrac><mml:mrow><mml:mfenced open="|" close="|" separators="|"><mml:mrow><mml:mfenced separators="|"><mml:mrow><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mo>&#x002dc;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mn>6</mml:mn></mml:mrow></mml:msub><mml:mo>(</mml:mo><mml:mi>m</mml:mi><mml:mo>,</mml:mo><mml:mi>n</mml:mi><mml:mo>)</mml:mo><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mn>6</mml:mn></mml:mrow></mml:msub><mml:mo>(</mml:mo><mml:mi>m</mml:mi><mml:mo>,</mml:mo><mml:mi>n</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mfenced></mml:mrow></mml:mfenced></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>(</mml:mo><mml:mi>m</mml:mi><mml:mo>,</mml:mo><mml:mi>n</mml:mi><mml:mo>)</mml:mo><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mfrac><mml:mo>=</mml:mo><mml:mn>11.756</mml:mn></mml:mrow></mml:mrow></mml:mrow></mml:mrow></mml:math>
</disp-formula></p></sec><sec id="S15"><label>4.2</label><title>Multiple Linear Regression:</title><p id="P38">Multiple linear regression relates transition matrix <inline-formula><mml:math id="M128" display="inline"><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mn>6</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> to <inline-formula><mml:math id="M129" display="inline"><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x02026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mn>6</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> in one linear functional form. Results in least squares coefficients are presented in <xref rid="T2" ref-type="table">Table 2</xref>.</p><p id="P39">The prediction equation is
<disp-formula id="FD13">
<mml:math id="M130" display="block"><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>T</mml:mi><mml:mo>&#x002c6;</mml:mo></mml:mover></mml:mrow><mml:mn>6</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mn>0.5</mml:mn><mml:msub><mml:mi>T</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>+</mml:mo><mml:mn>6.357</mml:mn><mml:msub><mml:mi>T</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>&#x02212;</mml:mo><mml:mn>0.561</mml:mn><mml:msub><mml:mi>T</mml:mi><mml:mn>3</mml:mn></mml:msub><mml:mo>+</mml:mo><mml:mn>12.287</mml:mn><mml:msub><mml:mi>T</mml:mi><mml:mn>4</mml:mn></mml:msub><mml:mo>&#x02212;</mml:mo><mml:mn>8.645</mml:mn><mml:msub><mml:mi>T</mml:mi><mml:mn>5</mml:mn></mml:msub></mml:mrow></mml:math>
</disp-formula>
<disp-formula id="FD14">
<mml:math id="M131" display="block"><mml:mrow><mml:mi>d</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>T</mml:mi><mml:mo>&#x002c6;</mml:mo></mml:mover></mml:mrow><mml:mn>6</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mn>6</mml:mn></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msubsup><mml:mstyle><mml:mo>&#x02211;</mml:mo></mml:mstyle><mml:mrow><mml:mi>m</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mn>6</mml:mn></mml:msubsup><mml:msubsup><mml:mstyle><mml:mo>&#x02211;</mml:mo></mml:mstyle><mml:mrow><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mn>6</mml:mn></mml:msubsup><mml:mfrac><mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>T</mml:mi><mml:mo>&#x002c6;</mml:mo></mml:mover></mml:mrow><mml:mn>6</mml:mn></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>m</mml:mi><mml:mo>,</mml:mo><mml:mi>n</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mn>6</mml:mn></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>m</mml:mi><mml:mo>,</mml:mo><mml:mi>n</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>|</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>m</mml:mi><mml:mo>,</mml:mo><mml:mi>n</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mfrac><mml:mo>=</mml:mo><mml:mn>5.887</mml:mn></mml:mrow></mml:math>
</disp-formula>
which is smaller than <inline-formula><mml:math id="M132" display="inline"><mml:mi>d</mml:mi><mml:mfenced separators="|"><mml:mrow><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mo>&#x002dc;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mn>6</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mn>6</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mfenced></mml:math></inline-formula> The better predictor, therefore, of <inline-formula><mml:math id="M133" display="inline"><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mn>6</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> is <inline-formula><mml:math id="M134" display="inline"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mi>&#x002c6;</mml:mi></mml:mover></mml:mrow><mml:mrow><mml:mn>6</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> Consequently, the predictor of <inline-formula><mml:math id="M135" display="inline"><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mn>7</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> is
<disp-formula id="FD15">
<label>(9)</label>
<mml:math id="M136" display="block"><mml:msubsup><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mn>7</mml:mn></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mn>0.02</mml:mn><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mn>6</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mn>0.98</mml:mn><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mi>&#x002c6;</mml:mi></mml:mover></mml:mrow><mml:mrow><mml:mn>6</mml:mn></mml:mrow></mml:msub></mml:math>
</disp-formula></p><p id="P40"><xref rid="F4" ref-type="fig">Figure 4</xref> shows the estimation of the optimal parameter <inline-formula><mml:math id="M137" display="inline"><mml:mi>&#x003b1;</mml:mi></mml:math></inline-formula>, and the prediction accuracy of relative risks for interval 7 across different models. Panel (a) shows the determination of the optimal <inline-formula><mml:math id="M138" display="inline"><mml:mi>&#x003b1;</mml:mi></mml:math></inline-formula> by minimizing the distance between estimated and observed relative risk values. Panel (b) compares observed versus predicted values (on a log scale) for the proposed Markov chain model and the Exponential Smoothing. The predictions of the relative risks look promising but have room for improvement. The proposed model demonstrates superior predictive performance compared to the Exponential Smoothing method, as evidenced by a lower Mean Squared Error (164.8 vs. 258.7).</p></sec></sec><sec id="S16"><label>5</label><title>Discussion</title><p id="P41">This study represents an initial step toward enabling statistical spatial cluster analysis to predict future spatial clusters effectively. By applying our developed method to U.S. COVID-19 mortality data across seven time intervals, we demonstrate that predictive accuracy ranges from moderate to high (<xref rid="F4" ref-type="fig">Figure 4</xref>), indicating its potential utility in disease surveillance, resource allocation, and pandemic preparedness. The length of time intervals can be adapted based on the characteristics of specific infectious diseases and data availability. Despite these promising findings, several limitations warrant consideration.</p><p id="P42">First, it is important to recognize that predicting relative risks for subsequent intervals does not ensure that spatial clusters will consistently align with those of prior intervals. As shown in <xref rid="T3" ref-type="table">Table 3</xref>, the overlap between high-risk regions from one interval to the next (i.e., interval i+1 and interval i) is often limited, varying between 17% and 51%. Nevertheless, these relative risk predictions can provide health officials and policymakers with an anticipatory perspective on likely future conditions, offering a gauge of infection severity that may aid in decision-making.</p><p id="P43">Second, the accuracy of the correction phase in the Markov chain model could be enhanced by incorporating additional techniques, such as autoregressive or moving average models, or by accounting for seasonality when relevant, as with influenza. Although the current model&#x02019;s relative risk predictions show acceptable accuracy, future research should focus on refining these predictions to improve reliability.</p><p id="P44">Third, and perhaps most critically, the effectiveness of relative risk predictions would be substantially enhanced by integrating predictions of spatial cluster radii. Our attempts to incorporate radii into the same predictive matrices as relative risks did not yield high accuracy. We propose that future models should separate radius and relative risk predictions into two interconnected submodels to improve prediction precision and practical applicability.</p><p id="P45">Our Future research will aim to advance the proposed methodology and enhance its applicability by addressing a few key areas as follows. Overcoming inconsistent spatial cluster alignment remains a significant challenge, as variations between consecutive intervals can limit the reliability of predictions. Developing algorithms that dynamically adjust the time intervals to overcome inconsistencies, such as machine learning-based alignment models [<xref rid="R28" ref-type="bibr">28</xref>], could significantly improve the robustness of our methodology. Additionally, enhancing the accuracy of the correction phase through hybrid models that combine Markov chains with differential equations [<xref rid="R10" ref-type="bibr">10</xref>], neural networks [<xref rid="R48" ref-type="bibr">48</xref>] or ensemble techniques [<xref rid="R20" ref-type="bibr">20</xref>] could improve the predictions. Furthermore, integrating spatial cluster radii predictions with relative risks in a cohesive framework will require innovative approaches, such as leveraging spatiotemporal modeling [<xref rid="R49" ref-type="bibr">49</xref>] and advanced simulation techniques [<xref rid="R17" ref-type="bibr">17</xref>], to achieve precise and reliable outputs. These enhancements could broaden the scope of applications for disease surveillance, calculation of the basic reproduction number <inline-formula><mml:math id="M140" display="inline"><mml:msub><mml:mrow><mml:mi>R</mml:mi></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> for each spatial cluster [<xref rid="R11" ref-type="bibr">11</xref>], and health resource management, ultimately contributing to more effective public health interventions.</p><p id="P46">In conclusion, this study highlights the potential of modified Markov chain models in predicting the relative risks of spatial clusters. This approach could be further developed to predict not only the relative risks but also the location and radius of each spatial cluster, enhancing its value for public health planning and intervention.</p></sec><sec sec-type="supplementary-material" id="SM1"><title>Supplementary Material</title><supplementary-material id="SD1" position="float" content-type="local-data"><label>1</label><media xlink:href="NIHMS2053985-supplement-1.pdf" id="d67e3618" position="anchor"/></supplementary-material></sec></body><back><ack id="S17"><title>Acknowledgement</title><p id="P47">This study was partially supported by the Center for Disease Control and Prevention under grant number 5U01CK000671-02.</p></ack><ref-list><title>References</title><ref id="R1"><label>[1]</label><mixed-citation publication-type="journal"><name><surname>Abboud</surname><given-names>CS</given-names></name>
<etal/>
<article-title>A space-time model for carbapenemase-producing klebsiella pneumoniae in hospitals</article-title>. <source>Epidemiology and Infection</source>, <year>2015</year>.</mixed-citation></ref><ref id="R2"><label>[2]</label><mixed-citation publication-type="journal"><name><surname>AlQadi</surname><given-names>H</given-names></name> and <name><surname>Bani-Yaghoub</surname><given-names>M</given-names></name>. <article-title>Incorporating global dynamics to improve the accuracy of disease models: Example of a covid-19 sir model</article-title>. <source>Plos one</source>, <volume>17</volume>(<issue>4</issue>):<fpage>e0265815</fpage>, <year>2022</year>.<pub-id pub-id-type="pmid">35395018</pub-id>
</mixed-citation></ref><ref id="R3"><label>[3]</label><mixed-citation publication-type="journal"><name><surname>AlQadi</surname><given-names>H</given-names></name>, <name><surname>Bani-Yaghoub</surname><given-names>M</given-names></name>, <name><surname>Balakumar</surname><given-names>S</given-names></name>, <name><surname>Wu</surname><given-names>S</given-names></name>, and <name><surname>Francisco</surname><given-names>A</given-names></name>. <article-title>Assessment of retrospective covid-19 spatial clusters with respect to demographic factors: case study of kansas city, missouri, united states</article-title>. <source>International Journal of Environmental Research and Public Health</source>, <volume>18</volume>(<issue>21</issue>):<fpage>11496</fpage>, <year>2021</year>.<pub-id pub-id-type="pmid">34770012</pub-id>
</mixed-citation></ref><ref id="R4"><label>[4]</label><mixed-citation publication-type="journal"><name><surname>AlQadi</surname><given-names>H</given-names></name>, <name><surname>Bani-Yaghoub</surname><given-names>M</given-names></name>, <name><surname>Wu</surname><given-names>S</given-names></name>, <name><surname>Balakumar</surname><given-names>S</given-names></name>, and <name><surname>Francisco</surname><given-names>A</given-names></name>. <article-title>Prospective spatial-temporal clusters of covid-19 in local communities: case study of kansas city, missouri, united states</article-title>. <source>Epidemiology &#x00026; Infection</source>, <volume>151</volume>:<fpage>e178</fpage>, <year>2023</year>.</mixed-citation></ref><ref id="R5"><label>[5]</label><mixed-citation publication-type="journal"><name><surname>Amin</surname><given-names>S</given-names></name>
<etal/>
<article-title>Applications of spatial scan statistics in disease cluster detection</article-title>. <source>International Journal of Health Geographics</source>, <year>2014</year>.</mixed-citation></ref><ref id="R6"><label>[6]</label><mixed-citation publication-type="journal"><name><surname>Andrade</surname><given-names>LA</given-names></name>
<etal/>
<article-title>Covid-19 mortality in an area of northeast brazil</article-title>. <source>Epidemiology &#x00026; Infection</source>, <volume>148</volume>, <fpage>2020</fpage>.</mixed-citation></ref><ref id="R7"><label>[7]</label><mixed-citation publication-type="journal"><name><surname>Anwar</surname><given-names>A</given-names></name>, <name><surname>Malik</surname><given-names>M</given-names></name>, <name><surname>Raees</surname><given-names>V</given-names></name>, and <name><surname>Anwar</surname><given-names>A</given-names></name>. <article-title>Role of mass media and public health communications in the covid-19 pandemic</article-title>. <source>Cureus</source>, <volume>12</volume>(<issue>9</issue>), <fpage>2020</fpage>.</mixed-citation></ref><ref id="R8"><label>[8]</label><mixed-citation publication-type="journal"><name><surname>Rikke Baastrup</surname><given-names>N</given-names></name>
<etal/>
<article-title>Use of spatial scan statistics for cancer epidemiology</article-title>. <source>Environmental Research</source>, <year>2013</year>.</mixed-citation></ref><ref id="R9"><label>[9]</label><mixed-citation publication-type="journal"><name><surname>Badaloni</surname><given-names>C</given-names></name>
<etal/>
<article-title>Space-time clustering of covid-19 in europe</article-title>. <source>Geospatial Health</source>, <year>2020</year>.</mixed-citation></ref><ref id="R10"><label>[10]</label><mixed-citation publication-type="journal"><name><surname>Bani-Yaghoub</surname><given-names>Majid</given-names></name>, <name><surname>Elhomani</surname><given-names>Abdellatif</given-names></name>, and <name><surname>Catley</surname><given-names>Delwyn</given-names></name>. <article-title>Effectiveness of motivational interviewing, health education and brief advice in a population of smokers who are not ready to quit</article-title>. <source>BMC Medical Research Methodology</source>, <volume>18</volume>:<fpage>1</fpage>&#x02013;<lpage>10</lpage>, <year>2018</year>.<pub-id pub-id-type="pmid">29301497</pub-id>
</mixed-citation></ref><ref id="R11"><label>[11]</label><mixed-citation publication-type="journal"><name><surname>Bani-Yaghoub</surname><given-names>Majid</given-names></name>, <name><surname>Gautam</surname><given-names>Raju</given-names></name>, <name><surname>Shuai</surname><given-names>Zhisheng</given-names></name>, <name><surname>Van Den Driessche</surname><given-names>P</given-names></name>, and <name><surname>Ivanek</surname><given-names>Renata</given-names></name>. <article-title>Reproduction numbers for infections with free-living pathogens growing in the environment</article-title>. <source>Journal of biological dynamics</source>, <volume>6</volume>(<issue>2</issue>):<fpage>923</fpage>&#x02013;<lpage>940</lpage>, <year>2012</year>.<pub-id pub-id-type="pmid">22881277</pub-id>
</mixed-citation></ref><ref id="R12"><label>[12]</label><mixed-citation publication-type="journal"><name><surname>Baygents</surname><given-names>G</given-names></name> and <name><surname>Bani-Yaghoub</surname><given-names>M</given-names></name>. <article-title>Cluster analysis of hemorrhagic disease in missouri&#x02019;s white-tailed deer population: 1980&#x02013;2013</article-title>. <source>BMC Ecology</source>, <volume>18</volume>:<fpage>1</fpage>&#x02013;<lpage>8</lpage>, <year>2018</year>.<pub-id pub-id-type="pmid">29347979</pub-id>
</mixed-citation></ref><ref id="R13"><label>[13]</label><mixed-citation publication-type="journal"><name><surname>Benjamin-Chung</surname><given-names>J</given-names></name> and <name><surname>Reingold</surname><given-names>A</given-names></name>. <article-title>Measuring the success of the us covid-19 vaccine cam-paign&#x02014;it&#x02019;s time to invest in and strengthen immunization information systems</article-title>. <source>American Journal of Public Health</source>, <volume>111</volume>(<issue>6</issue>):<fpage>1078</fpage>&#x02013;<lpage>1080</lpage>, <year>2021</year>.<pub-id pub-id-type="pmid">33600253</pub-id>
</mixed-citation></ref><ref id="R14"><label>[14]</label><mixed-citation publication-type="journal"><name><surname>Besag</surname><given-names>J</given-names></name> and <name><surname>Clifford</surname><given-names>P</given-names></name>. <article-title>Sequential monte carlo p-values</article-title>. <source>Biometrika</source>, <volume>78</volume>:<fpage>301</fpage>&#x02013;<lpage>330</lpage>, <year>1991</year>.</mixed-citation></ref><ref id="R15"><label>[15]</label><mixed-citation publication-type="journal"><name><surname>Block</surname><given-names>R</given-names></name>. <article-title>Software review: scanning for clusters in space and time: a tutorial review of satscan</article-title>. <source>Social Science Computer Review</source>, <volume>25</volume>(<issue>2</issue>):<fpage>272</fpage>&#x02013;<lpage>278</lpage>, <year>2007</year>.</mixed-citation></ref><ref id="R16"><label>[16]</label><mixed-citation publication-type="journal"><name><surname>Bonnet</surname><given-names>E</given-names></name>. <etal/>
<article-title>The covid-19 pandemic in francophone west africa</article-title>. <source>Research Square</source>, <year>2020</year>.</mixed-citation></ref><ref id="R17"><label>[17]</label><mixed-citation publication-type="journal"><name><surname>Caprarelli</surname><given-names>G</given-names></name> and <name><surname>Fletcher</surname><given-names>S</given-names></name>. <article-title>A brief review of spatial analysis concepts and tools used for mapping, containment and risk modelling of infectious diseases and other illnesses</article-title>. <source>Parasitology</source>, <volume>141</volume>(<issue>5</issue>):<fpage>581</fpage>&#x02013;<lpage>601</lpage>, <year>2014</year>.<pub-id pub-id-type="pmid">24476672</pub-id>
</mixed-citation></ref><ref id="R18"><label>[18]</label><mixed-citation publication-type="journal"><name><surname>Christie</surname><given-names>A</given-names></name>, <name><surname>Mbaeyi</surname><given-names>SA</given-names></name>, and <name><surname>Walensky</surname><given-names>RP</given-names></name>. <article-title>Cdc interim recommendations for fully vaccinated people: an important first step</article-title>. <source>JAMA</source>, <volume>325</volume>(<issue>15</issue>):<fpage>1501</fpage>&#x02013;<lpage>1502</lpage>, <year>2021</year>.<pub-id pub-id-type="pmid">33688914</pub-id>
</mixed-citation></ref><ref id="R19"><label>[19]</label><mixed-citation publication-type="journal"><collab>CDC Team VBCI Covid</collab>, <name><surname>Birhane</surname><given-names>Martha</given-names></name>, <name><surname>Bressler</surname><given-names>Samantha</given-names></name>, <name><surname>Chang</surname><given-names>Grace</given-names></name>, <name><surname>Clark</surname><given-names>Teresa</given-names></name>, &#x02026;, and <name><surname>Trujillo</surname><given-names>Alfonso</given-names></name>. <article-title>Covid-19 vaccine breakthrough infections reported to cdc&#x02014;united states, january 1&#x02212;april 30, 2021</article-title>. <source>Morbidity and Mortality Weekly Report</source>, <volume>70</volume>(<issue>21</issue>):<fpage>792</fpage>, <year>2021</year>.<pub-id pub-id-type="pmid">34043615</pub-id>
</mixed-citation></ref><ref id="R20"><label>[20]</label><mixed-citation publication-type="journal"><name><surname>Dietterich</surname><given-names>Thomas G.</given-names></name>. <article-title>Ensemble methods in machine learning</article-title>. <source>Handbook of Learning and Approximation</source>, <volume>1</volume>:<fpage>1</fpage>&#x02013;<lpage>16</lpage>, <year>2000</year>.</mixed-citation></ref><ref id="R21"><label>[21]</label><mixed-citation publication-type="journal"><name><surname>Gautam</surname><given-names>R</given-names></name>, <name><surname>Srinath</surname><given-names>I</given-names></name>, <name><surname>Clavijo</surname><given-names>A</given-names></name>, <name><surname>Szonyi</surname><given-names>B</given-names></name>, <name><surname>Bani-Yaghoub</surname><given-names>M</given-names></name>, <name><surname>Park</surname><given-names>S</given-names></name>, and <name><surname>Ivanek</surname><given-names>R</given-names></name>. <article-title>Identifying areas of high risk of human exposure to coccidioidomycosis in texas using serology data from dogs</article-title>. <source>Zoonoses and Public Health</source>, <volume>60</volume>(<issue>2</issue>):<fpage>174</fpage>&#x02013;<lpage>181</lpage>, <year>2013</year>.<pub-id pub-id-type="pmid">22856539</pub-id>
</mixed-citation></ref><ref id="R22"><label>[22]</label><mixed-citation publication-type="journal"><name><surname>Ghaderpour</surname><given-names>Ebrahim</given-names></name>, <name><surname>Pagiatakis</surname><given-names>Spiros D</given-names></name>, and <name><surname>Hassan</surname><given-names>Quazi K</given-names></name>. <article-title>A survey on change detection and time series analysis with applications</article-title>. <source>Applied Sciences</source>, <volume>11</volume>(<issue>13</issue>):<fpage>6141</fpage>, <year>2021</year>.</mixed-citation></ref><ref id="R23"><label>[23]</label><mixed-citation publication-type="journal"><name><surname>Gostin</surname><given-names>LO</given-names></name>, <name><surname>Cohen</surname><given-names>IG</given-names></name>, and <name><surname>Koplan</surname><given-names>JP</given-names></name>. <article-title>Universal masking in the united states: the role of mandates, health education, and the cdc</article-title>. <source>JAMA</source>, <volume>324</volume>(<issue>9</issue>):<fpage>837</fpage>&#x02013;<lpage>838</lpage>, <year>2020</year>.<pub-id pub-id-type="pmid">32790823</pub-id>
</mixed-citation></ref><ref id="R24"><label>[24]</label><mixed-citation publication-type="journal"><name><surname>Gostin</surname><given-names>LO</given-names></name> and <name><surname>Wiley</surname><given-names>LF</given-names></name>. <article-title>Governmental public health powers during the covid-19 pandemic: stay-at-home orders, business closures, and travel restrictions</article-title>. <source>JAMA</source>, <volume>323</volume>(<issue>21</issue>):<fpage>2137</fpage>&#x02013;<lpage>2138</lpage>, <year>2020</year>.<pub-id pub-id-type="pmid">32239184</pub-id>
</mixed-citation></ref><ref id="R25"><label>[25]</label><mixed-citation publication-type="journal"><name><surname>Horst</surname><given-names>MA</given-names></name> and <name><surname>Coco</surname><given-names>AS</given-names></name>. <article-title>Spread of illnesses in a community using gis</article-title>. <source>Journal of the American Board of Family Medicine</source>, <volume>23</volume>:<fpage>32</fpage>&#x02013;<lpage>41</lpage>, <year>2010</year>.<pub-id pub-id-type="pmid">20051540</pub-id>
</mixed-citation></ref><ref id="R26"><label>[26]</label><mixed-citation publication-type="journal"><name><surname>Huang</surname><given-names>L</given-names></name>, <name><surname>Huang</surname><given-names>L</given-names></name>, <name><surname>Tiwari</surname><given-names>R</given-names></name>, <name><surname>Zuo</surname><given-names>J</given-names></name>, <name><surname>Kulldorff</surname><given-names>M</given-names></name>, and <name><surname>Feuer</surname><given-names>E</given-names></name>. <article-title>Weighted normal spatial scan statistic for heterogeneous population data</article-title>. <source>Journal of the American Statistical Association</source>, <volume>104</volume>:<fpage>886</fpage>&#x02013;<lpage>898</lpage>, <year>2009</year>.</mixed-citation></ref><ref id="R27"><label>[27]</label><mixed-citation publication-type="journal"><name><surname>Huang</surname><given-names>SS</given-names></name>
<etal/>
<article-title>Automated outbreak detection in hospitals</article-title>. <source>PLoS Medicine</source>, <volume>7</volume>, <fpage>2010</fpage>.</mixed-citation></ref><ref id="R28"><label>[28]</label><mixed-citation publication-type="book"><name><surname>Ji</surname><given-names>D</given-names></name>, <name><surname>Dunn</surname><given-names>E</given-names></name>, and <name><surname>Frahm</surname><given-names>J</given-names></name>. <part-title>Spatio-temporally consistent correspondence for dense dynamic scene modeling</part-title>. In <source>Computer Vision&#x02212;ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11&#x02013;14, 2016, Proceedings, Part VI 14</source>, pages <fpage>3</fpage>&#x02013;<lpage>18</lpage>. <publisher-name>Springer International Publishing</publisher-name>, <year>2016</year>.</mixed-citation></ref><ref id="R29"><label>[29]</label><mixed-citation publication-type="book"><name><surname>Kemeny</surname><given-names>JG</given-names></name> and <name><surname>Snell</surname><given-names>LJ</given-names></name>. <source>Finite Markov Chains</source>. <publisher-name>Springer</publisher-name>, <publisher-loc>New York</publisher-loc>, <year>1976</year>.</mixed-citation></ref><ref id="R30"><label>[30]</label><mixed-citation publication-type="journal"><name><surname>Kleinman</surname><given-names>K</given-names></name>
<etal/>
<article-title>Space-time scan statistic in syndromic surveillance</article-title>. <source>Epidemiology and Infection</source>, <volume>133</volume>:<fpage>409</fpage>&#x02013;<lpage>419</lpage>, <year>2005</year>.<pub-id pub-id-type="pmid">15962547</pub-id>
</mixed-citation></ref><ref id="R31"><label>[31]</label><mixed-citation publication-type="journal"><name><surname>Kulldorff</surname><given-names>M</given-names></name>. <article-title>A spatial scan statistic</article-title>. <source>Communications in Statistics - Theory and Methods</source>, <volume>26</volume>(<issue>6</issue>):<fpage>1481</fpage>&#x02013;<lpage>1496</lpage>, <year>1997</year>.</mixed-citation></ref><ref id="R32"><label>[32]</label><mixed-citation publication-type="journal"><name><surname>Kulldorff</surname><given-names>M</given-names></name>, <name><surname>Huang</surname><given-names>L</given-names></name>, <name><surname>Pickle</surname><given-names>L</given-names></name>, and <name><surname>Duczmal</surname><given-names>L</given-names></name>. <article-title>An elliptic spatial scan statistic</article-title>. <source>Statistics in Medicine</source>, <volume>25</volume>:<fpage>3929</fpage>&#x02013;<lpage>3943</lpage>, <year>2006</year>.<pub-id pub-id-type="pmid">16435334</pub-id>
</mixed-citation></ref><ref id="R33"><label>[33]</label><mixed-citation publication-type="journal"><name><surname>Kulldorff</surname><given-names>M</given-names></name>, <name><surname>Mostashari</surname><given-names>F</given-names></name>, <name><surname>Duczmal</surname><given-names>L</given-names></name>, <name><surname>Yih</surname><given-names>K</given-names></name>, <name><surname>Kleinman</surname><given-names>K</given-names></name>, and <name><surname>Platt</surname><given-names>R</given-names></name>. <article-title>Multivariate spatial scan statistics for disease surveillance</article-title>. <source>Statistics in Medicine</source>, <volume>26</volume>:<fpage>1824</fpage>&#x02013;<lpage>1833</lpage>, <year>2007</year>.<pub-id pub-id-type="pmid">17216592</pub-id>
</mixed-citation></ref><ref id="R34"><label>[34]</label><mixed-citation publication-type="journal"><name><surname>Leveau</surname><given-names>CM</given-names></name>. <article-title>Variations of covid-19 mortality in buenos aires</article-title>. <source>Scielo Preprints</source>, <year>2020</year>.</mixed-citation></ref><ref id="R35"><label>[35]</label><mixed-citation publication-type="journal"><name><surname>Murthy</surname><given-names>N</given-names></name>. <article-title>Advisory committee on immunization practices recommended immunization schedule for adults aged 19 years or older&#x02014;united states, 2022</article-title>. <source>MMWR. Morbidity and Mortality Weekly Report</source>, <volume>71</volume>, <fpage>2022</fpage>.</mixed-citation></ref><ref id="R36"><label>[36]</label><mixed-citation publication-type="journal"><name><surname>Natale</surname><given-names>A</given-names></name>
<etal/>
<article-title>Detection of antimicrobial resistance clusters</article-title>. <source>Eurosurveillance</source>, <volume>22</volume>(<issue>30484</issue>), <fpage>2017</fpage>.</mixed-citation></ref><ref id="R37"><label>[37]</label><mixed-citation publication-type="journal"><name><surname>Raamkumar</surname><given-names>AS</given-names></name>, <name><surname>Tan</surname><given-names>SG</given-names></name>, and <name><surname>Wee</surname><given-names>HL</given-names></name>. <article-title>Measuring the outreach efforts of public health authorities and the public response on facebook during the covid-19 pandemic in early 2020: cross-country comparison</article-title>. <source>Journal of Medical Internet Research</source>, <volume>22</volume>(<issue>5</issue>):<fpage>e19334</fpage>, <year>2020</year>.<pub-id pub-id-type="pmid">32401219</pub-id>
</mixed-citation></ref><ref id="R38"><label>[38]</label><mixed-citation publication-type="book"><name><surname>Rohatgi</surname><given-names>VK</given-names></name>. <source>An Introduction to Probability Theory and Mathematical Statistics</source>. <publisher-name>Wiley</publisher-name>, <publisher-loc>New York</publisher-loc>, <year>1976</year>.</mixed-citation></ref><ref id="R39"><label>[39]</label><mixed-citation publication-type="journal"><name><surname>Sen-Crowe</surname><given-names>B</given-names></name>, <name><surname>McKenney</surname><given-names>M</given-names></name>, and <name><surname>Elkbuli</surname><given-names>A</given-names></name>. <article-title>Social distancing during the covid-19 pandemic: Staying home save lives</article-title>. <source>The American Journal of Emergency Medicine</source>, <volume>38</volume>(<issue>7</issue>):<fpage>1519</fpage>, <year>2020</year>.<pub-id pub-id-type="pmid">32305155</pub-id>
</mixed-citation></ref><ref id="R40"><label>[40]</label><mixed-citation publication-type="journal"><name><surname>Shekhar</surname><given-names>R</given-names></name>, <name><surname>Garg</surname><given-names>I</given-names></name>, <name><surname>Pal</surname><given-names>S</given-names></name>, <name><surname>Kottewar</surname><given-names>S</given-names></name>, and <name><surname>Sheikh</surname><given-names>AB</given-names></name>. <article-title>Covid-19 vaccine booster: to boost or not to boost</article-title>. <source>Infectious Disease Reports</source>, <volume>13</volume>(<issue>4</issue>):<fpage>924</fpage>&#x02013;<lpage>929</lpage>, <year>2021</year>.<pub-id pub-id-type="pmid">34842753</pub-id>
</mixed-citation></ref><ref id="R41"><label>[41]</label><mixed-citation publication-type="journal"><name><surname>Silva</surname><given-names>I</given-names></name>, <name><surname>Assun&#x000e7;&#x000e3;o</surname><given-names>RM</given-names></name>, and <name><surname>Costa</surname><given-names>M</given-names></name>. <article-title>Power of the sequential monte carlo test</article-title>. <source>Sequential Analysis</source>, <volume>28</volume>:<fpage>163</fpage>&#x02013;<lpage>174</lpage>, <year>2009</year>.</mixed-citation></ref><ref id="R42"><label>[42]</label><mixed-citation publication-type="journal"><name><surname>Souris</surname><given-names>M</given-names></name>
<etal/>
<article-title>Poultry farm vulnerability and avian influenza risk</article-title>. <source>International Journal of Environmental Research and Public Health</source>, <volume>11</volume>:<fpage>934</fpage>&#x02013;<lpage>951</lpage>, <year>2014</year>.<pub-id pub-id-type="pmid">24413705</pub-id>
</mixed-citation></ref><ref id="R43"><label>[43]</label><mixed-citation publication-type="journal"><name><surname>Tchole</surname><given-names>AI</given-names></name>
<etal/>
<article-title>Epidemic and control of covid-19 in niger</article-title>. <source>Journal of Global Health</source>, <volume>10</volume>, <fpage>2020</fpage>.</mixed-citation></ref><ref id="R44"><label>[44]</label><mixed-citation publication-type="webpage"><collab>The New York Times</collab>. <source>Coronavirus (covid-19) data in the united states</source>, <year>2021</year>. <comment>Retrieved</comment>
<date-in-citation>June 21, 2023</date-in-citation>, <comment>from <ext-link xlink:href="https://github.com/nytimes/covid-19-data" ext-link-type="uri">https://github.com/nytimes/covid-19-data</ext-link>.</comment></mixed-citation></ref><ref id="R45"><label>[45]</label><mixed-citation publication-type="journal"><name><surname>Tran</surname><given-names>Thao</given-names></name>, <name><surname>Bani-Yaghoub</surname><given-names>Majid</given-names></name>, and <name><surname>DeLisle</surname><given-names>James R</given-names></name>. <article-title>Non-emergency responses in the 311 system during the early stage of the covid-19 pandemic: a case study of kansas city</article-title>. <source>Disaster Prevention and Resilience</source>, <year>2023</year>.</mixed-citation></ref><ref id="R46"><label>[46]</label><mixed-citation publication-type="journal"><name><surname>Whittaker</surname><given-names>JA</given-names></name>, <name><surname>Rekab</surname><given-names>K</given-names></name>, and <name><surname>Thomason</surname><given-names>MG</given-names></name>. <article-title>A markov chain model for predicting the reliability of multi-build software</article-title>. <source>Information and Software Technology</source>, <volume>42</volume>(<issue>12</issue>):<fpage>889</fpage>&#x02013;<lpage>894</lpage>, <year>2000</year>.</mixed-citation></ref><ref id="R47"><label>[47]</label><mixed-citation publication-type="journal"><name><surname>Whittaker</surname><given-names>JA</given-names></name> and <name><surname>Thomason</surname><given-names>MG</given-names></name>. <article-title>Markov chain model for statistical software testing</article-title>. <source>IEEE Transactions on Software Engineering</source>, <volume>20</volume>(<issue>10</issue>):<fpage>812</fpage>&#x02013;<lpage>824</lpage>, <year>1994</year>.</mixed-citation></ref><ref id="R48"><label>[48]</label><mixed-citation publication-type="journal"><name><surname>Xiang</surname><given-names>R</given-names></name>, <name><surname>Chen</surname><given-names>L</given-names></name>, and <name><surname>Han</surname><given-names>X</given-names></name>. <article-title>Hybrid markov chain and neural network models for spatial prediction</article-title>. <source>IEEE Transactions on Neural Networks and Learning Systems</source>, <volume>33</volume>(<issue>8</issue>):<fpage>3623</fpage>&#x02013;<lpage>3635</lpage>, <year>2022</year>.</mixed-citation></ref><ref id="R49"><label>[49]</label><mixed-citation publication-type="journal"><name><surname>Yang</surname><given-names>H</given-names></name>, <name><surname>Zou</surname><given-names>Y</given-names></name>, <name><surname>Wang</surname><given-names>Z</given-names></name>, and <name><surname>Wu</surname><given-names>B</given-names></name>. <article-title>A hybrid method for short-term freeway travel time prediction based on wavelet neural network and markov chain</article-title>. <source>Canadian Journal of Civil Engineering</source>, <volume>45</volume>(<issue>2</issue>):<fpage>77</fpage>&#x02013;<lpage>86</lpage>, <year>2018</year>.</mixed-citation></ref><ref id="R50"><label>[50]</label><mixed-citation publication-type="journal"><name><surname>Zhang</surname><given-names>Y</given-names></name>
<etal/>
<article-title>Avian influenza a (h7n9) cluster analysis</article-title>. <source>International Journal of Environmental Research and Public Health</source>, <volume>12</volume>:<fpage>816</fpage>&#x02013;<lpage>828</lpage>, <year>2015</year>.<pub-id pub-id-type="pmid">25599373</pub-id>
</mixed-citation></ref></ref-list></back><floats-group><fig position="float" id="F1"><label>Figure 1:</label><caption><p id="P48">Significant spatial clusters of U.S. COVID-19 mortality. High-risk clusters (relative risk &#x0003e; 1) indicated by red circles, and low-risk clusters (relative risk &#x0003c; 1) shown in blue. (a) the overall significant spatial clusters; (b) the distribution of high-risk clusters was predominantly concentrated in the southern U.S. (c)-(h) notable temporal changes observed across the intervals representing the effects of vaccination efforts (d); the rise of the Delta variant (e), and the increasing significance of the Omicron variant (g), a substantial reduction of high-risk clusters by the end of the study period in panel (h).</p></caption><graphic xlink:href="nihms-2053985-f0001" position="float"/></fig><fig position="float" id="F2"><label>Figure 2:</label><caption><p id="P49">Time series of mortality data and spatial clusters (a) Time series of U.S. COVID-19 mortality data across seven defined time intervals: (1) May 24, 2020 &#x02212; September 13, 2020; (2) September 13, 2020 &#x02212; March 14, 2021; (3) March 14, 2021 &#x02212; June 13, 2021; (4) June 13, 2021 &#x02212; October 31, 2021; (5) October 31, 2021 &#x02212; March 13, 2022; (6) March 13, 2022 &#x02212; October 16, 2022; (7) October 16, 2022 &#x02212; March 12, 2023. The figure illustrates trends in mortality over these periods, highlighting significant fluctuations in COVID-19-related deaths; (b) the percentage of U.S. area covered by significant high-risk (i.e. relative risk greater than 1) and low-risk (relative risk less than 1) clusters for each interval of 1 through 7; (c) persistence of high-risk and low-risk clusters from the time interval i to interval i+1; (d) percent transition of high-risk to low-risk clusters and vice versa from the time interval i to interval i + 1.</p></caption><graphic xlink:href="nihms-2053985-f0002" position="float"/></fig><fig position="float" id="F3"><label>Figure 3:</label><caption><p id="P50">Accuracy of identified spatial clusters. (a) observed/expected mortality versus estimated relative risk for intervals 1&#x02013;7; (b) box plots of residuals, (c) plot of mortality data within each cluster, (d) plot of estimated relative risks of spatial clusters associated with each interval.</p></caption><graphic xlink:href="nihms-2053985-f0003" position="float"/></fig><fig position="float" id="F4"><label>Figure 4:</label><caption><p id="P51">Optimal value and the accuracy of predictions for interval 7. (a) Estimating the optimal value of <inline-formula><mml:math id="M8" display="inline"><mml:mi>&#x003b1;</mml:mi></mml:math></inline-formula> based on the distance between the estimated and observed relative risk values. (b) log scale observed versus predicted by the proposed Markov chain model for interval 7. The proposed model predictions are superior to those of Exponential Smoothing (Mean Squared Error 164.8 versus 258.7)</p></caption><graphic xlink:href="nihms-2053985-f0004" position="float"/></fig><table-wrap position="float" id="T1"><label>Table 1:</label><caption><p id="P52">US COVID-19 Control and Preventive Measures by Time Interval</p></caption><table frame="hsides" rules="groups"><colgroup span="1"><col align="left" valign="middle" span="1"/><col align="left" valign="middle" span="1"/><col align="left" valign="middle" span="1"/><col align="left" valign="middle" span="1"/></colgroup><thead><tr><th align="left" valign="top" rowspan="1" colspan="1">Interval</th><th align="left" valign="top" rowspan="1" colspan="1">Variants</th><th align="left" valign="top" rowspan="1" colspan="1">Control and Preventive Measures</th><th align="left" valign="top" rowspan="1" colspan="1">Ref.</th></tr></thead><tbody><tr><td align="left" valign="top" rowspan="1" colspan="1">1. 5/24/2020 &#x02013; 9/13/2020</td><td align="left" valign="top" rowspan="1" colspan="1">Original strain (Wuhan)</td><td align="left" valign="top" rowspan="1" colspan="1">Social distancing measures, use of face masks, limitations on gatherings, travel restrictions</td><td align="left" valign="top" rowspan="1" colspan="1">[<xref rid="R39" ref-type="bibr">39</xref>, <xref rid="R24" ref-type="bibr">24</xref>]</td></tr><tr><td align="left" valign="top" rowspan="1" colspan="1">2. 9/13/2020 &#x02013; 3/14/2021</td><td align="left" valign="top" rowspan="1" colspan="1">D614G variant</td><td align="left" valign="top" rowspan="1" colspan="1">Expanded mask mandates, introduction of COVID-19 vaccines (Pfizer, Moderna), enhanced testing and contact tracing</td><td align="left" valign="top" rowspan="1" colspan="1">[<xref rid="R23" ref-type="bibr">23</xref>, <xref rid="R13" ref-type="bibr">13</xref>]</td></tr><tr><td align="left" valign="top" rowspan="1" colspan="1">3. 3/14/2021 &#x02013; 6/13/202</td><td align="left" valign="top" rowspan="1" colspan="1">Alpha variant (B.1.1.7)</td><td align="left" valign="top" rowspan="1" colspan="1">Vaccination campaigns ramped up, continued mask-wearing in high-transmission areas</td><td align="left" valign="top" rowspan="1" colspan="1">[<xref rid="R19" ref-type="bibr">19</xref>, <xref rid="R18" ref-type="bibr">18</xref>]</td></tr><tr><td align="left" valign="top" rowspan="1" colspan="1">4. 6/13/2021 &#x02013; 10/31/2021</td><td align="left" valign="top" rowspan="1" colspan="1">Alpha variant predominant</td><td align="left" valign="top" rowspan="1" colspan="1">CDC recommendations for vaccinated individuals, return to in-person learning in colleges and schools, ongoing vaccination efforts</td><td align="left" valign="top" rowspan="1" colspan="1">[<xref rid="R18" ref-type="bibr">18</xref>, <xref rid="R35" ref-type="bibr">35</xref>]</td></tr><tr><td align="left" valign="top" rowspan="1" colspan="1">5. 10/31/2021 &#x02013; 3/13/2022</td><td align="left" valign="top" rowspan="1" colspan="1">Delta variant (B.1.617.2)</td><td align="left" valign="top" rowspan="1" colspan="1">Booster shots recommended, reinstated mask mandates in certain areas, increased indoor venue capacity restrictions</td><td align="left" valign="top" rowspan="1" colspan="1">[<xref rid="R35" ref-type="bibr">35</xref>, <xref rid="R40" ref-type="bibr">40</xref>]</td></tr><tr><td align="left" valign="top" rowspan="1" colspan="1">6. 3/13/2022 &#x02013; 10/16/2022</td><td align="left" valign="top" rowspan="1" colspan="1">Omicron variant (B.1.1.529)</td><td align="left" valign="top" rowspan="1" colspan="1">Vaccination updates including boosters, recommendations for masks in crowded indoor settings, continued public health surveillance</td><td align="left" valign="top" rowspan="1" colspan="1">[<xref rid="R39" ref-type="bibr">39</xref>, <xref rid="R35" ref-type="bibr">35</xref>]</td></tr><tr><td align="left" valign="top" rowspan="1" colspan="1">7. 10/16/2022 &#x02013; 3/12/2023</td><td align="left" valign="top" rowspan="1" colspan="1">Omicron subvariants (BA.1, BA.2)</td><td align="left" valign="top" rowspan="1" colspan="1">Focus on public health awareness campaigns, emphasis on personal responsibility regarding health measures</td><td align="left" valign="top" rowspan="1" colspan="1">[<xref rid="R7" ref-type="bibr">7</xref>, <xref rid="R37" ref-type="bibr">37</xref>]</td></tr></tbody></table></table-wrap><table-wrap position="float" id="T2"><label>Table 2:</label><caption><p id="P53">Summary of spatial high-risk and low-risk COVID-19 mortality clusters across intervals 1&#x02013;7, showing the number of clusters and the percentage of U.S. area covered by each.</p></caption><table frame="hsides" rules="groups"><colgroup span="1"><col align="left" valign="middle" span="1"/><col align="left" valign="middle" span="1"/><col align="left" valign="middle" span="1"/><col align="left" valign="middle" span="1"/><col align="left" valign="middle" span="1"/><col align="left" valign="middle" span="1"/><col align="left" valign="middle" span="1"/></colgroup><thead><tr><th align="center" valign="middle" rowspan="1" colspan="1">Interval</th><th align="center" valign="middle" rowspan="1" colspan="1">High Risk</th><th align="center" valign="middle" rowspan="1" colspan="1">Low Risk</th><th align="center" valign="middle" rowspan="1" colspan="1">Area of High Risk</th><th align="center" valign="middle" rowspan="1" colspan="1">Area of Low Risk</th><th align="center" valign="middle" rowspan="1" colspan="1">Overlap Area</th><th align="center" valign="middle" rowspan="1" colspan="1">Total Area</th></tr></thead><tbody><tr><td align="center" valign="middle" rowspan="1" colspan="1">1</td><td align="center" valign="middle" rowspan="1" colspan="1">25</td><td align="center" valign="middle" rowspan="1" colspan="1">38</td><td align="center" valign="middle" rowspan="1" colspan="1">1,170,531.55 (12.80%)</td><td align="center" valign="middle" rowspan="1" colspan="1">4,457,499.54 (48.73%)</td><td align="center" valign="middle" rowspan="1" colspan="1">4,830.30 (0.05%)</td><td align="center" valign="middle" rowspan="1" colspan="1">5,623,200.79 (61.47%)</td></tr><tr><td align="center" valign="middle" rowspan="1" colspan="1">2</td><td align="center" valign="middle" rowspan="1" colspan="1">38</td><td align="center" valign="middle" rowspan="1" colspan="1">34</td><td align="center" valign="middle" rowspan="1" colspan="1">3,661,450.98 (40.03%)</td><td align="center" valign="middle" rowspan="1" colspan="1">2,065,006.01 (22.57%)</td><td align="center" valign="middle" rowspan="1" colspan="1">7,781.26 (0.09%)</td><td align="center" valign="middle" rowspan="1" colspan="1">5,718,675.73 (62.52%)</td></tr><tr><td align="center" valign="middle" rowspan="1" colspan="1">3</td><td align="center" valign="middle" rowspan="1" colspan="1">40</td><td align="center" valign="middle" rowspan="1" colspan="1">39</td><td align="center" valign="middle" rowspan="1" colspan="1">1,447,800.92 (15.83%)</td><td align="center" valign="middle" rowspan="1" colspan="1">3,189,789.31 (34.87%)</td><td align="center" valign="middle" rowspan="1" colspan="1">5,915.07 (0.06%)</td><td align="center" valign="middle" rowspan="1" colspan="1">4,631,675.16 (50.63%)</td></tr><tr><td align="center" valign="middle" rowspan="1" colspan="1">4</td><td align="center" valign="middle" rowspan="1" colspan="1">30</td><td align="center" valign="middle" rowspan="1" colspan="1">25</td><td align="center" valign="middle" rowspan="1" colspan="1">4,667,511.42 (51.03%)</td><td align="center" valign="middle" rowspan="1" colspan="1">1,733,777.69 (18.95%)</td><td align="center" valign="middle" rowspan="1" colspan="1">24,583.89 (0.27%)</td><td align="center" valign="middle" rowspan="1" colspan="1">6,376,705.22 (69.71%)</td></tr><tr><td align="center" valign="middle" rowspan="1" colspan="1">5</td><td align="center" valign="middle" rowspan="1" colspan="1">29</td><td align="center" valign="middle" rowspan="1" colspan="1">36</td><td align="center" valign="middle" rowspan="1" colspan="1">2,928,563.46 (32.02%)</td><td align="center" valign="middle" rowspan="1" colspan="1">2,024,862.64 (22.14%)</td><td align="center" valign="middle" rowspan="1" colspan="1">14,785.95 (0.16%)</td><td align="center" valign="middle" rowspan="1" colspan="1">4,938,640.15 (53.71%)</td></tr><tr><td align="center" valign="middle" rowspan="1" colspan="1">6</td><td align="center" valign="middle" rowspan="1" colspan="1">29</td><td align="center" valign="middle" rowspan="1" colspan="1">25</td><td align="center" valign="middle" rowspan="1" colspan="1">2,760,890.45 (30.18%)</td><td align="center" valign="middle" rowspan="1" colspan="1">2,214,642.66 (24.21%)</td><td align="center" valign="middle" rowspan="1" colspan="1">3,259.16 (0.04%)</td><td align="center" valign="middle" rowspan="1" colspan="1">4,972,273.95 (54.36%)</td></tr><tr><td align="center" valign="middle" rowspan="1" colspan="1">7</td><td align="center" valign="middle" rowspan="1" colspan="1">28</td><td align="center" valign="middle" rowspan="1" colspan="1">30</td><td align="center" valign="middle" rowspan="1" colspan="1">1,556,809.44 (17.02%)</td><td align="center" valign="middle" rowspan="1" colspan="1">2,316,998.49 (25.33%)</td><td align="center" valign="middle" rowspan="1" colspan="1">5,328.31 (0.06%)</td><td align="center" valign="middle" rowspan="1" colspan="1">3,868,479.62 (42.29%)</td></tr></tbody></table></table-wrap><table-wrap position="float" id="T3"><label>Table 3:</label><caption><p id="P54">Transition and persistence rates of high-risk and low-risk COVID-19 mortality clusters across consecutive intervals.</p></caption><table frame="hsides" rules="groups"><colgroup span="1"><col align="left" valign="middle" span="1"/><col align="left" valign="middle" span="1"/><col align="left" valign="middle" span="1"/><col align="left" valign="middle" span="1"/><col align="left" valign="middle" span="1"/></colgroup><thead><tr><th align="center" valign="middle" rowspan="1" colspan="1">Intervals</th><th align="center" valign="middle" rowspan="1" colspan="1">High Risk Overlap</th><th align="center" valign="middle" rowspan="1" colspan="1">Low Risk Overlap</th><th align="center" valign="middle" rowspan="1" colspan="1">High-Low Transition</th><th align="center" valign="middle" rowspan="1" colspan="1">Low-High Transition</th></tr></thead><tbody><tr><td align="center" valign="middle" rowspan="1" colspan="1">1 &#x02192; 2</td><td align="center" valign="middle" rowspan="1" colspan="1">630,606.00 53.87%, 17.22%</td><td align="center" valign="middle" rowspan="1" colspan="1">1,451,535.93 32.56%, 70.29%</td><td align="center" valign="middle" rowspan="1" colspan="1">93,533.80 7.99%, 4.53%</td><td align="center" valign="middle" rowspan="1" colspan="1">11,311,032.80 29.41%, 35.81%</td></tr><tr><td align="center" valign="middle" rowspan="1" colspan="1">2 &#x02192; 3</td><td align="center" valign="middle" rowspan="1" colspan="1">727,798.55 19.88%, 50.27%</td><td align="center" valign="middle" rowspan="1" colspan="1">1,255,175.77 60.78%, 39.35%</td><td align="center" valign="middle" rowspan="1" colspan="1">979,045.90 26.74%, 30.69%</td><td align="center" valign="middle" rowspan="1" colspan="1">100,420.49 4.86%, 6.94%</td></tr><tr><td align="center" valign="middle" rowspan="1" colspan="1">3 &#x02192; 4</td><td align="center" valign="middle" rowspan="1" colspan="1">583,839.38 40.33%, 17.27%</td><td align="center" valign="middle" rowspan="1" colspan="1">696,469.49 21.83%, 40.17%</td><td align="center" valign="middle" rowspan="1" colspan="1">168,040.31 11.61%, 9.69%</td><td align="center" valign="middle" rowspan="1" colspan="1">1,042,772.11 32.69%, 30.85%</td></tr><tr><td align="center" valign="middle" rowspan="1" colspan="1">4 &#x02192; 5</td><td align="center" valign="middle" rowspan="1" colspan="1">1,146,392.24 33.92%, 39.15%</td><td align="center" valign="middle" rowspan="1" colspan="1">667,377.54 38.49%, 32.96%</td><td align="center" valign="middle" rowspan="1" colspan="1">476,609.88 14.10%, 23.54%</td><td align="center" valign="middle" rowspan="1" colspan="1">1,391,524.37 21.28%, 12.60%</td></tr><tr><td align="center" valign="middle" rowspan="1" colspan="1">5 &#x02192; 6</td><td align="center" valign="middle" rowspan="1" colspan="1">1,178,061.49 40.23%, 42.67%</td><td align="center" valign="middle" rowspan="1" colspan="1">811,950.67 40.09%, 36.66%</td><td align="center" valign="middle" rowspan="1" colspan="1">337,960.92 11.54%, 15.26%</td><td align="center" valign="middle" rowspan="1" colspan="1">335,357.67 16.56%, 12.15%</td></tr><tr><td align="center" valign="middle" rowspan="1" colspan="1">6 &#x02192; 7</td><td align="center" valign="middle" rowspan="1" colspan="1">802,151.60 29.05%, 51.53%</td><td align="center" valign="middle" rowspan="1" colspan="1">895,447.97 40.43%, 38.65%</td><td align="center" valign="middle" rowspan="1" colspan="1">315,871.16 11.44%, 13.63%</td><td align="center" valign="middle" rowspan="1" colspan="1">257,680.39 11.64%, 16.55%</td></tr></tbody></table></table-wrap></floats-group></article>