<?xml version="1.0" encoding="utf-8" ?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.2 20190208//EN"
                  "JATS-archivearticle1.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="1.2" article-type="other">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">jrepstat</journal-id>
<journal-title-group>
<journal-title>Journal of Reproducible Statistics</journal-title>
</journal-title-group>
<publisher>
<publisher-name>FAIR Press</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<title-group>
<article-title>Normalizing RNA concentrations in wastewater using a control without public health data</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<contrib-id contrib-id-type="orcid">https://orcid.org/0000-0002-7308-3084</contrib-id>
<name>
<surname>Faraway</surname>
<given-names>Julian</given-names>
</name>
<aff>Department of Mathematical Sciences, University of Bath</aff>
</contrib>
</contrib-group>
<pub-date date-type="pub" publication-format="electronic">
<year>2026</year>
</pub-date>
<volume>1</volume>
<elocation-id>1</elocation-id>
<history>
<date date-type="received" iso-8601-date="2026-08-28"><day>28</day><month>08</month><year>2026</year></date>
<date date-type="accepted" iso-8601-date="2026-09-17"><day>17</day><month>09</month><year>2026</year></date>
</history>
<permissions>
<license license-type="open-access" xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This is an open access article distributed under the terms of the Creative Commons Attribution 4.0 International License (CC-BY 4.0), which permits unrestricted use, distribution and reproduction in any medium provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="html" xlink:href="https://jrepstat.fairpressjournals.com/articles/normalizing-rna-concentrations-in-wastewater-using-a-control-without-public-health-data/"/>
<abstract>
<p>PCR-based measurement of RNA targets in wastewater epidemiology is subject to substantial variation, resulting in considerable uncertainty in disease prevalence estimates. Often, control RNA targets, which are either naturally present or introduced into wastewater samples, are used to normalize the estimated RNA concentration for the primary target. We consider methods for normalization that do not rely on the availability of disease prevalence from public health sources. We demonstrate why direct normalization is unlikely to be effective and may make variation worse. We propose a Kalman filter method for normalization and find that, given the quality of available controls, it offers only limited benefit. We discuss other methods of reducing variation.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Wastewater-based surveillance</kwd>
<kwd>Normalization</kwd>
<kwd>PCR</kwd>
<kwd>Controls</kwd>
<kwd>Pepper Mild Mottle Virus</kwd>
<kwd>Bovine Coronavirus</kwd>
</kwd-group>
<custom-meta-group>
<custom-meta>
<meta-name>data-availability-route</meta-name>
<meta-value>open</meta-value>
</custom-meta>
<custom-meta>
<meta-name>founding-article</meta-name>
<meta-value>true</meta-value>
</custom-meta>
</custom-meta-group>
</article-meta>
</front>
<body>
<p>Department of Mathematical Sciences<sup>1</sup>, University of Bath,
Bath, BA2 7AY, United Kingdom</p>
<sec id="abstract">
  <title>Abstract</title>
  <p>PCR-based measurement of RNA targets in wastewater epidemiology is
  subject to substantial variation, resulting in considerable
  uncertainty in disease prevalence estimates. Often, control RNA
  targets, which are either naturally present or introduced into
  wastewater samples, are used to normalize the estimated RNA
  concentration for the primary target. We consider methods for
  normalization that do not rely on the availability of disease
  prevalence from public health sources. We demonstrate why direct
  normalization is unlikely to be effective and may make variation
  worse. We propose a Kalman filter method for normalization and find
  that, given the quality of available controls, it offers only limited
  benefit. We discuss other methods of reducing variation.</p>
  <p><italic>Keywords:</italic> <italic>Wastewater-based surveillance,
  Normalization, PCR, Controls, Pepper Mild Mottle Virus, Bovine
  Coronavirus</italic></p>
</sec>
<sec id="introduction">
  <title>Introduction</title>
  <p>In wastewater-based epidemiology (WBE), the concentration of a
  target substance is measured at a wastewater treatment plant (WWTP) to
  learn about disease prevalence or lifestyle habits in the population
  in the catchment area of the WWTP. For example, the concentration of
  SARS-CoV2 RNA might be used to predict the prevalence of Covid-19 in
  the catchment area. A systematic review of the large literature on
  wastewater surveillance using SARS-CoV2 RNA may be found in (Deák,
  Lupu and Prangate, 2026). RNA targets associated with influenza,
  norovirus and many other diseases are now measured around the
  world.</p>
  <p>We need a precise estimate of the concentration of the RNA in a
  sample because this is assumed proportional to the number of infected
  individuals in a WWTP catchment area. Unfortunately, this estimate is
  subject to three substantive sources of variation. The first source is
  human-related – the rate at which RNA is excreted over the course of
  the infection, movement of people in and out of the catchment area
  etc. – we do not address this here. The second source concerns what
  happens to the RNA in the sewer between the human and the influent
  point at the WWTP where the sample is drawn and transported to the
  laboratory. The third source of variation concerns the preparation,
  extraction and the PCR measurement that occurs in the laboratory. The
  methods we discuss here relate mostly to laboratory variation and
  partly to the sewer variation.</p>
  <p>The exponential amplification through the cycles of a PCR
  measurement unavoidably leads to substantial variation in the
  estimated concentration of RNA in a sample. A common method for
  reducing this variation and other sources of variation in the process
  involves the measurement of a control target which is subject to
  similar perturbations during the process. There are two kinds of
  control. One relates to a substance which we assume is excreted at a
  constant rate by all individuals in the catchment area. Commonly used
  examples include pepper mild mottle virus (PMMoV) and CrAssphage.
  These controls will address variation due to the sewer as well as the
  laboratory. Another type of control is a known fixed quantity of RNA
  which is added into the sample after collection. Examples of this
  include porcine reproductive and respiratory syndrome virus and bovine
  coronavirus (BCoV). Any observed variation in the measurement of these
  controls can be attributed to at least some parts of the laboratory
  measurement process. Since we might reasonably assume that the control
  is subject to many of the same perturbations that apply to the target
  and will be correlated, we can attempt to use this information using
  normalization to reduce variation in the target. This type of
  adjustment is commonplace in applications beyond WBE but there are
  considerations particular to this situation. A systematic review of
  normalization methods used in WBE is found in (Ahmed <italic>et
  al.</italic>, 2026).</p>
  <p>Another type of normalization uses information about the sample,
  such as the ammonia content of the wastewater, that is not directly
  related to the PCR measurement. This normalization is aimed more at
  reducing bias than variation. In contrast, the controls are assumed to
  be approximately constant so normalization using these is focused on
  variance reduction. Our focus here is variance reduction using
  controls.</p>
  <p>We start with an examination of the most common way of using the
  control information to adjust the estimate of the target and explain
  why this method is unlikely to work in the circumstances that are
  likely to apply. The fundamental objection is that the measurement of
  the control, which uses the same technology as the measurement of the
  target, is too variable.</p>
  <p>We present a superior method of making the adjustment using the
  Kalman filter. The methodology is demonstrated on some WBE data from
  California. We also simulate data with properties like the observed
  data but with some changed features. These simulations illustrate
  different scenarios where this method of adjustment is more or less
  likely to be effective.</p>
  <p>Many researchers have used control adjustment where prevalence
  information about the disease for the targeted pathogen is available.
  In these situations, one can propose a model of the general form:</p>
  <p><italic>Prevalence = f(target RNA, control RNA, other
  variables)</italic></p>
  <p>There is a wide selection of statistical and machine learning
  models for estimating the function f() that would take advantage of
  the additional information supplied by the control and other
  variables. Most examples of this approach have prevalence of COVID19
  as the response where this is measured using public health monitoring
  independent of WBE. For recent examples, see (Darling <italic>et
  al.</italic>, 2025; Lamm <italic>et al.</italic>, 2025; Verani
  <italic>et al.</italic>, 2025). This modelling approach usually
  indicates some predictive value associated with some control variates.
  Virtually all published research shows some success in predicting the
  prevalence using the target RNA. While this is encouraging for the
  utility of WBE, there are some drawbacks.</p>
  <p>Many published articles start with the claim that WBE has
  advantages over standard public health indicators. These articles then
  proceed to use prevalence data obtained from public health sources to
  develop their models. For WBE to be truly useful, it must function
  without requiring public health data. One might argue that models
  calibrated using prevalence data can subsequently be used to predict
  prevalence in new situations. Successful extrapolation to different
  times, locations and targets is uncertain. Published models are often
  internally validated on their own data but robust external validation
  in truly new scenarios is problematic. Another difficulty with many
  published models is that they have been built retrospectively once all
  the data has been collected. For WBE to be useful as an early warning
  system, it needs to work in real time.</p>
  <p>For these reasons, we focus our attention on methods to normalize
  the target RNA measure using a control without any public health data.
  Our aim is to produce an improved RNA-based indicator of disease using
  auxiliary information but not using the prevalence.</p>
</sec>
<sec id="data">
  <title>Data</title>
  <p>We demonstrate our methods on data used by (Schenk <italic>et
  al.</italic>, 2024) and presented in (Boehm <italic>et al.</italic>,
  2024) where details of the data collection may be found. We focus on a
  single site, Sunnyvale, California and only three targets: the
  primary, SARS-CoV2 RNA and two controls, Pepper Mild Mottle Virus
  (PMMoV) RNA and Bovine Coronavirus (BCoV). The data was collected
  daily between 1<sup>st</sup> January 2022 and 30<sup>th</sup> June
  2024, a period of 911 days. The data is high quality, collected daily
  (excepting 14 missing days), consists of long time series and is
  publicly available making it a good representative example of current
  WBE data. Sunnyvale is an ordinary mid-sized city. In the
  supplementary materials, we compute statistics on the other 190 sites
  in the data indicating that Sunnyvale is not exceptional. The
  laboratory method here used wastewater solids whereas other methods
  use liquid influent. The effectiveness of normalization depends on the
  joint distribution of the targets. Our main purpose is not to make any
  specific claims about this chosen dataset, but to address the general
  problem.</p>
  <fig>
    <caption><p>Daily measured SARS-CoV2 and PMMoV RNA in gene copies
    per gram dry weight on a log scale for 897 days in Sunnyvale,
    California. Solid line indicates a spline smoothed
    fit.</p></caption>
    <graphic mimetype="image" mime-subtype="svg+xml" xlink:href="media/image2.svg" />
  </fig>
  <p>Two of the targets are shown in Figure 1. Details of the
  construction of these plots and subsequent analysis are provided in
  the supplementary material. We see a clear signal in the SARS-CoV2
  although the daily measurement varies considerably around this trend.
  The RNA counts for PMMoV are about 7,000 times higher than SARS-CoV2.
  In contrast, the standard deviation of the residuals around the trend
  for PMMoV is 27% smaller than that seen in SARS-CoV2. PMMoV has a much
  stronger concentration but can be measured with only modestly greater
  relative precision. We also see a downward trend in the PMMoV
  concentration. The trend may be due to external causes such as lower
  pepper consumption or temporal changes in the sewerage or the
  measurement process. This has consequences for the normalization but
  identifying the cause is difficult. This is explored in (Rosengart
  <italic>et al.</italic>, 2025). However, this problem is not the focus
  of this article.</p>
</sec>
<sec id="direct-normalization">
  <title>Direct Normalization</title>
  <p>Suppose we measure a target <inline-formula><alternatives>
  <tex-math><![CDATA[T_{t}]]></tex-math>
  <mml:math display="inline" xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:msub><mml:mi>T</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:math></alternatives></inline-formula> and
  a control variable <inline-formula><alternatives>
  <tex-math><![CDATA[C_{t}]]></tex-math>
  <mml:math display="inline" xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:msub><mml:mi>C</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:math></alternatives></inline-formula>
  where <italic>t</italic> is the time index. We assume that the
  expected value of the control, <inline-formula><alternatives>
  <tex-math><![CDATA[\mathbb{E}C_{t}]]></tex-math>
  <mml:math display="inline" xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:mrow><mml:mi mathvariant="double-struck">𝔼</mml:mi><mml:msub><mml:mi>C</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></alternatives></inline-formula> is
  constant in <inline-formula><alternatives>
  <tex-math><![CDATA[t]]></tex-math>
  <mml:math display="inline" xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:mi>t</mml:mi></mml:math></alternatives></inline-formula> and
  that there are common factors which
  cause <inline-formula><alternatives>
  <tex-math><![CDATA[T_{t}]]></tex-math>
  <mml:math display="inline" xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:msub><mml:mi>T</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:math></alternatives></inline-formula> and <inline-formula><alternatives>
  <tex-math><![CDATA[C_{t}]]></tex-math>
  <mml:math display="inline" xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:msub><mml:mi>C</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:math></alternatives></inline-formula> to
  positively covary. The mechanisms for this vary according to the type
  of control variate.</p>
  <p>Direct normalization computes:</p>
  <p><disp-formula><alternatives>
  <tex-math><![CDATA[T_{t}^{A} = c\frac{T_{t}}{C_{t}}
   \qquad(1)]]></tex-math>
  <mml:math display="block" xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:mrow><mml:msubsup><mml:mi>T</mml:mi><mml:mi>t</mml:mi><mml:mi>A</mml:mi></mml:msubsup><mml:mo>=</mml:mo><mml:mi>c</mml:mi><mml:mfrac><mml:msub><mml:mi>T</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:msub><mml:mi>C</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mfrac><mml:mspace width="2.0em"></mml:mspace><mml:mo stretchy="false" form="prefix">(</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false" form="postfix">)</mml:mo></mml:mrow></mml:math></alternatives></disp-formula></p>
  <p>where <inline-formula><alternatives>
  <tex-math><![CDATA[c]]></tex-math>
  <mml:math display="inline" xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:mi>c</mml:mi></mml:math></alternatives></inline-formula> could
  be set to maintain the mean level of <inline-formula><alternatives>
  <tex-math><![CDATA[T]]></tex-math>
  <mml:math display="inline" xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:mi>T</mml:mi></mml:math></alternatives></inline-formula> by
  using <inline-formula><alternatives>
  <tex-math><![CDATA[c = \overline{C}]]></tex-math>
  <mml:math display="inline" xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:mrow><mml:mi>c</mml:mi><mml:mo>=</mml:mo><mml:mover><mml:mi>C</mml:mi><mml:mo accent="true">¯</mml:mo></mml:mover></mml:mrow></mml:math></alternatives></inline-formula> or
  similar. This method is widely used in the WBE and other applications.
  Recent examples include (Sweetapple <italic>et al.</italic>, 2023;
  Pellett <italic>et al.</italic>, 2024; Schenk <italic>et al.</italic>,
  2024). Sometimes more than one control variate is considered.</p>
  <p>If <inline-formula><alternatives>
  <tex-math><![CDATA[T]]></tex-math>
  <mml:math display="inline" xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:mi>T</mml:mi></mml:math></alternatives></inline-formula> and <inline-formula><alternatives>
  <tex-math><![CDATA[C]]></tex-math>
  <mml:math display="inline" xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:mi>C</mml:mi></mml:math></alternatives></inline-formula> are
  normally distributed, <inline-formula><alternatives>
  <tex-math><![CDATA[T^{A}]]></tex-math>
  <mml:math display="inline" xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:msup><mml:mi>T</mml:mi><mml:mi>A</mml:mi></mml:msup></mml:math></alternatives></inline-formula> will
  have a normal ratio distribution. In the worst case,
  when <inline-formula><alternatives>
  <tex-math><![CDATA[T]]></tex-math>
  <mml:math display="inline" xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:mi>T</mml:mi></mml:math></alternatives></inline-formula> and <inline-formula><alternatives>
  <tex-math><![CDATA[C]]></tex-math>
  <mml:math display="inline" xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:mi>C</mml:mi></mml:math></alternatives></inline-formula> are
  mean centered, <inline-formula><alternatives>
  <tex-math><![CDATA[T^{A}]]></tex-math>
  <mml:math display="inline" xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:msup><mml:mi>T</mml:mi><mml:mi>A</mml:mi></mml:msup></mml:math></alternatives></inline-formula> would
  have a Cauchy distribution. This distribution has notoriously heavy
  tails and not even its expectation is defined. This would be a
  terrible choice for an indicator monitoring the prevalence of a
  disease. One can avoid this pitfall by not mean centering and ensuring
  that <inline-formula><alternatives>
  <tex-math><![CDATA[C_{t}]]></tex-math>
  <mml:math display="inline" xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:msub><mml:mi>C</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:math></alternatives></inline-formula>
  is bounded well away from zero i.e. in practice we must have a strong
  control that is persistently present at high concentrations. This is
  true for our Sunnyvale example where no such problems arise. The
  normal ratio distribution is known for having long tails so outliers
  may be generated and care must be taken. The properties of this
  distribution are discussed in (Marsaglia, 2006) who provides formulae
  to compute the variance of the ratio under different scenarios.</p>
  <p>It is more convenient to consider the measurements on a log scale.
  Furthermore, since the errors are multiplicative, given the nature of
  PCR, a log scale is even preferable. One can compare unadjusted
  variance on a log scale, <inline-formula><alternatives>
  <tex-math><![CDATA[var\left( \log(T) \right)]]></tex-math>
  <mml:math display="inline" xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:mrow><mml:mi>v</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mrow><mml:mo stretchy="true" form="prefix">(</mml:mo><mml:mi mathvariant="normal">log</mml:mi><mml:mo stretchy="false" form="prefix">(</mml:mo><mml:mi>T</mml:mi><mml:mo stretchy="false" form="postfix">)</mml:mo><mml:mo stretchy="true" form="postfix">)</mml:mo></mml:mrow></mml:mrow></mml:math></alternatives></inline-formula>
  with the comparable adjusted variance using the standard formula for
  the difference of two random variables:</p>
  <p><disp-formula><alternatives>
  <tex-math><![CDATA[\log{T^{A} = \log{T - \log{C + \log c}}} \qquad(2)]]></tex-math>
  <mml:math display="block" xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:mrow><mml:mrow><mml:mi mathvariant="normal">log</mml:mi><mml:mo>&#8289;</mml:mo></mml:mrow><mml:mrow><mml:msup><mml:mi>T</mml:mi><mml:mi>A</mml:mi></mml:msup><mml:mo>=</mml:mo><mml:mrow><mml:mi mathvariant="normal">log</mml:mi><mml:mo>&#8289;</mml:mo></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mo>−</mml:mo><mml:mrow><mml:mi mathvariant="normal">log</mml:mi><mml:mo>&#8289;</mml:mo></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mo>+</mml:mo><mml:mrow><mml:mi mathvariant="normal">log</mml:mi><mml:mo>&#8289;</mml:mo></mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:mrow></mml:mrow><mml:mspace width="2.0em"></mml:mspace><mml:mo stretchy="false" form="prefix">(</mml:mo><mml:mn>2</mml:mn><mml:mo stretchy="false" form="postfix">)</mml:mo></mml:mrow></mml:math></alternatives></disp-formula></p>
  <p><disp-formula><alternatives>
  <tex-math><![CDATA[{var(logT}^{A}) = var(\log{T)} + var(log\ C) - 2\rho\sqrt{var(logT).var(logC)} \qquad(3)]]></tex-math>
  <mml:math display="block" xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:mrow><mml:msup><mml:mrow><mml:mi>v</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mo stretchy="false" form="prefix">(</mml:mo><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:mi>T</mml:mi></mml:mrow><mml:mi>A</mml:mi></mml:msup><mml:mo stretchy="false" form="postfix">)</mml:mo><mml:mo>=</mml:mo><mml:mi>v</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mo stretchy="false" form="prefix">(</mml:mo><mml:mrow><mml:mi mathvariant="normal">log</mml:mi><mml:mo>&#8289;</mml:mo></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mo stretchy="false" form="postfix">)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mi>v</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mo stretchy="false" form="prefix">(</mml:mo><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:mspace width="0.222em"></mml:mspace><mml:mi>C</mml:mi><mml:mo stretchy="false" form="postfix">)</mml:mo><mml:mo>−</mml:mo><mml:mn>2</mml:mn><mml:mi>ρ</mml:mi><mml:msqrt><mml:mrow><mml:mi>v</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mo stretchy="false" form="prefix">(</mml:mo><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:mi>T</mml:mi><mml:mo stretchy="false" form="postfix">)</mml:mo><mml:mi>.</mml:mi><mml:mi>v</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mo stretchy="false" form="prefix">(</mml:mo><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:mi>C</mml:mi><mml:mo stretchy="false" form="postfix">)</mml:mo></mml:mrow></mml:msqrt><mml:mspace width="2.0em"></mml:mspace><mml:mo stretchy="false" form="prefix">(</mml:mo><mml:mn>3</mml:mn><mml:mo stretchy="false" form="postfix">)</mml:mo></mml:mrow></mml:math></alternatives></disp-formula></p>
  <p>The correlation between <inline-formula><alternatives>
  <tex-math><![CDATA[logT]]></tex-math>
  <mml:math display="inline" xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:mrow><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:mi>T</mml:mi></mml:mrow></mml:math></alternatives></inline-formula>
  and <inline-formula><alternatives>
  <tex-math><![CDATA[logC]]></tex-math>
  <mml:math display="inline" xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:mrow><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:mi>C</mml:mi></mml:mrow></mml:math></alternatives></inline-formula>
  is given by <inline-formula><alternatives>
  <tex-math><![CDATA[\rho.]]></tex-math>
  <mml:math display="inline" xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:mrow><mml:mi>ρ</mml:mi><mml:mi>.</mml:mi></mml:mrow></mml:math></alternatives></inline-formula>
  The <italic>c</italic> is a constant so does not affect the variance.
  We see that for variation in the normalized target to be smaller than
  that in the variation in unnormalized target, the control variation
  <inline-formula><alternatives>
  <tex-math><![CDATA[var(log\ C)]]></tex-math>
  <mml:math display="inline" xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:mrow><mml:mi>v</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mo stretchy="false" form="prefix">(</mml:mo><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:mspace width="0.222em"></mml:mspace><mml:mi>C</mml:mi><mml:mo stretchy="false" form="postfix">)</mml:mo></mml:mrow></mml:math></alternatives></inline-formula>
  needs to be small and the correlation needs to be large. The problem
  is that, in WBE applications, T and C are both measured using PCR so
  the variation in measurement will tend to be similar while the
  correlation tends to be moderate at best. The adjustment can easily
  make the variance higher. Some evidence of this is reported in
  (Darling <italic>et al.</italic>, 2025).</p>
  <p>Let’s consider the Sunnyvale data. To estimate
  <inline-formula><alternatives>
  <tex-math><![CDATA[var(\log{T)}]]></tex-math>
  <mml:math display="inline" xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:mrow><mml:mi>v</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mo stretchy="false" form="prefix">(</mml:mo><mml:mrow><mml:mi mathvariant="normal">log</mml:mi><mml:mo>&#8289;</mml:mo></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mo stretchy="false" form="postfix">)</mml:mo></mml:mrow></mml:mrow></mml:math></alternatives></inline-formula>
  and <inline-formula><alternatives>
  <tex-math><![CDATA[var(log\ C)]]></tex-math>
  <mml:math display="inline" xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:mrow><mml:mi>v</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mo stretchy="false" form="prefix">(</mml:mo><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:mspace width="0.222em"></mml:mspace><mml:mi>C</mml:mi><mml:mo stretchy="false" form="postfix">)</mml:mo></mml:mrow></mml:math></alternatives></inline-formula>,
  we need to make some assumptions. In Figure 1, we don’t expect the
  true value of the SARS-CoV2 measure to vary substantially from one day
  to the next because</p>
  <list list-type="bullet">
    <list-item>
      <p>SARS-CoV2 in the wastewater derives from individuals whose
      COVID infection lasts several days. These individuals contribute a
      rising then falling amount of SARS-CoV2 RNA over these days.</p>
    </list-item>
    <list-item>
      <p>The number of people infected with COVID will not change
      extremely rapidly from one day to the next (either up or down)</p>
    </list-item>
    <list-item>
      <p>We are observing a convolution of multiple SARS-CoV2 sources in
      the wastewater</p>
    </list-item>
  </list>
  <p>The reasons that the RNA measures vary so much from one day to the
  next is due to variability in the measurement system, particularly in
  PCR, and external factors such as rainfall and water quality. For
  these reasons, it is reasonable to assume that the RNA counts are
  smoothly varying over time and that the smooths presented in Figure 1
  are plausible estimates of the underlying RNA counts.</p>
  <p>We can compute the residuals from the fits in Figure 1 and use
  these to compute the correlations and standard deviations shown in
  Table 1.</p>
  <table-wrap>
    <caption>
      <p>Correlations and Standard Deviation of the measurement errors
      of SARS-CoV2 (SC2), PMMoV and BCoV. Final row shows 95% confidence
      intervals.</p>
    </caption>
    <table>
      <colgroup>
        <col width="15%" />
        <col width="15%" />
        <col width="13%" />
        <col width="13%" />
        <col width="14%" />
        <col width="15%" />
        <col width="15%" />
      </colgroup>
      <thead>
        <tr>
          <th><p>Correlation</p>
          <p>SC2 with</p>
          <p>PMMoV</p></th>
          <th><p>Correlation</p>
          <p>SC2 with</p>
          <p>BCoV</p></th>
          <th><p>SD</p>
          <p>SC2</p></th>
          <th><p>SD</p>
          <p>PMMoV</p></th>
          <th><p>SD</p>
          <p>BCoV</p></th>
          <th><p>Normalized</p>
          <p>SD SC2 by</p>
          <p>PMMoV</p></th>
          <th><p>Normalized</p>
          <p>SD SC2 by</p>
          <p>BCoV</p></th>
        </tr>
      </thead>
      <tbody>
        <tr>
          <td>0.23</td>
          <td>-0.08</td>
          <td>0.56</td>
          <td>0.41</td>
          <td>0.43</td>
          <td>0.61</td>
          <td>0.73</td>
        </tr>
        <tr>
          <td>(0.17,0.29)</td>
          <td>(-0.14,-0.01)</td>
          <td>(0.51,0.61)</td>
          <td>(0.36,0.46)</td>
          <td>(0.40,0.46)</td>
          <td></td>
          <td></td>
        </tr>
      </tbody>
    </table>
  </table-wrap>
  <p>We see that the correlation between the SARS-CoV2 and PMMoV RNA
  measurements is only modestly positive while that between SARS-CoV2
  and BCoV is slightly negative. This is not helpful for normalization.
  The standard deviation for the SARS-CoV2 measurement is large. Since
  the measurement is on a log scale so <inline-formula><alternatives>
  <tex-math><![CDATA[\exp(0.56) = 1.75]]></tex-math>
  <mml:math display="inline" xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:mrow><mml:mrow><mml:mi mathvariant="normal">exp</mml:mi><mml:mo>&#8289;</mml:mo></mml:mrow><mml:mo stretchy="false" form="prefix">(</mml:mo><mml:mn>0.56</mml:mn><mml:mo stretchy="false" form="postfix">)</mml:mo><mml:mo>=</mml:mo><mml:mn>1.75</mml:mn></mml:mrow></mml:math></alternatives></inline-formula>,
  we have about 75% relative variation. We need to reduce this variation
  for this indicator of disease prevalence to be more useful. The
  standard deviations of PMMoV and BCoV are smaller reflecting their
  much greater concentration in the sample. We compute the normalized
  standard deviations using the equation above. We find that
  normalization makes the variation larger and not smaller as we might
  have hoped. In this example, the correlation between SARS-CoV2 and
  PMMoV would need to be at least 0.37 for the normalization to be
  beneficial.</p>
  <p>One may object that these calculations depend on the particular
  smoother we have applied to the time series of data as seen in Figure
  1. An alternative argument uses the mean absolute successive
  differences (MASD) on the log scale as a measure of smoothness.</p>
  <p><disp-formula><alternatives>
  <tex-math><![CDATA[d_{i} = {|log}\left( T_{i} \right) - log(T_{i - 1})|\ and\ MASD = \ \sum_{i = 2}^{n}{d_{i}/(n - 1)} \qquad(4)]]></tex-math>
  <mml:math display="block" xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo stretchy="false" form="prefix">|</mml:mo><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="true" form="prefix">(</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="true" form="postfix">)</mml:mo></mml:mrow><mml:mo>−</mml:mo><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:mo stretchy="false" form="prefix">(</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>−</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo stretchy="false" form="postfix">)</mml:mo><mml:mo stretchy="false" form="prefix">|</mml:mo><mml:mspace width="0.222em"></mml:mspace><mml:mi>a</mml:mi><mml:mi>n</mml:mi><mml:mi>d</mml:mi><mml:mspace width="0.222em"></mml:mspace><mml:mi>M</mml:mi><mml:mi>A</mml:mi><mml:mi>S</mml:mi><mml:mi>D</mml:mi><mml:mo>=</mml:mo><mml:mspace width="0.222em"></mml:mspace><mml:munderover><mml:mo>∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>2</mml:mn></mml:mrow><mml:mi>n</mml:mi></mml:munderover><mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mi>/</mml:mi><mml:mo stretchy="false" form="prefix">(</mml:mo><mml:mi>n</mml:mi><mml:mo>−</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false" form="postfix">)</mml:mo></mml:mrow><mml:mspace width="2.0em"></mml:mspace><mml:mo stretchy="false" form="prefix">(</mml:mo><mml:mn>4</mml:mn><mml:mo stretchy="false" form="postfix">)</mml:mo></mml:mrow></mml:math></alternatives></disp-formula></p>
  <p>For the unnormalized SARS-CoV2 series, we compute MASD=0.55. This
  is similar to the previous estimated SD of 0.56 although the
  estimators are not exactly equivalent. We also compute MASD for the
  normalized series finding values of 0.59 and 0.72 for PMMoV and BCoV
  respectively. We see again that normalization has made the variation
  larger.</p>
  <p>Normalization is also intended to achieve bias reduction. A control
  spiked into the sample after collection may not achieve this. In this
  example, the BCoV measure has very little correlation with the
  SARS-CoV2 target. Since a constant amount of the control is added to
  the sample, there is little opportunity for bias reduction. This is a
  vindication of well-controlled laboratory procedures, but no
  normalization (of the kind described here) should be used.</p>
  <p>There is more potential for bias reduction for a control that
  passes through the sewer as well as the laboratory. Given the
  substantial day-to-day variation in PMMoV control, the pointwise
  direct normalization we discuss in this section will not be the best
  approach. We should use some local smoothing as we will discuss later.
  There is also the difficulty in ascertaining the cause of longer-term
  variation in PMMoV.</p>
  <p>The laboratory procedures used to collect the data discussed in
  this study resulted in relatively low correlation between the PMMoV
  control and the target. It is possible that other procedures would
  exhibit higher correlation and so more effective normalization. We
  shall see how much higher correlation is necessary.</p>
  <p>How can we do better? The low correlation and high variance of the
  PCR-measured control appear inherent to the measurement technology. It
  seems unlikely we would find something better than BCoV for this
  purpose. We might consider other properties of the sampled wastewater
  that are not measured using PCR for use as a control. For example,
  ammonia content has often been used for this purpose. This can be
  measured more accurately but the correlation with the target RNA will
  be lower because a different measurement process is being used. The
  prospects for improvement are not strong.</p>
  <p>Bias reduction is also a worthwhile objective for a water quality
  measure such as ammonia. PMMoV and particularly BCoV are intended as
  constant controls so are not intended for bias reduction.</p>
  <p>We might also consider a different method of normalization. The
  normalization adjustment takes the form (on the log scale):</p>
  <p><disp-formula><alternatives>
  <tex-math><![CDATA[log\ T^{A} = logT - \gamma logC + \log c \qquad(5)]]></tex-math>
  <mml:math display="block" xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:mrow><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:mspace width="0.222em"></mml:mspace><mml:msup><mml:mi>T</mml:mi><mml:mi>A</mml:mi></mml:msup><mml:mo>=</mml:mo><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:mi>T</mml:mi><mml:mo>−</mml:mo><mml:mi>γ</mml:mi><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:mi>C</mml:mi><mml:mo>+</mml:mo><mml:mrow><mml:mi mathvariant="normal">log</mml:mi><mml:mo>&#8289;</mml:mo></mml:mrow><mml:mi>c</mml:mi><mml:mspace width="2.0em"></mml:mspace><mml:mo stretchy="false" form="prefix">(</mml:mo><mml:mn>5</mml:mn><mml:mo stretchy="false" form="postfix">)</mml:mo></mml:mrow></mml:math></alternatives></disp-formula></p>
  <p>For the direct normalization, <inline-formula><alternatives>
  <tex-math><![CDATA[\gamma]]></tex-math>
  <mml:math display="inline" xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:mi>γ</mml:mi></mml:math></alternatives></inline-formula>
  takes the value of one. For a weakly correlated control, this
  adjustment is too large, and we need a smaller value of
  <inline-formula><alternatives>
  <tex-math><![CDATA[\gamma]]></tex-math>
  <mml:math display="inline" xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:mi>γ</mml:mi></mml:math></alternatives></inline-formula>.
  We formalize this idea in the next section.</p>
</sec>
<sec id="kalman-filter">
  <title>Kalman Filter</title>
  <p>We observe data over time. We have an imperfect current measurement
  of the target to be combined with auxiliary information from controls
  and past measurements. We need a normalization that would work in real
  time as we collect the data without needing to wait until data
  collection is complete. The Kalman filter is well suited for this
  task. For background on the method from a statistical perspective, see
  (Durbin and Koopman, 2012).</p>
  <p>We have an observation equation:</p>
  <p><disp-formula><alternatives>
  <tex-math><![CDATA[y_{t} = Z\alpha_{t} + \epsilon_{t}
   \qquad(6)]]></tex-math>
  <mml:math display="block" xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mi>Z</mml:mi><mml:msub><mml:mi>α</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>ϵ</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mspace width="2.0em"></mml:mspace><mml:mo stretchy="false" form="prefix">(</mml:mo><mml:mn>6</mml:mn><mml:mo stretchy="false" form="postfix">)</mml:mo></mml:mrow></mml:math></alternatives></disp-formula></p>
  <p><inline-formula><alternatives>
  <tex-math><![CDATA[y_{t}]]></tex-math>
  <mml:math display="inline" xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:msub><mml:mi>y</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:math></alternatives></inline-formula> is
  a <inline-formula><alternatives>
  <tex-math><![CDATA[p]]></tex-math>
  <mml:math display="inline" xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:mi>p</mml:mi></mml:math></alternatives></inline-formula>-dimensional
  vector of observations at time <inline-formula><alternatives>
  <tex-math><![CDATA[t]]></tex-math>
  <mml:math display="inline" xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:mi>t</mml:mi></mml:math></alternatives></inline-formula>,
  <inline-formula><alternatives>
  <tex-math><![CDATA[Z]]></tex-math>
  <mml:math display="inline" xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:mi>Z</mml:mi></mml:math></alternatives></inline-formula> is
  a fixed system matrix (we use an identity matrix in our case),
  <inline-formula><alternatives>
  <tex-math><![CDATA[\alpha_{t}]]></tex-math>
  <mml:math display="inline" xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:msub><mml:mi>α</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:math></alternatives></inline-formula> is
  the <inline-formula><alternatives>
  <tex-math><![CDATA[m]]></tex-math>
  <mml:math display="inline" xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:mi>m</mml:mi></mml:math></alternatives></inline-formula>-dimensional
  vector of the underlying states at time <inline-formula><alternatives>
  <tex-math><![CDATA[t]]></tex-math>
  <mml:math display="inline" xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:mi>t</mml:mi></mml:math></alternatives></inline-formula>
  (we set <inline-formula><alternatives>
  <tex-math><![CDATA[m = p\ ]]></tex-math>
  <mml:math display="inline" xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:mrow><mml:mi>m</mml:mi><mml:mo>=</mml:mo><mml:mi>p</mml:mi><mml:mspace width="0.222em"></mml:mspace></mml:mrow></mml:math></alternatives></inline-formula>in
  our examples), the measurement errors, <inline-formula><alternatives>
  <tex-math><![CDATA[\epsilon_{t}]]></tex-math>
  <mml:math display="inline" xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:msub><mml:mi>ϵ</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:math></alternatives></inline-formula>
  are multivariate normal, <inline-formula><alternatives>
  <tex-math><![CDATA[N(0,H)]]></tex-math>
  <mml:math display="inline" xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:mrow><mml:mi>N</mml:mi><mml:mo stretchy="false" form="prefix">(</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mi>H</mml:mi><mml:mo stretchy="false" form="postfix">)</mml:mo></mml:mrow></mml:math></alternatives></inline-formula>.
  We use a bivariate (p=2) with <inline-formula><alternatives>
  <tex-math><![CDATA[y_{t}]]></tex-math>
  <mml:math display="inline" xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:msub><mml:mi>y</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:math></alternatives></inline-formula> consisting
  of the logged SARS-CoV2 RNA and one of the two control RNA measures
  (also log scale). The covariance of the measurement errors,
  <inline-formula><alternatives>
  <tex-math><![CDATA[H]]></tex-math>
  <mml:math display="inline" xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:mi>H</mml:mi></mml:math></alternatives></inline-formula>,
  is composed of the SDs and correlations we have considered previously
  but will be estimated in the Kalman filtering.</p>
  <p>The states <inline-formula><alternatives>
  <tex-math><![CDATA[\alpha_{t}]]></tex-math>
  <mml:math display="inline" xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:msub><mml:mi>α</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:math></alternatives></inline-formula>represent
  the true unknown values of the RNA measures and are connected by a
  state equation:</p>
  <p><disp-formula><alternatives>
  <tex-math><![CDATA[\alpha_{t + 1} = T\alpha_{t} + R\eta_{t}
   \qquad(7)]]></tex-math>
  <mml:math display="block" xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:mrow><mml:msub><mml:mi>α</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>T</mml:mi><mml:msub><mml:mi>α</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:mi>R</mml:mi><mml:msub><mml:mi>η</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mspace width="2.0em"></mml:mspace><mml:mo stretchy="false" form="prefix">(</mml:mo><mml:mn>7</mml:mn><mml:mo stretchy="false" form="postfix">)</mml:mo></mml:mrow></mml:math></alternatives></disp-formula></p>
  <p><inline-formula><alternatives>
  <tex-math><![CDATA[T]]></tex-math>
  <mml:math display="inline" xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:mi>T</mml:mi></mml:math></alternatives></inline-formula> is
  a system matrix (set to the identity in our example),
  <inline-formula><alternatives>
  <tex-math><![CDATA[R]]></tex-math>
  <mml:math display="inline" xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:mi>R</mml:mi></mml:math></alternatives></inline-formula> is
  another system matrix (also the identity in our example), the state
  disturbances are <inline-formula><alternatives>
  <tex-math><![CDATA[\eta_{t}]]></tex-math>
  <mml:math display="inline" xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:msub><mml:mi>η</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:math></alternatives></inline-formula> are
  multivariate normal, <inline-formula><alternatives>
  <tex-math><![CDATA[N(0,Q)]]></tex-math>
  <mml:math display="inline" xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:mrow><mml:mi>N</mml:mi><mml:mo stretchy="false" form="prefix">(</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mi>Q</mml:mi><mml:mo stretchy="false" form="postfix">)</mml:mo></mml:mrow></mml:math></alternatives></inline-formula> where
  the covariance <inline-formula><alternatives>
  <tex-math><![CDATA[Q]]></tex-math>
  <mml:math display="inline" xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:mi>Q</mml:mi></mml:math></alternatives></inline-formula>
  controls how the state varies from one day to the next.</p>
  <p>We need to specify a prior guess for the initial state with:
  <inline-formula><alternatives>
  <tex-math><![CDATA[\alpha_{1} \sim N(a_{1},P_{1})]]></tex-math>
  <mml:math display="inline" xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:mrow><mml:msub><mml:mi>α</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>∼</mml:mo><mml:mi>N</mml:mi><mml:mo stretchy="false" form="prefix">(</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo stretchy="false" form="postfix">)</mml:mo></mml:mrow></mml:math></alternatives></inline-formula>.<inline-formula><alternatives>
  <tex-math><![CDATA[\ ]]></tex-math>
  <mml:math display="inline" xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:mspace width="0.222em"></mml:mspace></mml:math></alternatives></inline-formula>We
  usually have prior WBE experience in these applications allowing us to
  specify a reasonably strong prior but the Kalman filter updating will
  eventually make this choice unimportant in the long run. We have
  intentionally chosen a simple formulation to focus on the
  normalization.</p>
  <p>We implement this using the KFAS R package of (Helske, 2017). There
  are other Kalman filter packages written in R, Python and other
  languages. The algorithm is not complex and could be coded directly
  according to local requirements.</p>
  <p>We fit three models to the data:</p>
  <list list-type="order">
    <list-item>
      <p>Univariate response of logged SARS-CoV2 RNA. This serves as the
      baseline unnormalized performance. The Kalman filter combines the
      current observation with information about the previous
      observation and the estimated variation to produce a filtered
      estimate of the state.</p>
    </list-item>
    <list-item>
      <p>Bivariate response of the SARS-CoV2 RNA and PMMoV RNA, both on
      the log-scale. Information about the control is used to improve
      the estimate of the target state.</p>
    </list-item>
    <list-item>
      <p>Bivariate response of the SARS-CoV2 RNA and BCoV RNA, both on
      the log-scale. The use of the laboratory control is assessed.</p>
    </list-item>
  </list>
  <p>In all three models, normal diagnostics were checked and can be
  seen in the supplementary materials. The 14 missing values were
  dropped and the resulting series had length 897.</p>
  <p>The estimates are shown in Table 2. The first three numerical
  columns refer to the state covariance <inline-formula><alternatives>
  <tex-math><![CDATA[Q]]></tex-math>
  <mml:math display="inline" xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:mi>Q</mml:mi></mml:math></alternatives></inline-formula>
  and the second set of three columns refer to the observation
  covariance <inline-formula><alternatives>
  <tex-math><![CDATA[H.]]></tex-math>
  <mml:math display="inline" xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:mrow><mml:mi>H</mml:mi><mml:mi>.</mml:mi></mml:mrow></mml:math></alternatives></inline-formula>
  The method produces a filtered estimate of the state at each time,
  <inline-formula><alternatives>
  <tex-math><![CDATA[\widehat{\alpha_{t}}]]></tex-math>
  <mml:math display="inline" xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:mover><mml:msub><mml:mi>α</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo accent="true">̂</mml:mo></mml:mover></mml:math></alternatives></inline-formula>,
  as seen in Figure 2. We show the final<inline-formula><alternatives>
  <tex-math><![CDATA[,]]></tex-math>
  <mml:math display="inline" xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:mo>,</mml:mo></mml:math></alternatives></inline-formula>
  <inline-formula><alternatives>
  <tex-math><![CDATA[{SD(\widehat{\alpha}}_{n})]]></tex-math>
  <mml:math display="inline" xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:mrow><mml:msub><mml:mrow><mml:mi>S</mml:mi><mml:mi>D</mml:mi><mml:mo stretchy="false" form="prefix">(</mml:mo><mml:mover><mml:mi>α</mml:mi><mml:mo accent="true">̂</mml:mo></mml:mover></mml:mrow><mml:mi>n</mml:mi></mml:msub><mml:mo stretchy="false" form="postfix">)</mml:mo></mml:mrow></mml:math></alternatives></inline-formula>in
  the last column of the table.</p>
  <fig>
    <caption><p>Kalman filtered fit to SARS-CoV2 RNA time series
    normalized with PMMoV RNA. Data is grey and fit is
    black.</p></caption>
    <graphic mimetype="image" mime-subtype="svg+xml" xlink:href="media/image4.svg" />
  </fig>
  <p>We see that the filtered fit, shown in Figure 2, is smoother than
  the data but not as smooth as the fit in Figure 1. Since we believe
  the state series to be smooth, we could achieve a smoother fit using
  additional autoregressive terms in the state model for the Kalman
  filter. We would recommend this in practice, but we have not done it
  here because we are focussed on the effects of normalization by a
  control.</p>
  <table-wrap>
    <caption>
      <p>Estimates from 3 models, using SARS-CoV2 RNA alone or
      additionally using one control, PMMoV or BCoV.</p>
    </caption>
    <table>
      <colgroup>
        <col width="17%" />
        <col width="10%" />
        <col width="11%" />
        <col width="15%" />
        <col width="10%" />
        <col width="11%" />
        <col width="15%" />
        <col width="10%" />
      </colgroup>
      <thead>
        <tr>
          <th>Model</th>
          <th>State SD target</th>
          <th><p>State SD</p>
          <p>control</p></th>
          <th><p>State</p>
          <p>Correlation</p></th>
          <th><p>Obs. SD</p>
          <p>target</p></th>
          <th>Obs. SD control</th>
          <th><p>Obs</p>
          <p>Correlation</p></th>
          <th><p>Filter</p>
          <p>SE of</p>
          <p>state</p></th>
        </tr>
      </thead>
      <tbody>
        <tr>
          <td>SC2</td>
          <td>0.136</td>
          <td></td>
          <td></td>
          <td>0.522</td>
          <td></td>
          <td></td>
          <td>0.249</td>
        </tr>
        <tr>
          <td>SC2+PMMoV</td>
          <td>0.133</td>
          <td>0.025</td>
          <td>0.088</td>
          <td>0.523</td>
          <td>0.410</td>
          <td>0.231</td>
          <td>0.246</td>
        </tr>
        <tr>
          <td>SC2+BCoV</td>
          <td>0.135</td>
          <td>0.027</td>
          <td>0.312</td>
          <td>0.522</td>
          <td>0.434</td>
          <td>0.000</td>
          <td>0.249</td>
        </tr>
      </tbody>
    </table>
  </table-wrap>
  <p>From Table 2, we see that the state SD of the target suggests a
  relative variation from day-to-day of about 14% across all three
  models. We believe this is higher than the truth and can be reduced by
  using autoregressive terms. We expect the true level of PMMoV to vary
  due to sewer effects but the state of BCoV should be constant. We
  shall show shortly why the estimated state correlation for BCoV should
  not be interpreted strongly. The observation error SD for the target
  is similar for the three models and similar to the ad hoc estimate
  obtained in the previous section. There is a mild correlation between
  the measurement errors for PMMoV model but practically zero
  correlation for the BCoV model. We claim that this contrast indicates
  sewer effects for PMMoV while BCoV reflects only laboratory
  effects.</p>
  <p>The SE of the final filtered state estimate from Table 2 shows only
  small differences. Using the direct normalization method of the
  previous section made the variance larger. We have avoided this fate,
  but we did not achieve much improvement. The adjustment in the
  filtered estimate due to normalization is about 1%. It is generally an
  adjustment in the same direction of the direct adjustment of the
  previous section, but significantly smaller in size. Simulation
  evidence from the next section suggests we cannot be confident of even
  this small improvement. The Kalman updating step takes a form
  analogous to Equation (5) but the size of the step is reduced. Despite
  the possibility of some bias reduction, the results are disappointing
  as normalization using the controls has yielded no clear benefit.
  Under what circumstances might they be more effective?</p>
</sec>
<sec id="simulation">
  <title>Simulation</title>
  <p>We use simulation to explore the generality of the conclusions from
  the previous section and discover scenarios where the Kalman filtering
  normalization is clearly beneficial. We consider four simulated
  scenarios:</p>
  <list list-type="order">
    <list-item>
      <p>Data generated from the model for SARS-CoV2 and PMMoV RNA with
      parameters set to those estimated from the observed data in the
      previous section. We use this to check the stability of our
      conclusions.</p>
    </list-item>
    <list-item>
      <p>Data generated from the same model except with an observation
      error correlation of 0.8, much higher than the 0.23 observed. In
      this scenario, the target and control measurement are more tightly
      linked.</p>
    </list-item>
    <list-item>
      <p>Data generated from the first model except with a 10 times
      smaller control variation. In this scenario, the control can be
      measured with greater precision.</p>
    </list-item>
    <list-item>
      <p>Data generated from the first model except with a much stronger
      correlation between the states of 0.8 compared to the 0.09
      observed. In this scenario, the covariation of the target and
      control through the sewer are more tightly linked.</p>
    </list-item>
  </list>
  <p>In each scenario, we generate 1000 simulated datasets and report
  the mean estimates in Table 3.</p>
  <table-wrap>
    <caption>
      <p>Mean estimated parameters in four simulated scenarios using a
      Kalman filter model.</p>
    </caption>
    <table>
      <colgroup>
        <col width="15%" />
        <col width="9%" />
        <col width="12%" />
        <col width="15%" />
        <col width="11%" />
        <col width="12%" />
        <col width="15%" />
        <col width="10%" />
      </colgroup>
      <thead>
        <tr>
          <th>Model</th>
          <th>State SD target</th>
          <th><p>State SD</p>
          <p>control</p></th>
          <th><p>State</p>
          <p>Correlation</p></th>
          <th><p>Obs. SD</p>
          <p>target</p></th>
          <th>Obs. SD control</th>
          <th><p>Obs.</p>
          <p>Correlation</p></th>
          <th><p>Filter</p>
          <p>SE</p></th>
        </tr>
      </thead>
      <tbody>
        <tr>
          <td>Observed Data</td>
          <td>0.132</td>
          <td>0.024</td>
          <td>0.063</td>
          <td>0.523</td>
          <td>0.411</td>
          <td>0.232</td>
          <td>0.244</td>
        </tr>
        <tr>
          <td>High Obs. Correlation</td>
          <td>0.131</td>
          <td>0.024</td>
          <td>0.037</td>
          <td>0.520</td>
          <td>0.408</td>
          <td>0.792</td>
          <td>0.203</td>
        </tr>
        <tr>
          <td>Low control SD</td>
          <td>0.137</td>
          <td>0.026</td>
          <td>0.087</td>
          <td>0.519</td>
          <td>0.128</td>
          <td>0.236</td>
          <td>0.244</td>
        </tr>
        <tr>
          <td>High State Correlation</td>
          <td>0.135</td>
          <td>0.023</td>
          <td>0.723</td>
          <td>0.519</td>
          <td>0.410</td>
          <td>0.235</td>
          <td>0.245</td>
        </tr>
      </tbody>
    </table>
  </table-wrap>
  <p>In the case where we simulate from the fitted model, the mean
  results are close to the fitted estimates indicating a lack of bias in
  the fitting process. We can also use the simulated values to estimate
  the uncertainty in the parameter estimates (see supplementary
  materials for details). Most of the estimates have low enough
  variation to support our previous conclusions, but the estimate of the
  state correlation was found to be highly variable indicating that it
  is difficult to estimate this parameter well. This indicates why the
  values in Table 2 for this parameter must be viewed with caution.</p>
  <p>For the three scenarios where we perturb the parameters, we see
  that the estimated parameters are obtained without substantial bias.
  Only in the high observed error correlation situation do we observe
  worthwhile improvement in the filtered state SD. To realize this
  improvement, we would need to find a control whose measurement is much
  more strongly correlated with the target. Given current PCR
  technology, there is no candidate for such a control.</p>
  <p>We have used a simple bivariate Kalman filter in these examples to
  focus on the issue of variance reduction via a control. We could use
  autoregressive terms to improve the fit, but this would not improve
  the value of the controls. The controls could also be introduced as
  regressors in the model, but the weak correlation would remain an
  obstacle. We have used a Gaussian Kalman Filter but for targets that
  appear intermittently or with very low prevalence, we can adapt the
  method to use other distributions. An example of this can be seen in
  (Cluzel <italic>et al.</italic>, 2022). The Kalman filter is an
  example of dynamic linear model with wider applications in WBE found
  in (Ouyang <italic>et al.</italic>, 2025).</p>
</sec>
<sec id="alternatives-for-variance-reduction">
  <title>Alternatives for Variance Reduction</title>
  <p>We have seen that it is difficult to achieve worthwhile variance
  reduction by using a control in situations typical of PCR measurement
  of targets in WBE. There are some alternative avenues for improvement.
  The high variation in the measurement of the laboratory control BCoV
  suggests that the PCR measurement is the single largest component in
  the overall variation of the target measure. One might hope for
  technical improvements in this technology.</p>
  <p>The most used tool for variance reduction is replication. One
  additional true replicated measurement would reduce the standard error
  for the mean of the now two measurements by a factor of
  <inline-formula><alternatives>
  <tex-math><![CDATA[\sqrt{2}.]]></tex-math>
  <mml:math display="inline" xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:mrow><mml:msqrt><mml:mn>2</mml:mn></mml:msqrt><mml:mi>.</mml:mi></mml:mrow></mml:math></alternatives></inline-formula>
  This would be superior to any improvement seen above, even in the
  simulated scenarios. There are difficulties with this approach.
  Replication is already used in WBE – in our example data, at least six
  replicates were used, although only the mean was reported. Further
  improvements by using more replications are subject to diminishing
  returns as the standard error is inversely proportional to the square
  root of the number of replications. A further difficulty is that a
  true replicate means repeating the collection and processing steps
  from the source at the WWTP. A replicate derived from division of the
  sample at a later stage will be less effective.</p>
  <p>Smoothing is related to replication in that recent measurements are
  combined with the current measurement. Various methods including
  rolling averages are used. The Kalman filter method described earlier
  can incorporate autoregressive terms. This is an effective method of
  variance reduction. Because we can only smooth over a small number of
  days, the variance reduction is limited.</p>
  <p>In (Schenk <italic>et al.</italic>, 2024), RNA measurements for
  other common pathogens, such as Norovirus or Influenza are available.
  We can incorporate all these targets into a single Kalman filter
  model. Since these other targets are subject to many of the same
  sources of variation, the statistical effect of <italic>borrowing
  strength</italic> will occur, improving the estimates and reducing the
  variation. Another idea is to measure a different gene target for the
  same RNA.</p>
  <p>We may also return to models that use prevalence information of the
  form seen in Equation (1), but with the requirement that the control
  or adjusting variables are used in an explicit process model that
  allows the knowledge of appropriate normalization to be extrapolated
  to new scenarios with different targets and locations. Perhaps,
  something like the parameterization of the normalization seen in
  (Leisman <italic>et al.</italic>, 2024) might be of benefit.</p>
</sec>
<sec id="conclusion">
  <title>Conclusion</title>
  <p>We have shown that using controls to normalize PCR measurements in
  WBE for the purposes of variance reduction is challenging. Meanwhile,
  these controls have other uses. A control that reflects variation in
  the sewer contains useful information about the population in the
  catchment area of the WWTP. A control which is introduced into the
  sample after collection is still useful for quality control purposes.
  Thus, we are not claiming such controls have no value, but researchers
  should consider whether measuring other pathogen RNA markers is a
  better use of limited resources. The Kalman filter model used here
  presents a framework where control and other auxiliary information can
  be used to potential improve the estimate of the primary target that
  can be used in online settings as new data becomes available.</p>
</sec>
<sec id="acknowledgements">
  <title>Acknowledgements</title>
  <p>Funding from the Centre of Excellence in Water-Based Early-Warning
  Systems for Health Protection (CWBE) at the University of Bath</p>
</sec>
<sec id="conflicts-of-interest">
  <title>Conflicts of interest</title>
  <p>There are no conflicts of interest</p>
  <sec id="data-and-code-availability">
    <title>Data and Code Availability</title>
    <p>Data and code supplied in the supplementary materials.</p>
  </sec>
</sec>
<sec id="references">
  <title>References</title>
  <p>Ahmed, T. <italic>et al.</italic> (2026) ‘Assessing normalization
  methods in wastewater based epidemiology: a systematic review’,
  <italic>Environmental Science: Water Research &amp;
  Technology</italic>, 12(5), pp. 1374–1382. Available at:
  <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1039/D5EW01049G">https://doi.org/10.1039/D5EW01049G</ext-link>.</p>
  <p>Boehm, A. <italic>et al.</italic> (2024) ‘Data for Human pathogen
  nucleic acids in wastewater solids from 191 wastewater treatment
  plants in the United States.’ Stanford Digital Repository. Available
  at:
  <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.25740/hj801ns5929">https://doi.org/10.25740/hj801ns5929</ext-link>.</p>
  <p>Cluzel, N. <italic>et al.</italic> (2022) ‘A nationwide indicator
  to smooth and normalize heterogeneous SARS-CoV-2 RNA data in
  wastewater’, <italic>Environment International</italic>, 158, p.
  106998. Available at:
  <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.envint.2021.106998">https://doi.org/10.1016/j.envint.2021.106998</ext-link>.</p>
  <p>Darling, A. <italic>et al.</italic> (2025) ‘Comparative Assessment
  of Wastewater-Based Surveillance Normalization Methods to Improve
  Pathogen Monitoring in Rural Sewersheds’, <italic>Environmental
  Science &amp; Technology</italic>, 59(22), pp. 11095–11107. Available
  at:
  <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1021/acs.est.4c14485">https://doi.org/10.1021/acs.est.4c14485</ext-link>.</p>
  <p>Deák, G., Lupu, L. and Prangate, R. (2026) ‘A Systematic Review of
  Methodological Approaches to SARS-CoV-2 Wastewater Surveillance’,
  <italic>Viruses</italic>, 18(2), p. 205. Available at:
  <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.3390/v18020205">https://doi.org/10.3390/v18020205</ext-link>.</p>
  <p>Durbin, J. and Koopman, S.J. (2012) <italic>Time series analysis by
  state space methods</italic>. Oxford university press.</p>
  <p>Helske, J. (2017) ‘KFAS: Exponential Family State Space Models in
  <italic>R</italic>’, <italic>Journal of Statistical Software</italic>,
  78(10). Available at:
  <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.18637/jss.v078.i10">https://doi.org/10.18637/jss.v078.i10</ext-link>.</p>
  <p>Lamm, E.D. <italic>et al.</italic> (2025) ‘Inclusion of
  Physical–Chemical Water Quality Measurements Can Improve Associations
  between SARS-CoV-2 RNA Levels in Wastewater and COVID-19 Cases within
  Smaller Sewersheds’, <italic>Journal of Environmental
  Engineering</italic>, 151(9), p. 04025049. Available at:
  <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1061/JOEEDU.EEENG-8138">https://doi.org/10.1061/JOEEDU.EEENG-8138</ext-link>.</p>
  <p>Leisman, K. <italic>et al.</italic> (2024) ‘A modeling pipeline to
  relate municipal wastewater surveillance and regional public health
  data’, <italic>Water Research</italic>, 252, p. 121178. Available at:
  <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.watres.2024.121178">https://doi.org/10.1016/j.watres.2024.121178</ext-link>.</p>
  <p>Marsaglia, G. (2006) ‘Ratios of Normal Variables’, <italic>Journal
  of Statistical Software</italic>, 16(4). Available at:
  <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.18637/jss.v016.i04">https://doi.org/10.18637/jss.v016.i04</ext-link>.</p>
  <p>Ouyang, D. <italic>et al.</italic> (2025) ‘Dynamic Linear Models
  for Wastewater-Based Epidemiology with Missing Values: an Application
  to Covid-19 Surveillance’, <italic>Data Science in Science</italic>,
  4(1), p. 2562199. Available at:
  <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1080/26941899.2025.2562199">https://doi.org/10.1080/26941899.2025.2562199</ext-link>.</p>
  <p>Pellett, C. <italic>et al.</italic> (2024) ‘Multi-factor
  normalisation of viral counts from wastewater improves the detection
  accuracy of viral disease in the community’, <italic>Environmental
  Technology &amp; Innovation</italic>, 36, p. 103720. Available at:
  <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.eti.2024.103720">https://doi.org/10.1016/j.eti.2024.103720</ext-link>.</p>
  <p>Rosengart, A.L. <italic>et al.</italic> (2025) ‘Spatiotemporal
  Variability of the Pepper Mild Mottle Virus Biomarker in Wastewater’,
  <italic>ACS ES&amp;T Water</italic>, 5(1), pp. 341–350. Available at:
  <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1021/acsestwater.4c00866">https://doi.org/10.1021/acsestwater.4c00866</ext-link>.</p>
  <p>Schenk, H. <italic>et al.</italic> (2024) ‘SARS-CoV-2 surveillance
  in US wastewater: Leading indicators and data variability analysis in
  2023–2024’, <italic>PLOS ONE</italic>. Edited by D.L. Wannigama,
  19(11), p. e0313927. Available at:
  <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1371/journal.pone.0313927">https://doi.org/10.1371/journal.pone.0313927</ext-link>.</p>
  <p>Sweetapple, C. <italic>et al.</italic> (2023) ‘Dynamic population
  normalisation in wastewater-based epidemiology for improved
  understanding of the SARS-CoV-2 prevalence: a multi-site study’,
  <italic>Journal of Water and Health</italic>, 21(5), pp. 625–642.
  Available at:
  <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.2166/wh.2023.318">https://doi.org/10.2166/wh.2023.318</ext-link>.</p>
  <p>Verani, M. <italic>et al.</italic> (2025) ‘Evaluating Population
  Normalization Methods Using Chemical Data for Wastewater-Based
  Epidemiology: Insights from a Site-Specific Case Study’,
  <italic>Viruses</italic>, 17(5), p. 672. Available at:
  <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.3390/v17050672">https://doi.org/10.3390/v17050672</ext-link>.</p>
</sec>
<sec sec-type="data-availability">
<title>Data and code availability</title>
<p>All code and data are openly available at the submitted reproducibility bundle. Running `make` regenerates the reported results.</p>
</sec>
</body>
<back><ref-list><ref id="ref1"><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Ahmed</surname><given-names>Tahmina</given-names></name><name><surname>Philo</surname><given-names>Sarah E.</given-names></name><name><surname>Boehm</surname><given-names>Alexandria B.</given-names></name><name><surname>Halden</surname><given-names>Rolf U.</given-names></name><name><surname>Bibby</surname><given-names>Kyle</given-names></name><name><surname>Delgado Vela</surname><given-names>Jeseth</given-names></name></person-group><article-title>Assessing normalization methods in wastewater based epidemiology: a systematic review</article-title><source>Environmental Science: Water Research &amp; Technology</source><year>2026</year><volume>12</volume><issue>5</issue><fpage>1374</fpage><lpage>1382</lpage><publisher-name>Royal Society of Chemistry (RSC)</publisher-name><pub-id pub-id-type="doi">10.1039/D5EW01049G</pub-id></element-citation></ref><ref id="ref2"><element-citation publication-type="data"><person-group person-group-type="author"><name><surname>Boehm</surname><given-names>Alexandria</given-names></name><name><surname>Bidwell</surname><given-names>Amanda</given-names></name><name><surname>Wolfe</surname><given-names>Marlene</given-names></name><name><surname>Zulli</surname><given-names>Alessandro</given-names></name><name><surname>Duong</surname><given-names>Dorothea</given-names></name><name><surname>Shelden</surname><given-names>Bridgette</given-names></name><name><surname>White</surname><given-names>Bradley</given-names></name></person-group><article-title>Data for Human pathogen nucleic acids in wastewater solids from 191 wastewater treatment plants in the United States</article-title><year>2025</year><publisher-name>Stanford Digital Repository</publisher-name><pub-id pub-id-type="doi">10.25740/hj801ns5929</pub-id></element-citation></ref><ref id="ref3"><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Cluzel</surname><given-names>Nicolas</given-names></name><name><surname>Courbariaux</surname><given-names>Marie</given-names></name><name><surname>Wang</surname><given-names>Siyun</given-names></name><name><surname>Moulin</surname><given-names>Laurent</given-names></name><name><surname>Wurtzer</surname><given-names>Sébastien</given-names></name><name><surname>Bertrand</surname><given-names>Isabelle</given-names></name><name><surname>Laurent</surname><given-names>Karine</given-names></name><name><surname>Monfort</surname><given-names>Patrick</given-names></name><name><surname>Gantzer</surname><given-names>Christophe</given-names></name><name><surname>Guyader</surname><given-names>Soizick Le</given-names></name><name><surname>Boni</surname><given-names>Mickaël</given-names></name><name><surname>Mouchel</surname><given-names>Jean-Marie</given-names></name><name><surname>Maréchal</surname><given-names>Vincent</given-names></name><name><surname>Nuel</surname><given-names>Grégory</given-names></name><name><surname>Maday</surname><given-names>Yvon</given-names></name></person-group><article-title>A nationwide indicator to smooth and normalize heterogeneous SARS-CoV-2 RNA data in wastewater</article-title><source>Environment International</source><year>2022</year><volume>158</volume><fpage>106998</fpage><publisher-name>Elsevier BV</publisher-name><pub-id pub-id-type="doi">10.1016/j.envint.2021.106998</pub-id></element-citation></ref><ref id="ref4"><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Darling</surname><given-names>Amanda</given-names></name><name><surname>Davis</surname><given-names>Benjamin C.</given-names></name><name><surname>Byrne</surname><given-names>Thomas</given-names></name><name><surname>Deck</surname><given-names>Madeline</given-names></name><name><surname>Maldonado Rivera</surname><given-names>Gabriel E.</given-names></name><name><surname>Price</surname><given-names>Sarah</given-names></name><name><surname>Amaral-Torres</surname><given-names>Amber</given-names></name><name><surname>Markham</surname><given-names>Clayton</given-names></name><name><surname>Gonzalez</surname><given-names>Raul A.</given-names></name><name><surname>Vikesland</surname><given-names>Peter J.</given-names></name><name><surname>Krometis</surname><given-names>Leigh-Anne H.</given-names></name><name><surname>Pruden</surname><given-names>Amy</given-names></name><name><surname>Cohen</surname><given-names>Alasdair</given-names></name></person-group><article-title>Comparative Assessment of Wastewater-Based Surveillance Normalization Methods to Improve Pathogen Monitoring in Rural Sewersheds</article-title><source>Environmental Science &amp; Technology</source><year>2025</year><volume>59</volume><issue>22</issue><fpage>11095</fpage><lpage>11107</lpage><publisher-name>American Chemical Society (ACS)</publisher-name><pub-id pub-id-type="doi">10.1021/acs.est.4c14485</pub-id></element-citation></ref><ref id="ref5"><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Deák</surname><given-names>György</given-names></name><name><surname>Lupu</surname><given-names>Laura</given-names></name><name><surname>Prangate</surname><given-names>Raluca</given-names></name></person-group><article-title>A Systematic Review of Methodological Approaches to SARS-CoV-2 Wastewater Surveillance</article-title><source>Viruses</source><year>2026</year><volume>18</volume><issue>2</issue><fpage>205</fpage><publisher-name>MDPI AG</publisher-name><pub-id pub-id-type="doi">10.3390/v18020205</pub-id></element-citation></ref><ref id="ref6"><element-citation publication-type="book"><person-group person-group-type="author"><name><surname>Durbin</surname><given-names>James</given-names></name><name><surname>Koopman</surname><given-names>Siem Jan</given-names></name></person-group><article-title>Time Series Analysis by State Space Methods</article-title><year>2012</year><publisher-name>Oxford University Press</publisher-name><pub-id pub-id-type="doi">10.1093/acprof:oso/9780199641178.001.0001</pub-id></element-citation></ref><ref id="ref7"><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Helske</surname><given-names>Jouni</given-names></name></person-group><article-title>KFAS: Exponential Family State Space Models in &lt;i&gt;R&lt;/i&gt;</article-title><source>Journal of Statistical Software</source><year>2017</year><volume>78</volume><issue>10</issue><publisher-name>Foundation for Open Access Statistic</publisher-name><pub-id pub-id-type="doi">10.18637/jss.v078.i10</pub-id></element-citation></ref><ref id="ref8"><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Lamm</surname><given-names>Erik D.</given-names></name><name><surname>Babler</surname><given-names>Kristina M.</given-names></name><name><surname>Sharkey</surname><given-names>Mark E.</given-names></name><name><surname>Amirali</surname><given-names>Ayaaz</given-names></name><name><surname>Beaver</surname><given-names>Cynthia</given-names></name><name><surname>Boone</surname><given-names>Melinda M.</given-names></name><name><surname>Comerford</surname><given-names>Samuel</given-names></name><name><surname>Cooper</surname><given-names>Daniel</given-names></name><name><surname>Currall</surname><given-names>Benjamin</given-names></name><name><surname>Grills</surname><given-names>George S.</given-names></name><name><surname>Kobetz</surname><given-names>Erin</given-names></name><name><surname>Kumar</surname><given-names>Naresh</given-names></name><name><surname>Laine</surname><given-names>Jennifer</given-names></name><name><surname>Lamar</surname><given-names>Walter E.</given-names></name><name><surname>Lyu</surname><given-names>Jiangnan</given-names></name><name><surname>Kennedy</surname><given-names>Althea</given-names></name><name><surname>Perritano</surname><given-names>Stefan</given-names></name><name><surname>Mason</surname><given-names>Christopher E.</given-names></name><name><surname>Reding</surname><given-names>Brian D.</given-names></name><name><surname>Roca</surname><given-names>Matthew</given-names></name><name><surname>Schürer</surname><given-names>Stephan C.</given-names></name><name><surname>Shukla</surname><given-names>Bhavarth</given-names></name><name><surname>Solle</surname><given-names>Natasha Schaefer</given-names></name><name><surname>Tallon</surname><given-names>John J.</given-names></name><name><surname>Thomas</surname><given-names>Collette</given-names></name><name><surname>Tierney</surname><given-names>Braden T.</given-names></name><name><surname>Torres</surname><given-names>Belkis</given-names></name><name><surname>Venkatapuram</surname><given-names>Sreeharsha</given-names></name><name><surname>Vidović</surname><given-names>Dušica</given-names></name><name><surname>Williams</surname><given-names>Sion L.</given-names></name><name><surname>Yin</surname><given-names>Xue</given-names></name><name><surname>Zarnegarnia</surname><given-names>Yalda</given-names></name><name><surname>Solo-Gabriele</surname><given-names>Helena M.</given-names></name></person-group><article-title>Inclusion of Physical–Chemical Water Quality Measurements Can Improve Associations between SARS-CoV-2 RNA Levels in Wastewater and COVID-19 Cases within Smaller Sewersheds</article-title><source>Journal of Environmental Engineering</source><year>2025</year><volume>151</volume><issue>9</issue><publisher-name>American Society of Civil Engineers (ASCE)</publisher-name><pub-id pub-id-type="doi">10.1061/JOEEDU.EEENG-8138</pub-id></element-citation></ref><ref id="ref9"><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Leisman</surname><given-names>Katelyn Plaisier</given-names></name><name><surname>Owen</surname><given-names>Christopher</given-names></name><name><surname>Warns</surname><given-names>Maria M.</given-names></name><name><surname>Tiwari</surname><given-names>Anuj</given-names></name><name><surname>Bian</surname><given-names>George (Zhixin)</given-names></name><name><surname>Owens</surname><given-names>Sarah M.</given-names></name><name><surname>Catlett</surname><given-names>Charlie</given-names></name><name><surname>Shrestha</surname><given-names>Abhilasha</given-names></name><name><surname>Poretsky</surname><given-names>Rachel</given-names></name><name><surname>Packman</surname><given-names>Aaron I.</given-names></name><name><surname>Mangan</surname><given-names>Niall M.</given-names></name></person-group><article-title>A modeling pipeline to relate municipal wastewater surveillance and regional public health data</article-title><source>Water Research</source><year>2024</year><volume>252</volume><fpage>121178</fpage><publisher-name>Elsevier BV</publisher-name><pub-id pub-id-type="doi">10.1016/j.watres.2024.121178</pub-id></element-citation></ref><ref id="ref10"><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Marsaglia</surname><given-names>George</given-names></name></person-group><article-title>Ratios of Normal Variables</article-title><source>Journal of Statistical Software</source><year>2006</year><volume>16</volume><issue>4</issue><publisher-name>Foundation for Open Access Statistic</publisher-name><pub-id pub-id-type="doi">10.18637/jss.v016.i04</pub-id></element-citation></ref><ref id="ref11"><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Ouyang</surname><given-names>Difan</given-names></name><name><surname>Chung</surname><given-names>Lappui</given-names></name><name><surname>Osborn</surname><given-names>Mark</given-names></name><name><surname>Schacker</surname><given-names>Timothy W.</given-names></name><name><surname>Doss</surname><given-names>Charles R.</given-names></name></person-group><article-title>Dynamic Linear Models for Wastewater-Based Epidemiology with Missing Values: an Application to Covid-19 Surveillance</article-title><source>Data Science in Science</source><year>2025</year><volume>4</volume><issue>1</issue><publisher-name>Informa UK Limited</publisher-name><pub-id pub-id-type="doi">10.1080/26941899.2025.2562199</pub-id></element-citation></ref><ref id="ref12"><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Pellett</surname><given-names>Cameron</given-names></name><name><surname>Farkas</surname><given-names>Kata</given-names></name><name><surname>Williams</surname><given-names>Rachel C.</given-names></name><name><surname>Wade</surname><given-names>Matthew J.</given-names></name><name><surname>Weightman</surname><given-names>Andrew J.</given-names></name><name><surname>Jameson</surname><given-names>Eleanor</given-names></name><name><surname>Cross</surname><given-names>Gareth</given-names></name><name><surname>Jones</surname><given-names>Davey L.</given-names></name></person-group><article-title>Multi-factor normalisation of viral counts from wastewater improves the detection accuracy of viral disease in the community</article-title><source>Environmental Technology &amp; Innovation</source><year>2024</year><volume>36</volume><fpage>103720</fpage><publisher-name>Elsevier BV</publisher-name><pub-id pub-id-type="doi">10.1016/j.eti.2024.103720</pub-id></element-citation></ref><ref id="ref13"><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Rosengart</surname><given-names>AnnaElaine L.</given-names></name><name><surname>Bidwell</surname><given-names>Amanda L.</given-names></name><name><surname>Wolfe</surname><given-names>Marlene K.</given-names></name><name><surname>Boehm</surname><given-names>Alexandria B.</given-names></name><name><surname>Townes</surname><given-names>F. William</given-names></name></person-group><article-title>Spatiotemporal Variability of the Pepper Mild Mottle Virus Biomarker in Wastewater</article-title><source>ACS ES&amp;T Water</source><year>2025</year><volume>5</volume><issue>1</issue><fpage>341</fpage><lpage>350</lpage><publisher-name>American Chemical Society (ACS)</publisher-name><pub-id pub-id-type="doi">10.1021/acsestwater.4c00866</pub-id></element-citation></ref><ref id="ref14"><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Schenk</surname><given-names>Hannes</given-names></name><name><surname>Rauch</surname><given-names>Wolfgang</given-names></name><name><surname>Zulli</surname><given-names>Alessandro</given-names></name><name><surname>Boehm</surname><given-names>Alexandria B.</given-names></name></person-group><article-title>SARS-CoV-2 surveillance in US wastewater: Leading indicators and data variability analysis in 2023–2024</article-title><source>PLOS ONE</source><year>2024</year><volume>19</volume><issue>11</issue><fpage>e0313927</fpage><publisher-name>Public Library of Science (PLoS)</publisher-name><pub-id pub-id-type="doi">10.1371/journal.pone.0313927</pub-id></element-citation></ref><ref id="ref15"><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Sweetapple</surname><given-names>Chris</given-names></name><name><surname>Wade</surname><given-names>Matthew J.</given-names></name><name><surname>Melville-Shreeve</surname><given-names>Peter</given-names></name><name><surname>Chen</surname><given-names>Albert S.</given-names></name><name><surname>Lilley</surname><given-names>Chris</given-names></name><name><surname>Irving</surname><given-names>Jessica</given-names></name><name><surname>Grimsley</surname><given-names>Jasmine M.S.</given-names></name><name><surname>Bunce</surname><given-names>Joshua T.</given-names></name></person-group><article-title>Dynamic population normalisation in wastewater-based epidemiology for improved understanding of the SARS-CoV-2 prevalence: a multi-site study</article-title><source>Journal of Water and Health</source><year>2023</year><volume>21</volume><issue>5</issue><fpage>625</fpage><lpage>642</lpage><publisher-name>IWA Publishing</publisher-name><pub-id pub-id-type="doi">10.2166/wh.2023.318</pub-id></element-citation></ref><ref id="ref16"><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Verani</surname><given-names>Marco</given-names></name><name><surname>Federigi</surname><given-names>Ileana</given-names></name><name><surname>Angori</surname><given-names>Alessandra</given-names></name><name><surname>Pagani</surname><given-names>Alessandra</given-names></name><name><surname>Marvulli</surname><given-names>Francesca</given-names></name><name><surname>Valentini</surname><given-names>Claudia</given-names></name><name><surname>Atomsa</surname><given-names>Nebiyu Tariku</given-names></name><name><surname>Conte</surname><given-names>Beatrice</given-names></name><name><surname>Carducci</surname><given-names>Annalaura</given-names></name></person-group><article-title>Evaluating Population Normalization Methods Using Chemical Data for Wastewater-Based Epidemiology: Insights from a Site-Specific Case Study</article-title><source>Viruses</source><year>2025</year><volume>17</volume><issue>5</issue><fpage>672</fpage><publisher-name>MDPI AG</publisher-name><pub-id pub-id-type="doi">10.3390/v17050672</pub-id></element-citation></ref></ref-list>
</back>
</article>
