Data Historians and Water Quality: Turning Sensor Streams into Records

Technical introduction · Engineering perspective · Published: September 29, 2026

Key Takeaways

  • A data historian is the system of record for everything that changes in a process over time. It collects time-stamped process data from sensors, PLCs, DCS and SCADA systems — commonly over protocols such as OPC UA or Modbus — and records each measurement as a tag-value pair with a timestamp and a quality code [1].
  • Compression is what makes years of high-frequency storage affordable. At 10,000 tags logging once per second, raw storage runs to roughly 27 GB per day (about 10 TB per year uncompressed), and swinging-door compression typically reduces that by a factor of 10-20 [2].
  • In water plants, SCADA historian databases log continuous process data — turbidity, chlorine residuals, CT values, filter performance — at the intervals required under EPA drinking-water rules in 40 CFR Part 141, which is why inspectors increasingly ask for the historian record rather than the logbook [3].
  • A gap in the archive has to stay visible. An interval where no valid values were collected — collector down, communications lost, source stopped answering — must never be quietly interpolated as if the signal simply held steady [4].
  • The exit format already exists. EPA’s Water Quality Exchange (WQX) is the standard framework through which more than 900 federal, state, tribal and other organisations submit water quality monitoring data to the national data warehouse [5].
  • Scale check: a mid-size utility with 100,000 smart meters generating hourly reads produces roughly 8.7 million meter rows per day, and total daily ingest including SCADA telemetry can exceed 10 million rows [6].

What a Historian Is — and What It Is Not

A data historian is specialised software designed to collect, store and retrieve time-stamped process data from industrial equipment — sensors, PLCs, distributed control systems, SCADA. It connects to plant devices over industrial protocols such as OPC UA, OPC DA and Modbus, and records each measurement as a tag-value pair with a timestamp and a quality code, applying industrial compression so that years of high-frequency time series stay storable and queryable. In process industries including water treatment, the historian is the system of record for everything that changes in the process over time [1].

Three clarifications prevent most confusion. First, a historian is not the control system: the PLC decides and acts in milliseconds, while the historian remembers. Second, it is not a general-purpose business database. Transactional databases prioritise per-record accuracy; a historian is built for sequential, high-speed ingestion of time-ordered values and fast time-windowed retrieval. Third, it is not a data lake. The historian holds the raw process truth in engineering units, and analytics platforms build on top of it rather than replacing it.

For water quality, this distinction matters because the measurements themselves — pH, conductivity, turbidity, chlorine residual, dissolved oxygen — are the evidence base for both process control and regulatory reporting. Where that evidence lives, and how honestly it is stored, determines what the plant can prove later.

Compression, Timestamps, and Quality Codes: The Machinery of Evidence

Historians earn their keep with a two-stage filtering scheme that is worth explaining to anyone who has ever wondered why the trend looks “too smooth” or why the archive did not grow for a decade.

Stage one — exception reporting. At the interface, a new value is transmitted upstream only when it differs from the last transmitted value by more than a configured deadband. A flat signal produces almost no traffic; a changing signal reports every meaningful move.

Stage two — swinging-door compression. Inside the archive, algorithms such as the swinging-door trending method store a new point only when the deviation from the straight line connecting previously stored points exceeds a compression deviation. The result is a piecewise-linear approximation whose error is bounded by the configured tolerance. Published engineering arithmetic shows the scale of the effect: 10,000 tags logging once per second generate on the order of 27 GB per day of raw data, about 10 TB per year, and swinging-door compression typically cuts that by 10-20 times [2].

Two settings decide whether this machinery strengthens or weakens a water quality record:

  • Compression deviation is a fidelity decision. For compliance-critical tags, many plants store every value losslessly and reserve compression for fast-moving non-critical signals. The wrong deadband on a chlorine residual tag can flatten a brief excursion that a regulator would want to see.
  • Interpolation semantics decide what a trend display means. Most historians can return either recorded values or interpolated values between recorded points. Both are legitimate views, but a report that silently treats interpolated values as measurements is a report waiting to embarrass its author.

Quality codes are the third pillar. Each stored value carries a flag — good, bad or uncertain [1] — and a disciplined plant makes those flags part of the record rather than an internal detail. When a sensor was being calibrated at 02:00, the data from that window should say so forever.

Where the Historian Sits in the Water-Quality Data Chain

The modern water plant is a layered system, and the historian occupies a specific layer in it. Field sensors measure; PLCs and RTUs control and buffer; the SCADA layer alarms and supervises; the historian archives; reporting and analytics layers consume. In practice, water SCADA systems connect to instruments and controllers over Modbus TCP/IP, DNP3 and OPC UA, and SCADA historian databases log continuous process data — turbidity, chlorine residuals, CT values, filter performance — at the intervals required by EPA drinking-water rules under 40 CFR Part 141 and comparable state rules [3].

The volumes justify the architecture. A mid-size utility with 100,000 AMI endpoints generating hourly reads produces roughly 8.7 million rows of meter data per day, and with SCADA telemetry from plants and pump stations included, daily ingest can exceed 10 million rows — before counting water quality sensors or laboratory samples [6].

Data chain layer Responsibility Selection points that matter
Field sensors and analyzers Measure and convert to signals; the origin of data quality Native digital output (Modbus RTU/TCP or comparable); simultaneous multi-parameter sampling from one probe; self-reporting fault states
PLC / RTU layer Real-time control; communications buffering Store-and-forward on communication loss so no archive gap is created downstream
SCADA / HMI layer Alarms, operator view, setpoints Alarm management discipline; time synchronization across sites
Data historian Time-stamped archive; compression; quality codes Resolution (raw vs compressed per tag), retention horizon, deadband strategy, quality-code visibility
Reporting, analytics, cloud Averages, compliance exports, AI models Standard export paths (SQL, REST); WQX-compatible data structures; governed cloud sync rather than ad-hoc uploads

Bad Values, Gaps, and the Rules That Keep Records Honest

Two failure patterns account for most of the pain in historian-based water quality records, and both have well-established correct answers.

The first is the gap: an interval where the archive holds no valid data because collection stopped — a collector down, communications lost, the instrument offline. The correct handling is explicit. The archive should mark the absence rather than fill it, and retrieval across a gap must never quietly interpolate as though the signal held steady [4]. A gap that is honestly visible is an operations signal: what failed, when, and for how long. A gap that is smoothed over is a fabricated record.

The second is the bad value: a reading that arrived but is wrong, whether from bubbles at the sensor, a drifting probe mid-calibration, or a frozen communications value. The historian’s quality code is the correct vehicle [1]. Mark the value uncertain or bad, keep it in the archive, exclude it from averages knowingly, and record the maintenance action that explained it. Hand-editing historical process data to fill gaps or fix values violates basic data integrity expectations and would not survive an audit; the fix belongs in configuration — redundant sensors, automatic source switching, quality thresholds — not in the data.

Getting Records Back Out: Audits, Reports, and the Cloud

The historian’s payoff arrives when someone asks a question: an inspector, a process engineer, an AI model. Inspectors increasingly want continuous monitoring records rather than spot readings, because a handful of manual entries per shift cannot demonstrate continuous compliance. A historian turns the monthly report from a manual compilation task into an export job [3].

For ambient and environmental reporting, the destination format already exists. EPA’s Water Quality Exchange defines a standard set of data elements to which data partners map their records before submitting them to the national STORET data warehouse, and water quality data submitted from more than 900 federal, state, tribal and other organisations flows through it [5]. A plant whose internal data model — units, timestamps, result qualifiers — mirrors that discipline finds the export trivial. One that invents its own conventions finds it a project.

Cloud platforms fit downstream of the historian, not beside it. The archive stays close to the control system where its writers live; governed, summarised or full-fidelity copies flow outward for analytics. The design goal is best stated in one line: sensors that feed your AI water model, not just your dashboard. That requires the model to see synchronised, quality-flagged, multi-parameter data with honest timestamps — exactly the contract a well-configured historian provides. Shanghai ChiMay’s multi-parameter in-line analyzers are built as sources for that contract, publishing Modbus digital output with simultaneous parameter sampling; the product documentation is collected on Shanghai ChiMay’s online multiparameter water analyzer page.

A Selection Checklist

  • Resolution per tag class. Decide which tags are compliance-critical and store them losslessly; compress the rest deliberately rather than by default.
  • Retention horizon. Match or exceed the record-retention period your regulator requires, and verify that archived (not just online) data remains queryable.
  • Interfaces. OPC UA or DA and Modbus on the collection side [1]; SQL, REST and file exports on the consumption side; WQX-mappable structures for environmental reporting [5].
  • Time discipline. One clock across SCADA, historians and instruments. Timestamp uncertainty is data corruption that no compression algorithm can repair.
  • Quality-code transparency. If the historian knows a value is bad, every consumer — trend, report, model — must be able to know it too.
  • Gap behavior. Confirm the system surfaces gaps as gaps [4], and that exception plus compression settings are documented per tag [2].

The historian is unglamorous infrastructure, which is precisely why it decides outcomes. When the audit comes or the model trains, the plant that can produce a complete, honest, time-stamped record wins on facts.

References

[1] boTec — Process Historian (glossary). Defines the data historian as specialized software that collects, stores, and retrieves time-stamped process data from sensors, PLCs, DCS, and SCADA over protocols such as OPC UA, OPC DA, and Modbus, recording each measurement as a tag-value pair with timestamp and quality code.
https://botec.com/glossary/process-historian

[2] SYMESTIC — Industrial Data Historian: Time-Series DBs & Compression. Works the storage arithmetic for 10,000 tags at one-second logging: roughly 27 GB per day raw and 10 TB per year uncompressed, reduced by a typical factor of 10-20 with swinging-door compression.
https://www.symestic.com/en-us/what-is/industrial-data-historian

[3] NFM Consulting — How SCADA Systems Work in Water Treatment Plants. Describes water SCADA protocols (Modbus TCP/IP, DNP3, OPC UA) and notes that SCADA historian databases log continuous process data — turbidity, chlorine residuals, CT values, filter performance — at intervals required by EPA 40 CFR Part 141 and state rules.
https://www.nfmconsulting.com/knowledge/how-scada-works-water-treatment-plants

[4] Merobix — What Are Data Gaps in Historians? Explains that a data gap is an interval where the archive holds no valid data because collection stopped, and that gaps must be recorded explicitly rather than quietly interpolated as if the signal held steady.
https://www.merobix.com/blog/what-is-a-data-gap-historian

[5] U.S. EPA — Water Quality Data / Water Quality Exchange (WQX). Describes WQX as the universal format for sharing water quality data, with submissions from over 900 federal, state, and tribal agencies and other organizations.
https://19january2025snapshot.epa.gov/waterdata/water-quality-data

[6] TigerData — Water Utilities Database: How to Store and Query SCADA, AMI, and Quality Data at Scale. Estimates a mid-size utility with 100,000 AMI endpoints at roughly 8.7 million meter rows per day, with total daily ingest exceeding 10 million rows when SCADA telemetry is included.
https://www.tigerdata.com/learn/water-utilities-database-how-to-store-query-scada-ami-quality-data-at-scale

About the author: Written by the Shanghai ChiMay Technical Editorial Team — the group behind Shanghai ChiMay’s in-line water quality analyzer documentation, working daily on pH, conductivity, and multi-parameter sensing projects where Modbus-level integration and historian-ready data are part of the product, not an afterthought.