Sporala red del conocimiento
Página 1 de 137Sistema de estimación del nivel del agua de alerta temprana para mitig…
p. 1

Sistema de Estimaci´on del Nivel del Agua de Alerta Temprana Para Mitigar Riesgos Por Inundaciones en Comunidades Ribere˜nas Del Departamento Del Choc´o.

Jackson Berney Renteria Mena Universidad Tecnol´ogica de Pereira Ph.D. in Engineering Facultad de Ingenier´ıa El´ectrica, Electr´onica, F´ısica y Ciencias de la Computaci´on Pereira, Colombia

2025

p. 2

Early Warning Water Level Estimation System For Flood Risk Mitigation In Riparian Communities In The Department Of Choc´o.

Jackson Berney Renteria Mena Thesis submitted as a partial requirement for the degree of: PhD in Engineering Director:

Eduardo Giraldo-Su´arez, Ph.D.

Codirector:

Douglas Plaza Guinglia, Ph.D.

Working area:

Inverse dynamic problems Research group:

Automatic control Universidad Tecnol´ogica de Pereira Ph.D. in Engineering - Automatic and Electronic Line Pereira, Colombia

2025

p. 3

Acknowledgments As I conclude this significant stage of my academic life, I would like to express my most sincere gratitude to all the people and institutions that accompanied me throughout these years of effort and dedication.

First of all, to my wife, whose love, patience and strength sustained me during these four years away from home. Thank you for your unconditional support, for enduring the distance and for always being by my side, despite the difficulties. To my parents and my entire family, I am deeply grateful for your constant support and for being my source of inspiration. Their words of encouragement and their confidence in me have been a fundamental pillar to achieve this goal.

I would also like to express my gratitude to my thesis director, Ph.D. Eduardo Giraldo, and my co-director, Ph.D. Douglas Plaza, for their invaluable guidance during the development of this research. Thank you for your guidance, for helping me to meet the proposed objectives, and for sharing with me your vast knowledge, which has been crucial to successfully complete this thesis.

I also thank the Automatic Control Group of the Electrical Engineering Faculty of the Universidad Tecnol´ogica de Pereira, for welcoming me and making me feel part of this great academic community. To the Universidad Tecnol´ogica de Pereira, my deepest gratitude for receiving me as one of its members and giving me the necessary support for my professional development.

To the Universidad Tecnol´ogica del Choc´o, for its support during these years, providing me with the study commission that allowed me to dedicate myself full time to this doctorate. Without their support, it would not have been possible to complete this path.

Finally, my thanks to the Ministry of Science, Technology and Innovation (Minciencias), for the financial support provided through the Bicentennial Scholarships. Their contribution was vital for the completion of my doctoral studies and has been key in my formation as a researcher.

To all of you, my sincere thanks for your support and contribution to make this achievement a reality.

iii

p. 4

Abstract This dissertation research focuses on the development, implementation and analysis of advanced models for short- and long-term multivariate water level prediction in hydrological systems in the department of Choc´o, Colombia, a region characterized by high rainfall and vulnerability to flooding.

The study combines both nonlinear neural network techniques and traditional linear methods to address the complexity of the hydrological dynamics of the region and improve the accuracy of predictions over various time horizons.

Implemented linear models, such as autoregressive (AR) models and their multivariate variants, are used to establish a baseline for comparison in water level prediction. These models assume linear relationships between hydrological variables such as precipitation, flow and previous water levels, which allows for quick and easy-to-interpret predictions. However, their performance is limited in complex hydrological systems such as those of the Choc´o, where interactions between variables are often nonlinear and more difficult to model accurately.

To overcome these limitations, more advanced models based on neural networks are introduced, such as NARX (Nonlinear Autoregressive with Exogenous Inputs) models and long- and short-term memory (LSTM) neural networks. NARX models allow the inclusion of multiple exogenous variables and capture the inherent nonlinear dynamics of the system, which improves the accuracy of short- and long-term predictions. For their part, LSTM models are especially useful for long-term prediction, as they can retain important information over long time intervals, coping with the complex and seasonal characteristics of the region’s hydrological systems. In addition, data assimilation techniques that combine various sources of hydrological and meteorological information, such as flow, precipitation and water levels, have been applied to reduce uncertainty in predictions. This multivariate integration is essential to improve the accuracy of predictive models, both in the short and long term, in the context of the Choc´o, where climate variability is high and floods can have severe impacts.

A comparative analysis between linear and nonlinear methods highlights that, while linear models can provide reasonable short-term predictions in simple scenarios, nonlinear models, such as neural networks, are considerably more accurate and robust, especially in long-term prediction.

The latter are able to handle multiple variables iv

p. 5

and capture nonlinear interactions, which is crucial for a hydrological environment as dynamic as that of the Choc´o.

v

p. 6

Contents Acknowledgments iii Abstract iv

1

Introduction

1

1.1

Problem statement . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

2

1.2

Justification . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

5

1.3

Objectives . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

8

1.3.1

General Objective . . . . . . . . . . . . . . . . . . . . . . . . . .

8

1.3.2

Specific Objectives . . . . . . . . . . . . . . . . . . . . . . . . .

8

1.4

Study Area

. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

9

1.4.1

Hydrological station measurements . . . . . . . . . . . . . . . .

12

1.5

Hydrologic Modeling . . . . . . . . . . . . . . . . . . . . . . . . . . . .

13

1.5.1

Hydrologic Cycle . . . . . . . . . . . . . . . . . . . . . . . . . .

15

1.5.2

Processes of the hydrological cycle. . . . . . . . . . . . . . . . .

15

1.5.3

Hydrologic Model . . . . . . . . . . . . . . . . . . . . . . . . . .

16

1.6

Thesis Structure . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

17

1.7

Contributions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

17

1.7.1

Chapter 2 Contributions . . . . . . . . . . . . . . . . . . . . . .

18

1.7.2

Chapter 3 Contributions . . . . . . . . . . . . . . . . . . . . . .

18

1.7.3

Chapter 4 Contributions . . . . . . . . . . . . . . . . . . . . . .

19

1.7.4

Chapter 5 Contributions . . . . . . . . . . . . . . . . . . . . . .

20

vi

p. 7

Linear models for water level forecasting:

implementation and analysis

21

2.1

Multivariable AR Data Assimilation for Water Level, Flow, and Precipitation Data

. . . . . . . . . . . . . . . . . . . . . . . . . . . . .

22

2.1.1

AR multivariable model

. . . . . . . . . . . . . . . . . . . . . .

22

2.1.2

Multivarible AR Regularized Solution . . . . . . . . . . . . . . .

23

2.1.3

AR Hydrological model . . . . . . . . . . . . . . . . . . . . . . .

25

2.1.4

Experimental setup . . . . . . . . . . . . . . . . . . . . . . . . .

26

2.1.5

Regularized AR multivariable estimation results . . . . . . . . .

32

2.2

Multivariable ARX based forecasting for an early warning flood system

36

2.2.1

ARX multivariable model

. . . . . . . . . . . . . . . . . . . . .

37

2.2.2

Recursive Data Forecasting

. . . . . . . . . . . . . . . . . . . .

38

2.2.3

Experimental setup ARX data forecasting

. . . . . . . . . . . .

39

2.3

Academic discussion

. . . . . . . . . . . . . . . . . . . . . . . . . . . .

43

2.4

Summary

. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

43

2.5

Partial conclusions

. . . . . . . . . . . . . . . . . . . . . . . . . . . . .

43

3

Nonlinear models for water level forecasting:

implementation and analysis

45

3.1

Multivariable NARX-based neural network models for short-term water level forecasting.

. . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

46

3.1.1

Hydrological Variables . . . . . . . . . . . . . . . . . . . . . . .

46

3.1.2

NARX Based Neural Network Structure

. . . . . . . . . . . . .

47

3.1.3

Experimental Setup . . . . . . . . . . . . . . . . . . . . . . . . .

48

3.1.4

Estimation Results . . . . . . . . . . . . . . . . . . . . . . . . .

48

3.2

Comparative Analysis of Nonlinear Methods for Multivariable Water Level Prediction: The Case Study of the Atrato River . . . . . . . . . .

52

3.2.1

Theoretical framework . . . . . . . . . . . . . . . . . . . . . . .

52

3.2.2

Back-Propagation Nonlinear based Neural Network Structure. .

54

3.2.3

Regression metrics for the estimation of the quadratic error. . .

56

3.2.4

Experimental setup . . . . . . . . . . . . . . . . . . . . . . . . .

57

vii

p. 8

3.2.5

Estimation results . . . . . . . . . . . . . . . . . . . . . . . . . .

58

3.3

Academic discussion

. . . . . . . . . . . . . . . . . . . . . . . . . . . .

62

3.4

Summary

. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

62

3.5

Partial conclusions

. . . . . . . . . . . . . . . . . . . . . . . . . . . . .

62

4

Water Level forecasting by Hybrid Models

64

4.1

Multivariate Hydrological Modeling Based on Long Short Term Memory Networks for Water Level Forecasting . . . . . . . . . . . . . . . . . . .

65

4.1.1

Experimental setup . . . . . . . . . . . . . . . . . . . . . . . . .

66

4.1.2

LSTM (Long Short-Term Memory) Network . . . . . . . . . . .

66

4.1.3

Estimation Results. . . . . . . . . . . . . . . . . . . . . . . . . .

69

4.1.4

Forecasting Future Time Steps Based on Predictions and Forecasting Future Time Steps Based on Measurements . . . . .

71

4.2

Water Level Forecasting based on an Ensemble Kalman Filter with NARX Neural Network Model . . . . . . . . . . . . . . . . . . . . . . .

75

4.2.1

NARX neural network model for hydrologycal level forecasting .

76

4.2.2

EnKF forecasting . . . . . . . . . . . . . . . . . . . . . . . . . .

78

4.2.3

Novelty of the proposed model NARX-EnKF . . . . . . . . . . .

79

4.2.4

Experiment setup . . . . . . . . . . . . . . . . . . . . . . . . . .

80

4.2.5

NARX-EnKF Neural Network Validation Results

. . . . . . . .

81

4.3

Academic discussion

. . . . . . . . . . . . . . . . . . . . . . . . . . . .

92

4.4

Summary

. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

92

4.5

Partial conclusions

. . . . . . . . . . . . . . . . . . . . . . . . . . . . .

92

5

Mechanism of Social Appropriation and Knowledge Transfer for the Riverside Communities of Choc´o through a web platform and Virtual Environment

94

5.1

Choc´o Early Warning Web Estimation System (SEAT-Web) . . . . . .

95

5.1.1

SEAT-Web Key Functionalities . . . . . . . . . . . . . . . . . .

95

5.1.2

Elements of SEAT-Web

. . . . . . . . . . . . . . . . . . . . . .

95

5.2

Virtual Environment of an Early Warning System to Mitigate Flood Risks in the Department of Choc´o . . . . . . . . . . . . . . . . . . . . .

99

viii

p. 9

5.2.1

Design and Construction of the 3D Model of the Hydrological Station . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

100

5.2.2

General aspect of the environment

. . . . . . . . . . . . . . . .

107

5.3

Software registration . . . . . . . . . . . . . . . . . . . . . . . . . . . .

108

5.4

Summary

. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

109

5.5

Partial conclusions

. . . . . . . . . . . . . . . . . . . . . . . . . . . . .

109

6

Conclusions and Final Remarks

110

6.1

Conclusions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

110

6.2

Future works

. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

111

6.3

Academic Discussion . . . . . . . . . . . . . . . . . . . . . . . . . . . .

113

References114 ix

p. 10

List of Figures

1.1

Description of the entire course of the Atrato River in the Department of Choc´o - Colombia.

. . . . . . . . . . . . . . . . . . . . . . . . . . .

10

1.2

The location of hydrological stations in the Atrato River

. . . . . . . .

11

1.3

The location satelital of hydrological stations in the Atrato River

. . .

11

1.4

Level measurements at the three stations . . . . . . . . . . . . . . . . .

12

1.5

Water flow measurements at the three stations . . . . . . . . . . . . . .

12

1.6

Precipitation measurements at the three stations . . . . . . . . . . . . .

13

1.7

Results of the implemented hybrid hydrologic models.

. . . . . . . . .

17

2.1

Selection of the regularization parameter λ by using the GCV method for Level variable. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

26

2.2

Comparison of the estimated signals by using the real level data, ΘL LS and the ΘL Tikh, order 10.

. . . . . . . . . . . . . . . . . . . . . . . . . .

27

2.3

Selection of the regularization parameter λ by using the GCV method for Flow variable. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

28

2.4

Comparison of the estimated signals by using the real Flow data, ΘF LS and the ΘF Tikh, order 10 . . . . . . . . . . . . . . . . . . . . . . . . . . .

29

2.5

Selection of the regularization parameter λ by using the GCV method for Precipitation variable.

. . . . . . . . . . . . . . . . . . . . . . . . .

29

2.6

Comparison of the estimated signals by using the real Precipitation data, ΘP LS and the ΘP Tikh, order 10. . . . . . . . . . . . . . . . . . . . . . . . .

30

2.7

Comparison of the estimated signals by using the real Level data, ΘP LS and the ΘP Tikh, order 30.

. . . . . . . . . . . . . . . . . . . . . . . . . .

31

2.8

Comparison of the estimated signals by using the real Flow data, ΘP LS and the ΘP Tikh, order 30.

. . . . . . . . . . . . . . . . . . . . . . . . . .

31

x

p. 11

2.9

Comparison of the estimated signals by using the real Precipitation data, ΘP LS and the ΘP Tikh, order 30. . . . . . . . . . . . . . . . . . . . . . . . .

32

2.10 Selection of the regularization parameter λ by using the GCV method

for Level variable . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

33

2.11 Selection of the regularization parameter λ by using the GCV method

for Flow variable

. . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

33

2.12 Selection of the regularization parameter λ by using the GCV method

for Precipitation variable . . . . . . . . . . . . . . . . . . . . . . . . . .

34

2.13 Level variable estimation error comparison for the univariable AR model

with the least squares method (ELS) and the Tikhonov method (ELST), and for the multivariable AR model with the least squares method (ELM), and the Tikhonov method (ELMT). . . . . . . . . . . . . . . .

35

2.14 Flow variable estimation error comparison for the univariable AR model

with the least squares method (EFS) and the Tikhonov method (EFST), and for the multivariable AR model with the least squares method (EFM), and the Tikhonov method (EFMT). . . . . . . . . . . . . . . .

36

2.15 Precipitation variable estimation error comparison for the univariable AR

model with the least squares method (EFS) and the Tikhonov method (EFST), and for the multivariable AR model with the least squares method (EFM), and the Tikhonov method (EFMT).

. . . . . . . . . .

36

2.16 Response Recursive ARX coupled model for a fourth order system for

water level of E1. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

39

2.17 Response Recursive ARX coupled model for an order 4 system for water

level of E2.

. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

39

2.18 Response Recursive ARX coupled model for an order 4 system for level

of E3.

. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

40

2.19 Example of parameter evolution for some elements of A1 matrix . . . .

40

2.20 Quadratic estimation error response of the coupled recursive ARX model

(6 inputs, 3 outputs), for system of order 1, order 2, order 3 and order 4.

41

2.21 Computational time of algorithm execution by each ARX Coupled

system orders. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

41

2.22 Quadratic estimation error response of the coupled recursive ARX model

(6 inputs, 3 outputs) and the decoupled recursive ARX model (2 inputs, 1 outputs) for system of order 4. . . . . . . . . . . . . . . . . . . . . . .

42

xi

p. 12

2.23 Quadratic estimation error response of the coupled recursive ARX model

(6 inputs, 3 outputs), compared to the AR model ( 3 outputs) for system of order 4. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

42

3.1

NARX based Neural Networks Structure . . . . . . . . . . . . . . . . .

47

3.2

ARX Structure . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

47

3.3

First water level output and short-term estimation based on a multivariable ARX model. . . . . . . . . . . . . . . . . . . . . . . . . .

49

3.4

First water level output and short-term estimation based on a multivariable ARX model. . . . . . . . . . . . . . . . . . . . . . . . . .

49

3.5

First water level output and short-term estimation based on a multivariable NARX model. . . . . . . . . . . . . . . . . . . . . . . . .

50

3.6

First water level output and short-term estimation based on a multivariable NARX model. . . . . . . . . . . . . . . . . . . . . . . . .

50

3.7

Training error with the ARX multivariable model. . . . . . . . . . . . .

51

3.8

Training error with the NARX multivariable model. . . . . . . . . . . .

51

3.9

Structure of Neural Networks Based on NARX Nonlinear Model . . . .

53

3.10 Structure of Neural Networks Based on ARX linear Model . . . . . . .

54

3.11 Structure of Neural Networks Based on Back-Propagation nonlinear Model 55

3.12 Methodological architecture Backpropagation, NARX and ARX . . . .

56

3.13 The output for water level 1 and the short-term estimation are

determined using a linear multivariate least squares ARX model.

. . .

58

3.14 The output for water level 2 and the short-term estimation are

determined using a linear multivariate least squares ARX model.

. . .

59

3.15 Water level output 1 and short-term estimation based on a nonlinear

multivariate NARX model.

. . . . . . . . . . . . . . . . . . . . . . . .

59

3.16 Water level output 2 and short-term estimation based on a nonlinear

multivariate NARX model.

. . . . . . . . . . . . . . . . . . . . . . . .

60

3.17 Water level output 1 and short-term estimation based on a nonlinear

multivariate Back-Propagation model.

. . . . . . . . . . . . . . . . . .

60

3.18 Water level output 2 and short-term estimation based on a nonlinear

multivariate Back-Propagation model.

. . . . . . . . . . . . . . . . . .

61

xii

p. 13

4.1

Multivariate schematic of the long short-term memory (LSTM) neural network.

. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

66

4.2

Architecture of a regression LSTM neural network:summarizes a visual depiction detailing the structure of a Long Short-Term Memory (LSTM) neural network designed for regression tasks. It visually communicates the network’s components, including input data, LSTM layers for capturing temporal patterns, potential hidden layers, and the output layer for continuous prediction, facilitating the understanding of the network’s design and function in your study.

. . . . . . . . . . . . . .

67

4.3

Data flow structure in time unit t:this diagram illustrates the flow of data in time unit t. This diagram shows how the gates forget, update and generate the cell and hidden states. . . . . . . . . . . . . . . . . . .

67

4.4

Short-term Estimation of water level Bel´en of Bajira at hydrological station 1 using an LSTM model.

. . . . . . . . . . . . . . . . . . . . .

69

4.5

Short-term Estimation of water level Quibd´o at hydrological station 2 using an LSTM model.

. . . . . . . . . . . . . . . . . . . . . . . . . .

69

4.6

Water level 1 output from hydrological station 1, located in the town of Bel´en de Bajir´a time step response based on future predictions.

. . . .

72

4.7

Water level 2 output from hydrological station 2, located in the town of Bel´en de Bajir´a time step response based on future predictions.

. . . .

72

4.8

Water level 1 output from hydrological station 1, located in the town of Bel´en de Bajir´a response of future forecasts of time steps based on measurements.

. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

73

4.9

Water level 2 output from hydrological station 2, located in the town of Quibd´o response of future forecasts of time steps based on measurements.

73

4.10 Water level 1 output from hydrological station 1, located in the town of

Bel´en de Bajir´a time step response based on future predictions.

. . . .

74

4.11 Water level 2 output from hydrological station 2, located in the town of

Bel´en de Bajir´a time step response based on future predictions.

. . . .

74

4.12 Methodological architecture EnKF-NARX . . . . . . . . . . . . . . . .

77

4.13 Architecture of NARX, NARX-EnKF . . . . . . . . . . . . . . . . . . .

78

4.14 NARX neural network diagram without EnKF.

. . . . . . . . . . . . .

81

4.15 NARX neural network diagram coupled EnKF.

. . . . . . . . . . . . .

81

4.16 NARX-EnKF Feedforward model response . . . . . . . . . . . . . . . .

81

4.17 NARX-EnKF Best model performance response . . . . . . . . . . . . .

82

xiii

p. 14

4.18 NARX-EnKF Visual representation of the model fitted to the data. . .

82

4.19 NARX-EnKF model validation. . . . . . . . . . . . . . . . . . . . . . .

83

4.20 Description of the Diagram Taylor. . . . . . . . . . . . . . . . . . . . .

83

4.21 Output 1 water level forecast for an order 4 system of a multivariate

NARX model coupled to an EnKF filter with noise variance parameter of 0.01. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

84

4.22 Output 2 water level forecast for an order 4 system of a multivariate

NARX model coupled to an EnKF filter with noise variance parameter of 0.01. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

84

4.23 Output 1 water level forecast for an order 4 system of a multivariate

NARX model coupled to an EnKF filter with noise variance parameter of 0.1.

. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

85

4.24 Output 2 water level forecast for an order 4 system of a multivariate

NARX model coupled to an EnKF filter with noise variance parameter of 0.1.

. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

85

4.25 Output 1 water level forecast for an order 4 system of a multivariate

NARX model coupled to an EnKF filter with noise variance parameter of 0.4.

. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

86

4.26 Output 2 water level forecast for an order 4 system of a multivariate

NARX model coupled to an EnKF filter with noise variance parameter of 0.4.

. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

86

4.27 Output 1 water level forecast for an order 4 system of a multivariate

NARX neural network model without recursive nonlinear EnKF with noise variance parameter of 0.01.

. . . . . . . . . . . . . . . . . . . . .

87

4.28 Output 2 water level forecast for an order 4 system of amultivariate

NARX neural network model without recursive nonlinear EnKF with noise variance parameter of 0.01.

. . . . . . . . . . . . . . . . . . . . .

87

4.29 Output 1 water level forecast for an order 4 system of a multivariate

NARX neural network model without recursive nonlinear EnKF with noise variance parameter of 0.1. . . . . . . . . . . . . . . . . . . . . . .

88

4.30 Output 2 water level forecast for an order 4 system of amultivariate

NARX neural network model without recursive nonlinear EnKF with noise variance parameter of 0.1. . . . . . . . . . . . . . . . . . . . . . .

88

4.31 Output 1 water level forecast for an order 4 system of a multivariate

NARX neural network model without recursive nonlinear EnKF with noise variance parameter of 0.4. . . . . . . . . . . . . . . . . . . . . . .

89

xiv

p. 15

4.32 Output 2 water level forecast for an order 4 system of amultivariate

NARX neural network model without recursive nonlinear EnKF with noise variance parameter of 0.4. . . . . . . . . . . . . . . . . . . . . . .

89

5.1

SAT Control Panel. . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

97

5.2

Symbology of elements on the map. . . . . . . . . . . . . . . . . . . . .

97

5.3

Atrato river water level monitoring station 1. . . . . . . . . . . . . . . .

98

5.4

Atrato river water level monitoring station 2. . . . . . . . . . . . . . . .

98

5.5

Atrato river water level monitoring station 2. . . . . . . . . . . . . . . .

99

5.6

Atrato river water level monitoring station 3. . . . . . . . . . . . . . . .

99

5.7

3D Model Hydrological Station. . . . . . . . . . . . . . . . . . . . . . .

101

5.8

Base Hydrological Station. . . . . . . . . . . . . . . . . . . . . . . . . .

102

5.9

3D Model Ultrasonic Sensor. . . . . . . . . . . . . . . . . . . . . . . . .

102

5.10 3D Design of the Flow Sensor. . . . . . . . . . . . . . . . . . . . . . . .

103

5.11 3D Rain Gauge Design. . . . . . . . . . . . . . . . . . . . . . . . . . . .

104

5.12 3D Data Logger design. . . . . . . . . . . . . . . . . . . . . . . . . . . .

105

5.13 3D Design Atrato River - Unity. . . . . . . . . . . . . . . . . . . . . . .

105

5.14 Incorporation of Hydrological Station in Unity. . . . . . . . . . . . . . .

106

5.15 Data Visualization Center. . . . . . . . . . . . . . . . . . . . . . . . . .

106

5.16 Hydrologic Station Assembly and Disassembly. . . . . . . . . . . . . . .

107

5.17 Control Station. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

107

5.18 Assembly and Disassembly Area.

. . . . . . . . . . . . . . . . . . . . .

108

5.19 Implementation Zone.

. . . . . . . . . . . . . . . . . . . . . . . . . . .

108

xv

p. 16

List of Tables

1.1

Information of the used data sets. . . . . . . . . . . . . . . . . . . . . .

12

1.2

Location of the hydrological stations

. . . . . . . . . . . . . . . . . . .

13

1.3

Comparative Hydrological Models.

. . . . . . . . . . . . . . . . . . . .

16

1.4

Comparative performance of hydrological models

. . . . . . . . . . . .

16

3.1

Mean squared Estimation error for several nodes configurations

. . . .

49

3.2

Mean squared Estimation error

. . . . . . . . . . . . . . . . . . . . . .

51

3.3

Output 1

. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

61

3.4

Output 2. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

61

3.5

Model parameters of the nonlinear neural networks and linear model . .

61

4.1

RMSE, NSE for a system long short-term and computational execution time, Tic operates with the Toc function to measure the elapsed time. .

70

4.2

Model parameters of the nonlinear neural networks. . . . . . . . . . . .

71

4.3

RMSE and NSE for an LSTM system of future forecasts and future forecasts based on measurements.

. . . . . . . . . . . . . . . . . . . .

75

4.4

Parameters of the NARX model, NARX-EnKF. . . . . . . . . . . . . .

89

4.5

Root mean square estimation (RMSE) with Gaussian noise for NARX neural network coupled with EnKF and without EnKF.

. . . . . . . .

90

4.6

Nash-Sutcliffe model efficiency coefficient (NSE) with Gaussian noise for the NARX neural network coupled with EnKF and decoupled without EnKF. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

91

5.1

Description of the elements of the early warning system.

. . . . . . . .

96

xvi

p. 17

Chapter 1 Introduction Flooding constitutes one of the most critical natural hazards affecting riparian communities, particularly in regions like the Department of Choc´o, with a dense fluvial network and a humid tropical climate [1]. Populations residing along the rivers in this area are especially vulnerable due to the recurrent occurrence of flash floods, often intensified by extreme weather events and the growing effects of climate change [2]. These flood events pose not only a threat to human life, but also cause extensive damage to infrastructure, agriculture, and ecosystems—resulting in substantial economic losses and a decline in the quality of life [3]. Historically, flood response strategies in Choc´o have been predominantly reactive, based on local knowledge and direct observation of river behavior [4]. While these practices hold value, they are often insufficient to anticipate flooding with the necessary lead time, leading to delayed interventions and increased adverse impacts [5]. In light of this, implementing a system for estimating river water levels is proposed as a fundamental tool to enhance flood risk mitigation, aiming to improve preparedness and reduce associated damages [6]. The proposed system is designed as an early warning mechanism, enabling real-time monitoring of river water levels and providing accurate forecasts of potential flood events [7].

It integrates data from local meteorological stations with in situ data acquisition technologies—including water level, flow, and rainfall sensors—deployed through automated hydrological stations [8]. When combined with advanced predictive algorithms, this information allows the system to go beyond issuing alerts, delivering detailed forecasts that empower local authorities and communities to take preventive action in a timely and informed manner [9].

The expected impact of this system on Choc´o’s riparian communities is considerable.

It is anticipated to reduce flood vulnerability and enhance adaptive capacity and resilience in the face of climate-related threats [10]. Moreover, by delivering accurate, real-time information, the system will support more effective planning and risk management, facilitating better coordination among local authorities and emergency response agencies [11].

p. 18

1.1

Problem statement When designing and operating water resource systems, engineering always considers the variability and uncertainty of hydrological factors.

Precipitation, river flow, transpiration, and groundwater runoff are unpredictable processes [12]. The sequence of hydrological events rarely repeats itself.

Faced with the imminent decision to design reliable water supply structures, engineers traditionally rely on the so-called frequency analysis, whereby, by assuming the independence of events, and using different parametric or non-parametric methods, the probabilities of significant events can be obtained. This probability, or average recovery period, helps to select events to design control engineering structures through estimation algorithms. The hydrological systems analyst extends the above concepts to mathematical formulas for handling uncertain input, posing estimation algorithms, and linear programming with refinement models.

These models are designed to determine the optimal configurations for developing and operating irrigation engineering facilities across the basin.

They are based on the same concepts of repeatability and reliability as in traditional frequency analysis.

The Choc´o department, located in western Colombia, is considered one of the natural regions with the highest annual rainfall levels in the world [13]. Three large rivers and their tributaries cross it, Atrato, Baud´o, and San Juan, which are the main sources of surface water, communication routes, and areas of human settlement, as the three rivers pass through the 31 municipalities [14]. Historically, the population has settled on the banks of these rivers, exposing them to the natural phenomena associated with bodies of water, especially flooding.

The stories of indigenous and Afro-Colombian people settled along the banks of the principal rivers of the Choc´o department and their tributaries attest to the way floods have affected and shaped the way of life of the people who inhabit them [15], especially anxiety due to the unexpectedness of a flood that usually leaves severe structural damage to homes and material losses in terms of belongings and other materials, and sometimes even life itself. Some river flow monitoring points in some municipalities in the Choc´o geography department were manual, automatic, and centralized by the Institute of Hydrology, Meteorology, and Environmental Studies (IDEAM)[16].

This system manages the information from the periodic measurements taken on the rivers’ behavior in the Choc´o department; therefore, they do not provide constant and immediate information, but every day or every 12 hours. This means that the system is not the most suitable for early warning systems for floods, since data is not obtained in real time, but for a period that is not optimal for decision-making. This is why, in addition to developing early warning systems, these systems must make their own estimates of future data according to the data provided by the measurement station [17], through fast algorithms to estimate robust data, such as algorithms, models, and methods for system identification, such as least squares. This model type is significant in adaptive control systems and parameter estimation, in which plant parameters must be estimated online to calculate the corresponding controller [18].

p. 19

Another relevant approach is the autoregressive model (AR) or autoregressive with exogenous variables (ARX) regularized by the Tikhonov method. This type of modeling seeks to estimate the parameter vector using regularization techniques applicable to ill-conditioned problems.

In particular, Tikhonov regularization imposes a uniform penalty on all parameters, which mitigates the multicollinearity problem characteristic of linear regressions with many variables.

As a result, a more robust and efficient estimate of the parameters is obtained, at the cost of accepting a controlled and tolerable margin of error in the accuracy of the model [19]. The Expectation Maximization (EM) algorithm finds maximum likelihood or maximum a posteriori estimates of parameters in statistical models, where the model depends on unobserved latent variables [20]. The Bayesian algorithm allows us to calculate the explicit probabilities for each hypothesis [21]. The Kalman filter allows us to estimate parameters and decrease the noise of the real signal; in addition, it produces estimates of hidden variables based on inaccurate and uncertain measurements and provides a prediction of the future state of the system, based on previous estimates [22].

Accurate modeling of hydrological variables is crucial for effective flood forecasting and the design and operation of water resource systems, as indicated in the reference [12]. To achieve this, two types of techniques are typically used: white-box algorithms, which are based on mathematical modeling, and black-box algorithms, which employ nonlinear neural network techniques based on artificial intelligence.

The latter technique has been used with great success in early warning systems, as highlighted in [23], which discusses the use of artificial neural networks (ANN) for prediction and forecasting. In addition, the combination of ANN with the Soil and Water Assessment Tool (SWAT) has been applied to predict runoff and water resources, as described in [24]. Hydrological models have been developed to support flood early warning systems through estimation and prediction algorithms. In [25], the authors focus on the development of a flood forecasting system (FFS) capable of providing early warning to UDS managers of potential flooding, using a nonlinear autoregressive neural network with exogenous input (NARX) to predict the impact of a storm.

Meanwhile, in [26], a short-term memory model of a neural network (LSTM) is proposed for flood prediction, using daily flow and rainfall as input data, and analyzing the features that can affect model performance.

In [27], the authors present flood early warning systems using machine learning (ML) techniques, comparing the performance of five ML classification techniques for short-duration flood forecasting. Finally, [28] describes an IoT-based flood monitoring and prediction system based on artificial neural networks (ANN), which aims to monitor the humidity, temperature, pressure, rainfall and the level of the water of rivers and analyze their temporal correlation for flood prediction. The system is designed to improve the scalability and reliability of the flood management system. Several flood prediction and monitoring studies have also been developed using various ANN techniques. For example, in [29], the authors evaluated the bias correction of real-time precipitation data and improved hydrological models using the ANN bias correction method for real-time flood prediction. In [30], the authors focus on developing five different ANN models for flood prediction and comparing their performance. In [31], the authors use a multilayer perceptron to design a flood prediction model with flow

p. 20

as input-output variables, and the effectiveness of the proposed model is demonstrated through intensive experiments. In [32], the authors designed a flood monitoring system that integrates flow and water level sensors and uses a two-class neural network to predict flood status from data stored in the database.

Lastly, in [33], the authors employ a convolutional neural network (CNN) to predict time series variables such as water level in a flood model. However, CNNs are typically used for two-dimensional image classification with transfer learning. In addition, in [34], a flood prediction model is designed and evaluated that predicts future flood occurrence by building a hybrid deep learning algorithm called ConvLSTM, which integrates the predictive merits of CNN and the long-term memory network (LSTM). In [35], a fuzzy neural network is proposed using fuzzy numbers to account for uncertainty in the results and model parameters to predict peak flow in an urban river. In [36], the potential of the AI computational paradigm for modeling river flow is explored by developing nine different flood prediction models using all available training algorithms from ANN, fuzzy logic and adaptive neuro-fuzzy inference systems (ANFIS). Finally, in [37], a deep neural network is used to predict floods as a function of temperature and rainfall intensity, and its precision and error are compared with other machine learning models, such as the support vector machine (SVM), the nearest neighbor K (KNN) and the Naive Bayes.

While purely physical models (white box) and artificial intelligence models (black box) have proven effective in various hydrological prediction applications, their integration through hybrid approaches offers a promising alternative, especially in complex contexts such as the Choc´o department. Among these emerging approaches is gray box modeling, which combines knowledge of the physical system with machine learning capabilities to capture nonlinear relationships from the data [38, 39]. This modeling approach is beneficial when the available information is incomplete or uncertain, but the partial physical structures of the system are known [40]. Its application could, for example, incorporate existing hydrological knowledge of the Atrato river basin while dynamically adjusting parameters with neural networks trained in real time. Similarly, a recent and exciting trend in hydrology is using physics-informed machine learning (PIML). These techniques incorporate physical laws such as conservation of mass or energy directly into the training process of AI models [41]. Unlike traditional neural networks, PIML models are not limited to fitting empirical data but are guided by physical constraints, improving their generalization, robustness, and interpretability [42]. In scenarios where data availability is limited, such as rural or hard-to-access areas like Choc´o, these techniques could generate more accurate and reliable predictions, even under extreme or unseen conditions. The integration of these advanced approaches would not only broaden the methodological scope of the work but also enable the construction of more reliable and adaptive early warning systems capable of anticipating extreme events in highly variable hydrometeorological environments. Given the region’s historical vulnerability to flooding, their incorporation represents a valuable opportunity to strengthen risk management and territorial planning. Based on the above, the following question can be formulated: How can an early warning

p. 21

water level estimation system for flood risk mitigation in riparian communities in the Choc´o Department be developed using data estimation algorithms?

1.2

Justification This proposal is intended to provide an alternative solution to the problem posed, which seeks to develop an innovative estimation system to measure and provide information on river floods in real time, thus facilitating immediate decision-making by government and community entities in the event of a possible flood. Secondly, monitoring the river and its behavior ceases to be an exclusive activity for a few technicians. It becomes a social community activity to acquire technology that involves most of the population, who are usually affected by floods. Therefore, we want to implement an early warning system that includes computer language algorithms, either in MATLAB or Python software, along with methods and data estimation models such as: Auto-Regressive model (AR), least squares method, expectation maximization (EM) algorithm, and naive Bayes algorithm, among others. In addition, an early warning system is implemented with the latest communication technologies in the field of IoT, such as LoRaWAN and the General Packet Radio Service (GPRS).

It should be noted that developing a hydrological model involves a structured process of steps.

In the literature, different methodological approaches can be found to create hydrological models; one of the most influential is the methodological approach presented by Beven [43].In it, the modeling process is divided into five steps, in a workflow that can be iterative, depending on the success of the final modeling. The hydrological modeling process begins with the conceptual model, which allows us to decide which processes will be the most important to represent, which could be left out, and what level of detail is necessary to simulate the dynamics of a system. These perceptions are then formalized in a conceptual model, a mathematical abstraction representing the hydrologic processes in equations. In the next stage, called the procedural model, the equations of the conceptual model are translated into an algorithm to solve them. This stage involves the development of the equation-solving code and its execution to achieve numerical results.

After having decided on the estimation algorithm that implements the equations, we proceed to calibrate the equations - proceed to calibrate the parameters of the model [44]—which describe and summarize the properties of the area of the hydrological system. Finally, the performance of a calibrated model is evaluated in a process called validation [45]. With the calibrated model, one or more simulations are performed, which are compared with precipitation, flow, and level observations.

The modeling of hydrological variables is considered of great relevance for meteorological forecasting, especially when designing and operating water resource systems. Engineers are always aware of the variability and uncertainty of hydrological interactions, using different parametric or nonparametric methods [12]. The development of a hydrological

p. 22

model involves a structured process of steps. As reported in the literature, various methodological approaches are applied to formulate hydrological models. For example, in [43], the modeling process is divided into a workflow that can be iterative, depending on the success of the final modeling objective [43]. Moriasi et al. [46] establish widely adopted guidelines for the evaluation of hydrological model performance, proposing key metrics such as NSE, PBIAS and RSR with acceptance thresholds. Their approach has been instrumental in standardizing and quantifying accuracy in watershed simulations and is considered a methodological reference in applied hydrology. The hydrological modeling process begins with the perceptual model that allows us to decide which methods will be the most important to represent, which might be omitted, and what level of detail might be necessary to simulate the dynamics of a system. These perceptions are then formalized in a conceptual model, a mathematical abstraction of the physical processes.

This mathematical abstraction involves representing the hydrological processes in equations. In the next stage, called the procedural model, the equations of the conceptual model are translated into an algorithm to solve the problem. This stage involves the development of the solution code and its execution to obtain numerical results. After deciding on the estimation algorithm that implements the equations, the equations are calibrated by calibrating the model parameters [44], which describe and summarize the properties of the hydrological system zone. Finally, we evaluate the performance of the calibrated model using validation scenarios. With the calibrated model, simulations are carried out [45] to analyze the behavior of the models.

Hydrologic models have long been developed to provide a timely solution to flood early warning systems using black-box and white-box estimation and prediction algorithms. In [47], a data-driven approach is used to model river level behavior and develop river flood prediction using a piecewise linear system. However, the proposed model is not used for a multivariate system. In [27], flood early warning systems using Machine Learning (ML) techniques are presented, where a comparison of the performance of five ML classification techniques is shown for short-duration flood prediction. The study of [48] presents flood forecasting with machine learning models in an operational framework, where ML models are evaluated to predict river and flood conditions. In addition, the study presents a new ML methodology for modeling flood extent and depth.

The study of [49] proposes the evaluation of four hydrological models for operational flood forecasting in a Canadian prairie watershed. Here, four hydrological models (WATFLOOD, HBV-EC, HSPF, and HEC-HMS) were used, calibrated, and tested for the study of performance and precision used in operational flood forecasting tools.

In [50], the estimation of the water balance of the Colombian Pacific region is addressed.

In this research, the surface water balance equation is proposed and the average annual runoff is estimated in the continental zone of the geographic area.

In [51], the impacts of intervariable correlation of the climate model output are presented.

Specifically, the study aims to evaluate the effects of correcting the intervariate correlation of climate model outputs on hydrological modeling. The study [52] uses sequential data assimilation for real-time probabilistic flood mapping. This research uses real-time probabilistic flood mapping techniques for risk alerts and

p. 23

decision-making. The research of [53] deals with the design of real-time early warning systems for flash floods. This study reviews the major parties of early warning system techniques and a basic structure of an EWS for flash pluvial floods. In the review article [54], the application of remotely detected data to constrain operational rainfall flood forecasting, this article reviews recent advances in the integration of remotely sensed precipitation and soil moisture with rainfall-rainfall models for rainfall flood forecasting [54].

Following trends in the modeling approach, a review of studies using real-time black-box modeling is presented below. In [55], the architecture and use of a system are described to meet design requirements and enable model-based control, thus optimizing system predictability. In [56], a data-driven approach is presented as an operational model for real-time flood forecasting.

The study adopts an adaptive network-based fuzzy inference system (ANFIS), a data-driven approach to forecast daily river water levels using AR models. In [57], flood flow modeling in a river system using an adaptive neurofuzzy inference system, the paper presents the application of a data-driven model, the adaptive neuro-fuzzy inference system (ANFIS), in flood flow forecasting in a river system. In [58], early warning and flood forecasting for large rivers with the Lower Mekong as an example, this research describes the components of an early warning system. Provides details for designing a data-driven flood forecasting model. In [59], with a participatory approach, early warning systems are analyzed for risk management in Colombia. The paper reflects on various national and international approaches and experiences in warning systems. One of the main conclusions of [59] was that there are technological and operational limitations to implement state-of-the-art real-time early warning systems in the country under study.

An alternative solution to model the behavior of hydrological variables is presented in [60], where a land data assimilation system is proposed for food and water security applications in sub-Saharan Africa.

However, the proposed approach is based on satellite remote sensing and land surface models. On the other hand, in [61], a predictive control of a latent variable iterative learning model is presented for multivariate control of batch processes that can be applied to several hydrological models and variables. In [62], a weighted parameter estimation for nonlinear self-regressive Hammerstein systems with eXogenous inputs (ARX) is proposed that can be used to model the nonlinear characteristics of hydrological processes. Finally, in [63], a Multiple Input Single Output (MISO) and Autoregressive Moving Average with eXogenous Inputs (ARMAX) ARX model of a flood forecasting system is discussed, in which the proposed approach adequately follows the dynamics of the system.

In addition, in [64], an adaptive multivariate ARMAX model is also used to describe the dynamic behavior of a wastewater treatment plant. The above methods show a suitable option for modeling the behavior of hydrological variables to build an early warning system for the short term.

Climate phenomena, such as El Ni˜no and La Ni˜na, permanently impact the communities of the Colombian Pacific. Creating uncertainty due to flooding processes caused by the growth of the surface levels of the department rivers. This behavior is worth predicting

p. 24

to foresee possible disasters and help river communities be prepared. Betting on systems of immediate or constant data collection of natural phenomena and their behavior allows monitoring river levels and phenomena associated with their gradual or untimely growth, such as river flow, level, and precipitation. Meanwhile, an increase in fragility has been observed in the territories and sectors of the department of Choc´o; after the catastrophic floods, recovery is increasingly complex and repetitive, which leads to the need to develop an early warning system for floods: placing an adequate barrier and reducing that matches these needs of the community, based on an analysis of vulnerability as a risk assessment. It should be noted that floods are a recurring problem in the Choc´o department, generating the loss of human lives and material goods to the river communities in the department. For this reason, we want to develop an early warning prototype that consists of an innovative research proposal, compared to the existing ones on the market, and that fulfills the functions of predicting and alerting the population to imminent dangers due to overflowing of rivers. The selection of this problem as a project is motivated by our professional projection to the communities, presenting process automation solutions in an environment that lacks the latest technology and the development of applications of this technology, as is the Choc´o department, where none of its municipalities or river communities have this type of system for monitoring the department rivers and predictive estimator algorithms for the variables of flow, level and precipitation.

1.3

Objectives

1.3.1

General Objective Develop an Early Warning Water Level Estimation System to Mitigate Flood Risks in Riparian Communities in the Department of Choc´o.

1.3.2

Specific Objectives

1. Design predictive models from on-line and off-line measurements for linear

multivariate models through estimation algorithms such as Autoregressive Models (AR), Autoregression Models with exogenous variables (ARX) and Regularized Models.

2. Establish

predictive models from on-line and off-line measurements for multivariate nonlinear models through estimation algorithms such as nonlinear autoregressive models with exogenous inputs (NARX) where the nonlinearity is modeled using neural networks and through a hydrologic model.

3. Evaluate the performance of predictive models developed with real data acquired

from the database of the Institute of Hydrology, Meteorology and Environmental

p. 25

Studies (IDEAM), through its hydrological stations in order to mitigate flood risks.

1.4

Study Area The choice of the Atrato River in the Choc´o department as a study area to investigate flooding is based on several crucial factors.

Historically, the region has experienced significant and recurrent flooding, significantly affecting local communities, infrastructure, and the natural environment.

This persistent threat highlights the urgency of understanding the causes and patterns of flooding in this area to implement effective prevention and mitigation measures.

Choc´o is one of Colombia’s most vulnerable regions to flooding due to its mountainous topography and high rainfall.

This combination of natural factors increases the region’s susceptibility to extreme rainfall events and flash floods, exacerbating the local population and infrastructure risks. Therefore, the study of the Atrato River is relevant to understanding the specific hydrological phenomena of the area, but is also essential to strengthening the capacity to manage disaster risks and promote community resilience to future floods.

The Atrato River is Colombia’s third most navigable river, after the Magdalena and the Cauca rivers. It rises in the Cerro del Plateado in the municipality of El Carmen de Atrato, in the western mountain range of the Andes, and flows into the Gulf of Urab´a, in the Caribbean Sea, near the border with Panama. It runs through a large part of the Choc´o department, and in two sections of its course, it serves as a departmental border between Choc´o and Antioquia; due to its navigability, it is one of the means of transportation in the region. It is also part of the biogeographic Choc´o, considered the most biodiverse area on the planet and one of the rainiest; hence the high flow shown by this river. The river has a length of 750 km and a width varying between 150 and

500 m and a depth of 38 to 31 m. It flows into the Gulf of Urab´a through 18 mouths

that make up the river’s delta. Throughout its course, it receives about 150 rivers and

3000 streams. The World Wildlife Fund considers it one of the richest genetic banks

in the world.

The meteorological data selected primarily consist of three types of hydrological information:

flow rates, water levels, and precipitation from two stations. Each hydrological station provides measurements for these three variables. It’s important to note that there is a sampling frequency of one data point every 12 hours for each variable, resulting in 1578 data points over 789 days.

Figure 1.1 illustrates the course of the Atrato River through the Department of Choc´o, Colombia territory. This river, one of the most important in the country, flows from its source in the Western Cordillera to its mouth in the Gulf of Urab´a, in the Caribbean

p. 26

Sea. Along its course, the Atrato River crosses several regions, providing essential water resources and a crucial transportation route for local communities. The geography of the Choc´o, characterized by its dense tropical rainforest and high rainfall, significantly influences the flow and behavior of the river, making its study and monitoring a fundamental task for environmental management and sustainable development in the region.

Figure 1.1: Description of the entire course of the Atrato River in the Department of Choc´o - Colombia.

Figure 1.2 shows the location of the hydrological stations and the exact place where the data collection of the proposed case study is carried out. Hydrological stations are located at Colombia, Department of Choc´o .

p. 27

Quibdó Belén de Bajirá Hydrologycal Station Figure 1.2: The location of hydrological stations in the Atrato River Figure 1.3 shows the location of the hydrological stations and the exact place where the data collection of the study is performed in satellite form in the department of Choc´o, Colombia, Atrato river.

Figure 1.3: The location satelital of hydrological stations in the Atrato River In addition, in Table 1.1 the information period of the data sets used and the specified hydrological variables of the system can be observed.

p. 28

Table 1.1: Information of the used data sets. Type Organization Available Data Dataset Station 1(S1)

IDEAM

2021-2023

Flow, Level, Precipitation.

Station 2(S2)

IDEAM

2021-2023

Flow, Level, Precipitation.

Station 3(S3)

IDEAM

2021-2023

Flow, Level, Precipitation.

1.4.1

Hydrological station measurements The following figures describe the real measurements used for data forecasting, for the hydrological variables of water level, water flow and precipitation. Figure 1.4 shows the water level measurements at the three hydrologycal stations. Data sample Figure 1.4: Level measurements at the three stations Figure 1.5 shows the water flow measurements at the three corresponding stations. Data sample

3

3

3

Figure 1.5: Water flow measurements at the three stations Figure 1.6 shows the precipitation measurements at the three corresponding stations.

p. 29

Data sample Figure 1.6: Precipitation measurements at the three stations In Table 1.2 is shown in detail the location of the hydrological study stations of the Institute of Hydrology, Meteorology and Environmental Studies (IDEAM); which are located in Choc´o, department of Colombia; Station 1 (E1), Station 2 (E2) and Station 3 (E3). And also shows the altitude in meters above the sea level (MASL). Table 1.2: Location of the hydrological stations Station 1 Station 2 Station 3 Longitude 76° 40’ 10.75” W 76° 39’ 44.13” W 76° 39’ 43.8” W Latitude 5° 45’ 53.38” N 5° 41’ 52. 77” N 5° 41’ 24” N Altitude

20.579 MASL

20.83 MASL

20.83 MASL

City Bel´en de Bajir´a Quibd´o Quibd´o

1.5

Hydrologic Modeling Hydrologic modeling is an essential tool for analyzing, simulating, and predicting the behavior of the hydrological cycle, which is vital for water resource planning, disaster prevention, and water infrastructure design.

With the advance of science and technology, this discipline has evolved from empirical approaches to hybrid methodologies that combine physical models with artificial intelligence algorithms, which have significantly expanded its predictive capacity [43]; [65]. Hydrologic models are classified according to the degree of representation of physical processes:

• Empirical or black-box models, such as artificial neural networks or regression

models, are based exclusively on observational data, without explicitly considering the underlying physical processes [66].

p. 30

• Conceptual models (gray box) represent hydrological processes through simplified

structures, such as reservoirs and flows between compartments, allowing one to capture essential dynamics without needing detailed physical data [43].

• Physically-based models (white box) employ differential equations that describe

physical processes such as runoff, infiltration, and evapotranspiration in detail. They are suitable for complex studies but require considerable geophysical information [67].

• Hybrid models and Physics-Informed Machine Learning (PIML) integrate physical

knowledge with machine learning algorithms, such as LSTM or NARX neural networks, incorporating physical constraints in their structure to improve generalization and reduce uncertainty [65].

Several hydrological implementation and simulation platforms exist:

MATLAB,

Python, HEC-HMS, and SWAT (Soil and Water Assessment Tool). The construction of accurate hydrological models requires reliable and representative data. The primary data used are as follows:

The construction of reliable hydrological models mainly depends on the availability and quality of the data used. Among the fundamental variables are meteorological variables, such as precipitation, temperature, solar radiation, and evapotranspiration, which make it possible to represent the energy input and transfer processes in the basin. However, hydrometric variables, such as flow rates and water levels, provide direct information on the hydrological behavior of the system and are collected through field monitoring stations. In addition, spatial data, such as land use, slope, soil type, and digital elevation models, are essential to characterize the physical properties of the terrain and their influence on water dynamics.

To ensure the quality of the models, it is necessary to apply various data preprocessing techniques, including interpolation of missing values, smoothing of time series (e.g., using Savitzky-Golay or Rauch-Tung-Striebel filters), normalization of variables, and consistency analysis. These actions allow reducing noise, eliminating systematic errors and preparing the data for subsequent modeling in a robust and efficient way [68]. Validate model results, widely recognized statistical metrics are used. According to the guidelines of Moriasi et al [69], a model is considered acceptable if it meets: Nash-Sutcliffe Efficiency (NSE) > 0.50.

Kling-Gupta Index (KGE) close to 1.

Percentage Bias (PBIAS) between ±25%.

RSR < 0.70.

These indicators allow evaluating the fit between observations and simulations, and are essential for operational applications such as early warning systems.

p. 31

1.5.1

Hydrologic Cycle The hydrologic cycle describes the continuous movement of water within the Earth-atmosphere system, driven primarily by solar energy. Precipitation, flow, and water level are among the most relevant components for hydrologic modeling, as they characterize the processes of water inflow, transit, and storage in a watershed. Precipitation is the main input to the system and manifests itself in rain, snow, or hail. Its temporal and spatial measurement is critical for estimating runoff generation and aquifer recharge [70]. Once water reaches the ground, some infiltrates and some flow over the surface, generating surface runoff, which feeds water bodies and contributes to river flow.

The flow (expressed in m³/s) is the amount of water that passes through a section of the river in a given time interval and depends both on the intensity and distribution of precipitation and the physiographic characteristics of the basin (slope, vegetation cover, soil type, etc.) [71]. The water level reflects the height of water in a specific section of the channel or water body, indirectly indicating the volume stored or transported. The relationships between these three variables are fundamental to hydrological modeling. In general, the models transform precipitation time series into flow and level series, which allows simulating and forecasting the system’s behavior under different climatic and hydrological scenarios [72].

These simulations are essential in water resources planning, flood risk management, and early warning system design.

1.5.2

Processes of the hydrological cycle.

The hydrological cycle describes the continuous water path between the atmosphere, the Earth’s surface, and underground bodies. This cycle is driven by solar radiation and gravity, forming the basis of a watershed’s water balance. Its understanding is fundamental for modeling phenomena such as river flow, water level, and precipitation distribution [73, 74].

The main processes of the hydrologic cycle are:

• Evaporation: transfer of water from liquid surfaces to the environment in the

form of vapor, driven by solar energy.

• Transpiration: Water loss is caused by vapor from plants through the stomata

in their leaves. Combined with evaporation, it gives rise to evapotranspiration, a key variable in hydrological models.

• Condensation: Cooling of water vapor in the atmosphere that generates liquid

droplets and forms clouds.

p. 32

• Precipitation: Return of water to the earth’s surface from the atmosphere, either

as rain, snow, or hail.

It is the main source of recharge for watersheds and aquifers [75].

• Infiltration: Water penetration into the ground, from where it may feed aquifers

or slowly return to surface streams.

• Surface runoff: Part of the water that flows over the surface without infiltrating,

flowing into natural channels such as rivers and streams. It is one of the main components of flow.

• Percolation: Movement of infiltrated water to deeper subsurface areas.

• Storage: Water may temporarily accumulate in various reservoirs, such as rivers,

lakes, aquifers, soils, ice, or snow.

1.5.3

Hydrologic Model Based on the articles published during the development of this thesis, a comparative technical analysis of the hydrological models implemented in these investigations is carried out to evaluate their methodological approaches, performance, and applicability in the context of water level forecasting.

Table 1.3: Comparative Hydrological Models.

Article Type Model 1 Components Hydrological Cycle Complementary Technique Article 1 NARX + EnKF Flow, Precipitation, Level hybrid model Article 2 LSTM + Kalman Smoother Flow, Precipitation, Level hybrid model Accurate water level prediction requires integrating essential components of the hydrological cycle: precipitation represents the primary input to the system, while flow reflects the surface transport of water in rivers. Water level depends on the joint dynamics between infiltration, runoff, accumulation, and surface retention, so modeling it implies a holistic understanding of hydrological behavior [73]. Implementation of evaluation metrics for the validation of hybrid hydrological models: NSE, KGE, PBIAS, and RSR.

Table 1.4: Comparative performance of hydrological models Model

NSE

KGE

PBIAS (%)

RSR

NARX + EnKF (Atrato River)

0.9918

0.9903

+3.8

0.48

LSTM + Kalman Smoother (San Juan River)

0.9762

0.9730

4.6

0.54

p. 33

This Table 1.4 presents the comparative evaluation of the performance of two hybrid hydrological models applied to water level prediction in the Atrato and San Juan rivers, using statistical metrics widely accepted in hydrology.

Figure 1.7 shows the comparative result between the two models, the NARX + EnKF (Atrato River) and LSTM + Kalman Smoother (San Juan River), in terms of the comparative metrics, KGE, NSE, PBIAS, and RSR.

Figure 1.7: Results of the implemented hybrid hydrologic models.

1.6

Thesis Structure This document is organized as follows: in Chapter 2 are presented the linear models, in Chapter 3 are given the nonlinear models, in Chapter 4 are presented the hybrid models, in Chapter 5 is given the virtual environment for social appropriation of the results in riverside communities, and finally in Chapter 6 are given the conclusions and final remarks.

1.7

Contributions The main contributions developed in each of the chapters of this doctoral thesis are described in the following. Each chapter represents a specific advance in the study, modeling, and forecasting of hydrological variables through multivariable approaches and artificial intelligence techniques, emphasizing the behavior of water level in fluvial systems of the department of Choc´o.

p. 34

1.7.1

Chapter 2 Contributions The chapter provides relevant contributions to designing and strengthening Early Warning Systems (EWS) for floods by presenting linear models that allow for more accurate anticipation of hydrological behavior in vulnerable areas. Based on data analysis from three hydrological stations of the Atrato River, two different modeling approaches incorporating flow, water level, and precipitation information are applied and evaluated.

The first approach is based on a multivariate autoregressive (AR) model. It simultaneously assimilates the three hydrological variables and uses their natural correlations to generate more consistent and accurate estimates. This model is adjusted using regularization techniques, specifically with the Tikhonov method, and parameter selection is performed with generalized cross-validation, which avoids overfitting and improves predictive capacity. This type of multivariate data assimilation is beneficial for SAT, as it allows a coherent integration of various sources of information in real time.

The second approach implemented corresponds to an ARX (autoregressive with exogenous variables) model, which introduces a recursive structure that relates precipitation as input to system responses in terms of flow and water level. This technique uses least-squares identification algorithms and adapts dynamically to the available data. Its strength lies in representing the delayed and linear effects of the rainfall regime on the river system, a key feature for issuing timely early warnings. Both models were developed and validated with data provided by IDEAM, which allows their application in authentic contexts. The combination of these modeling strategies provides a solid quantitative basis for formulating early warnings, facilitating a more efficient response to flood hazards.

1.7.2

Chapter 3 Contributions This chapter strengthens flood warning systems (EWS) by applying advanced nonlinear models in water level prediction, integrating multiple hydrological variables. Previous research on the Atrato River shows how using neural networks with specific structures, such as NARX and BackPropagation, improves the ability to anticipate critical events in complex hydrological contexts.

The first fundamental contribution consists of implementing a multivariate NARX model based on neural networks, which allows a more accurate representation of the nonlinear relationships between precipitation, flow, and water level. This model has shown an outstanding ability to capture temporal dynamics and cross-dependencies between variables, providing short-term predictions with a smaller margin of error than traditional linear models such as ARX. This improvement in the accuracy and adaptability of the NARX model makes it a valuable tool for feeding flood warning

p. 35

systems with greater reliability.

The second contribution is related to the comparative analysis of different nonlinear methodologies, where the performance of a BackPropagation neural network is evaluated against NARX and ARX models.

This study shows that, depending on the structure and input data, neural network-based models can outperform linear approaches regarding accuracy and generalization capability. Including such models in the SAT allows for more effective capture of the complexities of the river system and more accurately anticipates sudden variations in water level. Overall, the nonlinear models developed and analyzed in this chapter represent an advance in incorporating artificial intelligence in flood risk management. Their practical application in real hydrological stations of the Atrato River reinforces their usefulness as a support for real-time decision-making, offering a more robust technological base to prevent adverse impacts on communities exposed to extreme events.

1.7.3

Chapter 4 Contributions This chapter provides key elements for strengthening Early Warning Systems (EWS) against floods through the application of nonlinear hybrid models oriented to water level prediction. Based on two recent investigations developed in the hydrological context of the Atrato River, it is demonstrated how advanced machine learning models, combined with data assimilation techniques, can significantly improve the capacity to anticipate extreme events.

The first contribution focuses on implementing LSTM (Long Short-Term Memory) neural networks, designed to handle multivariate data sequences, including precipitation, flow, and water level.

This model allows capturing complex temporal relationships between hydrological variables, improving short-term water level prediction.

In regions such as the Colombian Choc´o, where high rainfall and lack of monitoring infrastructure increase vulnerability, this model is an essential tool to make up for the shortcomings of the current EWS, providing more accurate and earlier warnings.

The second contribution derives from a hybrid approach that couples a NARX model (autoregressive neural network with exogenous inputs) with an Ensemble Kalman filter (EnKF), which allows uncertainty and robustness to be incorporated in the predictions. This model has been evaluated against different noise levels, demonstrating stability and accuracy even under adverse conditions. The NARX-EnKF combination provides a reliable two-day horizon water level estimate and responds effectively to external disturbances, which is essential for warning systems in dynamic and complex environments.

Both models, validated with performance metrics such as RMSE and Nash-Sutcliffe coefficient of efficiency (NSE), show a crucial technical advance in real-time multivariate prediction. Their application in real stations of the Atrato river demonstrates their

p. 36

operational potential to be integrated in local EWS, improving the response capacity to flood events.

This chapter proposes predictive modeling tools that strengthen decision-making in water risk management, making early warning systems more efficient and accurate in regions vulnerable to extreme events.

1.7.4

Chapter 5 Contributions This chapter provides a comprehensive approach to strengthening Early Warning Systems (EWS) for floods by incorporating social appropriation and knowledge transfer strategies adapted to the riparian communities of Choc´o. In a region where hydrometeorological phenomena are frequent and monitoring infrastructure is limited, including accessible and culturally relevant technologies represents a key advance for risk management.

One of the main contributions is developing a web platform oriented to local communities. This platform is conceived as an interactive and accessible space where users can consult real-time information related to flood warnings, weather forecasts, risk maps, and action protocols. This digital tool centralizes technical information and acts as a communication bridge between the authorities and the inhabitants, strengthening the community’s capacity to respond to emergencies.

In addition, an educational virtual environment is introduced, designed to simulate risk scenarios using real data and predictive models. Through interactive experiences, users can observe how different factors -such as rainfall intensity, river flow, or local geography- influence the level of risk. This tool promotes an active understanding of the natural environment and the processes that trigger floods, enabling people to learn to interpret warnings and act in advance.

The social appropriation process is another fundamental axis of the chapter. The platform and virtual environment were designed in collaboration with the communities of Choc´o, integrating their knowledge, languages, and ways of life. This ensured not only the acceptance of the tools but also their adaptation to local realities. The active participation of the community in developing and implementing these solutions reinforces their sustainability and usefulness over time.

These technological and social mechanisms transform EWS into more inclusive, decentralized, and efficient systems. The combination of predictive models, access to information, and digital education processes allows riverside communities to receive warnings, understand them, trust them, and respond promptly, thus strengthening their resilience to natural hazards.

p. 37

Chapter 2 Linear models for water level forecasting: implementation and analysis This Chapter presents two linear flood estimation methods for three hydrological stations in the Atrato River, measuring three hydrological variables: flow, level, and precipitation. The first method uses a Multivariable AR Data Assimilation for Water Level, Flow, and Precipitation Data model. The proposed autoregressive multivariate model considers correlations between water level, flow, and precipitation, and is estimated directly using real measurements.

To estimate the system parameters, a regularized estimation of the model is performed using the Tikhonov regularization method with generalized cross-validation to select the regularization parameters. On the other hand, the second study uses a Multivariable ARX-based Forecasting for an Early Warning Flood System. This method offers an alternative approach to existing linear models that represent the dynamics of water flow, precipitation, and river level by using a recursive ARX model consisting of an autoregressive exogenous structure that employs least squares identification using specific strategies that alternate between data assignment and parameter estimation. ARX models allow for the representation of the linearities and lags of the pluviometric process. The linear model estimation techniques used in Chapter 2 were applied using a sample of data provided by the Institute of Hydrology, Meteorology, and Environmental Studies (IDEAM) from three hydrological stations in the Atrato River. These data are represented in Figures 1.4 to 1.6, which show the values of the three hydrological variables: water flow, water level, and precipitation. In addition, Table 1.2 shows the geographic coordinates of the location of these stations.

p. 38

2.1

Multivariable AR Data Assimilation for Water Level, Flow, and Precipitation Data This section presents a method for data assimilation of a multivariable system that describes the behavior of water level, flow, and precipitation variables. The proposed multivariable AR model considers correlations between the water level, flow, and precipitation and is directly estimated using real measurements. In order to estimate system parameters, a regularized estimation of the model is performed using Tikhonov regularization method with generalized cross-validation for regularization parameter selection.

The proposed approach is evaluated by using data measured from a Colombian river located in the Choc´o department in Colombia and data from the meteorological information center of Argentina. The proposed multivariable autoregressive regularized estimated model is compared with three simultaneous univariable models.

An additional comparison is performed by considering a least squares solution for parameter estimation.

2.1.1

AR multivariable model Consider a AR multivariable model described as follows: y[k] + A1y[k −1] + A2y[k −2] + · · · + Apy[k −p] = e[k]

(2.1)

where Ai ∈Rm×m are the model matrix parameters, with i = 1, . . . , p being p the order of the system, e[k] ∈Rm×1 the noise with m the number of outputs, and y[k] ∈Rm×1 the measurement defined as:

y[k] =   y1[k] y2[k]

...

ym[k]  

(2.2)

Equation (2.1) can be rewritten as follows:

yT[k] + yT[k −1]AT

1 + · · · + yT[k −p]AT

p = eT[k]

(2.3)

and then yT[k] = −yT[k −1]

· · ·

yT[k −p]   AT

1...

AT p  + eT[k]

(2.4)

p. 39

By considering the values of k = 0, . . . , K, being K the total number of samples, the following matrix relation can be obtained:

  yT[1]

...

yT[k]

...

yT[K]   | {z } Y =   −yT[0]

· · ·

0

...

−yT[k −1]

· · ·

−yT[k −p]

...

−yT[K −1]

· · ·

−yT[K −p]   | {z } M   AT

1...

AT p   | {z } Θ +   eT[1]

...

eT[k]

...

eT[K]   | {z } ϵ

(2.5)

Adopting the approach proposed in [76], we obtain a discrete-time measurement equation, as shown below:

Y = MΘ + ϵ

(2.6)

where matrix Y ∈RK×m holds the measurements, M ∈RK×(m×p) is the Hankel matrix that holds the past measurements, and Θ ∈R(m×p)×m is the matrix that include the AR model parameters, and ϵ ∈RK×m represents the non-modeled features of the system, i.e. observation noise, and is assumed to be additive, white and Gaussian with zero mean and with covariance matrix defined by Cϵ.

2.1.2

Multivarible AR Regularized Solution The naive solution of an inverse problem associated to (2.6) can be achieved by the least squares solution. This can be performed by defining a functional given by

JLS = ∥Y −MΘ∥2

2,Cϵ

(2.7)

or

JLS = (Y −MΘ)TC−1

ϵ (Y −MΘ)

(2.8)

and

∂JLS

∂Θ = M TC−1 ϵ MΘ −M TC−1 ϵ Y

(2.9)

p. 40

by equaling (2.9) to zero, the following equation is obtained: bΘLS = (M TC−1 ϵ M)−1M TC−1 ϵ Y

(2.10)

being ΘLS the least-squares solution of the AR model of (2.6). If Cϵ = I the solution for ΘLS proposed in (2.10) can be simplified to bΘLS = (M TM)−1M TY

(2.11)

However, when the problem is rank-deficient, ill-posed or ill conditioned [76], the application of the Tikhonov Regularization Method can be performed, by defining a functional as follows:

JTikh = ∥Y −MΘ∥2 2,Cϵ + λ2 ∥Θ∥2

2

(2.12)

or JTikh = (Y −MΘ)TC−1 ϵ (Y −MΘ) + λ2ΘTΘ

(2.13)

and ∂JTikh ∂Θ = M TC−1 ϵ MΘ −M TC−1 ϵ Y + λ2Θ

(2.14)

by equaling (2.14) to zero, the following equation is obtained: bΘTikh = (M TC−1 ϵ M + λ2I)−1M TC−1 ϵ Y

(2.15)

being λ the regularization parameter, being bΘTikh the regularized AR solution of (2.6). Equation (2.15) can be simplified by defining Cϵ = I, resulting in bΘTikh = (M TM + λ2I)−1M TY

(2.16)

As indicated in [77], the regularization parameter λ is determined using the generalized cross-validation (GCV) method, which minimizes the following functional: Γ(λ) = ∥Y −MΘλ∥2

2

(trace(I −Mλ))2

(2.17)

being the influence matrix Mλ defined by Mλ = M(M TM + λ2I)−1M T

(2.18)

and Θλ defined as Θλ = (M TM + λ2I)−1M TY

(2.19)

It is worth noting that for the selection of the regularization parameter λ, the function (2.17) must be evaluated several times for several λ values.

p. 41

2.1.3

AR Hydrological model In order to consider the correlation of water level, water flow and precipitation, a multivariable AR model structure is selected as presented in (2.1), being y[k] defined as follows:

y[k] =   yL[k] yF[k] yP[k]  

(2.20)

where yL[k] is the water level at sample k, yF[k] is the water flow, and yP[k] is the precipitation. Therefore, the multivariable AR model can be defined as y[k] = − p X j=1 Ajy[k −j] + e[k]

(2.21)

being Aj ∈R3×3 the model parameters.

The model parameters are estimated by using (2.15), resulting in a regularized AR multivariable estimated model represented by ΘTikh, as follows: ΘTikh =   AT

1...

AT p  

(2.22)

The model (2.22) is updated for each new measurement, by performing the data assimilation task.

An univariable AR model can also be defined for each variable as follows: yL[k] = − p X j=1 aL j y[k −j] + eL[k]

(2.23)

yF[k] = − p X j=1 aF j y[k −j] + eF[k]

(2.24)

yP[k] = − p X j=1 aP j y[k −j] + eP[k]

(2.25)

being aL j ∈R the model parameters for the water level variable, aF j ∈R the model parameters for the water flow variable, and aP j ∈R the model parameters for the precipitation variable. The model parameters for each variable are estimated by using (2.15), resulting in a regularized AR univariable estimated model represented by ΘL, ΘF and ΘP, as follows:

ΘL Tikh =   aL

1...

aL p  , ΘF Tikh =   aF

1...

aF p  , ΘP Tikh =   aP

1...

aP p  

(2.26)

p. 42

The model parameters for each variable presented in (2.26) are also updated for each new measurement, by performing the data assimilation task.

It is worth mentioning that the parameters can also be estimated by using the least squares method as described in (2.10). In that case, the resulting parameters for the multivariable AR model described in (2.22) are defined as ΘLS, and the parameters (2.26) for level, flow and precipitation variables are defined as ΘL LS, ΘF LS an ΘP LS respectively.

2.1.4

Experimental setup In order to validate the multivariable AR regularized estimation method for data assimilation, a data set of hydrological variables of Level, Flow and Precipitation are analyzed.

The data set is measured at a hydrological station of the Institute of Hydrology, Meteorology and Environmental Studies (IDEAM). The IDEAM hydrological station number 11047010 is located at the following coordinates: Longitude -76.67, Latitude 5.76, Altitude 26.00. This station is located in Colombia country, in the Choc´o Department, Municipality of Quibdo, at the Atrato river. The sample time is 12 hours, and a total amount of 1578 samples are considered. In Fig. 2.1 is presented the selection of the regularization parameter by using GCV method for the Level variable.

Figure 2.1: Selection of the regularization parameter λ by using the GCV method for Level variable.

The selected value for λ regularization parameter is λ = 42.1861. By using this value, the vector of estimated parameters ΘL for the Level variable by using the regularized AR estimation method can be computed. It is worth noting that the regularization parameter implies a smoothing effect in the estimated signal, where an increase in the regularization parameter can be viewed as a smoother estimated signal. In (2.27) are shown the vectors for estimated parameters by using the regularized AR

p. 43

method ΘL Tikh and the least squares method ΘL LS for a system of order 10.

ΘL LS =  

0.8174

−0.1804

0.1300

−0.0007 −0.007

0.1326

−0.0918

0.1929

−0.1506

0.1525

  , ΘL Tikh =  

0.6396

−0.0095

0.0715

0.0298

0.0159

0.0878

−0.0094

0.1083

−0.0463

0.1053

 

(2.27)

By considering the estimated parameters of (2.27) for the regularized AR model ΘL Tikh, and the least squares AR model ΘL LS for a system of order 10, a comparison with the real data can be performed. In Fig. 2.2 is presented the comparison of the estimated signals by using the real data, ΘL LS and the ΘL Tikh is presented.

Figure 2.2: Comparison of the estimated signals by using the real level data, ΘL LS and the ΘL Tikh, order 10.

In Fig. 2.3 is presented the selection of the regularization parameter by using GCV method for the Flow variable.

p. 44

Figure 2.3: Selection of the regularization parameter λ by using the GCV method for Flow variable.

The selected value for λ regularization parameter is λ = 223.3113. By using this value, the vector of estimated parameters ΘF for Flow variable by using the regularized AR estimation method can be computed.

In (2.28) are shown the vectors for estimated parameters by using the regularized AR method ΘF Tikh and the least squares method ΘF LS for a system of order 10.

ΘF LS =  

0.6132

0.0008

0.0325

0.0221

0.0022

0.0811

0.0259

0.0676

−0.0169

0.1388

  , ΘF Tikh =  

0.4613

0.0827

0.0417

0.0314

0.0218

0.0660

0.0474

0.0605

0.0264

0.1169

 

(2.28)

By considering the estimated parameters of (2.28) for the regularized AR model ΘF Tikh, and the least squares AR model ΘF LS for a system of order 10, a comparison with the real Flow data can be performed. In Fig. 2.4 is presented the comparison of the estimated signals by using the real data, ΘF LS and the ΘF Tikh is presented.

p. 45

Figure 2.4: Comparison of the estimated signals by using the real Flow data, ΘF LS and the ΘF Tikh, order 10 In Fig. 2.5 is presented the selection of the regularization parameter by using GCV method for the Precipitation variable.

Figure 2.5: Selection of the regularization parameter λ by using the GCV method for Precipitation variable.

The selected value for λ regularization parameter is λ = 33.1963. By using this value, the vector of estimated parameters ΘF for Precipitation variable by using the regularized AR estimation method can be computed.

In (2.29) are shown the vectors for estimated parameters by using the regularized AR

p. 46

method ΘP Tikh and the least squares method ΘP LS for a system of order 10.

ΘP LS =  

0.1548

0.1467

0.0332

0.1137

0.1068

0.0442

0.0940

0.0983

0.0513

0.0292

  , ΘP Tikh =  

0.11852

0.1138

0.0537

0.0947

0.0899

0.0589

0.0833

0.0843

0.0611

0.0461

 

(2.29)

By considering the estimated parameters of (2.29) for the regularized AR model ΘP Tikh, and the least squares AR model ΘP LS for a system of order 10, a comparison with the real Precipitation data can be performed. In Fig. 2.6 is presented the comparison of the estimated signals by using the real data, ΘP LS and the ΘP Tikh is presented.

Figure 2.6: Comparison of the estimated signals by using the real Precipitation data, ΘP LS and the ΘP Tikh, order 10.

An additional comparison of the estimated results for the regularized AR method and the least squares AR method for a system of order 30 are presented in Fig. 2.7, Fig. 2.8, and Fig. 2.9.

p. 47

Figure 2.7: Comparison of the estimated signals by using the real Level data, ΘP LS and the ΘP Tikh, order 30.

Figure 2.8: Comparison of the estimated signals by using the real Flow data, ΘP LS and the ΘP Tikh, order 30.

p. 48

Figure 2.9: Comparison of the estimated signals by using the real Precipitation data, ΘP LS and the ΘP Tikh, order 30.

It is worth mentioning that the selection of the regularization parameter λ is directly related to how much we want to penalize or adjust the flexibility of our model.

2.1.5

Regularized AR multivariable estimation results The regularized AR multivariable solution by using Tikhonov is also compared with the real data and the least squares AR estimation. The selection of the regularization parameter λ is performed for each data-set by using the GCV method. A system of order 10 is selected in order to exemplify the behavior of proposed approach. Three λ values are obtained by using the GCV method related to each of the variables analyzed: Level, Flow and Precipitation. In Fig. 2.10 is presented the selection of the regularization parameter by using GCV method for the Level variable.

p. 49

Figure 2.10: Selection of the regularization parameter λ by using the GCV method for Level variable In Fig. 2.11 is presented the selection of the regularization parameter by using GCV method for the Flow variable.

Figure 2.11: Selection of the regularization parameter λ by using the GCV method for Flow variable In Fig. 2.11 is presented the selection of the regularization parameter by using GCV method for the Precipitation variable.

p. 50

Figure 2.12: Selection of the regularization parameter λ by using the GCV method for Precipitation variable The selected values of λ for each variable are λL = 32.3677, λF = 149.51, λP = 34.9909. By considering these values the mean of λL, λF and λP is selected as the regularization parameter for the regularized AR multivariable estimated solution, as λ = 72.28. In (2.30) and (2.31) are shown the matrices of estimated parameters by using the regularized AR method ΘTikh and the least squares method ΘLS for a system of order

10.

ΘLS =  

0.5898

0.0098

−0.0033

0.1733

0.8223

0.0076

0.5969

0.1230

0.1225

0.0281

−0.0012

0.0041

−0.3180 −0.1342

0.0183

0.3011

0.1065

0.1182

0.0160

0.01959

−0.0044

0.2094

0.1111

−0.0021

0.0956

0.02219

0.0028

0.0127

−0.0087

0.0021

−0.0091

0.02068

−0.0073

0.2143

0.0307

0.0832

0.0168

−0.0002

0.0007

0.3114

0.1271

0.0118

−0.3450

0.1539

0.0809

 

(2.30)

p. 51

ΘTikh =  

0.5883

0.0199

−0.0023

0.1445

0.7242

0.0103

0.2978

0.0747

0.0664

0.0245

−0.0069

0.0040

−0.2332 −0.0351

0.01676

0.1756

0.0635

0.0649

0.0202

0.0203

−0.0041

0.1527

0.0854

0.0010

0.0748

0.0293

0.0148

0.0093

−0.0074

0.0018

0.0403

0.0450

−0.0040

0.1136

0.0299

0.0483

0.0198

−0.0009

0.0006

0.2794

0.1266

0.0117

−0.1349

0.0839

0.0464

 

(2.31)

A comparison of the estimated results for the univariable and multivariable AR model estimated by Tikhonov regularization and least squares is also presented. The mean squared error is used for this comparison by considering models of orders 2 to 30. Fig. 2.13 shows the error comparison analysis for the Level variable, estimated for the univariable AR model with the least squares method (ELS) and the Tikhonov method (ELST), and the multivariable AR model with the least squares method (ELM), and the Tikhonov method (ELMT).

Figure 2.13: Level variable estimation error comparison for the univariable AR model with the least squares method (ELS) and the Tikhonov method (ELST), and for the multivariable AR model with the least squares method (ELM), and the Tikhonov method (ELMT).

A similar comparison is presented in Fig. 2.14 for the flow variable.

p. 52

Figure 2.14: Flow variable estimation error comparison for the univariable AR model with the least squares method (EFS) and the Tikhonov method (EFST), and for the multivariable AR model with the least squares method (EFM), and the Tikhonov method (EFMT).

A similar comparison is presented in Fig. 2.15 for the precipitation variable. Figure 2.15: Precipitation variable estimation error comparison for the univariable AR model with the least squares method (EFS) and the Tikhonov method (EFST), and for the multivariable AR model with the least squares method (EFM), and the Tikhonov method (EFMT).

2.2

Multivariable ARX based forecasting for an early warning flood system This section focuses on utilizing the recursive multivariable ARX approach applied to forecasting water levels along the Atrato river in Colombia. In-situ measurements of three meteorological stations are used. The data from the three measurement stations

p. 53

are analyzed in a decoupled estimation structure in the first part of the study. In this system, the algorithm is developed and validated with real measurements, where an uncoupled recursive ARX model is used for each station where there is a system of

2 inputs and 3 outputs. In turn, the response of the proposed approach is compared

to an AR model with three outputs where the estimated squared error for the three models is evaluated as the performance index.

In the study’s second part, three hydrological stations’ measurements are correlated in one multivariable ARX model with 6 inputs and 3 outputs. The multivariate data forecasting system uses the linear dynamic autoregression model with exogenous variables in a recursive way for three interconnected stations sequentially located in the Atrato river, department of Choc´o, western Colombia. These stations belong to the Institute of Hydrology, Meteorology, and Environmental Studies (IDEAM). The measurements are sampled every 12 hours of hydrological variables such as water flow, precipitation, and water level. The sampled data provided by the IDEAM is 800 days, where the water flow and precipitation are the inputs of each station, and the water levels of each station are the outputs.

2.2.1

ARX multivariable model Consider an ARX multivariable model described as follows:

y[k] + p X j=1 Ajy[k −j] = p X j=1 Bju[k −j] + e[k]

(2.32)

where Aj ∈Rm×m and Bj ∈Rm×n are the parameters of the model matrix, where y[k] ∈Rm×1 are the outputs and u[k] ∈Rn×1 are the inputs, being j = 1, . . . , p and p the order of the system, e[k] ∈Rm×1 the non-modeled features of the system and the data noise. It is worth noting that m the number of outputs and n the number of inputs of the system. The y[k] and u[k] are defined as:

y[k] =   y1[k] y2[k]

...

ym[k]  , u[k] =   u1[k] u2[k]

...

un[k]  

(2.33)

Equation (3.9) can be rewritten as follows:

yT[k] + p X j=1 yT[k −j]AT j = p X j=1 uT[k −j]BT j + eT[k]

(2.34)

p. 54

and then, yT[k] = ϕT[k −1]   AT

1...

AT p BT

1...

BT p   | {z } Θ[k−1] + eT[k]

(2.35)

being ϕ[k −1] =   −y[k −1]

...

−y[k −p] u[k −1]

...

u[k −p]  

(2.36)

2.2.2

Recursive Data Forecasting A recursive least squares multivariable method is proposed for data forecasting based on the multivariable ARX model described by (2.35), as follows: yT[k] = ϕT[k −1]Θ[k −1]

(2.37)

where the parameters of the hydrologycal model Θ[k −1] are computed iteratively as follows:

Θ[k] = Θ[k −1] + P[k −1]ϕ[k −1]

1 + ϕT[k −1]P[k −1]ϕ[k −1](yT[k] −yT

e [k])

(2.38)

being ye[k] defined as yT e [k] = ϕT[k −1]Θ[k −1]

(2.39)

and where P[k] is computed as P[k] = P[k −1] −P[k −1]ϕ[k −1]ϕT[k −1]P[k −1]

1 + ϕT[k −1]P[k −1]ϕ[k −1]

(2.40)

p. 55

2.2.3

Experimental setup ARX data forecasting Figure 2.16 shows the response of output 1, which corresponds to water level 1 of the 6-input, 3-output multivariable system of the recursive ARX model of fourth order, where the trained model performs data assimilation and predicts the next sample data, where the response of the real signal with respect to the estimated one is obtained. It is worth mentioning that the intial value of Θ is Θ[0] = 0, and the initial value of P is P[0] = 100I.

Figure 2.16: Response Recursive ARX coupled model for a fourth order system for water level of E1.

Figure 2.17 shows the response of output 2, which corresponds to water level 2 of the 6-input, 3-output multivariable system of the recursive ARX model of furth order, where the trained model performs data assimilation and predicts the next sample data, where the response of the real signal with respect to the estimated one is obtained. Figure 2.17: Response Recursive ARX coupled model for an order 4 system for water level of E2.

Figure 2.18 shows the response of output 3, which corresponds to water level 3 of the

p. 56

6-input, 3-output multivariable system of the recursive ARX model of fourth order, where the trained model performs data assimilation and predicts the next sample data, where the response of the real signal with respect to the estimated one is obtained. Figure 2.18: Response Recursive ARX coupled model for an order 4 system for level of E3.

Figure 2.19 shows an example of the parameter evolution. It can be seen that the parameters are updated at each sample. In general, the response of the other estimated parameters are similar to the one observed in Figure 2.19.

Figure 2.19: Example of parameter evolution for some elements of A1 matrix Figure 2.20 shows the response of the quadratic estimation error for a system of a first, second, third and fourth order of the coupled recursive ARX model, in which it can be observed how the system of order 4 has a better response of the quadratic error estimation, since it has a small percentage of error, compared to the other systems of first, second and third order.

It should also be noted that the system of fourth order, in terms of its computation time, takes a little longer in its execution, due to the magnitude of the operations, as shown in Figure 2.21.

p. 57

Figure 2.20: Quadratic estimation error response of the coupled recursive ARX model (6 inputs, 3 outputs), for system of order 1, order 2, order 3 and order 4. Figure 2.21 shows the response of the computed execution time for the recursive ARX coupled model of first, second, third and fourth order. It can be stated that for the fourth order recursive ARX model, the algorithm requires more execution time due to the amount of matrix operations to be performed in comparison with systems of less order. For example, for the system of first order the execution time is 9.4376µs, meanwhile for the system of fourth order the execution time is 14.0862µs, which represent an increment of 4.86µs.

Taking into account that the sample time is 12 hours, an increase of 4.86µs is not relevant. However, as shown in Figure 2.20, a lower error is obtained for the fourth order system Figure 2.21: Computational time of algorithm execution by each ARX Coupled system orders.

Figure 2.22 shows the response of the quadratic estimation error for a system of order

4 of the coupled recursive ARX model and the decoupled recursive ARX model. In this

response we can observe how the coupled recursive ARX obtains the best response in

p. 58

terms of the quadratic error estimation, where it is observed how this estimation error is the tendency to 0.

Figure 2.22: Quadratic estimation error response of the coupled recursive ARX model (6 inputs, 3 outputs) and the decoupled recursive ARX model (2 inputs, 1 outputs) for system of order 4.

Figure 2.23 shows the response of the quadratic estimation error for a system of order 4 of the coupled recursive ARX model and the AR model. In this response we can observe how the coupled recursive ARX obtains the best response in terms of the quadratic error estimation, In this response we can see how the coupled recursive ARX model obtains the best response in terms of the estimation of the quadratic error, in which the average of the squared errors of the two models is measured and we can see how the coupled ARX model obtains a better response than the decoupled ARX model in each of the outputs of the water level variable models.

Figure 2.23: Quadratic estimation error response of the coupled recursive ARX model (6 inputs, 3 outputs), compared to the AR model ( 3 outputs) for system of order 4.

p. 59

2.3

Academic discussion The results obtained in this section have been published in the following journals:

• Jackson

B.

Renteria-Mena, and Eduardo Giraldo, ”Multivariable AR Data Assimilation for Water Level, Flow, and Precipitation Data,”IAENG International Journal of Computer Science, vol.

50, no.

1, pp263-273, 2023.

ISSN: 1819-9224. [78]

• J. B. Renteria-Mena, and E. Giraldo, ”Real-time Adaptive Level Control of a

Multivariable Waste Water Treatment Plant,” Engineering Letters, vol. 30, no.2, pp444-452, 2022. ISSN: 1816-0948. [79]

2.4

Summary The chapter demonstrates that linear models, specifically AR and ARX, are practical tools for flood forecasting in vulnerable areas.

The AR model, by simultaneously integrating flow, water level, and precipitation, improves the accuracy of predictions due to its ability to capture correlations between variables. Its adjustment by regularization and cross-validation avoids over-fitting, making it reliable for real-time use. On the other hand, the ARX model allows understanding how precipitation influences the fluvial system with a certain time lag, a key aspect for issuing warnings sufficiently in advance. Both approaches, validated with real IDEAM data, confirm that proper modeling can significantly strengthen Early Warning Systems, facilitating a more timely and efficient response to flood risk.

2.5

Partial conclusions Using multivariate autoregressive models (AR) effectively improves the estimation of hydrological behavior by simultaneously incorporating flow, water level, and precipitation. This integration captures the natural relationships between variables, generating more coherent predictions adjusted to the local hydrological reality. The implementation of techniques such as Tikhonov regularization and generalized cross-validation helps to avoid over-fitting of the AR model, strengthening its capacity to generalize results in different hydrological scenarios and increasing its operational usefulness within an EWS By incorporating precipitation as an exogenous variable, the ARX model is beneficial for modeling the delayed effects of rainfall events on river flow and level. This anticipation capability is essential for issuing warnings sufficiently in advance in risk contexts.

p. 60

Both approaches have been validated with real data provided by IDEAM, which evidences their potential for implementation in real operational environments. This validation ensures the models can be reliable tools within the country’s monitoring and warning systems.

The combination of AR and ARX models provides a solid quantitative basis for formulating early warnings. This represents a relevant contribution to making timely and effective decisions regarding flood risk in vulnerable areas, such as the Atrato River basin.

p. 61

Chapter 3 Nonlinear models for water level forecasting: implementation and analysis This Chapter explores the implementation and analysis of nonlinear models applied to water level forecasting, based on two research studies published in international journals.

The first research, entitled Multivariable NARX-based Neural Networks Models for Short-term Water Level Forecasting, introduces an innovative application for multivariable prediction of hydrological variables using an NARX model. This approach is applied to two hydrological stations in the Atrato River, Colombia, where water level, flow, and precipitation variables are correlated using an NARX model structured in a neural network.

This design allows for the capture of complex dynamics and cross-correlations between hydrological variables, developing a short-term prediction of water level that can be used in flood early warning systems. The validation of this model is performed by comparing its estimation error with that of an ARX model, demonstrating that the NARX structure is more suitable for forecasting water level. The second investigation, entitled Comparative Analysis of Nonlinear Methods for Multivariable Water Level Prediction: The Case Study of the Atrato River, presents an innovative application of nonlinear models for multivariable prediction of hydrological variables, using a BackPropagation (BP) neural network model. The effectiveness of this model is evaluated in comparison with NARX and ARX models. This approach is implemented at two hydrological stations in the department of Choc´o, Colombia, along the Atrato River.

The BP neural network establishes correlations between water level, flow, and precipitation, handling these variables’ dynamic complexity and interconnections.

In addition, a short-term water level prediction is developed to integrate it into an early warning system for floods.

In Chapter 3, nonlinear model estimation techniques were applied using data provided by the Institute of Hydrology, Meteorology, and Environmental Studies (IDEAM) from two hydrological stations in the Atrato River, which are described in Section 1.4.

p. 62

3.1

Multivariable NARX-based neural network models for short-term water level forecasting.

The application of a multivariable NARX model based on neural network for short-term water level forecasting along the Atrato river in Colombia is presented and evaluated. To this end, the present work uses data from two hydrological stations located in the Atrato river, which are monitored by the Institute of Hydrology, Meteorology and Environmental Studies (IDEAM). The data includes measurements of flow, precipitation, and water level sampled every 12 hours over a period of 789 days. The multivariable NARX model is trained to predict the water levels of each station based on the inputs of water level, water flow and water precipitation by considering the inherent dynamic and correlation of the process. The performance of the models is evaluated based on the mean square error of the estimated outputs compared to the actual data. The performance of the proposed approach is compared to a multivariable ARX system.The main contribution of this paper is a general design of a multivariable NARX model structure based on neural networks for short-term water level forecasting.

3.1.1

Hydrological Variables In order to perform a water level forecasting based on the dynamic of a river, two hydrological stations are located on a river in two different positions. In order to consider the correlation among all the variables of the system and their corresponding nonlinearities, a nonlinear dynamical model is proposed. In (3.5) are shown the inputs and outputs of the proposed model y[k] = yL1[k] yL2[k]  , u[k] =   uF1[k] uPT1[k] uF2[k] uPT2[k]  

(3.1)

where yL1[k], yL2[k] correspond to the 2 outputs of the level variable of the two stations, uF1[k], uF2[k], uPT1[k], uPT2[k] are the 4 inputs of the neural network system corresponding to the two stations of the multivariable system, ie, the uFj[k] represents the j −th 2 inputs of the flow variable of the two stations and uPTj[k] represents j −th two more inputs of the rainfall variable of the two stations already mentioned, thus obtaining a multivariable system with 4 inputs and two outputs. The dynamic of the hydrological variables is defined by considering a Nonlinear function with an Auto-Regressive and exogenous inputs (NARX) as follows: y[k] = f(y[k −1], . . . , y[k −n], u[k −1], . . . , u[k −n]) + η[k]

(3.2)

being n the order of the NARX model and f(.) the nonlinear function, and η[k] the additive noise at time instant k.

p. 63

3.1.2

NARX Based Neural Network Structure In order to consider the NARX model of (4.8), the inputs are selected as u[k −j] and y[k −j], with j = 1, . . . , n, which correspond to a n −th order model In this work a 4−th order model (n = 4) is considered according to [80] where an analysis of the order selection is performed and the lowest estimations error is obtained for 3 −rd order model or higher. Therefore, by considering the variables described in (3.5), the proposed NARX model consists of 24 inputs and 2 outputs.

In order to approximate the nonlinear function of (4.8), the nonlinear function f(.) is approximated by using a neural network structure f ∗(.), as depicted in Fig. 4.1. Figure 3.1: NARX based Neural Networks Structure where the NARX model can be defined as follows:

y[k] = f ∗(y[k −1], . . . , y[k −n], u[k −1], . . . , u[k −n]) + η[k]

(3.3)

To this end, 24 input activation functions with one hidden layer and 2 outputs are considered. A feed-forward network is selected as a candidate for the NARX model in order to speed-up the training process. The training of the NARX model is performed offline by considering the data sample.

A linear ARX structure can be obtained by neglecting the hidden layer as depicted in Fig. 4.12.

Figure 3.2: ARX Structure

p. 64

where the ARX model can be defined as follows:

y[k] = −

4

X j=1 ajy[k −j] +

4

X j=1 bju[k −j]

(3.4)

where Aj ∈Rm×m and Bj ∈Rm×m are the parameters of the model matrix, where y are the outputs and u are the inputs; with j = 1, being p the order of the system, e[k] the noise with m, the number of outputs and inputs of the system, y[k] ∈Rm×1 and u[k] ∈Rm×1.

3.1.3

Experimental Setup In order to validate the proposed approach, a comparison analysis of the proposed NARX approach based on neural networks (3.7) is performed with a multivariable ARX model (3.8). A visual comparison of the real and estimated signals is presented for the ARX and NARX methods and also a quantitative evaluation based on the mean squared error is performed. Both systems (ARX and NARX) are trained offline by considering the measurement data. An evaluation in terms of the training error is also presented. The feed-forward network structure for the NARX approach considers one hidden layer with 256 units. The Rectified Linear Unit (ReLU) Activation Function is selected for the proposed approach, where the ReLU is a piecewise linear function that will output the input directly if it is positive, or zero otherwise. The implementation of the proposed NARX model based on neural networks and also the ARX model is performed in Python by using Tensorflow, which is an open-source machine learning library developed by Google.

3.1.4

Estimation Results The following are presented the estimation results for the ARX method for each of the two water level outputs. In Fig. 3.3 is shown the short-term estimation of the first water level output and also the real measurements. In Fig. 3.4 is depicted the short-term water level forecasting of the second output as well as the real corresponding measurements.

p. 65

Figure 3.3: First water level output and short-term estimation based on a multivariable ARX model.

Figure 3.4: First water level output and short-term estimation based on a multivariable ARX model.

In Table 3.1 is shown an analysis for the nonlinear neural network NARX in terms of the number of nodes in the hidden layer and their corresponding mean squared estimation error.

Table 3.1: Mean squared Estimation error for several nodes configurations NARX hidden layer nodes Level 1 Level 2 Total

2

1.6867

1.6859

3.3726

4

0.0266

0.0275

0.0541

8

0.0254

0.0272

0.0526

16

0.0227

0.0257

0.0484

32

0.0232

0.0201

0.0433

64

0.0235

0.0180

0.0415

128

0.0176

0.0133

0, 0309

256

0.0085

0.0108

0.0193

512

0, 0078

0.0101

0.0179

p. 66

From Table 3.1 it can be seen that the total estimation error is reduced by increasing the number of nodes in the hidden layer. It is noticeable that a there is no significant reduction in the total estimation error between 256 and 512 nodes. Therefore, in this work, a number of 256 nodes in the hidden layer are used for evaluation of the NARX models.

Presents the estimation results for the proposed multivariable NARX method for each of the two water level outputs by using 256 nodes at the hidden layer according to the results presented in Table 3.1. In Fig. 3.5 is presented the short-term estimation of the first water level output and also the real measurements. In Fig. 3.6 is presented the short-term water level forecasting of the second output as well as the real corresponding measurements.

Figure 3.5: First water level output and short-term estimation based on a multivariable NARX model.

Figure 3.6: First water level output and short-term estimation based on a multivariable NARX model.

Considering the forecast results presented in the Figures for the ARX and NARX models, the proposed multivariable NARX and ARX approaches show an adequate performance by visual inspection. In order to determine which approach tracks the dynamics of the hydrology variables more adequately, a quantitative evaluation is performed.

To this end, the mean squared error is computed in order to compare

p. 67

the real measurements and their corresponding forecasting for each of the considered methods. As a result, in Table 3.2 are presented the mean squared error for the proposed multivariable NARX approach and also the multivariable ARX method. It can be seen that the estimation error for the proposed NARX model is lower than the ARX model. It is worth noting that the reduction of estimation error for the NARX approach in comparison to the ARX approach is over the 50% for all variables. Table 3.2: Mean squared Estimation error Neural network model Level 1 Level 2 Total

ARX

0.0280

0.0263

0.0543

NARX

0.0085

0.0108

0.0193

The estimation errors during training for each of the methods considered are shown below. It can be seen that the multivariable ARX approach converge faster than the proposed multivariable NARX approach. However, the training error is lower for the NARX approach in comparison to the ARX approach. This behaviour validate the fact that there are nonlinear dynamics inherent to the measured hydrological variables and therefore the proposed NARX model forecast more adequately the data behaviour. Figure 3.7: Training error with the ARX multivariable model. Figure 3.8: Training error with the NARX multivariable model.

p. 68

3.2

Comparative Analysis of Nonlinear Methods for Multivariable Water Level Prediction: The Case Study of the Atrato River The aim of this investigation is to assess the efficacy of three nonlinear neural network models for a multivariate system with 4 inputs and 2 outputs to predict short-term water levels in the Atrato River, situated in the Choc´o department of Colombia. The nonlinear models under consideration include NARX, Back-Propagation neural networks and model linear ARX. The study relies on data samples from 2 hydrological stations along the Atrato River, has been meticulously monitored by the Institute of Hydrology, Meteorology, and Environmental Studies. The collected dataset includes measurements of river flow, precipitation, and water levels, consistently documented at 12-hour intervals over a duration of 789 days.

Training nonlinear models involves forecasting water levels by integrating input data on water level, water flow, and water precipitation. This process considers the intrinsic traits and interconnections among these variables, facilitating a thorough examination of the dynamics and correlations within the hydrological process. The multivariate neural network structure utilizes the flow and precipitation data from two stations as inputs and predicts the water levels of these stations as outputs. The novelty of this research is that the comparative approach between a backpropagation neural network, a NARX model and a linear ARX model in the context of online training as is the case, lies in the combination of different modeling and forecasting approaches with a dynamic and adaptive training methodology. While most previous studies have focused on static comparisons using batch training techniques, this research offers a more dynamic perspective by continuously adapting the model as new data arrives. This methodology allows for better capture of emerging patterns in the hydrologic data and improved accuracy and robustness of long-term forecasts. In addition, by incorporating nonlinear least squares techniques, this study more effectively addresses the complexity and nonlinearity of relationships between hydrologic variables. This not only improves the ability of models to forecast extreme hydrologic events in real time, but also contributes to a deeper and more accurate understanding of hydrologic systems. Taken together, this research offers a valuable tool for water management and disaster risk reduction by providing an innovative and adaptive approach to address hydrological challenges in continuous data stream environments.

3.2.1

Theoretical framework To predict water levels in a river considering its dynamic behavior, two hydrological stations are strategically positioned at different locations along the river. In order to account for the correlation among all variables within the system and address their

p. 69

nonlinearities, three proposed nonlinear dynamic models are employed. In (3.5) the inputs and outputs of the model are illustrated as follows: y[k] = yL1[k] yL2[k]  , u[k] =   uF1[k] uPT1[k] uF2[k] uPT2[k]  

(3.5)

where yL1[k] and yL2[k] are the water level outputs of the three stations of the multivariate system at samples k, and uF1[k], uF2[k] and uPT1[k], uPT2[k] are the inputs of the multivariable system, where the uFj[k] represent the j −th inputs of the water flows of the multivariable system and the outputs of the multivariable system. And uPTj[k] represents the j −th inputs and precipitation of the multivariable system. The dynamics of hydrological variables are characterized by incorporating a Nonlinear Auto-Regressive with exogenous inputs (NARX) function, expressed as: y[k] = f(y[k −1], . . . , y[k −n], u[k −1], . . . , u[k −n])

(3.6)

being n the order of the NARX model and f(.) the nonlinear function. NARX y ARX Nonlinear based Neural Network Structure.

In order to incorporate the NARX model from Equation (4.8), the inputs are chosen as u[k −j] and y[k −j], with j = 1, . . . , that correspond to a first-order model (n = 1). Thus, utilizing the variables outlined in Equation (3.5), the suggested NARX model involves 4 inputs and 2 outputs, with 10 nodes on its hidden layer. To approximate the nonlinear function described in (4.8), approximation of the nonlinear function f(.) is achieved through the utilization of a neural network structure.f ∗(.), as depicted in Fig. 3.9.

Figure 3.9: Structure of Neural Networks Based on NARX Nonlinear Model

p. 70

which can be established as follows:

y[k] = f ∗(y[k −1], . . . , y[k −n], u[k −1], . . . , u[k −n])

(3.7)

For this purpose, 4 input activation functions are utilized with one hidden layer and 2 outputs. A nonlinear least squares system identification structure is employed in this setup. A feed-forward network is chosen as a suitable candidate for the NARX model to enhance the efficiency of the training process. The training of both NARX and ARX models is conducted online, considering the available data sample. A linear ARX structure is derived by omitting the hidden layer, as illustrated in Fig. 3.10 Figure 3.10: Structure of Neural Networks Based on ARX linear Model which can be established as follows:

y[k] = −

4

X j=1 ajy[k −j] +

4

X j=1 bju[k −j]

(3.8)

3.2.2

Back-Propagation Nonlinear based Neural Network Structure.

Rumelhart and McClelland proposed the back-propagation neural network (BPNN) in 1986[81] It is a multi-layer feedforward neural network (Figure 2 is an example) trained by the error back-propagation algorithm and is one of the most widely used artificial neural network models[82, 83] rule is to use the steepest descent method and repeatedly modify weights and biases of the network through reverse iteration so that the sum of squared errors is minimized.Where Xi : i = 1, 2, . . . ni input nodes; Xj : j = 1, 2, . . . nj hidden layer neurons; Xfj, Hidden layer neurons evaluated in the activation function; Xk : i = 1, 2, . . .nk output nodes; Wij, weights between the input nodes and the neurons of the hidden layer; Wjk, weights between the neurons of the hidden layer Xfj and the output nodes.

Therefore, we deduce the following equation for the nonlinear Back Propagation model. Xj = ni X i WijXi

(3.9)

p. 71

Xfj = f(Xj) = f( ni X i WijXi)

(3.10)

Xk = nj X i WjkXfj

(3.11)

A nonlinear backpropagation structure with 10 nodes in the hidden layer, 4 inputs and

2 outputs can be obtained as follows Fig. 4.12

Figure 3.11: Structure of Neural Networks Based on Back-Propagation nonlinear Model For a better understanding, a flowchart using a Backpropagation, NARX and ARX neural network structure is presented as follows in Fig. 3.12. The diagram represents the sequence of steps to implement and compare Backpropagation, NARX and ARX neural network structures in a given context. Each step includes data preparation, neural network configuration, model training, and model performance evaluation. Comparison of results allows determining which of the neural network structures is most suitable for the problem at hand.

p. 72

Figure 3.12: Methodological architecture Backpropagation, NARX and ARX The water flow and precipitation in the Atrato river from the hydrological data of the two stations are used as input data, and the water level of the two hydrological stations is used as prediction for model training, thus having an order 1 system with a training model structure of 4 inputs and 2 outputs.

3.2.3

Regression metrics for the estimation of the quadratic error.

In this work, the evaluation of the prediction performance is the key to evaluate the quality of the model and optimize the efficiency. This study uses a series of indicators to evaluate the performance of the prediction of the outputs of the three models of nonlinear neural networks proposed in the case study, such as: NARX, ARX and back -Propagation, where the metrics of regression to calculate error,as root mean square error (RMSE), mean squared relative error (MSRE), mean absolute error (MAE), mean absolute relative error (MARE), Nash–Sutcliffe efficiency coefficient (NSE) and the Kling-Gupta Efficiency Coefficient (KGE)[84].

The following are the mathematical formulas with which the error regression metrics will be calculated.

RMSE =

r

1

nΣn i=1(y0i −yi)2

(3.12)

p. 73

MSRE =

s

1

nΣn i=1 (y0i −yi)2 y2 0i

(3.13)

MAE = 1

nΣn i=1[y0i −yi]

(3.14)

MARE = 1

nΣn i=1 [y0i −yi] y0i

(3.15)

NSE = 1 −Σn i=1(y0i −yi)2 Σn i=1(y0i −y0)

2

(3.16)

KGE = 1 −

p (r −1)2 + (α −1)2 + (β −1)2

(3.17)

where:

r = correlation between observations and predictions α = σp σo (ratio of standard deviations) β = µp µo (ratio of averages)

3.2.4

Experimental setup To assess the validity of the proposed methodology, a thorough comparative analysis is undertaken. The evaluation involves testing the NARX neural network approach against a nonlinear ARX model that utilizes the least squares system identification framework. Additionally, a comparison is made with a nonlinear neural network using Back Propagation. Both visual and quantitative assessments are conducted, comparing the actual and estimated signals for each model. This comprehensive evaluation employs five machine learning regression metrics.

All three systems, namely Back Propagation, ARX, and NARX, undergo online training using the measurement data. The computational training time for each model is also considered in the assessment. The feed-forward network structure for both the NARX and Back Propagation neural network approaches incorporates a hidden layer with 10 nodes. The NARX approach uses the Rectified Linear Unit (ReLU) activation function, a piecewise linear function that outputs the input directly if positive or zero otherwise.

p. 74

In contrast, the Back Propagation approach employs the sigmoid activation function, transforming input values to a (0,1) scale. The ARX model focuses on a nonlinear structure through least squares identification.

The implementation of the proposed NARX and ARX models based on neural networks is carried out using the Python algorithm with Tensorflow-Keras, an open-source machine learning library developed by Google. For online training, the Train on Batch methodology is employed. The Back Propagation network is implemented in Matlab Online, initialized with half of the data and trained with the remaining half of the sample.

3.2.5

Estimation results The following are the results of the estimation of linear models of nonlinear neural networks ARX, NARX and Back-Propagation proposed in the case study, therefore we have the results of the short-term prediction of the two outputs of the system that correspond to the hydrological variable of the water level. Figure 3.13 presents the short-term estimation water level output 1 and also the actual measurements of the ARX nonlinear neural network with a least-squares system identification structure.

Estimated ARX Real Output Level 1 [meters]

0

2

4

6

8

Time [days]

0

200

400

600

800

Time [days]

0

50

100

Output Level 1 [meters]

0

2

4

6

8

Figure 3.13: The output for water level 1 and the short-term estimation are determined using a linear multivariate least squares ARX model.

It can be seen in Figure 3.13 that there is a small difference in the first 50 days for the level output 1, and after that the signals are quite similar. This difference is due to the estimation process of the ARX nonlinear network which is performed by the least-squares identification structure.

Figure 3.14 presents the short-term estimate of water level output 2 and also the actual measurements of the ARX nonlinear neural network with a least-squares system identification structure.

p. 75

Estimated ARX Real Output Level 2 [meters] Output Level 2 [meters] Time [days]

0

0

200

400

600

800

2

2

4

6

0

4

6

Time [days]

0

50

100

Figure 3.14: The output for water level 2 and the short-term estimation are determined using a linear multivariate least squares ARX model.

As shown in Figure 3.13 for the level output 1, the level output 2 shown in Figure 3.14 shows a similar behaviour. I is worth noting small differences in the first 50 days, and after that the real and the estimated signals are almost the same. Figure 3.15 presents the short-term estimate of water level output 1 and also the actual measurements of the NARX nonlinear neural network.

Output Level 1 [meters] Time [days]

0

200

400

600

800

Time [days]

0

50

100

Output Level 1 [meters]

0

2

4

6

8

0

2

4

6

8

Estimated NARX Real Figure 3.15: Water level output 1 and short-term estimation based on a nonlinear multivariate NARX model.

In Figure 3.15 is shown that the estimation results for the level output 1 and the real output are quite similar, there is almost no difference between the two signals. However, it can be observed a small difference in the first 50 days. Apparently, from a qualitative point of view, the NARX method shown in this Figure works better than the ARX method shown in Figure 3.13.

Figure 3.16 presents the short-term estimation water level output 2 and also the actual measurements of the NARX nonlinear neural network.

p. 76

Estimated NARX Real Time [days]

0

200

400

600

800

Time [days]

0

50

100

Output Level 2 [meters]

0

2

4

6

Output Level 2 [meters]

0

2

4

6

Figure 3.16: Water level output 2 and short-term estimation based on a nonlinear multivariate NARX model.

The same behaviour described for Figure 3.15 is shown in Figure 3.16 for the level output 2. There is almost no difference between the real and the estimated signals after the first 50 days by using the NARX method.

Figure 3.17 presents the short-term estimation water level output 1 and also the actual measurements of the Back-Propagation nonlinear neural network. Real Estimated BP Output Level 1 [meters]

0

2

4

6

8

Time [days]

0

200

400

600

800

Time [days]

0

50

100

Output Level 1 [meters]

0

2

4

6

8

Figure 3.17: Water level output 1 and short-term estimation based on a nonlinear multivariate Back-Propagation model.

The comparison between the real and estimated signals by using the back-propagation method shows for level output 1 shown in Figure 3.17, presentes small differences in the first 100 days, which is quite similar to the ARX method shown in Figures 3.14 and

3.13.

Figure 3.18 presents the short-term estimation water level output 2 and also the actual measurements of the Back-Propagation nonlinear neural network.

p. 77

Output Level 2 [meters] Output Level 2 [meters] Time [days] Estimated BP Real

0

0

200

400

600

800

2

2

4

6

0

4

6

Time [days]

0

50

100

Figure 3.18: Water level output 2 and short-term estimation based on a nonlinear multivariate Back-Propagation model.

The same behaviour described in Figure 3.17 is observed in Figure 3.18. In order to obtain a quantitative analysis comparison of the methods, in Table 3.3 and Table 3.4 regression metric data for the most relevant squared error estimates of output 1 and output 2 of the nonlinear, ARX, NARX Back-Propagation models are presented. Table 3.3: Output 1 nonlinear models

RMSE

MSRE

MAE

MARE

NSE

KGE

Back Propagation(BP)

0.1926

0.0436

0.0868

0.0150

0.9986

0.9987

NARX

0.1660

0.0381

0.0872

0.0200

0.9862

0.9983

ARX

0.2480

0.0571

0.0869

0.0201

0.9684

0.9970

Table 3.4: Output 2.

nonlinear models

RMSE

MSRE

MAE

MARE

NSE

KGE

Back Propagation(BP)

0.0742

0.0157

0.0457

0.0079

0.9966

0.9992

NARX

0.0985

0.0216

0.0615

0.0140

0.9958

0.9988

ARX

0.1265

0.0291

0.0610

0.0142

0.9914

0.9984

In Table 3.5 are shown the Model parameters of the nonlinear neural networks NARX, ARX, Back Propagation used for training and testing the models. Table 3.5: Model parameters of the nonlinear neural networks and linear model nonlinear Learning # input # hidden # output models rate nodes nodes nodes Back Propagation(BP)

0.02

4

10

2

NARX

0.02

4

10

2

ARX

0.02

4

-

p. 78

3.3

Academic discussion The results obtained in this section have been published in the following journals:

• J.B. Renteria-Mena, E. Giraldo and Douglas Plaza, ”Multivariable NARX based

Neural Networks Models for Short-term Water Level Forecasting”, International conference on Time Series and Forecasting. July 12th-14th, 2023, Gran Canaria, Spain. [85]

• J.B. Renteria-Mena, D. Plaza, and E. Giraldo, ”Comparative Analysis of

Nonlinear Methods for Multivariable Water Level Prediction: The Case Study of the Atrato River”. Journal of Electrical and Computer Engineering, 2024, vol. 2024, no 1, p. 2894031. https://doi.org/10.1155/2024/2894031. ISSN: 2090-0147, 2090-0155. [86]

3.4

Summary The chapter provides evidence that nonlinear models based on neural networks, such as NARX and BackPropagation, significantly improve water level prediction in complex hydrological contexts such as the Atrato River. First, the multivariate NARX model proved superior to linear approaches (such as ARX) by capturing more accurately the nonlinear and temporal relationships between precipitation, flow, and water level, reducing the margin of error in short-term predictions. Secondly, a comparison was made between different nonlinear models, showing that, depending on the configuration and data, neural networks can offer greater accuracy and generalization capacity than traditional models.

Finally, the practical application of these models in real stations validates their potential as valuable tools to improve Early Warning Systems, facilitating a faster and more accurate response to flood events.

3.5

Partial conclusions Models based on neural networks, especially NARX, significantly improve over traditional linear methods (such as ARX) by more accurately representing the complex relationships between precipitation, flow, and water level. This adaptive capability makes them highly effective tools for real-time hydrological forecasting. The multivariate structure of the NARX model allows capturing temporal dynamics and cross-dependencies between hydrological variables, achieving short-term water level predictions with reduced margins of error. This precision is key to establishing reliable early warnings in vulnerable contexts such as the Atrato River.

p. 79

The comparative analysis between models shows that Back Propagation networks can also offer high performance, especially when properly adjusted to the data structure. Their generalization capability makes them suitable for incorporation as complementary components in prediction systems.

The incorporation of nonlinear models based on machine learning represents a relevant advance in the modernization of the EWS. It allows a deeper interpretation of hydrological processes and a greater capacity to anticipate critical variations. The validation of these models with real data from the Atrato River confirms their potential integration into operational environments, improving decision-making and community preparedness for extreme events such as flash floods.

p. 80

Chapter 4 Water Level forecasting by Hybrid Models This Chapter focuses on implementing and analyzing nonlinear hybrid models applied to water level forecasting, based on two research studies published in international journals. The first of these studies, Multivariate Hydrological Modelling Based on Long Short-Term Memory Networks for Water Level Forecasting, addresses the recurrent problem of floods in regions characterized by heavy rainfall and a dense river network. The lack of adequate infrastructure for flood prevention and forecasting exacerbates the vulnerability of these areas. The lack of early warning systems, the limited number of meteorological and hydrological stations, and deficiencies in urban planning aggravate the situation of affected communities. Given this scenario, investing in infrastructure that allows for better flood prediction and prevention, including advanced monitoring systems, developing predictive hydrological models, and constructing hydraulic works, is essential. This study introduces an innovative approach for multivariate hydrological variables, focusing on water level prediction at two hydrological stations located along the Atrato River in Colombia.

The second research presents an innovative application for short-term multivariate prediction of hydrological variables, using a multivariate NARX model coupled with a nonlinear and recursive Ensemble Kalman Filter (EnKF). This approach is implemented in two hydrological stations of the Atrato River in Colombia, where correlations between water level, flow, and precipitation are analyzed. The NARX model, based on neural networks, is designed to capture the complex dynamics of these hydrological variables and their cross-correlations. Short-term water level prediction, with a two-day horizon, is performed using a fourth-order NARX model. It is observed that coupling the NARX model with the EnKF improves the robustness of the proposed approach to external perturbations. Furthermore, the validity of the approach is tested by subjecting the NARX-EnKF model to five levels of additive white noise.

Regression metrics are employed to evaluate the performance of the proposed model using the root mean square error (RMSE) and the Nash-Sutcliffe model efficiency coefficient (NSE).

p. 81

In Chapter 4, estimation techniques for hybrid models were employed, using data provided by the Institute of Hydrology, Meteorology, and Environmental Studies (IDEAM), from two hydrological stations in the Atrato River. The location and characteristics of these stations are described in section 1.4.

4.1

Multivariate Hydrological Modeling Based on Long Short Term Memory Networks for Water Level Forecasting This work, explores and evaluates the implementation of a multivariate LSTM model based on recurrent neural networks (RNN) for predicting water levels in the Atrato River, located in the Choc´o Department, Colombia, over both short and long-term periods. The research utilizes data from two hydrological stations on the Atrato River, monitored by the Institute of Hydrology, Meteorology, and Environmental Studies (IDEAM). These data include measurements of flow rate, precipitation, and water level sampled every 12 hours over a span of 789 days with a data sample of 1578 data. The proposed study model makes predictions every 12 hours, since the flow, precipitation and level data from the two hydrological stations are recorded every 12 hours due to the limited memory of the embedded systems of the early warning system under study. The primary contribution of this research involves formulation of a multivariate LSTM model based on recurrent neural networks for short and long-term water level prediction.

It is noteworthy that LSTM networks have the ability to ”remember” relevant information from the sequence and retain it over multiple time steps, resembling the way our brain analyzes sequences. For instance, when reading a buyer’s review to make decisions, the LSTM network mimics our approach by focusing on words deemed relevant and discarding non-essential information, therefore, highlights that these networks have the ability to remember and process sequences of hydrological data over time, making them particularly useful for predicting hydrological variables such as water levels, flows or precipitation. As in other contexts, LSTMs can capture complex temporal dependencies in hydrological data, making them a valuable tool for improving the accuracy of hydrological models and the prediction of extreme events such as floods. Another importance of the model is that the LSTM model, in the development on the study of solutions to early warning system problems, is the method that has achieved the highest accuracy as indicated in the study by Atashi et al. [87] compared to other models such as ARX and NARX. This is crucial for finding optimal algorithms in terms of computational time and estimation accuracy for flood early warning systems using memory limited embedded systems. This study also compares the computational time of the proposed models NARX, ARX and LSTM, showing that although LSTM requires more computational time, it provides the highest accuracy among these models as indicated in the study by Mu˜noz et al. [88].

p. 82

4.1.1

Experimental setup To validate the proposed approach, two analyses are conducted in this study. The first aspect involves comparing the short and long-term recurrent neural network (LSTM) approach for a 4th-order system.

This validation includes a visual comparison of real and estimated signals using NARX and ARX methods, along with a quantitative evaluation based on Root Mean Squared Error (RMSE) and the Nash–Sutcliffe model efficiency coefficient (NSE). The second aspect relates to a forecasting approach for future time steps based on predictions and measurements. It’s important to note that, for this analysis, 90% of the data sample is chosen to train the LSTM recurrent neural network (RNN), corresponding to 1417 out of 1578 total data points. This training is validated through a test, using the remaining 10%, equivalent to 158 data points from the total sample.

4.1.2

LSTM (Long Short-Term Memory) Network An LSTM is a special functional block of recurrent neural networks (RNN) with a short-term to long-term memory [89]. It is an evolution of RNNs and helps to solve the evanescent gradient problem, where during training the gradients of the weights become smaller and smaller and thus the network stops storing useful information [90]. LSTM cells have three types of gates - an input gate, a memory and forgetting gate, and an output gate - for storing memories of past experiences. Short-term memory is retained for a long time and network [91] behavior is encoded in the weights. LSTM networks are particularly well suited for making predictions based on time-series data, such as handwritten text recognition and speech recognition. Moreover, LSTM is an architecture that emerged to alleviate these problems, because in RNNs there is the problem that past data cannot be remembered for a long time, there are six parameters in total. Through the four-gate structure, not only short-term memory but also long-term memory can be solved. for our multivariable model we employ the structure in Figure 4.1.

LTSM

LTSM

LTSM

LTSM

objective objective objective objective yt yt+1 yN-1 yN1 yN2 u1 u2 u3 u4 un u1 u2 u3 u4 un ~ ~ ~ ~ ~ Bias Bias u1 u2 u3 u4 un ~ ~ ~ ~ ~ Bias u1 u2 u3 u4 un ~ ~ ~ ~ ~ Figure 4.1: Multivariate schematic of the long short-term memory (LSTM) neural network.

In fig 4.2, the neural network starts with an input layer of sequences followed by an LSTM layer. The neural network ends with a fully connected layer and an output

p. 83

regression layer.

Secuence Input Secuence Input Secuence Input Secuence Input

LSTM

Fully Connected Regression Output Regression Output Figure 4.2: Architecture of a regression LSTM neural network:summarizes a visual depiction detailing the structure of a Long Short-Term Memory (LSTM) neural network designed for regression tasks.

It visually communicates the network’s components, including input data, LSTM layers for capturing temporal patterns, potential hidden layers, and the output layer for continuous prediction, facilitating the understanding of the network’s design and function in your study.

In fig 4.3, the diagram shows how the gates forget, update and generate the cell and hidden states. These components control the state of the cell and the hidden state of the layer, among these are: entrance door (i), door of oblivion (f), Cell candidate (g), exit door (o).

Forget Update Output Figure 4.3: Data flow structure in time unit t:this diagram illustrates the flow of data in time unit t. This diagram shows how the gates forget, update and generate the cell and hidden states.

The weights that can be learned from the LSTM layer are input weights W (InputWeights), recurrence weights R (RecurrentWeights) and bias b (Bias). The matrices W, R and b are the combination of input weights, cycle weights and offsets for each component, respectively. This layer connects matrices according to the following Eq. (4.1):

p. 84

W =   Wi Wf W g W 0  , R =   Ri Rf Rg R0  , b =   bi bf bg bo  

(4.1)

,where i, f, g and o determine the entry gate, forgetting gate, cell candidate and exit gate, respectively.

The cell state at time unit t is given by Eq. (4.2) Ct = ft ⊙Ct−1 + it ⊙gt,

(4.2)

,where ⊙determines the Hadamard product (element-level vector multiplication). The hidden state at time unit t is given by Eq. (4.3) ht = ot ⊙σc + Ct,

(4.3)

,where σc determines the state activation function. By default, the lstmLayer function uses the hyperbolic tangent function (tanh) to calculate the state activation function. The Eq.s of the components in time unit t are described below. In Eq. (4.4),we describe the gateway Eq. entrance door, in Eq. (4.5) oblivion door , in Eq. (4.6) cell candidate, in Eq. (4.7) open door. it = σg(Wiut[k] + Riht−1 + bi)

(4.4)

ft = σg(Wfut[k] + Rfht−1 + bf)

(4.5)

gt = σc(Wgut[k] + Rght−1 + bg)

(4.6)

ot = σg(Wout[k] + Roht−1 + bo)

(4.7)

p. 85

4.1.3

Estimation Results.

This subsection showcases the outcomes of short- and long-term forecasting using the multivariate LSTM model for a system of order 4.

Additionally, it examines the short-term response of the multivariate NARX and ARX models. The results of the LSTM method for both water level outlets are depicted in Figure 4.4 and Figure 4.5, both Figures show the short-term estimation results for the first water level outlet, compared to the actual measurements. Similarly, it shows the short-term forecast results for the second water level output, juxtaposed with the corresponding actual measurements, therefore, it is shown that the actual measurements are efficiently estimated using the multivariate LSTM model.

0 200 400 600 800 1000 1200 1400 1600

Samples Output Level 1 LSTM Station 1[m] Y-Real Y-Estim LSTM

10

9

8

7

6

5

4

3

2

Figure 4.4: Short-term Estimation of water level Bel´en of Bajira at hydrological station 1 using an LSTM model.

0 200 400 600 800 1000 1200 1400 1600

Samples Output Level 2 LSTM Station 2[m]

7

6

5

4

3

2

1

0

Y-Real Y-Estim LSTM Figure 4.5: Short-term Estimation of water level Quibd´o at hydrological station 2 using an LSTM model.

Evaluating the prediction outcomes depicted in Figures 8-10, the proposed LSTM multivariate nonlinear neural network (NNN) exhibits satisfactory performance

p. 86

upon visual inspection, mathematical computation of data analysis technique and the computational runtime of the algorithm.

To discern which approach most accurately captures the dynamics of the proposed model, a quantitative assessment is conducted. The mean square error is calculated for each considered method to compare actual measurements with their corresponding forecasts. The results of the mean square error for the multivariate LSTM approach are presented in the table. It is evident that the estimation error of the proposed LSTM model is lower than that of other models. Furthermore, the NSE coefficient value is closest to 1, aligning with the proposed theory that estimation is excellent when its value approaches 1.

In Table 4.1, is shown an analysis Regression metric data for the most relevant squared error estimates of the nonlinear recurrent neural network LSTM and the nonlinear neural network NARX, aditionality, shows the computational execution time of the nonlinear neural network algorithm, in order to show that the LSTM recurrent neural network (RNN) model validated with the NARX and ARX neural networks do not require much computational work from the machine’s memory.

Table 4.1:

RMSE, NSE for a system long short-term and computational execution time, Tic operates with the Toc function to measure the elapsed time. nonlinear models RMSE Output 1 RMSE Output 2 NSE Output 1 NSE Output 2 Tic-Toc

LSTM

0.0067

0.0028

0.9990

0.9991

0.0089

NARX

0.0052

0.0060

0.9990

0.9983

0.0051

ARX

0.0275

0.0071

0.9972

0.9980

0.0054

In Table 4.1, the analysis reveals that the RMSE of the estimation of both outputs from the LSTM recurrent neural network, corresponding to hydrological stations 1 and 2, is lower than that of the NARX neural network and the ARX model with a nonlinear structure. Specifically, the RMSE for the water level output from the LSTM model for station 2 is 0.0028, compared to 0.0060 for the NARX model and 0.0071 for the ARX model. This indicates an improvement in the RMSE of 53.4% of the water level in meters compared to the NARX model and 60.5% of the water level in meters compared to the ARX model. Moreover, the Nash–Sutcliffe efficiency coefficient for both outputs of the LSTM recurrent neural network approaches 1, suggesting superior performance compared to the other models considered. This coefficient signifies that the closer it is to 1, the more accurate the estimation provided by the proposed LSTM neural network model. It also illustrates the response of the computational execution time for the (RNN) LSTM model in the short term in comparison with the feedforward neural network (NARX) model and the ARX model. It is evident that the (RNN) LSTM model requires more execution time due to its complex structure and operation. Specifically, the RNN LSTM model takes longer to execute because of its recurrent nature, with an execution time of 0.0089 s. In contrast, the NARX model, with its feedforward structure, has a shorter execution time of 0.0051 s, reflecting a 57.3% reduction in computational time compared to the RNN LSTM. The ARX model also shows a relatively short execution time when obtaining algorithm results.

p. 87

In Table 4.2, the model parameters of the nonlinear LSTM, NARX, and ARX neural networks that were used to train and test the models are shown. Table 4.2: Model parameters of the nonlinear neural networks. nonlinear models Learning rate # input nodes # hidden nodes # output nodes

LSTM

0.02

4

64

2

NARX

0.02

4

64

2

ARX

0.02

4

-

2

Table 4.2 shows the parameters of the LSTM, NARX, and ARX linear neural networks. The first of these models is an LSTM network, which has a specialized architecture for handling sequences of data, such as time series. With 64 nodes in its hidden layer, the LSTM network has the ability to learn and model complex relationships between inputs and outputs over time. Its learning rate of 0.2 controls the rate of adjustment of model parameters during training, influencing its ability to converge to the optimal solution. The training of the neural networks was performed with 50 epochs for both the LSTM neural network and the NARX neural network. The second model is an NARX network, which also uses a neural network structure for prediction. Like the LSTM network, the NARX network has four inputs and two outputs. However, its architecture and inner workings may differ, as it is designed to incorporate feedback information from previous outputs into the current prediction. Finally, the third model is a linear ARX model with a nonlinear least squares structure. Although it uses a different structure from that of neural networks, this model is still able to capture nonlinear relationships between inputs and outputs. With four inputs and two outputs, the ARX model employs both linear and nonlinear modeling methods to make predictions.

4.1.4

Forecasting Future Time Steps Based on Predictions and Forecasting Future Time Steps Based on Measurements Time series are used for chronologically ordered sequences of data that are, in principle, equally spaced in time. Forecasting is the process of predicting future values of a time series based on previously observed patterns (autoregressive) or by including external variables and forecasting future time-series values based on measurements from an LSTM recurrent neural network (RNN) model.

In Figure 4.6, the estimation of the first output of the water level with future time-series values of the LSTM recurrent neural network (RNN) model is presented. In Figure 4.7, the estimation of the second output of the water level with future time-series values of the LSTM recurrent neural network (RNN) model is presented.

p. 88

Y-Estim Y-Train Train Y-Forecast Y-Test Test

0 200 400 600 800 1000 1200 1400 1600

Samples Output Level 1 Forecast [m]

10

9

8

7

6

5

4

3

2

Figure 4.6: Water level 1 output from hydrological station 1, located in the town of Bel´en de Bajir´a time step response based on future predictions.

0 200 400 600 800 1000 1200 1400 1600

Samples Output Level 2 Forecast [m]

7

6

5

4

3

2

1

0

-1

Y-Estim Y-Train Y-Forecast Y-Test Train Test Figure 4.7: Water level 2 output from hydrological station 2, located in the town of Bel´en de Bajir´a time step response based on future predictions. The Figure 4.6 and Figure 4.7, presents the forecasting of future time-series values of the LSTM recurrent neural network (RNN) model based on previously observed patterns (autoregressive) or by including relevant external variables. This forecasting technique allows the LSTM model to analyze and learn from historical data to identify patterns and trends, enabling it to make accurate forecasts of future time-series values. In addition, the ability to incorporate external variables into the model provides additional information that can further improve forecast accuracy by capturing external influences and complex relationships in the data. This flexibility and adaptability make LSTM models widely used in a variety of time-series forecasting applications, such as financial forecasting, sales analysis, and weather forecasting.

In Figure 4.8 and Figure 4.9 outputs forecasting response of future time steps based on measurements. Long-term estimation on measurements response of the two water level outputs of the LSTM model for the hydrological stations, the Figure shows Y-Real which corresponds to the real data of the study sample and Y-Estim which corresponds

p. 89

to the estimation of this real data by the LSTM model during the train stage, and Y-Test corresponds to the real data and Y-Forecast corresponds to the predictions during the test stage.

Output Level 1 Forecast [m]

10

9

8

7

6

5

4

3

2

Y-Estim Y-Train Y-Forecast Y-Test Train Test

0 200 400 600 800 1000 1200 1400 1600

Samples Figure 4.8: Water level 1 output from hydrological station 1, located in the town of Bel´en de Bajir´a response of future forecasts of time steps based on measurements.

0 200 400 600 800 1000 1200 1400 1600

Samples Output Level 2 Forecast [m]

7

6

5

4

3

2

1

0

Y-Estim Y-Train Y-Forecast Y-Test Train Test Figure 4.9: Water level 2 output from hydrological station 2, located in the town of Quibd´o response of future forecasts of time steps based on measurements. Figure 4.8 and Figure 4.9, visualizes the process of forecasting future values over a measured time series by using a recurrent neural network model (RNN) LSTM. This technique uses historical data to learn complex patterns and make accurate predictions about how the series will evolve in the future. The figure shows how the model interprets past measurements and generates projections for the next time steps. In Figure 4.10 and Figure 4.11, is presented a detailed comparison of the forecasting of future time series values based on measurements from the LSTM recurrent neural network (RNN) model and response of time step based on future predictions. In Figure 4.10 are presented the output 1 and output 2 forecasting responses based on

p. 90

future predictions. In Figure 4.11 are presented the output 1 and output 2 forecasting responses of future time steps based on measurements.

1418 1498 1578

Samples

1418 1498 1578

Samples Output Level 2 Forecast [m] Output Level 1 Forecast [m]

10

8

6

4

2

4

3

2

1

0

-1

Y-Forecast Y-Test Y-Forecast Y-Test Figure 4.10: Water level 1 output from hydrological station 1, located in the town of Bel´en de Bajir´a time step response based on future predictions.

1418 1498 1578

Samples

1418 1498 1578

Samples Output Level 2 Forecast [m] Output Level 1 Forecast [m]

10

8

6

4

2

4

3

2

1

Y-Forecast Y-Test Y-Forecast Y-Test Figure 4.11: Water level 2 output from hydrological station 2, located in the town of Bel´en de Bajir´a time step response based on future predictions. Long-term Forecasting on measurements response of the two water level outputs of the LSTM model for the hydrological stations. In Figure 4.10, output 1 and output 2 forecasting responses based on future predictions. In Figure 4.11, output 1 and output 2 forecasting responses of future time steps based on measurements. Forecasting future time series values based on measurements from the LSTM recurrent neural network (RNN) model involves using historical data to train the model and generate predictions for future time steps. The LSTM model analyzes past patterns and trends in the data to make accurate predictions about future values. Once the model is trained, it can be used to forecast future time series values by feeding it with input data representing past measurements.

p. 91

Additionally, the response of each time step based on future predictions involves evaluating the model’s performance by comparing its predicted values with the actual values observed in the future.

This assessment helps determine the accuracy and reliability of the LSTM model in forecasting future time series values. By analyzing the response of each time step, one can identify any discrepancies or errors in the predictions and refine the model accordingly to improve its forecasting capabilities. In Table 4.3, An analysis of two metric regression measures for the LSTM neural network of future forecasts and future forecasts based on the measurements is shown. Table 4.3: RMSE and NSE for an LSTM system of future forecasts and future forecasts based on measurements.

nonlinear model RMSE Output 1 RMSE Output 2 NSE Output 1 NSE Output 2 LSTM Forecast-Future

1.2325

3.4520

0.3849

0.1234

LSTM Forecast-Measurements

0.0620

0.0160

0.9692

0.9894

From Table 4.3, it is observed that the total estimation error of the LSTM recurrent neural network of future forecasts is larger than the estimation error of the LSTM neural network of future forecasts based on actual measurements, furthermore, it is observed that the Nash-Sutcliffe Efficiency coefficient of the LSTM recurrent neural network of future forecasts based on actual measurements closer to 1, which means that the proposed LSTM neural network model is very good for short-term predictions and very regular in the long term when the data for future predictions are made up of very few data samples as in the case of the study, since this type of neural networks require a lot of training data for future predictions.

4.2

Water Level Forecasting based on an Ensemble Kalman Filter with NARX Neural Network Model This research, presents a novel application for multivariate short-term forecasting for hydrological variables based on a multivariate NARX model coupled to an EnKF, the importance of the hybrid NARX-EnKF model for water level prediction lies in its ability to handle the complexity and uncertainty of hydrological systems. The NARX model (Nonlinear Autoregressive Model with Exogenous Inputs) is effective in capturing the nonlinear relationships between multiple hydrological variables, such as water level, flow and precipitation. This allows the model to make accurate shortand medium-term predictions, considering the internal dynamics of the system. To this end, two hydrological stations are considered located in the Atrato River in Colombia, where the hydrological variables of water level, water flow, and water precipitation are considered. It is worth noting that the hydrological stations are monitored by the Institute of Hydrology, Meteorology, and Environmental Studies (IDEAM) of Colombia.

p. 92

A multivariable NARX structure is used to describe the hydrological model by using the data measurements. The data include measurements that are sampled every 12 hour during a period of 789 days. In order to evaluate the performance of the proposed approach, a comparison analysis is developed by considering several additive noise disturbances and the NARX forecasting with and without the EnKF. The performance is evaluated in terms of the Mean-Square Error (RMSE) and the Nash-Sutcliffe model efficiency coefficient (NSE).

4.2.1

NARX neural network model for hydrologycal level forecasting The dynamic of the hydrological variables is defined by considering as follows: yk = f(xk) + ϵk

(4.8)

being n the order of the NARX model and f(.) the nonlinear function, and ηk the additive noise at time instant k, where xk xk =   yk−1

...

yk−n uk−1

...

uk−n  

(4.9)

Consider the following state space nonlinear model of hydrological activity: xk+1 =   f(xk)   I

0

. . .

0

0

0

I

. . .

0

0

0

0

. . .

0

0

0

. . .

I

0

0

0

. . .

0

I

0

    +  

0

...

0

I

...

0

  uk + ηk

(4.10)

xk+1 = F (xk, uk) + ηk

(4.11)

In order to consider the NARX model of (4.8), the inputs are selected as u[k −j] and y[k −j], with j = 1, . . . , n, which correspond to a n −th order model, in this work a 4−th order model (n = 4) is considered according to [80] where an analysis of the order selection is performed and the lowest estimations error is obtained for 3 −rd order model or higher. Therefore, by considering the variables described in (3.5), the proposed NARX model consists of 24 inputs and 2 outputs.

For a better understanding, a flowchart is presented using a NARX neural network structure, coupled to an EnKF filter and decoupled as shown in Fig. 4.12 and Fig. 4.13.

p. 93

Figure 4.12: Methodological architecture EnKF-NARX

p. 94

Figure 4.13: Architecture of NARX, NARX-EnKF

4.2.2

EnKF forecasting The state estimation xk can be obtained by applying the EnKF in two stages [92]. First, the forecast stage, where an ensemble of q forecasted state estimates is computed xfi k = F (¯xa k, uk) + ωi k

(4.12)

being i = 1, . . . , q, ωi k ∈Rn×1 is a zero mean random variable with normal distribution an covariance Ωk ∈Rn×n.

The sample error covariance matrix computed from ωi k converges to Ωk as q →∞. The forecasted measurement is given by yfi k = f ∗(xfi k )

(4.13)

where fi denotes the i−th forecast ensemble member. Then, the ensemble mean for states ¯xf k ∈Rn×1 is defined by ¯xf k = 1 q q X i=1 xfi k

(4.14)

and the ensemble mean for measurements ¯yf k is defined by ¯yf k = 1 q q X i=1 yfi k

(4.15)

The ensemble error matrix around the ensemble mean is defined as Ef k = h xf1 k −¯xf k

· · ·

xfq k −¯xf k i

(4.16)

being Ef k ∈Rn×q, and the ensemble output error Ef yk = h yf1 k −¯yf k

· · ·

yfq k −¯yf k i

(4.17)

p. 95

being Ef yk ∈Rd×q, and where the forecast covariances are approximated as P f k =

1

q −1Ef k  Ef k ⊤ P f xyk =

1

q −1Ef k Ef yk ⊤ P f yyk =

1

q −1Ef yk Ef yk ⊤ being P f k ∈Rn×n, P f xyk ∈Rn×d and P f yyk ∈Rd×d.

The second stage is the analysis stage, where an ensemble of perturbed observations yi k is obtained as follows yi k = yk + υi k

(4.18)

where υi k ∈Rd×1 is a zero mean random variable with normal distribution an covariance Υk ∈Rd×d. The sample error covariance matrix computed from υi k converges to Υk as q →∞. Also, the EnKF gain matrix is approximated as Kk = P f xyk P f yyk −1

(4.19)

where it is worth noting that P f yyk −1 is an inverse of d × d, being d << n.

Therefore, an ensemble of data assimilation cycles is obtained as follows: xai k = xfi k + Kk  yi k −f ∗(xfi k ) 

(4.20)

where ai denotes the i−th analysis ensemble member. And finally, the ensemble mean of the analysis stage is computed as ¯xa k = 1 q q X i=1 xai k

(4.21)

where ¯xa k is the estimated activity at sample k.

4.2.3

Novelty of the proposed model NARX-EnKF The use of a coupled NARX-EnKF (Nonlinear Autoregressive Model with Exogenous Inputs - Ensemble Kalman Filter) model to predict water level is crucial because of its ability to improve prediction accuracy and handle uncertainty effectively. The NARX model captures complex nonlinear relationships between hydrological variables, while the EnKF allows real-time data assimilation, continuously adjusting the model state with new observations. This results in more accurate and robust predictions, especially in dynamic environments with noisy data and variable conditions.

p. 96

In addition, the combination of NARX and EnKF is particularly useful in complex hydrologic systems with multiple stations and interrelated variables. This approach improves the management and planning of disaster risks, such as floods, facilitating the implementation of early warning systems and the optimization of resources. Overall, the NARX-EnKF hybrid model provides a valuable tool to improve the accuracy of hydrological predictions, adapt to external disturbances, and support informed decision-making in water management and extreme event preparedness.

4.2.4

Experiment setup To validate the proposed approach, a comparative analysis of the proposed NARX approach coupled and decoupled to a nonlinear recursive filter EnKF is performed through a state space model of the nonlinear autoregressive model with exogenous variables in order to validate the prediction of the nonlinear model for a system of order 4 for two hydrological stations with 24 inputs and two outputs using online training, it should be noted that this model of river level prediction is entered 5 noise perturbations as variance parameters in order to check the robustness of the model in terms of level prediction given by the model, It should be noted that this river level prediction model is given 5 noise perturbations as variance parameters in order to check the robustness of the model in terms of the level prediction given by the model, therefore, a visual comparison of the real and estimated signals for the nonlinear model mentioned for each of the 5 noise parameters of the NARX model with the EnKF filter and the NARX model without EnKF filter is presented, also a quantitative evaluation based on 2 techniques of the machine learning regression metrics already described above is performed.

The feedforward network structure for the NARX neural network approach considers a hidden layer with 20 nodes, the NARX rectified linear unit (ReLU) activation function is selected for the proposed approach, where a ReLU is a piecewise linear function that will output the input directly if it is positive or zero otherwise.A NARX model combined with an Ensemble Kalman Filter (EnKF) is essential for hydrological forecasting because of its ability to handle external disturbances such as white noise, thus improving the accuracy of the estimates.

The EnKF continuously adjusts NARX model predictions, allowing dynamic adaptation to rapid changes in hydrologic and meteorological conditions. This hybrid approach also facilitates uncertainty assessment and constant optimization of model parameters, providing more reliable and accurate predictions. These features are particularly valuable for risk management and flood early warning systems, improving decision-making in water resources management and disaster mitigation. The implementation of the NARX model and the proposed EnKF filter based on neural networks is performed in an algorithm developed in matlab under the sim function for its machine learning training, it should be noted that the members with which the training of the NARX is done with the filter and without the EnKF filter is 80 members.

p. 97

4.2.5

NARX-EnKF Neural Network Validation Results The results obtained from the proposed model are shown graphically below: Fig 4.15 shows the box plot of the proposed model and the NARX neural network model, while Fig 4.15 shows the NARX model coupled to the EnKF filter. Figure 4.14: NARX neural network diagram without EnKF. Figure 4.15: NARX neural network diagram coupled EnKF. the validation of the proposed model is shown, obtaining the results of the histogram and the best reference value of the RMSE for the proposed model.In Fig. 4.16 shows the distribution of the values of the weights in the neural network, which can help identify any irregularities or patterns that may be affecting the performance of the model. In Fig. 4.17 best model performance response of the neural network model 24 inputs for being an order 4 model, with a hidden layer of 40 nodes and two outputs.

0

500

1000

1500

2000

Instances Error Histogram with 20 Bins

-1.114

-1.003

-0.8922

-0.7811

-0.6701

-0.559

-0.448

-0.3369

-0.2259

-0.1148

-0.00375

0.1073

0.2184

0.3294

0.4405

0.5515

0.6626

0.7736

0.8847

0.9958

Errors = Targets - Outputs Training Validation Test Zero Error Figure 4.16: NARX-EnKF Feedforward model response

p. 98

12 Epochs

10-3

10-2

10-1

100

101

102

Mean Squared Error (mse) Train Validation Test Best Figure 4.17: NARX-EnKF Best model performance response In the following Figures shows a visual representation of how well the model fits the data and the training state of a NARX network. In Fig. 4.18 The training status of a NARX network is displayed, usually showing graphs that include the evolution of the training error, the validation error, and the test error over the training epochs.In Fig. 4.19 provides a comprehensive view of the training process and are useful for diagnosing problems such as overfitting and inadequate convergence.

2

4

6

8

Target

1

2

3

4

5

6

7

8

9

Output ~= 1*Target + 0.0068 Training: R=0.99929 Data Fit Y = T

2

4

6

8

Target

1

2

3

4

5

6

7

8

9

Output ~= 0.99*Target + 0.019 Validation: R=0.99861 Data Fit Y = T

2

4

6

8

Target

1

2

3

4

5

6

7

8

9

Output ~= 1*Target + 0.0078 Test: R=0.99844 Data Fit Y = T

2

4

6

8

Target

1

2

3

4

5

6

7

8

9

Output ~= 1*Target + 0.0087 All: R=0.99908 Data Fit Y = T Figure 4.18: NARX-EnKF Visual representation of the model fitted to the data.

p. 99

gradient Gradient = 0.020572, at epoch 12

10-5

10-4

10-3

mu Mu = 0.0001, at epoch 12

0

2

4

6

8

10

12

12 Epochs

0

2

4

6

val fail Validation Checks = 6, at epoch 12 Figure 4.19: NARX-EnKF model validation.

In Fig. 4.20 The taylor diagram is shown for the two models NARX and NARX coupled to an EnKF filter, there we can observe the analysis for the standard deviation of the two models, the Centered Root Mean Square Difference and Correlation, as a result we observe how the NARX model coupled to the EnKF filter presents a better correlation coefficient in terms of the estimation of the two models. Standard deviation

0.5

1

1.5

2

2.5

0

0.5

0

1

0

1.5

0

2

0

2.5

1

0.99

0.95

0.9

0.8

0.7

0.6

0.5

0.4

0.3

0.2

0.1

0

Co r r e l a t i o n Co e f f i c i e n t R M S D Real

NARX

EnKf Figure 4.20: Description of the Diagram Taylor.

Estimation Results.

In the following we present the graphical response of the real signal estimation with the estimated one of the two models NARX and NARX -EnKF, in reference to the external white noise perturbations of the two proposed models.

For a white noise of 0.01 is observed. In Fig. 4.21 the estimation of the first water level output for an order 4 system is presented, in which the real signal with noise, the estimated signal and the real signal without noise will be observed. In Fig. 4.22 the estimation of the second water level output for an order 4 system is presented, in which

p. 100

the real signal with noise, the estimated signal and the real signal without noise will be observed.

0

200

400

600

800

1000

1200

1400

1600

Time(days)

2

3

4

5

6

7

8

9

10

Output Level 1 Y-real noise Y-Estim EnKF Y-real noiseless Figure 4.21: Output 1 water level forecast for an order 4 system of a multivariate NARX model coupled to an EnKF filter with noise variance parameter of 0.01.

0

200

400

600

800

1000

1200

1400

1600

Time(days)

0

1

2

3

4

5

6

7

Output Level 2 Y-real noise Y-Estim EnKF Y-real noiseless Figure 4.22: Output 2 water level forecast for an order 4 system of a multivariate NARX model coupled to an EnKF filter with noise variance parameter of 0.01. For a white noise of 0.1 is observed.

In Fig. 4.23 the estimation of the first water level output for an order 4 system is presented, in which the real signal with noise, the estimated signal and the real signal without noise will be observed. In Fig. 4.24 the estimation of the second water level output for an order 4 system is presented, in which the real signal with noise, the estimated signal and the real signal without noise will be observed.

p. 101

1000

1500

Time(days)

2

3

4

5

6

7

8

9

10

Output Level 1 Y- real noise Y-Estim EnKF Y-real noiseless Figure 4.23: Output 1 water level forecast for an order 4 system of a multivariate NARX model coupled to an EnKF filter with noise variance parameter of 0.1.

0

200

400

600

800

1000

1200

1400

1600

Time(days)

0

1

2

3

4

5

6

7

Output Level 2 Y-real noise Y-Estim EnKF Y-real noiseless Figure 4.24: Output 2 water level forecast for an order 4 system of a multivariate NARX model coupled to an EnKF filter with noise variance parameter of 0.1. For a white noise of 0.4 is observed.

In Fig. 4.25 the estimation of the first water level output for an order 4 system is presented, in which the real signal with noise, the estimated signal and the real signal without noise will be observed. In Fig. 4.26 the estimation of the second water level output for an order 4 system is presented, in which the real signal with noise, the estimated signal and the real signal without noise will be observed.

p. 102

1000

1200

1400

1600

Time(days)

2

3

4

5

6

7

8

9

10

Output Level 1 Y-real noise Y-Estim EnKF Y-real noiseless Figure 4.25: Output 1 water level forecast for an order 4 system of a multivariate NARX model coupled to an EnKF filter with noise variance parameter of 0.4.

0

200

400

600

800

1000

1200

1400

1600

Time(days)

0

1

2

3

4

5

6

7

Output Level 2 Y-real noise Y-Estim EnKF Y-real noiseless Figure 4.26: Output 2 water level forecast for an order 4 system of a multivariate NARX model coupled to an EnKF filter with noise variance parameter of 0.4. The following is presented the results of the estimation of the multivariate NARX neural network model without recursive nonlinear EnKF filter for each of the two water level outputs.For a white noise of 0.01 is observed. In Fig. 4.31 the estimation of the first water level output for an order 4 system is presented, in which the real signal with noise, the estimated signal and the real signal without noise will be observed. In Fig. 5.6 the estimation of the second water level output for an order 4 system is presented, in which the real signal with noise, the estimated signal and the real signal without noise will be observed.

p. 103

1000

1500

Time(days)

3

4

5

6

7

8

9

10

Output Level 1 Y-real noise Y-Estim NARX Y-real noiseless Figure 4.27: Output 1 water level forecast for an order 4 system of a multivariate NARX neural network model without recursive nonlinear EnKF with noise variance parameter of 0.01.

0

200

400

600

800

1000

1200

1400

1600

Time(days)

0

1

2

3

4

5

6

7

Output Level 2 Y-real noise Y-Estim NARX Y-real noiseless Figure 4.28: Output 2 water level forecast for an order 4 system of amultivariate NARX neural network model without recursive nonlinear EnKF with noise variance parameter of 0.01.

For a white noise of 0.1 is observed.

In Fig. 4.29 the estimation of the first water level output for an order 4 system is presented, in which the real signal with noise, the estimated signal and the real signal without noise will be observed. In Fig. 4.30 the estimation of the second water level output for an order 4 system is presented, in which the real signal with noise, the estimated signal and the real signal without noise will be observed.

p. 104

1000

1200

1400

1600

Time(days)

2

3

4

5

6

7

8

9

10

Output Level 1 Y-real noise Y-Estim NARX Y-real noiseless Figure 4.29: Output 1 water level forecast for an order 4 system of a multivariate NARX neural network model without recursive nonlinear EnKF with noise variance parameter of 0.1.

0

200

400

600

800

1000

1200

1400

1600

Time(days)

0

1

2

3

4

5

6

7

Output Level 2 Y-real noise Y-Estim NARX Y-real noiseless Figure 4.30: Output 2 water level forecast for an order 4 system of amultivariate NARX neural network model without recursive nonlinear EnKF with noise variance parameter of 0.1.

For a white noise of 0.4 is observed.

In Fig. 4.31 the estimation of the first water level output for an order 4 system is presented, in which the real signal with noise, the estimated signal and the real signal without noise will be observed. In Fig. 5.6 the estimation of the second water level output for an order 4 system is presented, in which the real signal with noise, the estimated signal and the real signal without noise will be observed.

p. 105

1000

1200

1400

1600

Time(days)

3

4

5

6

7

8

9

10

Output Level 1 Y-real noise Y-Estim NARX Y-real noiseless Figure 4.31: Output 1 water level forecast for an order 4 system of a multivariate NARX neural network model without recursive nonlinear EnKF with noise variance parameter of 0.4.

0

200

400

600

800

1000

1200

1400

1600

Time(days)

-1

0

1

2

3

4

5

6

7

8

Output Level 2 Y-real noise Y-Estim NARX Y-real noiseless Figure 4.32: Output 2 water level forecast for an order 4 system of amultivariate NARX neural network model without recursive nonlinear EnKF with noise variance parameter of 0.4.

In Table 4.4 The parameters of the NARX nonlinear neural network model and the NARX-EnKF model are shown.

Table 4.4: Parameters of the NARX model, NARX-EnKF.

nonlinear models Learning rate # input nodes # hidden nodes # output nodes # Epoch EnKF

0.02

4

80

2

12

NARX

0.02

4

40

2

12

In addition, Table

4.5 show the RMSE estimates of the NARX model without the

EnKF filter and with the EnKF filter for the two outputs of the hydrologic variable level for the two hydrologic stations already described. It should be noted that the EnKF is an implementation of the Bayesian Monte Carlo update problem: given the probability density function of the modeled system state.

p. 106

In Table 4.5 a metric regression analysis of the data for the Root Mean Squared Error (RMSE) estimates of a NARX neural network with noise coupled to an EnKF and with decoupled noise without an EnKF is shown.

Table 4.5: Root mean square estimation (RMSE) with Gaussian noise for NARX neural network coupled with EnKF and without EnKF.

NARX coupled with EnKF Variance parameter RMSE Output 1 RMSE Output 2 Total

0.01

0.0158

0.0016

0.0174

0.05

0.0231

0.0075

0.0306

0.1

0.0389

0.0215

0.0604

0.2

0.0636

0.0327

0.0963

0.4

0.1238

0.1515

0.2753

NARX without EnKF

0.01

0.0250

0.0026

0.0276

0.05

0.0265

0.0081

0.0346

0.1

0.0486

0.0158

0,0644

0.2

0.0901

0.0659

0.1560

0.4

0.2484

0.2116

0.4600

From the Table 4.5 It is observed that the total estimation error of the Feedforward NARX neural network is greater than the estimation error of the NARX neural network coupled to the nonlinear recursive filter EnKF, it can be observed that the estimation of the two-station hydrological model used in this research presents good estimation of the model outputs, this can be observed in the calculated estimation error which is quite low and well below 1.Also It is observed that the total estimation error of the feedforward neural network is larger than the estimation error of the NARX neural network decoupled to the EnKF nonlinear recursive filter, it can be observed that the estimation of the two-station hydrological model used in this research presents good estimation but not as good as the NARX neural network coupled with the EnKF. A metric regression analysis of data for the Nash-Sutcliffe model efficiency coefficient (NSE) for the NARX neural network with noise coupled and decoupled to the EnKF filter is shown in Table 4.5 and Table 4.6.

p. 107

Table 4.6: Nash-Sutcliffe model efficiency coefficient (NSE) with Gaussian noise for the NARX neural network coupled with EnKF and decoupled without EnKF. NARX coupled with EnKF Variance parameter NSE Output 1 NSE Output 2

0.01

0.9918

0.9994

0.05

0.9917

0.9992

0.1

0.9916

0.9991

0.2

0.9913

0.9987

0.4

0.9892

0.9970

NARX without EnKF

0.01

0.9953

0.9996

0.05

0.9936

0.9978

0.1

0.9908

0.9949

0.2

0.9770

0.9814

0.4

0.9317

0.9401

Another aspect to highlight in the results obtained in this study is the analysis of the Nash-Sutcliffe Efficiency (NSE) coefficient, which indicates that the predictions are optimal if this value is 1 or tends to 1, that is, the model of the data obtained when making the predictions of the NARX model coupled and decoupled to the nonlinear recursive filter EnKF proposed in this study, it can be observed that for the NARX neural network without EnKF filter the coefficient as the noise of the variance parameter is higher, the NSE coefficient moves away from 1, therefore the estimation of the real signal is not so optimal with respect to the estimated one, this can be observed in Table 4.6, such is the case that for output 2 of the model proposed in the calculation of the NSE with a noise in the variance parameter of 0.4 a value of 0.4 is obtained.

4 a value of 0.9401 is obtained, the opposite case with the NARX model coupled to

the EnKF filter where it is observed how the NSE coefficient conserves optimal values close to 1 at the moment of doing the mathematical calculation of the NSE, we can observe it in Table 4.6, that for a noise of variance 0.4 a value of 0.9970 is obtained, in this case, the filtering of the NARX model contributes to an estimation improvement of 5.7% with respect to the NARX model without filtering. therefore, the EnKF filter contributes favorably to the NARX feedforward network model in terms of giving robustness to the hydrological model in terms of an external perturbation.The NARX model coupled with the Ensemble Kalman filter (EnKF) and subjected to external white noise perturbations has significant and promising future scope. This approach can greatly improve the accuracy and robustness of hydrologic predictions under complex and variable conditions.

With the advancement of data processing techniques and machine learning algorithms, the model could be integrated with real-time monitoring systems, providing more accurate and adaptive predictions. Its application could be extended to areas such as watershed management, water infrastructure planning and natural disaster mitigation.

The model’s ability to adapt to different hydrological conditions and its resistance to external disturbances make it ideal for use in global

p. 108

early warning systems, contributing to the resilience and sustainability of communities in the face of extreme hydrological events.

4.3

Academic discussion The results obtained in this section have been published in the following journals:

• J.

B.

Renteria-Mena,; D.

Plaza, and E.

Giraldo, ”Multivariate Hydrological Modeling Based on Long Short-Term Memory Networks for Water Level Forecasting”.

Information

15,

no.

6:

358,

2024.

https://doi.org/10.3390/info15060358. ISSN: 2078-2489. [93]

• J.

B.

Renteria-Mena,; D.

Plaza, and E.

Giraldo, ”Water-Level Forecasting Based on an Ensemble Kalman Filter with a NARX Neural Network Model”.

Engineering Proceedings

101,

no.

1:

2.

2025.

https://doi.org/10.3390/engproc2025101002. ISSN: 2673-4591. [94]

4.4

Summary The chapter shows that nonlinear hybrid models, based on advanced neural networks and data assimilation techniques, significantly improve the capacity of Early Warning Systems (EWS) to anticipate floods in complex contexts such as the Atrato River. Firstly, the use of LSTM networks, which can predict the short-term water level more accurately by capturing complex temporal relationships between hydrological variables, stands out. This is especially useful in regions with high vulnerability, such as the Colombian Choc´o. Second, a hybrid model that combines NARX with the Ensemble Kalman filter (EnKF) is implemented, achieving robust predictions even in the data’s presence of uncertainty or noise.

Both models were validated with metrics such as RMSE and NSE, showing superior performance over traditional approaches. These technological advances offer practical solutions to strengthen real-time EWS, optimizing decision-making and responsiveness to extreme events.

4.5

Partial conclusions The implementation of LSTM neural networks has proven effective in capturing complex temporal relationships between hydrological variables such as precipitation, flow, and water level. Their ability to process multivariate sequences greatly improves short-term water level prediction, which is crucial for vulnerable contexts with limited monitoring infrastructure, such as the Colombian Choc´o.

p. 109

The hybrid model that integrates NARX with the Ensemble Kalman Filter (EnKF) adds a layer of robustness to the predictive process, allowing to handle the uncertainty inherent in real data and changing conditions. Its stable performance against different noise levels makes it a viable option for complex hydrological scenarios. Both models have shown a superior ability to anticipate sudden variations in water level, even over prediction horizons of up to two days. This anticipatory capacity is essential for timely warnings and preventing impacts on communities exposed to extreme events. Evaluations performed with metrics such as RMSE and NSE confirm the good performance of these models.

In addition, their implementation in real stations of the Atrato River supports their viability as an operational part of the EWS, providing reliability to the monitoring and response systems.

Technological advances offer predictive tools that allow faster and more accurate decisions in the face of hydrometeorological hazards. By integrating artificial intelligence with data assimilation techniques, warning systems in particularly vulnerable regions are strengthened.

p. 110

Chapter 5 Mechanism of Social Appropriation and Knowledge Transfer for the Riverside Communities of Choc´o through a web platform and Virtual Environment The riverine communities of Choc´o, in the Pacific region of Colombia, live in an environment marked by rich biodiversity and a high vulnerability to natural phenomena such as floods and landslides.

Their proximity to bodies of water and the lack of robust early warning infrastructure increase their risks. In this context, implementing technological strategies that facilitate knowledge transfer and social appropriation is presented as a key solution to improving community resilience. Knowledge transfer is essential for equipping communities with the information they need to address environmental risks.

A web platform designed for riverine communities in Choc´o can act as an accessible resource center, providing crucial information on preventive measures and early warnings in real time. This web platform allows users to easily access weather forecasts, risk maps, and emergency action guides. It also facilitates direct communication between inhabitants and authorities, creating a two-way channel that improves the response to hazardous situations. By providing continuous access to this information, the platform informs and empowers the community to make informed decisions that can save lives.

p. 111

5.1

Choc´o Early Warning Web Estimation System (SEAT-Web) SEAT-Web is the user interface designed for the Sentify Early Warning System (EWS), an essential tool in environmental risk management.

This platform allows users to manage and monitor all components of the warning system, either locally or remotely, facilitating an efficient response to potential emergencies.

5.1.1

SEAT-Web Key Functionalities Local and Remote Management: SEAT-Web offers the ability to control elements of the early warning system from anywhere, as long as there is internet access. This flexibility is crucial to ensure that monitoring and response operations can be carried out regardless of the user’s physical location.

Station Monitoring: The interface allows continuous monitoring of the operation of the stations that make up the EWS. This includes real-time data collection and visualization, allowing users to always be aware of environmental conditions and detect any anomalies that may indicate an imminent risk.

Risk Detection through Advanced Models: SEAT-Web uses estimation techniques based on linear and nonlinear models to analyze the data collected by the stations. These models make it possible to predict and detect risks, such as rising river levels, with high accuracy, which is vital for making quick and accurate decisions. Multi-Channel Alert System: When a rise in river levels or any other significant risk is detected, SEAT-Web issues alerts through multiple channels. Users receive immediate notifications via text messages on Telegram, ensuring that critical information arrives quickly and effectively. In addition, the web platform activates sirens and other warning devices, providing an audible warning in affected areas.

5.1.2

Elements of SEAT-Web The SEAT-Web is composed of several key elements that work together to monitor and manage risks, especially in natural disaster contexts. The main components are detailed below:

Operator: The operator is the person responsible for monitoring the early warning system. This user is tasked with overseeing the operation of the EWS, ensuring that data is processed correctly and that alerts are issued in a timely manner when a risk is detected.

Risk Level: The risk level is a graphical or numerical representation that indicates the severity of a risk, such as a flood. At SEAT-Web, these levels are determined using

p. 112

both linear and nonlinear estimation models, which allow for an accurate assessment of the hazard and its potential impact.

Monitoring Station:

Monitoring stations are devices equipped with specialized sensors that capture environmental data, such as water level, wind speed, or soil moisture. This data is automatically sent to SAT for real-time analysis. The main elements of SEAT-Web are:

Elemento Descripci´on Zona Pol´ıgono o circunferencia georeferenciado que representa una regi´on geogr´afica definida por el usuario. Reglas Expresi´on matem´atica que relaciona umbrales sobre los valores reportados por los sensores y un nivel de riesgo. Nivel de riesgo Uno de los siguientes valores: rojo, naranja, verde. Estaci´on Dispositivo electr´onico que transmite informaci´on hacia el SAT.

Sensor Dispositivo que captura mediciones en una estaci´on de monitoreo.

Bocina Dispositivo electromec´anico que genera ruido en varios tonos.

Baliza Conjunto de l´amparas.

Acci´on manual Operaciones que ejecuta el SAT si una regla es verdadera previa acci´on del usuario.

Acci´on autom´atica Operaciones que ejecute el SAT si una regla es verdadera sin intervenci´on del usuario.

Table 5.1: Description of the elements of the early warning system. Control Panel The panel is composed of 5 sections that correspond to the SAT elements and the information they generate: notifications, monitoring, map, alerts, sirens.

p. 113

Figure 5.1: SAT Control Panel.

Map.

The map shows the stations and zones defined in the SAT. The position of the symbols indicates their location.

Description Figure 5.2: Symbology of elements on the map.

Notifications.

Messages related to the following events are displayed here in real time: Alerts: when a rule is met or not met in a zone. On/off of horns and beacons. Network: Change of connection status between SEAT-Web and the stations. Community:

Messages sent from the Telegram groups:

SAT-community, in SAT-communityY are the volunteers.

p. 114

System: Messages related to the operation of the different components of the SAT system.

Monitoring.

Lists the status of the stations and sensors. In the stations tab it is possible to see its type, if it is online or not and by clicking on the arrow you can see the equipment connected to the selected station, currently the SEAT-Web development has three hydrological stations located in the municipality of Bel´en de Bajir´a and two located in the city of Quibd´o, Each station measures with three hydrological sensors which measure flow, level and precipitation. At the same time, this SEAT-Web has two linear and nonlinear prediction models, the ARX and NARX models, where the level of the Atrato river is predicted for the three hydrological stations. The following is the monitoring of the Atrato river level in the three hydrological stations, where it is also shown how the estimation models make the forward prediction of the next river level data.

Figure 5.3: Atrato river water level monitoring station 1.

Figure 5.4: Atrato river water level monitoring station 2.

p. 115

Figure 5.5: Atrato river water level monitoring station 2.

Figure 5.6: Atrato river water level monitoring station 3.

5.2

Virtual Environment of an Early Warning System to Mitigate Flood Risks in the Department of Choc´o The development of a Virtual Environment of an Early Warning System for Flood Risk Mitigation in the Department of Choc´o is a critical response to the need to address natural hazards in this region. With the increasing frequency and intensity of floods, it is imperative to have effective tools to monitor, forecast and mitigate the impacts of these catastrophic events. The use of virtual reality (VR) and extended reality (XR) technologies offers a unique platform to create immersive interactive environments that allow users to better explore and understand flood risks in the Department of Choc´o. Unity, as one of the leading game and virtual environment development tools, provides the ability to build detailed models of terrain, water bodies and local structures to simulate realistic flood scenarios.

By implementing a digital model of the Choc´o environment, fed with real-time geospatial and meteorological data, users can virtually visualize and experience different flooding situations. This not only facilitates emergency planning and preparedness, but also allows the evaluation of the effectiveness of existing mitigation measures and the identification of high-risk areas that require intervention. In addition, the integration of XR technologies such as augmented reality (AR) and mixed reality (MR)

p. 116

extends the capabilities of the Virtual Early Warning Environment. AR can be used to overlay real-time contextual information over the physical environment, providing users with additional data on water levels, safe evacuation routes, and temporary shelter locations.

On the other hand, RM allows interaction between virtual and physical elements, facilitating collaboration between emergency responders and local authorities to coordinate mitigation and rescue actions.

The implementation of a Virtual Early Warning Environment not only has immediate practical applications in disaster management, but also serves as an invaluable educational tool.

Virtual reality-based training and simulation programs can help educate the population about flood risks, safety protocols and preparedness measures, thereby increasing community resilience and reducing vulnerability to future disasters.The creation of a Virtual Early Warning Environment for Flood Risk Mitigation in the Department of Choc´o represents an innovative convergence of VR, XR and geospatial data technologies to address critical natural hazard management challenges. By providing an interactive and collaborative platform, this solution has the potential to save lives, protect property and strengthen the resilience of flood-affected communities in Choc´o.

5.2.1

Design and Construction of the 3D Model of the Hydrological Station In this section we will talk about the final modeling of the hydrological station, this design was made with the help of Autodesk’s Fusion 360 software. The objective of this design is to obtain a .FBX file which could be uploaded to the Unity development engine.

p. 117

Figure 5.7: 3D Model Hydrological Station.

Base Hydrological Station For the base model of the hydrological station, the WXM WS1000 model of the WeatherXM company was taken as a reference, obtaining as a result the design shown in the Figure 5.8, this base was designed with the purpose of grouping different elements, in the first place it works as a support for the ultrasound sensor that is incorporated in the right lateral support shown in the Figure 5.8, On the other hand, the left side allows the incorporation of the rain gauge measurement system, this system was modified and coupled to the needs of the hydrological station, modules were incorporated for a battery, the necessary spaces for the communication system and a support system was designed at the bottom of the base.

p. 118

Figure 5.8: Base Hydrological Station. Ultrasonic Sensor For the design of the ultrasonic sensor we took as a reference the ultrasonic level meter SUL806, the final model is the one shown in the following Figure 5.9, and this would have a range of 10 meters.

Figure 5.9: 3D Model Ultrasonic Sensor.

p. 119

Flow Sensor The flow sensor was designed based on the OTT C31 model, which measures the river velocity by means of a rear propeller, in addition, this sensor would allow transmitting the measurement data to the data acquisition system and the final design implemented in the Unity development engine is shown in the Figure below. 5.10 Figure 5.10: 3D Design of the Flow Sensor.

Rain gauge The rain gauge was based on the design of the model implemented in the WXM WS1000 weather station, which incorporates a conventional pulse-based mechanism, as shown in the Figure below 5.11, In addition, this sensor allows to collect data of the amount of rainfall occurring in the environment and finally provide it to the datalogger, the final design is as shown in the Figure below 5.11.

p. 120

Figure 5.11: 3D Rain Gauge Design.

Data Logger The design of the Data Logger was made with the objective of storing the necessary systems for the storage, processing and evaluation of the data collected in the hydrological station, this data logger will be able to receive data from the flow sensor, rain gauge and level sensor, in addition, it must provide information on the status of the station batteries and finally transmit these data, finally the final design is shown in the Figure below5.12.

p. 121

Figure 5.12: 3D Data Logger design.

Construction of the Virtual Environment For the construction of the environment we proceeded to create a 3D project in Universal Render Pipeline (URP) mode, throughout this section we will see the different elements that were incorporated into the virtual environment.

Land Creation For the creation of the terrain we based ourselves on the topology presented in the city of Quibd´o - Choc´o, for the creation of this terrain we used internal tools of Unity, which allow us to deploy different shapes, colors and textures to define the shape and atmosphere of the terrain, finally we try to simulate the geography of the Atrato river and the final results are shown in the Figure below.5.13.

Figure 5.13: 3D Design Atrato River - Unity.

Incorporation of Hydrological Station Model Once the model of the hydrological station has been created, the .FBX is implemented in the previously created terrain. In addition, movements associated to the behavior that the station sensors should have are incorporated in order to provide a reliable operation to the real station. Finally, the environment with the implemented model is shown in the Figure below 5.14.

p. 122

Figure 5.14: Incorporation of Hydrological Station in Unity. Incorporation of the Data Analysis Center In this space the graphs and the station control center are incorporated, here different buttons were implemented to vary between the different graphs of each station, on the other hand, cameras were added to be able to visualize the parts of the station interacting with the environment. Finally, Figure ¨refgraphs shows the final result of the control center in which the graphics of the hydrological station are displayed. Figure 5.15: Data Visualization Center.

Interaction with the Parts of the Hydrological Station.

In this section we will be able to take and interact with the parts of the hydrological station, this will allow the user to know the operation of the parts of the station, and will be able to observe how the station is assembled and disassembled. Finally this section can be seen in the Figure 5.16

p. 123

Figure 5.16: Hydrologic Station Assembly and Disassembly.

5.2.2

General aspect of the environment In this section we will review the environment and its most important aspects, in this case we will talk about the assembly and disassembly system of the hydrological station (EH), in addition, we will review the estimation system implemented to analyze the level of the river shown in figure 1.4.

Control Station In this section you will be able to observe the river level estimation and measurement graphs, i.e., the station will capture data and deliver an estimated signal of the future behavior of the river level based on the precipitation in the area and the river flow velocity.

Figure 5.17: Control Station.

Assembly and Disassembly Area In this area you will be able to assemble and disassemble the hydrological station, interact with its parts and learn what each one of them works for.

p. 124

Figure 5.18: Assembly and Disassembly Area. Implementation Zone This section is designed so that you can learn about the context of the implementation of the hydrological station, i.e., you can learn about the place where the project was designed to be implemented, the research group in charge of developing this type of scenarios and the assembly of the hydrological station together with a brief description of its components.

Figure 5.19: Implementation Zone.

5.3

Software registration The software presented in this section has been registered in the ”DIRECCION NACIONAL DE DERECHO DE AUTOR” in Colombia, as follows:

• J.

B.

Renteria-Mena,; F.

Osorio-Arteaga, F.A.

Trujillo-Perdomo, S.

Velarde-Gomez; and E. Giraldo, ”ENTORNO VIRTUAL PARA UN SISTEMA

DE ESTIMACI´ON DEL NIVEL DEL AGUA DE ALERTA TEMPRANA

PARA MITIGAR RIESGOS POR INUNDACIONES EN COMUNIDADES

RIBERE˜NAS”.

DIRECCION NACIONAL DE DERECHO DE AUTOR.

Registro: 13-98-371. Fecha: 29-04-2024.

p. 125

5.4

Summary The chapter highlights how integrating social ownership and knowledge transfer strategies strengthens the Early Warning Systems (EWS) for floods in riverine communities in Choc´o. One of the main innovations is an interactive web platform that allows inhabitants to access real-time alerts, forecasts, risk maps, and action protocols, serving as a bridge between authorities and the community. In addition, an educational virtual environment is introduced that simulates risk scenarios with real data, facilitating learning about the causes and consequences of floods. A key aspect is that these tools were co-created with local communities, respecting their culture, language, and traditional knowledge. This not only favored their acceptance but also their effectiveness and sustainability.

Together, these initiatives transform EWS into more accessible, decentralized, and effective systems, allowing communities to not only receive warnings but also understand them and act promptly, strengthening their resilience to hydro-meteorological events.

5.5

Partial conclusions The development of an accessible interactive web platform has proven to be an effective tool for improving communication between authorities and communities, facilitating access to key information on flood risks and strengthening local emergency response capacity.

Incorporating a virtual educational environment, based on real data and simulations, promotes active understanding of hydrological phenomena and risk, enabling users to interpret warnings and adopt preventive measures promptly.

The tools’ participatory design, which integrated local knowledge, languages, and cultural contexts, was key to their acceptance and adaptation. This collaboration ensures that the solutions are sustainable, useful, and truly relevant to the communities involved.

The combination of technological resources with education and social participation processes allows for the transformation of EWS into more democratic, inclusive, and effective systems, where communities receive warnings and become active agents in risk management.

The proposed integrated approach enhances the autonomy and preparedness of Choc´o riverine communities for hydrometeorological events, helping to reduce vulnerability through empowerment, timely information, and education based on their real environment.

p. 126

Chapter 6 Conclusions and Final Remarks

6.1

Conclusions The water level estimation system for early warning in the main rivers of the Department of Choc´o, using linear, nonlinear and hybrid models, has proven to be an effective tool for risk mitigation in the riparian communities of the Department of Choc´o. In this work, linear models, such as the Autoregressive Model (AR) and the Autoregressive Model with Exogenous Inputs (ARX), and nonlinear models, including NARX, Backpropagation, LSTM and a hybrid nonlinear model NARX-EnKF, have been used Linear models, such as AR and ARX, provided reasonable short-term predictions. The AR model effectively captured internal water level dynamics from its own historical data, while the ARX, by including exogenous inputs such as precipitation and flow, improved predictions when considering relevant external variables. However, both models showed limitations when faced with conditions of high variability and nonlinear behaviors in the system On the other hand, nonlinear models -NARX, Backpropagation, LSTM, and the hybrid NARX-EnKF model—proved to be more effective in capturing the complexities of the Atrato River’s hydrological system. By taking advantage of advanced fitting and feedback techniques, these models significantly outperformed linear models in scenarios with high variability and nonlinear phenomena.

The NARX mode was particularly effective in short-term prediction, combining exogenous inputs with feedback from past outputs.

Backpropagation improved the performance of NARX by adjusting the neural network weights through iterations, allowing this model to better adapt to rapid changes in water level. Thanks to their ability to learn long-term dependencies, LSTM networks excel in medium- and long-term predictions. This is essential for risk management, as it allows anticipating hydrological events further in advance, providing valuable time to make

p. 127

strategic decisions.

One of the main innovations in this work was the implementation of the hybrid NARX-EnKF (Swarm Kalman Filter) model, which combines the predictive power of NARX with the ability of EnKF to handle uncertainty in the system’s state. This hybrid model proved particularly effective in improving water level prediction accuracy by integrating observational and dynamically modeled data. The NARX-EnKF could adjust to abrupt river-level variations, providing more robust predictions than traditional linear and nonlinear models.

The assimilation of multivariate data,including variables such as water level, flow, and precipitation, was critical to improving model performance, particularly in the case of the NARX-EnKF hybrid model, which optimized the integration of multiple sources of information to fit the model in real time.

Comparative analysis between the linear (AR and ARX), nonlinear (NARX, Backpropagation, LSTM), and hybrid nonlinear NARX-EnKF models showed that, although the linear models offer simplicity and computational efficiency, the nonlinear and hybrid models are significantly more effective in predicting complex phenomena. The NARX-EnKF, in particular, provides a robust solution to handle the uncertainty inherent in hydrological predictions, standing out as an innovative and powerful option in early warning systems.

Finally, validating these models with real data from the Atrato River confirmed their applicability in real scenarios. The NARX, Backpropagation, LSTM, and, especially, the hybrid NARX-EnKF and LSTM-EnKF models proved to be crucial tools for implementing early warning systems.

They provide greater accuracy in predictions and improve response capacity to adverse climatic events. This helps reduce the vulnerability of riparian communities, strengthening flood risk management.

6.2

Future works From the findings obtained in this study, several areas of improvement and future research can be identified that could enhance the water level estimation system for early warning in the rivers of the Department of Choc´o. Some possible lines of work are presented below:

• Exploration of more sophisticated hybrid models: Although the NARX-EnKF

hybrid model showed good results, it would be useful to investigate the combination of other hybrid approaches that optimize the capabilities of both linear and nonlinear models. This could include using deep neural networks (Deep Learning) in combination with classical time series techniques and integrating convolutional neural networks (CNN) better to address the complexity and nonlinear patterns in hydrological data.

p. 128

• Expansion and diversification of the data used: The dataset used in this study

could be expanded to increase the accuracy and generalization of the models. Incorporating more hydrological stations in different points of the Atrato River and other nearby basins and integrating data from various sources, such as satellites or climate models, would better capture the’ spatial and temporal variability of hydrological conditions. Also, considering additional factors such as temperature, humidity, or changes in land use could enrich predictions.

• Optimization of nonlinear models:

Although nonlinear models outperformed linear models in this study, it is still possible to improve their performance through finer adjustments in their calibration.

In future work, advanced parameter optimization techniques, such as evolutionary algorithms or machine learning approaches, could be explored to increase the predictive capacity of these models, especially in situations of high variability or extreme events.

• Real-time implementation: A major challenge for using these models is their

real-time operational early warning systems implementation. Future research could optimize the models to operate efficiently with continuous data and provide fast and accurate predictions, even under conditions of high computational load or during intense hydrological events.

• Evaluation of the impact on community management: A crucial aspect to be

addressed in future work is how integrating these models into early warning systems can improve risk management in riverine communities. Investigating how these warnings affect decision-making and disaster response policies, as well as the effectiveness of risk communication, would allow better adaptation of the systems to local needs.

• Research in hybrid models with artificial intelligence: As artificial intelligence

techniques advance, it would be valuable to explore the integration of hybrid approaches that combine neural networks with computational intelligence methods, such as optimization algorithms or reinforcement learning. These approaches could further improve the ability of models to anticipate complex hydrological phenomena and extreme events, such as flash floods. In summary, although the models employed in this work have proven effective for water level estimation in rivers in the Department of Choc´o, there are ample opportunities to improve their performance and applicability.

The use of emerging technologies, improvement in data quality and quantity, and the implementation of real-time operational systems are key areas that could take this early warning system to a higher level, contributing to more effective risk management in the region.

p. 129

6.3

Academic Discussion The results obtained during the development of this thesis have been published in the following journals and conferences:

• J. B. Renteria-Mena, and E. Giraldo, ”Real-time Adaptive Level Control of a

Multivariable Waste Water Treatment Plant,” Engineering Letters, vol. 30, no.2, pp. 444- 452, 2022. ISSN: 1816-0948. [79]

• Jackson B. Renteria-Mena, and Eduardo Giraldo, ”Multivariable AR Data

Assimilation for Water Level, Flow, and Precipitation Data,

”IAENG

International Journal of Computer Science, vol. 50, no. 1, pp. 263- 273, 2023.

ISSN: 1819-9224. [78]

• J. B. Renteria-Mena, E. Giraldo, and Douglas Plaza, ”Multivariable NARX-based

Neural Networks Models for Short-term Water Level Forecasting”, International Conference on Time Series and Forecasting. July 12th-14th, 2023, Gran Canaria, Spain. [85]

• J.

B.

Renteria-Mena,; D.

Plaza, and E.

Giraldo, ”Multivariate Hydrological Modeling Based on Long Short-Term Memory Networks for Water Level Forecasting”.

Information

15,

no.

6:

358,

2024.

https://doi.org/10.3390/info15060358. ISSN: 2078-2489. [93]

• J. B. Renteria-Mena, D. Plaza, and E. Giraldo, ”Comparative Analysis of

Nonlinear Methods for Multivariable Water Level Prediction: The Case Study of the Atrato River”. Journal of Electrical and Computer Engineering, 2024, vol. 2024, no 1, p. 2894031. https://doi.org/10.1155/2024/2894031. ISSN: 2090-0147, 2090-0155. [86]

• J. B. Renteria-Mena and E. Giraldo, ”Predictive Modeling of Water Level

in the San Juan River Using Hybrid Neural Networks Integrated with Kalman Smoothing Methods”.

Information 15, no.

12:

754.

2024.

https://doi.org/10.3390/info15120754. ISSN: 2078-2489. [95]

• J.

B.

Renteria-Mena, E.

Giraldo, and D.

Plaza, ”Water-Level Forecasting Based on an Ensemble Kalman Filter with a NARX Neural Network Model”.

Engineering Proceedings

101,

no.

1:

2.

2025.

https://doi.org/10.3390/engproc2025101002. ISSN: 2673-4591. [94]

p. 130

Bibliography [1] K. G. M. Pacheco, “Con el agua al cuello”: Una historia de batallas perdidas contra el agua y desastres por inundaciones en colombia, 1950-2011,” Agua y territorio= Water and Landscape, no. 22, pp. 77–91, 2023.

[2] H. Ayala Mosquera, “Extractivismo minero y sus efectos sobre el cambio clim´atico en el choc´o biogeogr´afico colombiano.,” Bolet´ın de Antropolog´ıa, vol. 37, no. 64,

2022.

[3] R. G´omez, J. L´opez, and P. Rodr´ıguez, “Impacto de las inundaciones en la infraestructura y econom´ıa local en colombia,” An´alisis de Riesgos Naturales, vol. 22, no. 2, pp. 76–89, 2020.

[4] A. R´ıos and M. P´erez, Respuestas comunitarias a eventos clim´aticos extremos. Editorial Acad´emica, 2020.

[5] M. ´Alvarez and J. Vargas, Gesti´on de riesgos en comunidades ribere˜nas. Editorial Universitaria, 2022.

[6] R. G´omez, Sistemas de alerta temprana para desastres naturales. Editorial Cient´ıfica, 2017.

[7] F. Jaramillo, Modelos predictivos en gesti´on de inundaciones. Instituto Nacional de Meteorolog´ıa, 2023.

[8] A. Castro, L. G´omez, and P. Rodr´ıguez, Tecnolog´ıas de monitoreo hidrol´ogico. Universidad de los Andes, 2019.

[9] J. M´endez and E. Su´arez, “Prevenci´on y mitigaci´on de riesgos de inundaci´on,” Revista de Ciencias Ambientales, vol. 28, no. 1, pp. 12–25, 2020. [10] D. Rojas and F. Mendoza, “Adaptaci´on al cambio clim´atico en comunidades ribere˜nas,” Revista de Gesti´on de Riesgos, vol. 30, no. 3, pp. 78–91, 2021. [11] C. P´erez, Planificaci´on y respuesta ante emergencias clim´aticas. Editorial Universitaria, 2019.

[12] R. L. Bras and I. Rodriguez-Iturbe, Random functions and hydrology. Courier Corporation, 1993.

p. 131

[13] K. I. Mart´ınez Calder´on et al., “Cr´onicas de viaje al choc´o,” 2018. [14] E. E. Cossio Roma˜na, “Fortalecimiento de la infraestructura en salud e integraci´on social en la zona rural y urbana de quibd´o, choc´o,” 2020. [15] S. P. Nore˜na Garc´ıa, “Responsabilidad del estado por el desplazamiento de comunidades por causas asociadas al cambio clim´atico en los departamentos de risaralda y choc´o durante la ola invernal de 2010-2011.,” 2019. [16] M. Institute of Hydrology and E. Studies, Consulta descarga datos Meterol´ogicos

- 2021. Bogot´a:: IDEAM dhime,, 2021.

[17] W. F. R. Mi˜nope, P. V. R. V. Liz´arraga, S. P. M. P´erez, V. T. Monteza, and H. I. M. Cabrera, “Modelamiento de procesos hidrol´ogicos aplicando t´ecnicas de inteligencia artificial: una revisi´on sistem´atica de la literatura,” ITECKNE: Innovaci´on e Investigaci´on en Ingenier´ıa, vol. 19, no. 1, p. 6, 2022. [18] R. W. Farebrother, Linear least squares computations. Routledge, 2018. [19] G. Pillonetto, T. Chen, A. Chiuso, G. D. Nicolao, and L. Ljung, “Regularized system identification: Learning dynamic models from data,” IEEE Control Systems Magazine, vol. 42, no. 1, pp. 24–48, 2022.

[20] G. J. McLachlan and T. Krishnan, The EM algorithm and extensions. John Wiley & Sons, 2007.

[21] S. Theodoridis, Machine learning:

a Bayesian and optimization perspective.

Academic press, 2015.

[22] D. Simon, Optimal state estimation:

Kalman, H infinity, and nonlinear approaches. John Wiley & Sons, 2006.

[23] E. Palchevsky, V. Antonov, R. Enikeev, and T. Breikin, “A system based on an artificial neural network of the second generation for decision support in especially significant situations,” Journal of Hydrology, vol. 616, p. 128844, 2023. [24] Z. LV, J. Zuo, and D. Rodriguez, “Predicting of runoff using an optimized swat-ann: A case study,” Journal of Hydrology: Regional Studies, vol. 29, p. 100688, 2020. [25] Y. Abou Rjeily, O. Abbas, M. Sadek, I. Shahrour, and F. Hage Chehade, “Flood forecasting within urban drainage systems using narx neural network,” Water Science and Technology, vol. 76, no. 9, pp. 2401–2412, 2017. [26] X.-H. Le, H. V. Ho, G. Lee, and S. Jung, “Application of long short-term memory (lstm) neural network for flood forecasting,” Water, vol. 11, no. 7, p. 1387, 2019. [27] P. Mu˜noz, J. Orellana-Alvear, J. Bendix, J. Feyen, and R. C´elleri, “Flood early warning systems using machine learning techniques: The case of the tomebamba catchment at the southern andes of ecuador,” Hydrology, vol. 8, no. 4, p. 183, 2021.

p. 132

[28] S. Bande and V. V. Shete, “Smart flood disaster prediction system using iot & neural networks,” in 2017 International Conference On Smart Technologies For Smart Nation (SmartTechCon), pp. 189–194, Ieee, 2017.

[29] A. Jabbari and D.-H. Bae, “Application of artificial neural networks for accuracy enhancements of real-time flood forecasting in the imjin basin,” Water, vol. 10, no. 11, p. 1626, 2018.

[30] R. Tabbussum and A. Q. Dar, “Comparative analysis of neural network training algorithms for the flood forecast modelling of an alluvial himalayan river,” Journal of Flood Risk Management, vol. 13, no. 4, p. e12656, 2020. [31] F. Y. Dtissibe, A. A. A. Ari, C. Titouna, O. Thiare, and A. M. Gueroui, “Flood forecasting based on an artificial neural network scheme,” Natural Hazards, vol. 104, pp. 1211–1237, 2020.

[32] S. I. Abdullahi, M. H. Habaebi, and N. A. Malik, “Flood disaster warning system on the go,” in 2018 7th International Conference on Computer and Communication Engineering (ICCCE), pp. 258–263, IEEE, 2018.

[33] N. Kimura, I. Yoshinaga, K. Sekijima, I. Azechi, and D. Baba, “Convolutional neural network coupled with a transfer-learning approach for time-series flood predictions,” Water, vol. 12, no. 1, p. 96, 2019.

[34] M. Moishin, R. Deo, R. Prasad, and N. Raj, “Designing deep-based learning flood forecast model with convlstm hybrid algorithm,” IEEE Access, vol. 9, pp. 43364–43377, 2021.

[35] U. T. Khan, J. He, and C. Valeo, “River flood prediction using fuzzy neural networks: an investigation on automated network architecture,” Water Science and Technology, vol. 2017, no. 1, pp. 238–247, 2018.

[36] R. Tabbussum and A. Q. Dar, “Performance evaluation of artificial intelligence paradigms—artificial neural networks, fuzzy logic, and adaptive neuro-fuzzy inference system for flood prediction,” Environmental Science and Pollution Research, vol. 28, no. 20, pp. 25265–25282, 2021.

[37] S. Sankaranarayanan, M. Prabhakar, S. Satish, P. Jain, A. Ramprasad, and A. Krishnan, “Flood prediction based on weather parameters using deep learning,” Journal of Water and Climate Change, vol. 11, no. 4, pp. 1766–1783, 2020. [38] P. Rodgers, Grey System Theory and Applications. Springer, 2000. [39] W. Yang and S. Liu, “Grey models in time series prediction: Theory and applications,” Expert Systems with Applications, vol. 94, pp. 302–317, 2018. [40] Z. Wan, W. Yu, J. Xu, X. Liu, and G. Zhang, “A review on flood forecasting technology based on deep learning models,” Journal of Hydrology, vol. 570, pp. 330–345, 2019.

p. 133

[41] M. Raissi, P. Perdikaris, and G. E. Karniadakis, “Physics-informed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations,” Journal of Computational Physics, vol. ??, no. ??, p. ??–??, 2019.

[42] J. Willard, X. Jia, S. Xu, M. Steinbach, and V. Kumar, “Integrating physics-based modeling with machine learning: A survey,” arXiv preprint arXiv:2003.04919v4, pp. 1–34, 2020.

[43] K. Beven, Rainfall-runoff modelling: the primer. John Wiley & Sons, 2011. [44] K. Beven, “How to make advances in hydrological modelling,” Hydrology Research, vol. 50, no. 6, pp. 1481–1494, 2019.

[45] K. Beven and P. Young, “A guide to good practice in modeling semantics for authors and referees,” Water Resources Research, vol. 49, no. 8, pp. 5092–5098,

2013.

[46] D. N. Moriasi, J. G. Arnold, M. W. Van Liew, R. L. Bingner, R. D. Harmel, and T. L. Veith, “Model evaluation guidelines for systematic quantification of accuracy in watershed simulations,” Transactions of the ASABE, vol. 50, no. 3, pp. 885–900,

2007.

[47] B. Hadid, E. Duviella, and S. Lecoeuche, “Data-driven modeling for river flood forecasting based on a piecewise linear arx system identification,” Journal of Process Control, vol. 86, pp. 44–56, 2020.

[48] S. Nevo, E. Morin, A. Gerzi Rosenthal, A. Metzger, C. Barshai, D. Weitzner, D. Voloshin, F. Kratzert, G. Elidan, G. Dror, et al., “Flood forecasting with machine learning models in an operational framework,” Hydrology and Earth System Sciences, vol. 26, no. 15, pp. 4013–4032, 2022.

[49] F. Unduche, H. Tolossa, D. Senbeta, and E. Zhu, “Evaluation of four hydrological models for operational flood forecasting in a canadian prairie watershed,” Hydrological Sciences Journal, vol. 63, no. 8, pp. 1133–1149, 2018. [50] M. Vel´asquez-Restrepo and G. Poveda, “Estimaci´on del balance h´ıdrico de la regi´on pac´ıfica colombiana,” Dyna, vol. 86, no. 208, pp. 297–306, 2019. [51] J. Chen, C. Li, F. P. Brissette, H. Chen, M. Wang, and G. R. Essou, “Impacts of correcting the inter-variable correlation of climate model outputs on hydrological modeling,” Journal of Hydrology, vol. 560, pp. 326–341, 2018. [52] K.

Jafarzadegan, P.

Abbaszadeh, and H.

Moradkhani, “Sequential data assimilation for real-time probabilistic flood inundation mapping,” Hydrology and Earth System Sciences, vol. 25, no. 9, pp. 4995–5011, 2021. [53] M. Acosta-Coll, F. Ballester-Merelo, M. Martinez-Peir´o, and E. De la Hoz-Franco, “Real-time early warning system design for pluvial flash floods—a review,” Sensors, vol. 18, no. 7, p. 2255, 2018.

p. 134

[54] Y. Li, S. Grimaldi, J. P. Walker, and V. R. Pauwels, “Application of remote sensing data to constrain operational rainfall-driven flood forecasting: A review,” Remote Sensing, vol. 8, no. 6, p. 456, 2016.

[55] E. A. Basha, S. Ravela, and D. Rus, “Model-based monitoring for early warning flood detection,” in Proceedings of the 6th ACM conference on Embedded network sensor systems, pp. 295–308, 2008.

[56] P. Khac-Tien Nguyen and L. Hock-Chye Chua, “The data-driven approach as an operational real-time flood forecasting model,” Hydrological Processes, vol. 26, no. 19, pp. 2878–2893, 2012.

[57] N. Ullah and P. Choudhury, “Flood flow modeling in a river system using adaptive neuro-fuzzy inference system,” Environ Manag Sustain Develop, vol. 2, no. 2, pp. 54–68, 2013.

[58] E. J. Plate, “Early warning and flood forecasting for large rivers with the lower mekong as example,” Journal of Hydro-environment Research, vol. 1, no. 2, pp. 80–94, 2007.

[59] J.-D. L´opez-Garc´ıa, Y. Carvajal-Escobar, and A.-M. Enciso-Arango, “Sistemas de alerta temprana con enfoque participativo: un desaf´ıo para la gesti´on del riesgo en colombia,” Luna azul, no. 44, pp. 231–246, 2017.

[60] A. McNally, K. Arsenault, S. Kumar, S. Shukla, P. Peterson, S. Wang, C. Funk,

C. D. Peters-Lidard, and J. P. Verdin, “A land data assimilation system for

sub-saharan africa food and water security applications,” Scientific Data, vol. 4, no. 1, 2017.

[61] X. Li, Z. Zhao, and F. Liu, “Latent variable iterative learning model predictive control for multivariable control of batch processes,” Journal of Process Control, vol. 94, pp. 1–11, 2020.

[62] J. Ding, Z. Cao, J. Chen, and G. Jiang, “Weighted parameter estimation for hammerstein nonlinear arx systems,” Circuits, Systems, and Signal Processing, vol. 39, no. 4, pp. 2178–2192, 2020.

[63] F. A. Ruslan, K. Haron, R. Adnan, et al., “Multiple input single output (miso) arx and armax model of flood prediction system: Case study pahang,” in 2017 IEEE 13th International Colloquium on Signal Processing & its Applications (CSPA), pp. 179–184, IEEE, 2017.

[64] J. B. Renteria-Mena and E. Giraldo, “Real-time adaptive level control of a multivariable waste water treatment plant,” Engineering Letters, vol. 30, no. 2,

2022.

[65] G. E. Karniadakis, I. G. Kevrekidis, L. Lu, P. Perdikaris, S. Wang, and L. Yang, “Physics-informed machine learning,” Nature Reviews Physics, vol. 3, no. 6, pp. 422–440, 2021.

p. 135

[66] D. P. Solomatine and D. B. Dulal, “Data-driven modelling: some past experiences and new approaches,” Journal of Hydroinformatics, vol. 6, no. 3, pp. 207–214,

2004.

[67] V. P. Singh and D. A. Woolhiser, Hydrological modeling: theory and practice. Springer, 2017.

[68] D. Solomatine and K. Dulal, “Model trees as an alternative to neural networks in rainfall-runoff modelling,” in Proceedings of the 6th International Conference on Hydroinformatics, pp. 2023–2028, World Scientific, 2004.

[69] D. N. Moriasi, J. G. Arnold, M. W. Van Liew, R. L. Bingner, R. D. Harmel, and T. L. Veith, “Model evaluation guidelines for systematic quantification of accuracy in watershed simulations,” Transactions of the ASABE, vol. 50, no. 3, pp. 885–900,

2007.

[70] S. Dingman, Physical Hydrology. Waveland Press, 3rd ed., 2015. [71] W. Brutsaert, Hydrology: An Introduction. Cambridge University Press, 2005. [72] K. Beven, Rainfall-Runoff Modelling: The Primer. John Wiley & Sons, 2nd ed.,

2012.

[73] V. T. Chow, D. R. Maidment, and L. W. Mays, Applied Hydrology. New York: McGraw-Hill, 1988.

[74] K. Beven, Rainfall-Runoff Modelling: The Primer. Wiley-Blackwell, 2nd ed., 2012. [75] V. P. Singh, Elementary Hydrology. New Jersey: Prentice Hall, 1992. [76] P. C. Hansen, Rank-deficient and discrete ill-posed problems: numerical aspects of linear inversion. SIAM, 1998.

[77] P. C. Hansen, J. G. Nagy, and D. P. O’leary, Deblurring images: matrices, spectra, and filtering. SIAM, 2006.

[78] J. B. Renteria-Mena and E. Giraldo, “Multivariable ar data assimilation for level, flow and of precipitation data,” IAENG International Journal of Computer Science (IJCS), vol. 50, no. 1, pp. 263–273, 2023.

[79] J. B. Renteria-Mena and E. Giraldo, “Real-time adaptive level control of a multivariable waste water treatment plant,” Engineering Letters, vol. 30, no. 2, pp. 444–452, 2022.

[80] J. B. Renteria-Mena, D. Plaza, and E. Giraldo, “Multivariable narx based neural networks models for short-term water level forecasting,” Engineering Proceedings, vol. 39, no. 1, p. 60, 2023.

[81] D. E. Rumelhart, G. E. Hinton, and R. J. Williams, “Learning representations by back-propagating errors,” nature, vol. 323, no. 6088, pp. 533–536, 1986.

p. 136

[82] F.-C. Chen, “Back-propagation neural networks for nonlinear self-tuning adaptive control,” IEEE control systems Magazine, vol. 10, no. 3, pp. 44–48, 1990. [83] A. T. Goh, “Back-propagation neural networks for modeling complex systems,” Artificial intelligence in engineering, vol. 9, no. 3, pp. 143–151, 1995. [84] M. Salarijazi, I. Ahmadianfar, and Z. M. Yaseen, “Prediction enhancement for surface water sodium adsorption ratio using limited inputs: Implementation of hybridized stacked ensemble model with feature selection algorithm,” Physics and Chemistry of the Earth, Parts A/B/C, vol. 134, p. 103561, 2024. [85] J. B. Renteria-Mena, D. Plaza, and E. Giraldo, “Multivariable narx based neural networks models for short-term water level forecasting,” Engineering Proceedings, vol. 39, no. 1, 2023.

[86] J. B. Renteria-Mena, D. Plaza, and E. Giraldo, “Comparative analysis of nonlinear methods for multivariable water level prediction: The case study of the atrato river,” Journal of Electrical and Computer Engineering, vol. 2024, no. 1,

p. 2894031, 2024.

[87] V. Atashi, H. T. Gorji, S. M. Shahabi, R. Kardan, and Y. H. Lim, “Water level forecasting using deep learning time-series analysis: A case study of red river of the north,” Water, vol. 14, no. 12, p. 1971, 2022.

[88] R. D. Pinzon Morales and Y. Hirata, “Bi-hemispherical neuronal network of the cerebellum with realistic climbing fiber reproduces asymmetrical motor learning during robot control,” Frontiers in Neural Circuits, vol. 8, 2014. [89] A.

Graves and J.

Schmidhuber, “Framewise phoneme classification with bidirectional lstm and other neural network architectures,” Neural networks, vol. 18, no. 5-6, pp. 602–610, 2005.

[90] S.

Hochreiter and J.

Schmidhuber, “Long short-term memory,” Neural computation, vol. 9, no. 8, pp. 1735–1780, 1997.

[91] M. Cho, C. Kim, K. Jung, and H. Jung, “Water level prediction model applying a long short-term memory (lstm)–gated recurrent unit (gru) method for flood prediction,” Water, vol. 14, no. 14, p. 2221, 2022.

[92] S. Gillijns, O. B. Mendoza, J. Chandrasekar, B. L. R. D. Moor, D. S. Bernstein, and A. Ridley, “What is the ensemble kalman filter and how well does it work?,” in 2006 American Control Conference, pp. 6 pp.–, June 2006. [93] J. B. Renteria-Mena, D. Plaza, and E. Giraldo, “Multivariate hydrological modeling based on long short-term memory networks for water level forecasting,” Information, vol. 15, no. 6, p. 358, 2024.

[94] J. B. Renteria-Mena, D. Plaza, and E. Giraldo, “Water-level forecasting based on an ensemble kalman filter with a narx neural network model,” in Engineering Proceedings, vol. 101, p. 2, 2025.

p. 137

[95] J. B. Renteria-Mena and E. Giraldo, “Predictive modeling of water level in the san juan river using hybrid neural networks integrated with kalman smoothing methods,” Information, vol. 15, no. 12, p. 754, 2024. Submission received: 22 Oct 2024; Revised: 20 Nov 2024; Accepted: 22 Nov 2024; Published: 26 Nov 2024.

Cita: Renteria Mena, Jackson Berney (2025), Sistema de estimación del nivel del agua de alerta temprana para mitigar riesgos por inundaciones en comunidades ribereñas del departamento del Chocó, Universidad Tecnológica de Pereira, p. N. https://hdl.handle.net/11059/16492