Papers
CITEEC is a leader in civil and building engineering research, focusing on innovative solutions for sustainable infrastructure and advanced materials.
Pioneering research for sustainable development
CITEEC carries out pioneering research in civil engineering, materials, sustainability and infrastructure. Thanks to multidisciplinary projects and international collaborations, scientific papers are being published that address key challenges facing the sector. This section presents the main papers and results produced by our research groups.
You can access a wide range of scientific publications, which can be filtered by subject, author, year and more. Use the filters provided to find relevant literature and learn more about each of our contributions to the advancement of knowledge.
2026
Montero-Lamas, Yaiza; Fernández-Sánchez, Alberto; Gestal-Pose, Marcos; Novales-Ordax, Margarita; Orro-Arcay, Alfonso
Bus travel time prediction in mixed traffic: a multisource big data approach with multiple modelling strategies Journal Article
In: European Transport Research Review, vol. 18, no. 1, pp. 57, 2026, ISSN: 1866-8887.
Abstract | Links | BibTeX | Tags: Bluetooth sensors, Bus travel time prediction, Inductive loops, Linear Regression, machine learning, Transit planning
@article{montero-lamas_bus_2026,
title = {Bus travel time prediction in mixed traffic: a multisource big data approach with multiple modelling strategies},
author = {Yaiza Montero-Lamas and Alberto Fernández-Sánchez and Marcos Gestal-Pose and Margarita Novales-Ordax and Alfonso Orro-Arcay},
url = {https://doi.org/10.1186/s12544-026-00815-3},
doi = {10.1186/s12544-026-00815-3},
issn = {1866-8887},
year = {2026},
date = {2026-07-01},
urldate = {2026-08-17},
journal = {European Transport Research Review},
volume = {18},
number = {1},
pages = {57},
abstract = {This study examines the prediction of bus travel times within urban corridors, using an extensive dataset from transit management databases and on-street sensors. The analysis focuses on a range of variables, including corridor configuration, general traffic conditions, and intrinsic bus transit factors. Employing Multiple Linear Regression (MLR), Artificial Neural Networks (ANN), Support Vector Regression (SVR), and Random Forest (RF), bus travel times are modelled across four distinct urban corridors in A Coruña, Spain, considering dynamic variables like traffic flow rate, general travel time, and average stream patronage per bus stop, with patronage showing the strongest influence on bus travel time in three of the four corridors, similar to the sum of all general traffic variables. Furthermore, these models are applied to a joint dataset encompassing all corridors, incorporating static variables such as bus stops per kilometre, and, for the first time in the field, the percentage of corridors with adjacent parking and percentage with one lane, both considered as an interaction. This allows predictions of bus travel time changes due to corridor modifications and travel times for new bus routes in unserved areas. Findings reveal that, for our dataset, none of the three ML approaches has consistently proven to be preferable in bus travel time predictions, while MLR provides competitive results, balancing accuracy and interpretability despite its flexibility constraints. The study underscores the importance of selecting models based on data range and transit context, advocating for simplicity in constrained scenarios. This developed methodology provides a valuable planning tool for transport agencies, adaptable to other urban contexts, and highlights the benefits of optimising corridor configurations to enhance bus travel time performance.},
keywords = {Bluetooth sensors, Bus travel time prediction, Inductive loops, Linear Regression, machine learning, Transit planning},
pubstate = {published},
tppubtype = {article}
}
Fernández-Sánchez, Alberto; Gestal-Pose, Marcos; Bolón-Canedo, Verónica; Dorado-De-La-Calle, Julián; Pazos-Sierra, Alejandro
Comparison of data set sample selection algorithms for data science: a systematic review Journal Article
In: Neural Computing and Applications, vol. 38, no. 10, pp. 408, 2026, ISSN: 1433-3058.
Abstract | Links | BibTeX | Tags: Data complexity, Dataset filtering, Instance hardness, Instance selection, machine learning
@article{fernandez-sanchez_comparison_2026,
title = {Comparison of data set sample selection algorithms for data science: a systematic review},
author = {Alberto Fernández-Sánchez and Marcos Gestal-Pose and Verónica Bolón-Canedo and Julián Dorado-De-La-Calle and Alejandro Pazos-Sierra},
url = {https://doi.org/10.1007/s00521-026-12124-w},
doi = {10.1007/s00521-026-12124-w},
issn = {1433-3058},
year = {2026},
date = {2026-05-01},
urldate = {2026-08-18},
journal = {Neural Computing and Applications},
volume = {38},
number = {10},
pages = {408},
abstract = {In the era of big data, selecting representative samples has become essential to mitigate overfitting, noise, and high computational cost in machine learning. This study systematically reviews the evolution of instance selection (IS) methods, highlighting the growing importance of instance hardness (IH) as a guiding criterion to improve training efficiency and model robustness. Through a comprehensive search in Scopus and Web of Science, fifty-five studies were identified and analyzed following strict inclusion and exclusion criteria. The reviewed works were classified according to their underlying rationale–error-based, geometric, heuristic, or explainability-driven–revealing that IH principles intersect these categories as a transversal perspective on data quality. Most studies focus on enhancing predictive accuracy (56%) and computational efficiency (36%), while bias reduction and privacy preservation remain secondary. Reported outcomes show significant dataset reductions (up to 97%) with minimal accuracy loss and, in some cases, notable performance gains (+32% accuracy, +67% improvement in MSE). Despite these advances, explicit references to IH are rare, though many methods implicitly rely on related metrics such as misclassification frequency or decision-boundary proximity. Overall, IS is gaining relevance across domains such as cybersecurity, biomedicine, and computer vision, yet the field still lacks standardized methodologies and benchmarking frameworks, underscoring the need for unified, IH-informed strategies for robust and generalizable instance selection.},
keywords = {Data complexity, Dataset filtering, Instance hardness, Instance selection, machine learning},
pubstate = {published},
tppubtype = {article}
}
Carro-Fidalgo, Humberto; Figuero-Pérez, Andrés; Sande-González-Cela, José; Peña-González, Enrique; Alvarellos-González, Alberto
Artificial intelligence to predict critical events in port operations Journal Article
In: Journal of Marine Science and Technology, 2026, ISSN: 1437-8213.
Abstract | Links | BibTeX | Tags: Downtimes, Field campaign, machine learning, Port operability
@article{carro_artificial_2026,
title = {Artificial intelligence to predict critical events in port operations},
author = {Humberto Carro-Fidalgo and Andrés Figuero-Pérez and José Sande-González-Cela and Enrique Peña-González and Alberto Alvarellos-González},
url = {https://doi.org/10.1007/s00773-026-01109-y},
doi = {10.1007/s00773-026-01109-y},
issn = {1437-8213},
year = {2026},
date = {2026-02-01},
urldate = {2026-02-01},
journal = {Journal of Marine Science and Technology},
abstract = {Port downtime is one of the most important economic and safety issues. The objective of this paper is the design of a predictive tool based on machine learning, capable of identifying downtimes. We have compared two approaches: one that uses the moored ship motions and follows a more theoretical approach. It is based on the physics of the problem, but has the difficulty of using movements that are obtained in a costly and complex way. The other option directly estimates the probability of downtime from ocean-meteorological data and ignores the moored ship motions. This is a simplification, as it omits movement constraints, but it is also more realistic, as it is very complex to obtain data on the state of moorings or estimated tension. The dataset uses ocean-meteorological forecast data and downtimes recorded during the port operations of 799 ships obtained during 8 years. We applied regression models to obtain the variables related to infragravity wave and moored ship motions. We performed a multiclass random forest classification by adjusting the weighting of the dataset to identify the most promising approach. We evaluated techniques such as random forest or gradient boosting machine, selecting the GBM for better validation performance, with a log-loss of 0.0103. We obtained F2 and F1 scores of 0.971 and 0.941 respectively, with which the model correctly classified 98% of the anchorages and 88% of the berthing disruptions. Error analysis by sea state and by stay record indicates that some of them are due to the subjectivity inherent in the anchoring phenomenon. The tool offers an alternative to the traditional analysis of significant movements through an easily exportable methodology.},
keywords = {Downtimes, Field campaign, machine learning, Port operability},
pubstate = {published},
tppubtype = {article}
}
Rodriguez, Jose L.; Ascencio, Estefania; Ageitos, Jose M.; Feijoo, Lucia; He, Shan; Daghighi, Amirreza; Casanola-Martin, Gerardo M.; Munteanu, Cristian R.; Rodríguez-Yáñez, Santiago; Pazos, Alejandro; Bediaga, Harbil; Ordiales, Brenda Durand; Esteban, Raquel; Heres, Ana-Maria; Yuste, Jorge Curiel; Epelde, Lur; Rasulev, Bakhtiyor; Gonzalez, Tomas; Arrasate, Sonia; González-Diaz, Humberto
In: Advanced Intelligent Systems, vol. 8, no. 4, pp. e202500587, 2026, ISSN: 2640-4567.
Abstract | Links | BibTeX | Tags: DNA sequences, machine learning, Perturbation Theory, Shannon Information Theory
@article{rodriguez_polymerase_2026,
title = {Polymerase Chain Reaction. Perturbation Theory and Machine Learning Artificial Intelligence-Experimental Microbiome Analysis: Applications to Ancient DNA and Tree Soil Metagenomics Cases of Study},
author = {Jose L. Rodriguez and Estefania Ascencio and Jose M. Ageitos and Lucia Feijoo and Shan He and Amirreza Daghighi and Gerardo M. Casanola-Martin and Cristian R. Munteanu and Santiago Rodríguez-Yáñez and Alejandro Pazos and Harbil Bediaga and Brenda Durand Ordiales and Raquel Esteban and Ana-Maria Heres and Jorge Curiel Yuste and Lur Epelde and Bakhtiyor Rasulev and Tomas Gonzalez and Sonia Arrasate and Humberto González-Diaz},
url = {https://onlinelibrary.wiley.com/doi/abs/10.1002/aisy.202500587},
doi = {10.1002/aisy.202500587},
issn = {2640-4567},
year = {2026},
date = {2026-01-01},
urldate = {2026-08-18},
journal = {Advanced Intelligent Systems},
volume = {8},
number = {4},
pages = {e202500587},
abstract = {The exponential growth of genomic data has created a pressing need for methods capable of interpreting complex biological information, especially in fields like paleogenomics and environmental metagenomics. Ancient DNA (aDNA) analysis faces challenges such as degradation and contamination, while soil metagenomic DNA (soDNA) analysis is hindered by microbial diversity and incomplete reference databases. To address these limitations, this study proposes the polymerase chain reaction (PCR).Perturbation Theory and Machine Learning (PTML) methodology, which integrates machine learning with Perturbation Theory to analyze genetic sequences without the need for alignment. Two models are developed: the first classifies bacterial aDNA sequences extracted from Miocene amber; the second predicts tree health using microbial gene abundance in forest soils. Both rely on entropy-based descriptors (θk) and structural differences (Δθk) between query and reference sequences, which serve as perturbation operators for supervised learning algorithms. This approach allows the detection of meaningful patterns even without complete genomic references. The aDNA model achieves 99.65% sensitivity and 99.81% specificity, while the soDNA model reaches 98.85% sensitivity and 92.56% specificity. These results confirm the robustness and applicability of PCR.PTML in diverse genomic contexts, presenting it as a valuable tool for ancient DNA classification and environmental metagenomics analysis.},
keywords = {DNA sequences, machine learning, Perturbation Theory, Shannon Information Theory},
pubstate = {published},
tppubtype = {article}
}
2025
Dominguez-Gortaire, Jose; Ruiz, Alejandra; Porto-Pazos, Ana Belén; Rodríguez-Yáñez, Santiago; Cedrón, Francisco
Alzheimer’s Disease: Exploring Pathophysiological Hypotheses and the Role of Machine Learning in Drug Discovery Journal Article
In: International Journal of Molecular Sciences, vol. 26, no. 3, pp. 1004, 2025, ISSN: 1422-0067.
Abstract | Links | BibTeX | Tags: AI applications, machine learning, mitochondrial dysfunctions, molecular docking, neuroinflammation, pathophysiology, therapeutic target, virtual screening
@article{dominguez-gortaire_alzheimers_2025,
title = {Alzheimer’s Disease: Exploring Pathophysiological Hypotheses and the Role of Machine Learning in Drug Discovery},
author = {Jose Dominguez-Gortaire and Alejandra Ruiz and Ana Belén Porto-Pazos and Santiago Rodríguez-Yáñez and Francisco Cedrón},
url = {https://www.mdpi.com/1422-0067/26/3/1004},
doi = {10.3390/ijms26031004},
issn = {1422-0067},
year = {2025},
date = {2025-01-01},
urldate = {2025-01-01},
journal = {International Journal of Molecular Sciences},
volume = {26},
number = {3},
pages = {1004},
publisher = {Multidisciplinary Digital Publishing Institute},
abstract = {Alzheimer’s disease (AD) is a major neurodegenerative dementia, with its complex pathophysiology challenging current treatments. Recent advancements have shifted the focus from the traditionally dominant amyloid hypothesis toward a multifactorial understanding of the disease. Emerging evidence suggests that while amyloid-beta (Aβ) accumulation is central to AD, it may not be the primary driver but rather part of a broader pathogenic process. Novel hypotheses have been proposed, including the role of tau protein abnormalities, mitochondrial dysfunction, and chronic neuroinflammation. Additionally, the gut–brain axis and epigenetic modifications have gained attention as potential contributors to AD progression. The limitations of existing therapies underscore the need for innovative strategies. This study explores the integration of machine learning (ML) in drug discovery to accelerate the identification of novel targets and drug candidates. ML offers the ability to navigate AD’s complexity, enabling rapid analysis of extensive datasets and optimizing clinical trial design. The synergy between these themes presents a promising future for more effective AD treatments.},
keywords = {AI applications, machine learning, mitochondrial dysfunctions, molecular docking, neuroinflammation, pathophysiology, therapeutic target, virtual screening},
pubstate = {published},
tppubtype = {article}
}
2024
Carro-Fidalgo, Humberto; Figuero-Pérez, Andrés; Sande-González-Cela, José; Alvarellos-González, Alberto; Costas, Raquel; Peña-González, Enrique
Estimation of moored ship motions using a combination of machine learning techniques Journal Article
In: Applied Ocean Research, vol. 153, pp. 104298, 2024, ISSN: 0141-1187.
Abstract | Links | BibTeX | Tags: machine learning, Moored ship motions, Ship movement prediction, Stacking
@article{carro_estimation_2024,
title = {Estimation of moored ship motions using a combination of machine learning techniques},
author = {Humberto Carro-Fidalgo and Andrés Figuero-Pérez and José Sande-González-Cela and Alberto Alvarellos-González and Raquel Costas and Enrique Peña-González},
url = {https://www.sciencedirect.com/science/article/pii/S014111872400419X},
doi = {10.1016/j.apor.2024.104298},
issn = {0141-1187},
year = {2024},
date = {2024-12-01},
urldate = {2026-07-28},
journal = {Applied Ocean Research},
volume = {153},
pages = {104298},
abstract = {The moored ship motions can cause problems for the efficiency of the operation, and for the people and equipment involved. Therefore, being able to predict movements and anticipate possible risk situations is of great interest to operators and the port community. This work presents a methodology applying different machine learning techniques that has allowed positive results to be obtained for this objective, with particular emphasis on the highest values (outliers), which are usually associated with problematic situations. The field campaigns carried out allowed 77 different vessels to be monitored in the outer port of A Coruña (Spain). The techniques used were gradient boosting (GBM), a neural network (DNN), a quantile regression (qReg) and several models generated by stacking (GBM*). The results indicate a lower root mean square error (RMSE) with the use of the latter technique (the validation on the swell is 0.13 m, while the DNN is twice as high), and a better performance on most motions in the outlier subset than those obtained with the individual models (the validation on the outlier subset for the pitch gives an RMSE of 0.12° and 0.2 for the GBM). Finally, the results show that this methodology can be extrapolated to other ports.},
keywords = {machine learning, Moored ship motions, Ship movement prediction, Stacking},
pubstate = {published},
tppubtype = {article}
}
Carro-Fidalgo, Humberto; Sande-González-Cela, José; Figuero-Pérez, Andrés; Alvarellos-González, Alberto; Peña-González, Enrique; Rabuñal-Dopico, Juan R.; Guerra, Andrés; Pérez, Juan Diego
Machine learning tool for wave overtopping prediction based on the safety-operability ratio Journal Article
In: Ocean Engineering, vol. 312, pp. 119006, 2024, ISSN: 0029-8018.
Abstract | Links | BibTeX | Tags: Field campaign, machine learning, Random forest, Wave overtopping
@article{carro_machine_2024,
title = {Machine learning tool for wave overtopping prediction based on the safety-operability ratio},
author = {Humberto Carro-Fidalgo and José Sande-González-Cela and Andrés Figuero-Pérez and Alberto Alvarellos-González and Enrique Peña-González and Juan R. Rabuñal-Dopico and Andrés Guerra and Juan Diego Pérez},
url = {https://www.sciencedirect.com/science/article/pii/S0029801824023448},
doi = {10.1016/j.oceaneng.2024.119006},
issn = {0029-8018},
year = {2024},
date = {2024-11-01},
urldate = {2026-04-28},
journal = {Ocean Engineering},
volume = {312},
pages = {119006},
abstract = {The phenomenon of wave overtopping affects both the safety of people and port facilities, as well as the operability of a port. This work aims to develop a predictive tool based on artificial intelligence to predict wave overtopping over a breakwater, and to reach a compromise between safeguarding people and port efficiency. The data used came from an extensive field campaign lasting 6.5 years. During this research, the overtopping events were recorded and integrated with the forecast data available from the Port. The model used is the random forest, which was applied with the metrics AUC-PR and F2 score to a weighted dataset. The validation results show a recall for alignments of 0.91 and 0.89, while in the test it reaches 0.94 and 0.95. The dangerousness of the wave overtopping was also analyzed, reducing the dangerous false negatives to 2 and to 14 in the two alignments studied. Operability per alignment was unnecessarily affected during validation by 4 and 10 days over 26.5 months, and in the test by 4 and 3 days over 6.5 months. The results show a proven tool with real data, capable of predicting wave overtopping, and a methodology that can be exported to other ports.},
keywords = {Field campaign, machine learning, Random forest, Wave overtopping},
pubstate = {published},
tppubtype = {article}
}
Nogueira-Garea, Xesús; Fernández-Fidalgo, Javier; Ramos-García, Lucía; Couceiro-Aguiar, Iván; Ramírez-Palacios, Luis
Machine learning-based WENO5 scheme Journal Article
In: Computers & Mathematics with Applications, vol. 168, pp. 84–99, 2024, ISSN: 0898-1221.
Abstract | Links | BibTeX | Tags: Euler equations, Finite difference, machine learning, Neural networks, WENO
@article{nogueira_machine_2024,
title = {Machine learning-based WENO5 scheme},
author = {Xesús Nogueira-Garea and Javier Fernández-Fidalgo and Lucía Ramos-García and Iván Couceiro-Aguiar and Luis Ramírez-Palacios},
url = {https://www.sciencedirect.com/science/article/pii/S0898122124002505},
doi = {10.1016/j.camwa.2024.05.031},
issn = {0898-1221},
year = {2024},
date = {2024-08-01},
urldate = {2026-04-21},
journal = {Computers & Mathematics with Applications},
volume = {168},
pages = {84–99},
abstract = {Machine learning (ML) is becoming a powerful tool in Computational Fluid Dynamics (CFD) to enhance the accuracy, efficiency, and automation of simulations. Currently, in the design of shock-capturing methods, there is still a heavy reliance on the expertise and scientific knowledge of each author, particularly in nonlinear components such as smoothness indicators and weighting functions. ML has the potential to reduce this dependency, since by leveraging large datasets, they can learn intricate patterns and make accurate predictions of these functions. In this work we present a neural network that compute the weighting functions in the WENO5 scheme. The proposed WENO5-NN scheme generalizes well for different resolutions, and in most of the cases tested, it outperforms the classical WENO5-JS scheme.},
keywords = {Euler equations, Finite difference, machine learning, Neural networks, WENO},
pubstate = {published},
tppubtype = {article}
}
Alvarellos-González, Alberto; Figuero-Pérez, Andrés; Rodríguez-Yáñez, Santiago; Sande-González-Cela, José; Peña-González, Enrique; Rosa-Santos, Paulo; Rabuñal-Dopico, Juan R.
Deep Learning-Based Wave Overtopping Prediction Journal Article
In: Applied Sciences, vol. 14, no. 6, pp. 2611, 2024, ISSN: 2076-3417.
Abstract | Links | BibTeX | Tags: Deep learning, machine learning, Neural networks, port management, port security, wave overtopping prediction
@article{alvarellos_deep_2024,
title = {Deep Learning-Based Wave Overtopping Prediction},
author = {Alberto Alvarellos-González and Andrés Figuero-Pérez and Santiago Rodríguez-Yáñez and José Sande-González-Cela and Enrique Peña-González and Paulo Rosa-Santos and Juan R. Rabuñal-Dopico},
url = {https://www.mdpi.com/2076-3417/14/6/2611},
doi = {10.3390/app14062611},
issn = {2076-3417},
year = {2024},
date = {2024-01-01},
urldate = {2026-04-20},
journal = {Applied Sciences},
volume = {14},
number = {6},
pages = {2611},
publisher = {Multidisciplinary Digital Publishing Institute},
abstract = {This paper analyses the application of deep learning techniques for predicting wave overtopping events in port environments using sea state and weather forecasts as inputs. The study was conducted in the outer port of Punta Langosteira, A Coruña, Spain. A video-recording infrastructure was installed to monitor overtopping events from 2015 to 2022, identifying 3709 overtopping events. The data collected were merged with actual and predicted data for the sea state and weather conditions during the overtopping events, creating three datasets. We used these datasets to create several machine learning models to predict whether an overtopping event would occur based on sea state and weather conditions. The final models achieved a high accuracy level during the training and testing stages: 0.81, 0.73, and 0.84 average accuracy during training and 0.67, 0.48, and 0.86 average accuracy during testing, respectively. The results of this study have significant implications for port safety and efficiency, as wave overtopping events can cause disruptions and potential damage. Using deep learning techniques for overtopping prediction can help port managers take preventative measures and optimize operations, ultimately improving safety and helping to minimize the economic impact that overtopping events have on the port’s activities.},
keywords = {Deep learning, machine learning, Neural networks, port management, port security, wave overtopping prediction},
pubstate = {published},
tppubtype = {article}
}
2023
Costas, Raquel; Carro-Fidalgo, Humberto; Figuero-Pérez, Andrés; Peña-González, Enrique; Sande-González-Cela, José
A Decision-Making Tool for Port Operations Based on Downtime Risk and Met-Ocean Conditions including Infragravity Wave Forecast Journal Article
In: Journal of Marine Science and Engineering, vol. 11, no. 3, pp. 536, 2023, ISSN: 2077-1312.
Abstract | Links | BibTeX | Tags: long waves prediction, machine learning, multidimensional operational threshold, port management, Port operability
@article{costas_decision-making_2023,
title = {A Decision-Making Tool for Port Operations Based on Downtime Risk and Met-Ocean Conditions including Infragravity Wave Forecast},
author = {Raquel Costas and Humberto Carro-Fidalgo and Andrés Figuero-Pérez and Enrique Peña-González and José Sande-González-Cela},
url = {https://www.mdpi.com/2077-1312/11/3/536},
doi = {10.3390/jmse11030536},
issn = {2077-1312},
year = {2023},
date = {2023-03-01},
urldate = {2026-08-13},
journal = {Journal of Marine Science and Engineering},
volume = {11},
number = {3},
pages = {536},
publisher = {Multidisciplinary Digital Publishing Institute},
abstract = {Port downtime leads to economic losses and reductions in safety levels. This problem is generally assessed in terms of uni-variable thresholds, despite its multidimensional nature. The aim of the present study is to develop a downtime probability forecasting tool, based on real problems at the Outer Port of Punta Langosteira (Spain), and including infragravity wave prediction. The combination of measurements from three pressure sensors and a tide gauge, together with machine-learning techniques, made it possible to generate long wave prognostication at different frequencies. A fitting correlation of 0.95 and 0.9 and a root mean squared error (RMSE) of 0.022 m and 0.012 m were achieved for gravity and infragravity waves, respectively. A wave hindcast in the berthing areas, met-ocean forecast data, and information on 15 real operational problems between 2017 and 2022, were all used to build a classification model for downtime probability estimation. The proposed use of this tool addresses the problems that arise when two consecutive sea states have thresholds above 3.97%. This is the limit for guaranteeing the safety of port operations and has a cost of just 0.6 unnecessary interruptions of operations per year. The methodology is easily exportable to other facilities for an adequate assessment of downtime risks.},
keywords = {long waves prediction, machine learning, multidimensional operational threshold, port management, Port operability},
pubstate = {published},
tppubtype = {article}
}