Proyecto de optimización de Energía Eólica con aplicación de inteligencia artificial

Abstract

Este trabajo desarrolla y compara modelos de aprendizaje automático para predecir la generación horaria del parque eólico Peñascal I, situado en el condado de Kenedy (Texas), con 84 turbinas Mitsubishi MWT92/2.4 y 202 MW de potencia instalada. En un sistema con alta penetración renovable como el texano, los errores de previsión obligan a activar reservas y generan sobrecostes por desvío, de ahí el interés de este tipo de predicciones. La base de datos combina dos fuentes públicas. Por un lado, los registros de generación de ERCOT, publicados en intervalos de quince minutos y agregados a resolución horaria. Por otro, las variables meteorológicas del modelo HRRR, obtenidas mediante la API de Pronósticos Históricos de Open-Meteo. Sobre estos datos se construyeron variables adicionales: el exponente de Hellmann, la densidad del aire, la velocidad del viento al cubo, codificaciones cíclicas del calendario y retardos y medias móviles de la propia serie de generación, calculados con un desplazamiento previo que evita fugas de información. Se compararon cinco modelos (regresión lineal, Random Forest, XGBoost, LightGBM y CatBoost) con una partición temporal que reserva el tramo final de la serie para la evaluación. CatBoost obtuvo el mejor resultado, con un R² de 0,7233 y un nMAE del 4,91 % de la capacidad instalada, aunque las diferencias entre modelos fueron pequeñas. Las variables autorregresivas concentran la mayor parte del poder predictivo. El curtailment ordenado por el operador introduce además un error irreducible que limita la precisión alcanzable con información puramente meteorológica.
This project develops and compares machine learning models to predict the hourly generation of the Peñascal I wind farm, located in Kenedy County (Texas), with 84 Mitsubishi MWT92/2.4 turbines and 202 MW of installed capacity. In a system with high renewable penetration such as the Texan grid, forecasting errors force the activation of reserves and generate imbalance costs, hence the interest in this type of prediction. The database combines two public sources. On one side, the generation records from ERCOT, published at fifteen-minute intervals and aggregated to hourly resolution. On the other, the meteorological variables from the HRRR model, obtained through the Open-Meteo Historical Forecast API. Additional variables were built on top of these data: the Hellmann exponent, air density, cubed wind speed, cyclic calendar encodings, and lags and rolling means of the generation series itself, all computed with a prior shift that prevents information leakage. Five models were compared (linear regression, Random Forest, XGBoost, LightGBM and CatBoost) using a temporal split that reserves the final stretch of the series for evaluation. CatBoost achieved the best result, with an R² of 0.7233 and an nMAE of 4.91% of installed capacity, although the differences between models were small. The autoregressive variables concentrate most of the predictive power. Curtailment ordered by the system operator also introduces an irreducible error that limits the accuracy attainable with purely meteorological information.
Ítem

Información detallada

Materias, derechos, colecciones e identificadores

Keywords

KTI-electronica (GITI-N), predicción eólica, aprendizaje automático, ERCOT, ingeniería de características, gradient boosting., wind generation forecasting, machine learning, ERCOT, feature engineering, gradient boosting.

Rights

Attribution-NonCommercial-NoDerivs 3.0 United States