Aug 2026· SPE Nigeria Annual International Conference and Exhibition· 0 citations· 21 references
Abstract
This study evaluates the prediction of flowing bottom-hole pressure (FBHP) in dry gas wells using machine learning techniques, specifically Random Forest and Artificial Neural Network (ANN) models. Unlike earlier work based on PROSPER-generated synthetic data, this study utilizes a real field dataset of 206 samples obtained from the ProBHP repository, originally compiled by Govier and Fogarasi (1975) and Asheim (1986). The dataset comprises 10 input variables, including production rates, well depth, tubing size, temperatures, and wellhead pressure, with measured bottom-hole pressure (MBHP) as the target. Feature importance analysis identified well depth, oil rate, water rate, and wellhead pressure as the most influential parameters. The data were split into 80% training and 20% testing sets, with Z-score-based outlier removal reducing the training data slightly. The Random Forest model showed strong predictive performance, achieving a test R2 of 0.81, MAE of 93.73 psig, and RMSE of 123.39 psig, with a cross-validation R2 of 0.72 ± 0.11. In contrast, the ANN model performed poorly, with a test R2 of 0.05 and MAE of 209.55 psig. Overall, the results highlight the reliability of Random Forest for FBHP prediction using real field data, while also showing the limitations of a simple ANN model on small, noisy datasets. The identified key parameters provide useful insights for well performance monitoring and production optimization in dry gas systems.
Accurate prediction of dew point pressure (Pd) is critical for managing gas condensate reservoirs, as liquid dropout near the wellbore creates condensate banking that reduces permeability and well productivity. This study proposes a data-driven approach using a Random Forest (RF) machine learning algorithm to predict Pd based on reservoir temperature and fluid composition. Utilizing a comprehensive dataset of 375 records, the model incorporates 13 predictors, including hydrocarbon fractions (C1 through C7+), non-hydrocarbons (N2, CO2, H2S), and heavy fraction properties (MC7+, γC7+).
The proposed RF model exhibits highly competitive accuracy, outperforming most traditional empirical correlations by achieving a superior overall data variance capture with a Coefficient of Determination (R2) of 0.8790 and an Average Absolute Percent Relative Error (APE) of 8.17%. While the Ahmadi-Elsharkawy model shows a marginally lower APE (7.90%), the RF framework avoids complex genetic programming equations and delivers superior global consistency across the entire pressure envelope. By capturing complex, non-linear thermodynamic interactions, the RF model provides a robust, fast, and cost-effective alternative to expensive laboratory PVT tests and complex equations of state, optimizing fluid characterization and production system design.
Alejandro Osorio Pozo, Lucio A. Perez, Karim Botan et al.· Romanian Journal of Petroleu...· 0 citations
Sand production has become a significant concern in the hydrocarbon recovery process from unconsolidated reservoirs which may result in equipment damage, flow restrictions, and costly operational downtime in vertical oil wells. Accurate prediction of sand production is vital to optimize well integrity. Conventional geomechanical and empirical models frequently fail to capture the highly non-linear interactions among reservoir pressure, multiphase flow rates, rock mechanical properties, and dynamic operating conditions. This study addresses the identified research gap by developing and systematically comparing four supervised machine learning classifiers for binary prediction of sand production occurrence using routine well-test data from a single vertical oil well in the Niger Delta basin. A total of 235 validated well test observations consisting of 19 recorded variables which include date and operational parameters such as production rates, pressure conditions, choke size, and fluid properties were pre-processed and analyzed. The target variable was formulated as a binary classification problem with the operational threshold sand rate > 0 lb/1000 bbl to enable early detection of any sanding event. Four machine learning algorithms including Logistic Regression, Decision Tree, Random Forest, and Support Vector Machine were developed and evaluated following feature optimization and model tuning. Model performance was assessed using accuracy, precision, recall, F1-score, and ROC–AUC metrics. The results show that ensemble and kernel-based methods significantly outperform linear and single-tree models, with the Random Forest classifier achieving the best model prediction accuracy of 93.62%. This strong performance demonstrates the model's robustness in capturing the complex, nonlinear interactions governing sand production behavior. This study demonstrates that machine learning classifiers can be effectively utilized in a manner that enables proactive sand management strategies, including choke adjustment, artificial lift optimization and selective sand control deployment, ensuring a minimal risk of equipment failure and increasing overall well productivity and hydrocarbon production.
S. E. Balogun, A. Joledo· SPE Nigeria Annual Internati...· 0 citations
Reliable estimation of liquid accumulation in gas-well tubing is important for characterizing liquid-loading conditions and supporting engineering assessment. Traditional mechanistic approaches commonly depend on an extensive set of wellbore descriptors and empirical parameters, while also requiring complicated solution procedures. This work addresses these constraints through a data-driven predictive framework that couples ensemble feature selection with ant-colony-optimized support vector regression (ACO-SVR). A majority-voting scheme was applied to 107 production-test records collected from a gas field. The scheme combined linear regression, grey relational analysis, random-forest mean decrease in impurity, the Pearson correlation coefficient, and SHAP attribution, and selected seven dominant factors from 11 candidate variables: casing pressure, tubing pressure, tubing depth, reservoir mid-depth, daily gas production, daily water production, and wellhead temperature. Ant colony optimization subsequently determined the SVR hyperparameters. Evaluation with 32 held-out well samples produced a root-mean-square error of 165.73 m, a mean absolute error of 103.26 m, a coefficient of determination of 0.94, and a mean relative error of 2.11%. Repeated five-fold cross-validation further yielded an average R2 of 0.91±0.04 and an RMSE of 181.6±24.8 m, indicating moderate variability across alternative data partitions. Relative to the untuned SVR, ACO-SVR lowered the root-mean-square error and mean absolute error by approximately 27.0% and 31.1%, respectively. Its mean relative error was also 3.66 percentage points below that of the PLATA model. The resulting framework provides accurate prediction of tubing liquid accumulation height from a small sample and offers quantitative information for liquid-loading assessment under the investigated operating conditions.
Improving drilling efficiency remains a central objective in petroleum operations due to its direct impact on time and cost. This study develops a predictive framework for estimating the Rate of Penetration (ROP) using supervised machine learning techniques combined with systematic data conditioning. The analysis is based on more than 11,000 measurements obtained from four horizontal wells. A rigorous preprocessing strategy was implemented to enhance data reliability, including removal of invalid entries and statistical outliers using the interquartile range method. This procedure reduced the dataset to 7,297 high-quality observations. In addition, target stabilization was introduced through Exponential Moving Average smoothing (spans of 5 and 10), which reduced short-term fluctuations and improved the learnability of the ROP signal. Three tree-based regression models—Decision Tree, Random Forest, and Gradient Boosting—were evaluated under both default configurations and optimized settings. Results show that model performance is strongly influenced by data conditioning. The Random Forest model achieved the highest accuracy, with a coefficient of determination (R2) of 0.96 and a mean squared error (MSE) of 26 when trained on the EMA-10 dataset. Gradient Boosting exhibited the largest improvement from hyperparameter tuning, with R2 increasing from 0.86 to 0.95. To bridge the gap between model development and practical use, the trained models were implemented in interactive applications for real-time prediction and parameter optimization. The outcomes demonstrate that careful preprocessing and noise-aware modeling significantly enhance predictive capability.
B. Elahifar, Thomas Philip Fagerli· Journal of Energy Resources...· 0 citations
Understanding permeability is essential for evaluating reservoir quality and field development planning. Reliable permeability estimation can reduce the uncertainty in reservoir characterization, particularly in intervals where core data are limited. As the industry relies on log-based interpretations and empirical correlations, the limitations of these approaches become apparent. Data-driven approaches offer a promising alternative to conventional empirical methods. The data set in this study comprises 252 samples with seven features derived from conventional well logs. Data preprocessing includes handling missing values, smoothing logs, feature engineering to add an extra input, and transformation with the Yeo-Johnson technique. A center moving average filter was used to reduce variance and improve data consistency. Ensemble machine-learning (ML) and baseline models were developed and evaluated using a 75-25 train-test split, four-fold cross validation, and model complexity assessment. Ensemble methods outperformed baseline models, with extremely randomized trees (ET) and random forest (RF) emerging as the most stable, achieving a mean R² of 0.93 and 0.89 and a low R² standard deviation (0.3). Multilinear regression (MLR) and artificial neural networks (ANNs) show limited accuracy, while gradient boosting (GB) and extreme gradient boosting (XGBoost) methods exhibit overfitting despite perfect training scores. Predicted kh values were compared with core data. Both linear and nonlinear empirical equations were derived using MLR, polynomial regression, and a power-law model. The power-law model (empirical equation) achieved an R² value of 0.79 and can therefore be used to estimate permeability. Additionally, a Gaussian mixture model (GMM) was used for unsupervised classification of hydraulic flow units (HFU) using the flow zone indicator (FZI), computed from the continuous permeability curve obtained from the best ML model. Thus, ML-based permeability prediction is an indispensable component of HFU modeling. The model identified three distinct flow zones, enabling HFU clustering and defining their corresponding petrophysical properties and depositional environments.
Vikram Kumar, Sayantan Ghosh, S. Maiti· Petrophysics· 0 citations
With increasing difficulty in oil and gas field development, accurate prediction of well productivity has become crucial. Traditional methods such as analytical solutions and numerical simulations have limited accuracy under heterogeneous and complex flow conditions. This study develops a CNN-LSTM model combining convolutional neural networks (CNN) and long short-term memory networks (LSTM) based on measured data from well J-1 in a shale oil block in eastern China for short-term multi-dimensional time series production forecasting. The model integrates CNN’s feature extraction with LSTM’s temporal modeling, using inputs including production rate, oil pressure, casing pressure, and production time. Compared with the traditional random forest (RF) model, the CNN-LSTM outperforms across R², MAE, MAPE, and RMSE metrics, achieving an R² of 0.9707 and MAPE below 5.1% on the test set. Results demonstrate strong fitting and predictive capabilities, indicating good applicability and potential for broader use in shale oil production forecasting.
Jiang Yao· Thermal Science· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.