Skip to content

Author

Yao-Ming Hong

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#software testing Open access Sep 2026

ConGen-R v0.3.0: Source-Aware Validation, Uncertainty Calibration, and Reliability Assessment for Concrete Compressive Strength Prediction

ConGen-R (Concrete Generalization and Reliability Tool) is bilingual research software developed to accompany the study “When High Accuracy Fails to Generalize: Source-Aware Validation and Uncertainty Calibration for Concrete Compressive Strength Prediction.” It addresses the risk that random observation-level data splitting may substantially overestimate the performance and engineering portability of machine-learning models for concrete compressive-strength prediction. The software implements publication-disjoint validation, mechanism-informed feature engineering, external dataset transfer assessment, split-conformal uncertainty calibration, publication-cluster bootstrap inference, training-support diagnostics, and material-domain warnings. HistGradientBoosting and a 300-tree Extra Trees model are provided for algorithmic comparison. The interactive application supports individual-mixture and batch CSV analysis and reports point predictions, an empirical 90% cross-publication prediction interval, inter-model disagreement, standardized nearest-neighbour distance, training-support categories, and application-risk warnings. This release contains: a reproducible Python 3.12 analysis package; command-line validation and prediction workflows; a Streamlit application; a standalone Traditional Chinese/English browser interface; trained HistGradientBoosting and Extra Trees deployment models; input auditing and physical-range checks; example input files; model-reconstruction and verification tools; automated core tests; citation, licensing, and reproducibility documentation. The primary analysis used 3,013 usable mixture–age observations from a single Zenodo benchmark containing records originating from 58 source publications. The DOI labels supplied with the benchmark were retained as publication-level grouping identifiers for leave-one-publication-out validation. The original benchmark is available as “Concrete Materials Data Extraction Benchmark, Version v1” at https://doi.org/10.5281/zenodo.22132837. Third-party row-level benchmark and UCI data are not redistributed in this software archive. Users wishing to reproduce the complete validation analysis must obtain the datasets from their original repositories and comply with the applicable licences. The browser application contains trained model parameters and training-support diagnostics required for local inference, but not the third-party row-level observations. ConGen-R is intended for research, teaching, mixture screening, and preliminary experimental planning. Its predictions, uncertainty intervals, support categories, and risk warnings are diagnostic aids rather than certified engineering acceptance criteria. The software does not replace trial batching, compressive-strength testing, mixture qualification, structural design, applicable standards, or professional engineering judgement. Predictions for oyster-shell concrete, recycled aggregates, or other materials not explicitly represented in the training benchmark should be treated as exploratory and confirmed experimentally. ConGen-R v0.3.0 is released under the MIT License.

Yao-Ming Hong · 0 citations
#software testing Dataset Open access Sep 2026

Global Rice Flowering-Stage Heat Exposure, 1961–2021: Derived Data and Reproducible Workflows

This repository provides the derived datasets, quality-control records, data dictionaries, analysis code, software environment specifications, and automated validation procedures supporting a global assessment of rice flowering-stage heat exposure from 1961 to 2021. The reproducibility package integrates rice-distribution information, crop-calendar data, daily temperature records, and time-varying harvested-area weights. It includes the main flowering-stage exposure estimates, sensitivity analyses using relative flowering centres of 0.65, 0.70, and 0.75, and nonlinear trend diagnostics. The nonlinear workflow includes LOWESS, cubic-spline modelling, segmented regression, temporal block cross-validation, and a 1,000-replicate moving-block bootstrap for breakpoint uncertainty. The package contains the data and workflows required to reproduce the reported derived results, including 424 ERA5-Land climate points, 530,862 rice-growing grid cells, 750,395 grid-season phenology records, and 77,775 cluster-year-centre sensitivity estimates. Newey–West heteroskedasticity- and autocorrelation-consistent inference uses a Bartlett kernel with a four-year lag. The flowering-centre sensitivity results are robust across the three tested settings. The nonlinear analysis identifies an estimated acceleration in warm-night exposure around 2001, with a bootstrap 95% interval of 1986–2010; this interval does not support interpreting 2001 as a uniquely determined breakpoint. Original ERA5-Land, GloRice, RiceAtlas, and FAOSTAT source data remain subject to their respective access conditions and licences. Users should consult the README, data dictionary, third-party data notice, and file manifest before reusing the materials.

Yao-Ming Hong · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.