Study on the Impact of Genetic Optimization Algorithm on the Sintering Process of Lithium Battery Cathode Materials
Abstract
# Surrogate-Assisted Genetic Algorithm for NCM91 Sintering Optimization This repository accompanies the paper "Study on the Impact of a GeneticOptimization Algorithm on the Sintering Process of Lithium Battery CathodeMaterials" (JoVE, 2026). It contains the coin-cell measurement recordscollected during the study, a back-propagation neural network (BPNN)ensemble surrogate, a real-coded genetic algorithm with adaptive Pc/Pm,four alternative optimization strategies used as benchmarks, fouralternative surrogate families (Ridge, SVR, Random Forest, GaussianProcess) used for the Table 3 model comparison, and the analysis scriptsthat produce every table and figure of the paper. ## Repository layout ```sintering_ga_optimization/├── README.md├── requirements.txt├── config/│ └── settings.py Seeds, sampling ranges, fitness references,│ constraints, GA/surrogate hyper-parameters├── src/│ ├── data_io.py CSV loading utilities│ ├── fitness.py Equation 3 (fitness) and the constraint penalty│ ├── surrogate.py BPNN ensemble (five 64-32-16 networks),│ repeated K-fold cross validation│ ├── alternative_models.py Ridge, SVR, RF, GPR (Matérn 5/2) with│ inner 5-fold grid search on shared partitions│ ├── genetic_algorithm.py SBX, polynomial mutation, tournament,│ elitism, adaptive Pc/Pm (Equations 4-7, 11, 12)│ ├── virtual_testbed.py Deterministic analytical testbed used for│ algorithm comparison│ └── benchmarks.py OED (L324), RSM (CCD + refinement),│ PSO, simulated annealing├── scripts/│ ├── run_pipeline.py End-to-end NCM91 analysis (seven steps)│ ├── run_ncm811_transfer.py Cross-chemistry NCM811 transfer│ └── make_figures.py Rebuild the fourteen data-driven figures├── data/│ ├── lhs_initial_100.csv 100 initial LHS coin-cell records│ ├── verification_stages_68.csv 68 verification coin-cell records│ ├── training_168.csv Concatenation of the two above│ ├── outside_loop_20.csv 20 outside-loop records│ ├── table4_cells.csv 18 Table 4 coin cells│ ├── ncm811_transfer_30.csv 30 NCM811 verification records│ └── ncm811_traditional_vs_optimized.csv 18 NCM811 two-process coin cells└── results/ ├── ga_ncm91_primary_log.csv Per-generation log of the primary GA run ├── surrogate_cv.json BPNN ensemble cross validation metrics ├── alternative_models_cv.json Table 3 Ridge/SVR/RF/GPR comparison ├── benchmarks.csv OED/RSM/PSO/SA/GA fitness on the testbed ├── multi_seed_stats.json Replicate GA statistics ├── regression_coefficients_capacity.csv Supp Table S6 ├── table4_statistics.json Replicate statistics and Welch tests ├── ncm811_parity.csv Surrogate parity on the 30 NCM811 experiments ├── ncm811_transfer_summary.json ├── paper_reference_benchmarks.json Numeric values reported in the paper └── figures/ Fourteen 800-DPI PDFs (vector)``` ## Data columns Each CSV in `data/` records one coin cell per row. Shared columns: | Column | Meaning || ------ | ------- || `label` | Record identifier || `stage`, `generation` | Verification stage and GA generation || `T_C`, `t_h` | Sintering temperature and holding time || `heating_rate_Cmin`, `Li_TM_ratio`| Heating rate and Li/TM molar ratio || `discharge_capacity_0p1C_mAhpg` | First-cycle discharge capacity at 0.1 C || `retention_70cycles_0p5C_pct` | Capacity retention after 70 cycles at 0.5 C || `rate_capacity_ratio_2C_pct` | Capacity retention at 2 C relative to 0.1 C || `specific_energy_kWh_per_kg` | Specific sintering energy per kg of on-spec product || `cation_mixing_pct` | Ni²⁺ occupancy of Li sites (from Rietveld refinement) || `D50_um` | Median secondary-particle grain size || `fitness` | Fitness computed by Equation 3 | Columns specific to Table 4 (`table4_cells.csv`) use shorter names(`C`, `R`, `discharge_2C`, `E`) and add `batch`/`cell`/`process`. TheNCM811 tables use the same schema minus the two constraint responses. ## Reproducibility Seeds are set with the derivation rules documented in the Protocol:* main seed `20260930`* surrogate ensemble members `20260930 + m` for `m` in `0..4`* 30 replicate GA runs `20260930 + s` for `s` in `0..29` Software versions used in the paper (`requirements.txt`): ```numpy>=1.24pandas>=2.0scipy>=1.11scikit-learn==1.4.0statsmodels>=0.14torch==2.2.1matplotlib>=3.7openpyxl>=3.1diptest>=0.8``` ## Installation ```bashgit clone https://github.com/ /sintering_ga_optimization.gitcd sintering_ga_optimizationpip install -r requirements.txt``` ## Running the analysis ```bash# Primary NCM91 pipeline (train surrogate, run GA, Table 3 model comparison,# benchmarks, quadratic regression, Table 4 statistics)python -m sintering_ga_optimization.scripts.run_pipeline # Cross-chemistry transfer to NCM811python -m sintering_ga_optimization.scripts.run_ncm811_transfer # Rebuild all fourteen data-driven figures from the pipeline outputspython -m sintering_ga_optimization.scripts.make_figures``` Pass `--fast` to either analysis script to shorten the surrogate trainingand reduce the number of outer-CV repetitions and replicate GA runs, sothat the pipeline finishes in roughly three minutes on a desktop CPU.In full mode the pipeline runs for about twenty minutes. ## Note on benchmark numbers The benchmark comparison runs each strategy on the deterministic analyticaltestbed described in `src/virtual_testbed.py`. The testbed reproduces thereported optimum (0.958 at `(878 °C, 16.5 h, 8.2 °C/min, Li/TM 1.08)`) butits exact ripple and local structure can be rebuilt from the trainingrecords in several ways, so the fitness values of the alternativestrategies may differ by a few percent from the values reported in Table 2of the paper. The qualitative hierarchy — the surrogate-assisted GA best,PSO second, OED/RSM/SA below — is preserved. For reference the paper'sreported values are stored in `results/paper_reference_benchmarks.json`. ## Pre-shipped versus regenerated outputs The files in `results/` and `results/figures/` are the production outputs of thefull-length pipeline reported in the paper (1000 GA generations, five surrogatefamilies cross-validated with 10×5 outer folds, 30 replicate GA runs, 300training epochs). Running `scripts/run_pipeline.py` without `--fast` reproducesthem from scratch in roughly twenty minutes on a desktop CPU; passing `--fast`shortens the training and the number of replicates so the pipeline finishes inabout three minutes, which is useful for a smoke test but gives shorter GAlogs and slightly different numbers. If you run in fast mode, expect`results/ga_ncm91_primary_log.csv` to shrink to 500 generations and the bestfitness reported in `results/figures/Figure 6.pdf` to shift accordingly; thepre-shipped files always reflect the full run.