Benchmark Datasets for Photopolymer Modelling and Machine Learning: Coverage, Leakage, Shift and Fitness for Use
Segurola, Juan
Benchmark datasets can accelerate photopolymer modelling and machine learning, but a benchmark is useful only when its split structure, coverage and metadata reflect the intended engineering deployment. Random row splits can leak batch, formulation, machine or specimen-family information across training and test sets and produce misleadingly optimistic performance. This review defines benchmark fitness in terms of target decision, coverage, independence, leakage control, distribution shift and uncertainty. It proposes split hierarchies and reporting requirements for photopolymer datasets whose dominant dependencies arise from shared material and process histories. ER-398.
Full text
Version DOI 10.5281/zenodo.23202645 · All versions in Zenodo