Benchmark Datasets for Photopolymer Modelling and Machine Learning: Coverage, Leakage, Shift and Fitness for Use
Segurola, Juan
Benchmark datasets can accelerate photopolymer modelling and machine learning, but a benchmark is useful only when its split structure, coverage and metadata reflect the intended engineering deployment. Random row splits can leak batch, formulation, machine or specimen-family information across training and test sets and produce misleadingly optimistic performance. This review defines benchmark fitness in terms of target decision, coverage, independence, leakage control, distribution shift and uncertainty. It proposes split hierarchies and reporting requirements for photopolymer datasets whose dominant dependencies arise from shared material and process histories. ER-398.
Texto completo
DOI de esta versión 10.5281/zenodo.23202645 · Todas las versiones en Zenodo