Data-driven multi-objective optimization of ML inference hardware configurations for energy, performance and cost

Selecting ML inference hardware requires balancing energy, performance, and cost, a core sustainable computing challenge. We present a data-driven framework that provides actionable decision support by identifying Pareto-optimal hardware configurations. We leverage a combined MLPerf™ Inference (v4.1...

Descripción completa

Detalles Bibliográficos
Autores: Castaño Fernández, Joel, Bustillo Ramírez, Jaime, Franch Gutiérrez, Javier|||0000-0001-9733-8830, Martínez Fernández, Silverio Juan|||0000-0001-9928-133X
Tipo de recurso: artículo
Fecha de publicación:2026
País:España
Institución:Universitat Politècnica de Catalunya (UPC)
Repositorio:UPCommons. Portal del coneixement obert de la UPC
Idioma:inglés
OAI Identifier:oai:dnet:upcommonspor::bd837069cf7acf68c8391dee21a066c5
Acceso en línea:https://hdl.handle.net/2117/461089
https://dx.doi.org/10.1016/j.suscom.2026.101347
Access Level:acceso abierto
Palabra clave:Sustainable computing
MLPerf
Energy efficiency
Multi-objective optimization
ML inference
Throughput
Benchmarking
Àrees temàtiques de la UPC::Informàtica::Intel·ligència artificial::Aprenentatge automàtic
Àrees temàtiques de la UPC::Informàtica::Impacte ambiental
Descripción
Sumario:Selecting ML inference hardware requires balancing energy, performance, and cost, a core sustainable computing challenge. We present a data-driven framework that provides actionable decision support by identifying Pareto-optimal hardware configurations. We leverage a combined MLPerf™ Inference (v4.1/v5.1) corpus, enriching it with standardized hardware descriptors, proxy energy (TDP-based with a conservative CPU factor), and indicative component cost. We then train task-specific regressors to predict throughput and energy per unit, using grouped cross-validation to avoid system-level leakage. The framework combines these predictions with cost to generate workload-specific, multi-objective recommendations. Across most tasks, results show divergent feature importance: accelerator scale and generation dominate throughput, while energy efficiency depends on architecture, memory subsystem, host CPU, and the Offline vs. Server scenario. Predictive models demonstrate strong accuracy: predicted Pareto sets match true fronts on held-out systems (median precision 0.92, recall 0.89). For computer vision, a balanced recommendation reduced energy per unit by 75.0% and cost by 86.5% compared to a max-throughput baseline, retaining 46.2% performance. Furthermore, our Top-3 predicted balanced configurations included a true Pareto-optimal system in 100% of test cases, versus 34.5% for random selection. For generative language tasks, prediction accuracy is robust for modern LLMs, though specialized tasks show variability due to unmodeled software features. Sensitivity analyses confirm recommendations remain robust to energy proxy and pricing uncertainty. This framework makes complex sustainable hardware selection trade-offs explicit and actionable.