Data-driven multi-objective optimization of ML inference hardware configurations for energy, performance and cost
Selecting ML inference hardware requires balancing energy, performance, and cost, a core sustainable computing challenge. We present a data-driven framework that provides actionable decision support by identifying Pareto-optimal hardware configurations. We leverage a combined MLPerf™ Inference (v4.1...
| Autores: | , , , |
|---|---|
| Tipo de recurso: | artículo |
| Fecha de publicación: | 2026 |
| País: | España |
| Institución: | Universitat Politècnica de Catalunya (UPC) |
| Repositorio: | UPCommons. Portal del coneixement obert de la UPC |
| Idioma: | inglés |
| OAI Identifier: | oai:dnet:upcommonspor::bd837069cf7acf68c8391dee21a066c5 |
| Acceso en línea: | https://hdl.handle.net/2117/461089 https://dx.doi.org/10.1016/j.suscom.2026.101347 |
| Access Level: | acceso abierto |
| Palabra clave: | Sustainable computing MLPerf Energy efficiency Multi-objective optimization ML inference Throughput Benchmarking Àrees temàtiques de la UPC::Informàtica::Intel·ligència artificial::Aprenentatge automàtic Àrees temàtiques de la UPC::Informàtica::Impacte ambiental |
| Sumario: | Selecting ML inference hardware requires balancing energy, performance, and cost, a core sustainable computing challenge. We present a data-driven framework that provides actionable decision support by identifying Pareto-optimal hardware configurations. We leverage a combined MLPerf™ Inference (v4.1/v5.1) corpus, enriching it with standardized hardware descriptors, proxy energy (TDP-based with a conservative CPU factor), and indicative component cost. We then train task-specific regressors to predict throughput and energy per unit, using grouped cross-validation to avoid system-level leakage. The framework combines these predictions with cost to generate workload-specific, multi-objective recommendations. Across most tasks, results show divergent feature importance: accelerator scale and generation dominate throughput, while energy efficiency depends on architecture, memory subsystem, host CPU, and the Offline vs. Server scenario. Predictive models demonstrate strong accuracy: predicted Pareto sets match true fronts on held-out systems (median precision 0.92, recall 0.89). For computer vision, a balanced recommendation reduced energy per unit by 75.0% and cost by 86.5% compared to a max-throughput baseline, retaining 46.2% performance. Furthermore, our Top-3 predicted balanced configurations included a true Pareto-optimal system in 100% of test cases, versus 34.5% for random selection. For generative language tasks, prediction accuracy is robust for modern LLMs, though specialized tasks show variability due to unmodeled software features. Sensitivity analyses confirm recommendations remain robust to energy proxy and pricing uncertainty. This framework makes complex sustainable hardware selection trade-offs explicit and actionable. |
|---|