AGAMOS: A graph-based approach to modulo scheduling for clustered microarchitectures

This paper presents AGAMOS, a technique to modulo schedule loops on clustered microarchitectures. The proposed scheme uses a multilevel graph partitioning strategy to distribute the workload among clusters and reduces the number of intercluster communications at the same time. Partitioning is guided...

Descripción completa

Detalles Bibliográficos
Autores: Aleta Ortega, Alexandre, Codina Viñas, Josep M., Sánchez Navarro, F. Jesús, González Colás, Antonio María|||0000-0002-0009-0996, Kaeli, D
Tipo de recurso: artículo
Fecha de publicación:2009
País:España
Institución:Universitat Politècnica de Catalunya (UPC)
Repositorio:UPCommons. Portal del coneixement obert de la UPC
Idioma:inglés
OAI Identifier:oai:upcommons.upc.edu:2117/91144
Acceso en línea:https://hdl.handle.net/2117/91144
https://dx.doi.org/10.1109/TC.2009.32
Access Level:acceso abierto
Palabra clave:Microprocessors
Graph theory
Clustered microarchitectures
ILP
Instruction replication
Modulo scheduling
Statically scheduled processors
Microprocessadors
Grafs, Teoria de
Àrees temàtiques de la UPC::Informàtica::Arquitectura de computadors
Descripción
Sumario:This paper presents AGAMOS, a technique to modulo schedule loops on clustered microarchitectures. The proposed scheme uses a multilevel graph partitioning strategy to distribute the workload among clusters and reduces the number of intercluster communications at the same time. Partitioning is guided by approximate schedules (i.e., pseudoschedules), which take into account all of the constraints that influence the final schedule. To further reduce the number of intercluster communications, heuristics for instruction replication are included. The proposed scheme is evaluated using the SPECfp95 programs. The described scheme outperforms a state-of-the-art scheduler for all programs and different cluster configurations. For some configurations, the speedup obtained when using this new scheme is greater than 40 percent, and for selected programs, performance can be more than doubled.