research-article

Open access

Efficient Generation of Compact Execution Traces for Multicore Architectural Simulations

Authors:

M. E. S. Elrabaa,

A. KhayyatAuthors Info & Claims

ACM Transactions on Architecture and Code Optimization (TACO), Volume 14, Issue 3

Article No.: 27, Pages 1 - 25

https://doi.org/10.1145/3106342

Published: 30 August 2017 Publication History

Abstract

Requiring no functional simulation, trace-driven simulation has the potential of achieving faster simulation speeds than execution-driven simulation of multicore architectures. An efficient, on-the-fly, high-fidelity trace generation method for multithreaded applications is reported. The generated trace is encoded in an instruction-like binary format that can be directly “interpreted” by a timing simulator to simulate a general load/store or x8-like architecture. A complete tool suite that has been developed and used for evaluation of the proposed method showed that it produces smaller traces over existing trace compression methods while retaining good fidelity including all threading- and synchronization-related events.

Supplementary Material

TACO1403-27 (taco1403-27.pdf)

Slide deck associated with this paper

Download
759.02 KB

References

[1]

S. B. (Intel). 2012. Pin - A dynamic binary instrumentation tool. Retrieved from https://software.intel.com/en-us/articles/pin-a-dynamic-binary-instrumentation-too

[2]

E. Argollo, A. Falc, P. Faraboschi, M. Monchiero, and D. Ortega. 2009. COTSon: Infrastructure for full system simulation. SIGOPS Oper. Syst. Rev. 43, 1, 52--61.

Digital Library

[3]

K. C. Barr and K. Asanovic. 2006. Branch trace compression for snapshot-based simulation. In Proceedings of the 2006 IEEE International Symposium on Performance Analysis of Systems and Software.

[4]

K. C. Barr, H. Pan, M. Zhang, and K. Asanovic. 2005. Accelerating multiprocessor simulation with a memory timestamp record. In Proceedings of the IEEE International Symposium on Performance Analysis of Systems and Software, 2005 (ISPASS’05).

Digital Library

[5]

S. Budanur, F. Mueller, and T. Gamblin. 2012. Memory trace compression and replay for SPMD systems using extended PRSDs. Comput. J. 55, 2, 206--217.

Digital Library

[6]

M. Burtscher, I. Ganusov, S. J. Jackson, J. Ke, P. Ratanaworabhan, and N. B. Sam. 2005. The VPC trace-compression algorithms. IEEE Trans. Comput. 54, 11, 1329--1344.

Digital Library

[7]

A. Butko, R. Garibotti, L. Ost, V. Lapotre, A. Gamatie, and G. Sassatelli. 2015. A trace-driven approach for fast and accurate simulation of manycore architectures. In Proceedings of the 2015 20th Asia and South Pacific Design Automation Conference (ASP-DAC'15).

[8]

T. E. Carlson, W. Heirman, and L. Eeckhout. 2011. Sniper: Exploring the level of abstraction for scalable and accurate parallel multi-core simulation. In Proceedings of the 2011 International Conference for High Performance Computing, Networking, Storage and Analysis (SC'11).

Digital Library

[9]

C.-W. Chen, C.-J. Ku, and T.-J. Liu. 2013. Efficient trace file compression design with locality and address difference. J. Inf. Sci. Eng. 29, 5, 1055--1070.

[10]

J. Edler and M. D. Hill. 1998. Dinero IV Trace-Driven Uniprocessor Cache Simulator. Retrieved from http://www.cs.wisc.edu/∼markhill/DineroIV/.

[11]

E. N. Elnozahy. 1999. Address trace compression through loop detection and reduction. In Proceedings of the 1999 ACM SIGMETRICS International Conference on Measurement and Modeling of Computer Systems.

Digital Library

[12]

A. Janapsatya, A. Ignjatovic, and J. Henkel. 2007. Instruction trace compression for rapid instruction cache simulation. In Proceedings of the Design, Automation 8 Test in Europe Conference 8 Exhibition, 2007 (DATE'07).

Digital Library

[13]

E, Johnson, J. Ha, and M. B. Zaidi. 2001. Lossless trace compression. IEEE Trans. Comput. 14, 158--173.

Digital Library

[14]

W. Jun, J. Beu, R. Bheda, T. Conte, D. Zhenjiang, and C. Kersey. 2014. Manifold: A parallel simulation framework for multicore systems. In Proceedings of the 2014 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS'14).

[15]

S. Kanev and R. Cohn. 2011. Portable trace compression through instruction interpretation. In Proceedings of the IEEE International Symposium on Performance Analysis of Systems and Software.

Digital Library

[16]

A. Ketterlin and P. Clauss. 2008. Prediction and trace compression of data access addresses through nested loop recognition. In Proceedings of the 6th Annual IEEE/ACM International Symposium on Code Generation and Optimization.

Digital Library

[17]

A. Khan, M. Vijayaraghavan, and S. Boyd-Wickizervind. 2012. Fast and cycle-accurate modeling of a multicore processor. In Proceedings of the IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS'12).

Digital Library

[18]

Y. Luo and L. K. John. 2004. Locality-based online trace compression. IEEE Trans. Comput. 53, 723--731.

Digital Library

[19]

J. Marathe, F. Mueller, T. Mohan, B. R. de Supinski, S. A. McKee, and A. Yoo. 2003. METRIC: Tracking down inefficiencies in the memory hierarchy via binary rewriting. Paper presented at the International Symposium on Code Generation and Optimization, 2003 (CGO’03).

Digital Library

[20]

MediaBench. 1997. http://euler.slu.edu/∼fritts/mediabench/.

[21]

A. Milenkovic and M. Milenkovic. 2003. Stream-based trace compression. Comput. Arch. Lett. 2, 4.

Digital Library

[22]

A. Milenkovic and M. Milenkovic. 2007. An efficient single-pass trace compression technique utilizing instruction streams. ACM Trans. Model. Comput. Simul. 17, 1, 2.

Digital Library

[23]

J. E. Miller, H. Kasture, G. Kurian, C. Gruenwald, N. Beckmann, and C. Celio. 2010. Graphite: A distributed parallel simulator for multicores. In Proceedings of the 2010 IEEE 16th International Symposium on High Performance Computer Architecture (HPCA'10).

[24]

S. Nilakantan, K. Sangaiah, A. More, G. Salvadory, B. Taskin, and M. Hempstead. 2015. Synchrotrace: Synchronization-aware architecture-agnostic traces for light-weight multicore simulation. In Proceedings of the 2015 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS'15).

[25]

M. Noeth, P. Ratn, F. Mueller, M. Schulz, and B. R. d. Supinski. 2009. Scalatrace: Scalable compression and replay of communication traces for high-performance computing. J. Parallel Distrib. Comput. 69, 8, 696--710.

Digital Library

[26]

PARSEC. 2007. from http://parsec.cs.princeton.edu/.

[27]

H. Patil, C. Pereira, M. Stallcup, G. Lueck, and J. Cownie. 2010. PinPlay: A framework for deterministic replay and reproducible analysis of parallel programs. In Proceedings of the 8th Annual IEEE/ACM International Symposium on Code Generation and Optimization.

Digital Library

[28]

R. Pengju, M. Lis, C. Myong Hyon, S. Keun Sup, C. W. Fletcher, and O. Khan. 2012. HORNET: A cycle-level multicore simulator. IEEE Trans. Comput.-Aided Design Integr. Circuits Syst. 31, 6, 890--903.

Digital Library

[29]

C. A. Prete, G. Prina, and L. Ricciardi. 1995. A trace-driven simulator for performance evaluation of cache-based multiprocessor systems. IEEE Trans. Parallel Distrib. Syst. 6, 9, 915--929.

Digital Library

[30]

A. Rico, A. Duran, F. Cabarcas, Y. Etsion, A. Ramirez, and M. Valero. 2011. Trace-driven simulation of multithreaded applications. Paper presented at the 2011 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS‘11).

Digital Library

[31]

Searchable Linux Syscall Table for x86 and x86_64. https://filippo.io/linux-syscall-table/.

[32]

Standard Performance Evaluation Corporation. 2000. https://www.spec.org/cpu2000/.

[33]

Z. Tan, A. Waterman, H. Cook, S. Bird, K. Asanovi, and D. Patterson. 2010. A case for FAME: FPGA architecture model execution. SIGARCH Comput. Archit. News 38, 3, 290--301.

Digital Library

[34]

TCgen 2.0: A Tool to Automatically Generate Lossless Trace Compressors (2006).

[35]

Valgrind Instrumentation Framework. 2000. http://valgrind.org.

[36]

S. C. Woo, M. Ohara, E. Torrie, J. P. Singh, and A. Gupta. 1995. The SPLASH-2 programs: characterization and methodological considerations. In Proceedings of the 22nd Annual International Symposium on Computer Architecture, 1995.

Digital Library

[37]

M. Xu, R. Bodik, and M. D. Hill. 2003. A “flight data recorder” for enabling full-system multiprocessor deterministic replay. SIGARCH Comput. Archit. News 31, 2, 122--135.

Digital Library

[38]

M. Xu, M. D. Hill, and R. Bodik. 2006. A regulated transitive reduction (RTR) for longer memory race recording. SIGOPS Oper. Syst. Rev. 40, 5, 49--60.

Digital Library

[39]

J. J. Yi, R. Sendag, L. Eeckhout, A. Joshi, D. J. Lilja, and L. K. John. 2006. Evaluating benchmark subsetting approaches. In Proceedings of the 2006 IEEE International Symposium on Workload Characterization.

[40]

X. Zhang and R. Gupta. 2005. Whole execution traces and their applications. ACM Trans. Arch. Code Optim. (TACO) 2, 3, 301--334.

Digital Library

Cited By

Brais HPanda P(2019)AlleriaACM Transactions on Embedded Computing Systems10.1145/335819318:5s(1-22)Online publication date: 8-Oct-2019
https://dl.acm.org/doi/10.1145/3358193

Index Terms

Efficient Generation of Compact Execution Traces for Multicore Architectural Simulations
1. Computer systems organization
  1. Architectures
    1. Parallel architectures
      1. Multicore architectures
2. Computing methodologies
  1. Modeling and simulation
    1. Simulation support systems
      1. Simulation tools

Recommendations

An iterative, multi-level, and scalable approach to comparing execution traces
ESEC-FSE '07: Proceedings of the the 6th joint meeting of the European software engineering conference and the ACM SIGSOFT symposium on The foundations of software engineering

In this paper, we overview a new approach to comparing execution traces. Such comparison can be useful for purposes such as improving test coverage and profiling system's users. In our approach, traces are compressed into different levels of compaction ...
Visualization, transformation, and analysis of execution traces with the eclipse TRACE4CPS trace tool
Abstract
An execution trace is a model of a single system behavior. Execution traces occur everywhere in the system’s lifecycle as they can typically be produced by executable models, by prototypes of (sub)systems, and by the system itself during its ...
An iterative, multi-level, and scalable approach to comparing execution traces
ESEC-FSE companion '07: The 6th Joint Meeting on European software engineering conference and the ACM SIGSOFT symposium on the foundations of software engineering: companion papers

In this paper, we overview a new approach to comparing execution traces. Such comparison can be useful for purposes such as improving test coverage and profiling system's users. In our approach, traces are compressed into different levels of compaction ...

Comments

Information & Contributors

Information

Published In

cover image ACM Transactions on Architecture and Code Optimization

ACM Transactions on Architecture and Code Optimization Volume 14, Issue 3

September 2017

278 pages

ISSN:1544-3566

EISSN:1544-3973

DOI:10.1145/3132652

Editor:
Koen De Bosschere
Ghent University

Issue’s Table of Contents

Copyright © 2017 ACM.

Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than ACM must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. Request permissions from [email protected]

Publisher

Association for Computing Machinery

New York, NY, United States

Publication History

Published: 30 August 2017

Accepted: 01 May 2017

Revised: 01 March 2017

Received: 01 October 2016

Published in TACO Volume 14, Issue 3

Permissions

Request permissions for this article.

Request Permissions

Check for updates

Author Tags

Qualifiers

Research-article
Research
Refereed

Contributors

Other Metrics

View Article Metrics

Bibliometrics & Citations

Bibliometrics

Article Metrics

1
Total Citations
View Citations
500
Total Downloads

Downloads (Last 12 months)143
Downloads (Last 6 weeks)10

Reflects downloads up to 22 Sep 2024

Other Metrics

View Author Metrics

Citations

Cited By

Brais HPanda P(2019)AlleriaACM Transactions on Embedded Computing Systems10.1145/335819318:5s(1-22)Online publication date: 8-Oct-2019
https://dl.acm.org/doi/10.1145/3358193

View Options

View options

PDF

View or Download as a PDF file.

eReader

View online with eReader.

Get Access

Login options

Check if you have access through your login credentials or your institution to get full access on this article.

Full Access

Get this Article

Media

Figures

Other

Tables

View Issue’s Table of Contents