TraceRTL: Agile Performance Evaluation for Microarchitecture Exploration

Authors

  • Zifei Zhang State Key Lab of Processors, Institute of Computing Technology, Chinese Academy of Sciences, 100190, Beijing, China , University of Chinese Academy of Sciences, 100190, Beijing, China Author
  • Yinan Xu State Key Lab of Processors, Institute of Computing Technology, Chinese Academy of Sciences, 100190, Beijing, China Author
  • Kaichen Gong School of Information Science and Technology, ShanghaiTech University, 200000, Shanghai, China Author
  • Sa Wang State Key Lab of Processors, Institute of Computing Technology, Chinese Academy of Sciences, 100190, Beijing, China , University of Chinese Academy of Sciences, 100190, Beijing, China Author
  • Dan Tang State Key Lab of Processors, Institute of Computing Technology, Chinese Academy of Sciences, 100190, Beijing, China , Beijing Institute of Open Source Chip, 100080, Beijing, China Author
  • Yungang Bao State Key Lab of Processors, Institute of Computing Technology, Chinese Academy of Sciences, 100190, Beijing, China , University of Chinese Academy of Sciences, 100190, Beijing, China Author

DOI:

https://doi.org/10.66834/5s84yp08

Keywords:

Trace-driven simulation, Performance evaluation, Cross ISA benchmarking

Abstract

While agile chip development methodologies have accelerated RTL design and simulation, performance evaluation remains constrained by challenges: 
(1) limited benchmarks availability due to incomplete peripheral/software simulation environments or unavailable source code;
(2) inefficient feature prototyping caused by the tight coupling between functional correctness and performance evaluation, particularly for large‑scale, error‑prone microarchitectures.
To address these challenges, we propose TraceRTL, an agile, trace-driven performance evaluation methodology that decouples the functional and performance components of CPU RTL designs.
It introduces three contributions to benchmarking community: (1) a trace‑driven exploration framework that bypasses full functional correctness while preserving performance behavior and supports to replay workload traces on RTL designs; (2) a quantitative analysis and mitigation methodology to identify and reduce trace-driven performance discrepancies; (3) a trace transformation technique, TraceBridge, that replays benchmark traces across different formats and instruction sets.
Using TraceRTL, we develop the first trace-driven RTL CPU derived from XiangShan, a high-performance out-of-order RISC-V processor.
TraceRTL achieves performance accuracies of 99.87% and 99.86% on SPECint2017 and SPECfp2017, respectively.
With TraceBridge, we evaluate x86-based Google workload traces on a RISC-V RTL CPU and reveal distinct memory-bound behavior.

References

1. D. Burger and T. M. Austin, “The simplescalar tool set, version 2.0,” SIGARCH Comput. Archit. News, vol. 25, no. 3, p. 13–25, Jun. 1997, doi:https://doi.org/10.1145/268806.268810. [Online]. Available: https://doi.org/10.1145/268806.268810 DOI: https://doi.org/10.1145/268806.268810

2. M. M. K. Martin et al., “Multifacet’s general execution-driven multiprocessor simulator (gems) toolset,” SIGARCH Comput. Archit. News, vol. 33, no. 4, p. 92–99, Nov. 2005, doi:https://doi.org/10.1145/1105734.1105747. [Online]. Available: https://doi.org/10.1145/1105734.1105747 DOI: https://doi.org/10.1145/1105734.1105747

3. N. Binkert, R. Dreslinski, L. Hsu, K. Lim, A. Saidi, and S. Reinhardt, “The m5 simulator: Modeling networked systems,” IEEE Micro, vol. 26, no. 4, pp. 52–60, 2006, doi:https://doi.org/10.1109/MM.2006.82. DOI: https://doi.org/10.1109/MM.2006.82

4. N. Binkert et al., “The gem5 simulator,” SIGARCH Comput. Archit. News, vol. 39, no. 2, p. 1–7, Aug. 2011, doi:https://doi.org/10.1145/2024716.2024718. [Online]. Available: https://doi.org/10.1145/2024716.2024718 DOI: https://doi.org/10.1145/2024716.2024718

5. T. E. Carlson, W. Heirman, and L. Eeckhout, “Sniper: exploring the level of abstraction for scalable and accurate parallel multi-core simulation,” in Proceedings of 2011 International Conference for High Performance Computing, Networking, Storage and Analysis, ser. SC ’11. New York, NY, USA: Association for Computing Machinery, 2011, doi:https://doi.org/10.1145/2063384.2063454. [Online]. Available: https://doi.org/10.1145/2063384.2063454 DOI: https://doi.org/10.1145/2063384.2063454

6. D. Sanchez and C. Kozyrakis, “Zsim: fast and accurate microarchitectural simulation of thousand-core systems,” in Proceedings of the 40th Annual International Symposium on Computer Architecture, ser. ISCA ’13. New York, NY, USA: Association for Computing Machinery, 2013, p. 475–486, doi:https://doi.org/10.1145/2485922.2485963. [Online]. Available: https://doi.org/10.1145/2485922.2485963 DOI: https://doi.org/10.1145/2485922.2485963

7. N. Gober et al., “The championship simulator: Architectural simulation for education and competition,” 2022. [Online]. Available: https://arxiv.org/abs/2210.14324

8. H. Golestani, R. Sen, V. Young, and G. Gupta, “Calipers: a criticality-aware framework for modeling processor performance,” in Proceedings of the 36th ACM International Conference on Supercomputing, ser. ICS ’22. New York, NY, USA: Association for Computing Machinery, 2022, doi:https://doi.org/10.1145/3524059.3532390. [Online]. Available: https://doi.org/10.1145/3524059.3532390 DOI: https://doi.org/10.1145/3524059.3532390

9. T. Nowatzki, J. Menon, C.-H. Ho, and K. Sankaralingam, “Architectural simulators considered harmful,” IEEE Micro, vol. 35, no. 6, pp. 4–12, 2015, doi:https://doi.org/10.1109/MM.2015.74. DOI: https://doi.org/10.1109/MM.2015.74

10. “Cbp2025 simulator framework,” https://ericrotenberg.wordpress.ncsu.edu/cbp2025-simulator-framework/, 2025.

11. S. Zhang, A. Wright, T. Bourgeat, and Arvind, “Composable building blocks to open up processor design,” in Proceedings of the 51st Annual IEEE/ACM International Symposium on Microarchitecture, ser. MICRO-51. IEEE Press, 2018, p. 68–81, doi:https://doi.org/10.1109/MICRO.2018.00015. [Online]. Available: https://doi.org/10.1109/MICRO.2018.00015 DOI: https://doi.org/10.1109/MICRO.2018.00015

12. T. Bourgeat, C. Pit-Claudel, A. Chlipala, and Arvind, “The essence of bluespec: a core language for rule-based hardware design,” in Proceedings of the 41st ACM SIGPLAN Conference on Programming Language Design and Implementation, ser. PLDI 2020. New York, NY, USA: Association for Computing Machinery, 2020, p. 243–257, doi:https://doi.org/10.1145/3385412.3385965. [Online]. Available: https://doi.org/10.1145/3385412.3385965 DOI: https://doi.org/10.1145/3385412.3385965

13. J. Bachrach et al., “Chisel: constructing hardware in a scala embedded language,” in Proceedings of the 49th Annual Design Automation Conference, 2012, pp. 1216–1225, doi:https://doi.org/10.1145/2228360.2228584. [Online]. Available: https://doi.org/10.1145/2228360.2228584 DOI: https://doi.org/10.1145/2228360.2228584

14. Verilator, “Verilator user’s guide,” https://www.veripool.org/guide/latest/, 2026.

15. H. Wang and S. Beamer, “Repcut: Superlinear parallel rtl simulation with replication-aided partitioning,” in Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 3, ser. ASPLOS 2023. New York, NY, USA: Association for Computing Machinery, 2023, p. 572–585, doi:https://doi.org/10.1145/3582016.3582034. [Online]. Available: https://doi.org/10.1145/3582016.3582034 DOI: https://doi.org/10.1145/3582016.3582034

16. K. Zhou, Y. Liang, Y. Lin, R. Wang, and R. Huang, “Khronos: Fusing memory access for improved hardware rtl simulation,” in Proceedings of the 56th Annual IEEE/ACM International Symposium on Microarchitecture, ser. MICRO ’23. New York, NY, USA: Association for Computing Machinery, 2023, p. 180–193, doi:https://doi.org/10.1145/3613424.3614301. [Online]. Available: https://doi.org/10.1145/3613424.3614301 DOI: https://doi.org/10.1145/3613424.3614301

17. H. Wang, T. Nijssen, and S. Beamer, “Don’t repeat yourself! coarse-grained circuit deduplication to accelerate rtl simulation,” in Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 4, ser. ASPLOS ’24. New York, NY, USA: Association for Computing Machinery, 2025, p. 79–93, doi:https://doi.org/10.1145/3622781.3674184. [Online]. Available: https://doi.org/10.1145/3622781.3674184 DOI: https://doi.org/10.1145/3622781.3674184

18. M. Emami, T. Bourgeat, and J. R. Larus, “Parendi: Thousand-way parallel rtl simulation,” in Proceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2, ser. ASPLOS ’25. New York, NY, USA: Association for Computing Machinery, 2025, p. 783–797, doi:https://doi.org/10.1145/3676641.3716010. [Online]. Available: https://doi.org/10.1145/3676641.3716010 DOI: https://doi.org/10.1145/3676641.3716010

19. S. Karandikar et al., “Firesim: Fpga-accelerated cycle-exact scale-out system simulation in the public cloud,” in Proceedings of the 45th Annual International Symposium on Computer Architecture, ser. ISCA ’18. IEEE Press, 2018, p. 29–42, doi:https://doi.org/10.1109/ISCA.2018.00014. [Online]. Available: https://doi.org/10.1109/ISCA.2018.00014 DOI: https://doi.org/10.1109/ISCA.2018.00014

20. ——, “Fireperf: Fpga-accelerated full-system hardware/software performance profiling and co-design,” in Proceedings of the Twenty-Fifth International Conference on Architectural Support for Programming Languages and Operating Systems, ser. ASPLOS ’20. New York, NY, USA: Association for Computing Machinery, 2020, p. 715–731, doi:https://doi.org/10.1145/3373376.3378455. [Online]. Available: https://doi.org/10.1145/3373376.3378455 DOI: https://doi.org/10.1145/3373376.3378455

21. M. Emami, S. Kashani, K. Kamahori, M. S. Pourghannad, R. Raj, and J. R. Larus, “Manticore: Hardware-accelerated rtl simulation with static bulk-synchronous parallelism,” in Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 4, 2023, pp. 219–237, doi:https://doi.org/10.1145/3623278.3624750. DOI: https://doi.org/10.1145/3623278.3624750

22. F. Elsabbagh, S. Sheikhha, V. A. Ying, Q. M. Nguyen, J. S. Emer, and D. Sanchez, “Accelerating rtl simulation with hardware-software co-design,” in Proceedings of the 56th Annual IEEE/ACM International Symposium on Microarchitecture, 2023, pp. 153–166, doi:https://doi.org/10.1145/3613424.3614257. DOI: https://doi.org/10.1145/3613424.3614257

23. C. Celio, D. A. Patterson, and K. Asanovic, “The berkeley out-of-order machine (boom): An industry-competitive, synthesizable, parameterized risc-v processor,” EECS Department, University of California, Berkeley, Tech. Rep. UCB/EECS-2015-167, 2015. [Online]. Available: https://www2.eecs.berkeley.edu/Pubs/TechRpts/2015/EECS-2015-167.html

24. J. Zhao, B. Korpan, A. Gonzalez, and K. Asanovic, “Sonicboom: The 3rd generation berkeley out-of-order machine,” in Fourth Workshop on Computer Architecture Research with RISC-V, vol. 5, 2020, pp. 1–7. [Online]. Available: https://people.eecs.berkeley.edu/∼krste/papers/SonicBOOM-CARRV2020.pdf

25. C. Celio, P.-F. Chiu, B. Nikolic, D. A. Patterson, and K. Asanovic, “BOOMv2: an open-source out-of-order RISC-V core,” in First Workshop on Computer Architecture Research with RISC-V (CARRV), 2017. [Online]. Available: https://www2.eecs.berkeley.edu/Pubs/TechRpts/2017/EECS-2017-157.pdf

26. K. Wang et al., “ XiangShan: An Open-Source Project for High-Performance RISC-V Processors Meeting Industrial-Grade Standards ,” in 2024 IEEE Hot Chips 36 Symposium (HCS). Los Alamitos, CA, USA: IEEE Computer Society, Aug. 2024, pp. 1–25, doi:https://doi.org/10.1109/HCS61935.2024.10665293. [Online]. Available: https://doi.ieeecomputersociety.org/10.1109/HCS61935.2024.10665293 DOI: https://doi.org/10.1109/HCS61935.2024.10665293

27. C. Chen et al., “Xuantie-910: A commercial multi-core 12stage pipeline out-of-order 64-bit high performance risc-v processor with vector extension: Industrial product,” in 2020 ACM/IEEE 47th Annual International Symposium on Computer Architecture (ISCA). IEEE, 2020, pp. 52– 64, doi:https://doi.org/10.1109/ISCA45697.2020.00016. DOI: https://doi.org/10.1109/ISCA45697.2020.00016

28. C. Bai, Q. Sun, J. Zhai, Y. Ma, B. Yu, and M. D. Wong, “Boom-explorer: Risc-v boom microarchitecture design space exploration framework,” in 2021 IEEE/ACM International Conference On Computer Aided Design (ICCAD). IEEE, 2021, pp. 1–9, doi:https://doi.org/10.1109/ICCAD51958.2021.9643455. DOI: https://doi.org/10.1109/ICCAD51958.2021.9643455

29. S. Gupta et al., “Imprecise store exceptions,” in Proceedings of the 50th Annual International Symposium on Computer Architecture, ser. ISCA ’23. New York, NY, USA: Association for Computing Machinery, 2023, doi:https://doi.org/10.1145/3579371.3589087. [Online]. Available: https://doi.org/10.1145/3579371.3589087 DOI: https://doi.org/10.1145/3579371.3589087

30. M. Ghaniyoun, K. Barber, Y. Xiao, Y. Zhang, and R. Teodorescu, “Teesec: Pre-silicon vulnerability discovery for trusted execution environments,” in Proceedings of the 50th Annual International Symposium on Computer Architecture, ser. ISCA ’23. New York, NY, USA: Association for Computing Machinery, 2023, doi:https://doi.org/10.1145/3579371.3589070. [Online]. Available: https://doi.org/10.1145/3579371.3589070 DOI: https://doi.org/10.1145/3579371.3589070

31. L. Wang et al., “Asynchronous memory access unit: Exploiting massive parallelism for far memory access,” ACM Trans. Archit. Code Optim., vol. 21, no. 3, Sep. 2024, doi:https://doi.org/10.1145/3663479. [Online]. Available: https://doi.org/10.1145/3663479 DOI: https://doi.org/10.1145/3663479

32. D. Wang et al., “A transfer learning framework for high-accurate cross-workload design space exploration of cpu,” in 2023 IEEE/ACM International Conference on Computer Aided Design (ICCAD), 2023, pp. 1–9, doi:https://doi.org/10.1109/ICCAD57390.2023.10323840. DOI: https://doi.org/10.1109/ICCAD57390.2023.10323840

33. H. Ando, “Performance improvement by prioritizing the issue of the instructions in unconfident branch slices,” in 2018 51st Annual IEEE/ACM International Symposium on Microarchitecture (MICRO), 2018, pp. 82–94, doi:https://doi.org/10.1109/MICRO.2018.00016. DOI: https://doi.org/10.1109/MICRO.2018.00016

34. Y. Xu et al., “Towards developing high performance risc-v processors using agile methodology,” in 2022 55th IEEE/ACM International Symposium on Microarchitecture (MICRO), 2022, pp. 1178–1199, doi:https://doi.org/10.1109/MICRO56248.2022.00080. DOI: https://doi.org/10.1109/MICRO56248.2022.00080

35. “Championship value prediction,” https://microarch.org/cvp1/, accessed: 2025-02-20.

36. “Google workload traces version 2,” https://console.cloud.google.com/storage/browser/external-traces-v2, accessed: 2025-02-20.

37. W. Su et al., “Dcperf: An open-source, battle-tested performance benchmark suite for datacenter workloads,” in Proceedings of the 52nd Annual International Symposium on Computer Architecture, ser. ISCA ’25. New York, NY, USA: Association for Computing Machinery, 2025, p. 1717–1730, doi:https://doi.org/10.1145/3695053.3731411. [Online]. Available: https://doi.org/10.1145/3695053.3731411 DOI: https://doi.org/10.1145/3695053.3731411

38. OpenXiangShan, “XiangShan,” https://github.com/OpenXiangShan/XiangShan, 2020.

39. SpinalHDL, “Scala based hdl,” https://github.com/SpinalHDL/SpinalHDL, 2024.

40. T. Sherwood, E. Perelman, G. Hamerly, and B. Calder, “Automatically characterizing large scale program behavior,” in Proceedings of the 10th International Conference on Architectural Support for Programming Languages and Operating Systems, ser. ASPLOS X. New York, NY, USA: Association for Computing Machinery, 2002, p. 45–57, doi:https://doi.org/10.1145/605397.605403. [Online]. Available: https://doi.org/10.1145/605397.605403 DOI: https://doi.org/10.1145/605397.605403

41. A. Sabu, H. Patil, W. Heirman, and T. E. Carlson, “Looppoint: Checkpoint-driven sampled simulation for multi-threaded applications,” in 2022 IEEE International Symposium on High-Performance Computer Architecture (HPCA). IEEE, 2022, pp. 604–618, doi:https://doi.org/10.1109/HPCA53966.2022.00051. DOI: https://doi.org/10.1109/HPCA53966.2022.00051

42. T. E. Carlson, W. Heirman, K. Van Craeynest, and L. Eeckhout, “Barrierpoint: Sampled simulation of multithreaded applications,” in 2014 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS). IEEE, 2014, pp. 2–12, doi:https://doi.org/10.1109/ISPASS.2014.6844456. DOI: https://doi.org/10.1109/ISPASS.2014.6844456

43. K. Asanovic et al., “The rocket chip generator,” EECS Department, University of California, Berkeley, Tech. Rep. UCB/EECS-2016-17, vol. 4, pp. 6–2, 2016. [Online]. Available: https://aspire.eecs.berkeley.edu/wp/wp-content/uploads/2016/04/Tech-Report-The-Rocket-Chip-Generator-Beamer.pdf

44. B. S´a, L. Valente, J. Martins, D. Rossi, L. Benini, and S. Pinto, “CVA6 RISC-V virtualization: Architecture, microarchitecture, and design space exploration,” IEEE Transactions on Very Large Scale Integration (VLSI) Systems, 2023, doi:https://doi.org/10.1109/TVLSI.2023.3302837. DOI: https://doi.org/10.1109/TVLSI.2023.3302837

45. R.-V. community, “Olympia,” https://github.com/riscv-software-src/riscv-perf-model, 2026.

46. C.-K. Luk et al., “Pin: building customized program analysis tools with dynamic instrumentation,” in Proceedings of the 2005 ACM SIGPLAN Conference on Programming Language Design and Implementation, ser. PLDI ’05. New York, NY, USA: Association for Computing Machinery, 2005, p. 190–200, doi:https://doi.org/10.1145/1065010.1065034. [Online]. Available: https://doi.org/10.1145/1065010.1065034 DOI: https://doi.org/10.1145/1065010.1065034

47. D. Bruening, E. Duesterwald, and S. Amarasinghe, “Design and implementation of a dynamic optimization framework for windows,” in 4th ACM workshop on feedback-directed and dynamic optimization (FDDO-4), 2001, p. 20.

48. N. Nethercote and J. Seward, “Valgrind: a framework for heavyweight dynamic binary instrumentation,” in Proceedings of the 28th ACM SIGPLAN Conference on Programming Language Design and Implementation, ser. PLDI ’07. New York, NY, USA: Association for Computing Machinery, 2007, p. 89–100, doi:https://doi.org/10.1145/1250734.1250746. [Online]. Available: https://doi.org/10.1145/1250734.1250746 DOI: https://doi.org/10.1145/1250734.1250746

49. DynamoRIO, “Dynamorio trace format,” https://dynamorio.org/sec drcachesim format.html, accessed: 2026-02-07.

50. F. Bellard, “QEMU, a fast and portable dynamic translator.” in USENIX annual technical conference, FREENIX Track, vol. 41, no. 46. California, USA, 2005, pp. 10–5555. [Online]. Available: https://www.usenix.org/legacy/event/usenix05/tech/freenix/full papers/bellard/bellard.pdf

51. S. Pandey, A. Yazdanbakhsh, and H. Liu, “Tao: Re-thinking dl-based microarchitecture simulation,” Proceedings of the ACM on Measurement and Analysis of Computing Systems, vol. 8, no. 2, pp. 1–25, 2024, doi:https://doi.org/10.1145/3656012. DOI: https://doi.org/10.1145/3656012

52. M. E. S. Elrabaa, A. Hroub, M. F. Mudawar, A. AlAghbari, M. Al-Asli, and A. Khayyat, “A very fast trace-driven simulation platform for chip-multiprocessors architectural explorations,” IEEE Transactions on Parallel and Distributed Systems, vol. 28, no. 11, pp. 3033–3045, 2017, doi:https://doi.org/10.1109/TPDS.2017.2713782. DOI: https://doi.org/10.1109/TPDS.2017.2713782

53. OpenXiangShan, “XS-gem5,” https://github.com/OpenXiangShan/GEM5, 2020.

54. J. L. Henning, “Spec cpu2006 benchmark descriptions,” ACM SIGARCH Computer Architecture News, vol. 34, no. 4, pp. 1–17, 2006, doi:https://doi.org/10.1145/1186736.1186737. DOI: https://doi.org/10.1145/1186736.1186737

55. J. Bucek, K.-D. Lange, and J. v. Kistowski, “SPEC CPU2017: Next-generation compute benchmark,” in Companion of the 2018 ACM/SPEC International Conference on Performance Engineering, 2018, pp. 41–42, doi:https://doi.org/10.1145/3185768.3185771. DOI: https://doi.org/10.1145/3185768.3185771

56. OpenXiangShan, “NEMU,” https://github.com/OpenXiangShan/NEMU, 2019.

57. A. Yasin, “A top-down method for performance analysis and counters architecture,” in 2014 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS), 2014, pp. 35–44, doi:https://doi.org/10.1109/ISPASS.2014.6844459. DOI: https://doi.org/10.1109/ISPASS.2014.6844459

58. A. Karpathy, “llama2.c: Inference Llama 2 in one file of pure C,” https://github.com/karpathy/llama2.c, accessed: 2026-02-07.

59. RISC-V, “RISC-V Instruction Set Manual,” https://github.com/riscv/riscv-isa-manual, accessed: 2026-02-07.

60. J. D. Bruguera, “Low-latency and high-bandwidth pipelined radix-64 division and square root unit,” in 2022 IEEE 29th Symposium on Computer Arithmetic (ARITH), 2022, pp. 10–17, doi:https://doi.org/10.1109/ARITH54963.2022.00012. DOI: https://doi.org/10.1109/ARITH54963.2022.00012

61. V. M. Vedula, J. A. Abraham, J. Bhadra, and R. Tupuri, “A hierarchical test generation approach using program slicing techniques on hardware description languages,” Journal of Electronic Testing, vol. 19, pp. 149–160, 2003, doi:https://doi.org/10.1023/A:1022885523034. DOI: https://doi.org/10.1023/A:1022885523034

62. L. Liu and S. Vasudevan, “Efficient validation input generation in rtl by hybridized source code analysis,” in 2011 Design, Automation & Test in Europe, 2011, pp. 1–6, doi:https://doi.org/10.1109/DATE.2011.5763253. DOI: https://doi.org/10.1109/DATE.2011.5763253

63. B. Mammo, J. Larimer, M. Morgan, D. Fan, E. Hennenhoefer, and V. Bertacco, “Architectural trace-based functional coverage for multiprocessor verification,” in 2012 13th International Workshop on Microprocessor Test and Verification (MTV), 2012, pp. 1–5, doi:https://doi.org/10.1109/MTV.2012.12. DOI: https://doi.org/10.1109/MTV.2012.12

64. S. Apostolakis, C. Kennelly, X. D. Li, and P. Ranganathan, “Necro-reaper: Pruning away dead memory traffic in warehouse-scale computers,” in Proceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2, ser. ASPLOS ’25. New York, NY, USA: Association for Computing Machinery, 2025, p. 689–703, doi:https://doi.org/10.1145/3676641.3716007. [Online]. Available: https://doi.org/10.1145/3676641.3716007 DOI: https://doi.org/10.1145/3676641.3716007

65. M. Khairy, Z. Shen, T. M. Aamodt, and T. G. Rogers, “Accel-sim: An extensible simulation framework for validated gpu modeling,” in 2020 ACM/IEEE 47th Annual International Symposium on Computer Architecture (ISCA), 2020, pp. 473–486, doi:https://doi.org/10.1109/ISCA45697.2020.00047. DOI: https://doi.org/10.1109/ISCA45697.2020.00047

66. A. Bakhoda, G. L. Yuan, W. W. L. Fung, H. Wong, and T. M. Aamodt, “Analyzing cuda workloads using a detailed gpu simulator,” in 2009 IEEE International Symposium on Performance Analysis of Systems and Software, 2009, pp. 163–174, doi:https://doi.org/10.1109/ISPASS.2009.4919648. DOI: https://doi.org/10.1109/ISPASS.2009.4919648

67. O. Mutlu, H. Kim, D. N. Armstrong, and Y. N. Patt, “An analysis of the performance impact of wrong-path memory references on out-of-order and runahead execution processors,” IEEE Transactions on Computers, vol. 54, no. 12, pp. 1556–1571, 2005, doi:https://doi.org/10.1109/TC.2005.190. DOI: https://doi.org/10.1109/TC.2005.190

68. S. Eyerman, S. Van den Steen, W. Heirman, and I. Hur, “Simulating wrong-path instructions in decoupled functional-first simulation,” in 2023 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS). IEEE, 2023, pp. 124–133, doi:https://doi.org/10.1109/ISPASS57527.2023.00021. DOI: https://doi.org/10.1109/ISPASS57527.2023.00021

69. B. R. Godala et al., “Correct wrong path,” arXiv preprint arXiv:2408.05912, 2024. [Online]. Available: https://doi.org/10.48550/arXiv.2408.05912

70. R. Sendag, A. Yilmazer, J. J. Yi, and A. K. Uht, “The impact of wrong-path memory references in cache-coherent multiprocessor systems,” Journal of Parallel and Distributed Computing, vol. 67, no. 12, pp. 1256–1269, 2007, best Paper Awards: 20th International Parallel and Distributed Processing Symposium (IPDPS 2006). doi:https://doi.org/https://doi.org/10.1016/j.jpdc.2007.03.005. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S0743731507000457 DOI: https://doi.org/10.1016/j.jpdc.2007.03.005

71. R. Sendag, A. Yilmazer, J. Yi, and A. Uht, “Quantifying and reducing the effects of wrong-path memory references in cache-coherent multiprocessor systems,” in Proceedings 20th IEEE International Parallel & Distributed Processing Symposium, 2006, pp. 10 pp.–, doi:https://doi.org/10.1109/IPDPS.2006.1639260. DOI: https://doi.org/10.1109/IPDPS.2006.1639260

72. S. R. Goldschmidt and J. L. Hennessy, “The accuracy of trace-driven simulations of multiprocessors,” ACM SIGMETRICS Performance Evaluation Review, vol. 21, no. 1, pp. 146–157, 1993, doi:https://doi.org/10.1145/166962.167001. DOI: https://doi.org/10.1145/166962.167001

73. K. Sangaiah et al., “Synchrotrace: Synchronization-aware architecture-agnostic traces for lightweight multicore simulation of cmp and hpc workloads,” ACM Trans. Archit. Code Optim., vol. 15, no. 1, Mar. 2018, doi:https://doi.org/10.1145/3158642. [Online]. Available: https://doi.org/10.1145/3158642 DOI: https://doi.org/10.1145/3158642

74. J. Weng et al., “Assassyn: A unified abstraction for architectural simulation and implementation,” in Proceedings of the 52nd Annual International Symposium on Computer Architecture, ser. ISCA ’25. New York, NY, USA: Association for Computing Machinery, 2025, p. 1464–1479, doi:https://doi.org/10.1145/3695053.3731004. [Online]. Available: https://doi.org/10.1145/3695053.3731004 DOI: https://doi.org/10.1145/3695053.3731004

75. J. Feliu, A. Perais, D. A. Jim´enez, and A. Ros, “Rebasing microarchitectural research with industry traces,” in 2023 IEEE International Symposium on Workload Characterization (IISWC), 2023, pp. 100–114, doi:https://doi.org/10.1109/IISWC59245.2023.00027. DOI: https://doi.org/10.1109/IISWC59245.2023.00027

approaches

Additional Files

Published

2026-04-12

Issue

Section

Full Length Articles/Research articles

How to Cite

TraceRTL: Agile Performance Evaluation for Microarchitecture Exploration. (2026). BenchCouncil Transactions on Benchmarks, Standards and Evaluations, 6. https://doi.org/10.66834/5s84yp08