
Z. Zhang et al.
3302837.
45. R.-V. community, “Olympia,” https://github.com/
riscv-software-src/riscv- perf-model, 2026.
46. C.-K. Luk et al., “Pin: building customized program
analysis tools with dynamic instrumentation,” in Pro-
ceedings of the 2005 ACM SIGPLAN Conference on
Programming Language Design and Implementation,
ser. PLDI ’05. New York, NY, USA: Associa-
tion for Computing Machinery, 2005, p. 190–200,
doi: https://doi.org/10.1145/1065010.1065034. [Online].
Available: https://doi.org/10.1145/1065010.1065034
47. D. Bruening, E. Duesterwald, and S. Amarasinghe, “Design
and implementation of a dynamic optimization framework
for windows,” in 4th ACM workshop on feedback-directed
and dynamic optimization (FDDO-4), 2001, p. 20.
48. N. Nethercote and J. Seward, “Valgrind: a framework
for heavyweight dynamic binary instrumentation,” in
Proceedings of the 28th ACM SIGPLAN Conference
on Programming Language Design and Implementation,
ser. PLDI ’07. New York, NY, USA: Association for
Computing Machinery, 2007, p. 89–100, doi: https:
//doi.org/10.1145/1250734.1250746. [Online]. Available:
https://doi.org/10.1145/1250734.1250746
49. DynamoRIO, “Dynamorio trace format,” https:
//dynamorio.org/sec drcachesim format.html, accessed:
2026-02-07.
50. F. Bellard, “QEMU, a fast and p ortable dynamic transla-
tor.” in USENIX annual technical conference, FREENIX
Track, vol. 41, no. 46. California, USA, 2005, pp. 10–5555.
[Online]. Available: https://www.usenix.org/legacy/event/
usenix05/tech/freenix/full
papers/bellard/bellard.pdf
51. S. Pandey, A. Yazdanbakhsh, and H. Liu, “Tao: Re-thinking
dl-based microarchitecture simulation,” Proceedings of the
ACM on Measurement and Analysis of Computing
Systems, vol. 8, no. 2, pp. 1–25, 2024, doi: https:
//doi.org/10.1145/3656012.
52. M. E. S. Elrabaa, A. Hroub, M. F. Mudawar, A. Al-
Aghbari, M. Al-Asli, and A. Khayyat, “A very fast
trace-driven simulation platform for chip-multiprocessors
architectural explorations,” IEEE Transactions on Parallel
and Distributed Systems, vol. 28, no. 11, pp. 3033–3045,
2017, doi: https://doi.org/10.1109/TPDS.2017.2713782.
53. OpenXiangShan, “XS-gem5,” https://github.com/
OpenXiangShan/GEM5, 2020.
54. J. L. Henning, “Spec cpu2006 benchmark descriptions,”
ACM SIGARCH Computer Architecture News, vol. 34,
no. 4, pp. 1–17, 2006, doi: https://doi.org/10.1145/
1186736.1186737.
55. J. Bucek, K.-D. Lange, and J. v. Kistowski, “SPEC
CPU2017: Next-generation compute benchmark,” in
Companion of the 2018 ACM/SPEC International
Conference on Performance Engineering, 2018, pp. 41–42,
doi: https://doi.org/10.1145/3185768.3185771.
56. OpenXiangShan, “NEMU,” https://github.com/
OpenXiangShan/NEMU, 2019.
57. A. Yasin, “A top-down method for performance analysis
and counters architecture,” in 2014 IEEE International
Symposium on Performance Analysis of Systems and
Software (ISPASS), 2014, pp. 35–44, doi: https:
//doi.org/10.1109/ISPASS.2014.6844459.
58. A. Karpathy, “llama2.c: Inference Llama 2 in one file of
pure C,” https://github.com/karpathy/llama2.c, accessed:
2026-02-07.
59. RISC-V, “RISC-V Instruction Set Manual,” https://github.
com/riscv/riscv-isa- manual, accessed: 2026-02-07.
60. J. D. Bruguera, “Low-latency and high-bandwidth
pipelined radix-64 division and square root unit,”
in 2022 IEEE 29th Symposium on Computer
Arithmetic (ARITH), 2022, pp. 10–17, doi:
https://doi.org/10.1109/ARITH54963.2022.00012.
61. V. M. Vedula, J. A. Abraham, J. Bhadra, and R. Tupuri,
“A hierarchical test generation approach using program
slicing techniques on hardware description languages,”
Journal of Electronic Testing, vol. 19, pp. 149–160, 2003,
doi: https://doi.org/10.1023/A:1022885523034.
62. L. Liu and S. Vasudevan, “Efficient validation input
generation in rtl by hybridized source code analysis,” in
2011 Design, Automation & Test in Europe, 2011, pp.
1–6, doi: https://doi.org/10.1109/DATE.2011.5763253.
63. B. Mammo, J. Larimer, M. Morgan, D. Fan, E. Hen-
nenhoefer, and V. Bertacco, “Architectural trace-based
functional coverage for multiprocessor verification,” in
2012 13th International Workshop on Microprocessor
Test and Verification (MTV), 2012, pp. 1–5, doi:
https://doi.org/10.1109/MTV.2012.12.
64. S. Apostolakis, C. Kennelly, X. D. Li, and P. Ranganathan,
“Necro-reaper: Pruning away dead memory traffic in
warehouse-scale computers,” in Proceedings of the 30th
ACM International Conference on Architectural Support
for Programming Languages and Operating Systems,
Volume 2, ser. ASPLOS ’25. New York, NY, USA:
Association for Computing Machinery, 2025, p. 689–703,
doi: https://doi.org/10.1145/3676641.3716007. [Online].
Available: https://doi.org/10.1145/3676641.3716007
65. M. Khairy, Z. Shen, T. M. Aamodt, and T. G.
Rogers, “Accel-sim: An extensible simulation framework
for validated gpu modeling,” in 2020 ACM/IEEE
47th Annual International Symposium on Computer
Architecture (ISCA), 2020, pp. 473–486, doi: https:
//doi.org/10.1109/ISCA45697.2020.00047.
66. A. Bakhoda, G. L. Yuan, W. W. L. Fung, H. Wong, and
T. M. Aamodt, “Analyzing cuda workloads using a detailed
gpu simulator,” in 2009 IEEE International Symposium
on Performance Analysis of Systems and Software, 2009,
pp. 163–174, doi: https://doi.org/10.1109/ISPASS.2009.
4919648.
67. O. Mutlu, H. Kim, D. N. Armstrong, and Y. N.
Patt, “An analysis of the p erformance impact of wrong-
path memory references on out-of-order and runahead
execution pro cessors,” IEEE Transactions on Computers,
vol. 54, no. 12, pp. 1556–1571, 2005, doi: https:
//doi.org/10.1109/TC.2005.190.
68. S. Eyerman, S. Van den Steen, W. Heirman, and
I. Hur, “Simulating wrong-path instructions in decoupled
functional-first simulation,” in 2023 IEEE International
Symposium on Performance Analysis of Systems and
Software (ISPASS). IEEE, 2023, pp. 124–133, doi:
https://doi.org/10.1109/ISPASS57527.2023.00021.
69. B. R. Godala et al., “Correct wrong path,” arXiv
preprint arXiv:2408.05912, 2024. [Online]. Available:
https://doi.org/10.48550/arXiv.2408.05912
70. R. Sendag, A. Yilmazer, J. J. Yi, and A. K. Uht,
“The impact of wrong-path memory references in cache-
coherent multiprocessor systems,” Journal of Parallel and
Distributed Computing, vol. 67, no. 12, pp. 1256–1269,
2007, best Paper Awards: 20th International Parallel
and Distributed Processing Symposium (IPDPS 2006).
16