Assembly Hall of Shame

hackernews

Assembly Hall of Shame

Overview

Instruction latency analysis usually focuses on performance optimization—making code run as fast as possible. The Assembly Hall of Shame takes the opposite approach: searching for the absolute floor of single-instruction performance.

🏆 Current Champions 🏆

Strategy: Use fxrstor64 to load 512-byte FPU/MMX/XMM state from a high-latency MMIO region in the PCIe fabric, then starve the fabric while the load is in flight — a fleet of hammer cores pounds a different high-latency MMIO register with tight 4-byte reads, saturating the PCIe root complex and endpoint with non-posted transactions, so CPU 0's 512-byte fxrstor64 must queue behind all that contending traffic.

Contender: AMD Ryzen 7 5800H

; CPU 0 — timed instruction
movl $0xfcc68830, %rsi
fxrstor64 %rsi

; CPUs 1..N — hammer loop against a different high-latency location
movl 0xfcc68858, %eax

🏆 Score: 198,002,498,236 cycles

🏆 Time: 62 seconds

Honorable Mentions

A spec-violating unaligned ymm0 load that forced non-posted dword transactions from stalled GPU registers was used to break the fundamental design of System Management Mode in smiiiiiiiiiiiiiiii.

vmovdqu 0xfcc003b1, %ymm0

Rules

  • Instructions may use whatever setup is necessary, but only a single instruction is eligible to be scored.
  • Trapped/emulated/virtualized instructions may only time the trap, not the handler.
  • Instructions must not be interruptible. rep movs, pause, etc. are disqualified.
  • Times are normalized based on the CPU base clock frequency.
  • All platforms must be in their factory stock configurations - no hardware modifications.

x86 Leaderboard

27. nop

Strategy: nop does nothing. It opens the leaderboard accordingly.

Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz

nop

Score: 1 cycles

Time: 0 nanoseconds

26. nop16

Strategy: Regular nop was too short, but how do we make nothing take longer? Try a lonnnnnng nop.

Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz

data16 data16 data16 data16 data16 data16 data16 nopl 0x00000000(%%eax,%%eax,1)

Score: 20 cycles

Time: 7 nanoseconds

25. rdtsc

Strategy: Just a reference instruction to get our bearings.

Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz

rdtsc

Score: 49 cycles

Time: 18 nanoseconds

24. idiv

Strategy: Use 128-bit dividend (rdx:rax=2:0) with small divisor to push the quotient above the ceiling imposed by sign-extension, driving the longest path through the divider microcode.

Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz

xorq %rax, %rax   ; rax = 0  (low 64 bits of dividend)
movq $2, %rdx     ; rdx = 2  (high 64 bits: full dividend = 2^65)
movq $5, %rbx     ; divisor → quotient = 2^65/5 ≈ 7.4×10^18
idivq %rbx

Score: 77 cycles

Time: 28 nanoseconds

23. enter

Strategy: Use maximum nesting depth (31) to force 30 display-pointer loads and pushes through the microcode display-walk path.

Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz

enter $0, $31       ; 0 bytes allocated, nesting depth 31 (maximum)

Score: 112 cycles

Time: 41 nanoseconds

22. fldl

Strategy: Try a small denormal to trigger an FP microcode assist.

Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz

    movabsq $0x0000000000000001, %rax
movq    %rax,-8(%rsp)
    fldl    -8(%rsp)

Score: 133 cycles

Time: 49 nanoseconds

Strategy: Just ensure the cache line is dirty.

Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz

clflush (%rax)          ; rax -> dirty cache line resident in L3

Score: 165 cycles

Time: 60 nanoseconds

20. fsin

Strategy: Use exponent 0x7ff to reach 'special value' processing in microcode; positive/negative, NaN/inf doesn't seem to make a difference, go with QNaN.

Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz

    movabsq $0x7fffffffffffffff, %rax
movq    %rax,-8(%rsp)
    fldl    -8(%rsp)
fsin

Score: 257 cycles

Time: 94 nanoseconds

19. mfence

Strategy: Saturate all write-combining line-fill buffers with movnti stores to distinct cache lines, forcing mfence to drain the full LFB write path to the uncore before retiring.

Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz

movnti %r9,0*64(%rdi)   ; ×16 distinct cache lines — saturate the write-combining LFBs
; …
movnti %r9,15*64(%rdi)
mfence                     ; must drain all pending LFB writes before retiring

Score: 326 cycles

Time: 120 nanoseconds

Strategy: Nothing for now, just check how long it takes to invalidate the TLB.

Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)

mov %rax, %cr3

Score: 352 cycles

Time: 110 nanoseconds

17. fadd

Strategy: Hit x87 FP microcode assist path by using denormal source operand.

Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz

fldl   subnorm    ; 1e-310: value < DBL_MIN, biased exponent = 0
faddl  subnorm    ; source is subnormal → FP microcode assist

Score: 677 cycles

Time: 249 nanoseconds

Strategy: Align lock-prefixed operand to straddle cache-line boundary, forcing CPU to assert the external bus lock rather than using the fast MESI cache-coherence path.

Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz

; split_ptr % 64 == 63 — dword spans bytes 63 (line N) and 64–66 (line N+1)
lock xaddl %r9d, (%rdi)

Score: 865 cycles

Time: 319 nanoseconds

15. fdiv -

Strategy: Use subnormal divisor, hardware hands control to microcode assist, assist normalizes operand, performs the division, then restores architectural state.

Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz

    movabsq $0x3ff0000000000000, %rax   ; 1.0 (normal dividend)
movq    %rax,-8(%rsp)
    fldl    -8(%rsp)                     ; ST(0) = 1.0

    movabsq $0x0000002000000000, %rax   ; 6.79e-313 (subnormal divisor)
movq    %rax,-8(%rsp)
    fdivl   -8(%rsp)                     ; ST(0) = 1.0 / subnormal → FP assist

Score: 883 cycles

Time: 325 nanoseconds

14. cpuid

Strategy: Use rakefield to find the highest latency CPUID leaves.

Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz

movl $6, %eax
cpuid

Score: 1248 cycles

Time: 460 nanoseconds

13. rdrand

Strategy: Execute in a tight loop to deplete the hardware entropy pool faster than it can be refilled, forcing subsequent calls to stall while the entropy source recovers.

Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz

rdrand %rax

Score: 5,579 cycles

Time: 2.057 microseconds

12. wrmsr

Strategy: Use project:nightshyft to identify high latency MSRs. MCG_CTL on Zen look like a winner: may be a microcode quiesce and synchronize on MCA error banks across hardware units, some potentially off-die, requiring fabric-level communication rather than a simple local register write.

Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)

movl $0x17b, %ecx       ; MCG_CTL
wrmsr

Score: 34,304 cycles

Time: 10.742 microseconds

11. out

Strategy: Target an I/O port that straddles a NIC device register boundary, triggering the device to quiesce its TX DMA engine on each write.

Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)

mov $0xf019, %dx
outl %eax, %dx

Score: 49,857 cycles

Time: 15.580 microseconds

10. rdmsr

Strategy: Use project:nightshyft to identify high latency model-specific-registers: VIA uses an undocumented register at 0x133 that gives wildly high response time. No idea what it does.

Contender: VIA Eden Processor 800MHz

movl $0x133, %ecx ; undocumented MSR
rdmsr

Score: 161,602 cycles

Time: 202.004 microseconds

Strategy: Fully load L1/L2/L3 caches with dirty lines to force DRAM writeback of entire hierarchy.

Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)

wbinvd

Score: 1,616,480 cycles

Time: 506.165 microseconds

8. in

Strategy: Target I/O port mapped to an ACPI PM block where an unaligned 4-byte read decodes into multiple non-posted loads from wherever this port goes.

Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)

mov $0x0413, %dx
inl %dx, %eax

Score: 12,524,415 cycles

Time: 3.921769 milliseconds

7. mov

Strategy: Use mmiotic to identify high-latency deadspace in PCIe fabric, hit unkown GPU register.

Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)

movl 0xfcc003b0, %esi

Score: 443,937,696 cycles

Time: 139.010268 milliseconds

6. mov rax -

Strategy: Search MMIO space for slowest registers in PCIe fabric, hit unknown GPU register, use 8-byte MMIO read to get two dword register accesses, which isn't technically allowed but works anyway.

Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)

movq0xfcc003b0, %rax

Score: 887,716,864 cycles

Time: 277.971228 milliseconds

Strategy: Search MMIO space for slowest registers in PCIe fabric, hit unknown GPU register, use 16-byte MMIO read to get four dword register accesses, which isn't technically allowed but works anyway.

Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)

vmovdqu 0xfcc003b0, %xmm0

Score: 1,774,555,776 cycles

Time: 555.664133 milliseconds

Strategy: Search MMIO space for slowest registers in PCIe fabric, hit unknown GPU register, use 32-byte MMIO read to get eight dword register accesses, which still isn't technically allowed but works anyway.

Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)

vmovdqu 0xfcc003b0, %ymm0

Score: 3,549,079,296 cycles

Time: 1.111345034 s

Strategy: Search MMIO space for slowest registers in PCIe fabric, hit unknown GPU register, use 32-byte unaligned MMIO read to get nine dword register accesses, which is even less allowed than the aligned version, but works anyway.

Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)

vmovdqu 0xfcc003b1, %ymm0

Score: 4,453,212,256 cycles

Time: 1.394428818 seconds

Strategy: Use mmiotic to identify high-latency deadspace in PCIe fabric, isolate region near 0's and offset state to avoid MXCSR corruption (possibly VGA buffer?), use fxrstor64 to load 512-byte FPU/MMX/XMM state from MMIO, forcing CPU to process 512 bytes of I/O transactions through slowest available memory aperture.

Contender: AMD Ryzen 7 5800H

movl $0xfcc68830, %rsi
fxrstor64 %rsi

Score: 74,584,168,512 cycles

Time: 23.354502677 seconds

1. 🏆 fxrstor64 🏆

Strategy: Extend fxrstor64 (baseline) by starving the fabric while the load is in flight — a fleet of hammer cores pounds a different high-latency MMIO register with tight 4-byte reads, saturating the PCIe root complex and endpoint with non-posted transactions, so CPU 0's 512-byte fxrstor64 must queue behind all that contending traffic.

Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)

; CPU 0 — timed instruction
movl $0xfcc68830, %rsi
fxrstor64 %rsi

; CPUs 1..N — hammer loop against a different high-latency location
movl 0xfcc68858, %eax

🏆 Score: 198,002,498,236 cycles

🏆 Time: 62 seconds

??. xrstor64 (AMX, MMIO) :finnadie:

Strategy: Leverage extended AVX state in Sapphire Rapids with MMIO approach from fxrstor64: xsave state area is 8KB vs 512 bytes, 16x size -> 1,000,000,000,000 cycles

Contender: TODO

; XCR0 must enable AMX components (bits 17-18); state area ~8KB
xrstor64 (%rsi)         ; rsi -> MMIO region, same technique as fxrstor64

ARM Leaderboard

  • T.B.D.

RISC-V Leaderboard

  • T.B.D.

Author

The assembly hall-of-shame is a research effort from Christopher Domas (@xoreaxeaxeax).

Source: hackernews

arrow_back Back to News