I wish they would just explain it in normal terms instead of this nasty LLM "engaging blog post" style
twothreeone 13 hours ago [-]
Chris usually takes an educational angle, I don't think this is LLM-generated content at all it's just his style. I highly encourage watching some of his DefCon or BlackHat talks, they're fun!
Brian_K_White 12 hours ago [-]
I looked at both links and don't see anything weird or annoying, and I hate overblown styles myself.
benmmurphy 7 hours ago [-]
‘The counters tell the story’ might be something they consider odd. But I’m not familiar with the author’s style so they could have had these tics pre-LLM or they picked up these tics from reading a lot of LLM content.
jonathrg 6 hours ago [-]
Yes, that's the point where I was like "wait a second" and started counting em dashes. A third alternative, perhaps the LLMs were trained hard on his prior work? :-)
kazinator 11 hours ago [-]
Bus cycles can be arbitrarily long on any processor that has memory cycles with a hand shake requiring an ack, with no timeout.
E.g. we can build a board around a MC68000 where we make it lock up forever in a bus cycle, waiting for a DTACK that doesn't arrive.
Some early microprocessors had clocked bus cycles without handshaking. They would put out an address on some address lines and signal some line together with a read/write indication, and then expect the transfer to be completed within some clock cycles. If nothing is attached to the address, they would read whatever values are on the bus, like maybe all 1's if it is an open drain system that requires the transmitting device to pull to ground to indicate zero.
I'd say that kind of thing belongs to a hall of shame; it requires software hacks to interface with anything that can't keep up with the prescribed bus cycle.
inigyou 7 hours ago [-]
Bus cycles can also be arbitrarily long on those microprocessors if they don't use dynamic logic - you can stop the clock.
Joel_Mckay 9 hours ago [-]
In general, more complex processors have latency issues, and in some ways modern chips have actually become worse with each design iteration.
Ultimately... failure modes must arise because naive gate design simulation models can't determine issues in a computationally feasible time frame.
The consequences of "fixing" CDC prone design flaws makes a processor many times slower (8 to 16 times slower on my dumb attempt), and develops weird alien design features very different from Von Neumann architectures.
This is why we can't have nice things. =3
kazinator 6 hours ago [-]
I smell Buridan's ass. :)
monocasa 17 hours ago [-]
It says in the rules
> Trapped/emulated/virtualized instructions may only time the trap, not the handler.
But I feel like that 12ms write to an ACPI IO port at current leaderboard position 8 is probably trapping to SMM and being handled there.
thyristan 4 hours ago [-]
Depending on his interpretation of the rules about trapped instructions, one could just build a loop in the x86 page tables. Those are usually a tree linked by pointers, and any page table lookup can create another page fault that creates another lookup that...
And the simplest thing you can do on such a system is just to loop indefinitely, thus creating a simple instruction with a memory access (mov or anything, doesn't really matter, even the instruction fetch for a nop would work) to take infinite time.
inigyou 3 hours ago [-]
Page tables are physically addressed, so can't recurse. I assume this thing actually works by causing a page fault on the first instruction of the page fault handler, which is a new instruction.
thyristan 3 hours ago [-]
> Page tables are physically addressed
Nope. Not on x86. You can use either physical or virtual addresses at your choosing. Consumer OSes use virtual ones, so you can swap out page tables (yes, really!). See https://wiki.osdev.org/X86_Paging "Page directory".
inigyou 2 hours ago [-]
Nothing in this section mentions them containing virtual addresses. In fact the word "physical" is written in bold. Are you a hallucinating LLM?
TomatoCo 19 hours ago [-]
This author also has other things like: A compiler that emits only `mov` instructions and another compiler that deliberately messes with the control flow so that, if disassembled, common debuggers will draw symbols like skulls or threats. https://github.com/xoreaxeaxeax/repsych
inigyou 18 hours ago [-]
He also bruteforced the entire opcode space to find undocumented instructions (sandsifter).
vanderZwan 9 hours ago [-]
Also came up with the original ..cantor.dust.. binary visualization tool, which is a tool I never used directly but Chris' presentation of it in 2012 is still one of the coolest talks I've ever seen.
Nop should be #1, because it is infinitely slow for what it does. ;)
jooops1 18 hours ago [-]
It increments rip by one.
EvanAnderson 14 hours ago [-]
I thought I remembered reading somewhere re: the 8086 microcode disassembly that NOP, which is encoded as XCHG AX,AX actually does run the XCHG microcode and uses an internal scratchpad register to do the exchange.
JoeAltmaier 14 hours ago [-]
There were several NOPs - XCHG BX,BX and so on. Those were taken later to be prefixes for new classes of opcodes.
bonzini 7 hours ago [-]
None was. XCHG AX,AX is special because XCHG AX,reg has a one-byte encoding.
You're probably confusing with:
- POP CS being broken and later becoming a prefix
- some opcodes being "reserved NOPs", i.e. reserved without generating #UD. They are used for instructions that may be defined in the future while guaranteeing backwards compatibility, for example new kinds of prefetches. MPX bounds checking instructions were also encoded in reserved NOPs.
fluoridation 15 hours ago [-]
No, that's done by the decoder. It actually does nothing.
dlcarrier 12 hours ago [-]
It's still part of the instruction to increment it by one, as opposed to write a value or offset to it, as jump instructions do.
fluoridation 12 hours ago [-]
See my sibling response.
loeg 13 hours ago [-]
The decoder is an implementation detail that is a subcomponent of NOP; GP was right, and your correction isn't.
phire 9 hours ago [-]
As specified by the spec, it arguably increments RIP by one.
The actual typical hardware implementation just fetches the next 16-32 bytes from icache, shifts it to the correct alignment, and slams it into a bunch of parallel decoders which each attempts to decode one x86 instruction per byte.
The next cycle, the first 1-6 non-overlapping valid instructions are accepted into a queue for further decoding. The NOP almost certainly takes up space in this queue.
At no point does RIP get incremented by one. There isn't even a single physical RIP register to increment, the CPU is "executing" dozens or even hundreds of RIPs in parallel.
fluoridation 13 hours ago [-]
It's not an implementation detail, because the decoder runs before the execution of every instruction. If we're going to say that NOP increments IP by one, then we should also say that ADD "stores in dst the addition of src and dst, as well as incrementing IP by the length of the instruction", and JMP imm "increments JMP by imm + the length of the instruction".
loeg 12 hours ago [-]
ADD does in fact do that.
fluoridation 11 hours ago [-]
I'm not disputing the total effect. I'm asking if you'd rather describe ADD and JMP in this manner, in order to say that NOP does not in fact do nothing.
dxdm 7 hours ago [-]
Love the pedantry here. I'm currently at: no op code does anything, it's all fancy effects in the hardware that can be described in arbitrary detail, which somehow allows me to cause these letters to appear on your screen.
fluoridation 6 hours ago [-]
Insightful and compelling. Tell me another one.
dxdm 4 hours ago [-]
With pleasure.
There was a boy
A very strange, enchanted boy
They say he wandered very far
Very far, over land and sea
A little shy and sad of eye
But very wise was he.
And then one day
A magic day he passed my way
And while we spoke of many things
Fools and kings
This he said to me:
"The greatest thing you'll ever learn
Is just to love and be loved in return."
russdill 15 hours ago [-]
I mean....there are several architectures out there which has a nop that is a jump forward. Kind of a tree forest issue imho
fluoridation 13 hours ago [-]
I don't know about other architectures in as much detail. I know x86 NOP does nothing.
mito88 18 hours ago [-]
Strategy: nop does nothing. It opens the leaderboard accordingly.
Score: 1 cycles
Time: 0 nanoseconds
layer8 17 hours ago [-]
It opens the leaderboard as #27, so in the last place.
hyperhello 12 hours ago [-]
It's a little faster than yep.
markus_zhang 17 hours ago [-]
Does that mean Chris Domas is ready for his next adventure?
codeshaunted 19 hours ago [-]
what im seeing from this chart is that we should be using the nop instruction for everything
bee_rider 19 hours ago [-]
Well the best code is no code. Nop could be second best though.
inigyou 18 hours ago [-]
Instructions unclear. Set the NX bit to ensure no code, and got a general protection fault.
rurban 6 hours ago [-]
62s for a single instruction! Wonder if an compilers cost tables knows that. But it's data dependent, and cost functions probably don't do that.
inigyou 3 hours ago [-]
Compilers don't even emit this instruction.
Retr0id 9 hours ago [-]
I wonder if you can do some damage with scatter/gather ops within a VM, such that each fetch is a TLB miss inside the VM, and every table walk fetch is a TLB miss outside of the VM (which gets you up to 24 "fetches per fetch").
vardump 19 hours ago [-]
A great resource for any performance deoptimization.
Very cool! Also, huh interesting. I’ve used rdtsc to measure cycle diffs but had no idea its execution takes that long. Is that common across architectures?
pbsd 17 hours ago [-]
The cycle count for RDTSC is ~25 cycles on Skylake-era microarchitectures. The 49 number shown in the OP seems off.
phire 6 hours ago [-]
It's actually benchmarking 1000 repetitions of the RDTSC instruction running in parallel.
My guess... On Skylake, multiple in-flight RDTSC instructions slow each other down for some reason?
Possibly because it's attempting to provide a strict monotonic guarantee, that no two RSTSC instructions will return the same timestamp. Intel's manual only claims monotonic, which theoretically allows for two RSTSC instructions to return the same timestamp.
17 hours ago [-]
inigyou 18 hours ago [-]
AFAIK it acts as some kind of execution barrier, to give meaningful timing.
PCIe is more like a packet-switched network than a bus, which is incidentally why things like Thunderbolt (effectively external PCIe) and sillier demonstrations like https://www.youtube.com/watch?v=q5xvwPa3r7M work.
Pretty dumb right? When latency gets this high, you need a more asynchronous design to get any reasonable performance. PCIe is clearly designed with the assumption of latencies a few hundred cycles at most (or usually) - this MMIO register is an extreme outlier. It might be unmapped, and timing out on the hardware side, or it might be converted to an access on some really slow configuration bus.
11 hours ago [-]
achierius 19 hours ago [-]
It'd be really interesting to see whether the winning (losing?) instructions/strategies would be different on other architectures. At least right now the top spot (`fxrstor64` on MMIO, starve PCIe) seems relatively architecture-independent, but maybe something about MMIO ordering rules on e.g. POWER would be different enough to change that -- or perhaps open up new avenues?
I wonder what the actual limit on this `fxrstor64` is right now. If you can stall the PCIe bus for that long, then why not indefinitely? Certainly there's no forward progress guarantee here.
inigyou 3 hours ago [-]
I think the idea was to find a long instruction on an ordinary PC. Of course by adding special hardware you can stall things.
baddash 15 hours ago [-]
just curious, how much do these actually discover useful practices or pitfalls, on top of just being for fun?
phire 6 hours ago [-]
The repo page says exactly why these long running instructions were researched.
Oh wow, glad to see Chris Domas active online again!
Telaneo 6 hours ago [-]
Domas manages to abuse x86 is ways that make me unsure whether or not I should be impressed or disgusted. I guess impressed, then disgusted over Intel (and AMD?).
IshKebab 17 hours ago [-]
Using MMIO is cheating and makes the results very boring.
It would be much more interesting to know the results if you're only allowed to use main memory.
arn3n 19 hours ago [-]
There’s definitely strategies here; A lot of the floating point operations use subnormals, and a lot of the worst instructions are slowed down by really, really fucking with MMIO.
56767865678 9 hours ago [-]
Sahil
eek2121 15 hours ago [-]
This is neat!
hncsiocp9x 2 hours ago [-]
[dead]
2_foos_in_a_bar 19 hours ago [-]
[dead]
metadat 19 hours ago [-]
It’s crazy how computers still seem to get perceivably slow every few years, given how many instructions can be executed in 1ms. Shameful, even..
What’s that law called about programmers wasting all the compute on abstraction?
AceJohnny2 14 hours ago [-]
There are two aspects to compute performance: latency and throughput.
There has been a ton of work to improve throughput, but that's often come at the cost of latency, because a great technique to improve throughput is batching, to avoid the per-task overhead cost.
We've also added many layers of abstraction.
A seminal example is Dan Luu's "computer latency" table (2017) https://danluu.com/input-lag/, which measures the time between a keypress and visible changes on the screen, and wherein an Apple IIe has latency 5x lower than a Lenovo X1 Carbon on Windows.
From some perspective, you could say that the Lenovo on Windows example is an impressive feat of engineering, considering all the subsystems involved (USB, interrupt dispatcher, input layer, windowing system, double-buffering...)
This throughput-over-latency tradeoff can be seen across the entire computer landscape design.
loeg 13 hours ago [-]
Idk man, my computers don't seem to get any slower over time -- no upgrades to any components, either.
metadat 13 hours ago [-]
Does responsiveness get better or worse with updates, in general?
loeg 12 hours ago [-]
Neither? I haven't noticed a perceivable change in a long time.
mwigdahl 18 hours ago [-]
Wirth's Law I believe.
inigyou 18 hours ago [-]
And remember to call him by name, not by value!
HappyPanacea 19 hours ago [-]
The new windows notepad is a disgrace
adamrezich 17 hours ago [-]
The new mspaint fucked, then unfucked, then refucked my decades-old muscle memory of Win+R mspaint Enter Ctrl+E 1 Tab 1 Enter Ctrl+V to open Paint, resize canvas to minimum, then paste from clipboard. When you press Ctrl+E now, the Units control is selected by default, for some completely asinine reason!!!
fluoridation 15 hours ago [-]
That's pretty cool. I've never needed to do that because Paint remembers the last canvas size set manually.
EvanAnderson 14 hours ago [-]
You aren't alone in being frustrated by decades of muscle memory being disrespected by MSFT.
sitzkrieg 15 hours ago [-]
i’m on LTSC for this (and many other) reasons lol! don’t touch my mspaint and notepad
inigyou 18 hours ago [-]
Andy and Bill's Law
summarybot 18 hours ago [-]
The OS should do less not more
RiverCrochet 15 hours ago [-]
vi should be a kernel-level system call, tunable with sysctls. For agentic management.
LoganDark 19 hours ago [-]
Huh? A millisecond is an eternity!
18 hours ago [-]
m463 18 hours ago [-]
I remember reading once somewhere:
If some app responds in 10ms or less, it is INTERACTIVE.
makes you think.
Xirdus 18 hours ago [-]
It is literally impossible to respond to input in 10ms on most platforms, for various reasons. The USB input lag of 12-30ms and the 60Hz refresh rate of most monitors being just the first two.
m463 16 hours ago [-]
I stand corrected.
I looked it up and it is .1 seconds (100ms)
The basic advice regarding response times has been about the same for thirty years [Miller 1968; Card et al. 1991]:
- 0.1 second is about the limit for having the user feel that the system is reacting instantaneously, meaning that no special feedback is necessary except to display the result.
- 1.0 second is about the limit for the user's flow of thought to stay uninterrupted, even though the user will notice the delay. Normally, no special feedback is necessary during delays of more than 0.1 but less than 1.0 second, but the user does lose the feeling of operating directly on the data.
- 10 seconds is about the limit for keeping the user's attention focused on the dialogue. For longer delays, users will want to perform other tasks while waiting for the computer to finish, so they should be given feedback indicating when the computer expects to be done. Feedback during the delay is especially important if the response time is likely to be highly variable, since users will then not know what to expect.
That’s the prevailing statistic, but your original number isn’t that wrong either:
Humans can perceive much smaller latencies.
If you look at the Card & Miller reference, at least some humans can perceive differences in ~50ms vs 100ms latencies when typing (in my limited testing, it’s likely you can!). There’s some newer research I don’t have handy that I believe found error rates decreased and NSAT improved until around at least 30ms (if not 20ms).
On that note, humans can definitely distinguish 60hz vs 120hz reliably (about 8ms faster per frame).
Even faster: with a reference (eg when dragging on a touchscreen), humans can distinguish down to at least 1ms vs 10ms of latency: https://m.youtube.com/watch?v=vOvQCPLkPt4
And you can probably distinguish metronomes that are off by about 1-2ms. Much smaller for other things (like metronomes that slightly slower or faster than one another).
60Hz monitors definitely prevent it, but I'm pretty sure USB lag is far less than 12-30ms. My USB mouse can make a round-trip to a remote server faster than that.
sitzkrieg 15 hours ago [-]
this is also why 60 hz refresh rate is all but dead outside of console gaming (not to mention 1000hz poll rate devices being the norm)
faresahmed 15 hours ago [-]
Sorry for nitpicking but the logical inverse of this statement is the stronger and more appropriate version:
If some app does not respond in 10ms or less, it is not interactive.
darksim905 11 hours ago [-]
Seems like spam from this creator since there are two things on the front page?
john_strinlai 11 hours ago [-]
submitted by two different people, both with year+ old accounts and decent karma. i dont think either is the author. the other submitter probably read this one, looked at the github, saw something else cool and posted it. (i almost did the same, but bookmarked it instead)
E.g. we can build a board around a MC68000 where we make it lock up forever in a bus cycle, waiting for a DTACK that doesn't arrive.
Some early microprocessors had clocked bus cycles without handshaking. They would put out an address on some address lines and signal some line together with a read/write indication, and then expect the transfer to be completed within some clock cycles. If nothing is attached to the address, they would read whatever values are on the bus, like maybe all 1's if it is an open drain system that requires the transmitting device to pull to ground to indicate zero.
I'd say that kind of thing belongs to a hall of shame; it requires software hacks to interface with anything that can't keep up with the prescribed bus cycle.
https://en.wikipedia.org/wiki/Metastability_(electronics)
Ultimately... failure modes must arise because naive gate design simulation models can't determine issues in a computationally feasible time frame.
The consequences of "fixing" CDC prone design flaws makes a processor many times slower (8 to 16 times slower on my dumb attempt), and develops weird alien design features very different from Von Neumann architectures.
This is why we can't have nice things. =3
> Trapped/emulated/virtualized instructions may only time the trap, not the handler.
But I feel like that 12ms write to an ACPI IO port at current leaderboard position 8 is probably trapping to SMM and being handled there.
Leads to x86 page table MMU magic being turing complete: https://github.com/jbangert/trapcc
And the simplest thing you can do on such a system is just to loop indefinitely, thus creating a simple instruction with a memory access (mov or anything, doesn't really matter, even the instruction fetch for a nop would work) to take infinite time.
Nope. Not on x86. You can use either physical or virtual addresses at your choosing. Consumer OSes use virtual ones, so you can swap out page tables (yes, really!). See https://wiki.osdev.org/X86_Paging "Page directory".
[0] https://github.com/Battelle/cantordust
[1] https://www.youtube.com/watch?v=4bM3Gut1hIk
You're probably confusing with:
- POP CS being broken and later becoming a prefix
- some opcodes being "reserved NOPs", i.e. reserved without generating #UD. They are used for instructions that may be defined in the future while guaranteeing backwards compatibility, for example new kinds of prefetches. MPX bounds checking instructions were also encoded in reserved NOPs.
The actual typical hardware implementation just fetches the next 16-32 bytes from icache, shifts it to the correct alignment, and slams it into a bunch of parallel decoders which each attempts to decode one x86 instruction per byte.
The next cycle, the first 1-6 non-overlapping valid instructions are accepted into a queue for further decoding. The NOP almost certainly takes up space in this queue.
At no point does RIP get incremented by one. There isn't even a single physical RIP register to increment, the CPU is "executing" dozens or even hundreds of RIPs in parallel.
There was a boy A very strange, enchanted boy They say he wandered very far Very far, over land and sea A little shy and sad of eye But very wise was he.
And then one day A magic day he passed my way And while we spoke of many things Fools and kings This he said to me: "The greatest thing you'll ever learn Is just to love and be loved in return."
Score: 1 cycles Time: 0 nanoseconds
[0]: https://en.wikipedia.org/wiki/Core_War
My guess... On Skylake, multiple in-flight RDTSC instructions slow each other down for some reason?
Possibly because it's attempting to provide a strict monotonic guarantee, that no two RSTSC instructions will return the same timestamp. Intel's manual only claims monotonic, which theoretically allows for two RSTSC instructions to return the same timestamp.
...and with things like https://en.wikipedia.org/wiki/ExpEther , you can get even higher latencies.
I wonder what the actual limit on this `fxrstor64` is right now. If you can stall the PCIe bus for that long, then why not indefinitely? Certainly there's no forward progress guarantee here.
A very long-running instruction can be used to break SMI: https://github.com/xoreaxeaxeax/smiiiiiiiiiiiiiiii
It would be much more interesting to know the results if you're only allowed to use main memory.
What’s that law called about programmers wasting all the compute on abstraction?
There has been a ton of work to improve throughput, but that's often come at the cost of latency, because a great technique to improve throughput is batching, to avoid the per-task overhead cost.
We've also added many layers of abstraction.
A seminal example is Dan Luu's "computer latency" table (2017) https://danluu.com/input-lag/, which measures the time between a keypress and visible changes on the screen, and wherein an Apple IIe has latency 5x lower than a Lenovo X1 Carbon on Windows.
From some perspective, you could say that the Lenovo on Windows example is an impressive feat of engineering, considering all the subsystems involved (USB, interrupt dispatcher, input layer, windowing system, double-buffering...)
This throughput-over-latency tradeoff can be seen across the entire computer landscape design.
If some app responds in 10ms or less, it is INTERACTIVE.
makes you think.
I looked it up and it is .1 seconds (100ms)
The basic advice regarding response times has been about the same for thirty years [Miller 1968; Card et al. 1991]:
- 0.1 second is about the limit for having the user feel that the system is reacting instantaneously, meaning that no special feedback is necessary except to display the result.
- 1.0 second is about the limit for the user's flow of thought to stay uninterrupted, even though the user will notice the delay. Normally, no special feedback is necessary during delays of more than 0.1 but less than 1.0 second, but the user does lose the feeling of operating directly on the data.
- 10 seconds is about the limit for keeping the user's attention focused on the dialogue. For longer delays, users will want to perform other tasks while waiting for the computer to finish, so they should be given feedback indicating when the computer expects to be done. Feedback during the delay is especially important if the response time is likely to be highly variable, since users will then not know what to expect.
from Jakob Nielsen:
https://www.nngroup.com/articles/response-times-3-important-...
less readable but the original paper:
https://www.yusufarslan.net/sites/yusufarslan.net/files/uplo...
Humans can perceive much smaller latencies.
If you look at the Card & Miller reference, at least some humans can perceive differences in ~50ms vs 100ms latencies when typing (in my limited testing, it’s likely you can!). There’s some newer research I don’t have handy that I believe found error rates decreased and NSAT improved until around at least 30ms (if not 20ms).
On that note, humans can definitely distinguish 60hz vs 120hz reliably (about 8ms faster per frame).
Even faster: with a reference (eg when dragging on a touchscreen), humans can distinguish down to at least 1ms vs 10ms of latency: https://m.youtube.com/watch?v=vOvQCPLkPt4
And you can probably distinguish metronomes that are off by about 1-2ms. Much smaller for other things (like metronomes that slightly slower or faster than one another).
This is a special interest of mine XD