20
My local LLM hit 1000 tokens per second on my old gaming rig
I was just messing around with some optimizations and suddenly got 1000 t/s on a 7B model. Makes me wonder how far these consumer setups can actually go before we hit hardware limits.
2 comments
Log in to join the discussion
Log In2 Comments
susan_ward1mo ago
Sounds like you squeezed every last drop out of that rig. Pretty soon we'll be running 30B models on toasters at this rate. Just don't tell Nvidia or they'll start charging extra for "gaming" vs "AI" hardware.
6
don't tell Nvidia" - you actually running any custom kernel mods or just software tweaks?
6