📢
20
c/ai-innovations•miller.jasonmiller.jason•1mo ago

My local LLM hit 1000 tokens per second on my old gaming rig

I was just messing around with some optimizations and suddenly got 1000 t/s on a 7B model. Makes me wonder how far these consumer setups can actually go before we hit hardware limits.
2 comments

Log in to join the discussion

Log In
2 Comments
susan_ward
susan_ward1mo ago
Sounds like you squeezed every last drop out of that rig. Pretty soon we'll be running 30B models on toasters at this rate. Just don't tell Nvidia or they'll start charging extra for "gaming" vs "AI" hardware.
6
patel.daniel
patel.daniel1mo agoMost Upvoted
don't tell Nvidia" - you actually running any custom kernel mods or just software tweaks?
6