

“sorry, best we can do is 2 shit jobs 7 days a week. AI took all the good ones and we’re not sharing a fucking thing with you”
- billionaire scum


“sorry, best we can do is 2 shit jobs 7 days a week. AI took all the good ones and we’re not sharing a fucking thing with you”


as long as it’s their hardware
This is the problem for me. AMD exist, Mac exists, Intel exists, and they rely on llama.cpp which now has the devs getting their paychecks from Nvidia… They’ll just “direct primary focus” onto cuda development and oops we didn’t touch SYCL support for 8 months… Whoopsiedoodle.
I literally just bought a b70… Fucking Nvidia…
Yes I know its open source, but the devs currently have a decent pace with updates. Relying on unpaid devs to care about SYLC when most people don’t run it anyway is probably worse than relying on Nvidia paid devs to eventually get to it :/


“together we will regulate model access, slow AMD, Mac, and Intel support on llama.cpp, and generally enshittify the experience like all corporations do! Welcome to the future, the same future as every other corporate acquisition!”


I think that no matter what happens we’re going to foot the bill. There’s no way in hell these assholes will ever face consequences of any sort. They already destroyed consumer electronics markets, and will bring the rest of the economy down with them if they fail, but they themselves and the parasite investors will never lose. We’ll have to pay…


It has been some time since my initial comment so at the time I was mainly using LM studio. Qwen 3.6 a3b is the MOE and it does work well on my card, but the dense model that is more intelligent/capable is the Qwen 3.6 27b which doesn’t fit on the card and does get offloaded, but offloading cuts the speed down to like 1/tps.
I have since found a version of the 27b model that is “quantized,” for lack of a better term, differently and has to be run through TabbyAPI which gets back to 30ish tps. It can’t offload so it must fit fully on the card which keeps the speed high. Might be worth a look if you’re interested, the only downside is that with my 16gb card the context limit has to be kept pretty low ~40k if I remember correctly


It’s the q4 quantization, but it requires 20+GB vram and my 5080 only has 16


Yeah I’m just beginning my local AI journey on a 5080, tried Qwen3.6 27b Q4 and was getting like 1tps because of the vram overflow. Ran it over night at it was still chewing on generating a prompt for a sub agent when I got up in the middle of the night until it simply ended in some kind of “fetch failure” lol. I think I gave it something too large to tackle, but either way 1tps is kinda garbage.
Except for the US where just like covid the conservative contrarians will conspiracy theory and hatemonger their way into going full steam ahead with absolutely no oversight. Their stock valuations depend on it…