~~Good article on this, but I can not get this to work on my computer yet. I've been trying all morning.~~ Okay, I finally got Qwen 3.8 27B working on my machine. I had to build a newer CUDA-enabled version of llama.cpp from source, adjust the GPU memory allocation, and limit the reasoning budget. It’s running now, and I’m still tweaking it for better speed. So far, so good.
this post was submitted on 22 Aug 2026
-3 points (38.5% liked)
AI discussions
48 readers
1 users here now
A community dedicated to respectful discussions about AI models. Meant to be a sister community to !stable_diffusion@lemmy.dbzer0.com Discussions can center around LLM chat models or image models, even audio.
This community is meant to be a resource for people looking to learn more about AI models, self-hosting them, or just how they work in general.
Rules:
- Be respectful to others (this means no ableism or transphobia)
- Do not promote capitalism or corporate AI models (discussing them is okay, don't encourage their use)
- No spam or trolling
- Provide links or sources when relevant
founded 4 months ago
MODERATORS