Recommendation for Local / open

Best Local LLM for Coding

The best local LLM for coding on consumer hardware is GLM 5.2, which multiple users report matches Claude Sonnet 4.6-level performance while running locally for complete privacy. If you have substantial hardware (128GB+ RAM), Kimi K2.6 is the benchmark leader among open-weights for coding reasoning tasks. For those with limited VRAM, Gemma 4 31B and Qwen3 Coder Next offer strong performance-to-size ratios that fit on a single high-end consumer GPU. The right choice depends heavily on your hardware constraints. Users with 64GB or less should look at GLM 5.2, Gemma 4 31B, or the smaller MoE models, while those with server-class memory can consider Kimi K2.6 despite its 600GB+ storage requirements at usable quantizations.

About this recommendation

Updated
Jul 17, 2026
Evidence through
Jul 17, 2026
Sources
32
Revision
v6
  1. GLM 5.2 is currently the best option for local coding if you want near-frontier performance without needing server-class hardware. Users consistently praise its capability, with one calling it the best open-weight model that beats proprietary Gemini and older Claude versions. It runs on more accessible hardware than the massive Kimi models while delivering Claude Sonnet 4.6-level quality. For developers prioritizing both quality and practical local deployment, this is the top choice.

    Best when: You want the best quality-to-hardware ratio and need a model that genuinely competes with proprietary coding assistants.

    Tips

    • Users report it matches Claude Sonnet 4.6 in capability, beating proprietary Gemini and older Claude versions for coding tasks.
      Source 1
      Also, GLM 5.2 seems to be the best open-weight model, and it beats proprietary Gemini and older versions of Claude which is amazing. You can have a model at level of Claude Sonnet 4.6 at home without sharing anything, and maybe even uncensor it.
    • Considered the best open-weight model by community members testing it against other local options.
      Source 1
      Also, GLM 5.2 seems to be the best open-weight model, and it beats proprietary Gemini and older versions of Claude which is amazing. You can have a model at level of Claude Sonnet 4.6 at home without sharing anything, and maybe even uncensor it.
    • Can run locally without sharing data, appealing for privacy-sensitive projects.
      Source 1
      Also, GLM 5.2 seems to be the best open-weight model, and it beats proprietary Gemini and older versions of Claude which is amazing. You can have a model at level of Claude Sonnet 4.6 at home without sharing anything, and maybe even uncensor it.
    • Users have found configurations that work in as little as 10GB RAM for certain setups.
      Source 2
      Someone posted this fork, which should work in just 10 GB RAM, at least for another big model. I thought of your request and came back to spread the news :) https: github.com ErikTromp colibri-hy3
    • Frequently mentioned alongside Kimi models as a top choice for coding assistance through OpenRouter or local deployment.
      Source 3
      > I have been moving more and more to K2.7 Code and GLM-5.2 the last few weeks. They are often good enough for assistance, very fast, and cheap. I've moved completely to local models that I run with my M1 Mac Studio (64gb ram) some time ago. But for the rare times when I feel the local, quantized Qwen3.6 isn't enough, I just connect to Openrouter and use something like Kimi, GLM or Deepseek for a fraction of the price of Anthropic et al.

    Watch out for

    • Performance varies significantly based on quantization and hardware setup.
      Source 4
      Just ran and scored 63 3d model generations (via code) across high and no reasoning. 3D Modeling benchmark quickly shows spatial, logic and code performance of the model so I think it's a very good indicator of the quality. Here are the results compared to Gemini 3.5 Flash: Model + config CodeErr gen Cost gen Median time Quality gemini-3.5-flash, low 0.71 $0.18 68s baseline GLM 5.2, reasoning high 0.61 $0.18 289s -6.0% GLM 5.2, reasoning off 1.52 $0.10 126s -13.6% Although it is cheaper, it is…
  2. Kimi K2.6 is the benchmark champion for open-weight coding models, ranking highly on SWE-Bench Verified and topping one-shot coding reasoning among open-weights. The catch is hardware requirements, with a usable Q8 quantization needing roughly 700GB of VRAM. This is a model for enterprises or enthusiasts with serious server hardware, not consumer gaming rigs. If you have the infrastructure, it delivers quality comparable to Gemini 3.1 Pro Preview.

    Best when: You have server-class hardware (128GB+ VRAM minimum, ideally 700GB+) and need the highest benchmark performance among fully open-weight models.

    Tips

    • Achieves approximately 80.2% on SWE-Bench Verified, putting it in the top tier of open-weight models and competitive with SOTA proprietary services.
      Source 5
      > As of May 2026, how much money do I need to spend to buy hardware to have a local model that is 80% as good as SOTA services for assisting me in writing code? https: llm-stats.com benchmarks swe-bench-verified SOTA (public proprietary models) would be Opus 4.7 at 0.876 80% of that would be around 0.7. These models qualify, and are upwards of 90% as good in benchmarks: DeepSeek-V4-Pro-Max - 1.6T (HuggingFace shows 862B, huh) - 0.806 Kimi K2.6 - 1.1T - 0.802 MiniMax M2.5 - 229B - 0.802 DeepSeek…
    • Currently the top open-weights model in one-shot coding reasoning, slightly better than GLM 5.1.
      Source 6
      Early benchmarks show tremendous improvement over Kimi K2 Thinking, which didn't perform well on our benchmarks (and we do use best available quantization). Kimi K2.6 is currently the top open weights model in one-shot coding reasoning, a little better than GLM 5.1, and still a strong contender against SOTA models from ~3 months ago (comparable to Gemini 3.1 Pro Preview). Agentic tests are still running, check back tomorrow. Open weights models typically struggle with longer contexts in agentic…
    • Strong contender against SOTA models from 3 months ago, comparable to Gemini 3.1 Pro Preview.
      Source 6
      Early benchmarks show tremendous improvement over Kimi K2 Thinking, which didn't perform well on our benchmarks (and we do use best available quantization). Kimi K2.6 is currently the top open weights model in one-shot coding reasoning, a little better than GLM 5.1, and still a strong contender against SOTA models from ~3 months ago (comparable to Gemini 3.1 Pro Preview). Agentic tests are still running, check back tomorrow. Open weights models typically struggle with longer contexts in agentic…
    • Users report no discernible loss of analytical quality compared to other top models for general coding tasks.
      Source 7
      And if you want to dial in a setting in between: I've switched to Kimi K2.6 (now K2.7) and DeepSeek through OpenRouter and Reasonix for pretty much everything, with no discernible loss of analytical quality or utility. However, like many commenters, I don't really believe in vibe-coding, long-horizon agentic one-shot agentic coding, etc. and do not use LLMs for huge generation tasks that involve designing things end-to-end. I also have an MBP with 128 GB of unified memory and do quite a bit of…
    • Available through privacy-respecting hosted services like Infomaniak for users who need API access without enterprise contracts.
      Source 8
      Yes, for client projects where privacy and security is important, but no enterprise contract: Open code against Infomaniak hosted OSS models: Qwen3.5-122B-A10B-FP8, Kimi-K2.6. I use API keys for billing. It performs like Dec 2025 in terms of my productivity back then.

    Watch out for

    • Q8 quantization requires around 600GB on disk and approximately 700GB of VRAM, making it inaccessible to most users.
      Source 9
      People thinking to self-host Kimi K2.6 had better be prepared for how big it is. Q8 K XL quantization for instance is around 600GB on disk. I would bet about 700GB of VRAM needed. Quantizations lower than Q8 are probably worthless for quality. Or 2.05TB on disk for the full precision GGUF. https: huggingface.co unsloth Kimi-K2.6-GGUF If you can afford the hardware to run Kimi K2.6 at any decent speed for more than 1 simultaneous user, you probably have a whole team of people on staff who are al…
    • Lower quantizations than Q8 are considered worthless for quality, leaving no middle ground for smaller setups.
      Source 9
      People thinking to self-host Kimi K2.6 had better be prepared for how big it is. Q8 K XL quantization for instance is around 600GB on disk. I would bet about 700GB of VRAM needed. Quantizations lower than Q8 are probably worthless for quality. Or 2.05TB on disk for the full precision GGUF. https: huggingface.co unsloth Kimi-K2.6-GGUF If you can afford the hardware to run Kimi K2.6 at any decent speed for more than 1 simultaneous user, you probably have a whole team of people on staff who are al…
  3. Gemma 4 31B hits a sweet spot for developers with consumer-grade hardware who still want strong coding performance. The 4-bit QAT version delivers nearly identical performance to full precision while fitting in 24GB VRAM. Users report it outperforms Qwen 3.6 27B for coding and C programming tasks in their testing. It's a practical choice for anyone canceling cloud subscriptions to go fully local.

    Best when: You have 24-32GB VRAM and want a model that runs efficiently at 4-bit quantization without significant quality loss.

    Tips

    • The 4-bit QAT version has practically identical performance to full precision in most benchmarks and user tests.
      Source 10
      A 4-bit quantization of either Qwen 3.6 27b or Gemma 4 31b will run on a 32GB Mac with a decent-sized, but not full-sized, context. 64GB gets you the full ~256k context and you don't need to quantize your KV cache (though 8-bit quantization of KV may be worth it for performance). The 4-bit QAT version of Gemma 4 has practically identical performance to the full size version or the 8-bit version in most benchmarks and my tests, so there's no reason to run anything else. The 4-bit Qwen is a littl…
    • Performs better than Qwen 3.6 27B for coding and C programming tasks according to user testing.
      Source 11
      I opted to buy a normal 32GB laptop for this very reason. I know how loud and hot the GPUs in my desktop run when running even smallish models like Qwen 27B or Gemma 4 31B (which is a better model for most than Qwen 3.6, despite the benchmarks). I also have a Strix Halo which doesn't get loud, because it has a single huge fan, but it does get hot. So, there's no way a laptop could work as hard as models make them work, and not be unbearable. Tiny fans trying to remove all that heat? They gotta…
      Source 12
      At 24GB, Gemma 4 31B QAT will be better and give more concise answers. This post is mostly about unquantized results, so it's less relevant and I can't say much about as I haven't tested Qwen or Gemma via cloud API or unquantized locally. All I can say is locally, quantized in a 24GB scenario, Gemma 4 31B is better in my tests which are mostly reasoning or C programming related. Gemma 4 is the only model series at this parameter scale I've seen correctly answer some of these. One of the answers…
    • Runs comfortably on an RTX 6000 PRO with 96GB at full precision, or on 32GB Macs with 4-bit quantization.
      Source 13
      in my experience running models that have been heavily quantized(q4) or altered to some extent has never made me say “wow, this is so amazing”. On the contrary, the model ended up in the thrash bin after a few prompts. I have an RTX 6000 PRO with 96GB, and what I can run comfortably is Qwen 3.6 27B or MoE, Gemma 4 31B. This is as far as it goes when you run the model at full precision and maximum context length. They perform well and you can use them for coding, doing research on the internet a…
      Source 10
      A 4-bit quantization of either Qwen 3.6 27b or Gemma 4 31b will run on a 32GB Mac with a decent-sized, but not full-sized, context. 64GB gets you the full ~256k context and you don't need to quantize your KV cache (though 8-bit quantization of KV may be worth it for performance). The 4-bit QAT version of Gemma 4 has practically identical performance to the full size version or the 8-bit version in most benchmarks and my tests, so there's no reason to run anything else. The 4-bit Qwen is a littl…
    • Good enough for users to cancel cloud subscriptions when paired with local inference via llama.cpp.
      Source 14
      For personal needs I connected VSCode with llama.cpp running Qwen 3.6 27B or Gemma 4 31B and it's good enough to cancel my cloud subscription. Qwen running on my 1st GPU at q4@176k context from 70 to 50 tok s with MTP, pretty good for coding. Gemma on the other hand is using both GPUs, running q8@64k context, doing document sentiment analysis, summarization, proofreading and translating, at consistent 25 tok s. Somewhat slow but usable for batched workflows. Might get some more once llama.cpp s…
    • Correctly answers reasoning and C programming questions that other models at this scale fail on.
      Source 12
      At 24GB, Gemma 4 31B QAT will be better and give more concise answers. This post is mostly about unquantized results, so it's less relevant and I can't say much about as I haven't tested Qwen or Gemma via cloud API or unquantized locally. All I can say is locally, quantized in a 24GB scenario, Gemma 4 31B is better in my tests which are mostly reasoning or C programming related. Gemma 4 is the only model series at this parameter scale I've seen correctly answer some of these. One of the answers…

    Watch out for

    • At Q8 quantization with full context, it requires both GPUs on some setups, reducing flexibility.
      Source 14
      For personal needs I connected VSCode with llama.cpp running Qwen 3.6 27B or Gemma 4 31B and it's good enough to cancel my cloud subscription. Qwen running on my 1st GPU at q4@176k context from 70 to 50 tok s with MTP, pretty good for coding. Gemma on the other hand is using both GPUs, running q8@64k context, doing document sentiment analysis, summarization, proofreading and translating, at consistent 25 tok s. Somewhat slow but usable for batched workflows. Might get some more once llama.cpp s…
    • Can run hot and loud on laptops when under load, making it unsuitable for some portable setups.
      Source 11
      I opted to buy a normal 32GB laptop for this very reason. I know how loud and hot the GPUs in my desktop run when running even smallish models like Qwen 27B or Gemma 4 31B (which is a better model for most than Qwen 3.6, despite the benchmarks). I also have a Strix Halo which doesn't get loud, because it has a single huge fan, but it does get hot. So, there's no way a laptop could work as hard as models make them work, and not be unbearable. Tiny fans trying to remove all that heat? They gotta…
  4. Kimi K2.7 Code benchmarks at 56.3% on the geometric mean, higher than Kimi K2.6's 48.2%, suggesting solid improvements. Community testing places it among the best open-source entries alongside GLM 5.2. However, it inherits the Kimi family's hardware demands. If you can run it, this is a strong coding-focused model, but expect similar (though slightly improved) infrastructure requirements as K2.6.

    Best when: You already have infrastructure for large models and want Kimi's latest coding-focused release with improved benchmarks.

    Tips

    • Achieves 56.3% on benchmark geometric mean, outperforming Kimi K2.6's 48.2% significantly.
      Source 15
      Benchmark geometric mean - GPT-5.5: 62.7% - Opus 4.8: 62.2% - Kimi K2.7 Code: 56.3% - Kimi K2.6: 48.2%
    • Listed alongside GLM 5.2 as a top open-source coding model in community benchmarks.
      Source 16
      Another round of my coding benchmark. This time it’s three open source entries: Kimi K2.7 Code, GLM 5.2, and the MiniMax M3 that got open weights but that I can’t run at home no matter how hard I squeeze. Before the numbers, the usual context for anyone who parachuted in here, plus an update to the data center soap opera, because there’s news.
      akitaonrailsOpen original ↗
    • Users report Kimi models are very good for many code tasks and the best quality among open weight models.
      Source 17
      The moat right now is model performance and what that means for how many tokens and additional time you spend. I say this as a relatively frequent user of Kimi models and generally a big fan. But on not-yet-gamed benchmarks like DeepSWE, Kimi K2.6 is beaten soundly by Claude Sonnet 4.6 ($3 $15) and even slightly by GPT 5.4 Mini ($0.75 $4.50). There's no question Kimi models are very good for a lot of code tasks. They're the best quality open weight model. But to get similar overall outcomes as…

    Watch out for

    • May require additional tokens and time compared to proprietary alternatives to achieve similar outcomes.
      Source 17
      The moat right now is model performance and what that means for how many tokens and additional time you spend. I say this as a relatively frequent user of Kimi models and generally a big fan. But on not-yet-gamed benchmarks like DeepSWE, Kimi K2.6 is beaten soundly by Claude Sonnet 4.6 ($3 $15) and even slightly by GPT 5.4 Mini ($0.75 $4.50). There's no question Kimi models are very good for a lot of code tasks. They're the best quality open weight model. But to get similar overall outcomes as…
  5. Qwen3 Coder Next is an 80B parameter MoE model with approximately 3B active parameters, making it surprisingly runnable on hardware like the Framework with 128GB RAM. Users report 36 tokens per second on a Strix Halo box, faster than smaller dense models. A strong option for those who can fit it, offering good coding performance in a MoE architecture that balances size and speed.

    Best when: You have 128GB+ RAM and want a fast MoE model specifically optimized for coding tasks.

    Tips

    • Runs at 36 tokens per second on a 128GB Strix Halo box, faster than denser models at similar sizes.
      Source 18
      For Qwen3.5-27b I'm getting in the 20 to 25 tok sec range on a 128GB Strix Halo box (Framework Desktop). That's with the 8-bit quant. It's definitely usable, but sometimes you're waiting a bit, though I'm not finding it problematic for the most part. I can run the Qwen3-coder-next (80b MoE) at 36tok sec - hoping they release a Qwen3.6-coder soon.
      UncleOxidantOpen original ↗
    • The 80B MoE architecture with 3B active parameters allows reasonable memory usage while maintaining capability.
      Source 19
      I've got a Framework with 128GB. Sure, I'm probably not going to run Deepseek V4 flash (though, I could run the 2-bit quant). But there are a lot of models (especially MOE models) that run fine. Even Qwen3.5-122B in a 4 or 5 bit quant can run. The Qwen3-coder-next (~80B IIRC) runs fine (IIRC I'm running a 6-bit quant of that one). The much vaunted Qwen3.6-27B is runnable, but kind of slow (~20t s with MTP), though that's not a memory limitation issue. Yeah, It's be great to have 256GB RAM, but…
      UncleOxidantOpen original ↗
      Source 20
      Are you talking about Qwen3 Coder 30b a3b Instruct from August 2025, which is a non-reasoning model? Or the more recent "Qwen3 Coder Next" from Feb this year with 80b params, 3b active? I found Qwen3 coder next to be quite good on openrouter [1], but couldn't run it locally. [1] https: openrouter.ai qwen qwen3-coder-next
    • Serves 6-8 concurrent requests with full 256k context in production deployments on DGX Spark hardware.
      Source 21
      In our company of 24 employees, we get by with two DGX Sparks. We don't use AI heavily, but each Spark can serve about 6-8 concurrent requests with a full context lenght of 256k, which is decent. We get about ~35 t s depending on the model we use (currently Qwen3.5 122B A10B and Qwen3 Coder Next), but we might set up a smaller model too for simpler tasks. This works for us and will work for years to come. It is not SOTA, but it works darn well for our purposes, and we control the compute and da…
    • Users found it quite good on OpenRouter before local deployment options were available.
      Source 20
      Are you talking about Qwen3 Coder 30b a3b Instruct from August 2025, which is a non-reasoning model? Or the more recent "Qwen3 Coder Next" from Feb this year with 80b params, 3b active? I found Qwen3 coder next to be quite good on openrouter [1], but couldn't run it locally. [1] https: openrouter.ai qwen qwen3-coder-next

    Watch out for

    • Requires 128GB RAM minimum for comfortable local running with reasonable quantization.
      Source 19
      I've got a Framework with 128GB. Sure, I'm probably not going to run Deepseek V4 flash (though, I could run the 2-bit quant). But there are a lot of models (especially MOE models) that run fine. Even Qwen3.5-122B in a 4 or 5 bit quant can run. The Qwen3-coder-next (~80B IIRC) runs fine (IIRC I'm running a 6-bit quant of that one). The much vaunted Qwen3.6-27B is runnable, but kind of slow (~20t s with MTP), though that's not a memory limitation issue. Yeah, It's be great to have 256GB RAM, but…
      UncleOxidantOpen original ↗
      Source 18
      For Qwen3.5-27b I'm getting in the 20 to 25 tok sec range on a 128GB Strix Halo box (Framework Desktop). That's with the 8-bit quant. It's definitely usable, but sometimes you're waiting a bit, though I'm not finding it problematic for the most part. I can run the Qwen3-coder-next (80b MoE) at 36tok sec - hoping they release a Qwen3.6-coder soon.
      UncleOxidantOpen original ↗
    • Some confusion exists between this model and the earlier Qwen3 Coder 30B A3B Instruct from 2025.
      Source 20
      Are you talking about Qwen3 Coder 30b a3b Instruct from August 2025, which is a non-reasoning model? Or the more recent "Qwen3 Coder Next" from Feb this year with 80b params, 3b active? I found Qwen3 coder next to be quite good on openrouter [1], but couldn't run it locally. [1] https: openrouter.ai qwen qwen3-coder-next
  6. Gemma 4 26B A4B is an MoE model designed for efficiency, capable of running on systems with as little as 32GB system RAM and no discrete GPU. Users report 30+ tokens per second on modest hardware. It excels at simpler programming tasks like text-to-SQL and synthetic data generation. However, it's not trusted for one-shot IDE changes, making it better suited as a supervised coding assistant rather than an autonomous one.

    Best when: You have limited hardware resources and want a fast, efficient model for simpler coding tasks under supervision.

    Tips

    • Can run on an i5-8500 with 32GB RAM and no GPU, making it accessible to basic hardware setups.
      Source 22
      Someone was able to run gemma-4-26B-A4B on an i5-8500 with 32 gb ram with NO GPU. Granted this is an extreme example these MoE models are value for money for a lot of use cases. https: www.reddit.com r LocalLLaMA s YontVNVRbL
    • Achieves 30 tokens per second on setups with 8GB VRAM plus 32GB system RAM.
      Source 23
      I have 8GB VRAM, but 32GB sys ram. I can run qwen 3.6 35B at 30 tok s. I also use pi, and it's smart enough to extend itself(multishot and maybe a few tries) For you, you could try gemma-4-26B-A4B
    • Good for simpler programming tasks like text-to-SQL and synthetic data generation.
      Source 24
      Not replaced but supplemented. For off-line coding current setup is pi + ds4-server + DeepSeek-V4-Flash REAP25 (on M2 Max 96gb). For simpler programming related (e.g. text2sql) as well as synthetic data generation, current best for me is llama.cpp + Gemma-4-26B-A4B (on gpu 7900xtx 24gb; sometimes nemotron-cascade-2-30b-a3b for 1M context). That and (dabbling now) auto-research uses lots of tokens. Used to get paused running out of token quotas all the time. The 1st local model I found somewhat…
    • Fast enough to serve as a coding co-pilot on small to medium context tasks with human oversight.
      Source 25
      I use Gemma 4 26B A4B on my Macbook (M4 Pro, 48 GB RAM) to study Rust (and ask other myriad questions). I don't trust it to do a good job in an IDE harness to one-shot anything but the most trivial of changes. Still, it's fast and good enough that it could handle being a "co-pilot" on small to medium context tasks where you've got your hands on the wheel and your eyes on the road — and are driving under the speed limit. That's remarkable given where we were a couple of years ago. I don't think…

    Watch out for

    • Not trustworthy for one-shot IDE changes beyond trivial modifications.
      Source 25
      I use Gemma 4 26B A4B on my Macbook (M4 Pro, 48 GB RAM) to study Rust (and ask other myriad questions). I don't trust it to do a good job in an IDE harness to one-shot anything but the most trivial of changes. Still, it's fast and good enough that it could handle being a "co-pilot" on small to medium context tasks where you've got your hands on the wheel and your eyes on the road — and are driving under the speed limit. That's remarkable given where we were a couple of years ago. I don't think…
    • Struggles with tool use and burns through tokens searching for the right tool in agent scenarios.
      Source 26
      I tried gemma-4-26B-A4B just to see if it could help me read sort my emails on a relatively under-powered setup (16GB VRAM + 32GB RAM) and it's not going well.. the model burns 24K tokens just on searching for the right tool and then dumps the email contents into context - i tried to get it to use code-mode to save context but the code-mode implementation can't save files so it was useless and im going to try to switch to "ssh-mode" into my devbox container. Still relatively new to this, so I'm…
      zaptheimpalerOpen original ↗
    • Code-mode implementation has limitations like inability to save files.
      Source 26
      I tried gemma-4-26B-A4B just to see if it could help me read sort my emails on a relatively under-powered setup (16GB VRAM + 32GB RAM) and it's not going well.. the model burns 24K tokens just on searching for the right tool and then dumps the email contents into context - i tried to get it to use code-mode to save context but the code-mode implementation can't save files so it was useless and im going to try to switch to "ssh-mode" into my devbox container. Still relatively new to this, so I'm…
      zaptheimpalerOpen original ↗
  7. DeepSeek V3 scores a strong 70.2% on the Aider polyglot coding benchmark, ranking #9 of 27, making it a solid choice for multi-language code editing. Users switching to it through OpenRouter report no discernible loss in analytical quality compared to pricier alternatives. However, the full model requires significant hardware resources, and users often access it via API rather than running locally. Good for those who want DeepSeek's capabilities but may prefer API access over local deployment.

    Best when: You want strong benchmark performance on multi-language coding tasks and are flexible about local vs API deployment.

    Tips

    • Scores 70.2% on the Aider polyglot coding benchmark, ranking #9 of 27 models tested.
      Source 27
      Scores 70.2% on the Aider polyglot coding benchmark (#9 of 27), which tests editing real code across many languages.
      Aider polyglot benchmarkOpen original ↗
    • Users report no discernible loss of analytical quality compared to other top models.
      Source 7
      And if you want to dial in a setting in between: I've switched to Kimi K2.6 (now K2.7) and DeepSeek through OpenRouter and Reasonix for pretty much everything, with no discernible loss of analytical quality or utility. However, like many commenters, I don't really believe in vibe-coding, long-horizon agentic one-shot agentic coding, etc. and do not use LLMs for huge generation tasks that involve designing things end-to-end. I also have an MBP with 128 GB of unified memory and do quite a bit of…
    • Available at a fraction of the price of Anthropic and OpenAI through OpenRouter.
      Source 3
      > I have been moving more and more to K2.7 Code and GLM-5.2 the last few weeks. They are often good enough for assistance, very fast, and cheap. I've moved completely to local models that I run with my M1 Mac Studio (64gb ram) some time ago. But for the rare times when I feel the local, quantized Qwen3.6 isn't enough, I just connect to Openrouter and use something like Kimi, GLM or Deepseek for a fraction of the price of Anthropic et al.

    Watch out for

    • Many users access via API rather than running locally due to size constraints.
      Source 3
      > I have been moving more and more to K2.7 Code and GLM-5.2 the last few weeks. They are often good enough for assistance, very fast, and cheap. I've moved completely to local models that I run with my M1 Mac Studio (64gb ram) some time ago. But for the rare times when I feel the local, quantized Qwen3.6 isn't enough, I just connect to Openrouter and use something like Kimi, GLM or Deepseek for a fraction of the price of Anthropic et al.
      Source 7
      And if you want to dial in a setting in between: I've switched to Kimi K2.6 (now K2.7) and DeepSeek through OpenRouter and Reasonix for pretty much everything, with no discernible loss of analytical quality or utility. However, like many commenters, I don't really believe in vibe-coding, long-horizon agentic one-shot agentic coding, etc. and do not use LLMs for huge generation tasks that involve designing things end-to-end. I also have an MBP with 128 GB of unified memory and do quite a bit of…
  8. MiniMax M2.7 offers an interesting alternative for users with substantial memory (128GB+) who are willing to experiment with aggressive quantizations. Users report that even IQ2_XXS quantization produces fewer hallucinated tool calls than Q8 Qwen 3.6 27B. However, it requires 6-bit quantization minimum for programming tasks, and runs slowly (<20 tps) in typical setups. A niche choice for those with specific memory configurations and tolerance for lower speeds.

    Best when: You have 128GB+ memory and want to experiment with an MoE model that tolerates aggressive quantization better than alternatives.

    Tips

    • Produces fewer hallucinated tool call parameters and bizarre invocations than Qwen 3.6 27B even at aggressive IQ2_XXS quantization.
      Source 28
      Strix Halo user here. While Qwen 3.6 27B exhibits remarkable intelligence density, I will still take unsloth's dynamic IQ2_XXS of Minimax M2.7 over Q8_0 Qwen 3.6 27B any day of the week, and this isn't just because of generation speed either. I wrote my own custom harness, and I get hallucinated tool call parameters and bizarre invocations with Q3.6 27B even at Q8_0, but no issues with the IQ2_XXS of M2.7.
    • Usable with 6-bit quantizations and fits in ~128GB memory configurations.
      Source 29
      Highy anecdotal: I have tried various self-hosted models using both vllm and llama.cpp. I am in a situation where I have access to large amount of memory (~320 GB). While experimenting with quantization I found that there is a non-trivial tradeoff between quality and memory footprint. Overall my experience follows the reported pattern of "2-bit is mwah, 4-bit half decent and 6-bit required for programming. Still, although MiniMax-m2.7 is useable with the 6-bit quantizations that unsloth provide…
      pet_the_birdOpen original ↗
      Source 30
      Until now I have not run models that do not fit in 128 GB. I have an Epyc server with 128 GB of high-throughput DRAM, which also has 2 AMD GPUs with 16 GB of DRAM each. Until now I have experimented only with models that can fit in this memory, e.g. various medium-size Qwen and Gemma models, or gpt-oss. But I am curious about how bigger models behave, e.g. GLM-5.1, Qwen3.5-397B-A17B, Kimi-K2.6, DeepSeek-V3.2, MiniMax-M2.7. I am also curious about how the non-quantized versions of the models wit…
    • Good value for coding tasks according to users running it alongside other models.
      Source 31
      >Being open is nice though, even though it doesn't matter that much for folks like me with a single consumer GPU. Of course it matters because that makes coding plans much cheaper than those from Anthropic and OpenAI. For personal use I have coding plans with GLM 5.1, Kimi K2.6, MiniMax M2.7 and Xiaomi MiMo V2.5 Pro and I am getting a lot of bang for the buck.

    Watch out for

    • Requires 6-bit quantization minimum for programming quality, 4-bit being only half-decent and 2-bit poor.
      Source 29
      Highy anecdotal: I have tried various self-hosted models using both vllm and llama.cpp. I am in a situation where I have access to large amount of memory (~320 GB). While experimenting with quantization I found that there is a non-trivial tradeoff between quality and memory footprint. Overall my experience follows the reported pattern of "2-bit is mwah, 4-bit half decent and 6-bit required for programming. Still, although MiniMax-m2.7 is useable with the 6-bit quantizations that unsloth provide…
      pet_the_birdOpen original ↗
    • Runs slowly at under 20 tokens per second on typical hardware configurations.
      Source 32
      I've been using MiniMax M2.7 with vllm on my dual Nvidia Spark cluster. Slow (<20 tps) but functional for most of my use cases.
    • Non-trivial tradeoff between quality and memory footprint requires experimentation.
      Source 29
      Highy anecdotal: I have tried various self-hosted models using both vllm and llama.cpp. I am in a situation where I have access to large amount of memory (~320 GB). While experimenting with quantization I found that there is a non-trivial tradeoff between quality and memory footprint. Overall my experience follows the reported pattern of "2-bit is mwah, 4-bit half decent and 6-bit required for programming. Still, although MiniMax-m2.7 is useable with the 6-bit quantizations that unsloth provide…
      pet_the_birdOpen original ↗

Frequently asked

What's the best local coding model for 24-32GB VRAM?
Gemma 4 31B runs well at 4-bit quantization on 24GB+ VRAM and outperforms Qwen 3.6 27B in real testing for coding and reasoning tasks. GLM 5.2 is another strong option that users report running on setups with as little as 10GB RAM for smaller models.
Can I run Kimi K2.6 on a consumer PC?
Probably not. A usable Q8 quantization requires around 600GB on disk and approximately 700GB of VRAM. Lower quantizations are considered worthless for quality, making this a data center or enterprise server option.
How does GLM 5.2 compare to Claude for coding?
Users report GLM 5.2 matches Claude Sonnet 4.6 in capability, beating proprietary Gemini and older Claude versions. It's considered the best open-weight model by some community members for local coding assistance.
Which model is fastest for coding on a MacBook?
Gemma 4 26B A4B runs quickly on MacBooks with 32-48GB RAM at 30+ tokens per second. Qwen3.6 27B is another popular option that users report running comfortably on M-series Macs with 64GB RAM.
What's the minimum RAM for a usable local coding model?
You can run useful models with 8-16GB VRAM plus system RAM. Gemma-4-26B-A4B has been run on systems with just 32GB system RAM and no GPU, though performance is limited. For comfortable coding, 24-32GB is a practical minimum.

Sources

  1. 1

    Also, GLM 5.2 seems to be the best open-weight model, and it beats proprietary Gemini and older versions of Claude which is amazing. You can have a model at level of Claude Sonnet 4.6 at home without sharing anything, and maybe even uncensor it.

    codedokode · Hacker News · Jul 9, 2026
  2. 2

    Someone posted this fork, which should work in just 10 GB RAM, at least for another big model. I thought of your request and came back to spread the news :) https: github.com ErikTromp colibri-hy3

    kzrdude · Hacker News · Jul 13, 2026
  3. 3

    > I have been moving more and more to K2.7 Code and GLM-5.2 the last few weeks. They are often good enough for assistance, very fast, and cheap. I've moved completely to local models that I run with my M1 Mac Studio (64gb ram) some time ago. But for the rare times when I feel the local, quantized Qwen3.6 isn't enough, I just connect to Openrouter and use something like Kimi, GLM or Deepseek for a fraction of the price of Anthropic et al.

    nozzlegear · Hacker News · Jun 30, 2026
  4. 4

    Just ran and scored 63 3d model generations (via code) across high and no reasoning. 3D Modeling benchmark quickly shows spatial, logic and code performance of the model so I think it's a very good indicator of the quality. Here are the results compared to Gemini 3.5 Flash: Model + config CodeErr gen Cost gen Median time Quality gemini-3.5-flash, low 0.71 $0.18 68s baseline GLM 5.2, reasoning high 0.61 $0.18 289s -6.0% GLM 5.2, reasoning off 1.52 $0.10 126s -13.6% Although it is cheaper, it is…

    ponyous · Hacker News · Jun 17, 2026
  5. 5

    > As of May 2026, how much money do I need to spend to buy hardware to have a local model that is 80% as good as SOTA services for assisting me in writing code? https: llm-stats.com benchmarks swe-bench-verified SOTA (public proprietary models) would be Opus 4.7 at 0.876 80% of that would be around 0.7. These models qualify, and are upwards of 90% as good in benchmarks: DeepSeek-V4-Pro-Max - 1.6T (HuggingFace shows 862B, huh) - 0.806 Kimi K2.6 - 1.1T - 0.802 MiniMax M2.5 - 229B - 0.802 DeepSeek…

    KronisLV · Hacker News · May 5, 2026
  6. 6

    Early benchmarks show tremendous improvement over Kimi K2 Thinking, which didn't perform well on our benchmarks (and we do use best available quantization). Kimi K2.6 is currently the top open weights model in one-shot coding reasoning, a little better than GLM 5.1, and still a strong contender against SOTA models from ~3 months ago (comparable to Gemini 3.1 Pro Preview). Agentic tests are still running, check back tomorrow. Open weights models typically struggle with longer contexts in agentic…

    gertlabs · Hacker News · Apr 20, 2026
  7. 7

    And if you want to dial in a setting in between: I've switched to Kimi K2.6 (now K2.7) and DeepSeek through OpenRouter and Reasonix for pretty much everything, with no discernible loss of analytical quality or utility. However, like many commenters, I don't really believe in vibe-coding, long-horizon agentic one-shot agentic coding, etc. and do not use LLMs for huge generation tasks that involve designing things end-to-end. I also have an MBP with 128 GB of unified memory and do quite a bit of…

    abalashov · Hacker News · Jun 16, 2026
  8. 8

    Yes, for client projects where privacy and security is important, but no enterprise contract: Open code against Infomaniak hosted OSS models: Qwen3.5-122B-A10B-FP8, Kimi-K2.6. I use API keys for billing. It performs like Dec 2025 in terms of my productivity back then.

    ozten · Hacker News · Jun 16, 2026
  9. 9

    People thinking to self-host Kimi K2.6 had better be prepared for how big it is. Q8 K XL quantization for instance is around 600GB on disk. I would bet about 700GB of VRAM needed. Quantizations lower than Q8 are probably worthless for quality. Or 2.05TB on disk for the full precision GGUF. https: huggingface.co unsloth Kimi-K2.6-GGUF If you can afford the hardware to run Kimi K2.6 at any decent speed for more than 1 simultaneous user, you probably have a whole team of people on staff who are al…

    walrus01 · Hacker News · May 3, 2026
  10. 10

    A 4-bit quantization of either Qwen 3.6 27b or Gemma 4 31b will run on a 32GB Mac with a decent-sized, but not full-sized, context. 64GB gets you the full ~256k context and you don't need to quantize your KV cache (though 8-bit quantization of KV may be worth it for performance). The 4-bit QAT version of Gemma 4 has practically identical performance to the full size version or the 8-bit version in most benchmarks and my tests, so there's no reason to run anything else. The 4-bit Qwen is a littl…

    SwellJoe · Hacker News · Jul 2, 2026
  11. 11

    I opted to buy a normal 32GB laptop for this very reason. I know how loud and hot the GPUs in my desktop run when running even smallish models like Qwen 27B or Gemma 4 31B (which is a better model for most than Qwen 3.6, despite the benchmarks). I also have a Strix Halo which doesn't get loud, because it has a single huge fan, but it does get hot. So, there's no way a laptop could work as hard as models make them work, and not be unbearable. Tiny fans trying to remove all that heat? They gotta…

    SwellJoe · Hacker News · Jun 29, 2026
  12. 12

    At 24GB, Gemma 4 31B QAT will be better and give more concise answers. This post is mostly about unquantized results, so it's less relevant and I can't say much about as I haven't tested Qwen or Gemma via cloud API or unquantized locally. All I can say is locally, quantized in a 24GB scenario, Gemma 4 31B is better in my tests which are mostly reasoning or C programming related. Gemma 4 is the only model series at this parameter scale I've seen correctly answer some of these. One of the answers…

    CMay · Hacker News · Jun 29, 2026
  13. 13

    in my experience running models that have been heavily quantized(q4) or altered to some extent has never made me say “wow, this is so amazing”. On the contrary, the model ended up in the thrash bin after a few prompts. I have an RTX 6000 PRO with 96GB, and what I can run comfortably is Qwen 3.6 27B or MoE, Gemma 4 31B. This is as far as it goes when you run the model at full precision and maximum context length. They perform well and you can use them for coding, doing research on the internet a…

    djx22 · Hacker News · Jul 4, 2026
  14. 14

    For personal needs I connected VSCode with llama.cpp running Qwen 3.6 27B or Gemma 4 31B and it's good enough to cancel my cloud subscription. Qwen running on my 1st GPU at q4@176k context from 70 to 50 tok s with MTP, pretty good for coding. Gemma on the other hand is using both GPUs, running q8@64k context, doing document sentiment analysis, summarization, proofreading and translating, at consistent 25 tok s. Somewhat slow but usable for batched workflows. Might get some more once llama.cpp s…

    Kostic · Hacker News · Jun 15, 2026
  15. 15

    Benchmark geometric mean - GPT-5.5: 62.7% - Opus 4.8: 62.2% - Kimi K2.7 Code: 56.3% - Kimi K2.6: 48.2%

    goldenarm · Hacker News · Jun 12, 2026
  16. 16

    Another round of my coding benchmark. This time it’s three open source entries: Kimi K2.7 Code, GLM 5.2, and the MiniMax M3 that got open weights but that I can’t run at home no matter how hard I squeeze. Before the numbers, the usual context for anyone who parachuted in here, plus an update to the data center soap opera, because there’s news.

    akitaonrails · Hacker News · Jun 14, 2026
  17. 17

    The moat right now is model performance and what that means for how many tokens and additional time you spend. I say this as a relatively frequent user of Kimi models and generally a big fan. But on not-yet-gamed benchmarks like DeepSWE, Kimi K2.6 is beaten soundly by Claude Sonnet 4.6 ($3 $15) and even slightly by GPT 5.4 Mini ($0.75 $4.50). There's no question Kimi models are very good for a lot of code tasks. They're the best quality open weight model. But to get similar overall outcomes as…

    DCKing · Hacker News · Jun 12, 2026
  18. 18

    For Qwen3.5-27b I'm getting in the 20 to 25 tok sec range on a 128GB Strix Halo box (Framework Desktop). That's with the 8-bit quant. It's definitely usable, but sometimes you're waiting a bit, though I'm not finding it problematic for the most part. I can run the Qwen3-coder-next (80b MoE) at 36tok sec - hoping they release a Qwen3.6-coder soon.

    UncleOxidant · Hacker News · Apr 22, 2026
  19. 19

    I've got a Framework with 128GB. Sure, I'm probably not going to run Deepseek V4 flash (though, I could run the 2-bit quant). But there are a lot of models (especially MOE models) that run fine. Even Qwen3.5-122B in a 4 or 5 bit quant can run. The Qwen3-coder-next (~80B IIRC) runs fine (IIRC I'm running a 6-bit quant of that one). The much vaunted Qwen3.6-27B is runnable, but kind of slow (~20t s with MTP), though that's not a memory limitation issue. Yeah, It's be great to have 256GB RAM, but…

    UncleOxidant · Hacker News · Jul 10, 2026
  20. 20

    Are you talking about Qwen3 Coder 30b a3b Instruct from August 2025, which is a non-reasoning model? Or the more recent "Qwen3 Coder Next" from Feb this year with 80b params, 3b active? I found Qwen3 coder next to be quite good on openrouter [1], but couldn't run it locally. [1] https: openrouter.ai qwen qwen3-coder-next

    kristianp · Hacker News · Jun 30, 2026
  21. 21

    In our company of 24 employees, we get by with two DGX Sparks. We don't use AI heavily, but each Spark can serve about 6-8 concurrent requests with a full context lenght of 256k, which is decent. We get about ~35 t s depending on the model we use (currently Qwen3.5 122B A10B and Qwen3 Coder Next), but we might set up a smaller model too for simpler tasks. This works for us and will work for years to come. It is not SOTA, but it works darn well for our purposes, and we control the compute and da…

    rsolva · Hacker News · May 15, 2026
  22. 22

    Someone was able to run gemma-4-26B-A4B on an i5-8500 with 32 gb ram with NO GPU. Granted this is an extreme example these MoE models are value for money for a lot of use cases. https: www.reddit.com r LocalLLaMA s YontVNVRbL

    sbmthakur · Hacker News · Jun 16, 2026
  23. 23

    I have 8GB VRAM, but 32GB sys ram. I can run qwen 3.6 35B at 30 tok s. I also use pi, and it's smart enough to extend itself(multishot and maybe a few tries) For you, you could try gemma-4-26B-A4B

    jboss10 · Hacker News · Jun 29, 2026
  24. 24

    Not replaced but supplemented. For off-line coding current setup is pi + ds4-server + DeepSeek-V4-Flash REAP25 (on M2 Max 96gb). For simpler programming related (e.g. text2sql) as well as synthetic data generation, current best for me is llama.cpp + Gemma-4-26B-A4B (on gpu 7900xtx 24gb; sometimes nemotron-cascade-2-30b-a3b for 1M context). That and (dabbling now) auto-research uses lots of tokens. Used to get paused running out of token quotas all the time. The 1st local model I found somewhat…

    ljosifov · Hacker News · Jun 16, 2026
  25. 25

    I use Gemma 4 26B A4B on my Macbook (M4 Pro, 48 GB RAM) to study Rust (and ask other myriad questions). I don't trust it to do a good job in an IDE harness to one-shot anything but the most trivial of changes. Still, it's fast and good enough that it could handle being a "co-pilot" on small to medium context tasks where you've got your hands on the wheel and your eyes on the road — and are driving under the speed limit. That's remarkable given where we were a couple of years ago. I don't think…

    argee · Hacker News · Jun 15, 2026
  26. 26

    I tried gemma-4-26B-A4B just to see if it could help me read sort my emails on a relatively under-powered setup (16GB VRAM + 32GB RAM) and it's not going well.. the model burns 24K tokens just on searching for the right tool and then dumps the email contents into context - i tried to get it to use code-mode to save context but the code-mode implementation can't save files so it was useless and im going to try to switch to "ssh-mode" into my devbox container. Still relatively new to this, so I'm…

    zaptheimpaler · Hacker News · Jun 15, 2026
  27. 27

    Scores 70.2% on the Aider polyglot coding benchmark (#9 of 27), which tests editing real code across many languages.

    Aider polyglot benchmark · Benchmark · Jul 21, 2026
  28. 28

    Strix Halo user here. While Qwen 3.6 27B exhibits remarkable intelligence density, I will still take unsloth's dynamic IQ2_XXS of Minimax M2.7 over Q8_0 Qwen 3.6 27B any day of the week, and this isn't just because of generation speed either. I wrote my own custom harness, and I get hallucinated tool call parameters and bizarre invocations with Q3.6 27B even at Q8_0, but no issues with the IQ2_XXS of M2.7.

    anonym29 · Hacker News · Jun 29, 2026
  29. 29

    Highy anecdotal: I have tried various self-hosted models using both vllm and llama.cpp. I am in a situation where I have access to large amount of memory (~320 GB). While experimenting with quantization I found that there is a non-trivial tradeoff between quality and memory footprint. Overall my experience follows the reported pattern of "2-bit is mwah, 4-bit half decent and 6-bit required for programming. Still, although MiniMax-m2.7 is useable with the 6-bit quantizations that unsloth provide…

    pet_the_bird · Hacker News · Jun 21, 2026
  30. 30

    Until now I have not run models that do not fit in 128 GB. I have an Epyc server with 128 GB of high-throughput DRAM, which also has 2 AMD GPUs with 16 GB of DRAM each. Until now I have experimented only with models that can fit in this memory, e.g. various medium-size Qwen and Gemma models, or gpt-oss. But I am curious about how bigger models behave, e.g. GLM-5.1, Qwen3.5-397B-A17B, Kimi-K2.6, DeepSeek-V3.2, MiniMax-M2.7. I am also curious about how the non-quantized versions of the models wit…

    adrian_b · Hacker News · Apr 23, 2026
  31. 31

    >Being open is nice though, even though it doesn't matter that much for folks like me with a single consumer GPU. Of course it matters because that makes coding plans much cheaper than those from Anthropic and OpenAI. For personal use I have coding plans with GLM 5.1, Kimi K2.6, MiniMax M2.7 and Xiaomi MiMo V2.5 Pro and I am getting a lot of bang for the buck.

    DeathArrow · Hacker News · May 3, 2026
  32. 32

    I've been using MiniMax M2.7 with vllm on my dual Nvidia Spark cluster. Slow (<20 tps) but functional for most of my use cases.

    mv4 · Hacker News · Jun 15, 2026

Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.