多个开源模型发布:LongCat-Flash-Lite等
原文:Uncensored Multi-Model Releases, LongCat-Flash-Lite-Sparse with MTPs and LSAs, Qwen3.8-27B with MTPs, Qwen3.5-122B-A10B with MTPs, Qwen3-Coder-Next and Laguna-S2.1 with Vision, All Available in GGUF Format! Bonus: Links to my llama.cpp Fork for LongCat-Flash-Lite Support and J-Wash Enhanced Fork!
过去一个月我非常努力地为社区带来了所有这些模型。最难做的肯定是 LongCat-Flash-Lite-Sparse,它需要大量的工作:首先我需要从零开始为它创建 Heretic 支持,还必须在 llama.cpp 上为它添加支持,这是相当困难和耗时的任务!它比我几周前发布的原版 LongCat-Flash-Lite 还要难做。它和原版 LongCat-Flash-Lite 一样仍是 69B-A3B 模型,但 LongCat-Flash-Lite-Sparse 新增了以下支持:- 稀疏注意力(LongCat-Flash-Lite 为密集注意力)- 1M 上下文长度(LongCat-Flash-Lite 为 256k)总之,LongCat-Flash-Lite-Sparse 在主线/上游 llama.cpp 中完全不受支持,因此要使用这些 GGUF,你需要从 GitHub 拉取我的 fork,地址在这里:https://github.com/erm14254/llama.cpp-minimax-m3-combined/tree/claude/longcat-win11 你需要通过 llama-server.exe 加载模型,并通过 llama-ui 与它交互。你有两个变体可选:无审查 Heretic(9/100 拒答率,KLD 为 0.0157)和 Ultra Uncensored HJeretic(4/100 拒答率,KLD 为 0.0779),两个变体都带有 MTP 和 LSA!模型链接如下:无审查 Heretic GGUF:https://huggingface.co/llmfan46/LongCat-Flash-Lite-Sparse-Uncensored-Heretic-Native-MTP-And-LSA-Preserved-GGUF Ultra Uncensored Heretic GGUF:https://huggingface.co/llmfan46/LongCat-Flash-Lite-Sparse-Ultra-Uncensored-Heretic-Native-MTP-And-LSA-Preserved-GGUF ---------------------------------------- LongCat 部分就到这里,接下来是带 MTP 的 Qwen3.8-27B Ultra Uncensored Heretic,拒答率 3/100,KLD 为 0.0244,链接如下:Safetensors:https://huggingface.co/llmfan46/Qwen3.8-27B-Ultra-Uncensored-Heretic-Native-MTP-Preserved GGUF:https://huggingface.co/llmfan46/Qwen3.8-27B-Ultra-Uncensored-Heretic-Native-MTP-Preserved-GGUF NVFP4:https://huggingface.co/llmfan46/Qwen3.8-27B-Ultra-Uncensored-Heretic-