Running Qwen3.8-Flash-Next on a 128 GB Mac: The Expert-Pruning Trap, and a Memory-Mapped n-gram Table That Gets You to 240K Tokens
Hello, everyone. Have you ever wanted to run the smartest model you can on your own Mac? I have. The...
Tag archive
Hello, everyone. Have you ever wanted to run the smartest model you can on your own Mac? I have. The...
Hello, everyone. Point a camera at something and get a sentence back describing what is happening...

Goose Swarm 3.0 runs local models on its own MLX engine, joins your Macs with LeanZero Link and serves an OpenAI-compatible goose endpoint.

Same 27B model as GGUF and MLX, every app graded by running it: 13 of 15 work on both. Overall speed is a wash; MLX was steadier and never hit the cap.

Qwen3.8-27B tuned for Forge and Jira: v0.5 beat v0.4 on every gate. GGUF for llama.cpp and Ollama, Q8_0 and Q6_K, now published alongside MLX.

How to fine-tune Qwen3.8-27B with LoRA on a Mac using MLX: 8-bit base with MTP kept, rank 32 vs 128, what a 96 GB M3 Ultra fits, the rule that picks the round.

Qwen3.8-27B tuned for Forge: compiling apps 12 to 21 of 25, valid manifests 14 to 23 of 25. Rejected first by our own rule on a loop that was an ASCII diagram.
Originally published on andrew.ooo — visit the original for any updates, code snippets that aged...

The MXFP8 MLX quant of Qwen3-Coder-Next (80B total, 3B active, 262K context) on a 128GB Mac: LM Studio 0.4.1 setup for Claude Code.
NVIDIA's Nemotron-3-Nano-Omni-30B-A3B is an open-weights model that sees, hears and reasons. There is...
「Ollamaが新しくMLXバックエンドに対応して、Apple SiliconでのローカルLLM推論が速くなった」という話をたまに見かける。手元はM1 Max...
MLX + Ollama 0.30.8 makes Apple Silicon competitive with CUDA. M4 Max 64GB runs 70B Q4. Ranked M3/M4/Pro/Max/Ultra RAM tiers for local LLM 2026.