
Aleph Alpha Kolibri 1: open German-English MoE model
Aleph Alpha released Kolibri 1 on October 3: a 78.1B-parameter open model with 3.46B active, built for German and English, under the Apache 2.0 license.
Tag archive

Aleph Alpha released Kolibri 1 on October 3: a 78.1B-parameter open model with 3.46B active, built for German and English, under the Apache 2.0 license.
DeepSeek R1: The Open-Source Reasoning Revolution That Changes Everything The...
FreeToken is a MoE serving engine for consumer hardware with CPU-GPU co-execution, expert caching, and runtime VRAM re-allocation, but the R

Alibaba's new omnimodal model reads text, images, audio and video inside a 1M-token context, and prices audio input more than 98% below its predecessor.
Quick Summary: π FreeToken is an edge-native serving engine designed to run large,...
Originally published on andrew.ooo β visit the original for any updates, code snippets that aged...
In April I wrote that your Intel laptop can run LLMs. That post was about 8B models β good little...
DeepSeek-V4-Flash-0731 claims to beat its own V4-Pro preview on all nine listed agentic benchmarks. I verify the table arithmetic row by row, reconcile the 284B vs 304B parameter figures from config.json, and size the local GGUF deployment honestly.
A deep dive into kimi-k3-in-c β a pure C99 inference engine that runs Kimi K3 on a single CPU with zero GPUs. Four architectural reductions achieve a 676Γ memory compression while preserving byte-identical output.
NVIDIA's first diffusion LLM: 60B total, only 3B active per tower. Real VRAM is 32-48GB, not 120GB. RTX 5090 32GB works with Q4; 5 GPUs ranked.
An open pull request to llama.cpp tracks which mixture-of-experts submodels get used most during inference and promotes them to GPU memory on the fly, roughly doubling decode speed on an 8GB card in the author's own tests - while slowing other models
An operator has the complete Kimi K3 checkpoint running across sixteen GB10 mini-workstations wired through a single 400G switch, producing roughly 21 to 25 tokens per second for one user, on hardware with a verifiable floor around $57,200.