Open-weights milestone
Thinking Machines releases Inkling, a 975B MoE open-weights model with 1M token context
Thinking Machines released Inkling, a 975B total / 41B active parameter open-source MoE model with up to 1M token context and native multimodal reasoning over text, images, and audio. The model features ShortConv attention, shared expert sink MoE, and day-one production support via SGLang with CUDA graph prefill, MXFP8 KV cache, and DFlash speculative decoding. Open weights are available immediately.



