Thinking Machines Lab Launches Inkling, Open-Weight AI Model with 975B Parameters
Thinking Machines Lab released Inkling on July 15, 2026, an open-weight mixture-of-experts model trained on 45 trillion tokens across text, image, audio, and video.
Inkling Arrives as Open-Weight Competitor
Thinking Machines Lab released its first proprietary AI model, Inkling, on Wednesday morning, July 15, 2026. The model represents the company’s bet against one-size-fits-all AI approaches through a mixture-of-experts architecture.
Inkling is an open-weight model, meaning outside developers and companies can download it and modify it directly. The system comprises 975 billion total parameters, drawing on approximately 41 billion parameters for any given task.
Training and Capabilities
Inkling was trained on 45 trillion tokens of text, image, audio, and video, and reasons natively across all three modalities. The model was trained entirely on Nvidia’s GB300 NVL72 systems.
In its development, Thinking Machines used other open-weight models, including Moonshot AI’s Kimi K2.5, to help generate some early post-training data for Inkling.
Performance and Positioning
On one benchmark, Inkling uses a third as many tokens as Nvidia’s Nemotron 3 Ultra to achieve the same coding performance. However, Thinking Machines explicitly states that Inkling is “not the strongest model available today, closed or open.”
Company Momentum
Thinking Machines now employs roughly 200 people, up from lower levels reported after departures earlier in 2026, including two co-founders who left for OpenAI in January. The rapid product launch reflects the company’s pace: OpenAI took roughly five years from founding to bring tech to market and show revenue; Anthropic took roughly three years; Thinking Machines did the same in approximately nine months.
This milestone comes following a strategic partnership with Nvidia in March 2026 to deploy a gigawatt of Vera Rubin computing capacity.
Open Models Show Financial Sector Promise
Separately, researchers from Bridgewater Associates and an unnamed company took an existing open-source model and trained it further on Bridgewater’s financial expertise. According to a joint evaluation published in late June 2026, the model achieved 84.7% on financial reasoning tests while costing roughly a fourteenth as much to run as top proprietary AI models.
Source: TechCrunch
Developments since publication
-
Moonshot released Kimi K3 on July 16, 2026 as a 2.8-trillion-parameter open-weight mixture-of-experts model with ~16 of 896 experts active per token, native vision, and a 1-million-token context windo Source
-
Kimi K3 is the largest open-weight model released to date, with full open weights scheduled for July 27. Source
-
Moonshot shipped Kimi K2.7 Code on June 12, 2026 with +21.8% improvement over K2.6 on Kimi Code Bench v2, plus a HighSpeed variant with ~6× faster inference. Source
-
Zhipu released GLM-5.2 on June 13, 2026 with MIT-licensed weights landing on Hugging Face around June 17. Source
-
GLM-5.2 scored 91.2% on GPQA Diamond and matches or beats GPT-5.5 on long-horizon coding at roughly 1/6 the cost. Source
-
MiniMax M3 posted 59.0% on SWE-bench Pro in June 2026, the top open-weight score, edging past Kimi K2.6's record of 58.6% from April. Source
-
Qwen family crossed 700 million Hugging Face downloads as of January 2026 with 113,000+ derivative models. Source
-
Qwen 3.7 Max shipped API-only on May 19, 2026 and is not open-weight, while the open Qwen line is Qwen 3.6 (April 2026, Apache 2.0). Source
-
DeepSeek V4 Pro was released on April 24, 2026 under MIT license with 1.6T total / 49B active parameters and a 1-million-token context window. Source
-
DeepSeek V4 Pro achieved SWE-bench Verified 80.6% as of its April 24, 2026 release, the top published open-source SWE-bench Verified score. Source
-
Mistral Large 3 was released on December 2, 2025 under Apache 2.0 license with MoE architecture of 675B total / 41B active and 128K context. Source
-
Moonshot AI unveiled Kimi K3 on July 16, 2026. Source
-
Kimi K3 boasts 2.7 trillion parameters, making it the largest open-weight large language model available at the time of release. Source
-
Moonshot claimed K3 performed competitively with Anthropic's Fable 5 and substantially outperformed Anthropic's Opus 4.8 and OpenAI's GPT 5.6 Sol and GPT 5.5. Source
-
K3 costs $15 per million output tokens compared to $4.40 per million output tokens for z.ai's GLM-5.2 and $0.87 for DeepSeek V4. Source
-
Anthropic's Fable costs $50 per million output tokens, compared to K3 at $15 per million output tokens. Source
-
Moonshot AI raised $2 billion in funding in May 2026, valuing the company at over $20 billion. Source
-
Moonshot's annual recurring revenue exceeded $200 million according to a statement from the company's financial advisor. Source
-
Analysts were not expecting China to produce a model as powerful as Fable until early 2027. Source
-
Inkling is a mixture-of-experts system with 975 billion total parameters, though it only draws on about 41 billion for any given task. Source
-
OpenAI took roughly five years to bring its tech to market and show revenue. Source
-
Anthropic took roughly three years to bring tech to market and show revenue. Source
-
Thinking Machines says it brought tech to market and showed revenue in about nine months. Source
-
Thinking Machines now employs roughly 200 people, up from levels reported after a wave of departures earlier in 2026. Source
-
Inkling is a mixture-of-experts system with 975 billion total parameters. Source
-
Inkling draws on about 41 billion parameters for any given task. Source
-
Inkling was trained on 45 trillion tokens of text, image, audio, and video. Source
-
Inkling reasons natively across text, image, audio, and video. Source
-
Thinking Machines now employs roughly 200 people. Source
-
Two co-founders of Thinking Machines left for OpenAI in January 2026. Source
-
In late June 2026, Bridgewater Associates and Thinking Machines published joint research showing their fine-tuned open-source model scored 84.7% on financial reasoning tests, beating top proprietary A Source
-
Thinking Machines says it brought AI to market and demonstrated revenue in roughly nine months, compared to OpenAI's roughly five years and Anthropic's roughly three years. Source