Tech News
← Home  ·  All topics

Open Weight

27 GoKawiil briefs on this topic

Mistral Unveils 1-Trillion Parameter Multimodal Model to Compete Globally

French AI firm Mistral AI announced the release of Mistral Large 4 (ML4), a multimodal model with one trillion parameters, aiming to surpass both Western and Chinese rivals. Currently accessible only through a controlled endpoint, the company plans to release the model's weights in three weeks after safety evaluations are completed.

Mistral Launches 'Le Chonk', a Competitive Open-Weight AI Model

French AI firm Mistral has unveiled Le Chonk, a one trillion-parameter model available for free, designed to rival top US and Chinese models. The model is currently in preview, with a final version expected soon, and is optimized for coding, cyberdefense, and niche industrial tasks. Mistral claims Le Chonk is the most capable open-weight model outside China, trained from scratch without relying on distillation techniques used by Chinese competitors.

Aleph Alpha releases Kolibri, a 78B-parameter open-weight LLM for German and English

Aleph Alpha, a German AI company, released Kolibri on 3 October 2026 under the Apache 2.0 license, with weights published on Hugging Face. It is a mixture-of-experts model with 78.1 billion total parameters but only 3.46 billion active per token, trained from scratch on 768 NVIDIA B200 GPUs across data centers in Germany and Finland on roughly 24 trillion tokens, over a fifth of them German.

AI startup releases Kolibri, an open-weight 'sovereign' language model

A new open-weight AI model called Kolibri has been released, positioned as a 'sovereign' alternative to existing large language models. The release includes benchmark comparisons against models such as Qwen3.6-35B-A3B, Qwen3-Next 80B-A3B, Nemotron 3 Super 120B-A12B and Mistral Small 4 119B-A6B, with metrics focused on non-hallucination rates and grounding accuracy when source documents lack answers.

Debate grows over 'abliterated' open-weight AI models stripped of safety guardrails

A growing discussion highlights how open-weight AI models can be modified—or 'abliterated'—to remove built-in safety restrictions, allowing users to bypass content moderation. Commentators note this concern extends beyond models originating from China to open-source AI broadly, raising questions about how the industry and regulators should respond.

Nvidia joins Microsoft, Meta, IBM in open-weight AI letter

On July 24, Nvidia signed 'Open Weights and American AI Leadership' alongside Microsoft, Meta, IBM, Hugging Face and Mistral. The letter argues that U.S. AI leadership requires not just powerful frontier models but an open ecosystem that spreads AI capabilities across the economy, distinguishing open-weight models—whose parameters can be downloaded, run and modified under licences like Apache 2.0—from closed models accessed only via API.

Alibaba releases 7B-parameter Qwen Image 2.1 model, restricts commercial resale

Alibaba Cloud has launched Qwen Image 2.1, an open-weight image generation model with only 7 billion parameters, small enough to run on older consumer GPUs like the RTX 3090. The model adds native transparent-background generation and can combine up to 10 reference images into a single consistent output, and Alibaba's own benchmarks claim it outperforms closed models including Google's Nano Banana 2.0. Unlike its predecessor, the new licensing terms bar commercial resale without a separate agreement from Alibaba.

Congressional briefing highlights China's lead in open-weight AI models

An AI researcher briefed members of Congress and staff on the current state of open-weight versus open-source language models, framed within U.S.-China competition. The briefing distinguished closed API-only models like GPT-4 and Claude from open-weight models such as Meta's Llama, Alibaba's Qwen and DeepSeek, and from fully open-source models like the Allen Institute's Olmo, noting Chinese firms have led in open-weight releases since around April 2025.

Study finds older NVIDIA GPUs like A100 retain long-term earning power via open-weight AI models

A new Ornn Data paper argues that NVIDIA's older Ampere-generation GPUs, particularly the A100, remain economically useful far longer than assumed because open-weight models can be run cheaply on them. The analysis found that self-hosted open-weight models like gpt-oss-120b produce output more cheaply on A100 chips than on newer H100 hardware, and that A100 rental prices have stayed unusually stable over a five-year contract period compared to Hopper and Blackwell chips.

MiMo-V2.6-Pro debuts as fast, high-scoring open-weight multimodal model

MiMo-V2.6-Pro is a new open-weight AI model supporting text, image, speech, and video inputs with a 1M-token context window, scoring 46 on the Artificial Analysis Intelligence Index versus a median of 18 among comparable models. It runs at 125 tokens per second, priced at $0.43 per million input tokens and $0.87 per million output tokens, with a total evaluation cost of $206.66.

Research finds data-weighting effects on LM training are non-monotonic across scale

A study examined how assigning different weights to training sequences affects loss reduction in language models of varying sizes, using both proprietary and open-weight models. It found that the relationship between sequence weight and learning is not linear across scale: small models learn general patterns regardless of weighting, medium-scale models learn patterns roughly proportional to their assigned weights, and large models again learn broadly regardless of weighting.

Independent researcher releases Laya, an open-source rival to TypeSafe AI's Jev decision engine

An independent developer has released Laya, an open-weight, non-autoregressive decision model that predicts structured outcomes in under 35 milliseconds using reinforcement learning with calibrated decisions (RLCD) and routing across more than 100 languages. The creator says the underlying approach dates back to a March 2025 arXiv paper and open Hugging Face weights, predating TypeSafe AI's closed, proprietary Jev model launched in September 2026 by AI figure Diogo Almeida's team.