A new ternary-quantized model, Ternary Bonsai 2 27B, has been released, built on Qwen3.8 27B and using {-1,0,+1} weights with FP16 group scaling to shrink the model to about 1.76 effective bits per weight and a 5.9GB footprint. Despite being over 9x smaller than its full-precision counterpart, it retains 98.2% of aggregate benchmark performance across reasoning, coding, vision and agentic tasks, and supports a 262K-token context window under an Apache 2.0 license.
prismml.com
· 2026-09-17
A VicOne-sponsored analysis highlights how physical AI systems—robots that use multimodal sensors and AI models to perceive and act—can be manipulated through corrupted training data, tampered infrastructure, or spoofed real-time sensor input, without any visible malfunction. It cites BadVLA, a NeurIPS 2025 study showing how a hidden trigger embedded in a Vision-Language-Action model can cause a robot to behave normally until the trigger appears, then subtly alter its physical movements.
spectrum.ieee.org
· 2026-09-16
DeepSeek has released V4.1-Flash on its API, replacing the earlier V4-Flash and V4-Flash-Vision-Exp models. The new model adds native multimodal capabilities and can be accessed by setting the model parameter to deepseek-flash.
twitter.com
· 2026-09-10
World Labs introduced Atlas, a multimodal autoregressive diffusion transformer pretrained from scratch to jointly handle text, images, video, and 3D data within a shared spatial context. The model can generate camera-controlled videos up to one minute at 1440p, reconstruct real-world scenes from sparse imagery into novel views and explicit 3D, and simulate space-time from video for effects and robotics workflows.
worldlabs.ai
· 2026-09-01
Hebbian Robotics, part of YC's Summer 2026 batch, has released HFlow, an open source SDK designed to manage large-scale multimodal data pipelines for robotics and physical AI. The tool automates orchestration, versioning, quality checks, and provenance tracking for datasets combining video, sensor state, and action data, using MCAP as its standard input/output format. The project is currently pre-v1 but functional end-to-end for local use.
github.com
· 2026-08-31
LAION has published LAION-BVD, a large open video dataset intended to support multimodal research on video, audio, and image models. The organization says the release is meant for non-commercial, academic use only, aimed at improving reproducibility and transparency in a field increasingly dominated by proprietary corporate datasets.
projects.laion.ai
· 2026-08-27
CNET ran a side-by-side comparison of ChatGPT and Gemini's ability to interpret images, including handwritten notes scribbled on an e-ink tablet using a passage from The Hound of the Baskervilles. In this test, ChatGPT struggled to correctly transcribe the messy handwriting, misreading key phrases and producing largely inaccurate output compared to what was actually written.
cnet.com
· 2026-08-24