Tech News
← Home  ·  All topics

Multimodal

7 GoKawiil briefs on this topic

Bonsai 2 27B compresses a 27B model to 5.9GB with 98.2% performance retained

A new ternary-quantized model, Ternary Bonsai 2 27B, has been released, built on Qwen3.8 27B and using {-1,0,+1} weights with FP16 group scaling to shrink the model to about 1.76 effective bits per weight and a 5.9GB footprint. Despite being over 9x smaller than its full-precision counterpart, it retains 98.2% of aggregate benchmark performance across reasoning, coding, vision and agentic tasks, and supports a 262K-token context window under an Apache 2.0 license.

Researchers warn AI-driven robots face new hidden 'backdoor' attack risks

A VicOne-sponsored analysis highlights how physical AI systems—robots that use multimodal sensors and AI models to perceive and act—can be manipulated through corrupted training data, tampered infrastructure, or spoofed real-time sensor input, without any visible malfunction. It cites BadVLA, a NeurIPS 2025 study showing how a hidden trigger embedded in a Vision-Language-Action model can cause a robot to behave normally until the trigger appears, then subtly alter its physical movements.

DeepSeek launches V4.1-Flash with built-in multimodal support via API

DeepSeek has released V4.1-Flash on its API, replacing the earlier V4-Flash and V4-Flash-Vision-Exp models. The new model adds native multimodal capabilities and can be accessed by setting the model parameter to deepseek-flash.

World Labs unveils Atlas, an omni world model for 3D generation and reconstruction

World Labs introduced Atlas, a multimodal autoregressive diffusion transformer pretrained from scratch to jointly handle text, images, video, and 3D data within a shared spatial context. The model can generate camera-controlled videos up to one minute at 1440p, reconstruct real-world scenes from sparse imagery into novel views and explicit 3D, and simulate space-time from video for effects and robotics workflows.

YC-backed Hebbian Robotics launches open-source HFlow SDK for robotics data pipelines

Hebbian Robotics, part of YC's Summer 2026 batch, has released HFlow, an open source SDK designed to manage large-scale multimodal data pipelines for robotics and physical AI. The tool automates orchestration, versioning, quality checks, and provenance tracking for datasets combining video, sensor state, and action data, using MCAP as its standard input/output format. The project is currently pre-v1 but functional end-to-end for local use.

LAION releases large-scale open video dataset for academic research

LAION has published LAION-BVD, a large open video dataset intended to support multimodal research on video, audio, and image models. The organization says the release is meant for non-commercial, academic use only, aimed at improving reproducibility and transparency in a field increasingly dominated by proprietary corporate datasets.

CNET Test Finds Gemini Outperforms ChatGPT on Visual Recognition Tasks

CNET ran a side-by-side comparison of ChatGPT and Gemini's ability to interpret images, including handwritten notes scribbled on an e-ink tablet using a passage from The Hound of the Baskervilles. In this test, ChatGPT struggled to correctly transcribe the messy handwriting, misreading key phrases and producing largely inaccurate output compared to what was actually written.