General Instinct released InstinctFlash, an open-source inference framework that speeds up robotics 'world-action' models like LingBot-VA and pi05 on Nvidia's Jetson Thor edge computer, RTX 4090 and RTX 5090 GPUs. The company reports benchmarks showing up to a 33.78x speedup over standard PyTorch inference, achieved through FP8 quantization, reduced sampling steps, and custom acceleration kernels, with no measured drop in task performance on real robots.
Wispr Advanced Interfaces Lab has released Canto, a new speech recognition model built specifically for real-time dictation in noisy, everyday conditions rather than clean studio audio. In tests using 10 hours of real Wispr Flow dictations from over 2,300 speakers, Canto posted the lowest word error rate compared with models from Google, OpenAI, AssemblyAI, and Deepgram.
Scientists led by Asor and colleagues developed a technique to directly observe individual virus-like particles as their protein shells form, publishing the results in Nature. The method allowed them to measure the speed and energy changes of shell assembly, something previously very difficult to track experimentally.
Researchers describe Delphy, a phylogenetics tool that encodes evolutionary trees as timed nodes carrying explicit mutation and 'missation' (missing-data) annotations relative to a reference sequence. This design lets the software efficiently reconstruct full sequences at any tree point and handle uncertain tip collection dates, aiming to make Bayesian phylogenetic analysis fast enough for near-real-time outbreak response.
Voicemod has released the Key Pocket, a USB-C hardware dongle that lets Android and iOS users apply real-time voice changing and sound effects during calls, streaming, or mobile gaming. It works as an external audio bridge, routing microphone input through Voicemod's app before passing it to other apps like Fortnite, Discord, or calling apps, sidestepping mobile OS restrictions on altering live mic audio. It costs $99.90 alone or $129.90 bundled with a lifetime Voicemod Pro license.
Researchers at UC San Francisco tested a surgically implanted 253-electrode array that reads brain activity to simultaneously produce on-screen text and animate a personalized avatar's movements. The system, tested with two participants with severe paralysis, is described as the first brain-computer interface to handle verbal and non-verbal communication together within seconds of intent.
Anthropic, founded by Dario Amodei to build AI cautiously and counterbalance riskier rivals, is now facing internal tension between its safety-first founding principles and the commercial pressure to compete aggressively in the AI market. The lab's original structure was meant to insulate it from the very competitive dynamics it now finds itself entangled in.
Linux developer Justin Schroeder shared a photo of an older Intel MacBook pointing its webcam at a mirror reflecting its own screen, letting an AI coding agent visually monitor its progress while tuning AMD Radeon GPU support for the Omarchy Linux distribution. Omarchy is built as an 'agent-first' OS with tools designed to let AI agents install, configure and debug the system autonomously.
Meta Superintelligence Labs has released Muse Voice Transcribe, a real-time speech-to-text model that streams transcription, detects sentence endpoints, and identifies more than 20 distinct speakers without a separate processing step. Trained on over 70 languages (25 thoroughly validated), the model handles long recordings beyond an hour and supports code-switching between languages, all offered via API at $0.18 per hour of audio.
Meta unveiled Muse Voice Transcribe, its first real-time audio perception model, capable of transcribing over 20 speakers simultaneously while switching between more than 20 validated languages, including mid-sentence code-switching. CEO Mark Zuckerberg demoed the tool on X, noting it uses adaptive delay to balance speed and accuracy. The model is now available through Meta's Mac AI app, Muse Code, and its Model API at $3 per 1,000 audio minutes.
Meta Superintelligence Labs has released Muse Voice Transcribe, a streaming speech recognition model that combines transcription, speaker diarization for over 20 voices, and endpointing in a single pass. It supports more than 70 languages, with 25 validated at launch, handles recordings over an hour long, and allows mid-sentence code-switching between languages. The model is live now in Meta AI for Mac, Muse Code, and via the Meta Model API at $3 per 1,000 audio-minutes.
Artie, a Y Combinator-backed startup building real-time data streaming infrastructure, has announced five open positions for Technical Account Executives. The company is seeking candidates focused on engineering quality, speed of execution, and measurable business impact.