Today we're launching Desert Ant Labs, a European frontier AI lab building opinionated on-device intelligence. We believe the best path to efficient intelligence starts on-device.
We're building small, specialized models for audio, vision, and text – each model answers in milliseconds, and costs nothing to run, so you can put intelligence in every product interaction, without being limited by token cost or inference speed. Small enough to run on a five-year-old phone, fast enough to use on every frame or keystroke, and better than the API call you're already paying for.
The first 18 models are live today (12 stable and six in beta), accessible via one SDK for Swift, Kotlin, and JavaScript. One model per task, each built to be the fastest way to complete that task on a device:
Voz: transcribe 10 minutes of audio in two seconds on an iPhone – 4.7x faster than Whisper – with a start and end time on every word.
Clear: a 9MB model that can turn a five-minute laptop recording into studio quality audio in one second.
Redact: mask names, addresses, and card numbers, in real time, in 27 languages, so they never reach your servers.
Tongue: identify 84 languages from three words, with a 2MB model.
Language ID accuracy, three words in Tongue · 2MB 0.933 293MB detector 0.887 Tongue names the language from three words, scoring 0.933 at 2MB against 0.887 for a 293MB detector.
And that's just to name a few. You can find full specs and benchmarks for the other fourteen, on desertant.com/models and Hugging Face. Every model is free up to 100k monthly active devices. No tokens, no logins.
Personal data caught, by system Redact · 12MB 88.8 GLiNER-PII · 2.3GB 91.1 Rampart · 14.7MB 61.4 OpenAI filter · 3GB 60.2 Redact catches 88.8% of the personal data in a text, close to the 2.3GB GLiNER-PII, from a 12MB model.
... continue reading