Jevstiller open-sourced: local model distills Jev with a set disagreement budget
A Show HN project called Jevstiller introduces a local CPU model trained on outputs from the Jev classification API, aiming to answer most requests in about 15 milliseconds instead of the roughly 300 milliseconds a network call to Jev takes. Users set a target agreement rate (e.g. 98%), and the system routes only enough traffic to the local model to keep its combined coverage and error rate within that budget, sending the rest to Jev.
GoKawiil's interpretation of the reporting above, not reported fact.
The approach reframes speed-versus-accuracy tradeoffs in a measurable way: rather than judging the local model against ground truth, it explicitly measures agreement with the original service across all traffic, which the author argues is what can actually be verified. This could matter for latency-sensitive uses like agent loops or real-time systems where every classification call currently costs a fixed round trip, though the benchmarks and guarantees are self-reported by the project's own scripts.
- Jevstiller distills a local model from Jev's outputs to cut per-call latency from ~300ms to ~15ms.
- It enforces a user-set 'disagreement budget' rather than a simple confidence threshold, routing uncertain cases back to Jev.
- The system measures agreement with Jev itself, not accuracy against ground truth, a distinction the project treats as central to its guarantee.
Source: jevstiller.pages.dev, 2026-09-29
Published there as: “Show HN: Jevstiller – Distill Jev into a local model, with a disagreement bound”
Read the original report → The summary and analysis above are GoKawiil's own, written from reporting by the source above. Facts and quotes belong to the original publisher.