Tools
hayamimi does real-time multilingual speech-to-text with speaker labels on CPU only, no GPU and no cloud
oboroge0/hayamimi (created 2026-08-25, 201 stars and 14 forks within a day, Python) runs live subtitles, a browser dashboard, speaker labels and translation entirely on CPU using sherpa-onnx, targeting Japanese and multilingual input. The claim that matters for builders is the hardware floor: most real-time diarization plus translation stacks assume a GPU, and a CPU-only pipeline changes where this can be deployed, including on the same machine already saturated by an agent workload. It is single-maintainer and one day old, so treat the latency claims as unverified until someone benchmarks it independently.
Source
↳ Follow the thread