Voice typing on hosts without speech input

The phone listens and recognizes; the PC, TV, or tablet only sees a Bluetooth keyboard typing the sentence you approved.

Source baseline: app 1.0.0 (8) · 23 September 2026

The idea. Speech recognition does not have to run on the machine that receives the text. Dotori Bluetooth Keyboard already makes the phone a standard Bluetooth HID keyboard, so it can recognize speech on the phone and deliver the result the same way a person would: as ordinary key presses. The host needs no microphone, no dictation feature, no driver, and no companion app.

Why it helps

Many devices that accept a keyboard cannot take dictation, or cannot take it in the language you want. The phone in your hand usually can. Moving recognition to the phone turns every keyboard-capable host into a voice-capable one.

Where it works

Phone OSAndroid 12 (API 31) or later. On Android 11 and earlier the microphone key is not drawn at all, because the on-device recognizer API does not exist there.
RecognizerThe phone must provide an on-device speech recognizer (SpeechRecognizer.isOnDeviceRecognitionAvailable). If it does not, the app says so instead of falling back to an online recognizer.
Language modelThe offline model for the chosen language must be installed. Recognition follows the app's current Korean/English state: ko-KR or en-US.
PermissionMicrophone permission, requested the first time you press the microphone key. The app does not lock the feature after a refusal; if Android stops asking, turn the permission on in the app's system settings.
HostAnything that pairs with the phone as a Bluetooth HID keyboard. Korean text additionally needs a host IME that the selected input profile can drive, for example Windows Korean 2-set.
Entry pointsThe microphone button before the help button on the portrait screen, the microphone key on the landscape custom keyboard, and the microphone button in sentence-input mode. All three open the same voice window, so the flow is identical in either orientation.

On Android 13 and later the app asks the recognizer before listening whether the language model is installed, missing, or unsupported. A model that is missing or still downloading is reported with a Download action that asks the system to fetch it. Android 12 and 12L (API 31–32) have no such query, so readiness is learned from the first real attempt; if the model is not ready, the app tells you to check the offline speech-recognition languages in the phone's settings.

Technical view: on-device only, by construction

The recognizer is created with SpeechRecognizer.createOnDeviceSpeechRecognizer, added in API 31. The app checks isOnDeviceRecognitionAvailable first and stops with an “engine unavailable” state when it is false. There is deliberately no code path to the ordinary createSpeechRecognizer, which may use a network service: sending what the user says off the device would change the premise of the feature, and the app has no backend of its own.

Supporting three API ranges (28–30, 31–32, 33+) in one binary needs care. The app's minSdk is 28, so phones on API 28–30 load the same classes. A method added in a newer API is resolved only when called and can sit behind an SDK_INT check; a newer type in a field or signature fails when the class loads, before any check runs. The controller therefore implements only RecognitionListener (API 1), and the API 33 types RecognitionSupport and RecognitionSupportCallback live in a single isolated object that is touched only inside an API 33 branch. A contract test keeps those names out of every other voice file.

Language tags are compared loosely. The recognizer may report ko_KR where the app asked for ko-KR, in any letter case. Comparing strings exactly would report an installed language as missing, so tags are normalized before the installed / pending / supported lists are matched. A failed support query is treated as “unknown” and the app simply tries, rather than declaring the language unsupported.

Why nothing is sent automatically

Partial recognition results are unstable: the recognizer revises earlier words as more audio arrives. On a normal phone keyboard that is harmless because the text is still editable. Over HID it is not — a key press that has reached the host cannot be recalled by the phone. So recognized text is shown, never typed, until the user confirms it.

At the same time, asking for confirmation after every sentence interrupts dictation. The controller resolves this with continuous listening:

From confirmed sentence to keystrokes

Voice adds no new transmission path. A confirmed sentence goes through the same sentence-send routine as typed sentence input: the text is normalized, split into Korean and English runs, modern Hangul syllables are decomposed into 2-set jamo keys, and the plan includes the IME switches the host profile needs. If any character cannot be encoded, nothing is sent — the host never receives half a sentence — and the text stays in the voice window so it does not have to be spoken again.

HID typing is paced key by key, so a long sentence takes seconds to arrive. While it is being sent the button shows “Sending…” and cannot be pressed again, which prevents the same sentence from being typed twice. Once the sentence is accepted for sending, the buffer is cleared and a fresh recognition session starts, so a partial that was still inside the recognizer is not delivered a second time. Clearing happens at acceptance, not at delivery: if the HID transfer later fails partway, that sentence does not return to the window.

The recognition language is taken from the Korean/English state the app already maintains for the host IME. The host remains the owner of IME state; if the display and the host drift apart, the existing fix — long-pressing the Korean/English key to realign the display — applies to voice as well, for the same reasons as typed input.

Session safety

Limits

Official sources