The idea. Speech recognition does not have to run on the machine that receives the text. Dotori Bluetooth Keyboard already makes the phone a standard Bluetooth HID keyboard, so it can recognize speech on the phone and deliver the result the same way a person would: as ordinary key presses. The host needs no microphone, no dictation feature, no driver, and no companion app.
Why it helps
Many devices that accept a keyboard cannot take dictation, or cannot take it in the language you want. The phone in your hand usually can. Moving recognition to the phone turns every keyboard-capable host into a voice-capable one.
- Hosts with no speech input at all. Smart TVs, set-top boxes, and many embedded or kiosk-style systems accept a Bluetooth keyboard but offer no dictation. Search boxes and login fields on a TV are exactly where typing with a remote hurts most — provided the TV lists the phone as a keyboard when pairing, which some Google TV models do not.
- PCs you cannot change. A shared or managed work computer may have no microphone, may block new software, or may not offer dictation in your language. The phone needs nothing installed on the host.
- Any app that accepts typing. Because the host only receives key presses, the text lands wherever the cursor is — a browser, a terminal, a document, a chat window — without per-app integration.
- Your voice stays on the phone. The app uses only Android's on-device recognition API and never falls back to a network recognizer. It sends no audio to the developer or to the host; the host receives the final text and nothing else.
- You review before it is typed. Recognized sentences collect in a window on the phone. Nothing reaches the host until you press Send, so a misheard word never becomes a keystroke you have to undo on the other machine.
Where it works
| Phone OS | Android 12 (API 31) or later. On Android 11 and earlier the microphone key is not drawn at all, because the on-device recognizer API does not exist there. |
|---|---|
| Recognizer | The phone must provide an on-device speech recognizer (SpeechRecognizer.isOnDeviceRecognitionAvailable). If it does not, the app says so instead of falling back to an online recognizer. |
| Language model | The offline model for the chosen language must be installed. Recognition follows the app's current Korean/English state: ko-KR or en-US. |
| Permission | Microphone permission, requested the first time you press the microphone key. The app does not lock the feature after a refusal; if Android stops asking, turn the permission on in the app's system settings. |
| Host | Anything that pairs with the phone as a Bluetooth HID keyboard. Korean text additionally needs a host IME that the selected input profile can drive, for example Windows Korean 2-set. |
| Entry points | The microphone button before the help button on the portrait screen, the microphone key on the landscape custom keyboard, and the microphone button in sentence-input mode. All three open the same voice window, so the flow is identical in either orientation. |
On Android 13 and later the app asks the recognizer before listening whether the language model is installed, missing, or unsupported. A model that is missing or still downloading is reported with a Download action that asks the system to fetch it. Android 12 and 12L (API 31–32) have no such query, so readiness is learned from the first real attempt; if the model is not ready, the app tells you to check the offline speech-recognition languages in the phone's settings.
Technical view: on-device only, by construction
The recognizer is created with SpeechRecognizer.createOnDeviceSpeechRecognizer, added in API 31. The app checks isOnDeviceRecognitionAvailable first and stops with an “engine unavailable” state when it is false. There is deliberately no code path to the ordinary createSpeechRecognizer, which may use a network service: sending what the user says off the device would change the premise of the feature, and the app has no backend of its own.
Supporting three API ranges (28–30, 31–32, 33+) in one binary needs care. The app's minSdk is 28, so phones on API 28–30 load the same classes. A method added in a newer API is resolved only when called and can sit behind an SDK_INT check; a newer type in a field or signature fails when the class loads, before any check runs. The controller therefore implements only RecognitionListener (API 1), and the API 33 types RecognitionSupport and RecognitionSupportCallback live in a single isolated object that is touched only inside an API 33 branch. A contract test keeps those names out of every other voice file.
Language tags are compared loosely. The recognizer may report ko_KR where the app asked for ko-KR, in any letter case. Comparing strings exactly would report an installed language as missing, so tags are normalized before the installed / pending / supported lists are matched. A failed support query is treated as “unknown” and the app simply tries, rather than declaring the language unsupported.
Why nothing is sent automatically
Partial recognition results are unstable: the recognizer revises earlier words as more audio arrives. On a normal phone keyboard that is harmless because the text is still editable. Over HID it is not — a key press that has reached the host cannot be recalled by the phone. So recognized text is shown, never typed, until the user confirms it.
At the same time, asking for confirmation after every sentence interrupts dictation. The controller resolves this with continuous listening:
- When the recognizer returns a final result and closes its session, the sentence is appended to a buffer (with a space between sentences) and listening restarts immediately.
- The window shows the buffer plus the current partial. The Send button is always there while listening and sends the whole buffer; a very long buffer may be shortened on screen, but it is sent in full.
- Silence and speech timeouts are normal pauses, not failures, and do not count toward the restart limit. Immediate repeated failures — for example another app holding the microphone — stop listening after five restarts without a result. Sentences already collected stay in the window and can still be sent; press the microphone again to resume listening.
- When the app restarts the session after a send, the cancel it issues was observed to produce an error callback of its own; the app ignores that one error so it cannot trigger a restart loop.
From confirmed sentence to keystrokes
Voice adds no new transmission path. A confirmed sentence goes through the same sentence-send routine as typed sentence input: the text is normalized, split into Korean and English runs, modern Hangul syllables are decomposed into 2-set jamo keys, and the plan includes the IME switches the host profile needs. If any character cannot be encoded, nothing is sent — the host never receives half a sentence — and the text stays in the voice window so it does not have to be spoken again.
HID typing is paced key by key, so a long sentence takes seconds to arrive. While it is being sent the button shows “Sending…” and cannot be pressed again, which prevents the same sentence from being typed twice. Once the sentence is accepted for sending, the buffer is cleared and a fresh recognition session starts, so a partial that was still inside the recognizer is not delivered a second time. Clearing happens at acceptance, not at delivery: if the HID transfer later fails partway, that sentence does not return to the window.
The recognition language is taken from the Korean/English state the app already maintains for the host IME. The host remains the owner of IME state; if the display and the host drift apart, the existing fix — long-pressing the Korean/English key to realign the display — applies to voice as well, for the same reasons as typed input.
Session safety
- Host changes discard the buffer, but do not interrupt. When the input session changes — another host, a new connection generation, or a protocol epoch — the sentences collected so far are dropped. Listening itself continues, because the connection becoming ready would otherwise cut the user off mid-sentence; a sentence being spoken at the moment of the change can therefore still land in the new buffer, and nothing is typed until you press Send.
- Leaving the app stops listening. The recognizer is released when the screen goes away and cancelled on
ON_STOP.ON_PAUSEis not used, because the microphone permission dialog pauses the activity and would cancel the request the user just approved. - Late results after leaving or cancelling are dropped. When the window is cancelled or the screen goes away, the recognizer's listener is detached before it is cancelled, so a result that arrives afterwards is never delivered.
Limits
- Accuracy, punctuation, and supported languages depend on the phone's on-device recognizer and its installed models, not on the app.
- The app does not know which field has focus on the host; text goes wherever the host cursor is.
- Only characters that the selected key map and the modern-Hangul plan can encode are sent.
- Typing speed is bounded by HID pacing, so dictating a long paragraph is slower to deliver than it is to speak.
- Phones without an on-device recognizer, or on Android 11 and earlier, cannot use the feature.