Scope and stack
Version 1 reads Korean and English receipts, one receipt per image, in multiple currencies without conversion. It keeps a summary per receipt rather than line items or categories. There is no account, server, or cloud sync.
| UI | Kotlin and Jetpack Compose, minSdk 26, targetSdk 36 |
|---|---|
| Input | ML Kit Document Scanner (one page) or Android Photo Picker (one image) |
| OCR | Bundled ML Kit Text Recognition v2, Korean model |
| Extraction | Korean and English label dictionaries plus a shared rule engine, owned by the app |
| Storage | Room database plus app-private image files |
| Shared code | libs/ai adapters for OCR and the scanner, and a pure-Kotlin core contract shared with the business card app |
The shared modules know nothing about receipts. Receipt labels, the money type, and the merge policy stay inside the app so that a second consumer cannot inherit receipt assumptions by accident.
Input and image lifecycle
The home screen has two separate buttons, Scan and Pick photo. Cancelling the scanner is not a failure and does not open the picker automatically.
- No dangerous permissions. The Document Scanner runs in Google Play services with its own camera access, and the Photo Picker hands over only the selected file. The app manifest declares neither camera nor storage permissions, and the final check is the merged release manifest including the ads SDK.
- Scanner and picker URIs are not stored. The image is copied into app-private staging immediately, because the scanner's file and the picker's grant are both temporary.
- The image is decoded upright from its EXIF orientation before OCR, so the recognition boxes and the displayed image share the same coordinates.
- Recognition is saved as a draft right away and draft edits are auto-saved shortly after each change. Edits to an already confirmed record are not auto-saved, so a half-finished change never overwrites or downgrades it.
- File writes and database writes are not one transaction, so the next start removes leftover staging files, temporary images from an interrupted review, and image files no record points to.
The first-run offline promise applies only to the bundled OCR. The scanner module is delivered by Play services, and the picker can list cloud photo providers. “The app does not send receipts to a server” and “obtaining the image never uses the network” are different statements.
Recognition and field states
One bundled Korean Text Recognition v2 model covers both Korean and Latin script, so English and mixed receipts do not need a second model. Supporting a script is not the same as being accurate on receipts, so English receipts are tested separately.
OCR lines are rebuilt into visual rows. ML Kit splits widely spaced text into separate lines. On a receipt the label and the amount sit at opposite ends of one row, so a synthetic Korean receipt measured on an Android 11 test phone produced a merchant and a date but no amount or tax. Pieces whose vertical overlap is at least half of the smaller height are now joined into one row before the rules run.
- Korean and English dictionaries are applied together to every receipt. The app does not guess a language first and pick one rule set.
- Amount labels are split by role. Totals such as
합계,결제금액,Grand total, andAmount paidare candidates. Components such as subtotal,Balance due,Amount due, tip and suggested tip, service charge, cash tendered, change, discounts, and points are not treated as the transaction amount. - Tax keeps its printed label (VAT, sales tax, GST,
부가세). When several tax rows exist without a stated tax total, the field stays for review instead of being summed. - A date that reads differently as MM/DD and DD/MM is flagged. A missing year or a missing date is never filled in from today or from the capture time.
- Only the last four digits of a card number are recognized as a field.
Each field is Missing, Needs review, or User confirmed. Only the user can produce the last state, and a field the user edited is not overwritten by later recognition. A record is Confirmed only when merchant, date, amount, and currency are all user confirmed and the user saves it as confirmed; everything else is a draft.
The design reserves a Gemini Nano (ML Kit GenAI Prompt) path that would return values plus the IDs of the OCR lines supporting them, validated by the app before use. It is not connected in this baseline: release builds do not include the Prompt module, and a debug-only availability probe on a Galaxy S25 returned UNAVAILABLE. Every result in this version comes from OCR and rules.
Money and currency
Amounts are stored as Money(currencyCode, unscaled: Long, scale: Int), so 12.34 USD is USD / 1234 / 2. The scale is saved with each record and is never reinterpreted from the device's current currency metadata, which can change with ICU and CLDR versions.
- Parsing goes from the printed string to
BigDecimalto the stored scale without passing through floating point and without rounding.12.345is not silently turned into12.35; a mismatch with the currency's usual digits is flagged for review. - The number parser checks grouping positions, spaces and non-breaking spaces, dots, commas, and parentheses, and requires the whole string to be consumed. When two readings are valid, both stay as candidates.
- Currency evidence is ranked: a code or unambiguous mark next to the amount, then the same transaction context, then ambiguous symbols, then hints.
$and¥stay ambiguous, and the app never maps English to USD or Korean to KRW. - The primary currency in Settings is only a suggestion and is not treated as evidence from the receipt.
- Amounts are stored as non-negative magnitudes, and the transaction type (purchase, refund, or unknown) decides the sign in totals, so a printed minus and a refund type are not applied twice.
Two rules were added after testing on a real device. When every explicit currency mark on a receipt is the same currency, that currency is used as context for unmarked total labels and is marked as a context decision for review (the total's 원 had been misread, but 부가세: 0원 on the same slip survived). And when the currency is explicitly zero-decimal, such as KRW or JPY, the reading of 13,500 as 13.5 is dropped; without that, nearly every won amount carried a false ambiguity. Neither rule applies to ambiguous symbols or to two-decimal currencies.
Totals and CSV export
The home screen shows the current month per currency, with spent, refunded, and net amounts, and no conversion between currencies. Only confirmed records whose transaction type is known count. Drafts and records still marked unknown stay out of the totals.
- Totals are added as
BigDecimalvalues. Summing the stored integers directly would treat 12.3 (scale 1) and 12.30 (scale 2) as a tenfold difference. - The CSV has one row per confirmed receipt and a
schema_versioncolumn. The amount is the exact decimal string from the stored scale, with a dot, no grouping, and no symbol, next to its ISO currency code. - Tax is written only when its currency matches the amount's currency; otherwise the cell is empty.
- Text cells that start with
=,+,-, or@are prefixed to block spreadsheet formula injection, and digit-only text with a leading zero is preserved the same way. The file is UTF-8 with a BOM and CRLF rows. - The file is written to the app cache and shared through a FileProvider URI with a temporary read grant, so no storage permission is involved. Export files older than 24 hours are removed at export time and at app start.
Storage, privacy, and backup
- Records live in Room and images in app-private storage. Nothing is written to the gallery.
- Images are kept by default. Deleting a receipt is a hard delete that also removes its image. When image keeping is turned off, receipts recognized afterwards use a temporary image that is deleted when review ends; images kept earlier are not deleted by the setting.
- Structured card data is limited to the last four digits, but a kept image still shows whatever was printed. The app does not promise to redact images.
- The full OCR text is not stored. Evidence highlights exist only during the review that follows recognition; reopening a saved record shows the image and saved fields.
- Images sit under
noBackupFilesDir, which Auto Backup skips. The database is excluded from cloud backup throughfullBackupContenton Android 11 and older anddataExtractionRuleson Android 12 and newer, and from device-to-device transfer as well, so a transfer cannot produce records without their images. - CSV is an export, not a backup. Version 1 has no restore or sync.
Verification and known limits
The money, parser, extractor, row assembly, merge, CSV, review form, and monthly total logic is covered by JVM unit tests using synthetic OCR text and fictional receipts; real receipts are kept out of the repository.
A release-only crash. Build 1.0.0 (1) crashed on launch while debug builds ran normally. ML Kit's firebase-components consumer rule keeps registrar classes without members, and R8 full mode removed their no-argument constructors. Component discovery then failed silently and the text recognizer received a null component. The libs/ai OCR and scanner modules now ship a keep rule for the registrar constructors, and the fix was built as 1.0.0 (2). That version code was used up in Play Console without a release, so 1.0.0 (3) is the build prepared for release.
- Version 1 has no search, no month history beyond the current month, and no duplicate-receipt warning.
- Line items, categories, multi-page or PDF input, currency conversion, and tax or household-ledger formats are outside the scope.
- A foreign card slip can print both the shop currency and the card currency. The app records one amount and currency pair chosen by the user and does not add or split them.
- The GenAI path remains a design until it can be measured on a supported device.
- A fixed month label that read “Month 9” in English was replaced by the locale's month name in 1.0.0 (3).