One photo of his room becomes his voice. Tap an object, and it speaks for him.
A communication aid for adults with aphasia after a stroke: they understand everything but can't find the words. No login, no install.
The 30-second path
- Open roomspeak.edycu.dev on a phone. A laptop works too. It starts in Bahasa Indonesia; tap English at the top first for English phrases.
- Tap one of the four example rooms, or take a photo of a room. In about 4 seconds, rings appear on up to 8 things he'd want to talk about, each with a short first-person phrase. The example rooms are AI-generated.
- Tap Selesai (Done in English). The photo fills the screen.
- Tap an object. The device speaks its phrase and the rest of the room dims. The bottom row (Ya, Tidak, Tolong, Sakit, Toilet) always speaks too.
To edit again, hold the gear button (top right) for 2 seconds. The scene is saved on the device, so the app reopens straight into this mode.
Receipts
Six AI-generated rooms sent to the live production function on 2026-10-04 at 13:03 UTC, one call each, encoded exactly as the app sends them:
- rooms answered
- 6 / 6
- spots found
- 43
- median wall clock
- 4.03 s
- slowest room
- 7.38 s
- provider cost
- $0.00
| Room (AI-generated) | Step | Spots | Time |
|---|---|---|---|
| Living room | 1 | 8 | 7.38 s |
| Stroke-recovery bedroom | 2 | 7 | 3.49 s |
| Dining area | 2 | 7 | 4.28 s |
| Family room at night | 2 | 7 | 3.38 s |
| Terrace | 2 | 8 | 4.07 s |
| Backlit kitchen | 2 | 6 | 3.99 s |
- The fallback ladder did real work.
gemini-3.8-flashanswered 1 room. It declined the other 5, andgemini-3.1-flash-liteanswered each of them in under 4.3 s.deepseek-flashwas not needed. - Cost. Gemini API free tier, so $0.00. If the DeepSeek step answers, one photo costs about $0.001 off-peak or $0.002 at peak. That is an estimate from DeepSeek's published price and one measured call (1,314 input and 1,374 output tokens).
- Tests, all without an API key. 25 Vitest tests, among them 3 regression tests each named for a defect the build hit, a property test over 100,000 generated model answers (0 invalid spots kept), and 3 checks that no API key reaches the browser. Plus 2 Playwright specs in CI and 5 browser scripts.
Every response, with each spot's phrase and box: receipt-2026-10-04.json. How the run was made: DEMO.md in the code repo.
Reproduce it
No clone and no key. This sends the example bedroom to the live function and prints the spots, the model that answered, and the time:
{ printf '{"lang":"en","image":"'; curl -s https://roomspeak.edycu.dev/samples/bedroom.jpg | base64 | tr -d '\n'; printf '"}'; } \
| curl -s https://roomspeak.edycu.dev/api/detect -H 'Content-Type: application/json' --data-binary @- -w '\n%{time_total} s\n'
The full receipt, from a clone of the repo:
npm ci && npx playwright install chromium
npm run receipt # the four example rooms → the live function, one call each
CI runs the tests and browser specs with detection and the device voice stubbed. That is the test suite, not the product. Both commands above call the real function.
Honest limitations
- No private photos on the free tier. Detection runs on the Gemini API free tier. Google's terms for unpaid use let it use submitted images to improve its products and let human reviewers read them, and ask that no personal information be sent. Fine for the AI-generated example rooms; a family's real room needs the paid tier first, where Google does not use prompts or images that way.
- Measured on AI-generated rooms only. The six receipt rooms and the four examples are AI-generated. There is no benchmark on real family photos yet. The caregiver can always remove, rename or add spots by hand.
- A proof of concept, not a medical device. One room, one device, the device's own voice. Without an Indonesian voice installed, phrases show as large text instead of speech. No clinical claims.
Links
- Live approomspeak.edycu.dev
- Storythe landing page
- Pitch deckten slides
- Code on GitHubMIT, planning docs in
devpost/ - Release v0.2.0the version on this page
- Demo videoyoutu.be/ijFneakCODk · 1:48