When the words won't come, his room speaks for him.

After a stroke, he understands every word said to him but can't find his own. A caregiver photographs his room once. Gemini rings up to 8 things he'd want to talk about and drafts a phrase for each, which the family can change. Then one tap on the kettle, and the device says it for him.

No accountThe scene stays in this browserAI is called once, at setup

His living room in afternoon light: a teak sideboard with a radio, an enamel kettle on a round table, a rattan chair, a walking frame and a standing fan. The room is dimmed; only the kettle stays lit inside its ring.
Aku mau kopi.I'd like coffee. Tap any ring to hear it, in Indonesian, the way he would.

Room photo AI-generated, so no family's home is published. Rings: Gemini's recorded answer for this photo, replayed (the page never calls the AI). The kettle's phrase was rewritten by hand, as a caregiver would. English glosses are ours.

8
things ringed from one photo, at most
shared/validate.ts
4.48s
from photo sent to answer back, in the recorded call below
media/detect-run.json
25
automated tests: ring maths, storage, the fallback ladder, no API key in the browser
tests/
46
checks in a real browser: speaking, editing, saving, failing, example rooms
e2e/

One photo in. Seven rings out.

This is the answer the app's own /api/detect returned for this photo, recorded on 4 October 2026 and replayed here. It came from the second model on the app's fallback ladder: in 4.48 s it named 7 things, wrote a phrase for each and drew a box around each one. The app turned every box into a ring.

The app's Setup screen: the same room with seven numbered terracotta rings, on the kettle, radio, fan, chair, window, walking frame and snack jar.
Setup, as the caregiver sees it. Ring numbers follow the order of the answer on the right.

The full response is public: media/detect-run.json. A second run on the same photo (detect-run-2.json) answered in 3.73 s with the same seven things.

Gemini drafts. The family decides.

For the kettle, Gemini guessed a drink: “Boleh minta minum?”, may I have a drink? The caregiver knows he means coffee. Tap the row, Ubah, type, Simpan, and the kettle now says “Aku mau kopi.”

Every phrase can be heard, rewritten or removed. A thing the AI missed is one tap on the photo away. A room holds up to 12 spots.

Gemini's draftSetup row 1, teko: “Boleh minta minum?”, with Ubah and Hapus buttons.
The caregiver's editThe same row being edited: the field reads Aku mau kopi, next to a Simpan button.

One tap, and the room goes quiet.

In Speak mode his room fills the screen. When he touches the kettle the rest of the photo dims, the kettle stays lit inside its ring, and the device speaks for him. No menus, no typing, nothing that times out.

Speak mode on a tablet: the room dimmed, the kettle lit inside an enlarged terracotta ring, and five large core-word buttons along the bottom: Ya, Tidak, Tolong, Sakit, Toilet.

The five words along the bottom never leave the screen, even when the photo has nothing useful. Tap one on the screen above, or tap the kettle.The five words along the bottom never leave the screen, even when the photo has nothing useful. Tap one:

No Indonesian voice on the device?

Setup says so once. In Speak mode the phrase then appears in large text instead, so the tap still says something.

Speak mode without a matching voice: “Aku mau kopi” shown in large text over the dimmed room.

Skip the setup. Open this room in the real app.

The same seven rings, the kettle already saying its phrase. Tap the kettle. To get back to Setup, hold the gear button in the corner for 2 seconds.

Saves this room in this browser, the way the app does, then opens the app. You'll be asked before it replaces a room you already saved. Or open the app and tap one of its four example rooms.

Every screen, straight from the app.

Six frames from the running app, in order: the recorded answer replayed, the kettle edited through the app's own editor, nothing drawn by hand.

First open: a welcome line, a language choice, one button to take or choose a room photo, and four example rooms labelled as AI-generated.
1First opena photo, or one of 4 example rooms
The photo with a calm Mencari benda (looking for objects) indicator.
2Looking“Mencari benda…”
Setup: seven numbered rings on the photo and the list of phrases, with the Selesai (done) button.
3Seven ringshear, edit, remove, add
A one-time card for the caregiver, Siap dipakai (ready to use), explaining the 2-second hold.
4Ready to useshown once, to the caregiver
Speak mode on a phone: the room with rings and the five core words.
5Speak modehis screen, every day
Speak mode after tapping the kettle: the room dimmed, the kettle lit inside its ring.
6“Aku mau kopi.”one tap, spoken

What happens between a photo and a voice.

The AI does the slow part once: finding the things in a room and writing the words. After that, everything he does happens on the device.

Setuponce, by the caregiver, online
Room photo shrunk to ≤1600 px POST /api/detect Gemini: label, phrase, box_2d checked, at most 8 box → ring, nudged apart saved on this device
Speakevery day, by him, on the device
a tap the ring under the finger, on release speechSynthesis in id-ID the room dims, the ring pulses

Rings on his real things

Gemini returns a box for each thing it finds. The app turns each box into a ring centred on it, sized to it within calm bounds, and nudges overlapping rings apart so each stays reachable.

box_2d [502, 575, 708, 685] x .630 y .605 r .033
box_2d → ring · boxToSpot + separateSpots

When Gemini is busy

The detect helper tries each model in turn, inside a 55 s budget. Times as measured on 6 test rooms.

  1. 1
    gemini-3.8-flashtightest boxes, ~4–12 s
  2. 2
    gemini-3.1-flash-lite~3–10 s · answered the call on this page
  3. 3
    deepseek-flashslightly looser boxes, ~6–10 s

Every answer is checked before the app sees it: four numbers in 0–1000, a real phrase, at most 8 items. Malformed items are dropped, never repaired.

api/detect.ts · shared/validate.ts

A tap counts on release

Only a finger that lifts inside the ring it started in speaks. A second finger cancels, so a resting palm stays quiet.

pointerup · nearestRing

The voice lives on the device

Phrases are spoken by the device's own voice, interrupting the last one. The scene is kept in this browser, not on a server.

speechSynthesis · localStorage

Five words that never leave

Yes, no, help, pain, toilet: always on screen, even when the AI found nothing useful.

CORE_WORDS

A way back he won't hit by accident

Setup opens only after a 2-second hold on the corner button. A short press just explains itself.

HoldButton · 2 s

Honest about what it is.

Lifted word for word from the limits section of the README.

A communication-aid prototype, not a medical device or therapy; no clinical claims.

README · Limits of this proof of concept

One room, one device, device voice only. A recorded family voice, more scenes and sharing between family members are out of scope.

README · Limits of this proof of concept

Detection quality depends on the photo; the caregiver can always remove, rename or add spots by hand.

README · Limits of this proof of concept

Planned in the open, before the first line of code.

Built with the Devpost Learn skill pack. The planning documents live in the repo, including every plan change the build forced.

Questions a judge would ask.

Does it work without the internet?

Speaking runs on the device: the tap, the device's voice and the saved scene. The internet is needed once, during setup, when Gemini looks at the photo. The app isn't installable for offline use yet; that's on the scope's "later" list.

Can I try it without a photo of a room?

Yes. The app's first screen offers four example rooms, labelled as AI-generated; tapping one runs the real setup on that photo. Or use the button further up this page to open this room, already set up.

Where does the photo go?

Once to the app's own detect helper, which asks Gemini for the objects and keeps the API key off the device. The scene (photo, rings, phrases) is then kept only in this browser's storage. No account, no server database.

What if Gemini rings the wrong things?

The caregiver removes, rewrites or adds spots by hand: a new spot is a tap on the photo and a typed phrase. If the AI finds nothing or the request fails, Setup says so plainly and adding spots still works. The five core words work with zero spots.

What if the device has no Indonesian voice?

Setup tells the caregiver once, and Speak mode shows the phrase in large text when a spot is tapped. Android Chrome and iOS Safari ship an Indonesian voice. English is available at setup.

Is this a medical device?

No. It's a communication-aid proof of concept, built on visual scene displays, a known approach in AAC (tools that speak for people who can't). It makes no clinical claims.

What on this page is real?

The app, every screen, and the seven rings: they are Gemini's recorded answer for this photo, replayed so the page never calls the AI. The room itself is AI-generated so that no family's home is published, and one phrase was rewritten by hand, as a caregiver would.

Aku mau kopi.

I'd like coffee. One photo, one tap, his words.