Personal productivity

Voice-First Daily Organizer

Can speaking a task beat typing it while walking, driving, or cooking — and when should the app just open the keyboard?

Scale
Small
Platforms
iOS · Android
Capabilities
Mobile · AI · Consumer

The experiment

Voice capture demos beautifully and fails on the sidewalk. The question this experiment asks is narrow and measurable: for tasks and short notes captured while walking, driving, or cooking, can a voice-first flow beat typing on total time to a correct, trusted entry — including the corrections? Not in a quiet room. In wind, over a running tap, with a toddler mid-sentence.

The success metric is deliberately unforgiving: an entry only counts when the user later confirms it captured what they meant. A fast capture the user has to re-read suspiciously that evening is a failed capture with extra steps.

Shape of the build

  • One-gesture capture from lock screen, watch, or headphone press; speak, done
  • On-device transcription first for latency and privacy; a language model then lifts the utterance into structure — task vs. note, due date, list — with confidence marks
  • Everything works offline; structuring catches up when it can
  • A daily review surface where captures are confirmed, corrected, or discarded in seconds

The repair problem

Misrecognition is guaranteed, so repair is the core UX, not an error state. Three rules under test: the app replays what it understood (“Tuesday — call dentist”), not what it heard; correcting one field never requires re-speaking the whole entry; and a low-confidence capture is filed as a raw voice memo with its audio attached rather than guessed into a wrong task. A visible “unsure” beats a confident mistake — the same principle as everywhere else in this portfolio, applied at two seconds per interaction.

When the answer is “open the keyboard”

Part of the experiment is mapping where voice honestly loses: precise times, names the model keeps mangling, anything whispered in an open-plan office, edits to existing entries. The app should detect these moments — repeated corrections, ambient noise, a third retry — and offer the keyboard without ceremony. A voice-first product that admits where voice is worse will be trusted in the moments where voice is better.

Evaluation plan

  • Time-to-trusted-entry, voice vs. keyboard, per context (walking, driving, kitchen)
  • Correction rate and repair time per capture; target hypothesis: median repair under five seconds
  • Abandonment: captures started by voice but finished by keyboard, as a map of where voice fails
  • Retention proxy: after two weeks, do users still choose voice when both are one gesture away? If not, the experiment has its answer.

Facing a similar problem for real?

This study's reasoning — discovery, architecture, evaluation — is exactly what a HummingByte engagement looks like. Bring us the real version.