Split the bills by texting a note.
Nestlit is an expense app for housemates. Type one line, the way you would in a group chat, and it becomes a shared expense. No forms, no keypad.
Now try it yourself on the phone. Type a note, or tap one:
The phone is a web replica with a simplified version of the app's instant parser. Send a note, mark a bill paid, then open the chart tab to see the balances move.
The form is the problem
Expense apps ask for an amount, a title, a category, a payer and a split. That is five decisions for one coffee, so people stop logging the small things and the balance stops being true. Housemates already describe expenses to each other in one line. Nestlit makes that line the input.



How it works
Two parsers work together: a deterministic one for speed, and an on-device language model for messy phrasing.
An instant parser reads the note
It finds the amount, payer, split, category and due date on every keystroke and shows a preview card. No network, no waiting.
The expense posts immediately
What you saw in the preview is what gets saved. Adding an expense never waits on the model.
The on-device model refines it
Apple's Foundation Models re-reads the note and patches the row only if it changes something meaningful. If the model is unavailable, nothing is lost.
What the model is trusted with
A language model is good at loose phrasing and unreliable at dates and identity. So each field has an owner, and the model's output is checked before it touches shared money.
| Field | Owner | Why |
|---|---|---|
| Due date, recurrence | Deterministic | A date detector is more reliable than the model on “due Friday” or “every 2 weeks until Dec”. |
| Payer | Model + code | The model extracts a name; code matches it against the real household. An unknown name falls back to you, never to a guessed person. |
| Name, category, split | Model | This is where phrasing actually varies. Category is a fixed list, so the model cannot invent one. |
| Amount | Parser first | The parser handles 2k, 124,50 and “2 pizzas 30”. A non-positive amount from the model is discarded. |
Product decisions
The interesting choices were mostly about what the AI should not do, and what the product should not have.
The model never blocks the user
An early build waited for the model before saving, and adds stalled for seconds on some phones. Speed on the core action mattered more than a slightly better first guess.
On-device, not cloud
These are a household's finances. The model runs locally: nothing leaves the phone, there is no per-request cost, and it works offline. The price is device coverage.
Working without AI is the default path
On older phones, or on any model error, the app keeps the parser's result. The product is complete without the model; the model only makes it more forgiving.
Show the interpretation first
A wrong parse is visible in the preview before it becomes a shared record, and every expense stays editable afterwards.
Cut: receipt scanning
Built as a stub, then removed. OCR that guesses wrong erodes trust in a money app, and a camera flow fights the text-first idea.
Cut: payments and trips
Settling up records that a debt was paid; the app never moves money. One household per person, focused on the everyday case.
What real use taught me
Using it in a real household through TestFlight surfaced things no design review did.
- “Playstation game 15th due 1000” was booked as 15. The parser read the ordinal as the price. Every misparse found in use becomes a permanent test; the parser now has 41.
- Silence looks like a crash. A note with no amount used to do nothing at all. It now says what is missing.
- Two phones disagree. A housemate who joined late was missing from balances, so the app said “all settled” while money was owed. Balance correctness across devices went ahead of every new feature.
- A green test suite is not a working product. One build shipped where every add failed, because no test touched the backend. Tests now run two simulated devices against one fake backend.
What's next
In order of priority.
- An eval set for the model path. The parser is well covered; the model's refinements are not measured yet. A labelled set of real notes, scored per field, so a prompt change is judged by numbers.
- Measure whether the model earns its place: how often its refinement changes the result, and how often the user then edits it back.
- A fallback for phones without Apple Intelligence, weighed against the privacy promise that makes on-device attractive.
- Accessibility: Dynamic Type and VoiceOver labels.