This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
The Person I Built This For
My thatha has been cooking sambar for about fifty years and has never once written a recipe down.
What he does instead is send voice memos. Every few weeks one lands in the family group chat, four or five minutes long, usually recorded while he's standing over a pot. "Add a little hing. Not that much. A bit less. You'll know." Then someone replies with a thumbs up, the chat moves on, and the recipe sinks under forty photos of a cousin's baby.
I've been scrolling back through that chat for a year looking for his sambar. This weekend I finally did something about it.
What I Built
Thatha's Kitchen takes one of his voice memos and turns it into a recipe card: title, ingredients, steps, and the little asides he's famous for ("don't rush the onions, they'll punish you").
The part I care about most is that it doesn't guess. When he says "a handful of dal," the card says "a handful of dal" and adds a flag: ask Thatha how big his hand is. Next time I'm at his place I can ask, and fix the card.
It also exports everything to a printable PDF, because he will never open an app but he will happily read a paper cookbook.
Code: https://github.com/your-username/thathas-kitchen
How It Works
Everything is open source and runs locally on my laptop:
- faster-whisper (large-v3, int8 on CPU) for transcription. His memos switch between Tamil and English mid-sentence, so I let Whisper keep the original language instead of forcing English.
- Qwen2.5 7B through Ollama to pull the recipe out of the transcript as JSON.
- Streamlit for a small page on top, and WeasyPrint for the PDF.
The prompt does most of the work. I ask the model for ingredients, steps and tips, plus a fourth list called unclear for anything vague or missing. My first version didn't have that list, and the model cheerfully invented "2 teaspoons of cumin" out of nowhere. For a dish my family has eaten for decades, that's a small betrayal.
import json
import ollama
from faster_whisper import WhisperModel
SYSTEM = """You turn a transcript of an elderly cook talking into a recipe.
Return ONLY JSON with these keys:
title (string), ingredients (list of strings), steps (list of strings),
tips (list of strings), unclear (list of strings).
Rules:
- Use the speaker's own words for quantities ("a handful", "a little").
- NEVER invent a quantity, ingredient or step that was not said.
- If something is vague or missing, put a short question in `unclear`.
- Keep his personal asides in `tips`."""
whisper = WhisperModel("large-v3", device="cpu", compute_type="int8")
def transcribe(path):
segments, _ = whisper.transcribe(
path,
initial_prompt="sambar, rasam, kuzhambu, tamarind, hing, curry leaves, toor dal",
)
return " ".join(s.text.strip() for s in segments)
def extract_recipe(transcript):
r = ollama.chat(
model="qwen2.5:7b",
format="json",
messages=[
{"role": "system", "content": SYSTEM},
{"role": "user", "content": transcript},
],
)
return json.loads(r["message"]["content"])
Things that went sideways:
- Whisper turned "kuzhambu" into something I can't print here. Giving it a vocabulary hint through
initial_promptfixed most of it. - The model kept "tidying up" his quantities into exact measurements. Adding the line "use the speaker's own words" stopped that.
- Long memos sometimes came back as broken JSON. Ollama's
format="json"mode fixed it.
Why Open Source Mattered Here
I want to be specific instead of just saying "privacy."
These are recordings of my grandfather's voice. Some are the last ones where he's telling a story and cooking at the same time. I wasn't comfortable uploading them to a service I can't see inside of, so everything runs on my laptop and nothing leaves it.
It also costs me nothing. There's no per-minute transcription bill and no API key to babysit, which matters when the project is a gift and not a startup.
And I could swap models. I tried Llama 3.1 8B and Qwen2.5 7B on the same memo, and Qwen was better at keeping the unclear list honest instead of filling gaps. I don't think I could have compared models like that, or pinned one, with a hosted API.
To be fair, a big hosted model would probably have been more accurate on the rare Tamil words. I decided that wasn't worth giving up the rest.
What He Said
I printed the first three recipes and gave him the pages on Sunday.
He put his glasses on, held the pages at arm's length, and read the sambar one slowly. Then he said:
"Who wrote this? It's exactly how I talk."
He found one thing wrong, which was that I'd put the tamarind in too early, and he corrected it in pen on the page. Then he asked if I could do his rasam next.
What's Next
I want to add photos of the finished dishes, let my cousins record their own memos, and build a read-aloud mode so he can have a recipe played back to him.
Thanks for reading. If someone in your family keeps their knowledge only in their head, go record them this weekend, even if you never build anything with it.-->












