Call transcripts and dictation, in detail

The short version is on the home page. Here is all of it: what a call leaves on disk and how long its transcript takes, how speaker labels and summaries work, and how dictation behaves.

01 Call transcripts

One call, one folder of ordinary files.

Start a recording and Saykeep captures your mic and what your Mac plays, straight from macOS: Zoom, Meet, Teams, a Slack huddle, a browser tab. No bot joins the call, and there’s no audio driver to install. After the call you get the transcript with who said what, as ordinary files.

One meeting = one folder in ~/Documents/Saykeep REC
mic.opusyour side~27 MB
system.opuseveryone else~25 MB
meeting.jsonwhen, how long, which tracks<1 KB
transcript.jsontimestamps and speakers~220 KB
transcript.mdreadable anywhere~70 KB
summary.mdsummary — (with a model you connect)3 KB

Demo is sped up. An hour of call on an M5: ~35 min with speaker labels, 3–7 min without.

Sizes for one hour of call, from real recordings. Opus audio (VLC, IINA or any player that handles Opus), JSON, Markdown — open them without Saykeep, no export step. Prefer transcripts only? A setting deletes the audio once the transcript is written.

When the call ends, Saykeep makes the transcript in the background

About 35 minutes per meeting-hour with speaker labels, which are on by default, on the M5 we measured; about 3–7 minutes with them switched off (Settings ▸ Transcription ▸ Speaker labels). Older Macs are slower, and we haven’t measured them yet.

Scaled from real 10–17-minute recordings. The transcription itself runs on your Mac’s GPU; speaker labelling runs on the CPU and finishes before the transcript is written, which is why it sets the wait.

Saykeep’s menu-bar panel during a call: Recording meeting, 12:34, the MIC and SYSTEM tracks both writing, a Stop meeting button with its shortcut, and the folder it is writing to.
The panel under the menu-bar icon while a call records: both tracks writing, the Stop meeting button, and the folder it is writing to.

Recording is visible by design

Whenever Saykeep is recording a meeting, the menu-bar icon shows it the whole time, and no setting turns that off — a test sweeps every setting to keep it so. Before your first meeting recording, a plain-language notice says recording may need every participant’s consent and names jurisdictions to be aware of. The indicator is on your screen only: Saykeep sends nothing into the call, so telling the other people is up to you.

One sentence that works, at the start of the call

“I’m recording this for my notes, on my computer — is that OK?”

Is it legal to record your calls? →

Audio you can let go of

Keep the audio (the default), have Saykeep delete it once the meeting’s transcript is written, or have it ask you each time you stop a meeting: Settings ▸ Privacy & Storage ▸ Recordings. Audio is deleted only after a transcript exists; transcripts and summaries are never deleted automatically.

02 Speaker labels

Who said what, labelled on your Mac.

Speaker labels are on by default. The first time, Saykeep downloads one small model (about 30 MB) from models.saykeep.app, with no account and no token; after that, labelling works without the network. It is most of the wait: switch it off in Settings ▸ Transcription ▸ Speaker labels and an hour of call takes about 3–7 minutes instead of about 35, on the M5 we measured.

In Recordings, press Speakers… on a meeting: each voice comes with two sample quotes and a name field. Rename “Speaker 1” to a real name, and Save names rewrites that meeting’s transcript with it. Voices on your own microphone are labelled “You”; for a meeting in a room, where your mic hears everyone, switch on Settings ▸ Transcription ▸ Label speakers on your microphone too.

Saykeep’s Recordings window, Library view, in dark mode: 4 recordings in ~/Documents/Saykeep. Each card carries chips such as ✓ Audio, ✓ Transcript, ✓ Summary, ✓ Speakers, ○ No summary, ○ Speakers off or ○ Not transcribed, and buttons such as Open transcript, Summarize now, Transcribe, Speakers… and Delete.
The Recordings window: each meeting says what it has (audio, transcript, summary, speakers), and Speakers… is where you name the voices.
03 After the call

The transcript is made in the background, after the call.

Stop the recording and carry on: Saykeep transcribes the call, labels the speakers and, if you connected a model, writes the summary, one job at a time, in order. Recordings ▸ Activity shows what it is doing and what is waiting. If your model can’t be reached, the summary waits and tries again with your next recording, after a restart, or when you press Retry now; the transcript is already saved.

A summary has three parts

Topics · Decisions, each marked Decided or Discussed · Action items, where the model is instructed to tag an owner on every item: the person at the microphone, a participant named out loud, or Unassigned. The text is written by the AI endpoint you connect; without one, the summary step is skipped.

## Topics- The beta ships once scope is locked.## Decisions- [Decided]: Lock the scope on Friday.## Action items- [Charlotte]: Write the announcement.- [Sebastian]: Have the setup guide ready by Friday.
Saykeep’s Recordings window, Activity view, in dark mode: 3 jobs · 1 running · 1 waiting · 1 deferred. A Summary job marked Waiting because the LLM endpoint could not be reached, with the line “The transcript is saved — this meeting is not lost.” and Retry now and Remove buttons; a Transcript job marked Running, Transcribing, with its progress bar; a third job marked Deferred (2/3).
Recordings ▸ Activity: one job at a time, in order. A summary whose model couldn’t be reached waits, and its transcript is already saved.
04 Summaries

Summaries come from a model you connect.

Saykeep doesn’t run a language model of its own; for summaries you point it at one. With a free model on your Mac, served by LM Studio or Ollama, the summary is made on your Mac too. With your own API key for an OpenAI-compatible service, the transcript text goes to that provider, with your name and “about you” note if you filled them in, and never your audio.

Settings ▸ Summaries & AI asks for the endpoint address, the model name and, if your service needs one, an API key, which is stored in your Mac’s keychain. Test tells you whether the endpoint answers. Until an address is set, summaries are off.

ChatGPT Plus doesn’t include API access, so it can’t be connected here. Without a model you still get the full transcript with speaker labels, and Summarize now in Recordings writes the summary later, once you connect one.

Saykeep Settings, Summaries & AI, in dark mode: the Summarise meetings automatically switch, on, and the Endpoint address field filled in with http://127.0.0.1:8080/v1, beside a Not tested label and a Test button.
Connecting a model is a form, not a terminal — the fields start empty. Here the address of a model server on the same Mac is filled in.
05 Dictation

Hold a key, speak, release.

Hold the key you chose and speak. About a second after you let go, once the model is warm, the text is on your clipboard, ready to paste into any app. Choose Settings ▸ Privacy & Storage ▸ Dictated text ▸ “Insert it where my cursor is” and Saykeep pastes it where you were typing.

It listens only to your microphone and keeps only the text — no audio file is written (a debugging option that keeps your last 20 clips is off by default). Dictation never starts the system-audio capture.

A custom vocabulary (a plain text file of your names and jargon, on your Mac) fixes their spelling in dictated text and in what the summariser reads, and hands Whisper the same list as a glossary hint while it transcribes. The saved meeting transcript is left as the model actually heard it.

Saykeep Setup, step 8 of 8: “Try it — hold your push-to-talk key”, with a transcribed sentence and chips reading 0.9 s, large-v3-turbo, mlx · Apple GPU, You're all set.
Setup ends with a live test: hold the key, say something, and see the text, the model and the backend it ran on. (Illustration at the measured median.)
0.89 s
median decode time, model warm

The on-device decode of a dictation of up to 30 seconds, once the model is warm: median 0.89 s across 101 dictations from the developer’s own logs on one Apple M5 Mac (32 GB), Whisper large-v3-turbo on the GPU — fastest 0.6 s, slowest 10.5 s, measured before the text is pasted. Longer dictations take proportionally longer; there is no cap. On a memory-pressured Mac the first dictation after a long idle has taken several seconds longer while macOS paged the model back in; since 0.3.7 that reload starts the moment you press the key (not yet re-measured).

~5%
words wrong on read speech, English and Polish — measured

The default models on a 100-sentence sample of the public FLEURS test set per language (read speech, not meeting audio), transcribed on one Apple M5 Mac with stock settings: roughly one word in twenty wrong, in both languages. Names and specialist terms are where any model slips; that is what the custom vocabulary is for.

large-v3-turbo
the dictation model, named

Whisper large-v3-turbo for dictation and Whisper large-v3 for meetings, both running on your Mac’s GPU through mlx-whisper. Dictation is the model’s own output plus your vocabulary fixes: punctuated, with no rewrite pass and no cloud. The language is detected for each recording unless you name yours in Settings ▸ Transcription ▸ Spoken language; we have measured English and Polish. Pick a smaller model there if you want speed over accuracy.

Download the free trial Buy — $39 30 days · all features · no credit card · no account · macOS 14.2+ (Apple Silicon)