AI Note-Takers Sell You Action Items. Not One of Eight Vendors Publishes an Accuracy Number for Them.
Jun 29, 2026 · 5 min read · by Jordan Kwan
TL;DR: I audited all eight major AI note-takers (Otter, Fireflies, Fathom, Granola, tl;dv, Read.ai, Avoma, Grain) against their own sites and docs. Transcription is close to solved and half the vendors will quote you a number for it. Action items are the tell: eight of eight sell automated action-item extraction, and zero of eight publish an accuracy figure for it, while the benchmark research on dialogue summarization still finds significant hallucination. Use these tools for memory. Do not wire them into judgment.
Every meeting tool now ships an AI notetaker that joins your calls, transcribes everything, and emails you "action items" like a helpful intern. Before you put one in every meeting, it's worth knowing which of its promises the vendors are willing to put a number on. On August 6, 2026 I went through all eight major vendors' homepages, docs, and help centers and kept score. (The bot-in-the-room problem this surfaces has become the category's main battleground; the Granola vs Fathom comparison covers where that fight landed.)
What is genuinely solved?
Transcription, nearly. This isn't vendor magic, it's a solved-research dividend: the Whisper paper trained on 680,000 hours of multilingual audio and reported models that approach human accuracy and robustness zero-shot, with no per-use-case fine-tuning. Every note-taker in the category is standing on that lineage.
The vendors are accordingly happy to quote numbers here, or at least four of eight are: Otter says 85-90% for clear audio, Fireflies claims "95% Accurate" on its homepage, Avoma claims 95% for transcription and speaker ID, and tl;dv cites 97.3% from its own blog benchmark. Read those as self-graded homework, none carries an independent audit, but the fact that this is the one feature anyone will attach a percentage to tells you it's the one they measure.
The honest asterisk is speakers, not words. On the standard AMI meeting corpus, current open-source diarization still posts error rates roughly in the 13-23% range depending on version and mic setup. "Who said it" remains the soft spot in "what was said," which matters the moment a summary attributes a commitment to a person.
What is the coin flip?
Action items. This is where the marketing outruns the measurement, and the audit gives you the shape of it in one line: eight of eight vendors sell automated action items ("automatically captured and assigned," "identifies action items and assigns ownership"), and zero of eight publish any accuracy number for that feature. The same companies quoting 95% for transcription go silent when the claim shifts from "we heard it" to "we understood it."
The research says the silence is earned. TofuEval, a benchmark for topic-focused dialogue summarization, found LLMs "hallucinate significant amounts of factual errors in the dialogue domain, regardless of the model's size." And the action-item-driven summarization work that represents the published state of the art reports scores on the AMI meeting corpus that are a solid research result and nowhere near "wire this into your task tracker." Extracting commitments means resolving hypotheticals ("we could send that over") versus promises, casual phrasing ("yeah I'll ping him") versus noise, and who actually owns a task versus who spoke last. That is judgment, not transcription.
The transcript is a record. The action items are an inference with a checkbox. Read them like a suspicious editor, not a to-do list.
Practical translation: treat every auto-extracted action item as a draft requiring confirmation. The failure mode that costs you isn't a missed note, it's a phantom commitment landing on someone's plate with your name on the assignment.
What does a bot in the room do to the conversation?
Here the audit's most interesting finding isn't a number, it's an engineering pattern. All eight vendors have built visible consent and disclosure machinery: Fireflies emails participants an hour ahead and pins a chat notice, Fathom sends consent emails 24 hours out and displays a recording banner its docs say "cannot be hidden," Avoma plays an audible recording announcement, tl;dv can replace your meeting link with a consent link that blocks recording if declined, and Fathom's help center states flatly that there is "no way to silently record".
Vendors don't build friction like that for fun. They build it because a recording bot in the participant list is a live social and legal variable, and because an entire subcategory (Granola's bot-free capture, Fathom's bot-free mode, tl;dv's "no bot required") now exists to make the observer less visible. The market itself is telling you the observer changes the meeting. If you sell for a living, assume candor drops when the machine joins, and decide call-by-call whether the archive is worth it.
The privacy footnote everyone ignores
A note-taker in every meeting means a searchable, permanent, cloud-hosted archive of every candid thing said in every meeting you take. That's an asset and a liability, and consent law is not a footnote: the Reporters Committee's recording guide counts about 11 states with all-party consent requirements, California, Illinois, Florida, Pennsylvania and Washington among them. Recording calls across state lines is your problem to solve before the bot joins, not after, which is precisely why the vendors built the consent tooling above; even Read.ai's own legal explainer walks buyers through it. Announce the bot, use the consent features, and kill it for anything sensitive. "Searchable forever" cuts both ways in a lawsuit or a leak.
What would I actually pay for?
The features with numbers attached, at full trust; the features without, as drafts:
- Transcription + search: yes. It's the measured feature, and turning a month of calls into a queryable record is real value.
- Recall summaries: yes, for catching up on a skipped call, with the hallucination literature in mind for anything load-bearing.
- Action items: only as suggestions a human confirms. Never auto-wired into a project tool. Zero of eight vendors will put a number on this feature; match their confidence.
Verdict
The category has quietly solved the hard technical problem (hearing) and is now overselling the harder one it can't yet measure (understanding commitment). The receipts pattern makes the boundary legible: numbers for transcription, silence for judgment. Buy the memory. Audit the judgment. And read the room, literally: eight vendors' worth of consent machinery exists because the bot in the participant list is never neutral. Some rooms it belongs in. Some rooms you close the laptop, look the person in the eye, and take notes like it's 2019.
No hype. Just the spec sheets, and what's missing from them.
Written by Jordan Kwan, founder of Reachium.
I build Reachium, the LinkedIn outreach platform behind the tactics you just read. Same brain, live product.
See what Reachium does ↗