Who said what, and what you owe the room
This is the part of the series where the material stops being about you and starts being about other people. So the obligations come first, and the capability second.
Start with what you owe
If you are recording a conversation, the people in it should know. Every time, said out loud, without waiting to be asked.
That sounds obvious. In practice it erodes fast: the tool joins automatically, the habit becomes invisible, and three months in nobody announces anything because it feels established. The person who joins the meeting for the first time has no idea.
A few things worth holding to:
Say it before, not after. "I record these so I do not have to take notes -- say if you would rather I did not."
Make refusal cheap. If declining is socially expensive, consent is not real. Someone saying no should cost them nothing and you should stop without discussion.
Do not record what is not yours to record. A conversation about a person, a conversation someone is having in confidence, a room you are a guest in. The fact that capture is easy is not a reason.
Own what you distribute. If you send round a summary drawn from a recording, you are asserting that this is what people said. That assertion is yours.
None of this is a compliance ritual. It is the condition that makes the rest of it usable, because a record people did not know was being made is a record you cannot actually use.
Why attribution matters
With that in place: attribution is the thing that turns an archive into work that moves.
"We decided the quote goes out before Friday."
That is a note. It is accurate. It drives nothing. Nobody wakes up on Thursday feeling it was theirs.
"Anna told Markus the quote goes out before Friday, and Markus said fine."
That is something else. It can be followed up. Someone can say "that was not mine to take" and someone else can say "yes, it was", and the conversation that follows is a real one.
A commitment without a sender is not a commitment. It is a phrasing. That, more than length or format, is why most meeting notes produce no follow-up.
The easy case and the hard case
Where this is easy: online meetings. Every participant has their own audio stream and their own login, so the transcript arrives already labelled. Nothing to work out.
Where it is hard: the ordinary room. Several people, one microphone, overlapping speech, different distances from the table, and nobody has recorded a voice sample in advance.
There is only the audio, and working out from a single track that there were four people and where one stopped and the next began is a substantially harder problem than reading a label that already exists.
It is worth knowing which case you are in, because the reliability is not comparable. Treat labels from an online meeting as solid. Treat speaker attribution from a room recording as a draft that a human should confirm before anything is sent anywhere.
Two things that sound alike
Here is the distinction that almost nobody makes, and it matters more than it sounds.
Speaker separation establishes that four different voices spoke and which of them said what. Speaker A, speaker B, speaker C. Nobody is named. No profile is stored. It is signal processing.
Speaker identification binds speaker B to a named person by matching the voice against a stored voice profile. That is a different operation, and a different category of processing: a voice profile used to single out a specific individual is handled as biometric data, the strictest class of personal data there is.
Both are called "knowing who said what" in ordinary speech, which is exactly why they get merged. But one is counting voices in a room. The other is measuring a body.
The order that follows once you have seen the difference:
Separate always. Speaker A, B, C, with no voice profile retained.
Never identify automatically. No matching against a biometric reference as a default behaviour.
Let the person bind the name. Someone who was in the meeting marks "speaker B" as themselves afterwards, if they choose to. Then a human made the connection, not a system that measured them.
That last step is also the right one for a reason that has nothing to do with law: whoever attaches a name to a commitment should be a person who can be held to account for getting it right.
If you are building this into something customers use, have the legal position checked properly rather than relying on a guide. The shape above is sound; the specifics depend on your jurisdiction and purpose.
An extracted action item is a claim
Most tools will now offer to pull decisions and action items out of a meeting automatically. This is genuinely useful and it is where the last mistake in this series happens.
What comes out is not a list of commitments. It is a list of claims about what was said, produced by inference over text. It can attribute a sentence to the wrong speaker, read a hypothetical as a decision, or miss the qualifier that reversed the meaning.
Sending that list round unreviewed means telling five people that they agreed to things, on the strength of a guess.
The fix costs a minute: read it before it goes out. An extracted item becomes a commitment when a person has accepted it, not when a model has produced it. Fast extraction plus a human acceptance step is a genuinely good workflow. Fast extraction alone is a way to generate confident misunderstandings at scale.
In some settings this matters more than convenience suggests. In construction, property and anything with a formal minute, the record of a meeting is a document that can be relied on later. An automatically extracted decision that nobody read is not a minute. It is an assertion.
Rules worth writing down
Write your own version of these and say them out loud in the meeting. Three sentences is enough.
- I record, and I say so before we start.
- Separated speakers, no voice profiles. If you want your name on your part, you add it.
- Nothing extracted goes out until I have read it.
If you can say that in a room without it sounding strange, you have the whole series working.
Where this leaves you
Four parts: the habit, the channels and what may go where, turning transcripts into context, and attribution with the obligations that come with it.
What you have at the end is not a tool. It is a way of working in which the things people actually say stop evaporating -- and in which the people who said them are treated properly on the way.
The model was never the difficult part. Knowing what you were missing is.