← Claude Code Fundamentals
Voice as Input Part 3 of 4 Beginner
6 min read

From transcript to context

From transcript to context

From transcript to context

You have the recording. You have the text. Now comes the step where most of the value is won or lost, and it takes about four seconds to get wrong.

You ask for a summary.

The summary is the least interesting thing you can do

A standard summary is optimised for being readable. That means it smooths.

It resolves "I think we could probably do that, though I am not sure about the licensing" into "agreed to proceed". It drops the concern someone raised once and did not repeat. It replaces the client's own words with neutral business phrasing.

Every one of those moves removes exactly what made the transcript worth having. The hedge was the risk. The concern that was not repeated is the one that resurfaces in month six. The client's own vocabulary is what makes your proposal sound like it was written for them.

This is not a quality problem you can solve with a better model. Smoothing is what summarising is. A more capable model summarises more elegantly and discards the same things.

So: read the summary if it helps you orient. Do not treat it as the output. The transcript is the raw material, and what you do with it is the work.

Feed it whole

The practical move is unglamorous: give the model the actual transcript, not a description of it, and ask it to do something specific with your own material.

Useful shapes:

Against a template. "Here are the transcripts from our four meetings with this client. Here is our standard scope document. Draft a first version using only what they actually said, and mark anything the template requires that they never mentioned." The final clause is the valuable one -- it surfaces the gaps rather than filling them with plausible invention.

Against a decision. "Here is the transcript. What did we commit to, what did we leave open, and what did someone object to that we did not resolve?" The third question is the one people never ask.

Against your own thinking. Record yourself talking through a problem for two minutes, then feed the transcript in and ask what you left out. You will be told about assumptions you did not know you were making.

Messy in, useful out

There is a counter-intuitive point here that takes a while to accept.

The messier your input, the better the output tends to be.

People instinctively tidy before feeding a model: they write a clean paragraph describing what they want, rather than handing over the tangled real thing. That feels like being helpful. It is actually throwing away the signal and keeping the wrapper.

Give it the unedited two minutes. The dead end you went down first tells it what not to suggest. The sentence you abandoned halfway tells it where you are uncertain. A clean paragraph contains none of that, and the answer comes back correspondingly generic.

The ratio decides the answer

A useful way to think about any single interaction: there are three quantities, not two. The context you put in, the prompt, and the output.

What determines whether an answer is grounded is mostly the relationship between them.

A small prompt sitting on top of a large body of your own material produces something specific to your situation. The same prompt with nothing underneath it produces something that sounds authoritative and could have been generated for anybody. And a long answer built on a thin context is the shape to be most suspicious of, because the words had to come from somewhere and they did not come from you.

This is why prompt craft is a smaller lever than it is usually made out to be. The wording matters some. What you filled the window with matters much more.

One transcript, or fifty

There is a threshold effect worth planning for.

A single transcript answers questions about one occasion. What was said, what was decided, what to follow up.

An accumulated set answers a different class of question entirely. What does this client keep coming back to. Where do deals consistently stall. Which objection do we meet regardless of who is in the room. What words do people actually use for the problem we think we are solving.

That second class is not available to anyone who keeps transcripts as loose files in folders. It requires that the material be kept, connected, and searchable as a body rather than as items.

Which is the point where this stops being a habit and starts being a small amount of structure. Gathering is the easy part. Knowing what it means is the work.

What to keep

Not everything deserves to be kept forever, and the classification from Part 2 does real work here.

A reasonable starting position:

Keep and accumulate: client conversations, project meetings, your own thinking-out-loud. This is the material where patterns over time are worth having.

Keep short, then let go: routine internal check-ins. Useful for a week, noise after a month.

Handle deliberately: anything about a person rather than a project. Some of it should not be kept at all, and some of it has retention limits that are not yours to choose.

The tension is real and worth stating: retention requirements and accumulated value pull in opposite directions. The answer is not to pretend otherwise. It is to know which material is which before you have a thousand files.

What comes next

You can now turn conversations into working context.

There is one thing the material still cannot do, and it is the thing that separates a good archive from something that makes work happen: it does not know who said what to whom.

That is Part 4, and it comes with obligations.

Before you move on 0 / 5
I understand why a standard summary discards the most useful part of a transcript
I can explain how the ratio between context and prompt affects whether an answer is grounded
I have used a full transcript as input for a real task rather than asking for a summary
I understand the difference between what one transcript can answer and what many can
I know which material I want to keep long enough to accumulate
Knowledge check 1 / 3

Try again