← Thinking

What voice agents are doing to content production

The most useful thing we learned from adding voice to our own systems was not about audio. It was that a long piece of content makes a far better conversation than it does a read. The audience's patience for long form has gone; its appetite for the thinking inside it has not. A conversation grounded on one piece closes that gap. This note is how we got there, what it costs, and what it turned up on the way.

Ask

The full essay, read aloud, with the words lit as they are spoken.

Talk it over, out loud, with an AI guide that has this essay in front of it and nothing else to do.

0:0023:26

Chapters

The thinking and the editing are Derek’s. The voice is an AI recreation of his own.

Ask puts you through to an AI speaking in a recreation of Derek’s voice. It sounds like him and isn’t him. ElevenLabs carries the conversation, and your microphone is only open while you are talking to it.

Nobody finishes the essay

This piece is about four thousand words long. As every publisher knows from their own numbers, most people who open something this length read the first few paragraphs, skim the headings, and leave. That is not a failing in readers. Long form asks for forty minutes of undivided attention, and the person who opens it usually arrived with a question, not with forty minutes.

For years the industry’s answer has been to make the content shorter: the summary, the key takeaways, the three bullets at the top. That throws away the thing that made the piece worth writing. The other answer, which we did not expect to reach by adding voice to our own systems, is to leave the content long and let the person have a conversation with it.

Press Ask at the top of this page and you are put through to an agent that has this essay in front of it and nothing else. Ask what it argues. Ask it to take you to the part about cost. Describe your situation and ask which section applies, and it will scroll the page to the section it is talking about. Five minutes of that gives most people what they came for, and the ones who want the whole thing can still read it or listen to it.

That is the thesis of this note. We set out to make our content audible. What we found was that long-form content is worth more as the grounding for a conversation than as a thing to be consumed: passive listening became active exploration, and once the content exists, the conversation is cheap to build.

The note is in three parts.

  • The conversation. Why long form is the substrate rather than the product; what that does to the job of an author; what changes when the voice moves from passive to active; and why the agents that carry most of the value are small ones, scoped to a single piece.
  • The production. Audio is two programmes with two budgets; what you ship is a contract rather than a file; the content model gets audited; and latency, accessibility and brand voice arrive as decisions rather than features.
  • The order. What taking the screen away does to your controls, the sequence to build in so the expensive mistakes come last, and three tests that tell you whether the conversation is the product.

Long form is the substrate, not the product

The obvious objection is that if nobody reads the essay, there is no point writing it. The answer is that the conversation is only as good as the piece underneath it. An agent grounded on a thin summary gives thin answers. An agent grounded on four thousand words of worked argument, with the caveats and the counter-examples left in, gives answers that hold up when the reader pushes back. The long form does not go away. It changes job, from the thing people consume to the thing the conversation draws on.

That reframing changes what publishing means. A piece now has three ways in, each shorter than the last: read it, listen to it, or ask it. The same holds for our training modules, where the learner can ask the tutor about the screen they are on, and for a workshop, where the record of a day the team already sat through is given back as something they can interrogate rather than a deck nobody opens. A workshop unit needs no narration at all: the recordings worth hearing from a day are the room’s own, and the agent covers everything else.

Diagram of one long-form piece as the substrate, feeding three ways in that narrow as they shorten: read it, forty minutes of undivided attention; listen to it, the same argument with your hands free; and ask it, the shortest, which gives the reader the part they came for.
The piece does not get shorter. Each way into it does, and most people take the shortest.

It also changes what finished means. An essay is finished when it can hold up a conversation about itself, which is a higher bar than reading well, and a better one.

The author’s job changes

Follow the thesis one step further and it reaches the people who write.

A book has always been a compressed method. The author worked out how to do something, and the reader’s job was to decompress it: read the whole thing, hold it in their head, and work out for themselves how the chapter on pricing applies to their situation on Monday. Most readers never finish that job, which is why most business books are bought, skimmed and shelved with the method still inside them.

So the question for an author now is which of two things they are making: the book, or the method for applying the book to the reader’s actual problem. The second is what a micro-contextual agent delivers. Give it the book and the reader’s situation, and it does the decompression the reader never would: which chapter applies, in what order, with what adjusted for their case.

We have watched this work on two books we did not write. Our training modules for a sales team are built on two well-known business books. The learner does not read either of them. On the final screen the tutor, with the book’s framework in front of it, interviews them about a real customer they are trying to win, and hands back a two-minute account of that customer in the framework’s terms and the learner’s own words, which they record and send to their manager. The book has been applied without being read.

That is the future we would bet on for authors. The book is still written, in full, because the method is only as good as the thinking under it. But what gets published alongside it, and what most people meet first, is the method: an agent that knows the book and asks about your problem. An author who ships only the artefact has, in the terms of this note, published in the passive voice.

The voice moves from passive to active

Narration is the passive voice. The system reads, the person receives, and however good the reading is, nothing they do changes what happens next. The first thing most teams build, and the first thing we built, is passive. Moving it to the active turned out to be a different product rather than a better version of the first one.

In our learning system, the voice that reads a screen can now be asked about it. It will turn the page, open a segment of a diagram, tick an item off a list, and, on the last screen of a module, interview the learner about a real customer story and hand back a two-minute script in their own words, which they then record in a voice of their choosing and send on. That last one is the thesis in miniature: the learner never reads the framework, they talk their way through it, and what they end up with is theirs.

Three rules came out of it, and none of them is technical.

One voice, one channel. We first built narration and conversation as two components at opposite ends of the page, and they talked at once, in the same voice, over each other. They are not two features. They are one voice that either reads to you or talks with you, and only one of those can be happening. So: one control with two modes, and choosing either silences the other. It sounds obvious written down, and it was still the first thing we got wrong.

The button is the invitation. When a person taps “do this with me”, the agent’s first words should be the work. Our first version greeted them, restated the task and asked whether they were ready, which is exactly what the tap had just said. The same instinct made the agent open every call by reading out its own disclosure, spending the first seconds of the conversation telling the listener what the notice beside the button had already told them. The written notice carries the disclosure. The agent says it plainly if asked who it is, and otherwise gets on with the job.

Say only what you can do. Put an agent on a screen full of recordings and it will happily offer to play one. It has no hands on the player, and the fastest way to lose a person’s trust in everything else it says is to promise something it then cannot do. So the grounding tells it, in words, what it cannot touch: it cannot play a recording, it cannot mark a module complete, and it can propose but never confirm.

Micro-contextual agents

The other thing that changed was our idea of what an agent is for.

The assistant everyone sets out to build is a general one: a voice over the whole product, with a tool catalogue, an audit trail and an approval gate, because it can reach real records. We built one, and it is the expensive half of everything in this note. What we did not expect was how much of the value would arrive from agents that are its opposite: small, scoped to one thing, and with authority over nothing.

A micro-contextual agent knows one unit of content and what the person is looking at within it. The tutor in our learning system is handed the exact text of the screen the learner is on, and handed it again on every screen change. The guide on this page is given the essay, section by section, and told which section you have scrolled to. A workshop gets an agent of its own, with the transcripts and board notes loaded into its knowledge base and retrieved per question rather than pushed into context on every connect. The rule that fell out is simple to state: push the context when the surface is the thing, and give the agent its own knowledge when the record is the thing.

Diagram contrasting one general assistant, which reaches records and so carries a tool catalogue, an audit trail and an approval gate, with three micro-contextual agents, each bound to one unit: a screen, an essay, a workshop. The small agents can read and point but not act, so they carry none of the apparatus.
One assistant that can act, and needs governing; many small agents that can only read, and do not.

Two findings about the grounding itself are worth writing down.

The first is that context accumulates, and the agent reads all of it in the present tense. Moving forward through a module this looks fine, because the latest update is obviously the live one. Moving back, it is not: the agent holds “screen seven” and then “screen six”, with nothing saying which won, and will cheerfully keep answering about seven. The fix is to say so, in words, that this update supersedes the last. One sentence, and a whole class of wrong answers gone.

The second is a test for whether an agent’s knowledge lives in the right place. Take away everything the page pushes at it and ask whether it still knows anything. Our shared tutor, stripped of its per-screen grounding, is a generic sales coach with the module’s voice. The workshop agent still holds the room, because the record is in the agent rather than in the page, and that is what lets the same agent answer somewhere else entirely, a chat channel say, where none of the page’s machinery exists.

What makes these agents cheap is what they are not allowed to do. An agent with no tool that writes needs no approval gate, no audit trail and no review pipeline, because there is nothing to approve. It is configuration rather than code: a persona, a knowledge base and a public identifier, and one can exist for every essay, module and workshop we publish. Where the general assistant costs governance, the small ones cost minutes.

Audio is two programmes, not one

Everything about cost and risk follows from one distinction, so it is worth making plainly: audio is two programmes, not one.

Narration is one-way and batch, and it produces an artefact. It costs money once, it is cacheable, and it fails safely. The worst outcome is a mispronounced word. What makes it newly worth doing is that synthetic reading is no longer something you endure. It is good enough that people pick it over reading, and the audio now knows where it is in the text word by word, which is what separates a recording from something you can follow along with, resume, and scrub by chapter.

Conversation is two-way and live, and it is an act rather than an artefact. Someone talks to your system and it does things. It costs money per user per minute, and it fails in ways that can reach real records. Its own threshold was turn-taking: a system can now be interrupted mid-sentence and pick the thread back up.

Nearly every mistake we have watched teams make here, including two of our own, comes from running them as one initiative, with one business case and one owner.

Diagram showing one ambition, audio-native, pointing outward to two separate programmes: narration, a one-way batch process that produces an artefact, and conversation, a live two-way act that reaches real records.
One ambition, two programmes, with different budgets, different risks, and different owners.

The budget splits the same way. Narration behaves like a build cost: one-off, predictable, cacheable, falling. Narrating four long essays comes to roughly 42,000 characters, or about four dollars, once. A full training module is about three-fifty. Conversation is the only cost in the picture that scales with users, which is a strategy fact rather than a billing detail. A per-user budget in minutes has to exist before launch, and the two halves cannot share a business case without one of them being wrong. Companies that plan this as a single “audio initiative” price it off the narration numbers, because those are the encouraging ones, and meet the conversation numbers in production. The micro-contextual agents above are the cheap end of the conversation half, and even they are not free: agent minutes remain the one cost here that scales with visitors, and nothing in a small agent bounds them on its own.

What you ship becomes a contract, not a file

Narration produces two files: the audio, and a timing sidecar carrying the version, the voice and model, the duration, a chapter list, and every word with a start and end time, plus a block index recording which element of the document each word came from.

Without that index, read-along is guesswork. With it, a player can highlight at block level and degrade gracefully when word matching drifts, and resume, chapter scrubbing and a podcast feed with real enclosures all become possible.

That is where the fork in the road sits: everything people want from audio is downstream of a contract, and none of it is downstream of a player. A company that buys the player has bought a feature. A company that defines the artefact has bought an asset every future surface can consume, and can build those surfaces before anything has been recorded. It also means the surfaces can be built before anything has been recorded. The same logic is what makes the conversation cheap: the text, cut at its headings, is the grounding contract, and the guide needs nothing else.

The contract has an editorial consequence too. A correction is no longer a correction. It is a correction plus a regeneration, and if nobody has decided what triggers that, you will either serve stale audio or re-narrate the library every time someone fixes a typo in a metadata field. We settled it with a rule rather than a process: re-narration is keyed to a hash of the narration script, not the source document. That reads as an implementation detail until you notice it is the only thing standing between you and a bill that scales with editorial activity.

Your content model gets audited

You cannot narrate a blob, and you cannot ground a conversation on one either. Say a document out loud and you find out how structured it actually is, which for most companies is less than they believed.

The obvious part is the spoken projection: skip figure blocks but speak the caption as “Figure: …”, turn headings into transitions, drop navigation and calls to action. Real work, and not tag-stripping.

The interesting part is anything that is not linear prose. Training modules hold charts, multiple-choice screens and checklists, and a chart has no reading order at all, so someone has to invent one. At least one of those decisions is pedagogy rather than formatting: on a quiz screen, the answer is not read until the learner has chosen. No automated conversion can make that call.

This is the change most likely to land on people who do not think they are in the audio project. Anything held as a table, a diagram, a form or an interactive step needs someone to decide what it sounds like. That is editorial work, and an honest plan staffs it.

Silence reads as breakage

In text, latency is an irritation. In voice it is a defect, and the threshold is unforgiving. More than about two seconds of silence reads as a fault.

Expect two to three and a half seconds before the first audio on a simple turn, and four to six when the assistant has to go and look something up. The levers are a faster model tier, fewer reasoning steps, a narrower tool catalogue, and pushing context in at the start of the call.

The most effective one is not an engineering lever. Prompt the assistant to say a sentence before it goes and looks, something like “let me pull that up for you”, so the audio covers the round trip. That is what a person would do, and it only occurs to you once you are listening rather than reading. A fair miniature of the whole exercise: audio problems are usually fixed by a change in manners, governance problems by a change in code.

The inverse matters just as much in the kind of system this is built for. In an enterprise workflow, long silence is often exactly what it should be: someone reading, someone working, someone thinking before they answer. An assistant that treats every pause as dead air and fills it with “is there anything else I can help with?” trains people to stop using it. What voice needs here is not urgency. It is an easy way to pause and resume. A person says “give me a minute”, or simply goes quiet, and the system waits without narrating its own patience.

Accessibility stops being separate

An audio mode that cannot be operated from a keyboard and does not announce itself is a contradiction, so the accessibility work and the audio work merge whether you plan for it or not.

Ours audited us on the way past: our training app had almost no ARIA and no live regions, so screen changes and quiz feedback were silent to a screen reader, and chart segments that could only be reached with a mouse. None of it was going to be found by the accessibility backlog it had sat on for a year. It was found by trying to make the app speak.

Audio is one of the few product investments where the accessibility dividend is structural rather than virtuous: you cannot ship the feature without collecting it.

The brand acquires a voice

Three questions arrive that a company will not have had before, and none is technical.

Which voice, and is it one? The harder version of the question is arriving already: if a listener can choose the voice, or the voice can be matched to who is listening, then the brand no longer owns the sound at the point of delivery. That is a decision about identity rather than configuration, and it is better made deliberately than discovered in a settings menu somebody shipped. Register sits underneath: measured and warm, or a bright explainer preset, is now a brand decision a build script has to be able to read.

Whose voice, and does the listener know? Ours is a professional clone of the founder’s. Hearing it narrate an essay he wrote is an easy case. Hearing it answer live is not: a person will reasonably believe they are speaking to him, which needs explicit disclosure at the start of every call. That is not a legal formality. It is the line between a product and a deception.

Who else is holding the recording? A voice vendor becomes a sub-processor holding client audio and transcripts, and any post-call webhook carries the full transcript. Under Australian privacy principles and GDPR alike that needs disclosure and probably consent at call start, which means a privacy policy that mentions it, on a surface that may not have had one.

Taking the screen away audits your controls

One more finding from the conversation half, and it deserves its paragraph even though it is no longer the point of this note. An adversarial review of the voice design for our client portal found six ways it broke, and not one was about audio. Every one was a place where a control had quietly assumed a person was looking at a screen: approval gates that left reversible actions open because a client would see the confirmation appear; a streaming layer that let a turn finish on disconnect, which is right for a closed tab and wrong for someone cutting in; an approval card left untapped on a phone in a pocket that blocked every turn after it; daily limits counted in turns, which a ten-minute call burns through in forty; and a data scope resolved from a session cookie that a server-to-server voice call never carries, so the fallback opened the widest scope rather than the narrowest.

The rule that came out of it: voice may propose a binding action and may never confirm one. The confirmation stays visual and tapped, enforced in code rather than in a prompt, because a rule that lives in a prompt is a request rather than a control. If your governance runs on approvals, confirmations and scopes derived from session state, some of it rests on the same unwritten assumption, and removing a modality is the cheapest way we know to find out where. It is also why the small agents are worth so much: none of this applies to an agent that cannot act, and a small agent needs no key and no server of its own, only a public identifier and the text of the page.

The order to do it in

Diagram of four steps ordered by blast radius: narrate what you already publish, ship the timing contract, a voice that can only read, and finally a voice with the authority to act.
Everything before the last step can be got wrong cheaply. That is the point of the order.

Narrate what you already publish. Smallest blast radius in the company, shareable within days, and it proves the pipeline against real content rather than a sample.

Ship the timing contract, not just the audio. Read-along, chapters and resume are what turn a novelty into something people use twice. The step most likely to be skipped, and the one that decides whether the rest is worth building.

Then the structured content: courses, documentation, anything non-linear, where the projection work and the accessibility debt both live, and both are cheaper to face on content you control.

Then a voice that can only read. This is where the conversations live, and it is the step that turned out to carry most of the value. It front-loads the voice, the persona, the disclosure script and the task set a fuller build reuses, while risking nothing. Most of the agents we now run stop here, on purpose.

Then a voice that can act. Last, deepest, and the only one that can carry work with consequences. By then you will have found the things that break, on paper rather than live.

The conversation is the product

Start from the reader rather than the format and the conclusion is hard to avoid. The reader arrives with a question and a few minutes. The essay, the module, the book, the record of a workshop: each is written in full, and each is met, by most of the people who meet it, through a conversation about it. The long form is the substrate. The conversation is the product.

That inverts where the effort goes. Most of what is sold as audio today is a play button, a vendor’s embed reading the page aloud, alt text and cookie notices included: the passive voice, automated. The work that matters is on the other side. The grounding. The persona. The small agent that knows one piece and what the reader is looking at, and the honesty to let it say only what it can do. None of that needs a studio, a key or a governance programme. It needs the content to be good enough to hold a conversation about itself, which is the bar every piece should now be written to.

Three tests, then, for your own work or for anything sold to you:

  • Ask it about the paragraph you are looking at. Does it know which one that is, or does it answer about the document as a whole? That is grounding, and it is the whole of the argument above.
  • Give it your situation and ask which part applies. Does it apply the piece to your problem, or summarise the piece back at you? That is the difference between a book and a method.
  • Ask it to do something that costs money. Does it refuse, say what it was about to do, and put a button in front of you? That is governance, and it is the smaller part of the story.

Pass the first two and the conversation is the product, whatever format the content was written in. Fail them and you have a play button.


Related: What “AI-First” Actually Means for Marketers. On the approval gate this design had to re-derive for a surface with no screen, and why “the harness assembles, the human commits” is the rule that survives a change of modality.

If this is the kind of problem your team is working through or you'd like to understand the technical implementation, we'd like to hear from you.

Talk to us