Audiobook production with AI – how much human involvement is still needed?
The audio world is undergoing a profound transformation. While artificial intelligence is becoming increasingly important across almost every industry, the audio and media landscape has also been significantly shaped by these developments in recent years.


Chris Kling
Copied to clipboard!
This article was first published in the audio@media channel of the Digital Publishing Report.
The ability to convert text into synthetic speech using text-to-speech (TTS) engines has now become so advanced that artificial voices are increasingly being seen as competitors to human voices in the audiobook industry.
These advances open up new possibilities, but they also raise fundamental questions: Will human-narrated content soon become the exception? There are good reasons to believe this will not happen — and that TTS may not even establish itself within the traditional audiobook store ecosystem.
Revolution and risk: The audio industry at a crossroads
Over the past year, we have witnessed rapid developments in artificial intelligence and the technologies associated with it. Products and start-ups built around AI-related business models are emerging and evolving at tremendous speed.
The publishing and media industries are right in the middle of this transformation. Generative AI applications, in particular, are at the center of many discussions about limits, ethics and copyright.
While the market for audio content such as audiobooks, audio dramas and podcasts has continued to reach new records in both audience numbers and revenue, some market participants have become uncertain or concerned that this new technology could negatively affect the development of their businesses.
Will synthetically narrated audio content created through text-to-speech (TTS) soon make human narration an anachronism — or at least a niche product?
This question concerns many people within the industry, including narrators, publishers, production companies and studios. And these concerns are not entirely unfounded: technological progress over the past few months has been remarkable.
The latest TTS engines are capable of understanding the tone and context of a text and generating voices that sound increasingly natural. The technology has now reached a point where it can adapt pitch and emphasis even when dealing with more complex material.
The latest TTS engines are capable of understanding the tone and context of a text and generating voices that sound increasingly natural.
Text-to-speech: Expectations versus reality
In practice, however, significant challenges remain despite these advances.
While many clients expect AI to generate high-quality audio at the push of a button, the reality of production looks very different.
If similar quality standards for accuracy and correct pronunciation are applied to TTS productions as they are to human productions, significant human intervention is still required to correct mistakes and algorithmic issues.
Although certain errors associated with human productions disappear, along with the cost of narrator fees, the amount of editing required and the currently limited market acceptance mean that consumers' willingness to pay for TTS productions often falls below their actual production costs.
Many large and small publishing houses are already experimenting with TTS audiobook productions released through conventional digital channels and stores. However, it remains to be seen whether this model will truly establish itself.
Is the future really one in which human-narrated and machine-narrated content compete side by side within the same established stores?
New players, new distribution channels
One major player has now opened the door to a very different distribution model.
Founded in 2022 and now valued at more than USD 1 billion, the start-up ElevenLabs set out to become the first company whose language model could understand the context of an entire sentence and adjust the tone and emotion of speech accordingly.
At the latest with the release of the ElevenLabs Reader App — followed shortly afterward by similar products from competitors — it became clear that this technology could find entirely different ways of reaching consumers.
The company offers a B2C solution that allows users to have their existing EPUB and PDF documents read aloud almost instantly, using a voice of their choice, in real time and completely free of charge.
In doing so, the company bypasses all conventional production and distribution processes and channels, including the entire traditional store ecosystem.
With its app, ElevenLabs bypasses all conventional production and distribution processes and channels, as well as the entire traditional store structure.
Of course, human fine-tuning and the associated quality control are absent from this model.
Consumers can now decide for themselves how well they can tolerate fully automated content and its potential shortcomings. Because the service is available free of charge, there is very little friction involved in simply trying it out without any further commitment.
It is certainly still too early to assess consumer acceptance of these products — particularly in a quality-conscious audiobook market such as Germany — but this ease of access is likely to continue increasing the acceptance of TTS.
I clearly remember how strange it still felt just a few years ago to speak to an automated telephone system, or even to assistants such as Siri or Alexa.
And because expectations are generally lower when it comes to free products, consumers are also likely to have a higher tolerance for imperfections.
The limitations of machines are the strengths of humans
So what consequences can we expect for the market for human-narrated productions?
As a production company, we began exploring this question two years ago.
Paradoxically, the more we researched and experimented with this new technology, the less concerned we became.
In fact, our findings ultimately encouraged us to invest even more heavily in our software and production workflows for human voices. They also contributed to the creation of the SaaS start-up Narrafix, which provides automated editing and proofing for human audiobook productions — making these productions significantly more competitive in relation to AI-generated voices.
The reason is that what even the best neural-network-based language models currently produce does not necessarily correspond to what many scientific definitions would describe as art or creativity.
The renowned cognitive scientist Margaret Boden, for example, defined creativity as the ability to produce ideas or artifacts that are new, surprising and valuable. (See: Margaret A. Boden (2004): The Creative Mind: Myths and Mechanisms. Routledge, London.)
The results we currently hear from AI increasingly resemble solid craftsmanship, but it is difficult to argue that they satisfy the qualities described in this definition.
While TTS represents an automated reproduction of speech, a moving human performance draws on an extensive repertoire of expressive possibilities and on the narrator's own interpretation — something that goes far beyond a purely functional reading.
More importantly, the art of successful interpretation requires empathy and sensitivity toward both the work and its audience.
To illustrate this when teaching as a guest lecturer, I sometimes begin presenting the same material to my students in a German rap dialect reminiscent of Haftbefehl.
The resulting laughter proves that our understanding of contextual and situational interpretation interacts with the audience's expectations — and does so in an individual, environment-specific way that a generic neural network cannot currently reproduce.
Having a feel for the text — but even more importantly, for the target audience and its emotional responses — is a core skill of an outstanding performer.
A skilled narrator subtly draws on situational influences that AI cannot yet perceive.
What distinguishes an artist is the ability to express something that many people feel but are unable to express themselves.
Many audiobook listeners therefore expect a voice that does more than simply reproduce words. They expect it to embody these qualities.
These are the aspects that make human narrators irreplaceable, because they enrich the listening experience on an emotional level.
By design, artificial intelligence operating according to current principles can only build upon what has been provided through its training data — in other words, it can create what a highly skilled copyist might be capable of producing.
This also means that we can only teach a machine what we ourselves are capable of understanding and expressing.
Put differently: if we assume that we fully understand and can communicate everything that moves us as human beings in art and aesthetics, then perhaps our fear of a machine created by humans is itself the product of arrogance — or, at the very least, an overestimation of our own species' abilities.
What distinguishes an artist is the ability to express something that many people feel but are unable to express themselves. Many audiobook listeners therefore expect a voice that does more than simply reproduce words — they expect it to embody these qualities.
Fortunately, there are still many aspects of our emotional responses that remain mysterious and cannot yet be fully explained scientifically.
According to some theories, reproduction and mortality lie behind all of our actions and motivations and therefore form the foundations of all artistic aesthetics. They motivate us to enable one and prevent the other.
As long as we are unable to translate the underlying reasons for these responses into program code, I believe human voice artists will continue, for a long time to come, to possess the ability to inspire us, move us, make us think and accompany us through difficult moments.
Solid craftsmanship isn't enough
Does this mean TTS will never establish itself in the market?
No. I believe the technology will find its place in certain areas.
Will we continue to find only human-narrated audiobooks on the market?
Only to a certain extent.
One conclusion we can draw from these considerations is that the simple fact that something has been narrated by a human is not automatically a guarantee that it possesses the human and artistic qualities needed to distinguish it from a machine-generated performance.
Artificial intelligence is likely to pose the greatest threat to narrators who deliver technically solid performances but are unable to draw intrinsically from their own sensitivity to the subtle nuances of a work or its target audience.
For these narrators, it may become extremely difficult to outperform machine-generated work either quantitatively or qualitatively.
As long as TTS cannot fully understand the complexity and subtleties of human emotions and cultural or subcultural contexts, however, human narrators who bring empathy and sensitivity to their work will continue to offer an irreplaceable value that goes far beyond simply reading text aloud.
Artificial intelligence is likely to pose the greatest threat to narrators who deliver technically solid performances but are unable to draw intrinsically from their own sensitivity to the subtle nuances of a work or its target audience.
What the future holds
TTS technologies will undoubtedly continue to improve over the coming years.
The audio boom and growing demand for affordable, quickly available audio content could lead to greater acceptance of TTS, particularly for specialist and nonfiction titles.
Recently launched apps that give consumers direct access to TTS tools could also influence listening behavior over the long term as media increasingly converge, establishing TTS as a complementary part of the audiobook market.
The technology also opens up exciting possibilities for interactive content.
At the same time, the streaming era has created greater pressure on costs. Although market growth means that revenue is increasing in many cases despite lower margins, TTS offers a more affordable alternative that may be particularly attractive for smaller productions.
Despite these developments, however, works created by humans will continue to play a key role — particularly wherever emotional depth and artistic quality matter.
In fact, could the growing volume of mediocre or automatically generated content ultimately lead to greater appreciation for exceptional human artistic performance?
Despite these developments, works created by humans will continue to play a key role — particularly wherever emotional depth and artistic quality matter.
After all, aren't we in the media industry always searching for the next product that moves people, stands out from the crowd and becomes something people talk about with their friends?
A great story, performed with empathy and a genuine sensitivity to the work, its context and the cultural moment, has exactly that power.
This becomes even more pronounced when combined with sound design, music and new formats that expand storytelling and immersion — allowing these productions to distinguish themselves even further from automatically generated formats.
The ability to interpret nuance and create a lasting emotional connection with the listener remains the exclusive territory of human beings.
For the expressive power of the spoken word, the artistically sensitive human performer will remain the benchmark for the foreseeable future.
