nomuraya — The Curious Operator

Hands-on accounts of building, breaking, and rebuilding AI agents.

Trying to create the ideal narration voice with AI, until you realize the limit

2026-06-17

"Why can't we reach the ideal voice in an era where AI can make a voice?"

After confronting this question for more than half a day, I realized that it was more essential than a technical answer.AIs can "mass-produce" voices, but cannot design "uniqueness."That is the conclusion.

Trigger: “I want that voice”

When creating a YouTube video, I was thinking about what to do with the voice of the narrator.

The barrier to speaking up for yourself is high.The quality of the microphone, the recording environment, the cost of editing, and above all, the psychological hurdles of “making your voice public.”It's costly to hire a professional narrator.A one-off video may cost a few thousand yen, but if you continue to make videos, each request is not realistic.

Then, it was a natural trend to think that it was AI.

I used the voiceover of a YouTube channel as a reference.High, clear, calm explanatory tempo, and not tired after listening to it for a long time.Specifically, it is a type that is often used in commentary channels - a voice that is not emotional or even mechanical, but has "well-organized humanity".

There was a clear image that "I want something close to this voice".From there, trial and error began.

I searched for an off-the-shelf product first: try the TTS tool from scratch

The first approach was to look for it.I tried to find something that matched what I already had.

There are several TTS tools that you can use for free.VOICEVOX, CoeFont, Google's Text-to-Speech, Amazon's Polly.I tried the one that supports Japanese from the beginning.

VOICEVOX is free to use and the quality is not bad.There were many character-type voices, and many voices were difficult to match for "explanation and narration" applications.Zundamon's personality is too strong to stop.Even the "normal" and "sexy" of Shikoku Medan have different directions.

CoeFont has a wide variety of voices, and it also has voices that are close to business.I tried it to the extent that I could try it on the free plan, but I couldn't find a voice saying "this is it".

Each time you try with each tool, try increasing the pitch, adjusting the talk speed, and changing the suppression.

The result was the same every time.Stop at "It may be close".Far from ideal.

It's hard to verbalize why you feel “different.”It is not a clear difference between a high voice, a low voice, a fast voice, or a slow voice, but a more vague sense that the atmosphere is different.Even if you try to analyze what is stabbing you with the reference voice, it is not a good word.The expression "well-developed humanity" was the closest in my mind, but it was a word that could not be converted into a parameter.

I spent a lot of time, but the impression that came out was "none of them are the same".

Switched from "Find" to "Make"

I felt the limitations of ready-made products and changed my policy.From "Search" to "Make".I changed my approach.

Tried Approach 1: Voice Transformation (RVC)

RVC (Retrieval-based Voice Conversion) is a technology that converts existing speech into another voice quality.The TTS output can be used as a source of conversion to the voice quality of the trained model.

In theory, it can be close to the ideal voice quality.However, when I actually tried it, I found that "it is bound to the limit of the transform source".If the original voice is A-like, no matter which model you convert it into, it will only be "A-like B".The habit of input remains in the output.

Specifically, we tried to convert the standard voice of VOICEVOX in RVC.The VOICEVOX output is slightly "synthetic voice-like".There is a tendency to suppress consonants too uniformly, and the processing of consonants is somewhat mechanical.Converting with RVC does not completely disappear to the extent that this habit is reduced.The personality of the voice quality of the target model also emerges.As a result, it became a subtle state that "VOICEVOX feeling remained and became a different voice".

If we improve the quality of the source, we may be able to improve it.However, in the end, "high-quality source audio" cannot be obtained unless it is recorded by myself or a professional.

! [Trial flow: find→ make→ notice] (assets/ai-voice-narration-limit-flow.png)

Tried Approach 2: Zero Shot TTS

When you pass the reference voice, it is a technology that reads with that voice.There was an expectation that if the quality of the reference voice improved, the quality of the output should also increase.

I actually tried several OSS models and APIs.As a result, there was a problem that the reference voice pulled too much.If there is a reference voice that is close to the ideal, it will approach, but if the reference voice itself is "only close", the output will also be "only close".

And in order to prepare the "reference voice of the ideal voice", it is necessary to have a person who has that voice.It's circulating.

Tried Approach 3: Voice Design

It is a function that can be used in OpenAI TTS, etc., and designs voice features with text prompts.Voices can be generated from instructions that are bright, clear, energetic, but not childlike.

This can create a "voice like that".However, it is only "those that anyone can make".The voice that can be designed at the prompt is the range of voice that can be designed at the prompt.In other words, everyone else gets the same voice at the same prompt.

No matter what I did, it was "just close"

Through trial and error, I noticed something in common.

Every time the parameter is adjusted, there is a sense of "closer than before".But the feeling of "this" never came.

Voice quality conversion, zero-shot, and voice quality design had different approaches, but the essence was the same - "manipulating parameters."Directionality of height, speed, reverberation, and voice quality.Adjust these to get closer.But you can't get there just by getting close.

I thought about why for a while.

The feeling that you might be able to get there if you make a few more adjustments makes you continue to trial and error.But at some point, I realized that the sensation itself was a trap."Close" does not mean "reachable".The adjustable range of the parameters and the ideal position may not overlap in the first place.

What I realized: You can't design for uniqueness

I realized that what I was looking for was not "parameters" but "uniqueness".

The ideal voice has a unique personality.The height of the voice, the way of speaking, the way of placing words, the way of taking between.These are "parameters" when viewed individually, but together they create the uniqueness of "this person's voice".Its uniqueness lies not in the combination of parameters itself, but in the context that the combination was "naturally born from the person's life and experience".

What can be changed by adjusting the parameters of AI are quantitative factors such as height, speed, and inhibition.But "uniqueness" is not there.Rather, it is in places where it cannot be quantified.

It is close to the feeling that the director said "this role can only be given to this person" at the scene of video production.It's not about specs or technology, it's about being.

When you are in the field of AI introduction support, you often see the same composition.In response to the request to automate this work with AI, it is often necessary to organize "what will be replaced by AI and what will not be replaced".

In one case, there was a request to fully automate customer correspondence with AI.If you actually delve into the requirements, "instant response to routine inquiries" can be handled without problems by AI.However, AI will not replace the relationship with customers who say, "I can leave it to you with peace of mind because I am the person in charge."If you confuse this, it will have the opposite effect of reducing customer satisfaction even if AI is included.

The story of the voice this time was a microcosm of that."Automation of voice generation" is good at AI, but "the personality of that voice" can not be released to AI.

Organize AI's strengths and weaknesses

! [map of AI's strengths and weaknesses] (assets/ai-voice-narration-limit-chart.png)

AI is good at:

- Automation to instantly voice text - Generate unlimited voice of a certain quality - Parameter adjustment to change direction - Compare multiple patterns in a short time

AI is not good at:

- Create a new, unique voice - Recreate the “personality” of a specific person (if no learning data is available) - Mimic personality born from context

Next Reality Solution: From “Solve it all with AI” to “Choose where to use AI”

In conclusion, I think that it is difficult at this time to "fully approach the ideal voice with AI".

What emerges as a realistic solution is the combination that "voice itself" is procured by means other than AI, and "automatic generation mechanism" is created by AI.

Specifically, it is a configuration in which a person with a voice close to the ideal is asked to record it once, and AI mass-produces based on that voice.It can be recorded once, and operation can be automated.Narrator costs are high because it is a "record every time, cost every time" model, so if you change it to a "record once and then deploy it with AI" model, the ongoing cost can be reduced.

Alternatively, there is a way to stop asking for the perfect ideal voice and output the voice "at a level that can be used for this purpose" with AI.If it is a narration of a YouTube video, even if you do not ask for "nature close to humans", AI voice at the level of "information can be heard and there is no stress" is often enough.Because viewers often want content, not voice quality.

Both of them take a step back from the idea of "trying to do everything with AI".

It is often said that AI is not a "universal tool", but a "tool to speed up certain tasks".The same is true in the generation of voices, and it is realistic to design a combination of other means to use "things that AI can do quickly" and "things that are difficult for AI".

This half-day trial and error was a time when I experienced it with a specific subject called voice.

"Choosing where to use AI" -- polishing that choice may be the essence of AI utilization.

Conclusion: What trial and error can teach you

The results were not perfect, but the trial and error was worth it.

The sense of “what is possible and what is impossible” cannot be grasped without actually trying it.Whether you read a spec or an article, you won't reach the resolution of your own experience of hitting a wall.

There are many evaluation articles on AI tools, but there are surprisingly few stories of "I noticed the limit when I tried it".Sometime the next person to try may find it helpful to see where the crowd is, rather than the success story.

Another byproduct of this trial and error is the increased resolution of “What do I want from my voice?”My own words of "well-equipped humanity" did not come out before trial and error.By trying dozens of different voices and thinking about what was different, the outline of my request became clear.

This is not a limitation of AI, but an aspect that deepens human recognition by using AI.I think that one of the unexpected values of using AI is the experience that "when I try to make AI do something, it becomes clear what I want to do."

When an AI realizes it can't or is difficult, it's not a failure, it's information.It's not over."This part is the starting point of the reorganization that you need to do yourself instead of AI, or you need to combine other means."I let go of my excessive expectations of technology and design a combination of my own judgment and AI's ability - this is close to the feeling of actually working with AI.

I hope this article is a reference for someone's trial and error in that sense.

About this article

This article is based on a trial-and-error experience looking for AI-narrated voice.We intentionally omit references to certain tools or services.The technology will continue to change, but I think the realization that "uniqueness cannot be designed" will not change.

---

*Machine-translated (MyMemory API) from a Japanese original at [nomuraya-hub.pages.dev](https://nomuraya-hub.pages.dev/). Pre-review draft. I am the same author writing under different pen names — "nomuraya / shimajima / 中翔" — depending on the medium.*

Subscribe

If you want occasional long-form posts about AI agents, FIRE, and how a curious operator thinks about both, drop your email below.

Subscribe via Substack →