Make Your Own Affirmation Audio — Without CapCut
If you have spent any time in the self-recorded audio corners of the internet, you have seen the same question asked over and over: what app do I use to make my own affirmation audio, and do I just record on my phone?
The answers are always a list of tools. Write your lines somewhere. Record or generate a voice. Drop it into CapCut. Layer it. Add rain or brown noise. Balance the volumes. Export. People who do this regularly report it takes 20–30 minutes the first time and 10–15 minutes once they know the steps.
That is a real method and it genuinely works. It is also four apps and a skill you did not set out to learn. This guide walks the DIY route honestly, shows you where it usually goes wrong, and then shows you the short version.
(One note: this is about making audio for a personal mindset practice. It is not treatment for any condition.)
The DIY method, step by step
Here is the workflow people actually describe, with the parts nobody mentions until you hit them.
- Write your lines. Usually in Notes or a doc. This is the step that decides whether any of it works, and it is the step most guides spend the least time on.
- Get a voice. Either record yourself on your phone's voice memo app, or paste the text into a text-to-speech site and download the file.
- Find a background. Rain, brown noise, an ambient track. Usually pulled from a video site, which brings its own licensing question.
- Layer them. This is the CapCut part — import the voice, import the background, stack them on separate tracks.
- Balance the volumes. Voice too loud and it is jarring; too quiet and it is inaudible. This is where most of the 20 minutes goes.
- Loop and export. Repeat the voice across 30–40 minutes, render, and move the file somewhere you can actually play it at night.
None of these steps is hard. There are just six of them, and every one has a way to go wrong.
Where it usually falls apart
Volume balance between voice and background. Your phone speaker, your headphones and your car will each disagree about what sounds right. Without loudness normalisation you end up re-exporting.
Loop seams. Copy-pasting the same voice clip end to end creates an audible click or a hard silence at every join. Smoothing those is fiddly.
Sample-rate mismatches. A 24 kHz text-to-speech file layered under a 44.1 kHz rain recording can come out thin or slightly off-pitch depending on how the editor resamples.
The file lives in the wrong place. Plenty of people finish, then realise the export is stuck in an app that will not let them download it, or in a private video upload they need a connection to reach.
That last one shows up constantly: people are happy to pay for a tool, as long as they end up owning the file.
The two steps that actually matter
Strip the workflow back and only two decisions change the outcome.
1. What the lines say
This is the whole thing. When researchers had people repeat "I am a lovable person," participants with low self-esteem felt worse — the statement sat too far from what they believed, so the mind argued back (Wood, Perunovic & Lee, Psychological Science, 2009).
No amount of layering rescues a line you do not believe. We cover the wording rules in How to Write Affirmations That Actually Work, and the quantity question in How Many Affirmations a Day?
2. Whose voice says them
Material processed in relation to yourself tends to be remembered better than the same material processed impersonally — the self-reference effect, which we unpack in this piece. Values-based lines engage self-related processing too, particularly when tied to your future self (Cascio et al., Social Cognitive and Affective Neuroscience, 2016).
Those studies were not run on VōxSōma or on any DIY audio. They describe research, not results you should expect.
If the sound of your own recorded voice makes you wince, that reaction is normal and has a boring physical explanation — see why your recorded voice sounds different.
Everything else is production, not practice
Layer counts, exact background choice, how many times the voice repeats — these are audio decisions, not practice decisions. They affect how pleasant the file is to listen to. They do not decide whether the practice does anything for you.
Which is worth knowing before you spend an evening learning a video editor to get them right. There is no right. There is only what you will actually press play on.
The short version
If you want the outcome without the production, that is what VōxSōma is. You describe what you want to work on, seven lines come back, you edit them until they sound like you, and you read them aloud once. The layering, the levels, the loop seams and the 36-minute structure are handled.
No text-to-speech setup. No video editor. No exporting between apps. Your recording stays on your device — more on that in our note on voice privacy. One payment, no subscription.
It is a real trade: you give up granular control over the mix, and you get back the evening you would have spent on it.
If you still want to build it yourself
Genuinely a good option if you enjoy the making. The minimum viable setup:
- One recorder — your phone's voice memo app is fine. Quiet room, phone about a hand-width away, speak past the microphone rather than into it.
- One editor — any that lets you put two audio tracks side by side. You do not need a video editor for an audio file.
- One background — pick it once and stop reopening the decision.
- Seven lines, not seventy. Short enough that re-recording after a month is not a project.
We walk through the recording side in more detail in the own-voice guide, and the layering logic in how to make an affirmation track.
Frequently asked questions
Can I make affirmation audio on my phone alone?
Yes. A voice memo app plus any editor that supports two audio tracks is enough. A video editor like CapCut works but is more tool than the job needs.
Is text-to-speech as good as my own voice?
It is easier and it is fine. Your own voice has one structural advantage: the words stay yours through to playback. We compare both in Your Own Voice vs an App Voice.
How long should the finished audio be?
Long enough to cover the wind-down you actually do. Most people land between 20 and 40 minutes. Beyond that you are mostly making a file you will never hear the end of.
How many affirmations should the audio contain?
Fewer than you think. Seven focused lines repeated beats two hundred generic ones, mostly because you will keep doing it.
Do I need binaural or a specific frequency?
No. The evidence for binaural beats is mixed, and nothing about the practice depends on a particular frequency. Choose a background you find pleasant.
VōxSōma is a personal wellness audio tool — not a medical device, not therapy, and not intended to diagnose, treat, cure, or prevent any condition. Individual experiences vary. If you have a sleep, attention, or mental-health condition, please speak with a qualified clinician.