# Voiceover, Music & the Audio Studio

Every film needs a voice and a score. CoreReflex narrates your lines in cinematic HD voices — or a voice you describe in plain words — composes original instrumental music from a single mood word, and gives you a full multitrack **Audio Studio** to mix it all: faders and meters, a waveform clip editor, automation lanes, mic recording, and a one-click bounce that flows straight back into your video.

> **Where this lives.** Everything on this page happens in the studio editor at [editor.corereflex.com](https://editor.corereflex.com). Voiceover and music generation live in the **composition inspector** (the right-hand panel, next to color grading, transitions, and export) and inside the Audio Studio's own **Create** panel. The Audio Studio itself is the **Audio** button in the editor's top bar, right next to **Video**. If you came here straight from directing a film, open your cut in the editor first — see [Editing & Mastering](editing-and-mastering.md) for the full timeline workflow.

**On this page**

- [Voiceover](#voiceover) — type a line, pick a voice, get narration
- [Design a voice](#design-a-voice) — describe a voice in words and speak with it
- [Voice cloning (consented)](#voice-cloning-consented)
- [Music](#music) — an original score from a prompt and a mood
- [How voice and music sit together](#how-voice-and-music-sit-together) — automatic ducking
- [The Audio Studio](#the-audio-studio) — sessions, mixer, clip editor, automation, recording, bounce
- [Credits and metering](#credits-and-metering)
- [Practical tips](#practical-tips)
- [Troubleshooting](#troubleshooting)
- [Where to go next](#where-to-go-next)

---

## Voiceover

Type a line, pick a voice, and CoreReflex narrates it for you in a second or two. The narration lands on your timeline as an audio clip you can move, trim, and stack like any other.

The **Voiceover** panel appears in two places: the video composition inspector, and the Audio Studio's **Create** panel (where generated clips land inside your open session instead — see [Generate into a session](#generate-voice-and-music-into-a-session)). At the top of the panel are two tabs: **Pick a voice** (this section) and **Design a voice** ([next section](#design-a-voice)).

### The voices

CoreReflex gives you two kinds of narration voice — a set of **premium Studio HD voices** for the most polished, cinematic read, and **CoreReflex Voice**, a natural narrator that runs on CoreReflex's own infrastructure for a fraction of the credits. You pick which one from the voice dropdown; the two engines are compared [just below](#studio-hd-or-corereflex-voice-picking-your-narrator).

The premium **Studio HD** voices are powered by **Chirp 3 HD** — Google's premium, natural-sounding text-to-speech, running first-party on Vertex AI. CoreReflex curates five of them, each tuned for a different kind of film:

| Voice | Tone | Best for |
|-------|------|----------|
| **Charon** | Deep, authoritative | Trailers, hero pieces, dramatic reveals |
| **Aoede** | Warm, bright | Brand films, founder stories, welcomes |
| **Kore** | Clear, confident | Explainers, how-tos, product walkthroughs |
| **Puck** | Upbeat, energetic | Social clips, ads, fast-paced cuts |
| **Fenrir** | Gritty, intense | Action, sports, high-energy spots |

**Charon** is the default. All five speak US English.

### Studio HD or CoreReflex Voice: picking your narrator

Above the voice dropdown you choose the **voice model** — the engine that actually speaks your line:

| Voice model | What it is | Credits per line |
|-------------|------------|------------------|
| **Studio HD** (the five voices above) | The premium, top-tier narration engine — the most cinematic, expressive read. | **5** |
| **CoreReflex Voice** | A natural, clear narrator that runs on CoreReflex's own infrastructure. Great for explainers, drafts, high-volume work, and anywhere you want a solid read for far fewer credits. | **1** |

Pick **CoreReflex Voice** and the same *type a line → generate → it lands on your timeline* flow works exactly the same — you just spend one credit instead of five. Reach for a **Studio HD** voice when you want the extra polish of a premium performance for a hero piece or final delivery, and lean on **CoreReflex Voice** for everything in between. It's the same studio; only the credit cost and the character of the read change.

> **Availability.** CoreReflex Voice is switched on by your operator. Where it hasn't been enabled yet, the dropdown simply shows the premium Studio HD voices — CoreReflex never quietly swaps one engine for another behind your back.

### Add a line of narration

1. Open your film in the editor and find the **Voiceover** panel in the inspector.
2. Make sure the **Pick a voice** tab is selected.
3. **Type the line** you want narrated into the text box.
4. **Pick a voice** from the dropdown.
5. Click **Generate voiceover.** You'll see it move through *Synthesizing…* and then *Adding to timeline…*
6. The narration clip lands on your audio track. Drag it to position it under the right shot.

To add another line, just type the next one and generate again — each line becomes its own clip, so you can voice a film one beat at a time and arrange the clips exactly where you want them.

### Good to know

- **One line at a time.** Each generation creates a single narration clip. Build a multi-line voiceover by generating each line and placing it, rather than pasting a whole script at once. This gives you precise control over timing and pacing.
- **Length limit.** A single voiceover line can be up to **5,000 characters** — far more than you'll ever want in one breath of narration, but worth knowing if you paste a long passage.
- **Speaking rate.** Voices narrate at a natural, even pace by default. You can nudge the speed between **0.5× and 1.5×** if a line needs to sit faster or slower against your footage — anything outside that range is clamped to it. For bigger pacing changes, break the line into shorter clips and space them out on the timeline.
- **Trim and layer freely.** Once a narration clip is on the timeline it behaves like any audio clip — trim its head and tail, nudge it frame by frame, or detach it to a separate track. See [Editing & Mastering](editing-and-mastering.md).
- **Master it right on the clip.** A generated narration clip carries its own **Audio FX** rack — the same compress/limit/de-ess/denoise/reverb chain as the master, plus presets and a one-click AI auto-master. Polish the voice on the clip itself instead of on the master bus. (Real processing is gated on the self-hosted audio worker — see [Audio FX & mastering](editing-and-mastering.md#per-clip-audio-fx-the-same-rack-on-every-clip).)
- **Your audio, your fleet.** Generated narration is stored in your own workspace storage and served back through a private signed link — it never lands on a third-party host.
- **It's metered.** A Studio HD line costs **5 credits**; a **CoreReflex Voice** line costs just **1** — see [Credits and metering](#credits-and-metering). Either way you'll see the cost and confirm it before it runs.

---

## Design a voice

Don't want any of the five studio voices? **Describe the voice you hear in your head** — *"warm, gravelly late-night radio host"*, *"bright, fast, youthful explainer"* — and CoreReflex speaks your line with it.

### Design and generate

1. In the **Voiceover** panel, switch to the **Design a voice** tab.
2. **Describe the voice** in the description box — gender, age, pace, energy, and texture all count. Up to **600 characters**.
3. **Type the line** to narrate in the text box below it.
4. Click **Design voice & generate.**
5. The clip lands on your timeline (or in your open audio session) exactly like a normal voiceover line.

### What actually happens — honestly

What "design" does depends on how your CoreReflex instance is set up, and the panel tells you which you'll get *before* you click:

- **Standard setup** — the panel says *"Your description is matched to the closest studio voice."* CoreReflex reads your description for gender, age, pace, energy, and timbre cues, scores it against the five curated voices, picks the closest match, and derives a natural speaking rate from your words (a "slow, measured" brief reads slower; a "snappy, energetic" one reads quicker). You always get a result.
- **Voice-design engine deployed** — when your operator has enabled a self-hosted voice-design engine, the panel says a bespoke voice is designed from your description, and your exact words shape a brand-new synthetic voice rather than mapping to a preset. The panel names the engine in use.

Either way, a designed voice **never clones a real person** — no reference recording is involved, so there is no consent step. To speak in a *specific real* voice you have rights to, use [voice cloning](#voice-cloning-consented) instead.

### Tips for a good description

- Lead with the strongest traits: *"deep, authoritative"* beats a paragraph of backstory.
- Say the gender explicitly (*"a woman's voice"*) if it matters — a bare "she" in passing is treated as a weak hint, so texture words can override it.
- Pace words work: *slow, measured, unhurried* vs. *fast, snappy, brisk*.
- Texture words work: *deep, gravelly, warm, bright, smooth, commanding*.

Designing a voice is metered the same as any voiceover line: **5 credits** per generation.

---

## Voice cloning (consented)

Beyond the studio and designed voices, CoreReflex can speak in a **specific voice you have the rights to use** — your own, or one you're authorized to clone. By default this runs on **CoreReflex's own infrastructure**, the same in-house engine behind CoreReflex Voice — your reference sample and the cloned voice stay on your fleet, not shipped off to a third-party cloning service. A cloned voice is attached to a **persona**, and once it's set, persona clips rendered through the configured persona engine speak in that cloned voice instead of a stock one.

> **Voice cloning lives in the Persona Studio,** opened from the **Clone a persona** card on your [dashboard](https://corereflex.com). It is a **Pro** feature, and it is gated behind a hard consent step — see below. Voice cloning never runs on the plain **Voiceover** panel; that panel only ever uses the five curated HD voices or a designed (non-cloned) voice.

### The consent gate

A persona — and therefore its cloned voice — **cannot be created or used until you record an ownership/likeness consent attestation.** You confirm, on the record, that you own or are authorized to use the voice and likeness. CoreReflex stamps who attested, when, and the exact statement onto the persona, and refuses to clone a voice or render a clip until that consent is granted. Revoke consent and the persona is disabled immediately. This is deliberate: it keeps voice cloning to voices you actually have permission to use.

### Two ways to clone

When you clone a persona's voice, CoreReflex supports two engines behind one consent-gated flow:

| Engine | How it clones | What you provide |
|--------|---------------|------------------|
| **Self-hosted** (default) | Zero-shot — the open voice engine matches a short reference sample at synthesis time, no enrollment step | A **6–10 second** reference sample, as a public link |
| **Chirp Instant Custom Voice** | Mints a Google cloning key from a reference plus a spoken consent recording | A reference recording **and** a consent recording reading the exact consent script |

The self-hosted, zero-shot path is the default when the operator has deployed the open voice engine. The Chirp Instant Custom Voice path is **allow-list gated by Google** and only available once your project has been approved for it; until then, choosing it returns a clear message rather than failing silently.

Once a voice is cloned, it's saved on the persona. From then on, rendering a persona clip that "speaks" a line uses that cloned voice automatically — you don't pick it again each time. The final talking-head video still needs the persona lip-sync engine to be configured; until then the studio keeps the persona, reference media, and consent record ready without pretending the final render is done.

---

## Music

CoreReflex writes **original, royalty-free instrumental music** for your film with **Vertex AI Lyria**. You describe the music you want and pick a mood, and Lyria composes a fresh score — no lyrics, no samples, nothing lifted from existing songs. It's yours to use.

### Score your film

1. Open the **Soundtrack** panel in the inspector (or in the Audio Studio's **Create** panel).
2. **Describe the soundtrack** you want in the prompt box — for example, *"slow, hopeful piano building to strings"* or *"driving synth pulse for a product reveal."*
3. **Pick a mood.** The mood shapes the overall feel and pairs with your prompt:

   | Mood | Feel |
   |------|------|
   | **Cinematic** | Sweeping orchestral, epic, film-trailer energy (default) |
   | **Upbeat** | Bright, energetic modern pop with a driving rhythm |
   | **Ambient** | Calm, atmospheric, spacious pads |
   | **Tense** | Dark, pulsing, dramatic suspense |
   | **Uplifting** | Warm, hopeful, corporate-positive acoustic |
   | **Lo-fi** | Relaxed, mellow lo-fi hip-hop with vinyl warmth |

4. Click **Generate soundtrack.** The track is composed in a few seconds, then placed on your timeline.

### Good to know

- **Clips are about 30 seconds.** Lyria produces roughly a **30-second** instrumental. For a longer film, generate a few tracks and lay them end to end, or loop and trim a single track to fit.
- **Prompt length.** The music prompt can be up to **500 characters** — keep it focused on instrumentation, energy, and feel.
- **Mood + prompt work together.** The mood you pick gives the composition a strong starting point even from a short prompt, so a one-word direction like *"Tense"* plus a brief phrase is enough to get a coherent score.
- **Always original.** Lyria is built to never reproduce existing songs. CoreReflex leads every request with an explicit originality instruction, and if a very specific prompt is ever refused for sounding too close to a real piece of music, it automatically retries once with a stronger, more original phrasing. If it's still refused, rephrase to describe the instrumentation and feel you want (rather than naming an artist or track) and generate again.
- **It's instrumental.** Music has no vocals, by design — so it never competes with your narration for the listener's attention.
- **Master the bed on the clip.** A generated soundtrack clip carries the same per-clip **Audio FX** rack as everything else — its AI auto-master recognizes a music clip and reaches for the *Music master* chain, so you can warm and glue the bed in one click before it goes under your voice. See [Audio FX & mastering](editing-and-mastering.md#per-clip-audio-fx-the-same-rack-on-every-clip).
- **Saved to your library, not the public feed.** Each track is stored in your own workspace and logged as a music asset. Music never appears in CoreReflex's public video showcase — only your finished, intentionally shared films can.
- **It's metered.** Each generated track costs **20 credits** when credit metering is on — see [Credits and metering](#credits-and-metering).

---

## How voice and music sit together

CoreReflex ducks music under voice at two levels — a quick automatic one on the video timeline, and a precise, per-region one in the Audio Studio.

### On the video timeline: automatic whole-clip duck

When you **generate a soundtrack** and it lands overlapping your voiceover (or a video clip carrying its own audio), CoreReflex **automatically lowers the whole music clip by 10 dB** so your narration stays clearly on top. You don't balance the mix by hand — the moment the generated score drops in over voice, it steps back.

- **It happens at generation time.** The duck is applied when the soundtrack is placed. If you generate music first and add voice later, the music won't retroactively duck — either regenerate the score after your voice is down, or lower the music clip's volume yourself.
- **It carries through to the final film.** The duck is an ordinary clip-volume change, so it's baked into your export — what you hear in the editor preview is what ships.
- **You stay in control.** It's a smart starting point, not a lock. Adjust the clip's volume on the timeline any time.

### In the Audio Studio: per-region ducking with Auto-duck

For the finer mix — music that dips *only while someone is speaking* and swells back between lines — use the **Auto-duck** button in the Audio Studio. One click writes a volume envelope on every music track in your session: the music drops **10 dB** under each voice region with smooth **quarter-second ramps** down and back up, and back-to-back voice clips merge into a single dip so the score never flutters. The result is ordinary, hand-tunable automation. See [Auto-duck](#auto-duck-dip-the-music-under-every-voice-line) below for the walkthrough.

---

## The Audio Studio

The Audio Studio is CoreReflex's multitrack audio room — the place to mix voice, music, and recordings properly, with picture on screen, before the result flows back into your video. It sits right next to the video editor: click **Audio** in the editor's top bar.

Work in the Audio Studio is organized into **audio sessions**. A session is a self-contained multitrack arrangement — its own tracks, clips, mixer settings, and automation — saved with your project like everything else. Sessions undo/redo with the rest of your edit, and everything you bounce out of one lands in your own media library.

You can start a session two ways:

- **From scratch** — create an empty session and build a mix (a podcast bed, a music cue, a full soundscape).
- **From your video** — select a clip on the video timeline and click **Edit in Audio**; CoreReflex seeds a session from that clip, keeps it linked, and later bounces your mix straight back into the video. See [the round-trip](#the-round-trip-edit-in-audio-and-hot-update) below.

### Create and manage sessions

1. Click **Audio** in the editor's top bar. With no session open, the center of the screen is the **session list**.
2. Click **+ New session** to create one — it opens immediately.
3. Back on the list, each session row shows its name and track count, plus:
   - **Open** — click the row.
   - **Rename** — double-click the name (or click **Rename**), type, press Enter.
   - **Delete** — click **Delete**, then click **Confirm?** to really delete the session and its tracks.
   - **Place in video** — appears once a session has been bounced; drops the mixed WAV onto your video timeline as a normal audio clip and switches you to the Video view. Un-bounced sessions show **Bounce first** instead.

### The session screen

Open a session and the studio arranges itself around it:

- **Header** — **← Sessions** takes you back to the list. Next to the session name, a badge tells you the session's place in your video: **Bounced** once a mix exists, or **N in video** when clips on your video timeline are bound to this session (those are the clips a re-bounce will update).
- **Center: the session player** — plays the session's audio, and shows **picture** when the session carries a picture-reference video track (see the round-trip below), so you mix against the image. The player is also a drop target: drag audio or video files onto it and they land on the session's tracks.
- **Bottom: the session timeline** — the same multitrack timeline you know from the video editor (drag, trim, split, snap to the grid), scoped to this session. Its toolbar carries the session's main actions: **Record**, the grid controls (timecode or BPM ruler — shared with the video timeline, so a BPM you set here is the same grid there), **Auto-duck**, quick **Soundtrack** and **Voiceover** generators, the **loudness target** selector, and **Bounce**.
- **Side panels** — **Media** (place any project asset into the session), **Create** (record, voiceover, soundtrack — full-size panels), **Mixer**, and **Clip editor**.

### Add sound to a session

- **From your library:** open the **Media** panel. Your project's audio assets are listed up front (everything else under *All media*). Click a row to place it in the open session — audio lands as a clip at the playhead; a video asset lands as a **muted picture-reference track**, so you can mix against footage without doubling its sound.
- **By drag and drop:** drop files onto the session player or the session timeline. They upload to your workspace and land on the session's tracks.
- **By recording or generating** — next two sections.

### Record into a track

Record a mic take — a scratch VO, a real narration, a podcast segment — straight into the session:

1. Park the playhead where the take should start.
2. Click **Record** in the session toolbar (or in the **Create** panel). Your browser asks for microphone permission the first time — allow it.
3. The button turns into a red elapsed-time counter while you record.
4. Click it again to stop. The take is placed as a clip **at the frame where you started recording**, on the first session track with room — recording into an empty session creates its first track automatically.
5. The take is kept exactly as recorded, cached locally for instant playback, and uploaded to your workspace storage in the background.

### Generate voice and music into a session

The same **Voiceover** (pick or design) and **Soundtrack** generators described above are available inside the Audio Studio, wired to your open session:

- The **Create** panel carries the full-size panels.
- The session toolbar carries compact **Voiceover** and **Soundtrack** popovers for quick one-off lines and cues.

Either way, the generated clip lands **on the open session's tracks at the playhead** — not on the video timeline. (With no session open, the Create panel tells you so, and generation falls back to the video timeline.) A soundtrack generated into a session auto-ducks against that session's voice clips, never against the video timeline — the two surfaces stay independent.

### The mixer

Open the **Mixer** panel with a session open: one **channel strip per track**, plus the **Master** strip.

Each channel strip gives you:

- **Pan** — the slider at the top, left/center/right with a readout (L50, C, R25…). *Honesty note:* pan is applied in the **bounce**, not the live preview — the session preview plays mono-summed, so you'll hear your panning in the bounced mix and the final video.
- **Gain fader** — a vertical fader from **−60 dB to +12 dB**. Drag it; it magnetizes to exactly **0 dB (unity)** near the detent, and a double-click resets it to 0. The dB readout sits underneath.
- **Level meter** — a live peak meter beside the fader, running while the session plays.
- **M / S** — mute and solo per track.

The **Master** strip carries the session's master **level meter** and the session's **master FX rack** — the same compressor/limiter/reverb rack you know from clips and the project master. Master FX is **baked into the bounce**, not the live preview (the rack says so on its face) — same contract as the project-master rack, which bakes into the export.

Mixer moves are undo-friendly: one drag of a fader is one undo step.

### The clip editor

Select any audio clip in the session timeline and open the **Clip editor** panel — a magnifier for that one clip:

- **Waveform** — a big, zoomable waveform of the clip, with **draggable fade ramps** right on the audio so you can set fade-in/fade-out by eye.
- **Clip gain** — the clip's own volume, separate from its track fader.
- **Fades** — numeric fade-in/fade-out controls, matching the ramps on the waveform.
- **Speed** — the clip's playback rate.
- **Split at playhead** — cuts the clip in two at the session playhead (park the playhead inside the clip first; the button stays disabled otherwise).
- **Clip FX** — the same per-clip Audio FX rack as everywhere else in the studio.

### Automation lanes

Every session track can carry **volume and pan automation** — values that change over time, drawn as a curve over the track:

1. In the session timeline, find the small **A** chip at the top-left of a track row. (It glows when the track already has automation.)
2. Click it to expand the automation lane over the row, with **Vol** and **Pan** buttons to switch which curve you're editing.
3. **Double-click empty space** to add a keyframe at that time and value.
4. **Drag a point** to move it — the curve follows live, and the move lands as one undo step when you release.
5. **Double-click a point** (or Alt-click it) to remove it.
6. Click the **A** chip again to collapse the lane.

Values between keyframes are smoothly interpolated. Volume automation multiplies the track's level (from silence up to 2×, with 1 = no change); pan automation sweeps left to right. Automation is audible in the live session preview and rendered identically in the bounce.

> While a lane is expanded, it sits on top of the row — collapse it (click **A**) to drag the clips on that row again.

### Auto-duck: dip the music under every voice line

The one-click version of a mixing chore every editor knows:

1. Get your voice clips and music clips onto the session's tracks.
2. Click **Auto-duck** in the session toolbar.
3. CoreReflex finds every voice region in the session and writes a **volume automation envelope on every music track**: the music drops **10 dB** during each voice region, with smooth **0.25-second ramps** down and back up. Back-to-back voice clips merge into one clean dip instead of a flutter. A toast confirms what it did — e.g. *"Ducked 1 music track under 3 voice regions."*

What makes this better than a baked-in effect:

- **It's just automation.** Open the music track's automation lane (the **A** chip) and you'll see the envelope as ordinary keyframes — drag any of them to hand-tune the depth or timing.
- **Audible immediately, rendered identically.** The live session preview and the bounce read the same automation.
- **One undo.** The whole envelope is a single Cmd+Z away.

How clips are classified: by their **names**. Everything CoreReflex generates is named for you (voiceover clips, soundtrack clips, mic recordings), so generated and recorded material classifies itself. For imported files, make sure the name says what it is — files with *voice, narration, speech, dialog, podcast, vocal,* or *recording* in the name count as voice; files with *music, soundtrack, song, beat, score, instrumental,* or *bgm* count as music. A track that carries any voice clip is never ducked, and a session with no recognizable voice clip or no music track gets a clear "Nothing to duck" message instead of a silent no-op.

### Bounce to WAV — with loudness targets

**Bounce** mixes the whole session down to a single high-quality stereo WAV (48 kHz) — the session's finished mix.

1. (Optional) Pick a **loudness target** from the dropdown next to the Bounce button:

   | Target | For |
   |--------|-----|
   | **Loudness: off** | No normalization — the mix ships at its natural level |
   | **−14 LUFS · streaming** | YouTube, Spotify, and most streaming platforms |
   | **−16 LUFS · social / podcast** | Social feeds and podcast distribution |
   | **−23 LUFS · broadcast** | Broadcast delivery |

   The target is saved on the session, so every re-bounce keeps hitting the same spec.
2. Click **Bounce.**
3. The mix renders offline — nothing plays out loud — and lands as a WAV asset in your media library, playable instantly.

What the bounce includes, in order: every clip's trim, speed, gain, fades, and clip FX; every track's fader, pan, mute/solo, and automation (including Auto-duck envelopes); the session's master FX rack; and finally loudness normalization to your target. Loudness is measured to the broadcast standard (BS.1770 integrated loudness), and the gain is **peak-capped at −1 dBFS** — a loud target on quiet, peaky material gets as close as the peak allows rather than clipping, and the confirmation tells you so (e.g. *"−18.2 → −14 LUFS (peak-capped)"*).

If any clips in your **video** are bound to this session, the bounce **updates them all instantly** — see the round-trip below. The whole bounce (new asset + updated clips) is one undo entry.

### The round-trip: Edit in Audio and hot-update

This is the Audio Studio's signature move — pull sound out of your video, mix it properly, and push the result back without ever exporting a file by hand:

1. **From an audio clip:** select it on the video timeline and click **♪ Edit in Audio** in the inspector. A session opens, seeded with a copy of the clip — same source, trim, level, fades, and FX.
2. **From a video clip with sound:** select it and click **♪ Edit audio in Session**. CoreReflex detaches the clip's audio onto its own timeline item, seeds a session from it, and adds a **muted picture-reference track** so the session player shows the footage while you mix against it.
3. **Mix.** Faders, automation, Auto-duck, clip FX, extra tracks, recordings — anything.
4. **Bounce.** The session's mix replaces the audio of every bound clip in the video, instantly — the confirmation reads *"Bounced — updated 1 clip in the video."* Your video composition is hot-updated in place; play it and you hear the new mix.
5. **Iterate freely.** The video clip stays linked to its session (the inspector shows *Bounced from "session name"*, and the session header shows *N in video*). Click **Edit in Audio** again any time to reopen the same session; every re-bounce updates every bound clip again.

You can also start from the other end: bounce a standalone session and use **Place in video** on the session list to drop the mix onto the video timeline as a bound clip — future re-bounces keep it fresh too.

---

## Credits and metering

Generating audio with AI is metered; mixing it is not.

| Operation | Credits per run |
|-----------|-----------------|
| Voiceover line — **Studio HD** voice (picked or designed) | **5** |
| Voiceover line — **CoreReflex Voice** | **1** |
| Soundtrack (one ~30-second track) | **20** |

Everything else on this page — the whole Audio Studio, mixing, automation, Auto-duck, recording, the clip editor, bouncing at any loudness target — draws **no credits**.

Credits are only charged when your instance has credit metering switched on, and CoreReflex stops you *before* an operation your balance can't cover — nothing is partially charged. See [Plans & Billing](plans-and-billing.md) for how credits and plans work.

---

## Practical tips

- **Voice first, then score.** Lay your narration down where you want it, then generate music. On the video timeline that's what makes the automatic duck kick in; in a session it's what gives Auto-duck its voice regions to work with.
- **Match the voice to the mood.** A *Cinematic* score under **Charon** reads as a trailer; *Uplifting* under **Aoede** reads as a warm brand film. Picking a voice and mood that agree makes the whole piece feel intentional.
- **Reach for Design a voice when the five presets don't fit.** It always produces something usable — on a standard setup you'll get the closest studio voice at a pace derived from your description, which is often exactly the nudge you needed.
- **Build long pieces from clips.** Both voice (per line) and music (about 30 seconds) come in clips. For anything longer, think in segments: voice each beat, score each section, and arrange the clips on the timeline.
- **Keep music prompts about feel, not songs.** Describe instruments, tempo, and energy — *"warm acoustic guitar, gentle build"* — rather than referencing real tracks. You'll get better, faster results and avoid originality refusals.
- **Name imported files honestly.** Auto-duck classifies clips by name — a music file called `track-01.mp3` won't be recognized as music. `bed-music.mp3` will.
- **Auto-duck, then hand-tune.** Let Auto-duck write the envelope, then open the automation lane and drag a keyframe or two where the music should breathe differently. It's much faster than drawing the whole curve yourself.
- **Set the loudness target once.** Pick −14 LUFS (or your platform's spec) on the session before your first bounce — it's remembered, so every revision ships at the same level.
- **Mix against picture.** Seeding a session from a video clip gives you the footage in the session player — timing music swells and dialogue dips to the image beats guessing with timecode.
- **Clone only voices you're cleared to use.** The consent gate is there for a reason. Keep a clean reference sample (6–10 seconds, clear, single speaker) and you'll get a far better clone.
- **Polish in the editor.** Trim, nudge, and layer your voice and music clips, then add color, transitions, and master-bus audio effects on the way out. See [Editing & Mastering](editing-and-mastering.md).

---

## Troubleshooting

- **"Nothing to duck — this session needs a voice clip and a music track."** Auto-duck found no recognizable voice region or no music track. Check that your session has both, and that imported files have voice-ish / music-ish names (see [Auto-duck](#auto-duck-dip-the-music-under-every-voice-line)). Generated and recorded clips are always recognized.
- **I can't hear my panning.** Pan (and the session's master FX rack) are applied in the **bounce**, not the live session preview. Bounce the session and play the bounced clip — the pan is there.
- **Split at playhead is greyed out.** The session playhead must be *inside* the selected clip. Move the playhead into the clip and try again.
- **I can't drag clips on a track.** An expanded automation lane sits over the row. Click the **A** chip to collapse it, then drag.
- **"Bounce first."** *Place in video* needs a mix to place — click into the session and **Bounce**, then place it.
- **"Bounce asset is missing."** The session's last mix was removed from the project. Bounce the session again to recreate it.
- **"Microphone access denied."** Your browser blocked the mic. Allow microphone access for the editor in your browser's site permissions and click **Record** again.
- **Music generation was blocked twice.** Your prompt likely resembles an existing piece too closely, even after the automatic originality retry. Rephrase to instruments, tempo, and energy — don't name artists, songs, or lyrics.
- **A designed voice sounds like one of the five studio voices.** On a standard setup, that's exactly what it is — the closest match to your description (the panel says so before you generate). A bespoke designed voice requires the operator-deployed voice-design engine.
- **"Cloned/designed voice synthesis needs the self-hosted voice engine."** A cloned (or engine-designed) voice was requested but the open voice engine isn't deployed on your instance. Ask your operator — CoreReflex fails honestly here rather than silently substituting a stock voice.
- **A voiceover or soundtrack won't generate and mentions credits.** Your workspace balance can't cover the run (voiceover 5, music 20). Top up or upgrade — see [Plans & Billing](plans-and-billing.md).

---

## Where to go next

- **[Editing & Mastering](editing-and-mastering.md)** — trim, layer, grade, and apply master audio effects on the way to a finished cut.
- **[Directing Films](directing-films.md)** — brief the agentic director to board and produce your shots, then build a consented persona that speaks in a cloned voice.
- **[Recording, Meetings & Podcasts](recording-and-podcasts.md)** — capture real audio and video, master it with one-click presets, and meet the real-time voice agent.
- **[Plans & Billing](plans-and-billing.md)** — how tiers, credits, and metering work.
- **[Coming Soon](coming-soon.md)** — what's shipping next.
