WEBVTT

00:00:00.000 --> 00:00:03.538
<v Drew>Hey. So somebody asked me the other
day... can you actually make a podcast

00:00:03.538 --> 00:00:07.076
<v Drew>with AI? Like, the whole thing? And
here's the thing — you're listening to

00:00:07.076 --> 00:00:10.759
<v Drew>one right now. This whole show is made
with AI. So... yeah. The answer's yes.

00:00:10.759 --> 00:00:13.328
<v Drew>But there's a whole spectrum of what
"yes" looks like.

00:00:13.328 --> 00:00:17.602
<v Drew>This is AI Honestly. I'm Drew. Each
episode I take one real question from a

00:00:17.602 --> 00:00:22.106
<v Drew>founder or business owner — and tell you
what's real, what's not, and what it'd

00:00:22.106 --> 00:00:24.474
<v Drew>actually take to build. No hype. No DOOM.

00:00:29.474 --> 00:00:29.474
<v Drew><break time="2s" />

00:00:29.474 --> 00:00:33.305
<v Drew>So I get this question because making a
podcast sounds like a lot of work.

00:00:33.305 --> 00:00:37.293
<v Drew>Recording is hard. Editing is hard.
Writing a script is hard. And most people

00:00:37.293 --> 00:00:41.438
<v Drew>who have something to say don't have a
free afternoon every week to sit down and

00:00:41.438 --> 00:00:45.111
<v Drew>do all that. So when they hear "AI can
make your podcast," what they're

00:00:45.111 --> 00:00:48.995
<v Drew>wondering is... can I skip the tedious
part and still end up with something

00:00:48.995 --> 00:00:49.257
<v Drew>good?

00:00:49.257 --> 00:00:53.251
<v Drew>And the answer is yes. With a catch. Let
me walk you through how this show

00:00:53.251 --> 00:00:57.245
<v Drew>actually gets made, because it's a pretty
good example of the whole range.

00:00:57.245 --> 00:01:00.972
<v Drew>Here's the flow - I talk. That's it,
that's my part. I open the voice memo app

00:01:00.972 --> 00:01:04.894
<v Drew>on my phone and I just... ramble. I could
be in traffic with sirens going. I could

00:01:04.894 --> 00:01:08.524
<v Drew>be in the car, I could be at my desk.
Doesn't matter. I go back and forth, I

00:01:08.524 --> 00:01:12.349
<v Drew>repeat myself, I lose my train of thought
— none of it matters, because the next

00:01:12.349 --> 00:01:13.220
<v Drew>steps clean it up.

00:01:13.220 --> 00:01:17.664
<v Drew>The voice memo app gives me an automatic
transcript. I take that transcript and I

00:01:17.664 --> 00:01:22.163
<v Drew>feed it into AI, and that's where it gets
turned into a script. Now, I've set up a

00:01:22.163 --> 00:01:26.552
<v Drew>bunch of guidelines ahead of time. What
an episode is, what it should cover, the

00:01:26.552 --> 00:01:30.162
<v Drew>tone, all that. So it's not just
summarizing me. It's reshaping my

00:01:30.162 --> 00:01:33.051
<v Drew>rambling into something that actually
holds together.

00:01:33.051 --> 00:01:37.302
<v Drew>But that script sounds like AI if you
don't push it the extra mile. You know

00:01:37.302 --> 00:01:41.666
<v Drew>the tells. The em-dash with the little
descriptor after it. The line that goes

00:01:41.666 --> 00:01:46.087
<v Drew>"it's not this - it's that." All that
stuff. People can smell it. So I've got a

00:01:46.087 --> 00:01:50.508
<v Drew>long list of things I tell it not to do,
and that gets me maybe 80% of the way.

00:01:50.508 --> 00:01:54.702
<v Drew>The last 20% is me, by hand, just
chopping. Cutting the AI-scent stuff out.

00:01:54.702 --> 00:01:58.840
<v Drew>And that part actually works for me,
because I'd rather a show be a little

00:01:58.840 --> 00:02:01.844
<v Drew>abrupt than a little eloquent. I don't
need the fluff.

00:02:01.844 --> 00:02:05.994
<v Drew>Then the script goes into ElevenLabs.
That's the voice part. It's a clone of my

00:02:05.994 --> 00:02:10.144
<v Drew>voice, and it reads the script back to
me. And ElevenLabs is also where I do my

00:02:10.144 --> 00:02:14.241
<v Drew>editing — I'm listening live, chopping
things, rearranging, and that's where I

00:02:14.241 --> 00:02:18.445
<v Drew>drop in my theme music too, which is a
little AI-generated track. I just tuck it

00:02:18.445 --> 00:02:20.094
<v Drew>under the front of the episode.

00:02:20.094 --> 00:02:24.289
<v Drew>Here's a thing about voice cloning,
though. You'd think you'd want the most

00:02:24.289 --> 00:02:28.598
<v Drew>polished, professional clone possible.
You can do that. But a really polished

00:02:28.598 --> 00:02:33.077
<v Drew>clone often has this... AI scent to it.
Sounds too clean. In my case?? My actual

00:02:33.077 --> 00:02:37.669
<v Drew>delivery is kind of monotone and flat. So
when I run it through AI, the quirks get

00:02:37.669 --> 00:02:42.261
<v Drew>masked, because flat is just how I talk.
It works in my favor. That's not gonna be

00:02:42.261 --> 00:02:46.570
<v Drew>true for everybody — but it's worth
knowing that "more polished" isn't always

00:02:46.570 --> 00:02:47.307
<v Drew>"more real.".

00:02:47.307 --> 00:02:49.815
<v Drew>Okay. What it'd take if you wanted to do
this yourself.

00:02:49.815 --> 00:02:53.680
<v Drew>The voice part — ElevenLabs — is the one
real cost. You're looking at their

00:02:53.680 --> 00:02:57.806
<v Drew>Creator plan, which is twenty-two dollars
a month. That's the tier that gets you

00:02:57.806 --> 00:03:01.933
<v Drew>the good voice cloning. And it runs on
credits. Here's the thing to watch: every

00:03:01.933 --> 00:03:05.902
<v Drew>time you edit and re-render the voice,
you burn credits. So if you're doing a

00:03:05.902 --> 00:03:09.872
<v Drew>weekly show with a lot of fiddly editing,
you can chew through your allotment

00:03:09.872 --> 00:03:13.685
<v Drew>faster than you'd think, and then you're
paying overages. Budget for that.

00:03:13.685 --> 00:03:17.808
<v Drew>The rest of my pipeline is basically
free. After ElevenLabs, I export two

00:03:17.808 --> 00:03:22.102
<v Drew>files — the MP3 of the episode, and a VTT
file, which is the transcript with

00:03:22.102 --> 00:03:26.511
<v Drew>timestamps. I hand those to my pipeline
in Claude Code and I just say, publish

00:03:26.511 --> 00:03:30.520
<v Drew>this episode. And it does the whole
thing. It writes the show notes. It

00:03:30.520 --> 00:03:35.158
<v Drew>builds the new entry in the feed, the one
Apple and Spotify and Amazon pick up. It

00:03:35.158 --> 00:03:39.452
<v Drew>creates a page on my website for the
episode — which is partly SEO, but it's

00:03:39.452 --> 00:03:43.919
<v Drew>more than that, people actually go back
and read these. And it publishes all of

00:03:43.919 --> 00:03:44.892
<v Drew>it to Cloudflare.

00:03:44.892 --> 00:03:48.668
<v Drew>And Cloudflare makes this nearly free.
Normally, hosting a podcast — unless

00:03:48.668 --> 00:03:52.700
<v Drew>you're locked into Spotify's ecosystem —
runs you at least twenty bucks a month.

00:03:52.700 --> 00:03:56.680
<v Drew>Cloudflare is so cheap that storing this
stuff is basically pennies. Maybe if I

00:03:56.680 --> 00:04:00.661
<v Drew>hit a hundred episodes and I'm storing
all of it, I'd start paying something. I

00:04:00.661 --> 00:04:02.447
<v Drew>don't know. But right now? Pennies.

00:04:02.447 --> 00:04:06.743
<v Drew>There's even some simple analytics baked
in. A little dashboard — filters out

00:04:06.743 --> 00:04:10.643
<v Drew>bots, deduplicates, shows me who
downloaded an episode, who viewed the

00:04:10.643 --> 00:04:14.656
<v Drew>page. Nothing fancy, but it's there. And
the system also pulls out clips

00:04:14.656 --> 00:04:19.008
<v Drew>automatically — finds the interesting
soundbites, drops them on a graphic with

00:04:19.008 --> 00:04:23.417
<v Drew>subtitles and the episode question up
top. So I've got social posts ready to go

00:04:23.417 --> 00:04:24.830
<v Drew>without doing extra work.

00:04:24.830 --> 00:04:27.616
<v Drew>Now — what to watch for, beyond the
credits thing.

00:04:27.616 --> 00:04:32.106
<v Drew>Building the pipeline itself takes real
time up front. The per-episode workflow

00:04:32.106 --> 00:04:36.539
<v Drew>is fast — ten minutes to talk into my
phone, then up to a half hour of editing

00:04:36.539 --> 00:04:41.030
<v Drew>before it's live. But the pipeline that
makes that possible? That was the work.

00:04:41.030 --> 00:04:45.635
<v Drew>So if you're picturing zero effort, it's
zero effort after you've done the setup.

00:04:45.635 --> 00:04:50.104
<v Drew>The other catch. The hard part of a
podcast was never the production. And AI

00:04:50.104 --> 00:04:54.871
<v Drew>just took the production part and made it
accessible. What it didn't solve — what

00:04:54.871 --> 00:04:59.578
<v Drew>nothing solves — is the marketing. The
distribution. Getting anybody to actually

00:04:59.578 --> 00:05:03.868
<v Drew>listen. That's still the real job. The
making is largely handled now. The

00:05:03.868 --> 00:05:05.000
<v Drew>being-heard is not.

00:05:05.000 --> 00:05:09.288
<v Drew>So... can you make a podcast with AI?
Yes. And it's a genuinely good setup,

00:05:09.288 --> 00:05:13.865
<v Drew>especially compared to cloning yourself
on video, which is way harder — voice is

00:05:13.865 --> 00:05:18.269
<v Drew>much easier than video. If you've got
something to say and you don't have the

00:05:18.269 --> 00:05:22.789
<v Drew>time to run a weekly show the old way,
you can replace most of that with AI and

00:05:22.789 --> 00:05:26.787
<v Drew>end up with something real. This is
actually episode two, so I'm still

00:05:26.787 --> 00:05:30.495
<v Drew>finding out how it holds up. But the
production? That part works.

00:05:30.495 --> 00:05:34.865
<v Drew>That's it for this episode. If you've got
a question you've been wondering about...

00:05:34.865 --> 00:05:39.129
<v Drew>can AI do this thing?? send it to me. The
address is in the show notes. I'm Drew.

00:05:39.129 --> 00:05:40.248
<v Drew>This was AI Honestly.

