Asset 20 8 2
Does AI recommend your business? Run the free check →

Join 15,000 business owners, marketers and entrepreneurs. The Sunday newsletter you'll be annoyed only arrives once a week.

Article

What Does Audio to Text Transcription Work Involve?

The short version: audio to text transcription work involves listening to a recording, running it through software that gives you a rough draft, then fixing that draft by hand for accuracy, speaker labels, punctuation and formatting before delivering a clean file. The typing bit is now the smallest part of the job. The judgment calls, working out who said what, cleaning up mumbled words, deciding whether to include the “ums”, are where the actual work and the actual money sit.

The bit everyone gets wrong before they even start

Most people picture transcription as sitting with headphones on, typing along in real time like a court stenographer. That’s not what it looks like in 2026, and it hasn’t looked like that for years. What it involves is more like editing than typing. You take a machine-generated draft that’s about 85 to 95 percent accurate, and you spend most of your time hunting down the wrong 5 to 15 percent.

That gap sounds small. It isn’t. On a one-hour interview with two speakers and a bit of background noise, a rough transcript can have forty or fifty errors, wrong names, swapped speaker tags, homophones like “there” and “their,” entire sentences that the software simply guessed at because someone spoke over someone else. Fixing that takes real concentration, not fast fingers.

What happens, step by step

Here’s the real workflow I use and that most professional transcribers I know follow, whether they’re doing it for a client directly or through a platform:

  • Step 1: Listen once through, no typing. You get a feel for accents, speaker count, audio quality, and any technical terms or names that will trip up the software.
  • Step 2: Run the machine draft. Tools like Otter.ai, Descript, or Whisper (OpenAI’s speech recognition model) produce a first pass in minutes rather than hours.
  • Step 3: Correct in short passes. Fifteen to twenty minutes of audio at a time, checking the draft against the recording, fixing misheard words and wrong speaker labels.
  • Step 4: Format to spec. This is where clients differ wildly. Some want verbatim (every “um,” every false start, every stutter). Some want clean verbatim (readable, filler words stripped out). Some want timestamps every thirty seconds for video captioning.
  • Step 5: Proofread against the original audio a second time. Skipping this step is the single most common reason transcripts get sent back for revision.
  • Step 6: Deliver in the requested file format. Word document, SRT for subtitles, plain text with timecodes, whatever the brief says.

A tidy one-hour recording with one clear speaker takes an experienced transcriber roughly two to three hours to turn into a clean, accurate file. Two speakers with crosstalk, background noise, or strong accents can push that to four or five hours for the same sixty minutes of audio. That ratio, three to four hours of work per hour of audio, is the number nobody puts in the job advert.

A real example from my own desk

A few months back I recorded a forty-minute conversation with a client for a case study. Two people, a decent microphone, a quiet room. I assumed it would take twenty minutes to clean up. It took closer to two hours.

The software mislabelled the speakers eleven times because our voices sat in a similar pitch range. It transcribed “Zapier” as “the paya” three separate times. It missed an entire eight-second stretch where the client interrupted me to correct a figure, the audio just vanished from the draft, no gap marker, nothing. If I’d sent that transcript straight to the client without checking it against the recording, the case study would have had a factual error in it that I’d have had no idea existed. That’s the part of the job that separates someone who can be trusted with client audio from someone who’s just fast at a keyboard.

The uncomfortable bit most guides skip over

Here’s the truth that a lot of “how to become a transcriptionist” content dances around: raw typing speed used to be the whole job, and now it’s worth almost nothing. Automatic speech recognition has become good enough that anyone can generate a rough transcript for free or nearly free in the time it takes to make a coffee. That means the market has already priced out people who only offer speed.

What’s left is the editing skill, the ear for context, the judgment about what a client needs versus what they asked for in one vague sentence, and the ability to handle audio that machines still struggle with, thick accents, overlapping speakers, technical or medical terminology, poor quality phone recordings. If you’re coming into this work expecting to get paid for typing quickly, you’ll be disappointed and probably underpaid. If you come in ready to be the quality control layer that catches what the machine missed, there’s still real work here.

This is also why some beginners feel misled once they start. I’ve written before about what a transcription job means for newcomers, and the honest short answer is that it’s closer to proofreading and light editing than to stenography.

The tools doing the heavy lifting (and their limits)

Most working transcribers in 2026 use a combination of an automatic speech recognition tool for the first draft and manual software for cleanup. Common choices include Otter.ai, Descript, Rev’s own transcription editor, and Whisper-based tools that plenty of freelancers run for free or near-free. None of these handle overlapping speech well. None of them reliably catch sarcasm-driven punctuation choices, and none of them know that “Kubernetes” isn’t a typo when your client is a software engineer.

That’s the gap a human fills. Platforms like Rev, TranscribeMe, GoTranscript and Scribie build their entire business model on that gap, they take the machine draft, hand it to a human editor, and charge the client for the accuracy the machine alone can’t deliver.

What it pays

Pay varies more than people expect, and it’s tied directly to how messy the audio is, not how long the file is. Rough current ranges:

  • General platforms like Rev or GoTranscript: roughly $0.30 to $1.10 per audio minute, so a $15 to $66 range for a one-hour file, depending on difficulty and turnaround speed.
  • Specialist legal or medical transcription: $15 to $25 per audio hour when starting out, more once you’re certified and fast.
  • Direct client work, the kind you build yourself rather than through a platform, can run $30 to $60 per audio hour once you have a portfolio and repeat clients who trust your accuracy.

None of that accounts for the three-to-four-times multiplier between audio length and actual working time, so when you’re comparing this to an hourly wage, divide the audio-hour rate by three or four to get closer to your real hourly earnings while you’re new. That’s the maths that job adverts conveniently leave out.

If you want the fuller breakdown of rates and platforms, I’ve laid it out in transcription jobs for beginners and what they pay, and for the wider landscape of home-based work, work from home jobs and how to spot the scams is worth a read before you sign up anywhere that asks for money upfront to “activate” your account. Real transcription platforms never charge you to start.

How to get started

If this sounds like something you want to try, the practical route is simpler than most guides make it sound:

  • Do a free timed test transcription (Rev and GoTranscript both offer these) to see your real accuracy and speed before committing.
  • Get a decent pair of closed-back headphones. Cheap earbuds miss the low-frequency detail that tells you whether someone said “can” or “can’t.”
  • Learn keyboard shortcuts for your playback software, rewinding five seconds without touching the mouse saves real time across a working day.
  • Start with general transcription before chasing legal or medical work, which usually needs a short paid course and sometimes certification.

For a more complete walkthrough of the first thirty days, how to get started with English transcription jobs covers the practical setup in more detail than I have room for here.

Where this fits into a bigger business

If you’re a business owner rather than someone job-hunting, transcription work matters for a different reason. Every podcast episode, webinar, or client call you record is raw material sitting unused unless it’s turned into text. That transcript becomes blog posts, LinkedIn updates, email content, and quote graphics. I’ve built an entire process around exactly this, and I walk through it in my content repurposing workflow for solopreneurs, one recording turned into seven usable assets, and none of it works if the transcript underneath it is sloppy.

And if transcription is one option on a longer list of things you’re weighing up as flexible income, it sits alongside a wider set of choices covered in legitimate work from home jobs for beginners, where it’s usually one of the lower barrier-to-entry options but not the highest paying one long term.

Who this suits and who it doesn’t

It suits people who can sit still and concentrate for long stretches, who don’t mind repetitive listening, and who have a good ear for language, including regional accents and technical jargon. It does not suit people who want quick money for minimal effort, or who assume the software will do the job for them. The software is a starting point, not a finish line, and clients can tell within the first paragraph whether you checked your work against the actual recording or just skimmed the auto-generated draft and hit send.

Frequently asked questions

Is audio to text transcription work still worth doing in 2026?

Yes, but only if you position yourself as the accuracy layer on top of AI drafts rather than as a fast typist. Raw typing speed has been priced out by free automatic transcription tools, but the editing and quality-checking skill is still paid work.

How long does it take to transcribe one hour of audio?

For clean, single-speaker audio, expect two to three hours of work. For multiple speakers, background noise, or strong accents, it can take four to five hours for the same sixty minutes of recording.

Do I need special software to start transcribing?

No expensive software is required to begin. Free or low-cost tools like Otter.ai or Whisper-based apps handle the first draft, and you edit it in a basic word processor or the platform’s own editor. A decent pair of headphones matters more than any paid software.

What’s the biggest mistake beginners make with transcription work?

Trusting the machine draft too much and skipping the second listen-through against the original audio. That’s the step where wrong speaker labels, missing words, and factual errors get caught before a client ever sees them.

Related reading: How to Spot Legitimate Transcription Jobs Versus Scams and What Does Repurposing Content Mean (And How Do You Do It Well).

Where to check the details

Published and maintained by the Lilach Bullock team, covering marketing, AI and business growth.
Your buyers are asking AI who to use. Does it say you?

See for free whether ChatGPT, Claude, Perplexity, Gemini and Google name you, and get the plan to become the answer.

Check my AI visibility →
Sundays only

Get the Sunday newsletter.

One email a week. AI experiments, marketing tactics, and the workflows Lilach is building right now in her own business.

Subscribe free

Let’s get your marketing running on AI.

Book a free 30-minute call

We figure out what you need, where AI fits in, and what working together would look like.

Book the call →

Or take the 30-second calculator

You’ll see the hours and the money quietly leaking out of your week, and the three workflows worth building first.

Take the calculator →

Or grab the free AI resource library

Prompt packs, templates, checklists, and swipe files. The exact tools I build for paying clients. Yours, free.

Get the library →
Keep reading

More from the blog.